On-chip cache-friendly dimension transformation device and neural network processor

By designing an on-chip cache-friendly dimension transformation device for neural networks, the problem of data dimension transformation increasing software complexity and memory access overhead in the prior art is solved, and efficient data dimension transformation and improved data throughput are achieved.

CN114840470BActive Publication Date: 2025-05-16CHENGDU DENGLIN TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210335890.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-05-16
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

When the prior art realizes data dimension transformation between different layers or different networks in a neural network, it increases the software complexity of AI algorithms, increases the memory access overhead of the chip, and reduces the throughput of the chip.

Method used

A dimension transformation device is designed for on-chip cache-friendly, including a control module, a data cache module, a write control module and a read control module. The device generates an input address through configuration information, reads data in the order of output data dimensions from low to high dimensions, and realizes dimension transformation while data transfer.

Benefits of technology

Without reducing the efficiency of on-chip cache access, data dimension transformation can be effectively completed, reduce the processor's calculation load, improve data throughput, and save software overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114840470B_ABST
    Figure CN114840470B_ABST
Patent Text Reader

Abstract

The present application provides a dimension transformation device that is friendly to on-chip cache, which includes a control module, a data cache module composed of multiple storage blocks, a write control module and a read control module. The control module obtains configuration information related to the instruction while receiving the data handling instruction to be processed, and in response to determining that the input data dimension and the output data dimension in the configuration information do not match, generates a corresponding input address according to the configuration information to read the corresponding data from the external storage unit in the order of the output data dimension information from low dimension to high dimension, writes the input data from the external storage unit into the data cache module through the write control module, and reads the data from the data cache module through the read control module for output. The dimension transformation device can effectively complete the data dimension transformation without reducing the on-chip cache access efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to data processing technology in a neural network, and in particular to a data dimension transformation device suitable for use in a neural network processor. Background Art

[0002] The statements in this section are merely intended to provide background information related to the technical solution of the present application to aid understanding, and they do not necessarily constitute prior art for the technical solution of the present application.

[0003] Artificial intelligence (AI) technology has developed rapidly in recent years and has penetrated into various fields such as visual perception, speech recognition, assisted driving, smart home, traffic dispatch, etc. Many AI algorithms involve learning and computing based on neural networks, such as convolutional neural networks (CNN), recurrent neural networks (RNN), deep neural networks (DNN), etc. These AI algorithms often involve multi-layered neural networks and learning models composed of multiple neural networks, requiring strong parallel computing capabilities to process massive amounts of data. Therefore, processors that support multi-core parallel computing, such as GPUs, GPGPUs, and AI acceleration chips, are usually used to perform related neural network operations. These processors often need to provide the same batch of data to different layers of the neural network or to different neural networks, and the formats and dimensions of the data required by different layers of these neural networks or between different neural networks are often different.

[0004] Taking a convolutional neural network as an example, each layer in the neural network uses the input feature data and the parameters related to the layer (for example, convolution parameters, etc.) to perform the relevant operations of the layer (for example, convolution operations, etc.). The feature data can also be called a feature map, which can be regarded as a data block with a certain width and height. The output feature data obtained by each layer can be provided to the next layer as the input feature data of the next layer, but due to the different number of nodes and operation methods of each layer, the dimensions, arrangement, format, etc. of the required input feature data are often different. The existing software programming method reads the data to be processed from the on-chip cache to the memory, performs software dimensional transformation, and then writes it from the memory to the on-chip cache, thereby realizing the data matching problem between different layers or different networks. However, this method not only increases the software complexity of the AI ​​algorithm, but also increases the memory access overhead of the chip and reduces the throughput of the chip.

[0005] It should be noted that the above content is only used to help understand the technical solution of the present application, and is not used as a basis for evaluating the prior art of the present application. Summary of the invention

[0006] The present application provides a dimension transformation device that is friendly to an on-chip cache and can effectively complete data dimension transformation without reducing the on-chip cache access efficiency.

[0007] The above purpose is achieved through the following technical solutions:

[0008] According to a first aspect of an embodiment of the present application, a dimension transformation module friendly to on-chip cache is provided, comprising a control module, a data cache module composed of multiple storage blocks, a write control module and a read control module. The control module obtains configuration information related to the instruction while receiving the data handling instruction to be processed, and the configuration information at least includes base address information of input data and output data, dimension information of input data and output data, data size and data step length of each dimension of input data and output data. The control module generates corresponding input addresses using the received configuration information to read corresponding data in order from low dimension to high dimension of output data dimension. The write control module writes input data from an external storage unit into the data cache module. The read control module reads data from the data cache module for output according to the instruction of the control module.

[0009] The dimension transformation device of the above embodiment can not only move data between different storages, but also realize dimension transformation or matching between different data while moving data. For example, after the processor completes the calculation of one layer of the neural network, the dimension transformation device can be used to move data and perform data dimension transformation at the same time in the process of moving the calculation result from the on-chip cache to the off-chip storage, so that when the processor starts to execute the next layer of operation, what is read is the data that has been transformed in advance. The processor no longer needs to occupy additional clock cycles and computing resources to perform the data dimension transformation process. Therefore, the dimension transformation device can not only reduce the computing load of the processor but also improve its data throughput. In addition, the dimension transformation device performs dimension transformation according to the configuration information in parallel during the data transfer process, and can read the corresponding data in sequence from low dimension to high dimension while outputting. Compared with the method of taking out all the data through software for dimension transformation and then writing, unnecessary software overhead is saved. And the dimension transformation is completed without sacrificing data transmission performance.

[0010] In some embodiments, each storage block in the data cache module is an on-chip random access memory (i.e., on-chip RAM), and the number of storage blocks should at least be sufficient to divide the preset bit width of the input data and the bit width of the output data; wherein the bit width of the input data is the same as the bit width of the output data. In the above embodiments, it is no longer necessary to have an expensive on-chip RAM that accommodates the entire bit width of the input data. Instead, by using a data cache structure consisting of multiple smaller width storage blocks, the input data is divided into multiple segments and stored in parallel in multiple storage blocks. This saves area overhead and reduces the cost and performance requirements for on-chip RAM.

[0011] In some embodiments, the device may include multiple output channels, and the number of the output channels should at least be sufficient to divide the preset bit width of the input data and the bit width of the output data. In some embodiments, each output channel corresponds to an output address, the data in each channel is continuous, and the depth of each storage block is at least equal to or greater than the ratio of the data bit width of each output channel to the bit width of a single data.

[0012] In the above embodiment, the internal cache overhead of the dimension transformation device is reduced by adopting multiple output channels, and output can be performed when the demand of only one of the output channels is met in the data cache, thereby improving data transmission efficiency.

[0013] In some embodiments, the data of each dimension of the input data and the output data are stored in a linear arrangement from low dimension to high dimension.

[0014] In some embodiments, the control module is further configured to include a group of input counters and multiple groups of output counters. When generating an input address to read data, the control module uses this group of input counters to indicate the dimension information of the input data currently being read. Each dimension corresponds to an input counter. After each data read request is completed, the count values ​​of this group of input counters are updated together. The multiple groups of output counters correspond to multiple output channels, and each channel corresponds to a group of output counters. The control module generates an output address for each channel, and uses a group of output counters corresponding to each channel to indicate the dimension information of the data currently being output, and each dimension corresponds to an output counter. After each output data of each channel is completed, the count values ​​of this group of output counters of the channel are updated together. Generally speaking, the number of dimensions of the input data should be the same as the number of dimensions of the output data.

[0015] When the control module requests to read data from the external storage unit, the generated input address is calculated based on the input data base address information, the data size information of each dimension, the data step information of each dimension and the count value of the input counter of each dimension.

[0016] In some embodiments, when the control module detects that the lowest dimension of the input data dimension and the output data dimension has changed, the input counter of each dimension is adjusted in the following manner to read the data: except for the lowest dimension of the output data, all data will be read back in sequence from low dimension to high dimension according to the input data dimension, and the input counter of the lowest dimension of the output data will increase by 1 after each data request until the lowest dimension of the output data is completed, and the input counter of the lowest dimension of the original input data will increase by the single input data amount, which is related to the amount of data in the input data read each time. For example, the single input data amount can be the ratio of the bit width of the input data to the bit width of a single data.

[0017] In some embodiments, when the input data dimension and the output data dimension do not change (for example, NDHWC-> NDHWC), or when the lowest dimension of the input data and the output data do not change (for example, NDHWC-> WHDNC), the control module reads the data in the order of the input data dimension from low dimension to high dimension. That is, after each input data request, the input counter of the lowest dimension corresponding to the input data (for example, dimension C in the above example) increases the single input data amount (that is, the amount of data contained in each input data), and when the count value of the input counter of the lowest dimension corresponding to the input data reaches the data size of the corresponding dimension, the input counter of the lowest dimension is cleared and the input counter of the upper dimension is incremented by 1 (for example, if the input data dimension is NDHWC, after the dimension C data is read, the read counter of the upper dimension W is incremented by 1), and then the above process is repeated until all the data are read. When outputting data, the control module adjusts the output counters of each dimension in the following manner to control each output channel to output data: each output channel writes data in order from low to high dimensions of the input data, that is, after each output data, the output counter of the lowest dimension corresponding to the input is increased by the single output data amount (the size of which is related to the width of each data and the number of channels). When the count value of the output counter corresponding to the lowest dimension of the input data reaches the data size of that dimension, the output counter of the lowest dimension is cleared and the output counter of the upper dimension is incremented by 1, and the above process is repeated until all data is output.

[0018] In some embodiments, when the lowest dimension of the input data and the output data changes (eg, NDHWC -> NDHCW or NDHWC -> CDHNW), the control module reads the data in the order of the output data dimension from low dimension to high dimension. That is, after each input data request, the input counter of the lowest dimension corresponding to the output data (such as dimension W in the above example) is increased by 1. When the count value of the input counter corresponding to the lowest dimension of the output data reaches the data size of this dimension, the input counter of this dimension is cleared and the input counter of the upper dimension in the output data dimension is carried forward. If the upper dimension (C) of the lowest dimension (W) in the output data dimension is the lowest dimension of the input data dimension (such as the input data dimension NDHWC to the output data dimension NDHCW in the above example), the number of carry is the single input data amount (which is related to the data width of each input data), and in other cases, the number of carry is 1 (for example, if the output data dimension is NDHCW, when the data of the lowest dimension W is read, the input counter of the upper dimension C is carried forward by the single input data amount B / b, because dimension C is exactly the lowest dimension of the input data; and if the output data dimension is CDHNW, when the data of the lowest dimension W is read, the input counter of the upper dimension N is carried forward by 1), and the above process is repeated until all the data are read. When outputting data, the control module adjusts the output counters of each dimension in the following manner to control the output data of each output channel: each output channel writes data in order from low to high output data dimensions, that is, after each output data, the output counter of the lowest dimension corresponding to the output data increases the single output data amount (whose size is related to the width of each data and the number of channels). When the count value of the output counter of the lowest dimension corresponding to the output data reaches the data size of this dimension, the output counter of the lowest dimension is cleared and the output counter of the upper dimension in the corresponding output data dimension is carried. If the upper dimension of the output data dimension happens to be the lowest dimension of the input data dimension, the carry number is the single output data amount (whose size is related to the data width of each output data), and in other cases the carry number is 1. The above process is repeated until all data are output.

[0019] In the above embodiment, by setting corresponding input counters and output counters in each dimension, a flexible input address and output address generation method is provided, which helps the dimension transformation device to read and output data more simply and conveniently.

[0020] In some embodiments, the write control module is configured to: when the input data dimension and the output data dimension have not changed, or when the lowest dimension of the input data and the output data has not changed, count the input data received each time, and generate a write address for each storage block based on the current count value modulo the depth of each storage block in the data cache module, and write the currently received input data to each storage block of the data cache module accordingly. The storage depth of the data cache module can be obtained using B / M / b, where B is the bit width of the output data, M is the number of output channels, and b is the number of bits of a single data. In this way, the write address for each storage block of the data cache module generated by the write control module can be obtained by the following formula:

[0021]

[0022] in is a count of the input data received by the write control module, which is a count value starting from 0, i is an integer starting from 0, and N is the number of storage blocks.

[0023] In some embodiments, the read control module is configured to: when the input data dimension and the output data dimension do not change, or when the lowest dimension of the input data and the output data do not change, generate a read address for each storage block in the data cache module according to the following formula according to the instruction of the control module:

[0024]

[0025] in is the count value of the read data by the read control module, which is a count value starting from 0; B is the bit width of the output data, M is the number of output channels, b is the number of bits of the read data, N is the number of storage blocks, and i is an integer starting from 0.

[0026] In some embodiments, the write control module is configured to: count each received input data when the lowest dimension of the input data dimension and the output data dimension changes; perform a cyclic bit right shift operation on the currently received input data, and the number of right shift bits is the current count value multiplied by B / N, where N is the number of storage blocks in the data cache module, and B is the bit width of the input data; modulo the depth of each storage block in the data cache module according to the current count value to generate a write address for each storage block, and write the processed data to each storage block of the data cache module accordingly. Generally, the storage depth of the data cache module can be obtained using B / M / b, where B is the bit width of the output data, M is the number of output channels, and b is the number of bits of a single data. In this way, the write address for each storage block of the data cache module generated by the write control module can be obtained by the following formula:

[0027]

[0028] in is a count of the input data received by the write control module, which is a count value starting from 0, i is an integer starting from 0, and N is the number of storage blocks.

[0029] In some embodiments, the read control module is configured to generate a read address for each storage block in the data cache module according to the following formula according to the instruction of the control module when the lowest dimension of the input data dimension and the output data dimension changes:

[0030]

[0031] in It represents the count value of the data read out, which is a count value starting from 0; B is the bit width of the output data, M is the number of output channels, b is the number of bits occupied by a single data, i is an integer starting from 0, and N is the number of storage blocks.

[0032] In the above embodiment, a flexible reading and writing method is provided for the data cache inside the dimension transformation device, so that data with the same output lowest dimension are stored in different storage blocks, and these data with the same output lowest dimension can be read out simultaneously during output.

[0033] According to a second aspect of an embodiment of the present application, a processor for a neural network is provided, comprising a dimensionality transformation device according to the first aspect of an embodiment of the present application, which is used to transfer data between an on-chip cache and an off-chip memory of the processor.

[0034] The technical solution of the embodiment of the present application may have the following beneficial effects:

[0035] The dimension transformation device can not only transfer data between different storages, but also realize dimension transformation or matching between different data while transferring data, which can reduce the computational load of the processor and improve its data throughput. In addition, the dimension transformation device uses a data cache composed of multiple small on-chip RAMs and multi-channel parallel output. Only part of the cached data is needed to meet the data volume of a single output channel before output can be started, without the need to read all the feature data transferred from one neural network to another or from one layer of a neural network to another layer into the cache and then convert it. Therefore, it not only saves the area overhead of the on-chip cache but also does not affect the efficiency of data transmission.

[0036] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0038] Figure 1 Schematic diagram of the module structure of a dimensionality transformation device according to an embodiment of the present application.

[0039] Figure 2 Schematic diagram of the relationship between the data storage address, dimension size and step size according to an embodiment of the present application.

[0040] Figure 3 The diagram is a schematic diagram of a process in which a dimensionality transformation device reads data from an external storage unit according to an embodiment of the present application.

[0041] Figure 4 The diagram is a schematic diagram of a process in which a dimensionality transformation device reads data from an external storage unit according to an embodiment of the present application.

[0042] Figure 5 It is a structural schematic diagram of a dimensionality transformation device according to another embodiment of the present application. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solutions and advantages of this application more clear, the present application is further described in detail by specific embodiments in conjunction with the accompanying drawings. It should be understood that the described embodiments are part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0044] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the application.

[0045] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0046] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0047] The data dimension information may be, but is not limited to, one or more of the following dimensions commonly used in artificial intelligence networks: height (H), width (W), depth (D), channel (C), number of samples (N), etc. Common data in NDHWC format has a dimension of 5, the lowest dimension is C, and the highest dimension is N. In the operation process of a neural network, the processor usually processes one of the layers of the neural network and passes the output of the layer as feature data to the next layer in the network for processing. The dimensional arrangement of feature data between different layers and different networks is often different. For example, some layers are arranged in NDHWC, while some layers are arranged in NCDHW. As mentioned above, dimensional transformation through software programming increases the memory access overhead of the chip and reduces the throughput of the chip. In practice, the inventor tried to use a dedicated hardware module to implement dimensional transformation in an on-chip cache, but found that in order to complete the dimensional transformation, an on-chip cache that can at least hold a large amount of data for dimensional transformation is required, which is a large overhead in terms of area and energy consumption.

[0048] In an embodiment of the present application, a dimensionality transformation device is provided that does not sacrifice data transmission performance and can significantly reduce on-chip cache usage. Figure 1A functional module structure diagram of a dimension transformation device according to an embodiment of the present application is given. The dimension transformation device includes a control module and a data path composed of a write control module, a data cache module and a read control module. The control module generates an input address according to the instructions and configuration information received from the external control unit to request read data from the previous module, and instructs the write control module to prepare to receive input data from the previous module. The received instructions refer to instructions that require the participation of the dimension transformation device, which may include but are not limited to data handling instructions such as storage instructions (STORE). The previous module here may be but is not limited to memory, on-chip random access memory (can be referred to as on-chip RAM), on-chip read-only memory (can be referred to as on-chip ROM), external memory, and any other data cache. The control module may also generate an output address according to the configuration information received from the external control unit, so as to request write data to the next module, and instruct the read control module to prepare to read data from the data cache module to transfer output data to the next module. The next module here may be but is not limited to memory, on-chip RAM, external memory, and any other data cache. When the write control module receives input data from the previous level module, it generates a write address for the data cache module according to its related storage state, and then transmits the write signal, write address and input data to the data cache module. The data cache module receives the write signal, write address and input data input from the write module, and caches the input data according to the write address. At the same time, the data cache module can also receive the read signal and read address from the read control module, and transmit the output data corresponding to the read address to the read control module. The read control module can generate a read address and a read signal according to the indication signal and configuration information from the control module and send them to the data cache module, and then transmit the output data returned by the data cache module together with the output address generated by the control module to the next level module.

[0049] In an embodiment of the present application, the control module instructs each module to read data in a specific order and output data in a corresponding set manner according to the received configuration information, so as to realize the dimensional transformation between the input data and the output data (which will be described in detail below). The dimensional transformation device can not only move data between different storages, but also realize the dimensional transformation or matching between different data while moving the data. For example, after the processor completes the calculation of one layer of the neural network, the dimensional transformation device can be used to move the data and perform the data dimensional transformation at the same time in the process of moving the calculation result from the on-chip cache to the off-chip storage, so that when the processor starts to execute the next layer of operation, the data read is the data that has been pre-transformed. The processor no longer needs to occupy additional clock cycles and computing resources to perform the data dimensional transformation process. Therefore, the dimensional transformation device can not only reduce the computing load of the processor but also improve its data throughput.

[0050] In an embodiment of the present invention, the bit widths of input data and output data are required to be the same, so that the data throughput can be unchanged during the dimension conversion. In the following text, B is used to represent the bit widths of input data and output data, and the bit range of input data and output data can be recorded as [B-1:0]. Commonly used bit widths of input and output data include but are not limited to 2048, 1024 and 512. However, in computer processing, data is not accessed in bits, but in data consisting of multiple bits as the basic unit, for example, 8 bits constitute 1 byte. In the following text, b is used to represent the number of bits occupied by a single data, and b should usually be a multiple of 8. The data types supported in the embodiments of the present application may include but are not limited to 8-bit integers, 16-bit integers, 32-bit integers, 16-bit floating point numbers, 32-bit floating point numbers and 64-bit floating point numbers. In this way, in this article, the data width that can be read from the external unit each time by the dimension conversion device is B / b, that is, B / b data can be read at one time (hereinafter, it can also be collectively referred to as "a set of data")

[0051] In an embodiment of the present application, a data cache module is composed of multiple on-chip RAMs. For the convenience of description, it is assumed that the data cache module contains N RAMs, which can also be referred to as N storage blocks (banks), namely bank0 to bankN-1. A single data as the basic unit of dimensional transformation can only be stored in one bank, and cannot be stored across multiple banks. Therefore, the bit width of each bank should be an integer multiple of the number of bits b occupied by a single data. In order to facilitate one-time reading and output of data to improve data throughput, input data with a bit width of B is evenly cached in N banks and can be read and output at one time, so N is set to a natural number that satisfies the bit width B of the input data and output data. That is, in the data cache module, the bit width of each bank is B / N, and B / N is an integer multiple of the number of bits b occupied by a single data. In this way, each bank corresponds to a write address, a read address, and data with a data bit width of B / N. That is, bank0 corresponds to write address 0, read address 0, and input data / output data with a bit range of [B / N-1:0]; bank1 corresponds to write address 1, read address 1, and input data / output data with a bit range of [2*B / N-1:B / N]; and so on, bankN-1 corresponds to write address N-1, read address N-1, and input data / output data with a bit range of [B-1: (N-1)*B / N].

[0052] The depth of each bank of the data cache module depends on the specific output requirements, that is, the depth of each bank at least meets the set minimum output data volume. In some embodiments of the present application, in order to reduce the internal cache overhead, the output data can be divided into M output channels, each output channel corresponds to an output address, that is, output address 0 to output address M-1. The data in each channel must be continuous, and the channels may be discontinuous. M should be a natural number that can divide the bit width B of the input data and the output data, so the data width of each output channel is B / M. It should be understood that a single data cannot be output across channels, so B / M cannot be too small, and should be an integer multiple of the number of bits b occupied by a single data. Accordingly, the depth of each bank of the data cache (that is, the number of data that can be stored in each bank) is B / M / b to meet the storage needs required for dimensional transformation. In some embodiments, to meet the continuous transmission of data, each bank can use a ping-pong RAM with a depth of 2*B / M / b. The size of M affects the time and resources required for the transformation from the lowest dimension to a higher dimension. The larger M is, the more on-chip storage resources are required and the longer the output delay is, but the result of the dimensional transformation is more complete. Therefore, the choice of M is the result of a compromise. The common parameters defined above can be taken, but not limited to, the reference values ​​provided below: B=2048, N=32, M=8. At the same time, to support all data formats, b takes the minimum value that can be obtained, here b=8, at this time the required depth of each bank is B / M / b=32 or 2*B / M / b=64, and the bit width of each bank is B / N=64.

[0053] In addition, it should be pointed out that, without loss of generality, the scenario targeted by the embodiments of the present application is the dimensional transformation of data stored in a linear arrangement format, that is, each dimensional data is stored sequentially from the lowest dimension to the highest dimension and the address is increasing. Taking the dimensional arrangement format NDHWC as an example, its dimension is 5, of which the lowest dimension is C and the highest dimension is N. Each piece of data can be uniquely represented by a dimensional coordinate in the dimensional arrangement of the data. For example, the data in the NDHWC format can be represented by a 5-dimensional coordinate, (0, 0, 0, 0, 0) represents the first data in the data; (0, 2, 0, 0, 1) represents the third data set in the D dimension and the second data in the first data set in the H and W dimensions. At the same time, as the coordinates increase, the address of the external storage where the data is located is increasing by default, such as Figure 2As shown, the data size of each dimension indicates the number of data on that dimension. For example, the data size of dimension 0 is L0, that is, there are L0 data on dimension 0, starting from data 0 to data L0-1. According to the above coordinate representation, the value on a certain dimension coordinate must be smaller than the corresponding dimension data size. For example, assuming that the data size of the C dimension in the specified NDHWC is 2, there are only two data coordinates in the C dimension direction: 0 and 1. When the coordinates of the lower dimension are completed, the data of the higher dimension will be carried. For example, assuming that the size of the C dimension in the specified NDHWC is 2 and the size of the W dimension is 2, the order in which the data is arranged in the memory should be (0,0,0,0,0), (0,0,0,0,1), (0,0,0,1,0), (0,0,0,1,1)... The step size of each dimension refers to the change in address caused by adding 1 to each dimension. The step size of each dimension must be greater than or equal to the data size of the lower dimension, and the step size of the lowest dimension is the size of 1 data. For example, the data arrangement in the NDHWC format mentioned above, from (0,0,0,1,0) to (0,0,0,1,1), will increase the step size of the W dimension in the address. For another example, Figure 2 The dimension data size of dimension 0 is L0, the dimension data size of dimension 1 is L1, and the address interval between data block 0 of dimension 1 and the next data block 1 of dimension 1 is the step size of dimension 1, which is greater than or equal to the dimension data size L0 of dimension 0. Figure 2 It can be seen that in the scenario of storage in the above-mentioned linear arrangement format, the storage address of a certain data in the corresponding dimension can be easily determined according to the given starting address, dimension data size and dimension step, thereby obtaining the corresponding data from the storage address.

[0054] Continue to refer Figure 1 When the control module receives a data movement instruction from an external control unit, it will simultaneously obtain configuration information related to executing the instruction. The configuration information includes but is not limited to: dimension information of input data, size information of each dimension of input data, step information of each dimension of input data, dimension information of output data, size information of each dimension of output data, step information of each dimension of output data, base address information of input data, base address information of output data, data format information, etc. The control module generates an input address based on the received configuration information and sends it together with an input request to the previous level module to request to read data. The input request and input address are sent once per clock cycle until all dimensions of data are read.

[0055] In some embodiments, the control module includes a group of input counters and multiple groups of output counters for generating input addresses and output addresses respectively. Among them, this group of input counters corresponds to the dimension information of the input data, and each dimension corresponds to an input counter to identify the dimension information of the data currently read. This group of input counters updates the count value together after each data read request. And multiple groups of output counters correspond to multiple output channels, and each channel corresponds to one group of output counters. The group of output counters corresponding to each output channel corresponds to the dimension information of the output data, and each dimension corresponds to an output counter to identify the dimension information of the data currently output. Each group of output counters updates the count value together after each output data. Generally, the number of dimensions of the input data should be the same as the number of dimensions of the output data. And in general, the initial values ​​of the above-mentioned counters are zero, and they are incremented when the corresponding counting conditions are met, until the count value reaches the data size of the corresponding dimension, and then it is cleared and carried to the counter of the higher dimension. That is to say, the counting condition of the counter of the higher dimension is that the value of the counter of the lower dimension reaches the size of its dimensional data. The self-increment value of the counter of the higher dimension is always 1, but the self-increment value of the counter of the lowest dimension is related to the data width of each data. The counting condition of the lowest dimension counter is the number of data in each requested data or each written data.

[0056] The input address calculated by the control module each time it requests to read data is the sum of the count results of the input counters of each dimension and the data step length of each dimension. When the control module starts reading, it starts an input counter with an initial value of 0 for each dimension to represent the dimensional coordinates used for each piece of data, and uses the dimensional coordinates and the step length information of each dimension to calculate the input address corresponding to each input request. For example, in the case of NDHWC, the requested input address can be calculated according to the following formula:

[0057]

[0058] in , , , as well as Represents the value of the input counter for each dimension in the NDHWC format data, and , , , as well as Indicates the data step size of each dimension. For NDHWC, the data step size of the lowest dimension C The value is 1. In an embodiment of the present application, the data requested by the control module to the previous level at each moment are the data of the lowest dimension of the input data dimension, so there will be no situation where an input request contains changes in the input data dimension other than the lowest dimension, that is, there will be no situation where the high-dimensional coordinates of the first data and the last data of an input request are different. If the remaining data of the lowest dimension is not enough for the amount of data requested once, only the existing amount of data will be requested, and the insufficient part will be left blank to ensure that the address calculation will not be confused. For example: For the case where the input data dimension is NDHWC, the data corresponding to the address of the first input request of the control module is (0,0,0,0,0) to (0,0,0,0,B / b-1), as mentioned above, B is the preset bit width of the input data, and b is the number of bits contained in a single data; but if the data size of the C dimension of the input data is only B / 2 / b, that is, there are only B / 2 / b data in the C dimension, then the data corresponding to the address of the first input request is (0,0,0,0,0) to (0,0,0,0,B / 2 / b-1).

[0059] The control module determines whether dimension transformation is required based on the input data dimension information and the output data dimension information in the received configuration information. When it is determined that dimension transformation is not required, data can be read sequentially from low dimension to high dimension in the above manner. When it is determined that dimension transformation is required, the input counters of each dimension can be adjusted so that the corresponding input addresses generated can read the corresponding data in the order from low dimension to high dimension according to the output data dimension.

[0060] More specifically, when the input data dimension and the output data dimension do not change, or when the lowest dimension of the input data and the output data do not change, the control module reads data sequentially from low dimension to high dimension according to the input data dimension. For example, if the input dimension is NDHWC, the control module will read the data of B / b C dimensions in sequence in each clock cycle until all the data of the C dimensions are obtained when NDHW is (0, 0, 0, 0), and then the W dimension is increased by 1 and the data of B / b C dimensions are read in sequence in each clock cycle until all the data of the C dimensions are obtained when NDHW is (0, 0, 0, 1), and so on.

[0061] Figure 3A flow chart of a dimension transformation device reading data from an external unit according to an example of the present application is given. During initialization, each chip of the dimension transformation device is reset, all counters in the control module are reset, and the internal control logic restores the initial device. When the control module receives instruction information and configuration information, if it is determined that the input data dimension and the output data dimension have not changed (for example, NDHWC-> NDHWC), or it is determined that the lowest dimension of the input data and the output data has not changed (for example, NDHWC-> WHDNC), the control module calculates the input address according to the count value of the input counter of each dimension and the step size of the data of each dimension and issues an input data request, and starts to read data in the order of the input data dimension from low dimension to high dimension. After each input data request, the input counter corresponding to the lowest dimension of the input data (such as dimension C in the above example) increases the single input data amount (i.e., the number of data corresponding to each data); when the count value of the input counter corresponding to the lowest dimension of the input data reaches the data size of the dimension, the input counter of the lowest dimension is cleared and the input counter corresponding to the previous dimension is increased by 1 (for example, if the input data dimension is NDHWC, then after the dimension C data is read, the input counter of the previous dimension W is increased by 1, that is, the input counter of the W dimension is increased by 1), and then the above process is repeated continuously, from low dimension to high dimension until all data are read. When the control module determines whether the count value of the input counter of each dimension has reached the data size of each corresponding dimension, if it has been reached, it means that all data of the current instruction have been read, and the internal state of the control module can be restored at this time to prepare for the execution of the next instruction.

[0062] When the lowest dimension of the input data dimension and the output data dimension changes, the control module needs to adjust the order of reading data. Taking the input dimension as NDHWC and the output dimension as NCDHW as an example, when the control module detects that the lowest dimension of the input data is C and the lowest dimension of the output data is W, it will generate an address to first read the data from (0, 0, 0, 0, 0) to (0, 0, 0, 0, B / b-1), and then read the data from (0, 0, 0, 1, 0) to (0, 0, 0, 1, B / b-1), until the W dimension is completed and then continue to read the data from (0, 0, 0, 0, B / b) to (0, 0, 0, 0, 2*B / b-1), and so on. In other words, when it is found that the lowest dimension of the input data dimension and the output data dimension has changed, the control module will still read all the data back in order from low dimension to high dimension according to the input data dimension except the lowest dimension of the output data. The input counter of the lowest dimension of the output data will increase by 1 after each data request until the lowest dimension of the output data is exhausted. The input counter of the lowest dimension of the original input data will then increase by the single input data amount. It should be understood that the above method is only an example and not a limitation. Figure 2In the scenario of storing in the linear arrangement format described above, the storage address of a data in the corresponding dimension can be easily determined according to the starting address, dimension size and dimension step given by the input data, and thus the corresponding data can be obtained from the storage address. Therefore, the control module can generate the corresponding input address according to the starting address, input data dimension size and dimension step given in the configuration information to read the corresponding data from the previous level module for caching according to the output data dimension arrangement in the configuration information, thereby completing the corresponding dimension transformation during the data movement process.

[0063] Figure 4 A schematic diagram of a process flow of a dimension transformation device according to an example of the present application for reading data from an external unit when the lowest dimension of the input data dimension and the output data dimension changes is given. The input address is still generated according to the input address formula given above, based on the count value of the input counter of each dimension and the data step length of each dimension, but with Figure 3 The difference is that the counting method of the input counters of each dimension has changed. When the lowest dimension of the input data and output data changes (such as NDHWC -> NDHCW or NDHWC -> CDHNW), the control module adjusts the counting method of the input counters of each dimension so that the corresponding input addresses generated can read the corresponding data in the order of the output data dimension from low dimension to high dimension. Figure 4 As shown, after each input data request, the input counter corresponding to the lowest dimension of the output data (such as dimension W in the above example) is increased by 1. When the count value of the input counter corresponding to the lowest dimension of the output data reaches the data size of this dimension, the input counter of this dimension is cleared and the input counter of the upper dimension in the output data dimension is carried forward; if the upper dimension (C) of the lowest dimension (such as W) in the output data dimension is the lowest dimension in the input data dimension (such as the input data dimension NDHWC to the output data dimension NDHCW in the above example), the number of carries is the single input data amount (that is, the input counter corresponding to the lowest dimension of the input data is increased by the data corresponding to each data entry). In some cases, the carry number of the input counter of the previous dimension is 1 (i.e., plus 1) (for example, assuming that the input data dimension is NDHWC, if the output data dimension is NDHCW, after the lowest dimension W data is read, the carry number of the input counter to the previous dimension C is the single input data amount B / b, because dimension C is exactly the lowest dimension of the input data; and if the output data dimension is CDHNW, after the lowest dimension W data is read, the carry number of the input counter to the previous dimension N is 1), and the above process is repeated until all the data are read.

[0064] Continue to refer Figure 1After the control module sends an input request and input address to the previous module, the previous module will feedback the input data and input indication after a certain delay. After receiving the input indication, the write control module will count the input data and divide the input data evenly into N parts according to the number of banks N in the data cache module, so that each piece of data can be stored in each bank of the data cache module accordingly. As mentioned above, bank0 corresponds to write address 0 and data with a bit range of [B / N-1:0] in the input data; bank1 corresponds to write address 1 and data with a bit range of [2*B / N-1:B / N] in the input data; and so on, bankN-1 corresponds to write address N-1 and data with a bit range of [B-1: (N-1)*B / N]. The write control module generates a write address for the input data based on the count of the input data. As the count of input data increases, the write address corresponding to each bank will also increase accordingly. For example, each segment of the first input data is stored in the address numbered 0 of each bank, while each segment of the second input data can be stored in the address numbered 1 of each bank in sequence. The maximum number of data that each bank can store is called the depth of the bank. The greater the depth of the bank, the larger the area occupied and the higher the cost; and the smaller the bank depth, the greater the access overhead, and multiple accesses are often required to complete data transfer. As mentioned above, the depth of each bank of the data cache module can be set according to the bit width of the input data and the number of output channels (for example, B / M / b or 2*B / M / b in the above example). When the count of the input data by the write control module is the same as the depth of the bank, the count returns to 0 and a new round of counting begins again.

[0065] When the input data dimension and the output data dimension do not change, or when the lowest dimension of the input data and the output data do not change, the write address of each storage block of the data cache module generated by the write control module can be obtained by the following formula:

[0066]

[0067] in It is the count of the input data received by the write control module. It is a count value starting from 0. i is an integer starting from 0, and N is the number of storage blocks. B / M / b indicates the storage depth of the data cache module (which can also be understood as the depth of each storage block). B represents the bit width of the input or output data, M is the number of output channels, and b is the number of bits of a single data.

[0068] In the case where the lowest dimension of the input data dimension and the output data dimension changes, the write control module counts the input data received each time; performs a cyclic bit right shift operation on the currently received input data, and the number of right shift bits is the current count value multiplied by B / N, where N is the number of storage blocks in the data cache module and B is the bit width of the input data; modulo the depth of each storage block in the data cache module according to the current count value to generate a write address for each storage block, and write the processed data to each storage block of the data cache module accordingly. Generally, the write address for each storage block of the data cache module generated by the write control module can be obtained by the following formula:

[0069]

[0070] in is the count of the input data received by the write control module, which is a count value starting from 0, i is an integer starting from 0, N is the number of storage blocks, and the other parameters are the same as above.

[0071] It can be seen that when the lowest dimension of the input data dimension and the output data dimension changes, the write control module adjusts the storage order of the input data in the data cache so that the output to the next-level module is also output data that is continuously and linearly arranged according to its dimension. Still taking the input dimension as NDHWC and the output dimension as NCDHW as an example, when the control module detects that the lowest dimension of the input data dimension and the output data dimension has changed, it will send a corresponding dimension change indication to the write control module. The write control module counts the received input data (for example, as mentioned above, modulo the depth of each bank), and performs a bit right shift operation on the current input data by multiplying the count value by B / N. This operation can put data with the same lowest output dimension into different banks, so that data with the same lowest output dimension can be read out at the same time.

[0072] Continue to refer Figure 1, the write control module will transmit the generated write address and write signal together with the processed input data to the data cache module, and notify the control module how much data has been transmitted to the data cache module. The control module can notify the read control module to read data from the data cache module for output when there is enough data to output. When there is no change in the lowest dimension of the input data dimension and the output data dimension, when a piece of data is written into the data cache, the control module can determine that there is enough data to output. Because in this case, the continuous data obtained from the previous stage according to the continuous linear arrangement of the input data can be directly output continuously. When the lowest dimension of the input data dimension and the output data dimension changes, since the data is read across dimensions, as mentioned above, when obtaining data from the previous module, the lowest dimension of the output data will increase the dimension coordinate by 1 after each data request until the number of data coordinates of the lowest dimension of the output data reaches the size of the data required by the output channel, and then it can be determined that there is enough data to output. In the embodiment where the dimension conversion device is provided with M output channels, in order to ensure a stable throughput rate, the control module will determine that there is enough data to output when at least B / M / b data are written into the data cache (that is, the data bit width of at least one channel is satisfied).

[0073] For the process of writing data cache, we still take the above input dimension of NDHWC and output dimension of NCDHW as an example, where the reference values ​​given above are used: B=2048, N=32, M=8, and b=8 is taken to support all data formats, the depth of each bank is 64, and the bit width of each bank is 64. The conversion process for the input dimension of NDHWC and the output dimension of NCDHW is as follows: for the NDHWC data written in the first clock, the data with dimension coordinates (0, 0, 0, 0, 0) to (0, 0, 0, 0, 7) is written to address 0 of bank0, and in the second clock cycle, the data (0, 0, 0, 1, 0) to (0, 0, 0, 1, 7) is written to address 1 of bank1... At the thirty-second clock cycle, the data (0, 0, 0, 31, 0) to (0, 0, 0, 31, 7) is written to bank31. At this time, the data meets the minimum data output requirement.

[0074] Continue to refer Figure 1 When the control module determines that there is enough data output, the control module triggers the read control module. The read control module generates N read addresses and read signals to read the data in the data cache module. Corresponding to the above write control module, when the input data dimension and the output data dimension do not change, or when the lowest dimension of the input data and the output data do not change, the read control module generates the read address for each storage block in the data cache module according to the following formula according to the instruction of the control module:

[0075]

[0076] in is the count value of the readout control module for the readout data, which is a count value starting from 0; B is the bit width of the output data, M is the number of output channels, b is the number of bits of the readout data, N is the number of storage blocks, and i is an integer starting from 0. When the lowest dimension of the input data dimension and the output data dimension changes, the readout control module generates the read address for each storage block in the data cache module according to the following formula (wherein each parameter is the same as above) according to the instruction of the control module:

[0077] .

[0078] Each read address uses a different offset, so that the output data is read and output in order from low to high in the output dimension.

[0079] For the process of reading data cache, the reference value given above is still used, and the conversion process for the input dimension of NDHWC and the output dimension of NCDHW is as follows: At the first moment of reading, the read control module will read the data at address 0 of bank0, that is, the data from (0, 0, 0, 0, 0) to (0, 0, 0, 0, 7) represented by the NDHWC dimension. At the same moment, the read control module will read the data at address 1 of bank1, that is, the data from (0, 0, 0, 1, 0) to (0, 0, 0, 1, 7) represented by the NDHWC dimension... At the same moment, the read control module will read the data at address 31 of bank31, that is, the data from (0, 0, 0, 31, 0) to (0, 0, 0, 31, 7) represented by the NDHWC dimension. Finally, the read control module will arrange the data according to the W dimension into 512 bits of data from (0,0,0,0,0) to (0,0,0,0,31) represented by the NCDHW dimension as the data corresponding to output address 0; 512 bits of data from (0,1,0,0,0) to (0,1,0,0,31) as the data corresponding to output address 1; 512 bits of data from (0,2,0,0,0) to (0,2,0,0,31) as the data corresponding to output address 2; 512 bits of data from (0,3,0,0,0) to (0,3,0,0,31) as the data corresponding to output address 3; The 512-bit data from (0,4,0,0,0) to (0,4,0,0,31) is used as the data corresponding to output address 4; the 512-bit data from (0,5,0,0,0) to (0,5,0,0,31) is used as the data corresponding to output address 5; the 512-bit data from (0,6,0,0,0) to (0,6,0,0,31) is used as the data corresponding to output address 6; the 512-bit data from (0,7,0,0,0) to (0,7,0,0,31) is used as the data corresponding to output address 7.

[0080] The control module instructs the readout control module to work and generates the output address corresponding to each output channel according to the count of the output counter of each dimension of each output channel and the data step of each dimension. When the control module starts the output, it starts an output counter with an initial value of 0 for each dimension to represent the dimensional coordinates used for each piece of data, and uses the dimensional coordinates and the step information of each dimension to calculate the corresponding output address. For example, in the case of using five dimensions of N, D, H, W, and C, each output address calculation can be given by the following formula:

[0081]

[0082] in , , , as well as Represents the output counter value corresponding to each dimension of the output data in the NDHWC format, and , , , as well as Indicates the step size of each output dimension. For the case where the output data dimension is NCDHW The value of is 1.

[0083] When the input data dimension and the output data dimension do not change, or when the lowest dimension of the input data and the output data do not change, the control module adjusts the output counters of each dimension in the following manner to control each output channel to output data when outputting data: each output channel writes data in the order of the input data dimension from low to high. That is, after each data output, the output counter of the lowest dimension corresponding to the input is increased by the single output data amount (the size of which is related to the width of each data and the number of channels). When the count value of the output counter corresponding to the lowest dimension of the input data reaches the data size of the dimension, the output counter of the lowest dimension is cleared and the output counter of the upper dimension is increased by 1, and then the above process is repeated until all data is output.

[0084] In the case where the lowest dimension of the input data dimension and the output data dimension changes, the control module adjusts the output counters of each dimension in the following manner when outputting data to control the output data of each output channel: each output channel writes data in the order of the output data dimension from low to high. That is, after each data is output, the output counter of the lowest dimension corresponding to the output data increases the single output data amount (the size of which is related to the width of each data and the number of channels). When the count value of the output counter of the lowest dimension corresponding to the output data reaches the data size of the dimension, the output counter of the lowest dimension is cleared and the output counter of the upper dimension in the corresponding output data dimension is carried. If the upper dimension of the output data dimension happens to be the lowest dimension of the input data dimension, the carry number is the single output data amount (the size of which is related to the data width of each output data), and in other cases, the carry number is 1. The above process is repeated until all data are output.

[0085] It should be understood that although the above introduction of dimensionality transformation takes the five dimensions of N, D, H, W, and C as an example for input data and output data, the input data and output data are not limited to five dimensions. Instead, any one or more combinations of the above dimensions (for example, a combination of four dimensions, a combination of three dimensions, or a combination of two dimensions, etc.) can be adopted. As long as the number of dimensions of the input data and the output data is the same and only the dimension arrangement is different, the dimensional transformation device of the embodiment of the present application introduced above is applicable.

[0086] Figure 5 A functional module structure diagram of a dimensionality conversion device according to another embodiment of the present invention is given. Figure 1 The difference between the device shown is that the control module is responsible for generating input addresses and output addresses respectively through a dedicated input address generator (also called input walker) and an output address generator (also called output walker). The control module controls and schedules other modules according to the configuration information received from the external control unit. The input walker can generate an input address based on the information provided by the control module in order to request read data from the previous level module and instruct the write control module to initialize the internal state in preparation for receiving input data from the previous level module. The output walker can generate an output address based on the information provided by the control module in order to request write data to the next level module. The remaining modules are combined with the above Figure 1 The modules described are similar and will not be described again here.

[0087] It can be seen that the dimension transformation device according to the embodiment of the present invention can realize data dimension transformation during data movement or data transmission, thereby reducing the computational load of the processor. In addition, the dimension transformation device adopts a data cache composed of multiple small on-chip RAMs. It is not necessary to read all the feature data transferred from one neural network to another or from one layer of a neural network to another layer into the cache and then convert them. Instead, it starts from each dimension of the feature data and can output part of the data while reading the data of the corresponding dimension according to the configuration information. This not only saves the area overhead of the on-chip cache but also does not affect the efficiency of data transmission. The above-mentioned dimension transformation device can be directly integrated on the processor chip, or can be integrated in the DMA module of the processor for use.

[0088] In some other embodiments of the present invention, a processor for a neural network is provided, which includes the dimension transformation device described above in conjunction with the accompanying drawings. In the processor, tasks of multiple threads are simultaneously run on different computing cores of the processor, and different computing cores perform different calculations according to instructions. The data processed by each computing core and the calculation results are temporarily stored in the internal on-chip cache, and the dimension transformation device described above is used to perform data transfer between the on-chip cache of the processor and the off-chip memory.

[0089] References in this specification to "various embodiments," "some embodiments," "one embodiment," or "an embodiment," etc. refer to a particular feature, structure, or property described in conjunction with the embodiment being included in at least one embodiment. Thus, the appearances of the phrases "in various embodiments," "in some embodiments," "in one embodiment," or "in an embodiment," etc., in various places throughout the specification do not necessarily refer to the same embodiment. Furthermore, particular features, structures, or properties may be combined in any suitable manner in one or more embodiments. Thus, particular features, structures, or properties shown or described in conjunction with one embodiment may be combined in whole or in part with features, structures, or properties of one or more other embodiments without restriction, as long as the combination is not illogical or inoperable.

[0090] The expressions of "including", "having" and similar meanings in this specification are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or devices. "One" or "an" does not exclude multiple situations. In addition, the elements in the drawings of this application are for illustrative purposes only and are not drawn to scale.

[0091] Although the present application has been described through the above-mentioned embodiments, the present application is not limited to the embodiments described herein, and also includes various changes and modifications made without departing from the scope of the present application.

Claims

1. A dimension transformation device that is friendly to on-chip cache, comprising a control module, a data cache module composed of multiple storage blocks, a write control module and a read control module, wherein: The control module is configured to obtain configuration information related to the instruction while receiving the data handling instruction to be processed, wherein the configuration information at least includes base address information of input data and output data, dimension information of input data and output data, data size and data step length of each dimension of input data and output data; The control module is configured to generate corresponding input addresses using the received configuration information to read corresponding data in order from low dimension to high dimension of the output data dimension, wherein the control module is also configured to include a set of input counters, and when the control module requests to read data from the external storage unit, the input address generated is calculated based on the base address information of the input data, the size information of each dimension data, the step information of each dimension data, and the count value of the input counter of each dimension; the control module is also configured to adjust the input counter to read data in order from low dimension to high dimension of the output data dimension when it is detected that the lowest dimension of the input data dimension and the output data dimension changes, including: after each input data request, the input counter of the lowest dimension corresponding to the output data increases by 1; when the count value of the input counter corresponding to the lowest dimension of the output data reaches the data size of the dimension, the input counter of the dimension is cleared and carried to the input counter of the upper dimension in the output data dimension; if the upper dimension is the lowest dimension of the input data dimension, the carry number is the amount of data input at a single time, otherwise the carry number is 1; repeat the above process until all data are read; The write control module is configured to write input data from an external storage unit into the data cache module; The readout control module is configured to read data from the data cache module and output the data according to the instruction of the control module.

2. The device according to claim 1, wherein each storage block in the data cache module is an on-chip random access memory, and the number of the storage blocks should at least be sufficient to divide the preset bit width of the input data and the bit width of the output data, wherein the bit width of the input data is the same as the bit width of the output data.

3. The device according to claim 1, wherein the device comprises a plurality of output channels, and the number of the output channels should at least be sufficient to divide a preset bit width of the input data and a preset bit width of the output data.

4. The device according to claim 3, wherein each of the output channels corresponds to an output address, the data within each of the output channels is continuous, and wherein the depth of each of the storage blocks is at least equal to or greater than the ratio of the data bit width of each of the output channels to the bit width of a single data.

5. The device according to claim 1, wherein the dimensional data of the input data and the output data are stored in a linear arrangement from low dimension to high dimension.

6. According to the device according to claim 1, the control module is also configured to include multiple groups of output counters. When the input data dimension and the output data dimension do not change, or when the lowest dimension of the input data and the output data do not change, the control module adjusts the output counters of each dimension in the following manner to control each output channel to output data when outputting data: each output channel writes data in the order of the input data dimension from low to high.

7. The apparatus according to claim 6, wherein each output channel writes out data in the order of input data dimensions from low to high, comprising: After each data is output, the output counter of the lowest dimension corresponding to the input is increased by the amount of single output data. When the count value of the output counter corresponding to the lowest dimension of the input data reaches the data size of this dimension, the output counter of the lowest dimension is cleared and the output counter of the upper dimension is incremented by 1. The above process is then repeated until all data is output.

8. According to the device according to claim 1, the control module is also configured to include multiple groups of output counters. When the lowest dimension of the input data dimension and the output data dimension changes, the control module adjusts the output counters of each dimension in the following manner when outputting data to control the output data of each output channel: each output channel writes data in order from low to high output data dimensions.

9. The apparatus according to claim 8, wherein each output channel writes out data in order of output data dimension from low to high, comprising: After each data is output, the output counter of the lowest dimension corresponding to the output data increases the single output data amount. When the count value of the output counter of the lowest dimension corresponding to the output data reaches the data size of this dimension, the output counter of the lowest dimension is cleared and the output counter of the previous dimension in the corresponding output data dimension is carried forward. If the previous dimension of the output data dimension happens to be the lowest dimension of the input data dimension, the carry number is the single output data amount, and in other cases the carry number is 1. The above process is repeated until all data are output.

10. The apparatus according to claim 1, wherein the write control module is further configured to: When the input data dimension and the output data dimension have not changed, or when the lowest dimension of the input data and the output data have not changed, each received input data is counted, and the depth of each storage block in the data cache module is modulo the current count value to generate a write address for each storage block, and the currently received input data is written to each storage block of the data cache module accordingly.

11. The apparatus according to claim 1, wherein the write control module is further configured to: When the lowest dimension of the input data dimension and the output data dimension changes, count the input data received each time; Perform a bit cyclic right shift operation on the currently received input data. The number of right shift bits is the current count value multiplied by B / N, where N is the number of storage blocks in the data cache module and B is the bit width of the input data; The depth of each storage block in the data cache module is modulo the current count value to generate a write address for each storage block, and the processed data is written into each storage block of the data cache module accordingly.

12. The device according to claim 3, wherein the read control module is further configured to generate a read address for each storage block in the data cache module according to the following formula according to the instruction of the control module when the lowest dimension of the input data dimension and the output data dimension changes: in It represents the count of the amount of data read out, which is a count value starting from 0; B is the bit width of the output data, M is the number of output channels, b is the number of bits of a single data, i is an integer starting from 0, and N is the number of storage blocks.

13. A processor for a neural network, comprising a dimensionality transformation device according to any one of claims 1 to 12, which is used to transfer data between an on-chip cache and an off-chip memory of the processor.

Citation Information

Patent Citations

  • Line transpose architecture design method based on two-dimensional FFT (Fast Fourier Transform) processor

    CN106021182A

  • Hardware accelerator engine

    CN108268943A