Direct memory access apparatus, operation method and data processing apparatus
By utilizing the instruction storage module, read module, and data conversion module of the direct memory access device, the problem of decreased CPU performance in shaping feature data is solved, thereby improving memory bandwidth utilization and the overall performance of the neural network processor.
Patent Information
- Application Number
- PCT/CN2024/129387
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-26
- Filing Date
- 2024-11-01
- Publication Date
- 2025-10-30
AI Technical Summary
Traditional CPU processing of feature data shaping operations leads to performance degradation and low memory bandwidth utilization, affecting the overall performance of neural network processors.
It employs a direct memory access device, including an instruction storage module, a read module, a data conversion module, and a write module. Through the internal data conversion module and instruction storage module, it utilizes an instruction-based operation mechanism to process feature data, reducing software design complexity and improving memory bandwidth utilization and the working efficiency of the CPU and NPU.
It significantly improves the efficiency of the CPU and NPU, reduces the demand for memory bandwidth, and enhances the efficiency of data processing.
Smart Images

Figure CN2024129387_30102025_PF_FP_ABST
Abstract
Description
Direct memory access device and its operating method, data processing device
[0001] This application claims priority to Chinese Patent Application No. 202410510747.5, filed on April 26, 2024, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0002] Embodiments of this disclosure relate to a direct memory access device and operating method, and a data processing device. Background Technology
[0003] In traditional neural networks, the shaping (or reorganization) of feature data is typically handled by the CPU. However, this approach significantly degrades CPU performance and has a substantial impact on memory bandwidth. Furthermore, the low data transformation efficiency severely affects the overall performance of the neural network processor. Therefore, a new technology is needed to address these issues and improve data processing efficiency and system performance.
[0004] Summary of the Invention
[0005] At least one embodiment of this disclosure provides a direct memory access (DMI) device, which includes an instruction storage module, a read module, a data conversion module, and a write module. The instruction storage module is configured to store operation instructions for the DMI device; the read module is configured to read first read data according to a first operation instruction stored in the instruction storage module; the data conversion module is configured to convert the first read data into first write data; and the write module is configured to output the first write data according to the first operation instruction.
[0006] For example, in at least one embodiment of the direct memory access apparatus provided in this disclosure, the direct memory access apparatus further includes an instruction fetching module. The instruction fetching module is configured to receive a first operation instruction to be operated and store the received first operation instruction in the instruction storage module.
[0007] For example, in a direct memory access apparatus provided in at least one embodiment of this disclosure, the read module is further configured to read a second operation instruction to be operated and store the read second operation instruction in the instruction storage module.
[0008] For example, in at least one embodiment of the direct memory access apparatus provided in this disclosure, the direct memory access apparatus further includes a data prefetch storage module. The data prefetch storage module is coupled between the read module and the instruction storage module and configured to store a second operation instruction read by the read module and to be stored in the instruction storage module.
[0009] For example, in at least one embodiment of the direct memory access apparatus provided in this disclosure, the direct memory access apparatus further includes a first multiplexing unit, wherein the instruction fetching module and the read module are coupled to the instruction storage module through the first multiplexing unit, and the first multiplexing unit is configured to perform storage management on operation instructions to be stored in the instruction storage module.
[0010] For example, in a direct memory access apparatus provided in at least one embodiment of this disclosure, the first operation instruction includes a source address, a destination address, a data length, and format parameters for data format conversion required for data transfer.
[0011] For example, in a direct memory access device provided in at least one embodiment of this disclosure, the format parameters include a data read step size and a data storage step size, the data read step size and the data storage step size are different, the read module reads the first read data from the first external device through the data read step size, the read module or the write module converts the first read data into the first write data based on the data storage step size, and the first read data and the first write data have different data dimension arrangement formats.
[0012] For example, in a direct memory access device provided in at least one embodiment of this disclosure, the first read-in data and the first write-out data include feature data used in a neural network, the dimensions of which include height, width, and number of channels.
[0013] For example, in at least one embodiment of the direct memory access device provided in this disclosure, the instruction fetch module receives an operation instruction to be operated via an advanced peripheral bus. The direct memory access device further includes a read address port and a read data port, as well as a write address port and a write data port. The read address port and the read data port are coupled to the read module, and the direct memory access device uses the read address port and the read data port to obtain the first read data from a first external device. The write address port and the write data port are coupled to the write module, and the direct memory access device uses the write address port and the write data port to write the first write data to a second external device.
[0014] For example, in a direct memory access device provided in at least one embodiment of this disclosure, the read module includes a read address state machine and a read data state machine. The read address state machine is configured to generate read address information of the first read data based on the first operation instruction, and the read data state machine reads the first read data from a first external device according to the read address information.
[0015] For example, in a direct memory access apparatus provided in at least one embodiment of this disclosure, the read address state machine is coupled to the instruction storage module and obtains the first operation instruction from the instruction storage module, and the read data state machine is coupled to the data conversion module and writes the read first read data into the data conversion module.
[0016] For example, in a direct memory access device provided in at least one embodiment of this disclosure, the write module includes a write address state machine and a write data state machine. The write address state machine is configured to generate write address information of the first write data based on the first operation instruction, and the write data state machine writes the first write data to a second external device according to the write address information.
[0017] For example, in a direct memory access apparatus provided in at least one embodiment of this disclosure, the write address state machine is coupled to the instruction storage module and obtains the first operation instruction from the instruction storage module, and the write data state machine is coupled to the data conversion module and obtains the first written data from the data conversion module.
[0018] For example, in a direct memory access apparatus provided in at least one embodiment of this disclosure, the data conversion module includes a read data storage unit coupled between the read module and the write module and configured to store data written by the read module.
[0019] For example, in a direct memory access apparatus provided in at least one embodiment of this disclosure, the data conversion module includes a data partitioning format conversion module, which is configured to partition the data length of at least one dimension of the first read data having multiple dimensions.
[0020] For example, in a direct memory access apparatus provided in at least one embodiment of this disclosure, the data conversion module further includes a first cache unit, an arbitration unit, a second multiplexing unit, and a second storage unit. The data partitioning format conversion module includes at least one conversion channel, each of the at least one conversion channel being configured to partition the data length of at least one dimension of the first read-in data. The first cache unit is coupled to the at least one conversion channel through the arbitration unit and is configured to store the first read-in data obtained from the read-in module. The second storage unit is coupled to the at least one conversion channel through the second multiplexing unit and is configured to store data obtained from the second multiplexing unit and write the stored data to the write-out module.
[0021] For example, in a direct memory access apparatus provided in at least one embodiment of this disclosure, each of the at least one conversion channel includes a first shift unit, a partitioning unit, a plurality of partitioned conversion storage units, a conversion state machine, a data merging unit, and a second shift unit. The first shift unit obtains data from an arbitration unit and writes the shifted data into the partitioning unit. The partitioning unit divides the data received from the first shift unit into different parts and writes them into the plurality of partitioned conversion storage units respectively. Each of the plurality of partitioned conversion storage units stores data at a specific position into the data merging unit according to the data partitioning information of the conversion state machine. The data merging unit performs a merging and combining operation on the data received from the plurality of partitioned conversion storage units and writes it into the second shift unit. The second shift unit shifts and compacts the data and sends it out of the conversion channel.
[0022] At least one embodiment of this disclosure also provides a method for operating a direct memory access device, comprising: in response to receiving a first operation instruction, storing the first operation instruction in an instruction storage module of the direct memory access device; based on the first operation instruction, acquiring first read-in data through the read-in module and converting the first read-in data into first write-out data through the data conversion module, and outputting the first write-out data through the write-out module.
[0023] At least one embodiment of this disclosure also provides a data processing apparatus, the data processing apparatus including the direct memory access apparatus of any of the above embodiments. Attached Figure Description
[0024] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure.
[0025] Figure 1A shows a schematic diagram of a direct memory access device provided in at least one embodiment of the present disclosure.
[0026] Figure 1B shows a detailed schematic diagram of an example of the direct memory access device in Figure 1A.
[0027] Figure 2 shows a schematic diagram of a data partitioning format conversion module provided in at least one embodiment of the present disclosure.
[0028] Figure 3 illustrates a schematic diagram of reading data from a predetermined area according to at least one embodiment of the present disclosure.
[0029] Figure 4 shows a schematic diagram of dividing channel data length according to at least one embodiment of the present disclosure.
[0030] Figure 5 shows a schematic diagram of a merging unit obtaining data from multiple partitioned conversion storage units T_MEM, according to at least one embodiment of the present disclosure.
[0031] Figure 6 shows a flowchart illustrating an operation method provided in at least one embodiment of this disclosure.
[0032] Figure 7 shows a schematic diagram of a data processing apparatus provided in at least one embodiment of the present disclosure.
[0033] Figure 8 shows a schematic diagram of a non-transitory storage medium provided in at least one embodiment of the present disclosure. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0035] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as “comprising” or “including” mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as “connected” or “linked” are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as “upper,” “lower,” “left,” and “right” are used only to indicate relative positional relationships; these relative positional relationships may change accordingly when the absolute position of the described objects changes.
[0036] To keep the following description of embodiments of the present disclosure clear and concise, detailed descriptions of certain known functions and components are omitted.
[0037] A neural network is an artificial intelligence model that mimics the structure and function of neurons in the human brain. It consists of a large number of artificial neurons (nodes) that transmit information to each other through connections (weights), forming a complex network structure. A neural network typically includes an input layer, hidden layers, and an output layer. The input layer receives raw data, the output layer produces predictions, and the hidden layers process information and extract features between the input and output layers. Neural networks perform prediction and classification tasks by learning patterns and regularities in a dataset. During training, the neural network continuously adjusts the connection weights using the backpropagation algorithm to minimize the error between the prediction and the true label. This learning method allows the neural network to gradually optimize its parameters and improve its generalization ability to new data.
[0038] Currently, neural networks are widely used in fields such as computer vision, natural language processing, and speech recognition. For example, convolutional neural networks (CNNs) perform exceptionally well in image recognition, while recurrent neural networks (RNNs) have important applications in natural language processing. Furthermore, the development of deep learning technology has also driven the development of neural networks, with deep neural networks (DNNs) demonstrating powerful capabilities in processing complex data and tasks.
[0039] The core idea of Convolutional Neural Networks (CNNs) is to construct complex functions by fitting simple functions layer by layer, thereby extracting features from the input image. Due to the nature of CNNs, a large amount of intermediate data (feature data) is generated during feature extraction, and the feature extraction of each layer depends on the feature data of the previous layer. To significantly improve the computing power of the CNN network processor (NPU), multi-channel and multi-convolutional kernel parallel processing methods are used in the hardware implementation. To ensure the computational efficiency of the NPU, the feature data of the previous layer needs to be shaped when calculating the features of the next layer to meet the requirements of the hardware design architecture.
[0040] Data shaping involves operations such as remapping, splitting, and merging feature data. Traditional CPU processing methods place high demands on CPU performance. Furthermore, discontinuous data access reduces memory (e.g., Double Data Rate, DDR) bandwidth utilization, increasing memory bandwidth requirements. Simultaneously, large-scale data transformations consume significant time, impacting the performance of neural processing units (NPUs, such as Tensor Processing Units (TPUs) and AI accelerators).
[0041] The inventors of this disclosure noted that when performing operations such as feature data remapping, partitioning, and merging during CNN processing, relying solely on the CPU often consumes significant CPU resources, leading to a substantial performance degradation. Simultaneously, these intensive data operations have high memory bandwidth requirements, easily creating bandwidth bottlenecks and further limiting the ability to quickly and efficiently transfer data between memory and the CPU. Furthermore, since the CPU is relatively inefficient at handling such tasks, this directly impacts the overall efficiency of the NPU system, significantly slowing down the inference and training speed of the entire deep learning model.
[0042] Therefore, a dedicated module for processing feature data shaping is needed to improve the efficiency of the CPU and NPU and achieve a more efficient data processing flow.
[0043] At least one embodiment of this disclosure provides a direct memory access (DMI) device, which includes an instruction storage module, a read module, a data conversion module, and a write module. The instruction storage module is configured to store operation instructions for the DMI device; the read module is configured to read first read data according to the first operation instruction stored in the instruction storage module; the data conversion module is configured to convert the first read data into first write data; and the write module is configured to output the first write data according to the first operation instruction.
[0044] In the direct memory access apparatus provided in the above embodiments of this disclosure, by setting up a data conversion module and an instruction storage module dedicated to processing feature data (e.g., shaping or reassembling) inside the direct memory access apparatus, and by adopting an instruction-based operation mechanism for the data conversion module, and storing operation instructions for the data conversion module in the instruction storage module, the design complexity at the software level can be effectively reduced, the actual utilization rate of memory bandwidth can be improved, the system's demand for memory bandwidth resources can be reduced, and the working efficiency of the CPU and NPU can be significantly improved.
[0045] At least one embodiment of this disclosure also provides an operation method for a direct memory access device, comprising: in response to receiving a first operation instruction, storing the first operation instruction in an instruction storage module of the direct memory access device; based on the first operation instruction, acquiring first read-in data through a read-in module and converting the first read-in data into first write-out data through a data conversion module, and outputting the first write-out data through a write-out module.
[0046] The technical effects of the operation method of the direct memory access device according to the above embodiments of this disclosure are the same as the technical effects of the direct memory access device described above, and therefore will not be repeated.
[0047] The various embodiments of this disclosure will now be described with reference to specific examples.
[0048] Figure 1A shows a schematic diagram of a direct memory access device provided in at least one embodiment of the present disclosure. Figure 1B shows a detailed schematic diagram of an example of the direct memory access device in Figure 1A.
[0049] As shown in Figure 1A, the Direct Memory Access (EDMA) device includes an instruction storage module 110, a read module 120, a data conversion module 130, and a write module 140. Here, the instruction storage module 110 is configured to store operation instructions for the direct memory access device; the read module 120 is configured to read first read data according to the first operation instruction stored in the instruction storage module 110; the data conversion module 130 is configured to convert the first read data into first write data; and the write module 140 is configured to output the first write data according to the first operation instruction.
[0050] For example, as shown in Figure 1A, the instruction storage module 110 can receive or store operation instructions (or descriptors) from outside the direct memory access device (EDMA) via a first bus (e.g., an Advanced eXtensible Interface, APB). For example, the read module 120 can be directly or indirectly coupled to the instruction storage module 110 and a second bus (e.g., an Advanced Peripheral Bus, AXI), and retrieves data from a storage area outside the EDMA via the second bus based on operation instruction (e.g., a first operation instruction) information obtained from the instruction storage module 110. For example, when the instruction storage module 110 has available storage space, the read module 120 can retrieve or prefetch instruction information from a storage area outside the EDMA via the second bus, and store the instruction information in the instruction storage module 110 for use by the EDMA.
[0051] For example, as shown in FIG1A, the read module 120 may be directly or indirectly coupled to the data conversion module 130, and based on the operation instruction (e.g., a first operation instruction) information obtained from the instruction storage module 110, data obtained from a storage area outside the direct memory access device (EDMA) is transmitted to the data conversion module 130. For example, the data conversion module 130 may convert (shape or reassemble) the data based on the first operation instruction. For example, the data conversion module 130 may transmit the processed or converted data to the write module 140 coupled to the data conversion module 130.
[0052] For example, the write module 140 may be coupled to the instruction storage module 110, the data conversion module 130, and a third bus (e.g., an Advanced Peripheral Bus (AXI)). For example, the write module 140 may transfer data obtained from the data conversion module 130 to a memory area outside the EDMA (Electronic Direct Memory Access Device) via the third bus based on operation instructions (e.g., a first operation instruction) obtained from the instruction storage module 110. For example, the first bus, second bus, and third bus may be of the same or different types; the embodiments of this disclosure do not impose any limitations on this, as long as the corresponding functions required by the embodiments of this disclosure can be achieved.
[0053] It should be noted that the connection lines between any modules shown in Figure 1A can be direct or indirect, wired or wireless, and the embodiments of this disclosure do not impose any limitations on this. Furthermore, although some modules in Figure 1A do not show connection lines, these modules may still be connected. For example, the instruction storage module 110 and the data conversion module 130 can be directly connected or communicate with each other, allowing the data conversion module 130 to obtain operation instructions for data conversion from the instruction storage module 110. The accompanying drawings, including Figure 1A, shown in the embodiments of this disclosure are illustrative and should not constitute any limitation on the embodiments of this disclosure.
[0054] Here, "first operation instruction" refers to one of the operation instructions stored in the instruction storage module for the read module, data conversion module, and write module; "first read data" refers to any data that is read in and will be processed by the data conversion module; "first write data" corresponds to "first read data".
[0055] For example, as shown in Figure 1B, a specific example of a direct memory access device could be EDMA (Enhanced Direct Memory Access); another example is that the instruction storage module could be a random access memory (RAM), which could be configured to store operation instructions for the EDMA device. For example, the operation instructions could be in descriptor form. For example, the read module could be configured to read first read data according to a first operation instruction stored in the RAM, the data conversion module could be configured to convert the first read data into first write data, and the write module could be configured to output the first write data according to the first operation instruction.
[0056] It should be noted that although the example in Figure 1B specifically shows that the instruction storage module is a descriptor random access memory with a storage capacity of 192 bits multiplied by (*)16 and a depth of 2, the embodiments of this disclosure are not limited to this. The instruction storage module in the embodiments of this disclosure can be any other memory capable of storing instructions and can be of any type and any capacity.
[0057] Similarly, although other storage areas or units for storing instructions or data shown in all the accompanying drawings, including FIG1B, provided in the embodiments of this disclosure are specifically shown as a certain type (e.g., a first-in-first-out queue FIFO) of storage areas or units, the embodiments of this disclosure do not impose any limitations on the type of these storage areas or units or the size and type of data that can be stored. The embodiments of this disclosure can use any type (e.g., any type other than a first-in-first-out queue FIFO or the same type as the storage area included in the external device (listed below)) and capacity of memory to store the corresponding operation instructions and data.
[0058] For example, the first input data and the first output data include feature data used in the neural network, and the dimensions of the feature data include height, width, and number of channels. For example, the first input data and the first output data include feature data used in a neural network (e.g., a convolutional neural network CNN), and the dimensions of the feature data include height H, width W, and number of channels C. It should be noted that the neural network here can also be other types of neural networks besides convolutional neural networks CNN, and the embodiments of this disclosure do not impose any limitations on this.
[0059] For example, the first operation instruction includes a source address, a destination address, a data length, and format parameters for data format conversion required for data transfer. For example, each descriptor may include a source address, a destination address, a data length, and format parameters for data format conversion required for data transfer by the Direct Memory Access Device (EDMA). For example, the source address indicates the starting position of the data transfer (or transport), the destination address indicates the target position of the data transfer, and the data length indicates the length of data to be transferred. For example, the EDMA controls the data transfer process based on the source address, destination address, data length, and format parameters for data format conversion required for data transfer. Although embodiments of this disclosure specifically exemplify several exemplary pieces of information included in the first operation instruction, embodiments of this disclosure are not limited thereto. Embodiments of this disclosure may also include any other additional information as needed, as long as it enables data shaping.
[0060] For example, the format parameters include the data reading step size and the data storage step size; for example, the data reading step size and the data storage step size are different. The reading module reads the first read data from the first external device through the data reading step size, and the reading module or the writing module converts the first read data into the first write data based on the data storage step size. The first read data and the first write data have different data dimension arrangement formats.
[0061] For example, format parameters used for data format conversion may include a data read step size for the Direct Memory Access Device (EDMA) to read data from the memory of an external device at a predetermined step size, and a data storage step size for storing the read data in a specified storage area at a predetermined step size. For example, the data read step size and the data storage step size may be the same or different. For example, if the EDMA does not remap the data, the data read step size and the data storage step size may be the same. For example, if the EDMA needs to remap the data, the data read step size and the data storage step size may be different. For example, remapping the data may involve reordering the dimensions of data with different dimensions.
[0062] For example, the first or second external device mentioned in the embodiments of this disclosure (see below) may include random access memory (RAM), read-only memory (ROM), flash memory, cache memory, disk storage, tape storage, optical storage, solid state drive (SSD), dynamic random access memory (DRAM), static random access memory (SRAM), dual-port memory, shared memory, virtual memory, etc. The embodiments of this disclosure do not limit the type of external device, as long as it can include a storage area for storing data or instructions and can implement the functions described in the embodiments of this disclosure.
[0063] For example, taking feature data in a neural network as an example, WHC and HWC are formats for representing image data, where W, H, and C represent the three dimensions of the image data, respectively. For example, in the WHC format, the image data is represented in the order of width, height, and number of channels. For example, in the HWC format, the image data is represented in the order of height, width, and number of channels. Although two specific data formats are exemplified here, the embodiments of this disclosure are not limited to these, and the data remapping techniques of the embodiments of this disclosure can be applied to the remapping of any data with different multiple dimensions.
[0064] For example, when a Direct Memory Access Device (EDMA) needs to remap data, the EDMA can convert data in a first format into data in a second format, different from the first format, based on data read step size and data store step size with different values. The EDMA can also convert WHC format data into HWC format data based on data read step size and data store step size with different values. For instance, the EDMA's data read module reads first read data from a first external device based on the data read step size, and the EDMA's data read module or write module converts the first read data into first write data based on the data store step size, where the first read data and the first write data have different data dimension arrangement formats.
[0065] For example, in a Direct Memory Access (DDMA) device, the data read module can first read first input data from a first external device with a data read step size, and then write the first input data to a second external device with a data storage step size different from the data read step size. This achieves the conversion of the first input data into the first output data stored in the second external device. Alternatively, the DDMA device can first read first input data from a first external device with a data read step size, and then write the first input data to a read data storage unit RD_FIFO (see below) with a data storage step size different from the data read step size. This achieves the conversion of the first input data into the first output data stored in the read data storage unit RD_FIFO, and finally, the write module stores the first output data in the second external device.
[0066] For example, in at least one embodiment, the read-in module includes a read address state machine and a read data state machine. The read address state machine is configured to generate read address information for the first read-in data based on a first operation instruction, and the read data state machine reads the first read-in data from a first external device according to the read address information.
[0067] For example, as shown in the example of Figure 1B, the read module may include a read address state machine AR_FSM and a read data state machine RD_FSM. The read address state machine AR_FSM can be configured to generate read address information for reading first read data from an external device (e.g., a first external device) based on a first operation instruction. The read data state machine RD_FSM can read the first read data from the first external device according to the read address information. For example, a storage unit RD_CMD_FIFO may be included between the read address state machine AR_FSM and the read data state machine RD_FSM. The storage unit RD_CMD_FIFO can be used to store relevant information for data transfer sent from the read address state machine AR_FSM to the read data state machine RD_FSM.
[0068] For example, the read module is also configured to read the second operation instruction to be operated and store the read second operation instruction in the instruction storage module. For example, as shown in Figure 1B, the read data state machine RD_FSM can be configured to read the second operation instruction to be operated from an external storage device and store the read second operation instruction in the descriptor ram.
[0069] For example, in at least one embodiment, the direct memory access device further includes a data prefetch memory module, which may be coupled between the read module and the instruction memory module and configured to store a second operation instruction read by the read module and to be stored in the instruction memory module. For example, as shown in FIG1B, the data prefetch memory module may be a prefetch FIFO, which may be coupled between the read data state machine RD_FSM and the descriptor ram and configured to store the second operation instruction read by the read data state machine RD_FSM and to be stored in the descriptor ram. For example, the second operation instruction may be first stored in memory outside the direct memory access device EDMA, and then read into the direct memory access device EDMA via the read data state machine RD_FSM when needed.
[0070] For example, in at least one embodiment, the data conversion module includes a read data storage unit coupled between the read-in module and the write-out module and configured to store data written by the read-in module.
[0071] For example, as shown in Figure 1B, the data conversion module may include a read data storage unit RD_FIFO. The read data storage unit RD_FIFO can be coupled between the read data state machine RD_FSM and the write data state machine WD_FSM (see description below) and configured to store data written by the read data state machine RD_FSM. For example, due to address misalignment and narrow transmission situations during EDMA data transfer, RD_FIFO is used as a buffer to compact the narrow data read from RD_FSM into RD_FIFO. This facilitates narrow transmissions and concurrent transmissions encountered during subsequent address-aligned transmissions by WD_FSM, improving bus operation efficiency.
[0072] For example, in at least one embodiment, the data conversion module includes a data partitioning format conversion module, which is configured to partition the data length of at least one dimension of the first input data having multiple dimensions.
[0073] For example, as shown in Figure 1B, the data conversion module may further include one or more (four in Figure 1B as an example) data partitioning format conversion modules N2E (e.g., modules that convert N-format data to E-format data). The data partitioning format conversion module N2E can be configured to partition the data length of at least one dimension of the first input data having multiple dimensions. For example, the data partitioning format conversion module N2E can be configured to partition the channel data length of the first input data having multiple dimensions. It should be noted that the data partitioning format conversion module N2E can also be configured to partition the data length of other dimensions of the first input data having multiple dimensions besides channels. For example, N-format and E-format can respectively represent the data length of C0 (see below) in the storage unit. For example, N-format can represent the data length of C0 (see below) in the storage unit of the data before conversion by the data partitioning format conversion module N2E, and E-format can represent the data length of C0 (see below) in the storage unit of the data after conversion by the data partitioning format conversion module N2E.
[0074] For example, in at least one embodiment, the data conversion module further includes a first cache unit, an arbitration unit, a second multiplexing unit, and a second storage unit. The data partitioning format conversion module includes at least one conversion channel. Each of the at least one conversion channel is configured to partition the data length of at least one dimension of the first read data. The first cache unit is coupled to at least one conversion channel through the arbitration unit and configured to store the first read data obtained from the read module. The second storage unit is coupled to at least one conversion channel through the second multiplexing unit and configured to store the data obtained from the second multiplexing unit and write the stored data to the write module.
[0075] For example, as shown in Figure 1B, the first cache unit can be a first-in-first-out queue, such as labeled as a bridged FIFO, the arbitration unit is labeled as ARBIT, the second multiplexing unit is labeled as MUX 2, and the second storage unit can be a first-in-first-out queue, such as labeled as T_FIFO. For example, the data partitioning format conversion module includes at least one conversion channel (e.g., at least one data partitioning format conversion module N2E), each of the at least one conversion channel being configured to partition the data length of at least one dimension of the first read-in data. For example, each of the at least one conversion channel is configured to partition the channel data length of the first read-in data. For example, the first cache unit (e.g., FIFO) is coupled to at least one conversion channel via the arbitration unit ARBIT and configured to store the first read-in data obtained from the read data state machine RD_FSM, and the second storage unit (e.g., T_FIFO) is coupled to at least one conversion channel via the second multiplexing unit and configured to store data obtained from the second multiplexing unit MUX 2 and write the stored data to the write-out module (e.g., the write data state machine WD_FSM mentioned later). For example, the converted data can be temporarily stored in a second memory unit to facilitate data transfer in the subsequent write data state machine WD_FSM. For example, the arbitration unit ARBIT can be responsible for determining which data has access to memory and when.
[0076] For example, in the conversion channel conversion of the Direct Memory Access Device (EDMA), if the N-format data (see below) is 32-byte aligned data, but the data bus width is 512 bits (64 bytes), the bridge FIFO realizes the alignment from 32 bytes to 64 bytes encountered in non-address aligned transmission, preparing for the data conversion of the subsequent data partitioning format conversion module N2E.
[0077] For example, since there is a brief data read vacuum period when the data partitioning format conversion module N2E performs data conversion, and there may be a delay in AXI bus data reading, the embodiments of this disclosure adopt a parallel conversion mode of multiple (e.g., 4-way) data partitioning format conversion modules N2E to improve the conversion efficiency of the direct memory access device EDMA.
[0078] For example, in at least one embodiment, the direct memory access device further includes an instruction fetch module. For instance, the instruction fetch module may be configured to receive a first operation instruction to be operated and store the received first operation instruction in an instruction storage module.
[0079] For example, as shown in Figure 1B, the instruction fetch module can be a control status register (CSR). For example, the control status register can be configured to receive a first operation instruction to be operated and store the received first operation instruction in the descriptor RAM. For example, the control status register can be responsible for parsing data from the Advanced Peripheral Bus (APB) and sending control signals to other modules, or obtaining the status of other modules and returning it to the control unit via the APB bus.
[0080] For example, as shown in Figure 1B, the instruction storage area Rd_trigger_fifo can also be coupled between the control status register and the read address state machine AR_FSM.
[0081] For example, as shown in Figure 1B, the instruction storage area wr_trigger_fifo can also be coupled between the control status register and the write address state machine AW_FSM.
[0082] For example, the instruction storage areas Rd_trigger_fifo and wr_trigger_fifo can be two 8-bit wide, 4-bit deep FIFOs (First-In-First-Out queues) used to store trigger information issued by the host computer. This trigger information may include whether a reset of the prefetch start address is required, the descriptor entry number, whether the data in the descriptor RAM is valid, and the trigger type. For example, when a trigger instruction from the host computer is detected, the trigger information is first stored in the instruction storage areas Rd_trigger_fifo and wr_trigger_fifo, and then RD_FSM and WR_FSM will read the corresponding information and execute the corresponding commands.
[0083] For example, the instruction storage area Rd_trigger_fifo can store trigger information from the control status register and then send this trigger information to the read address state machine for data transfer operations. Alternatively, the instruction storage area Rd_trigger_fifo can be omitted between the control status register and the read address state machine AR_FSM, and the control status register can be directly coupled to the read address state machine AR_FSM. Similarly, the instruction storage area wr_trigger_fifo can store trigger information from the control status register and then send the operation instruction to be performed to the write address state machine for data transfer operations. Again, the instruction storage area wr_trigger_fifo can be omitted between the control status register and the write address state machine AW_FSM, and the control status register can be directly coupled to the write address state machine WD_FSM.
[0084] For example, the instruction fetch module receives the operation instructions to be performed via the Advanced Peripheral Bus (APB). For instance, as shown in Figure 1B, the control status register can receive the first operation instruction to be performed via the APB. For example, the APB can be the control channel communication interface between the external CPU and the EDMA (External Memory Access Device), responsible for configuring the EDMA's operating mode and data reshape descriptor, as well as acquiring its operating status.
[0085] For example, in at least one embodiment, the write module includes a write address state machine and a write data state machine. The write address state machine is configured to generate write address information for the first write data based on a first operation instruction, and the write data state machine writes the first write data to a second external device according to the write address information.
[0086] For example, as shown in Figure 1B, the write module may include a write address state machine AW_FSM and a write data state machine WD_FSM. For example, the write address state machine AW_FSM may be configured to generate write address information for first write data based on a first operation instruction, and the write data state machine WD_FSM may write the first write data to a second external device according to the write address information. For example, a storage unit, such as a first-in-first-out queue WR_CMD_FIFO, may be included between the write address state machine AW_FSM and the write data state machine WD_FSM. This storage unit WR_CMD_FIFO can be used to store information related to data transfer sent from the write address state machine AW_FSM to the write data state machine WD_FSM.
[0087] For example, storage units such as the first-in-first-out queues RD_CMD_FIFO and WR_CMD_FIFO can serve as communication links between AR_FSM and RD_FSM, and AW_FSM and WD_FSM, respectively, to support AXI's outstanding available space statistics function. For example, the FIFO size can be 16 bits * 64 depth. For example, storage units RD_CMD_FIFO and WR_CMD_FIFO can be used to store transfer end identifiers, Direct Memory Access Device (EDMA) operating modes (e.g., data transfer modes, N2E conversion modes, etc.), data transfer sizes, etc. It should be noted that while the embodiments of this disclosure show various specific parameters and values, these are merely exemplary, and the embodiments of this disclosure do not impose any limitations on these parameters and values.
[0088] For example, in at least one embodiment, the direct memory access device further includes a first multiplexing unit, through which the instruction fetch module and the read module are coupled to the instruction storage module, and the first multiplexing unit is configured to perform storage management on operation instructions to be stored in the instruction storage module.
[0089] For example, as shown in Figure 1B, the first multiplexing unit is labeled MUX 1. The control status register and the read data state machine RD_FSM can be coupled to the descriptor RAM through the first multiplexing unit MUX 1. The first multiplexing unit MUX 1 can be configured to manage the storage of operation instructions to be stored in the instruction storage module. For example, the first multiplexing unit MUX 1 can be configured to manage how operation instructions from the control status register and the prefetch FIFO are stored in the descriptor RAM.
[0090] For example, in at least one embodiment, the direct memory access device further includes a read address port, a read data port, a write address port, and a write data port.
[0091] The read address port and read data port are coupled to the read module. The direct memory access device (DMI) uses the read address port and read data port to obtain the first read data from the first external device. The write address port and write data port are coupled to the write module. The DMI uses the write address port and write data port to write the first write data to the second external device. In the embodiments of this disclosure, there are no restrictions on the type and structure of the read address port, read data port, write address port, and write data port; they can be based on various existing applicable interface / bus standards, etc.
[0092] For example, a Direct Memory Access (EDMA) device can achieve high-speed data transmission and communication with external devices through the Advanced eXtensible Interface (AXI). In this case, the read address port, read data port, write address port, and write data port are all ports compliant with the AXI standard. This disclosure will not elaborate on the AXI standard. For example, the external device can be a Static Random Access Memory (SRAM), Double Data Rate (DDR), etc. The embodiments of this disclosure are not limited to these. The external device mentioned in the embodiments of this disclosure can be any device with a storage area.
[0093] For example, as shown in Figure 1B, the Direct Memory Access Device (EDMA) may further include a read address port AXI_AR, a read data port AXI_RD, a write address port AXI_WR, and a write data port AXI_WD, all conforming to the AXI standard. The read address port AXI_AR and the read data port AXI_RD can be coupled to the read address state machine AR_FSM and the read data state machine RD_FSM, respectively. The EDMA can acquire first read data from a first external device based on the read address port AXI_AR and the read data port AXI_RD. The write address port AXI_WR and the write data port AXI_WD can be coupled to the write address state machine AW_FSM and the write data state machine WD_FSM, respectively. The EDMA can write first write data to a second external device based on the cooperation of the write address port AXI_WR and the write data port AXI_WD.
[0094] For example, in at least one embodiment, the read address state machine is coupled to the instruction storage module and obtains a first operation instruction from the instruction storage module, and the read data state machine is coupled to the data conversion module and writes the first read data into the data conversion module.
[0095] For example, as shown in Figure 1B, the read address state machine AR_FSM can be coupled to the descriptor ram and obtain the first operation instruction from the descriptor ram, and the read data state machine RD_FSM is coupled to the read data storage unit RD_FIFO and the bridge FIFO and writes the first read data into the read data storage unit RD_FIFO and / or the bridge FIFO.
[0096] For example, in at least one embodiment, the write address state machine is coupled to the instruction storage module and obtains a first operation instruction from the instruction storage module, and the write data state machine is coupled to the data conversion module and obtains the first write data from the data conversion module.
[0097] For example, as shown in Figure 1B, the write address state machine AW_FSM can be coupled to the descriptor ram and obtain the first operation instruction from the descriptor ram, and the write data state machine WD_FSM can be coupled to the second memory unit T_FIFO and obtain the first write data from the second memory unit T_FIFO.
[0098] For example, in at least one embodiment, as shown in FIG1B, the Direct Memory Access Device (EDMA) may further include an interrupt controller configured to manage interrupts that may occur during EDMA transfers. For example, when an EDMA data transfer is completed or an error occurs, the interrupt controller generates an interrupt signal to notify the processor to handle the corresponding event. For example, upon receiving the interrupt signal, the processor may suspend its current task, handle the interrupt event, and then continue executing its original task.
[0099] Figure 2 shows a schematic diagram of a data partitioning format conversion module provided in at least one embodiment of the present disclosure.
[0100] For example, as described above, each of the at least one conversion channel includes a first shift unit, a partitioning unit, a plurality of partitioned conversion storage units, a conversion state machine, a data merging unit, and a second shift unit.
[0101] For example, as shown in Figure 2, a conversion channel may include a data partitioning format conversion module N2E. Each data partitioning format conversion module N2E may include a first shift unit (e.g., shift unit 1), a partitioning unit, multiple partitioning conversion storage units T_MEM, a conversion state machine T_FSM, a data merging unit (or merging unit), and a second shift unit (e.g., shift unit 2). For example, shift unit 1 can obtain data from the arbitration unit ARBIT in Figure 1B and write the shifted and compacted data into the partitioning unit. For example, the partitioning unit can divide the data received from shift unit 1 into different parts and write them into multiple partitioning conversion storage units T_MEM respectively. Each T_MEM in the multiple partitioning conversion storage units T_MEM can store the data at the corresponding position into the data merging unit (or merging unit) according to the data partitioning information (Gen_addr) of the conversion state machine T_FSM. The data merging unit can perform a merging and combination operation on the data received from the multiple partitioning conversion storage units T_MEM and then write it into shift unit 2. Shift unit 2 can shift and compact the data and send it out of the conversion channel. For example, shift unit 2 can shift and compact data before sending it to the second multiplexing unit MUX 2.
[0102] For example, if shift unit 1 obtains 512 bits of data from arbitration unit ARBIT, in order to quickly extract the required data from the 512 bits of data, the partitioning unit divides the data into 64 equal parts, each of which is 8 bits, and stores the 64 parts of data into 64 partitioning conversion storage units T_MEM respectively. For example, each partitioning conversion storage unit T_MEM retrieves the data at the corresponding address generated by the conversion state machine T_FSM and stores it in the merging unit. The merging unit then sends the data to unit 2 for data compaction before sending it to the second multiplexing unit MUX 2. It should be noted that the specific number of various units and the specific values shown in Figure 2 are merely exemplary, and the embodiments of this disclosure do not impose any limitations on them.
[0103] Figure 3 illustrates a schematic diagram of reading data from a predetermined area according to at least one embodiment of the present disclosure.
[0104] For example, in order to read data in a predetermined area in Figure 3 (e.g., the data in the channel area corresponding to 0 to 15 shown on the left side of Figure 3), each data read needs to be shifted to take into account the surface step size and line step size, so as to ensure that the required data is read each time, thereby reducing unnecessary readings and increasing the efficiency of data reading.
[0105] For example, the data partitioning format conversion module N2E performs a cyclic right shift operation on the input 512-bit data (compacted or uncompacted) to compact the data, and stores the data obtained from the right shift in the partitioning conversion memory unit T_MEM inside the N2E module. The shift operation is determined by the data format and C0 (see below). When the data format is int8, it shifts right by N*C0 bytes each time (e.g., N is the 0th clock cycle). When the data format is int16 or fp16, it shifts right by 2N*C0 bytes each time. For example, by controlling the read address of the partitioning conversion memory unit T_MEM, the data is written to T_FIFO, and finally written to the external memory unit by WD_FSM, completing the data conversion.
[0106] Figure 4 shows a schematic diagram of dividing channel data length according to at least one embodiment of the present disclosure.
[0107] As shown in Figure 4, the data partitioning and format conversion module N2E converts data from N-format data to E-format data. For example, N-format data has three dimensions: width W, height H, and channels C. For example, N-format data is compacted feature data, and the data in the storage unit follows the order C0->W->H->C, with C0 being 32 bytes in size. Dimension C0 changes the fastest, and dimension C changes the slowest. For example, E-format data in the storage unit follows the order C0->W->H->C, and C0 can be 1 byte, 2 bytes, 4 bytes, 8 bytes, or 16 bytes in size. It should be noted that although the embodiments of this disclosure specifically give the size of C0 for N-format data and E-format data, the embodiments of this disclosure are not limited to this. The size of C0 for N-format data or E-format data can be any applicable size (e.g., 1 byte, 2 bytes, 4 bytes, 8 bytes, 16 bytes, 64 bytes, etc.).
[0108] Figure 5 illustrates a schematic diagram of a merging unit acquiring data from multiple partition conversion storage units T_MEM according to at least one embodiment of the present disclosure. For example, Figure 5 shows an example of the data partition format conversion module N2E converting N-format data with C0 as 32 bytes into E-format data with C0 as 2 bytes.
[0109] For example, referring to Figures 2 and 5, the first column of the serial number in Figure 5 can represent the number of the different storage depths of each partition transition memory unit T_MEM, and the other columns can represent the data stored in the 64 partition transition memory units T_MEM. For example, each column represents the data stored in one corresponding partition transition memory unit T_MEM. For example, ram_0 to ram_63 can represent the 64 partition transition memory units T_MEM in Figure 2. For example, each partition transition memory unit T_MEM in Figure 2 extracts data from each different channel in a predetermined order according to the corresponding address or shift information generated by the transition state machine T_FSM based on the dimension information of the obtained data (e.g., shifting and reading or extracting data in the partition transition memory unit T_MEM according to the shift information generated by the transition state machine T_FSM), and stores the extracted data in the merging unit to merge the data. It should be noted that although the partitioning and conversion memory unit T_MEM is specifically represented as random access memory (RAM) in Figure 5, the embodiments of this disclosure do not impose any limitations on this, and the partitioning and conversion memory unit T_MEM can be any memory unit capable of storing data.
[0110] For example, Figure 5 illustrates dividing the data stored in channel C0 from 32 bytes to 2 bytes. Before the division, each channel C0 includes data C0-C31 (32 bytes in total). After the division, each channel C0 includes data C0-C1 (2 bytes in total). (That is, the data is stored by changing the channel C0 at a rate of 2 bytes per unit.) For example, as shown in Figure 5, based on the corresponding address or shift information generated by the transition state machine T_FSM based on the dimension information of the obtained data, the data from C0 to C63 in each channel C0 can be extracted in units of two adjacent bytes (e.g., C0-C1, C2-C3, ..., C62-C63). Then, the extracted data is sent to the merging unit to merge and reassemble the data (e.g., during storage, the data changes from changing the channel C0 at a rate of 32 bytes to changing the channel C0 at a rate of 2 bytes per unit) to achieve data division.
[0111] For example, when writing data into multiple partitioning and conversion memory cells T_MEM in Figure 5, it is necessary to use shifting to ensure that no two data Cx (e.g., x is any one from 0 to 63) are contained in the same memory column (e.g., ram column). This ensures that within one clock cycle, data Cx and Cx+1 (e.g., x is any one from 0 to 62) can be retrieved from multiple partitioning and conversion memory cells T_MEM and sent to the merging unit for merging, ultimately realizing the partitioning of the data length of channel C0.
[0112] Table 1 shows some parameters that may be included in an operation instruction (or descriptor) received by a direct memory access device according to at least one embodiment of the present disclosure.
[0113] Table 1
[0114] For example, in Table 1, MSB can represent the most significant bit, and LSB can represent the least significant bit. For example, the storage location of the corresponding parameter is determined by the MSB and LSB of the instruction or descriptor. For example, in the access type, RW represents a read / write operation. For example, R represents read, and W represents write. For example, for the default initialization value column, taking 16'd0 and 1'b0 as examples, 16'd0 can represent setting the initial value of the 16-bit decimal number used to store the corresponding parameter to 0, and 1'b0 can represent setting the initial value of the 1-bit binary number used to store the corresponding parameter to 0. For example, for the parameters in the default initialization value column, d can represent decimal, and b can represent binary. It should be noted that the data in Table 1 is only used to represent the parameters that the descriptor or instruction may include in the embodiments of this disclosure, but this disclosure does not impose any restrictions on which specific types of parameters are included or the specific values of each parameter, as long as the EDMA data shaping scheme mentioned in the embodiments of this disclosure can be implemented.
[0115] In the direct memory access apparatus provided in the above embodiments of this disclosure, by setting up a data conversion module and an instruction storage module dedicated to processing feature data shaping inside the direct memory access apparatus, and by adopting an instruction-based operation mechanism for the data conversion module, and storing operation instructions for the data conversion module in the instruction storage module, the design complexity at the software level can be effectively reduced, the actual utilization rate of memory bandwidth can be improved, the system's demand for memory bandwidth resources can be reduced, and the working efficiency of the CPU and NPU can be significantly improved.
[0116] Figure 6 shows a schematic flowchart of an operation method of the above-described direct memory access apparatus provided in at least one embodiment of the present disclosure.
[0117] As shown in Figure 6, in some embodiments of this disclosure, the operation method includes steps S101-S103.
[0118] Step S101: In response to receiving the first operation instruction, the first operation instruction is stored in the instruction storage module of the direct memory access device.
[0119] Step S102: Based on the first operation instruction, the first read data is obtained through the read module and converted into the first write data through the data conversion module.
[0120] Step S103: Output the first written data through the write module.
[0121] For example, the operation method may also include receiving a first operation instruction to be operated through an instruction acquisition module and storing the received first operation instruction in an instruction storage module.
[0122] For example, the operation method may also include reading the second operation instruction to be operated through the read module and storing the read second operation instruction into the instruction storage module.
[0123] For example, the operation method may also include storing a second operation instruction read by the read module and to be stored in the instruction storage module through a data prefetch storage module coupled between the read module and the instruction storage module.
[0124] For example, the operation method may also include storing and managing operation instructions to be stored in the instruction storage module through a first multiplexing unit.
[0125] For example, the operation method may also include reading first input data from a first external device by a data reading step size, and converting the first input data into first output data based on a data storage step size, wherein the first input data and the first output data have different data dimension arrangement formats, and the data reading step size and the data storage step size are different.
[0126] For example, the operation method may also include dividing the data length of at least one dimension of the first input data with multiple dimensions by using a data partitioning format conversion module.
[0127] In the operation method provided by at least one embodiment of the present disclosure, by setting up a data conversion module and an instruction storage module dedicated to processing feature data shaping inside the direct memory access device, and by adopting an instruction-based operation mechanism for the data conversion module, and storing operation instructions for the data conversion module in the instruction storage module, the design complexity at the software level can be effectively reduced, the actual utilization rate of memory bandwidth can be improved, the system's demand for memory bandwidth resources can be reduced, and the working efficiency of the CPU and NPU can be significantly improved.
[0128] At least some embodiments of this disclosure also provide a data processing apparatus, which includes any of the above-described direct memory access devices.
[0129] Figure 7 shows a schematic diagram of a data processing apparatus provided in at least one embodiment of the present disclosure.
[0130] As shown in FIG7, the data processing apparatus 700 according to an embodiment of the present disclosure includes one or more processors 701, one or more memories 702 and a direct memory access device (EDMA) 704. The processors 701, memories 702 and EDMA 704 can be interconnected via a bus 703.
[0131] EDMA 704 can be a direct memory access device mentioned in the embodiments of this disclosure. EDMA 704 can perform any of the operations mentioned in the embodiments of this disclosure (e.g., data transfer, data shaping, etc.) under the control of processor 701.
[0132] Processor 701 can perform various actions and processes according to the program or code stored in memory 702. Specifically, processor 701 can be an integrated circuit chip with signal processing capabilities. For example, processor 701 can be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), off-the-shelf programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the various methods and steps disclosed in the embodiments of this disclosure. For example, processor 701 can be a CPU, NPU, etc. The general-purpose processor can be a microprocessor or any conventional processor, such as an x86 architecture, ARM architecture, RISC-V architecture, etc.
[0133] The memory 702 is used for non-temporary storage of computer-executable instructions, and the processor 701 is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor 701, the operation method provided in at least one embodiment of this disclosure is implemented.
[0134] For example, memory 702 may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus random access memory (DRRAM). It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0135] The data processing apparatus of this disclosure can be implemented as a terminal or a server, for example, it can be used in mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (e.g., vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers, or it can be used as a local area network server, cloud server, etc.
[0136] In the data processing apparatus provided in at least one embodiment of the present disclosure, by setting up a data conversion module and an instruction storage module dedicated to processing feature data shaping inside the direct memory access device, and by adopting an instruction-based operation mechanism for the data conversion module, and storing operation instructions for the data conversion module in the instruction storage module, the design complexity at the software level can be effectively reduced, the actual utilization rate of memory bandwidth can be improved, the system's demand for memory bandwidth resources can be reduced, and the working efficiency of the CPU and NPU can be significantly improved.
[0137] At least one embodiment of this disclosure also provides a non-transitory storage medium for non-transitory storage of computer-executable instructions. For example, when the computer-executable instructions are executed by a processor, they implement the operating method provided in at least one embodiment of this disclosure.
[0138] Figure 8 is a schematic diagram of a non-transitory storage medium provided in some embodiments of the present disclosure. As shown in Figure 8, the non-transitory storage medium 800 can non-transitory store computer-executable instructions 810, which, when executed by a computer, implement the operation methods provided in any embodiment of the present disclosure.
[0139] Similarly, the non-transitory storage medium in the embodiments of this disclosure may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. It should be noted that the memory used in the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0140] The technical effects of the aforementioned non-transitory storage medium are the same as those of the aforementioned operating method, and will not be repeated here.
[0141] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing at least one executable instruction for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0142] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0143] In addition to the exemplary description above, the following points should be noted regarding this disclosure:
[0144] (1) The accompanying drawings of the embodiments of this disclosure only involve the structures involved in the embodiments of this disclosure, and other structures can be referred to the general design.
[0145] (2) For clarity, the thickness and dimensions of layers or structures are enlarged in the drawings used to describe embodiments of the present disclosure. It will be understood that when an element such as a layer, film, region or substrate is referred to as being “above” or “below” another element, the element may be “directly” located “above” or “below” the other element, or there may be intermediate elements present.
[0146] (3) Where there is no conflict, the embodiments of this disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.
[0147] The above description is only a specific embodiment of this disclosure, but the protection scope of this disclosure is not limited thereto. The protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. A direct memory access device, comprising: The instruction storage module is configured to store operation instructions for the direct memory access device; The read-in module is configured to read first read-in data according to a first operation instruction stored in the instruction storage module; The data conversion module is configured to convert the first read-in data into the first written-out data; The write module is configured to output the first write data according to the first operation instruction.
2. The direct memory access device according to claim 1, further comprising: The instruction acquisition module is configured to receive the first operation instruction to be operated and store the received first operation instruction in the instruction storage module.
3. The direct memory access apparatus according to claim 2, wherein, The read-in module is further configured to read the second operation instruction to be operated and store the read second operation instruction into the instruction storage module.
4. The direct memory access device according to claim 3, further comprising: A data prefetch storage module is coupled between the read module and the instruction storage module and configured to store a second operation instruction read by the read module and to be stored in the instruction storage module.
5. The direct memory access apparatus according to claim 2, further comprising a first multiplexing unit, in, The instruction acquisition module and the read-in module are coupled to the instruction storage module through the first multiplexing unit, which is configured to manage the storage of operation instructions to be stored in the instruction storage module.
6. The direct memory access device according to any one of claims 1-5, wherein, The first operation instruction includes the source address, destination address, data length, and format parameters for data format conversion required for data transfer.
7. The direct memory access apparatus according to claim 6, wherein, The format parameters include a data read step size and a data storage step size, which are different from each other. The read module reads the first read data from the first external device according to the data read step size. The read module or the write module converts the first read data into the first write data based on the data storage step size. The first read data and the first write data have different data dimension arrangement formats.
8. The direct memory access device according to any one of claims 1-7, wherein, The first input data and the first output data include feature data used in the neural network, and the dimensions of the feature data include height, width and number of channels.
9. The direct memory access apparatus according to claim 2, wherein, The instruction acquisition module receives operation instructions to be performed via an advanced peripheral bus, and the direct memory access device further includes: The read address port and read data port are coupled to the read module, and the direct memory access device obtains the first read data from the first external device based on the read address port and read data port; The write address port and write data port are coupled to the write module, and the direct memory access device writes the first write data to the second external device based on the write address port and the write data port.
10. The direct memory access device according to any one of claims 1-9, wherein, The read-in module includes a read address state machine and a read data state machine. The read address state machine is configured to generate read address information for the first read data based on the first operation instruction. The read data state machine reads the first read data from the first external device according to the read address information.
11. The direct memory access apparatus according to claim 10, wherein, The read address state machine is coupled to the instruction storage module and obtains the first operation instruction from the instruction storage module. The read data state machine is coupled to the data conversion module and writes the first read data into the data conversion module.
12. The direct memory access apparatus according to claims 1-11, wherein, The write module includes a write address state machine and a write data state machine. The write address state machine is configured to generate write address information for the first written data based on the first operation instruction. The write data state machine writes the first write data to the second external device according to the write address information.
13. The direct memory access apparatus according to claim 12, wherein, The write address state machine is coupled to the instruction storage module and obtains the first operation instruction from the instruction storage module; the write data state machine is coupled to the data conversion module and obtains the first written data from the data conversion module.
14. The direct memory access device according to any one of claims 1-13, wherein, The data conversion module includes a read data storage unit. The read data storage unit is coupled between the read module and the write module and is configured to store the data written by the read module.
15. The direct memory access device according to any one of claims 1-14, wherein, The data conversion module includes a data partitioning format conversion module. The data partitioning format conversion module is configured to partition the data length of at least one dimension of the first read-in data, which has multiple dimensions.
16. The direct memory access apparatus according to claim 15, wherein, The data conversion module further includes a first cache unit, an arbitration unit, a second multiplexing unit, and a second storage unit. The data partitioning format conversion module includes at least one conversion channel, and each of the at least one conversion channel is configured to partition the data length of at least one dimension of the first read-in data. The first cache unit is coupled to the at least one conversion channel via the arbitration unit and configured to store the first read-in data obtained from the read-in module. The second storage unit is coupled to the at least one conversion channel via the second multiplexing unit and configured to store data obtained from the second multiplexing unit and write the stored data to the write-out module.
17. The direct memory access apparatus according to claim 16, wherein, Each of the at least one conversion channel includes a first shift unit, a partitioning unit, multiple partitioned conversion storage units, a conversion state machine, a data merging unit, and a second shift unit. The first shift unit obtains data from the arbitration unit and writes the shifted data into the partitioning unit. The partitioning unit divides the data received from the first shift unit into different parts and writes them into the plurality of partitioning conversion storage units respectively. Each of the plurality of partitioning conversion storage units stores data at a specific position into the data merging unit according to the data partitioning information of the conversion state machine. The data merging unit performs a merging and combination operation on the data received from the plurality of partitioning conversion storage units and writes it into the second shift unit. The second shift unit shifts and compacts the data and sends it out of the conversion channel.
18. A method of operating a direct memory access device according to claim 1, comprising: In response to receiving a first operation instruction, the first operation instruction is stored in the instruction storage module of the direct memory access device; Based on the first operation instruction, the first read data is obtained through the read module and converted into the first write data through the data conversion module; as well as The first written data is output through the write module.
19. A data processing apparatus comprising a direct memory access device according to any one of claims 1-17.
Citation Information
Patent Citations
Direct memory access device, data transmission method and integrated circuit system
CN115168260A
Direct memory access controller, data transmission method and system
CN115905060A
Data transmission method and device, computer equipment and readable storage medium
CN116150057A
Direct memory access device, operation method and data processing device
CN118427136A
Memory Control Method, Memory Control Apparatus, and Image Forming Method That Uses Memory Control Method
US20200210123A1