Data processing method, DMA controller and system on chip

Through the asynchronous operation and parallel processing channels of the DMA controller and the acceleration engine module, task data is pre-read and data read and write instructions and calculation processes are decoupled, which solves the problem of long wait time for the acceleration engine module and improves data processing efficiency and computing performance.

CN120353578APending Publication Date: 2025-07-22PHYTIUM TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510362516.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

During the data transmission between the CPU and memory, the acceleration engine module in the prior art needs to wait for the reply of the previous data read/write instruction to return before performing the calculation task, resulting in small processing bandwidth and low computing performance.

Method used

Through the asynchronous operation of the DMA controller and the data reading and writing unit, command execution unit and computing unit of the acceleration engine module, task data is pre-read and data reading and writing instructions and calculation processes are decoupled, and a parallel processing channel is used to reduce the waiting time.

Benefits of technology

It improves data processing efficiency, reduces the idle time of the calculation unit, and ensures the correctness of the calculation timing and processing speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353578A_ABST
    Figure CN120353578A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses a data processing method, a DMA controller and a system on chip, the method is applied to the DMA controller, and the method comprises the following steps: receiving a data read / write instruction sent by a data read / write unit, and pre-reading task data from a memory according to the data read instruction; providing required task data for the command execution unit; and the data writing unit is used for receiving a calculation result which is sent by the command execution unit and aims at the required task data, and writing the calculation result into a memory according to a data writing instruction received from the data reading and writing unit. According to the technical scheme provided by one or more embodiments, the data processing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a data processing method, a DMA controller, and a system on a chip. Background Art

[0002] When data is transferred between a CPU and a memory, an engine and a DMA controller are usually used to transfer data instead of the processor core of the CPU, which can improve the usage efficiency of the CPU. The processor core receives the operation instructions of the CPU and transmits the operation instructions to the engine, and the engine interacts with the storage medium through the DMA controller.

[0003] In practical applications, when the acceleration engine module executes a data read / write instruction, it needs to wait for the response of the previous data read / write instruction to return before it can execute the calculation task and send a new data read / write instruction, resulting in a long waiting time and a small processing bandwidth of the acceleration engine module, thus leading to a low computing performance of the acceleration engine module.

[0004] In view of this, how to improve the data processing efficiency has become the focus of current research in the computer field. Summary of the Invention

[0005] This application provides a data processing method, a DMA controller, and a system on a chip, which can improve the data processing efficiency.

[0006] In a first aspect of this application, a data processing method is provided, which is applied to a DMA controller. The DMA controller is communicatively connected to an acceleration engine module. The acceleration engine module includes a data read / write unit, a command execution unit, and a calculation unit that are connected in sequence. The method includes: receiving a data read / write instruction sent by the data read / write unit, and prefetching task data from a memory according to a data read instruction. The data read / write instruction is generated by the data read / write unit based on the address information in command stream data; providing the required task data to the command execution unit, so that the command execution unit sends the required task data and a pre-generated calculation task to the calculation unit, where the calculation task is pre-generated by the command execution unit based on the task information in the command stream data; receiving a calculation result of the required task data sent by the command execution unit, and writing the calculation result into the memory according to a data write instruction received from the data read / write unit; where the calculation result is processed by the calculation unit according to the calculation task and the required task data, and is fed back by the calculation unit to the command execution unit.

[0007] The technical solution provided by this embodiment of the present application decouples the processing process of data read / write instructions and the calculation process of task data to perform parallel processing of instruction processing and data calculation. Among them, the DMA controller pre-reads the task data required for subsequent calculation tasks in response to the data read / write instructions, and at the same time commands the execution unit to generate calculation tasks according to the task information in the command stream data. With the help of the calculation unit, the calculation process of the task data is completed, and the calculation result is written to the memory. It can be seen that through the technical solution provided by this embodiment of the present application, the pre-reading of task data and the processing of calculation tasks are performed in an asynchronous working manner, avoiding the situation that the calculation unit is idle due to a large memory read delay, thereby improving the data processing efficiency.

[0008] In a possible implementation manner, the data read / write instruction has a timing identifier, and the timing identifier is used to represent the timing of data reading and data writing in the task flow of the current calculation task. The task flow of the current calculation task is parsed by the data read / write unit based on the received command stream data.

[0009] The technical solution provided by this embodiment of the present application describes the processing of the request sending order when multiple data read / write instructions are involved. The processing timing relationship of multiple data read / write instructions is represented by the timing identifier, so that when facing complex calculations, the asynchronous working method can improve the operation efficiency while still ensuring the correctness of the operation timing.

[0010] In a possible implementation manner, pre-reading task data from the memory according to the data read instruction includes: converting the data read instruction into a read control signal of the bus, and pre-reading the task data from the memory through the bus based on the read control signal; adding the timing identifier carried in the data read instruction to the task data, and caching the task data with the added timing identifier.

[0011] The technical solution provided by this embodiment of the present application describes how the DMA controller pre-reads task data from the memory. Adding the corresponding timing identifier during the pre-reading of task data can ensure that the subsequent command execution unit can read the correct data according to the timing identifier, thus ensuring that the calculation unit can process the data according to the task flow of the calculation task.

[0012] In a possible implementation manner, writing the calculation result into the memory according to the data write instruction received from the data read / write unit includes: caching each calculation result sent by the command execution unit; identifying the timing identifier carried in the currently to-be-processed data write instruction, and querying the calculation result with the timing identifier from the cached each calculation result; writing the queried calculation result into the memory according to the data output address specified by the data write instruction.

[0013] The technical solution provided by this embodiment of the present application describes how to write the calculation result into the memory. The DMA controller queries the corresponding calculation result through the timing identifier of the data write instruction and writes it into the memory to ensure the correct timing of data writing through the timing identifier.

[0014] In a possible implementation, providing the required task data to the command execution unit includes: receiving a data acquisition request sent by the command execution unit, where the data acquisition request includes one or more timing identifiers, and the timing identifier is used to characterize the timing of data reading and data writing in the task flow of the current calculation task, and the task flow of the current calculation task is parsed by the command execution unit based on the task information; and feeding back the task data with the timing identifier to the command execution unit.

[0015] The technical solution provided by this embodiment of the present application further describes the parsing process of the task information by the command execution unit and the timing identification of the task data to be calculated.

[0016] In a possible implementation, the calculation result carries a timing identifier added by the calculation unit, and the timing identifier carried in the calculation result is consistent with the timing identifier carried in the task data that has been processed.

[0017] The technical solution provided by this embodiment of the present application also adds a timing identifier to the calculation result. Subsequently, the data write instruction and the calculation result to be written are associated through the timing identifier, thereby ensuring that the data write instruction can write the correct calculation result.

[0018] In a possible implementation, the DMA controller includes parallel read request channels and read response channels. Among them, after the DMA controller receives a data read instruction sent by the data read / write unit, it converts the data read instruction into a read control signal for the bus, sends the read control signal through the read request channel, and receives the task data returned by the bus through the read response channel.

[0019] The technical solution provided by this embodiment of the present application describes the parallel processing logic of the read request channel and the read response channel of the DMA controller. The sending of the read request no longer needs to wait for the return of the previous read response before it can be sent.

[0020] In a possible implementation, the DMA controller further includes parallel write request channels and write response channels. Among them, the DMA controller generates a write control signal for the bus according to the data write instruction and the calculation result, sends the write control signal through the write request channel, and receives the write response signal returned by the bus through the write response channel.

[0021] The technical solution provided by this embodiment of the present application describes the parallel processing logic of the write request channel and the write response channel of the DMA controller. The sending of the write request no longer needs to wait for the return of the previous write response before it can be sent.

[0022] In a possible implementation, the DMA controller further includes a write data channel. When generating the write control signal for the bus, the DMA controller also generates the write data signal for the bus. After sending the write control signal through the write request channel, the DMA controller sends the write data signal associated with the write control signal through the write data channel.

[0023] The technical solution provided by this embodiment of the present application provides a write data channel in addition to the write request channel and the write response channel, thus providing a way to write the data to be written into the memory.

[0024] A second aspect of the present application provides a DMA controller. The DMA controller is communicatively connected to an acceleration engine module. The acceleration engine module includes a data read / write unit, a command execution unit, and a calculation unit connected in sequence. The DMA controller includes: a data receiving unit, configured to receive the data read / write instruction sent by the data read / write unit, where the data read / write instruction is generated by the data read / write unit based on the address information in the command stream data, and receive the calculation result sent by the command execution unit, where the calculation result is obtained by the calculation unit according to the calculation task and the required task data, and is fed back by the calculation unit to the command execution unit; a data processing unit, configured to prefetch the task data from the memory according to the data read instruction, and write the calculation result into the memory according to the data write instruction received from the data read / write unit; a data sending unit, configured to provide the required task data to the command execution unit, so that the command execution unit sends the required task data and the pre-generated calculation task to the calculation unit, where the calculation task is pre-generated by the command execution unit based on the task information in the command stream data.

[0025] A third aspect of the present application provides a system-on-chip. The system-on-chip includes a DMA controller and an acceleration engine module, where: the acceleration engine module is configured to process the command stream data and interact with the DMA controller based on the processing result; the DMA controller is configured to execute the above data processing method during the interaction with the acceleration engine module. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] To more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0027] FIG. 1(a) and FIG. 1(b) are schematic comparison diagrams before and after the parallel processing of the read request channel and the read response channel provided by an embodiment of the present application;

[0028] FIG. 2(a) and FIG. 2(b) are schematic comparison diagrams before and after the parallel processing of the write request channel and the write response channel provided by an embodiment of the present application;

[0029] Figure 3 It is a schematic diagram of the method for data prefetching and data processing in an embodiment of the present application;

[0030] Figure 4 It is a schematic diagram of the function of data prefetching and data write-back provided by another embodiment of the present application;

[0031] Figure 5 It is a schematic diagram of the function of data processing and calculation provided by another embodiment of the present application;

[0032] Figure 6 It is a schematic diagram of the composition of a DMA controller provided by an embodiment of the present application;

[0033] Figure 7 It is a schematic diagram of the structure of a system-on-chip provided by an embodiment of the present application. Specific Embodiments

[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0035] In addition, the descriptions involving "first", "second", etc. in this application are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the embodiments of this application, unless otherwise specified, the meaning of "a plurality" is two or more. In addition, the use of "based on" or "in accordance with" implies openness and inclusiveness because a process, step, calculation, or other action "based on" or "in accordance with" one or more of the stated conditions or values may, in practice, be based on additional conditions or values beyond those stated.

[0036] With the continuous development of the computer industry, the application of the acceleration engine module has gradually covered multiple industries and fields. For example, a cryptographic accelerator is used to quickly perform encryption and decryption operations to protect data transmission and storage security, a digital signal processor is used to accelerate signal processing tasks such as modulation and demodulation, voice and image coding, a wireless communication accelerator is used to improve data transmission efficiency and processing efficiency in 5G and 2G networks, and an edge computing accelerator is used to quickly process data at the place where data is generated in Internet of Things devices to reduce the dependence on cloud resources. In the application of the above-mentioned related acceleration engine modules, processing command stream data is the most critical step.

[0037] In the related art, it is usually the acceleration engine module that parses the instructions and then starts the DMA controller to read the parameters or data in the memory, waits for the data to return and then executes the calculation task, and after completing the calculation task, writes the calculation result back to the specified memory area through the DMA controller and waits for the write response to return. During this process, the calculation task cannot be executed until the read data returns, and the write-back of the calculation result needs to wait for the write response to return, resulting in a relatively long waiting time for the arithmetic unit of the accelerator and low data processing efficiency. In addition, in the related art, both the read and write tasks need to wait for the previous read return or write response before executing the next read or write task, resulting in a large read-write delay and thus low data processing efficiency.

[0038] Specifically, this application provides a scenario example of a cryptographic acceleration engine module. A cryptographic acceleration engine module is a hardware component used to improve the speed of encryption and decryption in the data processing process, and it is widely applied to various scenarios that require privacy protection, such as network communication, data storage, authentication, etc.

[0039] The password acceleration engine module can execute complex encryption algorithms quickly. First, it is necessary to determine the task information to be executed, including the type of cryptographic operation to be executed, algorithm parameters, and task identifier. Specifically, cryptography includes AES, RSA, etc., algorithm parameters include key size, initialization vector, etc., and the type of cryptographic operation includes encryption, decryption, signature generation, signature authentication, etc. Exemplarily, the password acceleration engine module receives the above task information, determines the operation type and algorithm parameters to be executed, and prepares a computing task according to the task information, such as setting algorithm parameters, loading keys, etc., obtains task data through the DMA controller, and the password acceleration engine module uses the task data and the computing task to execute cryptographic operations. After the operation is completed, it outputs computing results such as encrypted ciphertext, decrypted plaintext, hash value, etc., and returns the above computing results to the memory for storage.

[0040] It should be noted that the above-mentioned acceleration engine module is equivalent to the password acceleration engine module.

[0041] The above is only a scenario example provided in the specification and is not intended to limit the present invention. Any modifications, equivalent replacements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0042] In view of this, one or more embodiments of the present application provide a data processing method, a DMA controller, and a system-on-chip, which can improve the data processing efficiency.

[0043] One embodiment of the present application provides a data processing method, which is applied to a DMA controller. The DMA controller is communicatively connected to an acceleration engine module. The acceleration engine module includes a data read / write unit, a command execution unit, and a computing unit connected in sequence. Specifically:

[0044] S1: Receive a data read / write instruction sent by the data read / write unit, and prefetch task data from the memory according to the data read instruction. The data read / write instruction is generated by the data read / write unit based on the address information in the command stream data.

[0045] The above command stream data is a data stream sent by a task management module upstream of the acceleration engine module. It includes multiple task commands with a timing relationship. The above task management module can be understood as a processor core. Each of the above task commands carries its own address information. Through the above address information, the memory address accessed by each task command can be determined.

[0046] The above data reading and writing unit receives the command data stream issued by the task management module, parses the address information carried by each task command in the command data stream, determines the memory address to be accessed according to the address information carried by each task command, and generates corresponding data read / write instructions according to the task commands. The above data read / write instructions act on the DMA controller, and the DMA controller accesses the memory in advance through the bus network to read the required task data, and temporarily stores the pre-read task data in the cache area of the DMA controller.

[0047] Specifically, in response to the data read instruction sent by the data reading and writing unit, the DMA controller sends a read control signal to the memory to pre-read the task data from the memory, and returns the pre-read task data to the cache area of the DMA controller through the bus network. By pre-reading the task data, the task data of each task command can be temporarily stored in the cache area of the DMA controller, and the command execution unit can directly call the corresponding task data from the cache area and send it to the computing unit, thereby reducing the idle time of the computing unit.

[0048] S3: Provide the required task data to the command execution unit, so that the command execution unit sends the required task data and the pre-generated computing task to the computing unit, where the computing task is pre-generated by the command execution unit based on the task information in the command stream data.

[0049] When parsing the address information carried by each task command in the command stream data, the above data reading and writing unit also parses the task information carried by each task command, and sends the task information to the command execution unit. The above task information represents the computing task that the current task command needs to perform. The command execution unit generates a corresponding computing task according to the above task information, obtains the pre-read task data from the cache area of the DMA controller, and sends the above computing task and task data to the computing unit.

[0050] Exemplarily, the above task information may be a set of instructions or parameters for guiding the cryptographic acceleration engine module to perform cryptographic operations. Then, the above computing task may be related operation instructions, such as performing AES encryption and verifying RSA signatures. The corresponding pre-read task data may include input data such as the plaintext to be encrypted, the ciphertext to be decrypted, the message to be hashed, key data, and other auxiliary data such as initialization vectors and salt values. The present application does not impose too many restrictions on this.

[0051] S5: Receive the calculation result for the required task data sent by the command execution unit, and write the calculation result into the memory according to the data write instruction received from the data reading and writing unit; wherein, the calculation result is processed by the calculation unit according to the calculation task and the required task data, and is fed back by the calculation unit to the command execution unit.

[0052] In this embodiment, the calculation unit processes the task data according to the calculation task to obtain a calculation result, and sends the above calculation result to the DMA controller. Exemplarily, the cryptographic acceleration engine module uses the task data and the calculation task to perform cryptographic operations. When the calculation task is to perform AES encryption, the task data includes the plaintext to be encrypted, and the calculation result obtained after encryption processing is the encrypted plaintext. The above calculation result is returned to the DMA controller for temporary storage, and the DMA controller responds to the instruction to write the above calculation result back to the memory.

[0053] In response to the data write instruction sent by the data reading and writing unit, write the calculation result sent by the command execution unit back to the specified address in the memory. Since the command execution unit temporarily stores the calculation result in the DMA controller in advance, the DMA controller can immediately call the calculation result for writing back after receiving the data write instruction, thereby reducing the waiting time for the calculation unit each time.

[0054] In this embodiment, the processing process of the data write instruction and the calculation process of the task data are decoupled. By means of asynchronous operation, the DMA controller can pre-complete the reading of the task data, avoiding the situation that the calculation unit is idle when the data reading delay is large due to reading the data first and then issuing the calculation task, thereby improving the data processing efficiency.

[0055] In one embodiment, after the above data reading and writing unit receives the command stream data, the sending order of each request in the command stream data and the writing order of the calculation result can be constrained by timing. Specifically, the data reading and writing unit analyzes the command stream data to obtain the task flow of the calculation task, and the task flow is used to define the timing of data reading and data writing. Since there is a timing relationship in the above command stream data, a task flow representing the timing relationship can be obtained, and the sending timing of the data read / write instruction is determined according to the above task flow, and a respective timing identifier is added to each data read / write instruction generated by the data reading and writing unit. Among them, the above data reading and writing unit adds a timing identifier to the generated data read / write instruction according to the above task flow. For example, a timestamp or a unique serial number is added to the instruction structure of the data read / write instruction as the timing identifier, and the data read / write instruction carrying the timing identifier is sent to the DMA controller, so as to provide the timing information of the task command in the data processing flow.

[0056] In this embodiment, adding a timing identifier to each data read / write instruction can ensure the normal asynchronous operation of different tasks among multiple task commands. By using the timing identifier to represent the processing order of multiple data read / write instructions, when facing a large number of complex operations, the asynchronous operation method can improve the operation efficiency while still ensuring the correctness of the operation timing.

[0057] In one embodiment, while the command execution unit performs task parsing, the DMA controller pre-reads task data from the memory according to the received data read instruction. Specifically, the DMA controller converts the received data read instruction into a read control signal for the bus, and based on the above read control signal, pre-reads task data from the memory through the bus, adds the timing identifier carried in the data read instruction to the above task data, and caches the task data with the added timing identifier.

[0058] In this embodiment, adding the timing identifier to the task data can facilitate the command execution unit to retrieve the corresponding task data. When the command execution unit needs to continuously perform multiple calculation tasks, in order to obtain the task data corresponding to the current calculation task, the task data corresponding to the current calculation task can be retrieved by identifying the task data with the same timing identifier as the current task data. In addition, temporarily storing the pre-read task data with the added timing identifier in the cache area of the DMA controller can facilitate the command execution unit to call the task data required for each calculation task at any time.

[0059] In one embodiment, after obtaining the task information of the command stream data, the command execution unit controls the task flow according to the above task information. Specifically, the command execution unit parses the above task information to obtain the task flow of the current calculation task, and the above task flow is used to define the timing of data reading and data writing. Correspondingly, the above command execution unit identifies one or more timing identifiers represented by the above task flow, and obtains one or more task data with the same timing identifier from the DMA controller, and sends the current calculation task and its corresponding task data to the calculation unit.

[0060] In this embodiment, the command execution unit reads the correct data according to the timing identifier, ensuring that the calculation unit can process according to the task flow of the calculation task, and ensuring the correct rate of the calculation task execution while improving the calculation efficiency of the engine.

[0061] In one embodiment, after the computing unit completes the calculation based on the computing task and task data, it feeds back the obtained calculation result to the command execution unit. Specifically, the computing unit adds the timing identifier carried in the processed task data to the obtained calculation result, and feeds back the calculation result with the added timing identifier to the command execution unit. Subsequently, the command execution unit determines the timing identifier corresponding to the calculation result among one or more timing identifiers represented by the task flow, so as to determine the order in which each calculation result is written into the data, and sends each calculation result to the DMA controller.

[0062] In one embodiment, the DMA controller writes the calculation result sent by the command execution unit into the memory according to the data write instruction sent by the data read / write unit. Among them, the sending timing of the data write instruction is determined by the timing of data writing represented by the task flow of the data read / write unit. Specifically, the DMA controller caches each calculation result sent by the command execution unit, identifies the timing identifier carried in the currently to-be-processed data write instruction, and queries the calculation result with the same timing identifier from the cached calculation results. After the corresponding calculation result is queried, the DMA controller writes the queried calculation result into the memory according to the data output address specified by the current data write instruction.

[0063] In this embodiment, the calculation result to be written is associated with the data write instruction through the timing identifier, so as to ensure that the data write instruction can write the correct calculation result and ensure the correct timing of data writing. In addition, if an abnormal situation occurs during data reading and writing, such as the data read / write instruction fails to be executed successfully, or the corresponding task data or calculation result cannot be queried, the command execution unit reports the read / write error status.

[0064] In one embodiment, after each task flow of the computing task is completed, the data read / write unit receives the task execution status sent by the command execution unit. The above task execution status may include the write completion status and the read / write error status. The above task execution status is fed back to the sender of the command stream data, that is, the task management module. By reporting step by step from the command execution unit, the task management module can clarify the execution situation of the task.

[0065] Please refer to FIG. 1(a) and FIG. 1(b). In one embodiment, the DMA controller adopts parallel processing when performing data prefetching. The DMA controller includes parallel read request channels and read response channels. The above read request channels are used to send read control signals, and the above read response channels are used to transmit task data back to the DMA controller. Specifically, after the DMA controller receives the data read instruction sent by the data read / write unit, it converts the above data read instruction into a read control signal of the bus, sends the above read control signal through the read request channel, and receives the task data returned by the bus through the read response channel.

[0066] As shown in FIG. 1(a), for a normal read request channel, the next read request can only be sent after the response to the first read request is returned, resulting in a relatively long data reading time and delaying the processing progress of the acceleration engine module. In this embodiment, as shown in FIG. 1(b), by processing the read request channel and the read response channel in parallel, the sending of a read request no longer needs to wait for the response to the previous read request to be returned, improving the processing bandwidth and speed of the read channel, and thus enhancing the processing performance of the acceleration engine module.

[0067] Please refer to FIGS. 2(a) and 2(b). In one embodiment, the DMA controller also uses parallel processing when writing data. The DMA controller includes parallel write request channels and write response channels. The above-mentioned write request channels are used to send write control signals, and the above-mentioned write response channels are used to send write response signals to the DMA controller. Specifically, the DMA controller generates a write control signal for the bus based on the data write instruction and the calculation result, sends the write control signal through the write request channel, and receives the write response signal returned by the bus through the write response channel.

[0068] In addition, in this embodiment, the DMA controller further includes a write data channel for transmitting the calculation result to the memory. Specifically, when generating the write control signal for the bus, the DMA controller also generates a write data signal for the bus. After sending the write control signal through the write request channel, the DMA controller sends the write data signal associated with the write control signal through the write data channel.

[0069] As shown in FIG. 2(a), for a normal write request channel, the next write request can only be sent after the response to the first write request is returned, resulting in a relatively long data reading time and delaying the processing progress of the acceleration engine module. In this embodiment, as shown in FIG. 2(b), by processing the write request channel and the write response channel in parallel, the sending of a write request no longer needs to wait for the response to the previous write request to be returned, improving the processing bandwidth and speed of the write channel, and thus enhancing the processing performance of the acceleration engine module.

[0070] Exemplarily, please refer to Figure 3 , this application provides an embodiment of a command processing method for an acceleration engine module. The read request channel and the read response channel are processed in parallel. The sending of read command a2 no longer needs to wait for the read response signal a of the previous read command a1 11 to return, and the sending of write command b2 no longer needs to wait for the write response signal b of the previous write command b1 11 to return. Specifically, the DMA controller converts read command a1 / read command a2 into read control signal a 10 / read control signal a 20And send it to the memory to read task data. After the reading is completed, a read response signal a is returned 11 / Read response signal a 21 Similarly, the DMA controller converts the write command b1 / write command b2 into a write control signal b 10 / Write control signal b 20 And send it to the memory to write the calculation result, and then return a write response signal b 11 / Write response signal b 21 In addition, while the data is being read, the data processing and calculation are carried out synchronously, thereby reducing the waiting time of the calculation unit and further improving the execution efficiency of the accelerator.

[0071] In the above embodiment, by processing the read request channel and the read response channel in parallel, and processing the write request channel and the write response channel in parallel, changing the serial process in the traditional technical solution, the sending of read / write requests no longer requires the response of the previous request. The parallel pipeline mechanism can avoid the problem of the calculation unit being idle or blocked due to the large delay of read / write responses, thereby improving the execution efficiency of the accelerator.

[0072] In addition, the present application also provides an embodiment of another data processing method. In this embodiment, the data read / write instruction is decoupled from the data processing process to perform pre-reading of data, and its data reading process and calculation process adopt an asynchronous working mode.

[0073] Please refer to Figure 4 , the data read / write unit performs operations of data pre-reading and data writing back. Specifically, the upstream task management module issues command stream data, and the data read / write unit parses the above command stream data to obtain the data input address and data output address to be accessed. According to the above data input address, data output address and the timing information of the task, a data read / write instruction is generated and issued to the DMA controller. The DMA controller converts the above data read / write instruction into a read control signal or a write control signal that can be transmitted to the memory module through the bus to read task data or write the cached calculation result into the memory.

[0074] At the same time, please refer to Figure 5, the command execution unit obtains task information other than the data address from the data reading and writing unit, parses the above task information, and generates a control flow and calculation tasks. The control flow can be understood as the work flow for accelerating tasks, and the calculation tasks can be understood as the data calculation processes at each node in the above work flow. It can be understood that the command execution unit is the task control unit of the acceleration engine module, responsible for parsing the content to generate calculation tasks and controlling the process nodes of each unit of the acceleration engine module. The command execution unit sends the task data and calculation tasks obtained by the DMA controller from the memory to the calculation unit. The calculation unit processes the task data according to the calculation tasks and feeds back the obtained calculation results to the command execution unit. The command execution unit temporarily stores the obtained calculation results in the DMA controller to wait for the data to be written back.

[0075] In view of this, decouple the data reading and writing instructions from the data processing process. The data reading and writing unit extracts the data reading and writing instructions from the command stream data and directly sends the data reading and writing instructions to the DMA controller. The DMA controller can then pre-read the task data required for subsequent calculation tasks. In an asynchronous working mode, the command execution unit generates calculation tasks and obtains the pre-read task data from the DMA, and completes the calculation process results with the help of the calculation unit. It is passed from the command execution unit to the DMA controller, and the DMA controller combines the data write instructions received from the data reading and writing unit to write the calculation results to the memory. In this way, the situation where the calculation unit is idle due to a large memory read latency can be avoided, thereby improving the data processing efficiency.

[0076] Please refer to Figure 6 , this application also provides a DMA controller. The DMA controller is communicatively connected to the acceleration engine module. The acceleration engine module includes a data reading and writing unit, a command execution unit, and a calculation unit connected in sequence. The DMA controller includes:

[0077] A data receiving unit 100, configured to receive the data read / write instructions sent by the data reading and writing unit. The data read / write instructions are generated by the data reading and writing unit based on the address information in the command stream data, and to receive the calculation results sent by the command execution unit. The calculation results are obtained by the calculation unit according to the calculation tasks and the required task data and are fed back to the command execution unit by the calculation unit;

[0078] A data processing unit 200, configured to pre-read task data from the memory according to the data read instruction, and write the calculation results into the memory according to the data write instructions received from the data reading and writing unit;

[0079] A data sending unit 300 is configured to provide required task data to the command execution unit, so that the command execution unit sends the required task data and a pre-generated computing task to the computing unit, where the computing task is pre-generated by the command execution unit based on task information in the command stream data.

[0080] In one embodiment, the data read / write instruction has a timing identifier, which is used to represent the timing of data reading and data writing in the task flow of the current computing task, and the task flow of the current computing task is parsed by the data read / write unit based on the received command stream data.

[0081] In one embodiment, the data processing unit is specifically configured to, after receiving a data read instruction sent by the data read / write unit, convert the data read instruction into a read control signal for the bus, send the read control signal through the read request channel, and receive the task data returned by the bus through the read response channel.

[0082] In one embodiment, the data processing unit is further specifically configured to cache each computing result sent by the command execution unit; identify the timing identifier carried in the current data write instruction to be processed, and query the computing result with the timing identifier from the cached computing results; and write the queried computing result into the memory according to the data output address specified in the data write instruction.

[0083] In one embodiment, the data receiving unit is specifically configured to receive a data acquisition request sent by the command execution unit, where the data acquisition request includes one or more timing identifiers, and the timing identifiers are used to represent the timing of data reading and data writing in the task flow of the current computing task, and the task flow of the current computing task is parsed by the command execution unit based on the task information.

[0084] Correspondingly, the data sending unit is specifically configured to feedback the task data with the timing identifier to the command execution unit.

[0085] In one embodiment, the computing result carries a timing identifier added by the computing unit, and the timing identifier carried in the computing result is consistent with the timing identifier carried in the task data after processing is completed.

[0086] In one embodiment, the DMA controller includes parallel read request channels and read response channels. Among them, after receiving a data read instruction sent by the data read / write unit, the DMA controller converts the data read instruction into a read control signal for the bus, sends the read control signal through the read request channel, and receives the task data returned by the bus through the read response channel.

[0087] In one embodiment, the DMA controller further includes a parallel write request channel and a write response channel. Among them, the DMA controller generates a write control signal for the bus according to the data write instruction and the calculation result, sends the write control signal through the write request channel, and receives the write response signal returned by the bus through the write response channel.

[0088] In one embodiment, the DMA controller further includes a write data channel. Among them, when the DMA controller generates a write control signal for the bus, it also generates a write data signal for the bus. After the DMA controller sends the write control signal through the write request channel, it sends the write data signal associated with the write control signal through the write data channel.

[0089] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding foregoing embodiments, and will not be elaborated herein.

[0090] Please refer to Figure 7 , this application also provides a system-on-chip, which includes a DMA controller 20 and an acceleration engine module 10, where:

[0091] The acceleration engine module 10 is used to process command stream data and interact with the DMA controller 20 based on the processing result;

[0092] The DMA controller 20 is used to execute the data processing method in the foregoing embodiments during the interaction with the acceleration engine module.

[0093] In one embodiment, the system-on-chip further includes interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Each component communicates and connects with each other using different buses or other means, and can be installed on a common main board or installed in other ways as needed.

[0094] In one embodiment, the above-mentioned memory can be a random access memory, which is used to temporarily store and quickly access data. In one implementation, the memory can be the hardware for temporarily storing data and instructions when the CPU executes tasks, such as DRAM (Dynamic Random Access Memory) and SRAM (Static Random Access Memory).

[0095] Among them, the memory may include a program storage area and a data storage area. The program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory may optionally include a memory remotely provided with respect to the processor, and these remote memories can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0096] In one embodiment, the above-mentioned memory can be a storage medium for storing and holding data and instructions. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories.

[0097] The system or unit illustrated in the above embodiments can be specifically implemented by a computer chip or an entity, or by a product with a certain function. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0098] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0099] Those skilled in the art should understand that the embodiments of the present application can be provided as a method and a system on a chip. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.

[0100] This application is described with reference to the flowcharts and / or block diagrams of methods, DMA controllers, and systems-on-chip according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more processes of the flowchart and / or one or more blocks of the block diagram.

[0101] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions specified in one or more processes of the flowchart and / or one or more blocks of the block diagram.

[0102] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes of the flowchart and / or one or more blocks of the block diagram.

[0103] It should also be noted that the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.

[0104] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, they are described relatively simply, and the relevant parts can be referred to the description of the method embodiments.

[0105] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

[0106] Although the embodiments of the present application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations fall within the scope defined by the appended claims.

Claims

1. A data processing method, characterized in that, Applied to a DMA controller, the DMA controller is communicatively connected to an acceleration engine module, and the acceleration engine module includes a data read / write unit, a command execution unit, and a calculation unit connected in sequence. The method includes: Receiving a data read / write instruction sent by the data read / write unit, and prefetching task data from the memory according to the data read instruction. The data read / write instruction is generated by the data read / write unit based on the address information in the command stream data; Providing the required task data to the command execution unit so that the command execution unit sends the required task data and a pre-generated calculation task to the calculation unit, where the calculation task is pre-generated by the command execution unit based on the task information in the command stream data; Receiving the calculation result for the required task data sent by the command execution unit, and writing the calculation result into the memory according to the data write instruction received from the data read / write unit; where the calculation result is obtained by the calculation unit based on the calculation task and the required task data and is fed back by the calculation unit to the command execution unit.

2. The method according to claim 1, wherein The data read / write instruction has a timing identifier, which is used to represent the timing of data reading and data writing in the task flow of the current calculation task. The task flow of the current calculation task is parsed by the data read / write unit based on the received command stream data.

3. The method according to claim 1 or 2, characterized in that, Prefetching task data from the memory according to the data read instruction includes: Converting the data read instruction into a read control signal for the bus, and prefetching task data from the memory through the bus based on the read control signal; Adding the timing identifier carried in the data read instruction to the task data, and caching the task data with the added timing identifier.

4. The method according to claim 1 or 2, characterized in that, Writing the calculation result into the memory according to the data write instruction received from the data read / write unit includes: Caching each calculation result sent by the command execution unit; Identifying the timing identifier carried in the current data write instruction to be processed, and querying the calculation result with the timing identifier from the cached calculation results; Writing the queried calculation result into the memory according to the data output address specified by the data write instruction.

5. The method according to claim 1, wherein Providing the required task data to the command execution unit includes: Receiving a data acquisition request sent by the command execution unit, where the data acquisition request includes one or more timing identifiers, which are used to represent the timing of data reading and data writing in the task flow of the current calculation task. The task flow of the current calculation task is parsed by the command execution unit based on the task information; Feeding back the task data with the timing identifier to the command execution unit.

6. The method according to claim 1, wherein The calculation result carries a timing identifier added by the calculation unit, and the timing identifier carried in the calculation result is consistent with the timing identifier carried in the task data for which the processing is completed.

7. The method according to claim 1, characterized in that, The DMA controller includes parallel read request channels and read response channels. Among them, after receiving a data read instruction sent by the data reading / writing unit, the DMA controller converts the data read instruction into a read control signal for the bus, and sends the read control signal through the read request channels, and receives task data returned by the bus through the read response channels.

8. The method according to claim 1 or 7, characterized in that, The DMA controller further includes parallel write request channels and write response channels. Among them, the DMA controller generates a write control signal for the bus according to the data write instruction and the calculation result, and sends the write control signal through the write request channels, and receives a write response signal returned by the bus through the write response channels.

9. The method according to claim 8, characterized in that, The DMA controller further includes a write data channel. Among them, when generating a write control signal for the bus, the DMA controller also generates a write data signal for the bus. After sending the write control signal through the write request channels, the DMA controller sends a write data signal associated with the write control signal through the write data channel.

10. A DMA controller, characterized in that, The DMA controller is communicatively connected to an acceleration engine module. The acceleration engine module includes a data reading / writing unit, a command execution unit, and a calculation unit connected in sequence. The DMA controller includes: A data receiving unit, configured to receive data read / write instructions sent by the data reading / writing unit, where the data read / write instructions are generated by the data reading / writing unit based on address information in command stream data, and receive a calculation result sent by the command execution unit, where the calculation result is obtained by the calculation unit according to a calculation task and required task data, and is fed back by the calculation unit to the command execution unit; A data processing unit, configured to prefetch task data from a memory according to a data read instruction, and write the calculation result into the memory according to a data write instruction received from the data reading / writing unit; A data sending unit, configured to provide required task data to the command execution unit, so that the command execution unit sends the required task data and a pre-generated calculation task to the calculation unit, where the calculation task is pre-generated by the command execution unit based on task information in the command stream data.

11. An on-chip system, characterized in that, The system-on-chip includes a DMA controller and an acceleration engine module, where: The acceleration engine module is configured to process command stream data and interact with the DMA controller based on a processing result; The DMA controller is configured to execute the data processing method according to any one of claims 1 to 9 during the interaction with the acceleration engine module.