Data transmission method, apparatus and device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MOORE THREADS TECH CO LTD
- Filing Date
- 2026-04-03
- Publication Date
- 2026-06-30
Smart Images

Figure CN122309414A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a data transmission method, apparatus, and device. Background Technology
[0002] In related technologies, the Direct Memory Access (DMA) controller is a key component in computing systems, used to perform efficient data transfer between peripherals and memory. Traditional DMA controllers typically manage data transfer tasks using descriptors (DESCs), where a descriptor contains information such as the source address, destination address, and data length required for a single data transfer.
[0003] Each data transfer task requires configuration by the CPU via software. Specifically, when a data transfer is needed, the Central Processing Unit (CPU) performs the following operations: Configures a corresponding descriptor for the transfer operation, containing key parameters necessary for the operation, such as the source memory address, destination memory address, and the length of the data to be transferred. Configures the storage address of this descriptor in a specific channel register of the DMA controller. Starts the DMA channel by writing to the control register, thus initiating the data transfer task. After completing this configuration, the DMA controller generates an interrupt signal to notify the CPU. Upon responding to this interrupt, if a new data transfer requirement arises, the CPU must repeat the entire configuration process to initiate the next DMA operation.
[0004] Therefore, the above method has at least the following drawbacks: In scenarios that require continuous or batch data transfer, the CPU needs to frequently intervene in the configuration process of each independent transfer task, resulting in a large amount of interrupt handling overhead and low system efficiency. This performance bottleneck caused by the CPU's frequent configuration and DMA initiation is particularly prominent in scenarios that require frequent processing of a large number of small data packets (such as network communication and storage I / O), which seriously restricts the system's data throughput and real-time response capability. Summary of the Invention
[0005] This disclosure provides a data transmission method, apparatus, and device.
[0006] In a first aspect, this disclosure provides a data transfer method applied to a direct memory access (DMA) controller, the method comprising:
[0007] Get the channel descriptor corresponding to any DMA channel in the DMA controller;
[0008] When the channel descriptor is a preset descriptor in the descriptor chain, multiple descriptors in the descriptor chain are prefetched according to the channel prefetching strategy corresponding to any DMA channel.
[0009] The prefetched descriptors are cached in a shared buffer for the channel region corresponding to any DMA channel; wherein the shared buffer contains multiple logically isolated channel regions, each corresponding to a multiple DMA channel.
[0010] Control any DMA channel to perform data transfer operations based on multiple descriptors cached in the channel area.
[0011] Secondly, this disclosure provides a data transfer apparatus applied to a direct memory access (DMA) controller, the apparatus comprising:
[0012] The interface conversion module is used to receive access requests;
[0013] Multiple DMA channels are connected to the interface conversion module; each DMA channel is used to obtain the corresponding channel descriptor through the interface conversion module, and when the channel descriptor is a preset descriptor in the descriptor chain, it prefetches multiple descriptors in the descriptor chain according to the channel prefetching strategy configured for the DMA channel.
[0014] A shared buffer, connected to multiple DMA channels, is used to cache the descriptor corresponding to each DMA channel; the shared buffer contains multiple channel regions that are logically isolated from each other and each corresponds to multiple DMA channels.
[0015] The data transmission module, connected to multiple DMA channels and a shared buffer, is used to perform data transmission operations based on multiple descriptors cached in the channel area of any DMA channel when a data transmission request is received from any DMA channel.
[0016] Thirdly, this disclosure provides an electronic device, including:
[0017] At least one processor; and
[0018] Memory that is communicatively connected to at least one processor;
[0019] As described above, the Direct Memory Access (DMA) controller communicates with at least one processor and memory.
[0020] The processor is configured to initiate a data transfer request or configuration operation to the DMA controller so that the DMA controller executes the above-described method.
[0021] In the embodiments provided in this disclosure, when the channel descriptor is determined to be a preset descriptor in the descriptor chain, multiple descriptors in the descriptor chain are prefetched according to the channel prefetching strategy corresponding to any DMA channel, and cached in the channel area of the shared buffer corresponding to any DMA channel. This prefetching method allows for the continuous acquisition of multiple descriptors, avoiding frequent interrupt overhead. Furthermore, in this disclosure, multiple DMA channels share the same shared buffer, facilitating flexible configuration of the corresponding channel prefetching strategy and channel area for each DMA channel. This avoids the resource waste caused by fixed allocation of independent caches for each channel, enabling each DMA channel in the system to more fully utilize the total capacity of the shared buffer. This significantly reduces the problem of idle cache resources caused by a certain channel being idle or under low load, thereby achieving higher resource utilization at the hardware level and reducing chip area and cost.
[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0023] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:
[0024] Figure 1 A flowchart illustrating a data transmission method provided in this embodiment of the disclosure;
[0025] Figure 2 An overall block diagram of a DMA controller in an example is shown;
[0026] Figure 3 A schematic diagram of the descriptor structure is shown;
[0027] Figure 4 A flowchart illustrating the chain pattern is shown;
[0028] Figure 5 A schematic diagram of the prefetching process in blockchain mode is shown;
[0029] Figure 6 A schematic diagram of the descriptor prefetching process based on the water level control strategy in chained mode is shown.
[0030] Figure 7 This diagram illustrates the case where shared_buffer_pool = 0.
[0031] Figure 8 This diagram illustrates the case where shared_buffer_pool = 1.
[0032] Figure 9 The process diagram of each prefetch operation is shown;
[0033] Figure 10 A structural diagram of a data transmission device provided in an embodiment of this disclosure;
[0034] Figure 11 A block diagram of an electronic device is provided in this disclosure. Detailed Implementation
[0035] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0036] Unless otherwise specified, the various embodiments and features of this disclosure may be combined with each other. As used herein, the term "and / or" includes any and all combinations of one or more of the associated enumerated entries.
[0037] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, they specify the presence of features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0038] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0039] Figure 1This is a flowchart illustrating a data transmission method provided in an embodiment of this disclosure. The method is applied to a Direct Memory Access (DMA) controller, which can be integrated into a System-on-Chip (SoC) as a dedicated hardware engine responsible for efficient data transfer between peripherals and memory. The DMA controller is a key component for improving data throughput efficiency in a computing system, typically connected to the central processing unit (CPU), memory controller, and various peripheral interfaces via a system bus. It usually includes a set of slave interfaces for receiving configuration commands from the processor, and one or more master interfaces for actively initiating read and write accesses to system memory. The DMA controller in this solution supports multi-channel concurrent transmission and can be applied to various scenarios, such as network packet processing, storage controller data exchange, image processing unit data loading, and data scheduling in artificial intelligence accelerators. This application does not limit the specific peripheral types served by the DMA controller or the integrated SoC platform architecture.
[0040] Reference Figure 1 The method includes:
[0041] Step S110: Obtain the channel descriptor corresponding to any DMA channel in the DMA controller.
[0042] In this context, a DMA channel refers to a logical unit within the DMA controller that can independently perform data transfer tasks. The channel descriptor indicates the parameters required for the DMA data transfer task, typically including but not limited to source address, destination address, and data length.
[0043] Step S120: When the channel descriptor is a preset descriptor in the descriptor chain, prefetch multiple descriptors in the descriptor chain according to the channel prefetching strategy corresponding to any DMA channel.
[0044] In this context, a descriptor chain is a sequence of multiple descriptors linked by pointers. Correspondingly, a preset descriptor is typically a key node in the descriptor chain, meaning a specific descriptor located at a predetermined position within the chain. For example, a preset descriptor could be the first descriptor in the chain, from which the DMA controller obtains the start information for the entire chain. Alternatively, a preset descriptor could be the last descriptor in a descriptor block, containing navigation information pointing to the next descriptor block, thus supporting the processing of large-scale or circular data streams.
[0045] Different channel prefetching strategies can be configured for different DMA channels. Each DMA channel's prefetching strategy controls the prefetching rate, prefetching timing, and / or prefetching method. Specifically, the channel prefetching strategy can be implemented based on a water level triggering method, for example, by presetting at least one water level threshold to control the prefetching behavior of the corresponding channel. Additionally, the channel prefetching strategy can also be a priority-weighted strategy: the prefetching strategy can be combined with channel priorities, allowing higher-priority channels to obtain larger bandwidth quotas or shorter arbitration wait times when prefetching is triggered.
[0046] Those skilled in the art will understand that the channel prefetching strategy is not limited to the water level control or priority strategy listed above. Various triggering mechanisms based on time, event, or load prediction can also be flexibly adopted, as long as they can achieve the effect of controlling the prefetching behavior according to the channel status.
[0047] Step S130: Cache the prefetched descriptors into the channel region corresponding to any DMA channel in the shared buffer; wherein the shared buffer contains multiple channel regions that are logically isolated from each other and correspond to multiple DMA channels respectively.
[0048] As the name suggests, a shared buffer is shared by multiple DMA channels within the DMA controller. This shared buffer is used to cache descriptors and is typically located within the DMA controller, accessible by all DMA channels. Compared to equipping each channel with its own independent cache space, the multi-channel sharing approach in this disclosure significantly reduces chip area and hardware cost.
[0049] Therefore, the shared buffer in this application refers to a unified, physically contiguous storage area within the DMA controller, which is accessed and used by all DMA channels within the controller. A channel region can be a segment of address space logically allocated for each DMA channel within the shared buffer, using configurable address pointers and region size registers. Logically, each channel region can be isolated from others by pointer boundaries, ensuring that a channel can only access its corresponding region. Thus, the channel regions in this application have a dynamically configurable characteristic, meaning that the size of the channel region can be dynamically adjusted according to the channel load. Furthermore, the channel regions have a logical isolation characteristic, meaning that this isolation can be achieved through hardware-managed pointers, rather than physical isolation.
[0050] A DMA channel's channel region refers to a logically contiguous segment of address space allocated within a shared buffer for each DMA channel. The size of each channel's channel region is typically configurable. Multiple DMA channels can be mapped to multiple channel regions within the shared buffer using pointers or other management mechanisms, ensuring that each channel can only access its own channel region and cannot access other channels' regions, achieving hardware-level isolation. The allocation of channel regions across multiple DMA channels can be dynamically configured based on factors such as the historical load and task type of each channel.
[0051] Step S140: Control any DMA channel to perform a data transfer operation based on multiple descriptors cached in the channel area.
[0052] The DMA controller controls each DMA channel to sequentially obtain descriptors from the channel area and performs data transfer operations based on the obtained descriptors.
[0053] Therefore, in this application, when the channel descriptor is determined to be a preset descriptor in the descriptor chain, multiple descriptors in the descriptor chain are prefetched according to the channel prefetching strategy corresponding to any DMA channel, and cached in the channel area of the shared buffer corresponding to any DMA channel. Through prefetching, multiple descriptors can be continuously acquired, avoiding frequent interrupt overhead. Furthermore, in this disclosure, multiple DMA channels share the same shared buffer, facilitating flexible configuration of the corresponding channel prefetching strategy and channel area for each DMA channel. This avoids the resource waste caused by fixed allocation of independent caches for each channel, enabling each DMA channel in the system to more fully utilize the total capacity of the shared buffer, significantly reducing the problem of idle cache resources caused by a certain channel being idle or under low load, thereby achieving higher resource utilization at the hardware level and reducing chip area and cost.
[0054] In addition, those skilled in the art can make various modifications and variations to the above embodiments:
[0055] In one optional implementation, to improve descriptor acquisition efficiency, an address navigation field can be added to the descriptor. Accordingly, multiple descriptors in the descriptor chain are prefetched as follows: the address navigation field in the channel descriptor is obtained; based on the address navigation field in the channel descriptor and the channel prefetching strategy corresponding to any DMA channel, multiple descriptors in the descriptor chain are prefetched; wherein, the first field value in the address navigation field of the channel descriptor is used to characterize the starting address information of multiple descriptors in the descriptor chain. The address navigation field refers to a specific data area reserved in the descriptor data structure to guide the DMA controller on how to acquire subsequent descriptors. This field can be a collection containing multiple pieces of information. For example, this field may occupy specific words in the descriptor (such as WORD6 and WORD7) to store the complete system memory address. The first field value refers to the portion of the address navigation field used to store address information, pointing to the physical starting address of the next one or more descriptors in system memory. The DMA controller's prefetching module can directly use this address as the target address for the read operation. By using the address navigation field embedded in the descriptor, the DMA controller can automatically load subsequent tasks, significantly reducing the CPU intervention frequency, lowering system overhead, and ensuring the continuity of transmission and high throughput.
[0056] To facilitate control over the timing and rate of prefetching operations, the channel prefetching strategy for any DMA channel can include a water level control strategy. This strategy controls the number of available descriptors cached in the channel region of any DMA channel. The core of this strategy is to dynamically manage the descriptor inventory (i.e., the number of available descriptors) in the buffer by setting a water level threshold. Specifically, the buffer inventory status is quantified into different water levels, and prefetching operations are triggered or stopped based on the water level. Multiple descriptors in the descriptor chain can be prefetched in the following ways: the total number of descriptors in the descriptor chain is determined based on the address navigation field in the channel descriptor; and multiple descriptors in the descriptor chain are prefetched based on the total number of descriptors in the descriptor chain and the channel water level threshold in the water level control strategy. For example, the total number of descriptors in the descriptor chain can be determined based on the value of the second field in the address navigation field of the channel descriptor. In practice, if the total number of descriptors is small, they can be prefetched all at once; if the total number is large, exceeding the upper limit of the channel water level threshold, a batch prefetching method can be adopted, with the maximum number of descriptors prefetched each time limited by the water level threshold. Batch prefetching can combine multiple small memory access latencies into a single large burst transfer, significantly improving bus utilization efficiency. Furthermore, the water level threshold ensures the smoothness of prefetching behavior, avoiding sudden impacts on the memory bus, which helps guarantee fairness among multiple channels and balances high performance with system stability.
[0057] Optionally, the channel water level thresholds in the water level control strategy include: a first channel water level threshold (also called a low water level threshold) and a second channel water level threshold (also called a high water level threshold), where the first channel water level threshold is less than the second channel water level threshold. Correspondingly, multiple descriptors in the descriptor chain can be prefetched in the following way: if the first comparison result between the number of available descriptors cached in the channel region of any DMA channel and the first channel water level threshold is a first preset result, then multiple descriptors in the descriptor chain are prefetched based on the second comparison result between the total number of descriptors and the second channel water level threshold. The first preset result can be: the number of available descriptors is less than or equal to the first channel water level threshold. In this case, it indicates that the number of available, pending descriptors in the channel region has been consumed to the low water level or lower, thus satisfying the triggering condition for the prefetch operation. The second comparison result is used to determine the specific execution strategy for this prefetch operation. For example, it can be: prefetching all descriptors in the entire descriptor chain at once; or it can be: dividing the entire descriptor chain into multiple batches for batch prefetching. By using a comparison mechanism based on the total number of descriptors and a high-water mark threshold, the DMA controller can dynamically determine the prefetching method according to the task size and available resources. Small tasks can be completed quickly in one go, while large tasks can be processed in batches.
[0058] For example, based on a second comparison between the total number of descriptors and the second channel watermark threshold, the prefetch batches of multiple descriptors in the descriptor chain and the batch prefetch quantity within each prefetch batch can be determined. Based on the prefetch batches and the batch prefetch quantity within each prefetch batch, multiple descriptors in the descriptor chain are prefetched in batches. Here, a prefetch batch can refer to multiple consecutive execution units formed by decomposing the prefetch task of the entire descriptor chain. The batch prefetch quantity refers to the number of descriptors that need to be acquired for a single prefetch batch; this quantity is used to determine the amount of data accessed in a single memory access. The second channel watermark threshold plays two roles in the batch prefetching process: firstly, it determines the target upper limit of the batch prefetch quantity; secondly, it ensures that the buffer does not overflow.
[0059] Preferably, the sum of the number of descriptors prefetched in each batch and the number of available descriptors cached in the channel region of any DMA channel does not exceed the second channel watermark threshold. For example, before the prefetch operation of any batch is triggered, it needs to be ensured that the prefetch operation meets the following condition: current free buffer capacity + number of prefetch requests in this batch ≤ high watermark threshold.
[0060] In practice, the prefetch controller compares the total number of descriptors (N) with the high-water mark threshold (N²). If N <= (N² - current available descriptors), the descriptors are divided into one batch, and the batch prefetch quantity is N. If N > (N² - current available descriptors), the descriptors are divided into multiple batches. Except for the last batch, the prefetch quantity for each batch can be set to (N² - current available descriptors), thus filling the buffer to the high-water mark threshold as much as possible. This method achieves efficient utilization of memory bandwidth and improves bus transmission efficiency through the largest possible burst transfers. Optionally, the batch prefetch quantity can also be dynamically adjusted based on the real-time load of the system bus, the priority of the descriptor chain, or historical access patterns.
[0061] In the above approach, by breaking down large tasks into multiple small batches, long-term bus monopolies can be divided into multiple short-term bus occupancy periods, thereby improving transmission efficiency and pipeline performance. This allows each prefetch to fully utilize the available buffer space and initiate efficient burst transmissions, thereby maximizing the amortization of memory access latency and ensuring the efficiency and continuity of the DMA execution pipeline.
[0062] In one optional implementation, to facilitate the management of descriptor prefetching progress, multiple prefetch address registers can be pre-configured. Correspondingly, multiple DMA channels correspond to multiple prefetch address registers, each used to record the address information of the descriptor to be prefetched. Multiple descriptors in the descriptor chain are prefetched in batches as follows: after the current batch of descriptors is prefetched, the value of the prefetch address register corresponding to the DMA channel is updated; based on the updated value of the prefetch address register, the next batch of descriptors is prefetched. The prefetch address register can be a hardware register set up internally by the DMA controller for each channel, used to store the starting physical address in system memory of the next descriptor to be prefetched, thus facilitating the recording and management of the current progress of the descriptor chain prefetching task. The address information of the descriptor to be prefetched refers to the system memory physical address stored in the prefetch address register, which points to the starting position of the next batch of descriptors to be fetched in the descriptor chain. By introducing prefetch address registers and the corresponding automatic update mechanism, fully hardware-automated address management can be achieved, enabling address progression operations to be completed automatically within the hardware, thereby significantly reducing system overhead.
[0063] Optionally, the difference between the second channel water level threshold and the first channel water level threshold is greater than the first channel water level threshold. Furthermore, the first and second channel water level thresholds for each DMA channel can be dynamically adjusted based on the channel load, data transfer frequency, and transfer task type of the DMA channel. If the difference is too small, for example, N2-N1<=N1, it may cause the descriptor to be quickly consumed back to the low water level after the prefetch operation is completed, leading to frequent triggering of the prefetch operation, wasting bus bandwidth and power consumption, and failing to effectively hide memory latency. The above constraints can avoid inefficient jitter. Specifically, the DMA controller driver (such as driver software) can continuously monitor the performance counters of each channel (such as descriptor consumption rate and buffer idle level). When a change in load characteristics is detected (such as changing from low-frequency large packets to high-frequency small packets), the driver can dynamically calculate and configure updated high and low water level thresholds.
[0064] In this system, multiple DMA channels in the DMA controller can be virtual DMA channels. The area size of the channel region corresponding to each virtual DMA channel in the shared buffer can be dynamically adjusted according to the channel load, data transfer frequency, and transfer task type of the virtual DMA channel. A virtual DMA channel refers to multiple independent DMA channel instances simulated on a single physical DMA transfer engine or shared hardware resources using techniques such as time-division multiplexing and logical partitioning. Each virtual channel has complete channel functions (such as independent configuration registers, status, and interrupts), but multiple virtual channels share the underlying physical transfer path and buffer resources. Dynamic adjustment of the area size means that during system operation, the size of the channel region allocated to a certain virtual channel is dynamically changed based on real-time monitored performance indicators or preset strategies. This can be achieved by modifying parameters such as the buffer start address and depth defined in the channel's configuration register.
[0065] To facilitate the management of read and write progress for each channel, the channel region of any DMA channel can be identified by a corresponding channel pointer. Each virtual DMA channel corresponds to a set of channel pointers, including a read pointer pointing to the next descriptor to be read and a write pointer pointing to the next descriptor to be written. The movement range of the read and write pointers is determined based on the address range of the channel region of the virtual DMA channel. The channel pointers are configured internally by the DMA controller for each virtual channel to manage the descriptor access order within the corresponding channel region. Channel pointers can also be implemented using hardware address registers or state machines to record the current operation progress of the channel. The read pointer, within the channel pointer group, manages the descriptor read progress and points to the address of the next descriptor within the channel region that will be read and used for data transfer by the DMA engine. The write pointer, within the channel pointer group, manages the descriptor write progress and points to the next free address within the channel region that can be used to store a new prefetched descriptor.
[0066] The default descriptor in a descriptor chain can include the first descriptor in the chain. Furthermore, if the descriptor chain further includes multiple sub-chains, the default descriptor can also include the last descriptor in the first sub-chain. The address navigation field in the last descriptor of the first sub-chain indicates the total number of descriptors in the second sub-chain and its starting position. The second sub-chain is the next sub-chain after the first sub-chain. For example, if the first sub-chain is the i-th sub-chain, the second sub-chain is the (i+1)-th sub-chain, where i is a natural number. A sub-chain can also be called a descriptor block, meaning that a descriptor chain further includes multiple descriptor blocks. Correspondingly, when a descriptor chain further includes multiple sub-chains, it belongs to a specific block-chain pattern within the chained pattern. Therefore, a sub-chain can refer to a logically independent and continuous sequence of descriptors within a large descriptor chain. Each sub-chain can be stored contiguously in memory, but different sub-chains are usually not contiguous. Sub-chains can serve as the basic unit for block management.
[0067] In chained mode, the first descriptor in the descriptor chain typically refers to the first descriptor directly written to the DMA channel by software through the configuration interface. In block-chained mode, the last descriptor in a sub-chain refers to the last descriptor of that sub-chain. Unlike ordinary descriptors, the address navigation field of the last descriptor is not invalid; instead, it provides navigation information pointing to the next task unit (the next sub-chain), thus serving a linking and navigation function.
[0068] For example, in block-chain mode, the descriptor chain can be divided into N sub-chains. The internal descriptors of each sub-chain are stored contiguously. The address navigation field of the tail descriptor of the current sub-chain contains the starting address of the next sub-chain. After processing the current sub-chain, the DMA controller automatically loads the information for the next sub-chain and jumps to it. This method supports the movement of non-contiguous memory data. By breaking down a large task's single, high-latency prefetch into multiple small-batch, low-latency prefetches, it significantly reduces reliance on large blocks of contiguous physical memory and mitigates the impact of memory fragmentation.
[0069] Optionally, when the first subchain is the last subchain among multiple subchains, the address navigation field in the last descriptor of the first subchain is configured to point to the starting address of the first subchain. During configuration, the address navigation field of the tail descriptor of the last subchain can be set to the starting address of the first subchain. After processing the last subchain, the DMA controller automatically jumps back to the first subchain and begins a new loop without software intervention. This method enables cyclic data transfer, and is particularly suitable for scenarios requiring continuous data transfer, such as audio playback, video display, and real-time data stream acquisition.
[0070] In this application, the address navigation field refers to a key control information field stored in the descriptor that guides the DMA controller to prefetch subsequent descriptors. This field includes at least a quantity field and an address field. For brevity and to adapt to different contexts, the starting memory address indicated by the address field may be described as starting address information, starting position, or starting address in different parts of this application. It should be understood that these descriptions have the same technical meaning in this application, all referring to the starting address of the descriptor memory obtained through the address navigation field for initiating the prefetch operation. Specifically, depending on the context, the object referred to may be the starting address of the entire descriptor chain, the starting address of the next descriptor subchain, or the first address of a circular descriptor chain.
[0071] In the above approach, considering that descriptors can exist in multiple modes, to facilitate rapid determination of the descriptor's mode, after obtaining the channel descriptor corresponding to any DMA channel in the DMA controller, the following operations can be performed: Obtain the preset mode field from the channel descriptor; if the channel descriptor belongs to chained mode based on the field value of the preset mode field, determine whether the channel descriptor is a preset descriptor in the descriptor chain based on its position in the descriptor chain. The preset mode field can be a specific bit or field in the descriptor data structure, used to indicate the operating mode to which the descriptor belongs. For example, a bit of 1 indicates chained mode, and 0 indicates single-transfer mode. Accordingly, after reading the descriptor, the DMA controller first parses the preset mode field. If it is determined to be in chained mode, the chained processing logic is activated (such as checking whether it is a preset descriptor, initiating prefetching, etc.); if it is determined to be in non-chained mode, an interrupt notification software is generated after a single transfer. This approach enables a single DMA controller to be compatible with multiple operating modes, improving the versatility and utilization of hardware resources, and allowing the selection of the most suitable transfer mode for different tasks.
[0072] To facilitate understanding, an example is provided below to illustrate the specific implementation details of the data transfer method provided in this application. This example is used to implement a descriptor prefetching mechanism for DMA devices in high-latency and small-packet transmission scenarios, involving technical fields related to high-speed data transfer and DMA control optimization. In modern SoCs, heterogeneous accelerators (such as CPU + GPU), and high-performance computing systems, there is a frequent need for large-scale data exchange between peripherals and memory, which places higher demands on transmission latency, bandwidth utilization, and resource efficiency. This example relates to direct memory access technology in computer systems, particularly to implementing a descriptor prefetching mechanism for DMA devices in scenarios such as high-latency memory access or small-packet data transfer.
[0073] With the continuous improvement of data transfer rates in modern computer systems, especially in scenarios involving network interface controllers, storage controllers, and high-speed peripherals, DMA typically needs to process a large number of small data packets. Each data packet usually corresponds to a descriptor, indicating the starting address of memory and the data length. In related technologies, after processing each descriptor, DMA often needs to wait to fetch the next descriptor from memory, which may cause pipeline stalls or a decrease in throughput, especially in systems with high memory access latency. To address this issue, this example proposes a descriptor prefetching mechanism based on a watermark trigger. By prefetching multiple descriptors in batches when the remaining number in the DMA buffer (i.e., the shared buffer mentioned above) falls below a set low watermark (i.e., the first channel watermark threshold mentioned above), the continuity of the DMA pipeline is ensured, data transfer efficiency is improved, and performance loss caused by descriptor waiting is reduced. The DMA buffer is also called the shared descriptor buffer, which is the specific implementation of the shared buffer mentioned above. Furthermore, by triggering a descriptor prefetching pause mechanism when the number of prefetched descriptors in the buffer reaches the high watermark (i.e., the second channel watermark threshold mentioned above), it is possible to reduce bus bandwidth usage and ensure multi-channel fairness. It should be noted that this example only illustrates a watermark-triggered watermark control strategy. In practice, those skilled in the art can configure various other types of channel prefetching strategies for each channel, such as channel priority-based prefetching strategies, channel load-based prefetching strategies, etc.
[0074] In one related technology, the DMA engine manages data transfer through a descriptor table (or PRD table). Each descriptor records the starting memory address and data length required for a single data move. For example, the DMA engine internally sets up a descriptor prefetch buffer to temporarily store descriptors already read from memory. This allows the engine to directly fetch the next descriptor from the buffer when processing data transfer corresponding to the current descriptor, theoretically enabling continuous transfer and improving the overall throughput of the DMA. Specifically, when there is no available space in the prefetch buffer, regardless of the current DMA processing speed, the DMA control logic requests the maximum capacity of descriptors and fills the buffer from the memory table all at once, maximizing the prefetch amount and utilizing the cache space.
[0075] Therefore, the above method sets up a descriptor prefetch buffer inside the DMA controller and adopts a greedy prefetch strategy of "initiating prefetching at maximum capacity as long as there is any empty space in the buffer": when an available space appears in the buffer, the DMA engine requests descriptors from memory equal to the maximum capacity of the buffer to fill the buffer as much as possible, thus hoping that when processing the current descriptor, the next one or more descriptors will already be ready in the buffer, thereby achieving continuous transfer. Although this implementation can improve continuity under ideal conditions, it has several fundamental drawbacks:
[0076] (1) When the prefetching and consumption speeds are mismatched, a large number of redundant descriptors that have been retrieved but cannot be stored will be generated, directly causing a waste of memory / bus bandwidth and invalid work. For example, since the processing speed of DMA and the prefetching speed may not match, if the DMA consumes descriptors slowly while the prefetching speed is fast, the prefetch buffer may be filled up quickly, resulting in the inability to store the extra descriptors retrieved from memory, which in turn leads to the descriptors being discarded, thus wasting memory access bandwidth and system resources.
[0077] (2) In scenarios with high memory access (or bus) latency, if the prefetch latency is higher than the consumption rate, the DMA may need to wait for the next descriptor to arrive after completing the current descriptor, causing pipeline pauses and thus failing to guarantee continuous transmission. For example, if the DMA consumes descriptors quickly, but the prefetch operation latency is high, the next descriptor may not be prefetched into the buffer after the current descriptor data transmission is completed. The DMA must wait for the arrival of the next descriptor, resulting in a short pause in data transmission and failing to guarantee the continuity of the DMA pipeline.
[0078] (3) In the above global greedy prefetching method, there is a lack of back pressure and isolation mechanism for multi-channel shared resources (such as global descriptor buffer and memory bandwidth), which can easily lead to a single channel repeatedly occupying the cache or monopolizing the bandwidth for a long time, resulting in starvation of other channels or a decrease in fairness.
[0079] (4) The above-mentioned global greedy prefetch strategy lacks hysteresis and dynamic adjustment mechanisms (no high / low water level intervals), which can cause frequent start-stop and performance jitter under load changes or short burst scenarios. The above disadvantages are particularly prominent in DMA scenarios for high-latency or small packet (each descriptor corresponds to a small data block) transmission: small packets lead to frequent descriptor switching, memory latency amplifies the pause cost caused by insufficient prefetching, and greedy prefetching amplifies the negative effects of bandwidth waste and descriptor discarding. For example, the above technical solution cannot dynamically adjust according to the real-time status of the DMA pipeline or the water level of the prefetch buffer, but is limited to requesting the maximum capacity at a fixed time. This may result in DMA idle waiting or buffer resource waste under different loads or different transmission scenarios, reducing the overall transmission efficiency.
[0080] To address the aforementioned issues, this example proposes a descriptor prefetching method based on low-watermark and high-watermark control, combined with a shared descriptor buffer and a quota / hysteresis mechanism. The basic idea is to maintain programmable high-watermark and low-watermark levels on a globally shared descriptor buffer, using the number of descriptor entries as the unit of measurement. When the number of available descriptors in the buffer falls below the low-watermark, batch prefetching (prefetching multiple descriptors at once) is triggered. Conversely, when the buffer's usage reaches or exceeds the high-watermark, all new prefetch requests are stopped until the buffer falls back below the low-watermark. In scenarios where multiple DMA channels share the same global descriptor buffer, the minimum descriptor quota for each DMA channel can be combined, and a hysteresis interval between the low-watermark and high-watermark levels can be used to avoid frequent switching.
[0081] Key implementation points of the scheme in this example include: before triggering the prefetch operation, checking and limiting the number of descriptors after this prefetch to no more than the high watermark, and prefetching multiple descriptors at once when allowed to compensate for the round-trip overhead under high latency. Furthermore, a higher high watermark can be configured for high-priority channels or short-term exclusive scenarios. Since each DMA channel can correspond to different high and low watermarks, and the high and low watermarks of each DMA channel can be flexibly configured and adjusted programmatically, the watermark quota for each channel can be adjusted urgently through register tuning via firmware / driver, thereby adapting to systems with different latency and packet length distributions.
[0082] As can be seen, the proposed solution in this example introduces a water level triggering mechanism, which only performs a prefetch operation when the number of available descriptors in the prefetch buffer is lower than a set threshold, and prefetches multiple descriptors at once. This ensures that the DMA can continuously acquire subsequent descriptors after the current descriptor data transfer is completed, thereby guaranteeing that the DMA pipeline does not stop, while reducing the risk of descriptor discarding and resource waste, and achieving more efficient and reliable data transfer.
[0083] The solution in this example addresses the shortcomings of related technologies by making at least the following improvements:
[0084] (1) By initiating bulk prefetching only when the buffer is below the low water level and stopping prefetching when the high water level is reached, the number of descriptors that have been retrieved but have nowhere to be stored can be significantly reduced, fundamentally reducing memory / bus bandwidth waste and descriptor discarding caused by redundant prefetching.
[0085] (2) By using low-water triggering to prefetch multiple descriptors in batches, the workload of a single prefetch operation can be amplified to cover the round-trip latency of memory access, thereby ensuring that the subsequent descriptor is in place when the current descriptor is completed in a high-latency environment, eliminating or shortening the pause time of the DMA pipeline, and improving the continuity and throughput in small packet transmission scenarios.
[0086] (3) By adopting a global high-water back pressure of shared buffer and quota and priority constraints for each channel, the system effectively prevents a single channel from occupying the descriptor buffer or bandwidth without limit under high load, thereby improving system fairness and reducing the starvation risk of other channels. Since the shared descriptor buffer in this example is shared by multiple DMA channels, it is easy to flexibly configure the specific parameters corresponding to each channel, thereby achieving flexible adjustment to fully meet the different load requirements of different channels.
[0087] (4) By using a DMA descriptor prefetching mechanism based on low / high watermark hysteresis control and a shared buffer, the inefficiency and resource waste caused by an overly aggressive prefetching strategy can be resolved. For example, related technologies typically perform maximum capacity prefetching immediately when there is empty space in the buffer, which can easily cause frequent start-stop operations, wasting bus bandwidth and potentially leading to the accidental discarding of descriptors. This example introduces a dual-watermark hysteresis interval design: prefetching is triggered only when the number of available descriptors in the buffer is below the low watermark, and prefetching stops once the high watermark is reached, thus avoiding frequent start-stop operations. At the same time, emergency refill and short-term boost mechanisms are added to ensure rapid response and maintain low latency even under instantaneous high load scenarios. Furthermore, this example reduces the impact of high memory latency through batch prefetching and ensures fairness when multiple DMA channels share buffer resources through global backpressure control and channel quota management, preventing individual channels from monopolizing resources, thereby significantly improving the actual performance and resource utilization efficiency in high-latency and small-packet transmission DMA devices.
[0088] Therefore, this example can solve at least the following technical problems:
[0089] (1) This example solves the problem that traditional DMA is difficult to efficiently handle small data packet transmission: Many DMA designs are well optimized for large data packet transmission, but when the data consists of many small packets (each packet contains a small amount of data and is frequent), frequent descriptor acquisition leads to frequent task start-up and interruption. The switching in between results in low efficiency, high latency and poor throughput. Therefore, a DMA controller design method is needed to maintain high throughput and low latency even in the case of small packets (i.e. small transactions and frequent transactions).
[0090] (2) This example addresses the problems of inflexible resource allocation, wasted buffer resources, inefficient descriptor prefetching, and difficulty in balancing throughput and resource utilization in DMA controllers under multi-channel / multi-task / mixed load (large and small packets) conditions. It proposes a DMA design method with a shared buffer and adjustable buffer depth. By allowing multiple channels to share the same shared buffer, and dynamically adjusting the buffer depth / prefetch depth used by each channel based on its current task load, packet size, and transmission frequency, this approach ensures that, on the one hand, the DMA has sufficient resources to maintain high throughput under high load or large packet transmission; on the other hand, it avoids excessive prefetching under low load and small packet transmission, thereby saving memory, bandwidth, and hardware resources, improving overall system resource utilization efficiency, and reducing power consumption.
[0091] Figure 2 The overall structure diagram of the DMA controller in this example is shown. Figure 2 As shown, the DMA controller includes three sets of I / O interfaces based on the AXI 4.0 transfer protocol: a slave interface, a first set of master interfaces, and a second set of master interfaces. The slave interface, also called the AXI slave interface, is used to receive basic configuration operations and various control commands triggered by the upper-level module for the DMA controller. The first set of master interfaces, also called the AXI master 0 interface, connects to the SoC's external memory (Host side, abbreviated as H) and is used to write data from the DMA to or read data from the external memory. The second set of master interfaces, also called the AXI master 1 interface, connects to the SoC's on-chip memory system (Local DDR side, abbreviated as D) and is used to write data from the DMA controller to or read data from the on-chip memory. The two sets of master interfaces can operate in a multi-channel statistical time-division multiplexing mode. On the software side, a specific DMA channel can be enabled through the AXI slave interface, and the first descriptor can be configured for that DMA channel. The descriptor occupies 8 words, or 32 bytes in size.
[0092] like Figure 2 As shown, the slave interface receives configuration and commands from the system bus (such as the CPU or other master devices) and serves as the interaction entry point between the DMA controller and external control logic. The register configuration module connects to the slave interface via the interface conversion module, used to parse and store the global and channel-specific configuration parameters of the DMA (such as transfer mode, address, data length, etc.). The DMA controller coordinates the operation of multiple DMA channels based on the settings of the register configuration module and manages the allocation and prefetching strategy of the shared buffer. DMA channels 0, 1, and 2 can function as independent data transfer execution units, each capable of handling one transfer task independently. Figure 2This explanation uses three DMA channels as an example; in practice, the number of DMA channels is usually much greater. The shared buffer, serving as a data / descriptor cache area shared by multiple DMA channels, dynamically allocates resources based on the load of each channel to improve memory access efficiency and bandwidth utilization. The channel arbitration module schedules and arbitrates requests from multiple DMA channels simultaneously to access the shared buffer or main interface, based on preset priority or fairness policies. The data receiving and sending module, acting as the physical interface layer for data transmission, handles data exchange with external devices or memory through the first set of main interfaces and / or the second set of main interfaces.
[0093] Figure 3 A schematic diagram of the descriptor structure is shown. For example... Figure 3 As shown, the descriptor defines the source, destination, data volume, and control logic of data transmission, and supports automatic prefetching of descriptors through link addresses to support multiple data transmission modes. Figure 3 The descriptor in the system can be divided into four main areas: control and mode area, transmission parameter area, address area, and link address area.
[0094] First, regarding Figure 3 The meaning of each field in the document is explained below:
[0095] NXT_LB_DESC_ACNT: Number of Next Link Block Descriptors (Next Link Block DescriptorAmount);
[0096] TR_DATA_ACNT: Transfer Data Amount (bytes);
[0097] SRC_ADDR_L: Low 32 bits of the source address;
[0098] SRC_ADDR_H: Source Address High 32 bits;
[0099] DST_ADDR_L: Destination Address Low 32 bits;
[0100] DST_ADDR_H: Destination Address High 32 bits;
[0101] LB_ADDR_L: Low 32 bits of the link block address;
[0102] LB_ADDR_H: Link Block Address High 32 bits;
[0103] RESERVED: Reserved bit, indicating that this bit field is not currently in use and is reserved for future expansion;
[0104] LBA_TYPE_EN: Link block address type enabled.
[0105] Among them, LBA_TYPE_EN is a control bit used to indicate the type of the current descriptor:
[0106] 0: The current descriptor is a data transfer type descriptor.
[0107] 1: The current descriptor is a Link Block Address Type Descriptor.
[0108] All address fields (SRC / DST / LB) are divided into high and low 32 bits to support a 64-bit address space. The TR_DATA_ACNT field is valid when LBA_TYPE_EN=0 and indicates the number of data bytes to be transmitted. The NXT_LB_DESC_ACNT field is valid when LBA_TYPE_EN=1 and indicates the number of the next link block descriptors.
[0109] Control and mode areas correspond to Figure 3 In WORD0, the LBA_TYPE_EN field is the link block address type enable field, which is used to indicate the descriptor type. When this field is 0, it means that the current descriptor is a data transfer type descriptor; when this field is 1, it means that the current descriptor is a link block address type descriptor. That is, this descriptor is mainly used to provide the address of the next descriptor, and the data transfer related fields (such as source address, data volume) are invalid.
[0110] Additionally, WORD0 can also serve as a prefetch mode enable field. When enabled, the DMA controller will prefetch the next descriptor into the DMA's internal descriptor buffer via the AXI Master interface, using the address provided by the LB_ADDR field (WORD6, WORD7), before the current data transfer is complete, to achieve continuous operation. Furthermore, WORD0 can also represent the prefetch quantity (Bits 31:16). When prefetch mode is enabled, it specifies the number of descriptors to be prefetched consecutively starting from the address indicated by LB_ADDR, thus allowing multiple descriptors to be preloaded in batches to further improve efficiency. LB_ADDR represents the lower 32 bits of the link block address, and LB_ADDR_H represents the higher 32 bits of the link block address.
[0111] In particular, depending on the needs of the actual use case, the WORD0 field can simultaneously carry multiple semantics such as prefetch mode enable and prefetch quantity, depending on the use case of the descriptor. In addition, the WORD0 field can also represent the specific mode of the descriptor.
[0112] The TR_DATA_ACNT field of WORD1 is used to characterize the amount of data to be transferred (Bits 31:0), which defines the total number of bytes of data to be transferred in this descriptor operation.
[0113] In WORD2 to WORD5, the SRC_ADDR_L field is used to represent the lower 32 bits of the source address, the SRC_ADDR_H field is used to represent the higher 32 bits of the source address, the DST_ADDR_L field is used to represent the lower 32 bits of the destination address, and the DST_ADDR_H field is used to represent the higher 32 bits of the destination address.
[0114] In WORD6 and WORD7, the LB_ADDR_L field represents the lower 32 bits of the link block address, i.e., the lower 32 bits of the address where the next descriptor is stored. This field is valid only when prefetch mode or link block address type is enabled (e.g., Bit0=1 in WORD0) and is used by the DMA's AXI Master interface to prefetch the next descriptor. Similarly, the LB_ADDR_H field represents the higher 32 bits of the link block address, i.e., the higher 32 bits of the address where the next descriptor is stored, thus combining with LB_ADDR_L to form a complete 64-bit link address.
[0115] Depending on the descriptor configuration (such as the control bits of WORD0), the DMA channel can operate in the following three modes:
[0116] (1) Single Mode: Usually PREFETCH_EN is 0. DMA performs only one data transfer as defined for the current descriptor, and stops or waits to configure a new descriptor after completion.
[0117] (2) Chain Mode: PREFETCH_EN is 1. After the data transfer of the current descriptor is completed, the DMA automatically loads the next descriptor from the address pointed to by LB_ADDR and continues execution, forming a task chain.
[0118] (3) Block Chain Mode: PREFETCH_EN is 1 and PREFETCH_NUM is greater than 1. When the DMA is executing the current transfer, it will continuously prefetch a specified number of descriptors from the starting address of LB_ADDR into the shared buffer to form a descriptor block, and then execute them in sequence to further improve the efficiency of continuous processing.
[0119] Therefore, WORD2 and WORD3 in the descriptor indicate the source address for data transfer, meaning the DMA reads data from the source address back to the shared buffer inside the DMA for storage via the AXI master port. WORD1 indicates the amount of data transferred, i.e., the amount of data read from the source address. WORD4 and WORD5 indicate the destination address for data transfer, i.e., the DMA reads data from the shared buffer and writes it to the destination address via the AXI master interface. Bit 0 in WORD0 indicates whether it is in prefetch mode. In prefetch mode, new descriptors are read into the shared buffer inside the DMA via the addresses indicated by WORD6 and WORD7 using the AXI master. Thus, after the channel completes the data transfer of the previous descriptor, it can immediately read new descriptors from the shared buffer and enable a new round of data transfer. Bits [31:16] in WORD0 indicate the number of consecutive descriptors prefetched from the prefetch address. Based on the structure of the descriptor, this example divides the channel's data transfer mode into the three modes mentioned above: single-pass mode, chained mode, and block-chained mode.
[0120] In single-pass mode, when the software configures the first descriptor, WORD0 bit[0] is disabled (bit[0]=0). This channel only completes the data transfer of the current descriptor and does not prefetch new descriptors. After the current data transfer is completed, the software will be notified via a hardware interrupt. The software needs to reconfigure the descriptor according to the requirements to complete a new round of data transfer. In single-pass mode, from task startup to the completion of a complete data transfer, four core roles are involved: CPU (software), source (Src), DMA controller, and destination (Dst), which includes the following three stages:
[0121] Phase 1: Task Configuration and Startup. This phase involves software configuration and descriptor distribution. The software first prepares a descriptor in memory, containing all the key parameters for this transfer. Then, the software configures or distributes the address or content of this descriptor to the DMA controller via register writing or other methods, enabling the DMA controller to obtain the descriptor for this transfer task.
[0122] Phase Two: Data Transfer Execution. After obtaining the descriptor, the DMA controller automatically performs the following data transfer operations: First, the DMA controller parses the descriptor to obtain information such as the source address and the amount of data to be transferred, thereby initiating a read data transaction request to the source. Then, the DMA controller enters a waiting state. Upon receiving the read request, the source reads data from the specified address and returns the data to the DMA controller via the bus. The DMA controller receives and temporarily stores the data. The DMA controller performs necessary hardware-level processing on the raw data stream just received from the source. The most common processing includes data alignment, data splitting / merging, etc. Finally, the DMA controller, based on the destination address in the descriptor, initiates a write data transaction request to the destination to write the data to the target location. Next, the DMA controller enters a waiting state again to wait for the destination to return a completion response for all write transactions, thereby ensuring that all data has been successfully received and written by the destination.
[0123] Phase 3: Post-processing and Interrupt Generation. Once it's confirmed that all data has been successfully read from the source and written to the destination, the DMA controller determines that the single-mode transfer task is complete. Subsequently, post-processing operations are performed, such as triggering a completion interrupt to the CPU. In Single Mode, software-issued descriptors trigger only one complete transfer process. After the process ends, the DMA channel enters an idle state until the software configures a new descriptor.
[0124] Figure 4 The flowchart of the chained mode is shown, that is: while configuring the first descriptor on the software side, WORD0bit[0] is enabled, the addresses of the descriptors to be prefetched are given, namely the addresses of WORD6 and WORD7, as well as the number of descriptors to be prefetched, and the correct descriptors need to be placed in advance at the corresponding addresses. Thus, the enabled channel in the DMA will simultaneously initiate descriptor prefetching requests and data transfer requests, but during the arbitration process, descriptor prefetching will be prioritized. The advantage of this mode is that it ensures that after the current descriptor is consumed, the prefetched descriptors can be parsed immediately, ensuring that data can be processed continuously on the AXI master interface. Figure 4 As shown, the workflow of the chain pattern includes the following steps:
[0125] S401: Software preparation descriptor.
[0126] S402: Software configures the first descriptor into the DMA.
[0127] S403: DMA initiates a data transfer operation based on the descriptor.
[0128] S404: DMA is waiting for the data to be returned for the data transfer operation.
[0129] S405: DMA performs data alignment and splitting processes internally.
[0130] S406: The DMA writes the processed data to the destination address.
[0131] S407: DMA is waiting for a write response signal.
[0132] S408: DMA waits for the data transfer operation to complete and generates an interrupt.
[0133] S409: DMA uses the starting address to prefetch the descriptor.
[0134] S410: DMA is waiting for the result of the descriptor prefetch operation.
[0135] Therefore, in the DMA chained mode workflow, the execution entity switches between software (CPU side) and DMA controller hardware: steps S401 and S402 are executed by software, responsible for preparing and configuring the initial descriptor. From step S403 onwards, the process enters the DMA hardware operation stage. Steps S403 to S410 are all executed by DMA.
[0136] It should be noted that steps S409, S410, S403, and S404 can be executed concurrently. That is, by pre-fetching the next descriptor, the instruction fetching (acquiring the task) and execution (transferring data) stages are overlapped, hiding the latency of accessing memory to acquire the descriptor. This allows the DMA data transfer channel (AXI Master port) to remain continuously operational, thereby significantly improving the overall bandwidth and efficiency in continuous transfer scenarios.
[0137] Blockchain mode is an improvement on chained mode. Since each descriptor is 32 bytes, storing a large number of descriptors would occupy a large contiguous block of physical address space. Furthermore, if there are "circular" data movement requirements, or repeated movement requirements for a specific address block, the chained mode cannot be used. Additionally, in AI computing or heterogeneous computing, if there are massive data movement volumes, with the number of descriptors reaching thousands or even tens of thousands, the chained mode cannot meet these requirements. To solve these problems, in blockchain mode, the descriptor chain is further divided into multiple descriptor blocks (i.e., sub-chains). In the last descriptor in the prefetched descriptor (i.e., the last descriptor in the first sub-chain), bit 0 in WORD0 is reconfigured to enable it, and the starting address of the next prefetched descriptor (i.e., the first descriptor in the second sub-chain) and the number of descriptors to be prefetched in the next round (i.e., the counter value) are given. The DMA controller calculates the number of descriptors consumed using the counter value, and determines which descriptor among the descriptors to be prefetched is the last descriptor, i.e., the tail descriptor. Then, based on the new WORD0 indicator bit in the tail descriptor, it updates the number of descriptors and the starting address of the new descriptors to be prefetched.
[0138] Figure 5 A schematic diagram of the prefetching process in chain mode is shown. Figure 5 The meanings of the English abbreviations in this context are as follows: DESC: Descriptor. Block: Link block; for example, "descriptor block" means "descriptor block". Figure 5As shown, unlike the simple chained mode where each descriptor points to the next individual descriptor, the block-chained mode organizes descriptors into multiple descriptor blocks. Each descriptor block contains multiple descriptors, and the last descriptor within a block points to the next descriptor block. The software first configures the DMA controller with the starting address of descriptor block 0 (i.e., descriptor block 0) and the number of descriptors N to be prefetched. The DMA controller prefetches the entire descriptor block 0 (including DESC 0 to DESC N-1) from external memory at once. Then, the DMA controller executes the descriptors (DESC 0, DESC 1, ...) in descriptor block 0 sequentially. The link address fields (i.e., address navigation fields) in multiple descriptors within the same block are typically invalid prefetch addresses because the order of the fields is implicitly determined by the physical location within the block, eliminating the need for explicit linking and thus saving storage overhead. When the execution reaches the last descriptor (DESC N-1) of descriptor block 0, the last descriptor contains the following two key pieces of information: (1) the starting address of the next descriptor block: pointing to the starting location of descriptor block 1 in memory; (2) the prefetch count (M) of the next descriptor block: indicating the number of descriptors to be prefetched next time. Before the data transfer defined by DESC N-1 is completed, the DMA controller uses the provided new address and new count to initiate a prefetch request for descriptor block 1 in advance. When descriptor block 0 is completed, the descriptor of descriptor block 1 is ready, thus achieving seamless connection. This process can be repeated (for example, the tail descriptor of descriptor block 1 can point to and prefetch descriptor block 2). By grouping descriptors by blocks, the software overhead and hardware complexity of managing and finding individual descriptors in a very long task chain can be greatly reduced. It is particularly suitable for circular buffer operations or repeated movement of specific address regions. A loop can be formed simply by pointing the tail descriptor of the last block to the starting block.
[0139] When transmitting small data packets in block-chain and chained modes, the following problems exist under high latency: all current descriptors have been consumed, but newly prefetched descriptors have not yet returned to the shared buffer. Even if a certain number can be prefetched each time, the design still needs to wait for all descriptors to be consumed before starting a new round of prefetching, thus causing performance issues in small packet transmission. For example, continue to refer to... Figure 4 , Figure 4The prefetch operations performed in steps S409 and S410 completely overlap with the data transfer operations performed in steps S403 and S404. After the software configures the first descriptor, the DMA immediately initiates a prefetch request for the next descriptor, and the DMA executes the data transfer process defined for the current descriptor in parallel. Specifically, the prefetch operation initiated in S409 is completed before S408 (current transfer complete) (S410). Therefore, when the current transfer ends, the next descriptor is ready, and the DMA can immediately start the next round of transfer with zero wait. However, in real-world scenarios with small packet transfers and high access latency, the amount of data defined for each descriptor is very small, resulting in a very short data transfer phase (S403-S408). Consequently, the prefetch speed often cannot keep up with the consumption speed, leading to high latency issues. For example, the operation of prefetching a descriptor from memory (S409-S410) may take a long time, far exceeding the time for small packet transfers, resulting in a long waiting time for the descriptor. To address the above issues, this example designs a descriptor prefetching mechanism based on water level control.
[0140] Figure 6 The diagram illustrates the descriptor prefetching process implemented in chained mode based on a water level control strategy, specifically including the following steps:
[0141] S601: Configure the buffer depth corresponding to the channel descriptor of any DMA channel via hardware.
[0142] Specifically, the hardware configures the buffer depth of a channel descriptor in the shared buffer for a given channel. For example, buffer depth = N0 means that the default buffer depth of a DMA channel's channel descriptor is N0.
[0143] S602: Configure the channel water level threshold for any DMA channel.
[0144] For example, the channel watermark thresholds can be configured through the AXI slave interface. For instance, for a specific channel, the low watermark [31:0] and high watermark [31:0] registers can be set in the corresponding register (such as the prefetch address register). For example, low_watermark[31:0] = N1, high_watermark[31:0] = N2, and N2 > N1, N2 ≤ N0, and N2 - N1 > N1.
[0145] S603: Configure the first descriptor.
[0146] Specifically, the first descriptor is configured through the AXI slave interface. For example, the address navigation field (i.e., the WORD0 field) in the first descriptor must indicate that the current descriptor is in Chain mode, and the number of descriptors to be prefetched is N, i.e., the number to be prefetched = N. For example, the total number N of descriptors in the descriptor chain can be determined based on the value of the second field in the address navigation field of the channel descriptor.
[0147] S604: Compare the quantity to be pre-sampled with the high water level in the channel water level threshold, i.e., determine whether N is greater than N2.
[0148] S605: Determine that the prefetching will be divided into at least two batches, and perform the prefetching operation for the first batch.
[0149] If N > N2, then the prefetch size for the first batch is determined to be N2, meaning the first prefetch size is N2. When the prefetch is complete, the number of descriptors in the shared buffer may reach the high-water mark.
[0150] S606: Wait for the number of available descriptors in the shared buffer to drop to a low level in the channel water level threshold.
[0151] After the first batch of prefetch operations is completed, wait for the number of available descriptors in the shared buffer to be N1.
[0152] S607: Calculate the number of remaining unprefetched descriptors.
[0153] The number of remaining unprefetched descriptors is equal to the difference between the number of descriptors to be prefetched (N) and the number of descriptors already prefetched (N²). For example, the number of remaining unprefetched descriptors left_n = N - N².
[0154] S608: Determine if the number of remaining unprefetched descriptors is greater than N2-N1.
[0155] This step aims to determine whether the remaining capacity of the current buffer is sufficient to accommodate the remaining unprefetched descriptors at once. Since the current buffer currently has N1 available descriptors, to ensure that the number of available descriptors after the next prefetch does not exceed N2, the number of descriptors prefetched in the next batch cannot be greater than N2-N1.
[0156] S609: If the number of remaining unprefetched descriptors is greater than N2-N1, then determine the number of descriptors to be prefetched in the next batch as N2-N1.
[0157] S610: If the number of remaining unprefetched descriptors is no greater than N2-N1, then determine the number of descriptors to be prefetched in the next batch as left_n.
[0158] Steps S606 to S609 are operations that are executed cyclically until all descriptors are prefetched.
[0159] S611: Determine to prefetch all descriptors at once.
[0160] If N is not greater than N², then the prefetch quantity is determined to be N, so that all descriptors are prefetched at once.
[0161] S612: All descriptors have been prefetched; the prefetching operation is complete.
[0162] As data migration continues, the available descriptors in the shared buffer are continuously consumed. When the number of remaining available descriptors in the shared buffer reaches the low watermark, a second prefetch is initiated. Each prefetch process must ensure that the shared buffer does not overflow after prefetching. If, within a certain period, the number of descriptors consumed is found to be too fast, and the prefetched number is lower than the consumed number, `high_watermark[31:0]` can be dynamically adjusted to increase the number of descriptors prefetched each time. If, within a certain period, the number of descriptors prefetched is found to be higher than the consumed number, `low_watermark[31:0]` can be increased and `high_watermark[31:0]` can be decreased to reduce the number of descriptors prefetched each time, thereby reducing bus occupancy during prefetching. Dynamic configuration requires waiting for the current descriptor to be prefetched before it takes effect, and the number of descriptors prefetched in the next transaction must be greater than N2-N1; otherwise, the dynamic configuration is invalid and meaningless. Figure 6 The descriptor prefetching scheme based on the water level concept shown can ensure that data transfer will not be interrupted due to descriptor prefetching in scenarios with high latency and small packet transmission, so that the AXImaster bus can always be in pipeline, which greatly improves the transmission efficiency in this scenario.
[0163] Additionally, in Block Chain Mode, to address performance bottlenecks caused by untimely descriptor prefetching in scenarios with small data packets and high latency transmission, this example... Figure 6 Based on the Chain Mode water level control strategy shown, prefetching is performed in conjunction with the preset descriptors (i.e., the last descriptor of each descriptor block) in Block Chain Mode. This method can identify and respond to key link points in the descriptor chain, thereby enabling a non-stop prefetch pipeline across descriptor blocks and eliminating idle waiting time between batches.
[0164] exist Figure 6In the previous method, the next batch of prefetching was only triggered when the descriptors in the buffer were consumed to the low watermark (N1). However, in Block Chain Mode, it is necessary to identify the special Tail descriptor in the descriptor chain. This descriptor not only defines its own data transfer task, but also indicates the starting address (New LB_ADDR) and the total number (Nb1) of the next descriptor block through configuration information in WORD0 (such as bit0 enabling). Assuming the initial state in Block Chain Mode is the same as... Figure 6 Similarities: The buffer depth is N0, the high / low watermarks are N2 and N1 respectively, the total number of descriptors to be prefetched is N, and the entire prefetch process can be divided into the following two stages:
[0165] In the first stage, initial batch pre-fetching is performed based on water level. The process in this stage is similar to... Figure 6 Similarly, if N > N2, prefetch in batches. The first batch prefetches N2 descriptors. When the number of available descriptors in the waiting buffer drops to a low watermark N1 due to DMA consumption, calculate the remaining unprefetched number left_n = N - the cumulative prefetched number. Compare left_n with the maximum capacity that can be prefetched in a single batch (N2 - N1): if left_n > (N2 - N1), initiate the next batch prefetch with a quantity of (N2 - N1); if left_n ≤ (N2 - N1), initiate the final batch prefetch with a quantity of left_n. Repeat the above operation until all left_n descriptors are prefetched. The first phase aims to ensure the utilization of the buffer through batch prefetching.
[0166] In the second phase, cross-block prefetching is dynamically triggered based on preset descriptors (such as the Tail descriptor) in the descriptor chain. This second phase resolves the latency issue of batch switching. After the first phase is complete and all left_n descriptors are prefetched back to the buffer, if the WORD0 bit0 of the last descriptor (i.e., the Tail descriptor) is enabled, it is determined that the current chain needs to enter block chain mode, and the starting address of the next descriptor block and the total number Nb1 are parsed from that descriptor. Additionally, the current state of the shared buffer can be checked. Assuming that the number of descriptors remaining in the shared buffer that have not been consumed by DMA is left_b0, different prefetching strategies can be determined based on the comparison between the total inventory (left_n + left_b0) and the low-water mark N1:
[0167] Case A: left_n + left_b0 < N1. In this case, even if the DMA consumes all existing descriptors in the buffer (including the just prefetched left_n descriptors), the total amount will not reach the low watermark N1, so a low watermark-based prefetch cannot be triggered. Therefore, a prefetch request for the next descriptor block (Nb1) needs to be initiated immediately. To ensure that the buffer capacity N2 is not exceeded, the actual prefetch quantity is min(Nb1, N2 - (left_n + left_b0)). That is, the smaller value between Nb1 and the current remaining space in the buffer is taken.
[0168] Scenario B: left_n + left_b0 ≥ N1. In this case, the buffer's current capacity can support DMA consumption for a period of time and has the potential to trigger a low-watermark-based prefetch. Therefore, instead of prefetching immediately, we can continue to wait. When DMA continues to consume descriptors, causing the buffer capacity to drop to the low-watermark N1, the next prefetch is triggered. When triggering the next prefetch, the prefetch quantity is min(Nb1, N2 - N1). That is, the prefetch quantity is the smaller of Nb1 and the difference between the high and low watermarks (the maximum safe prefetch quantity per operation).
[0169] Regardless of the situation, once a new descriptor block (of number Nb1 or an adjusted value) is prefetched, the last descriptor in that new descriptor block may be a new Tail descriptor, thus repeating the second phase operation. This process continues until the entire descriptor chain is prefetched and executed.
[0170] It can be seen that in the Block Chain mode, the same water level design idea is adopted, and the second stage is repeatedly executed until left_n ≤ (N2 - N1), where left_n represents the number of descriptors prefetched in the (n + 1)-th time and is left_n. When all the descriptors with the quantity of left_n are prefetched, if it is judged that the WORD0 bit0 of the tail descriptor (i.e., the N-th one) is enabled, then it is judged that the prefetch mode is Block Chain Mode, and a new prefetch address and a new prefetch quantity Nb1 can be obtained from the tail descriptor. Suppose that when all the left_n descriptors are prefetched, the remaining number of descriptors in the shared buffer is left_b0. If left_n + left_b0 < N1 and the descriptors can never be consumed to the low water level at this time, then if N2 ≥ Nb1 + (left_n + left_b0), a prefetch of the (n + 2)-th descriptor is immediately initiated with the quantity of Nb1. If N2 < Nb1 + (left_n + left_b0), a prefetch of the (n + 2)-th descriptor is also immediately initiated, and the prefetch quantity is N2 - (left_n + left_b0); the above operations are repeatedly executed subsequently until all the prefetching and data transfer are completed. If left_n_n + left_b0 ≥ N1, when the descriptors are continuously consumed to the low water level, if Nb1 ≤ N2 - N1, a prefetch of the (n + 2)-th descriptor is initiated with the quantity of Nb1; if Nb1 > N2 - N1, the prefetch quantity of the (n + 2)-th descriptor is N2 - N1; the above operations are repeatedly executed subsequently until all the prefetching and data transfer are completed.
[0171] Based on the above descriptor prefetch mechanism, it is necessary to reserve descriptor storage buffer space for each channel. The DMA initiates a new descriptor prefetch through the internal descriptor quantity counter and the descriptor prefetch address register according to the consumption of the internal descriptors in the buffer of this channel and the prefetch quantity that can be stored in the current buffer. The precise control method in this example can solve greedy prefetching, and can also solve problems such as bandwidth waste and descriptor discarding.
[0172] In practice, computer systems have diverse data transfer needs, including multiple peripherals, multiple tasks, and multiple data streams. They simultaneously require high throughput (large data blocks) and low latency (small data packets and real-time requirements), while also needing the CPU to remain idle to process computational logic. Single-channel DMA cannot meet these complex demands. Therefore, multi-channel DMA controllers have become the mainstream design, capable of serving multiple channels concurrently or using statistical time-division multiplexing, performing intelligent scheduling and priority management, and providing the system with efficient, flexible, and scalable data transfer capabilities. In multi-channel DMA controllers, a separate descriptor buffer is typically statically reserved for each channel to store prefetched descriptors, supporting independent operation of each channel. The problem with this approach is that when the number of channels is large or channels are idle for extended periods / transfer tasks are sparse, the large buffer area required by all channels combined leads to wasted hardware resources. In the shared descriptor-buffer pool mechanism proposed in this example, all channels uniformly use the same static random-access memory buffer (SRAM buffer), instead of statically allocating independent buffers for each channel. The buffer depth occupied by each channel in the shared buffer pool is statically controlled through register configuration. The system can dynamically set the required buffer resources for each channel based on its actual task type (large data transfer or frequent small packet transmission), latency and throughput requirements, channel activity, etc. This ensures sufficient buffer resources to cover the latency of prefetching descriptors and maintain high throughput during small packet or high-frequency tasks, while also significantly saving hardware buffer space and bus bandwidth resources under large packet, low-frequency, or single-channel load conditions, avoiding resource idleness and unnecessary prefetch waste.
[0173] The hardware depth of the SRAM buffer for the total descriptors of n channels is set to N by hardware configuration, and a shared_buffer_pool_en register can be added to the DMA internal configuration register. If shared_buffer_pool_en=0, the descriptor buffer depth allocated to each channel is fixed at N / n, and each channel sets high_watermark[31:0] and low_watermark[31:0] within the range of N / n. If shared_buffer_pool_en=1, the buffer depth needs to be set for each channel, and the sum of the buffer depths of all channels must be less than or equal to the hardware depth N. Figure 7The diagram shows the case where shared_buffer_pool = 0. Figure 8 The diagram shows the case where shared_buffer_pool = 1.
[0174] Figure 7 This diagram illustrates the static, fixed-partition structure of the descriptor buffer in physical memory when the shared buffer pool function of the DMA controller is disabled. The system reserves a contiguous physical memory space for descriptors of all DMA channels, with a total address range from 0 to N-1 and a total size of N descriptor storage units. This memory segment is equally divided into n fixed-size blocks, each allocated to a specific DMA channel (channel 0 to channel n-1) as its private descriptor buffer. Because the partitioning is equal, each channel receives a buffer of exactly the same size, N / n descriptor storage units. For example, channel 0's buffer occupies addresses from 0 to 1×(N / n)-1, channel 1 occupies addresses from 1×(N / n) to 2×(N / n)-1, and so on. As shown in the figure, the low water mark [31:0] = N1, that is, low_watermark[31:0] = N1. Therefore, the low water mark value of each channel is fixed at N1. The high water mark [31:0] = N2, that is, high_watermark[31:0] = N2. Therefore, the high water mark value of each channel is fixed at N2.
[0175] Figure 8 This illustrates the dynamic, variable partitioning structure used by the descriptor buffer when the shared buffer pool feature is enabled. A contiguous block of physical memory from 0 to N-1 is reserved. Unlike the disabled state, this memory is no longer pre-divided but rather serves as a unified shared buffer pool. The memory in the pool can be dynamically and non-uniformly partitioned during system initialization or runtime based on the actual needs of each DMA channel. Figure 8 As shown, channel 0 is allocated blocks from 0 to n0-1 (i.e., the channel region corresponding to channel 0), channel 1 is allocated blocks from n0-1 to n0+n1-1, and so on. The size of each channel's buffer can be different. Therefore, compared to... Figure 7 In other words, Figure 8 The descriptor buffer size for each channel is not a fixed value but can be flexibly configured and adjusted. Furthermore, Figure 8 The low and high water level lines for each channel can also be flexibly configured, meaning that different water level thresholds can be configured for different channels. Figure 7 and Figure 8 The address range and depth values shown are merely examples and can be flexibly adjusted by those skilled in the art.
[0176] The following describes several specific scenarios in this example:
[0177] I. Control the number of prefetches per burst based on burst length and burst size (transmission bit width).
[0178] Regarding how to prefetch the specific number of descriptors, the following example illustrates this: The AXI master data width is 512 bits, the size of a single descriptor is 32 bytes, the maximum burst length is 32, the default descriptor buffer depth for channel 0 is 48, and the channel is configured with low_watermark[31:0]=N1=16 and high_watermark[31:0]=N2=32. The prefetch mode is Chain mode, and the total number of descriptors to be prefetched is 101. Assuming each descriptor requires 4KB of data to be moved, and given the small packet mode and significant latency (32KB of data to be transferred before a single descriptor prefetch is completed), the first counter (total_counter) counts the total number of descriptors to be prefetched, and the second counter (buffer_counter) calculates the number of descriptors already stored in the channel buffer. The AXI behavior on the bus and the internal counting behavior are as follows:
[0179] (1) Before prefetching: total_counter=101, buffer counter=0;
[0180] (2) The first prefetch occurs: the prefetch quantity is 48, the AXI burst size (AXI burst_size) = 6, the AXI burst length (AXI_burst_length) = 24, and after the prefetched descriptor is read back, update the two counters mentioned above: total_counter = 53, buffer_counter = 48.
[0181] (3) As data migration continues, descriptors are continuously consumed. When buffer_counter = N1 = 16, i.e., when the low watermark is reached, the channel initiates a second prefetch, with a prefetch quantity of N2-N1=16. At this time, AXIburst_size = 6 and AXI_burst_length = 8. The 16 descriptors below the low watermark are still being consumed. After the descriptors are read back, the two counters mentioned above are updated, total_counter = 37. Assuming that 4 out of the original 16 descriptors at the low watermark are consumed, then buffer_counter = 4 + 16 = 20.
[0182] (4) Repeat step (3) and continue with two prefetches of 16 descriptors each. The number of descriptors in the last prefetch is 5. Since 5 < N2 - N1, the number of descriptors in the fifth prefetch is 5. At this time, AXI burst_size = 6 and AXI_burst_length = 3. After the descriptors are prefetched, update total_counter = 0. Assume that there are 2 remaining descriptors out of the original 16 low-watermark descriptors. Then buffer_counter = 2 + 5 = 7. Since it is Chain mode, after the remaining 7 descriptors are consumed, all transmission tasks are completed. Figure 9 The figure shows the schematic diagram of the process of each prefetch operation.
[0183] Similarly, if it is Block Chain mode, update total_counter through the last prefetched descriptor to determine the next prefetch quantity. In the above entire mechanism, only the size of AXI_burst_length is adjusted. In other methods, the sizes of parameters such as AXI_burst_size and AXI_burst_length can also be adjusted simultaneously to achieve more precise prefetching. For example, in the above example, if the number of descriptors in the last descriptor is 1, then AXI_burst_size = 5 can be adjusted and read through the narrow transaction in the AXI protocol.
[0184] In addition, in the above example, the maximum value of AXI_burst_length is 32. If the system does not support a larger AXI_burst_length, descriptor prefetching can also be performed in the form of multiple bursts. For example, if the maximum value of AXI_burst_length is 8 and the prefetch quantity is 48 descriptors with AXI_burst_size = 6, the same purpose can be achieved through 3 bursts with AXI_burst_length = 8.
[0185] II. Control the prefetch process based on the descriptor counter and the descriptor prefetch address register.
[0186] It has been elaborated above how to count the number of remaining descriptors that have not been prefetched and the number of descriptors that have been prefetched and can be consumed and parsed in the current descriptor buffer through total_counter and buffer_counter. Through precise control of the counter, it is possible to calculate how much space is left in the buffer of this channel and initiate precise prefetching. At the same time, in this example, the prefetch address can be managed by adding the prefetch address register chain_adderss[47:0] to ensure that the precise storage address of the descriptor can be determined during each prefetch.
[0187] For example, continuing with the previous example, the descriptor prefetch address for the first prefetch is configured into the chain_address[47:0] register of the corresponding channel 0 through the AXI slave interface. Assuming the initial address is 0x00, multiple prefetches can be performed through the following process:
[0188] (1) The first prefetch is 48, and the prefetch address is 0x0. At the same time as the AXI master sends out the data prefetch, the address of chain_address[47:0] is updated to 0x0+20H×48=0x600. Then the address of the next prefetch is 0x600.
[0189] (2) The second prefetch has a prefetch quantity of 16 and a prefetch address of 0x600. At the same time as the AXI master sends out the data prefetch, the address of chain_address[47:0] is updated to 0x600+20H×16=0x800. Then the address of the next prefetch is 0x800.
[0190] (3) The third prefetch, the prefetch quantity is 16, the prefetch address is 0x800. At the same time as the AXI master sends out the data prefetch, the address of chain_address[47:0] is updated to 0x800+20H×16=0xA00, then the address of the next prefetch is 0xA00;
[0191] (4) The fourth prefetch is 16, and the prefetch address is 0xA00. At the same time as the AXI master sends out the data prefetch, the address of chain_address[47:0] is updated to 0xA00+20H×16=0xC00. Then the address of the next prefetch is 0xC00.
[0192] (5) The fifth prefetch, the number of prefetches is 13, the prefetch address is 0xC00. At the same time as the AXI master sends out the data prefetch, the address of chain_address[47:0] is updated to 0xC00+20H×5=0xCA0. Since the descriptor prefetching is completed in chain mode at this time, there is no need to perform the next prefetch.
[0193] Additionally, if the mode is Block Chain mode, then based on the fifth prefetch, if WORD0 bit0 in the fifth descriptor is enabled, the DMA controller determines that the mode is Block Chain mode. At this time, chain_address[47:0] is updated to the addresses of WORD6 and WORD7, and the total_counter value is updated to the value of the WORD0[31:16] register, so that a new round of prefetching and counting of the two counters can be performed.
[0194] III. Mechanism for reading descriptors based on First-In-First-Out (FIFO) pointers controlling the shared buffer.
[0195] In this example, all channels share a shared buffer to store descriptors, reducing hardware resources. Simultaneously, adjusting the descriptor buffer depth for each channel controls latency, thereby improving throughput for small packet transmissions. A FIFO design is used to control the depth of the corresponding channel in each SRAM. For example, in... Figure 8 In the context of channel 2, if the descriptor buffer depth is n2, the FIFO read pointer for this channel will jump between ptr1 = n0 + n1 - 1 and ptr2 = n0 + n1 + n2 - 1. The buffer_counter is calculated based on the relationship between the read pointer rd_ptr and the write pointer wr_ptr. If wr_ptr >= rd_ptr (i.e., no wrapback), then buffer_counter = wr_ptr - rd_ptr; if wr_ptr < rd_ptr (i.e., FIFO wrapback occurs), then buffer_counter = wr_ptr - rd_ptr + fifo_depth. Here, fifo_depth is the depth of SRAM occupied by this channel, i.e., n2. Because a separate descriptor buffer is not reserved for each channel, but rather all channels share the same buffer pool, the overall buffer area (SRAM capacity) can be significantly reduced, lowering chip cost and hardware resource consumption.
[0196] In summary, this example demonstrates that by sharing the same SRAM buffer pool for storing prefetch descriptors across all channels, and employing a per-channel FIFO + programmable buffer depth + watermark triggering mechanism, the required buffer resources for each channel can be dynamically and statically allocated, thereby significantly improving system resource utilization efficiency. Specifically, in complex scenarios involving multiple channels, multiple tasks, and dynamic load changes (large packets, small packets, high frequency, low frequency), it has at least the following beneficial effects: (1) It avoids reserving fixed and redundant descriptor buffer space for each channel, thereby saving chip area and reducing hardware resource costs; (2) When some channels are idle or only handle large block transmissions, buffer allocation can be reduced to avoid waste caused by excessively deep buffers; (3) For frequent small packet or high-concurrency tasks, a sufficiently deep shared buffer + water-mark mechanism ensures timely replenishment of prefetched descriptors, avoiding transmission pauses caused by descriptor buffer exhaustion, thereby improving DMA throughput and response performance; (4) By combining resource conservation and performance assurance, the DMA controller achieves a good balance between throughput, transmission response and latency, hardware resource usage (area and power consumption), and system flexibility (supporting mixed loads, multiple tasks, and multiple channels).
[0197] Therefore, this example has at least the following characteristics:
[0198] (1) Shared descriptor buffer pool: All channels share the same SRAM buffer prefetch buffer instead of statically allocating an independent buffer for each channel.
[0199] (2) Resource partitioning mechanism based on FIFO + per-channel buffer depth: Each channel corresponds to a FIFO segment in the shared SRAM buffer, and its buffer depth and the start and end range of FIFO read / write pointer are set by the configuration register.
[0200] (3) Prefetch triggering mechanism: When the available remaining space (or number of free descriptors) in the shared buffer is lower than the set low-watermark, a prefetch operation for the next set of descriptors is triggered to supplement descriptor resources, dynamically maintaining a constant number of usable descriptors in each channel, thus solving the performance problem of small packet transmission. Adjusting the high watermark prevents excessive prefetching, which can lead to bus resource occupation and descriptor discarding.
[0201] (4) Dynamically adjustable programmable buffer depth and resource allocation: Supports flexible adjustment of the buffer depth of each channel based on channel load, data packet characteristics (large / small / frequent / intermittent), channel priority, etc., so as to balance performance requirements and resource conservation.
[0202] (5) Balancing high throughput with hardware resource efficiency: It can meet the high requirements of prefetch buffer in high-frequency small packet transmission, and also prevent buffer resource waste in large packet and low-frequency scenarios, thus balancing throughput performance and resource cost (SRAM area, memory bandwidth, power consumption, etc.).
[0203] In summary, traditional multi-channel descriptor prefetching DMA controller designs typically allocate a separate descriptor buffer or prefetch cache space for each channel to store multiple prefetched descriptors. This approach wastes significant hardware resources when there are many channels or when channels are idle or under low load for extended periods. Furthermore, for scenarios with uneven loads and diverse task characteristics, its cache resource allocation lacks flexibility, making it difficult to balance resource utilization and performance efficiency.
[0204] This example effectively solves the aforementioned bottleneck by sharing a single descriptor buffer pool across all channels and combining a per-channel FIFO, programmable buffer depth, and watermark-triggered prefetching mechanism, bringing the following key advantages:
[0205] (1) Reduce hardware resource overhead: Since each channel is no longer statically allocated a dedicated buffer, but instead all channels share the same SRAM buffer pool, the total buffer capacity requirement and SRAM area are significantly reduced, thereby reducing chip resource overhead and cost.
[0206] (2) Adapt to multi-task, multi-channel and mixed load scenarios and improve resource utilization efficiency: buffer-depth can be configured on demand for each channel (based on channel activity, transmission frequency, data packet characteristics, etc.), making buffer allocation and resource use more flexible. Idle channels reduce resource occupation, while busy channels and small packet high-frequency channels can obtain sufficient buffer resources.
[0207] (3) Optimize the performance of small packet, high frequency and high latency transmission: With the help of watermark trigger prefetch + buffer-pool + FIFO, prefetching will automatically occur when the number of buffer descriptors is lower than the threshold, so that the descriptor supply is continuous and uninterrupted, thereby reducing the transmission pause or transmission delay caused by descriptor retrieval, and improving the stable throughput and response performance of small packets and high frequency tasks.
[0208] (4) Avoid resource waste and bandwidth and power consumption waste: When there are large data packets, low frequency and continuous transmission tasks, it is not necessary to reserve too deep buffers for each channel or prefetch a large number of descriptors for a long time, thereby reducing useless buffer occupation, prefetch overhead, bus bandwidth occupation and power consumption.
[0209] (5) The system has stronger scalability and is suitable for high-concurrency systems with a large number of channels: When the number of DMA channels increases (multiple peripherals, multiple tasks, multiple data streams), the shared buffer pool mechanism avoids the linear growth of buffer resources, the hardware resource overhead grows slowly, the system is easier to expand and the cost control is better.
[0210] In summary, the technical solution presented in this example balances performance and resource efficiency, and has significant advantages in complex systems with multiple channels, multiple tasks, and multiple load types (such as SoCs, multiple peripherals / I / O, heterogeneous accelerator platforms, etc.).
[0211] This disclosure also provides a data transmission device. For example... Figure 10 As shown, this device is used in a direct memory access (DMA) controller, and the device includes:
[0212] Interface conversion module 101 is used to receive access requests;
[0213] Multiple DMA channels 102 are connected to the interface conversion module; each DMA channel is used to obtain the corresponding channel descriptor through the interface conversion module, and when the channel descriptor is a preset descriptor in the descriptor chain, it prefetches multiple descriptors in the descriptor chain according to the channel prefetching strategy configured for the DMA channel.
[0214] Shared buffer 103 is connected to multiple DMA channels and is used to cache the descriptor corresponding to each DMA channel; wherein, the shared buffer contains multiple channel regions that are logically isolated from each other and correspond to multiple DMA channels respectively;
[0215] The data transmission module 104 is connected to multiple DMA channels and a shared buffer, and is used to perform data transmission operations based on multiple descriptors cached in the channel area of any DMA channel when a data transmission request is received from any DMA channel.
[0216] The data transmission module can also be used through Figure 2 The data receiving and sending modules in the present invention are implemented. Therefore, the data transmission device in this disclosure can be integrated into the Direct Memory Access (DMA) controller, and correspondingly, the functions performed by the data transmission device can also be implemented through... Figure 2The DMAC overall hardware architecture shown employs multiple hardware modules working collaboratively. Alternatively, the data transfer device can also be the Direct Memory Access (DMA) controller itself; this disclosure does not limit the specific implementation of the data transfer device.
[0217] Optionally, the data transmission device may also include at least one of the following modules:
[0218] The register configuration module, connected to multiple DMA channels, is used to configure the corresponding channel prefetch strategy for each DMA channel.
[0219] The channel arbitration module, connected between multiple DMA channels and the data transmission module, is used to arbitrate multiple DMA channels to authorize a single DMA channel to send a data transmission request to the data transmission module based on the arbitration result.
[0220] The access request can be any type of request sent to the DMA controller via the system bus, including but not limited to configuration requests to read or write the DMA controller's configuration registers, descriptor read requests to obtain descriptors, and data transfer requests to indicate the start of data migration. A data transfer request is an internal signal or command issued by the DMA channel to the data transfer module after obtaining a valid descriptor from the channel area, requesting the execution of the actual data migration operation defined by that descriptor.
[0221] The specific working principles and implementation details of each of the above modules can be found in the descriptions of the corresponding parts in the method embodiments, and will not be repeated here.
[0222] This disclosure also provides an electronic device, including:
[0223] At least one processor; and
[0224] Memory that is communicatively connected to at least one processor;
[0225] As described above, the Direct Memory Access (DMA) controller communicates with at least one processor and memory.
[0226] The processor is configured to initiate a data transfer request or configuration operation to the DMA controller so that the DMA controller executes the above-described method.
[0227] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0228] In addition, this disclosure also provides a request processing device, an electronic device, and a computer-readable storage medium, all of which can be used to implement any of the graphics processing methods provided in this disclosure. The corresponding technical solutions and descriptions are described in the corresponding section of the method and will not be repeated here.
[0229] Figure 11 This is a block diagram of an electronic device provided in an embodiment of the present disclosure.
[0230] Reference Figure 11 This disclosure provides an electronic device, which includes: at least one processor 701; at least one memory 702; and a DMA controller 703.
[0231] The DMA controller 703 is communicatively connected to at least one processor 701 and at least one memory 702. The memory 702 stores a computer program executable by at least one processor 701. When executed by at least one processor 701, the computer program configures and controls the DMA controller so that the DMA controller can perform the methods of any of the above embodiments. In some embodiments, the DMA controller may be integrated within the processor 701. In other embodiments, the DMA controller may be independent of the hardware module of the processor 701.
[0232] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0233] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable program instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0234] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0235] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0236] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0237] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0238] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.
Claims
1. A data transmission method, characterized in that, The method is applied to a direct memory access (DMA) controller, and the method includes: Obtain the channel descriptor corresponding to any DMA channel in the DMA controller; When the channel descriptor is a preset descriptor in the descriptor chain, multiple descriptors in the descriptor chain are prefetched according to the channel prefetching strategy corresponding to any DMA channel; The prefetched descriptors are cached in a shared buffer corresponding to the channel region of any of the DMA channels; wherein the shared buffer contains multiple logically isolated channel regions, each corresponding to a multiple DMA channel; Control any of the DMA channels to perform data transfer operations based on multiple descriptors cached in the channel region.
2. The method according to claim 1, characterized in that, The step of prefetching multiple descriptors in the descriptor chain according to the channel prefetching strategy corresponding to any DMA channel includes: Retrieve the address navigation field from the channel descriptor; Based on the address navigation field in the channel descriptor and the channel prefetching strategy corresponding to any DMA channel, prefetch multiple descriptors in the descriptor chain; The first field value in the address navigation field of the channel descriptor is used to characterize the starting address information of multiple descriptors in the descriptor chain.
3. The method according to claim 2, characterized in that, The channel prefetching strategy corresponding to any DMA channel includes: a water level control strategy, which is used to control the number of available descriptors cached in the channel region of any DMA channel; The step of prefetching multiple descriptors in the descriptor chain based on the address navigation field in the channel descriptor and the channel prefetching strategy corresponding to any DMA channel includes: The total number of descriptors in the descriptor chain is determined based on the value of the second field in the address navigation field of the channel descriptor; Based on the total number of descriptors in the descriptor chain and the channel water level threshold in the water level control strategy, multiple descriptors in the descriptor chain are prefetched.
4. The method according to claim 3, characterized in that, The channel water level thresholds in the water level control strategy include: a first channel water level threshold and a second channel water level threshold, wherein the first channel water level threshold is less than the second channel water level threshold. The step of pre-fetching multiple descriptors from the descriptor chain based on the total number of descriptors in the descriptor chain and the channel water level threshold in the water level control strategy includes: If the first comparison result between the number of available descriptors cached in the channel region of any DMA channel and the first channel water level threshold is a first preset result, then multiple descriptors in the descriptor chain are prefetched according to the second comparison result between the total number of the multiple descriptors and the second channel water level threshold.
5. The method according to claim 4, characterized in that, The step of prefetching multiple descriptors in the descriptor chain based on a second comparison result between the total number of the multiple descriptors and the second channel water level threshold includes: Based on the second comparison result between the total number of the multiple descriptors and the second channel water level threshold, the prefetch batch of the multiple descriptors in the descriptor chain and the batch prefetch quantity in each prefetch batch are determined. Based on the prefetch batches of the multiple descriptors and the batch prefetch quantity in each prefetch batch, the multiple descriptors in the descriptor chain are prefetched in batches. Wherein, the sum of the number of descriptors prefetched in each batch and the number of available descriptors cached in the channel region of any DMA channel is not greater than the second channel water level threshold.
6. The method according to claim 5, characterized in that, The multiple DMA channels correspond to multiple prefetch address registers, and each prefetch address register is used to record the address information of the descriptor to be prefetched; The batch prefetching of multiple descriptors in the descriptor chain includes: After the descriptor prefetching for the current batch is completed, update the value of the prefetch address register corresponding to the DMA channel; Prefetch the descriptors for the next batch based on the updated value of the prefetch address register.
7. The method according to claim 5, characterized in that, The difference between the water level threshold of the second channel and the water level threshold of the first channel is greater than the water level threshold of the first channel; Furthermore, the first channel water level threshold and the second channel water level threshold corresponding to each DMA channel are dynamically adjusted according to the channel load, data transmission frequency, and transmission task type of the DMA channel.
8. The method according to any one of claims 1-7, characterized in that, The multiple DMA channels in the DMA controller are virtual DMA channels. The area size of the channel region corresponding to each virtual DMA channel in the shared buffer is dynamically adjusted according to the channel load, data transmission frequency, and transmission task type of the virtual DMA channel.
9. The method according to claim 8, characterized in that, The channel region of any DMA channel is identified by the channel pointer corresponding to any DMA channel; Each virtual DMA channel corresponds to a set of channel pointers, including a read pointer pointing to the next descriptor location to be read and a write pointer pointing to the next descriptor location to be written; and the movement range of the read pointer and the write pointer is determined according to the area address range of the channel region of the virtual DMA channel.
10. The method according to any one of claims 2-7, characterized in that, The preset descriptor in the descriptor chain includes: the first descriptor in the descriptor chain; Furthermore, if the descriptor chain further includes multiple sub-chains, the preset descriptor in the descriptor chain also includes: The last descriptor in the first subchain; wherein the address navigation field in the last descriptor in the first subchain is used to indicate the total number of descriptors in the second subchain and the starting position; wherein the second subchain is the next subchain after the first subchain.
11. The method according to claim 10, characterized in that, When the first subchain is the last subchain among the plurality of subchains, the address navigation field in the last descriptor of the first subchain is configured to point to the starting address of the first subchain among the plurality of subchains.
12. The method according to any one of claims 1-7, characterized in that, After obtaining the channel descriptor corresponding to any DMA channel in the DMA controller, the method further includes: Obtain the preset mode field in the channel descriptor. If the channel descriptor belongs to the chained mode based on the field value of the preset mode field, determine whether the channel descriptor is a preset descriptor in the descriptor chain based on the position of the channel descriptor in the descriptor chain.
13. A data transmission device, characterized in that, The apparatus is used in a direct memory access (DMA) controller, and the apparatus includes: The interface conversion module is used to receive access requests; Multiple DMA channels are connected to the interface conversion module; wherein each DMA channel is used to obtain a corresponding channel descriptor through the interface conversion module, and when the channel descriptor is a preset descriptor in the descriptor chain, it prefetches multiple descriptors in the descriptor chain according to the channel prefetching strategy configured for the DMA channel; A shared buffer, connected to the plurality of DMA channels, is used to cache the descriptor corresponding to each DMA channel; wherein, the shared buffer contains multiple channel regions that are logically isolated from each other and correspond to the plurality of DMA channels respectively; The data transmission module, connected to the plurality of DMA channels and the shared buffer, is used to perform data transmission operations based on a plurality of descriptors cached in the channel area of any DMA channel when a data transmission request is received from any DMA channel.
14. The apparatus according to claim 13, characterized in that, The device further includes at least one of the following modules: A register configuration module, connected to the multiple DMA channels, is used to configure a corresponding channel prefetch strategy for each DMA channel. A channel arbitration module, connected between the plurality of DMA channels and the data transmission module, is used to arbitrate the plurality of DMA channels to authorize a single DMA channel to send a data transmission request to the data transmission module based on the arbitration result.
15. An electronic device, characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; The direct memory access (DMA) controller as described in any one of claims 1 to 12, wherein the DMA controller is communicatively connected to the at least one processor and the memory; The processor is configured to initiate a data transfer request or configuration operation to the DMA controller, so that the DMA controller performs the method of any one of claims 1 to 12.