A multi-level multi-priority dma arbitration circuit and arbitration method

CN122654049BActive Publication Date: 2026-09-29SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611160158.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-08-03
Publication Date
2026-09-29
Estimated Expiration
2046-08-03

AI Technical Summary

Technical Problem

由于总线读事务(需等待数据返回)与写事务(直接推流)的响应周期不对等,读事务的延迟容易对后续写事务产生反压(Backpressure),造成总线带宽利用率低下

Benefits of technology

本发明提出的一种多优先级多级DMA仲裁架构,通过将调度层级在逻辑和物理上进行正交解耦,有效克服了传统仲裁机制的瓶颈。本架构通过引入配置状态寄存器(CSR)接口实现了硬件支持、软件触发的“优先级反转”机制,赋予了系统极强的自适应流量整形能力,不仅解决了低优先级通道因固定优先级导致的“数据饥饿”问题,还坚决捍卫了在恶劣网络环境下接收链路的可用性和实时性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122654049B_ABST
    Figure CN122654049B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of integrated circuit design, and provides a multi-level multi-priority DMA arbitration circuit and arbitration method, which is decoupled into three orthogonal scheduling levels on the logic and physical levels, including: a first-level arbitration layer, a channel-level arbitrator for receiving concurrent requests of multiple independent DMA channels and screening out a unique authorized channel; a channel-level multiplexer for performing data and attribute gating of the authorized channel; a second-level arbitration layer, a function-level arbitrator for distinguishing request types, including descriptor requests and data requests, and dynamically adjusting arbitration priorities between different requests through a priority inversion signal provided by an internal state machine and an external configuration state register interface; a third-level arbitration layer connected with the output of the second-level arbitration layer, for mapping internal arbitration results to an external system bus, maximizing bus bandwidth utilization through read-write separation and response demultiplexing mechanism. The bottleneck of the traditional arbitration mechanism is effectively overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of integrated circuit design technology, and particularly relates to a multi-level, multi-priority DMA arbitration circuit and arbitration method. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] In the design of storage controller chips, a common approach is to employ a multi-channel design. Multiple peripherals on different channels simultaneously access shared resources within the chip. Therefore, an arbitrator is needed to arbitrate requests from each channel, ensuring that only one channel can access the chip at a time. When multiple channels simultaneously receive write requests from multiple requesting modules, there are typically two approaches: one is to input each request into the arbitrator for arbitration, and the other is to first arbitrate requests within each channel, followed by second-level arbitration for requests between channels. However, existing arbitration methods require arbitrating a large number of requests, leading to low efficiency in generating arbitration results and thus impacting data write speed.

[0004] Existing DMA controller arbitration logic mainly adopts the following typical schemes: Fixed priority arbitration, which is the simplest arbitration method, assigning a fixed priority level to each DMA channel. When multiple channels issue transmission requests simultaneously, the arbitrator always selects the channel with the highest priority for service. However, in existing DMA patents using pure hardware fixed priority arbitration, when a high-priority channel has a sudden surge in large or continuous data transmission demands, it will monopolize bus control for a long time. This causes the requests of low-priority channels to be suspended for a long time, unable to obtain service, which in turn leads to FIFO overflow or data loss of the corresponding peripherals. Polling arbitration, where the arbitrator checks the requests of each channel in a fixed order and provides service to each channel in a round-robin manner. The polling mechanism solves the "starvation" problem and ensures absolute fairness, but it treats all channels equally. For receive requests or control signal requests that are extremely sensitive to latency, such as in Ethernet MAC, the polling mechanism cannot provide a fast response, resulting in a decline in the service quality of real-time services. Hybrid arbitration mechanism, which combines the advantages of fixed priority and polling arbitration, groups channels, using fixed priority between groups and polling within groups. This method balances efficiency and fairness to a certain extent. Arbitration based on weighted fair queues assigns weights to each channel and allocates bandwidth according to the weight ratio.

[0005] The above scheme only schedules at the "channel" level, without distinguishing whether the data transmitted within the channel is "control information" (such as a transport descriptor) or "payload data". Furthermore, some existing architectures mix read and write requests on the same arbitration node before connecting to the system bus (such as AXI). Because the response cycles of bus read transactions (which require waiting for data return) and write transactions (which push data directly) are unequal, the latency of read transactions can easily create backpressure on subsequent write transactions, resulting in low bus bandwidth utilization. Summary of the Invention

[0006] To address at least one of the technical problems existing in the background art described above, the present invention provides a multi-level, multi-priority DMA arbitration circuit and arbitration method. To achieve the above objective, the present invention adopts the following technical solution: A first aspect of the present invention provides a multi-level, multi-priority DMA arbitration circuit, which is decoupled into three orthogonal scheduling levels at the logical and physical levels, including: The first-level arbitration layer includes a channel-level arbitrator and a channel-level multiplexer. The channel-level arbitrator is used to receive concurrent requests from multiple independent DMA channels and filter out the uniquely authorized channel. The channel-level multiplexer is used to perform data and attribute gating of the authorized channel. The second-level arbitration layer is connected to the output of the first-level arbitration layer and includes a functional-level arbitrator and a master arbitrator. The functional-level arbitrator is used to distinguish the request type, including descriptor requests and data requests, and dynamically adjusts the arbitration priority between different requests through the priority inversion signal provided by the internal state machine and the external configuration state register interface. The third-level arbitration layer, connected to the output of the second-level arbitration layer, is used to map the internal arbitration results to the external system bus, maximizing bus bandwidth utilization through read-write separation and response demultiplexing mechanisms.

[0007] Furthermore, the channel-level arbitrator in the first-level arbitration layer is a priority encoder with an enable terminal implemented using a deeply cascaded pure combinational logic network. It solidifies the channel priority order from high to low. When multiple channels pull up the request level in the same clock cycle, all signals lower than the current highest request bit are shielded by the logic gate array, and the authorization vector of a single hotspot is output.

[0008] Furthermore, a high-dimensional AND-OR logic tree is constructed inside the channel-level multiplexer in the first-level arbitration layer. Using the single-hotspot authorization signal as a mask, all unauthorized channel data is set to zero. Then, through bitwise OR operation, the data of the only selected channel is aggregated onto the one-dimensional output bus.

[0009] Furthermore, the channel-level multiplexer in the first-level arbitration layer also includes a pipelined register slice, which is used to pass authorized data requests through a first-level D flip-flop for sizing before sending them to the second-level arbitration layer, in order to eliminate combinational logic glitches and strictly limit the timing critical path within the first-level arbitration layer.

[0010] Furthermore, the second-level arbitration layer also includes two request generation engines: a descriptor engine and a data path engine. The descriptor engine is used to prefetch descriptor streams from system memory and parse out the physical storage location, packet length, and various offload flags of data packets. These requests are uniformly classified as descriptor requests by the system. The data path engine is used to move large amounts of Ethernet frame payload data, which are classified as data requests.

[0011] Furthermore, when priority replacement is effective, the state machine of the functional arbitrator forcibly places the decision threshold for data requests before descriptor requests during state transition evaluation.

[0012] Furthermore, the main arbiter is equipped with sub-arbiters that are orthogonal to each other for read and write channels. When the receive priority configuration input signal is enabled, all write memory requests from the receive data path have a higher preemption right than the send data path. When the read priority enable signal is enabled, memory requests to read the TSO header template are set as privileged operations, allowing them to seamlessly jump the queue between normal descriptor polling.

[0013] Furthermore, the third-level arbitration layer includes independent read crossbar switches, write crossbar switches, and response demultiplexers. The read crossbar switches and write crossbar switches are two completely independent arbitrators, which are used to map internal read requests to the read address channel of the system bus and to map internal write requests to the write address and write data channel of the system bus, respectively, so as to realize concurrent utilization of bidirectional bandwidth for reading and writing. The response demultiplexer is used to: when a request is sent to the system bus via a crossbar switch, tag each transaction with a transaction tag containing the identity information of the request source; when the system bus returns data or a completion status, parse the transaction tag and route the returned payload to the corresponding request source through an internal demultiplexing matrix.

[0014] Furthermore, the third-level arbitration layer also includes an asynchronous clock domain adaptation module, which includes a first asynchronous FIFO and a second asynchronous FIFO, serving read data and write data respectively. Internally, it adopts a Gray code pointer FIFO structure and has built-in data shifter and byte mask generation logic to splice the internal data stream into burst data blocks that conform to the bus protocol.

[0015] A second aspect of the present invention provides a multi-level, multi-priority DMA arbitration method, based on a multi-level, multi-priority DMA arbitration circuit as described in the first aspect, comprising: Receive concurrent requests from multiple independent DMA channels and select the only authorized channel; Perform data and attribute selection for the authorized channel; The system distinguishes between request types, including descriptor requests and data requests, and dynamically adjusts the arbitration priority between different requests through priority inversion signals provided by the internal state machine and the external configuration state register interface. The internal arbitration result is mapped to the external system bus, maximizing bus bandwidth utilization through read / write separation and response demultiplexing mechanisms.

[0016] Compared with the prior art, the beneficial effects of the present invention are: This invention proposes a multi-priority, multi-level DMA arbitration architecture that effectively overcomes the bottlenecks of traditional arbitration mechanisms by orthogonally decoupling the scheduling levels both logically and physically. This architecture implements a hardware-supported, software-triggered "priority inversion" mechanism through the introduction of a Configuration Status Register (CSR) interface, endowing the system with extremely strong adaptive traffic shaping capabilities. This not only solves the "data starvation" problem caused by fixed priorities in low-priority channels but also resolutely safeguards the availability and real-time performance of the receiving link in harsh network environments.

[0017] The architecture of this invention completely separates the read crossbar switch from the write crossbar switch, realizing complete physical orthogonality and asynchronization of read and write operations. This allows the write channel to still operate at full capacity during the latency period when the read channel is waiting for data to return, thereby achieving 100% bidirectional saturation utilization of the bus bandwidth.

[0018] At the underlying hardware implementation level, the L1 layer effectively filters out glitches generated by combinational logic through a pipelined register slicing mechanism and strictly confines timing-critical paths within the L1 layer, laying a solid physical foundation for increasing the operating frequency of the entire DMA controller. Furthermore, this design achieves precise hardware acceleration for complex protocols such as TCP Segment Offload (TSO), completely eliminating bus addressing latency during giant frame segmentation through seamless queueing of privileged-level operations. Simultaneously, utilizing an asynchronous FIFO structure based on Gray code pointers and a response demultiplexer, the system can safely overcome the metastability risks introduced by the clock domain, automatically concatenating fragmented data streams into optimal burst blocks conforming to external bus specifications, and accurately routing out-of-order loads back to the correct source module.

[0019] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0020] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0021] Figure 1 This is a diagram of a multi-level, multi-priority DMA arbitration circuit architecture provided in an embodiment of the present invention; Figure 2 This is a waveform diagram of the DMA interface write data function test provided in an embodiment of the present invention; Figure 3 This is a waveform diagram of the DMA interface write descriptor function test provided in an embodiment of the present invention; Figure 4 This is a waveform diagram of the DMA interface read data function test provided in an embodiment of the present invention; Figure 5 This is a waveform diagram of the DMA interface read descriptor function test provided in an embodiment of the present invention. Detailed Implementation

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0023] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0024] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0025] This invention proposes a multi-priority DMA arbitration architecture, primarily addressing the inherent contradiction between high real-time control flow and data flow fairness in traditional arbitration mechanisms. To address the issues of data starvation caused by existing fixed priorities and the inability of polling mechanisms to guarantee rapid response to control signals, this architecture independently distinguishes between control information and payload paths at the L2 functional level scheduling layer. By introducing a configuration status register interface, this design implements a dynamic priority inversion mechanism. When the system detects an abnormality in a specific buffer level, the threshold for data request can be forcibly increased via software triggering. This achieves adaptive flow shaping while ensuring low-latency response to critical control signaling, preventing peripheral FIFO overflow caused by long-term suspension of low-priority channels. At the system architecture coordination and bus transmission level, to address the backpressure defects of the system bus caused by traditional read-write hybrid scheduling, this architecture constructs completely independent read crossbar switches and write crossbar switches at the L3 layer. This mechanism enables asynchronous read and write operations, allowing for bidirectional saturation utilization of bus bandwidth. Simultaneously, the architecture introduces a privileged level mechanism for TCP segmentation offloading, eliminating bus addressing latency during frame header reading.

[0026] like Figure 1 As shown, as an embodiment of the present invention, this embodiment provides a multi-level, multi-priority DMA arbitration circuit. The circuit architecture is clearly decoupled into three orthogonal scheduling levels at the logical and physical levels, including: The first-level arbitration layer, L1, is located at the very front of the entire architecture, directly facing multiple concurrent channels of the underlying hardware (such as...). Figure 1 The 16-channel request (CH0-CH15) marked in the image has the core mission of filtering out the only legitimate visitor from a massive number of concurrent requests within the same clock cycle with the lowest possible gate delay.

[0027] The BARB_PRI (Base Arbiter Priority) module is the physical arbitration front of the L1 layer, receiving request signals from the request bus of 16 independent DMA channels. Internally, this module is implemented using a deeply cascaded pure combinational logic network. After logic synthesis, it is mapped to a high-speed priority encoder with an enable pin.

[0028] The design imposes a strict priority order on the physical circuitry: channel 15 (CH15) is given the highest priority, while channel 0 (CH0) is given the lowest priority.

[0029] When multiple channels pull the request level high in the same clock cycle, the logic gate array inside BARB_PRI will instantly shield all signals below the current highest request bit, and finally output a set of standard one-hot grant vectors to the next stage.

[0030] For example, if CH15 and CH0 are requested simultaneously, the output vector will only be 1 in the 15th bit, with the rest being 0. This single-hotspot encoding provides extremely high hardware execution efficiency for subsequent multiplexer control.

[0031] After obtaining the single hotspot channel authorization output by the BARB_PRI module, the BARB_MUX (Base ArbiterMultiplexer) module is responsible for performing the actual data and attribute gating.

[0032] Because each channel contains not only a single request enable signal, but also a wide-bit destination address, burst count, and cache control attributes, BARB_MUX internally constructs a high-dimensional AND-OR tree. It uses a single hotspot grant signal as a mask to zero out all ungranted channel data, and then uses a bitwise OR operation to aggregate the data of the uniquely selected channel onto a one-dimensional output bus. In high-speed clock domains, the combinational logic multiplexing of a wide bus is highly susceptible to timing violations.

[0033] To ensure smooth routing of the physical backend, BARB_MUX provides a pipelined register slicing option. Authorized data requests (selected channel requests) undergo timing shaping via a single-stage D-flip-flop before being sent to the L2 layer. The critical path is strictly confined to the L1 layer. This mechanism also effectively filters out combinational logic glitches generated by the priority encoder, laying the foundation for increased frequency across the entire DMA controller.

[0034] The second-level arbitration layer, L2, is the core "brain" of the architecture. It is responsible for conducting intelligent game based on state machines between the control flow (descriptors) and the data flow (payload data), and introduces the Configuration Status Register (CSR) to realize the intervention and reversal of dynamic priorities.

[0035] Layer L1 solves the physical problem of "who leaves first," while Layer L2 (functional arbitration and scheduling) is responsible for solving the business logic problem of "who is more important." For example... Figure 1 As shown, the L2 layer contains a complex request generation engine (Path A and Path B) and two major arbitration hubs responsible for overall coordination: the Function Arbiter (FARB) and the Memory Arbiter (MARB).

[0036] In advanced network applications, DMA data movement is not simply a contiguous copy of memory, but rather a highly structured linked list or circular queue operation.

[0037] Path A (Descriptor Engine): This part consists of a pipeline of submodules such as TDT_core (transmit descriptor core), TDT_ctrl (control module), tdt_data (data parsing), and tDT_rreq (read request generation). It is used to prefetch the descriptor stream from system memory, parsing the physical storage location of the data packets, packet length, and various offload flags. Since descriptors directly determine the subsequent behavior of DMA, these requests are uniformly classified as Type 0 (descriptor requests) by the system and enjoy a higher basic priority than ordinary data by default.

[0038] The first-level pipeline (address and pointer management) TDT_core (send descriptor core) module internally maintains the base address, head pointer, and tail pointer of the descriptor circular queue. When a new descriptor is detected to be waiting to be processed, TDT_core calculates the physical absolute address of the next target descriptor and pushes it into the next level; The second-level pipeline (request generation and distribution) tDT_rreq (read request generation) module receives the physical address from TDT_core. This module is responsible for translating the internal address into a standard memory read request and routing it to the L2 layer MARB, which then initiates the read operation to the external bus.

[0039] The third-level pipeline (data decoding and parsing) is responsible for hardware-level decoding of the received bit stream when the external system bus returns the raw data stream of the descriptor. It extracts the source address, packet length, and specific protocol offload flags (such as TSO enable bit, checksum and insertion control bits) of the network packet payload pointed to by the descriptor.

[0040] The fourth-level pipeline (control scheduling and delivery), TDT_ctrl (control module), serves as the pipeline coordination unit for Path A, receiving key information parsed from tdt_data. On one hand, it distributes payload transmission tasks to the data path; on the other hand, it packages the current descriptor's state update request into a descriptor request and sends it to the FARB functional level arbitrator in the L2 layer, using the aforementioned priority state machine for high-priority scheduling processing.

[0041] Path B (Data Path): This includes RDF (Receive Data FIFO), TDF (Transmit Data FIFO), and MARB_TSO (TCP Segment Offload Memory Interface) for acceleration specific protocol stacks. Requests generated in this part are designed to move large Ethernet frame payloads and are classified as Type 1 (Data Request).

[0042] Specifically, in Path B (data path), the RDF (receive data FIFO), TDF (transmit data FIFO), and the specific protocol acceleration module MARB_TSO generate data transfer requests in physical parallel, and complete dynamic coordination and resource arbitration interaction through the core hub of L2 layer MARB (memory arbitrator). Specifically, RDF actively generates write memory requests based on the accumulated data level at the receiving end, TDF generates continuous payload read memory requests based on descriptor sending instructions, and MARB_TSO independently calculates slices and generates privileged header read requests for large TCP packets. After these three concurrent requests are unified and converged to the MARB state machine, MARB performs closed-loop flow control feedback by sampling the empty and full states of the FIFO in real time, and responds to dynamic configuration signals when resource contention occurs, granting RDF write requests priority over send requests to prevent data packet loss at the receiving end. At the same time, it allows privileged header read requests of MARB_TSO to seamlessly jump in during the normal multi-shot payload read cycle of TDF. Finally, read and write requests that are authorized through unified arbitration are completely stripped and independently sent to the L3 layer for non-blocking bus transmission.

[0043] The FARB module is the core state machine network coordinating Type 0 and Type 1 requests. Internally, FARB maintains a master state machine called farb_fsm. This state machine includes states such as IDLE (idle), REQ_TYPE0 (processing descriptor requests), REQ_TYPE1 (processing data requests), and the corresponding WAIT (waiting for bus authorization). The design of this state machine strictly adheres to the latency characteristics of the bus handshake protocol, effectively absorbing the time difference in bus authorization return and avoiding invalid overlapping requests. Traditional DMA absolutely fixes descriptor priorities; if a sudden surge of massive data in the send queue occurs, continuous reading of descriptors may completely block the data packet transmission path. Figure 1 As shown, FARB provides an external configuration status register (CSR) interface, priority_alt Input. When the software detects an abnormality in a specific buffer level, this signal can be raised by writing to the register. Once priority_alt is effective, farb_fsm will forcibly place the judgment threshold for Type 1 (data request) before Type 0 during state transition evaluation. This hardware-supported, software-triggered "priority inversion" mechanism gives the system extremely strong adaptive traffic shaping capabilities.

[0044] Unlike FARB, which handles high-level logic, the MARB (M-Arbiter) module is the final convergence point and allocation center for all substantive read / write memory requests (RDF Req, TDF Req, TDT Req, etc.). MARB deploys sub-arbiters (marb_ch) that are orthogonal to the read / write channels. It not only handles concurrent access conflicts but also monitors the full / empty status of the underlying buffers, generating precise flow control feedback to control the request issuance rate. MARB introduces the receive priority configuration signal priority_rx CSR. When enabled, all write memory requests from the receive data path (RDF) gain preemption rights over transmit data (TDF). This mechanism ensures that the availability and real-time performance of the receive link are resolutely protected even in harsh network environments. TCP Segmentation Offload (TSO) technology allows the operating system to deliver large TCP packets, up to 64KB, to MAC hardware slices for transmission in a single pass. In this mode, the DMA hardware must repeatedly read the same TCP header and concatenate it with different slice payloads. The MARB receive read priority enable signal `priority_tso_rd` is designed for this purpose. When enabled, the MARB's internal state machine will set memory requests for reading TSO section templates as privileged operations, allowing them to seamlessly interleave between ordinary descriptor polling. This highly targeted microarchitectural design completely eliminates bus addressing latency during the giant frame segmentation process.

[0045] After layers of filtering by L1 physical filtering and L2 service scheduling, the "authorized M-Req (read / write)" and "authorized F-Req" ​​that finally break through need to complete the interface with the external high-speed system bus (such as the AXI interconnect matrix) at the L3 layer. The design concept of the L3 layer is: full decoupling, asynchronous and non-blocking transmission.

[0046] The third-level arbitration layer (L3) is responsible for seamlessly and non-blockingly mapping the internal scheduling results to the external high-speed system bus (such as AXI4 / AXI5). Through read-write separation and response demultiplexing mechanisms, it maximizes bus bandwidth utilization.

[0047] Layer 3 decouples two completely independent cross-arbiters: the read cross-connector (XARB_RD) and the write cross-connector (XARB_WR). XARB_RD focuses on mapping internal read requests to the bus's read address channels (such as the AXI AR channel) and is responsible for negotiating burst length and 4KB boundary segmentation. The XARB_RD mechanism independently interfaces with the bus's write address and write data channels (such as the AXI AW / W channel). Because read and write operations are completely physically orthogonal, XARB_WR can still operate at full capacity while the system is waiting for descriptor read data to return, continuously pushing network packet payloads from the internal receive FIFO into system memory, achieving 100% bidirectional saturation utilization of the bus bandwidth.

[0048] Layer L3 incorporates XARB_RSP_DMUX (Response Demultiplexer). When the DMA sends a request to the system bus via a crossbar switch, the hardware automatically assigns a unique transaction ID to each transaction. This ID embeds the identity information of the request source (e.g., whether it comes from the TDT descriptor engine or the TDF data path). When the system bus returns a data packet or a completion status (Response ID) through the response processing path, XARB_RSP_DMUX demultiplexes and compares these IDs in real time. Through an internal high-speed demultiplexing matrix (the DMUX funnel-shaped logic block in the diagram), out-of-order payloads are precisely routed back to the correct source—if it's descriptor data, it's pushed to the "descriptor engine write-back" path; if it's network payload, it's pushed into the corresponding FIFO buffer.

[0049] The far right of the architecture diagram shows the last line of defense at Layer 3: the XDF and XDW FIFO buffers and data alignment modules. The internal application clock (Core Clock) and external system bus clock (Bus Clock) of DMA often reside in different frequency domains, or are even completely asynchronous clock sources. The XDF (primarily for reading data) and XDW (primarily for writing data) internally employ an asynchronous FIFO structure based on Gray Code pointers to safely absorb and overcome the metastability risks introduced by clock domains. The XDF and XDW modules incorporate a Barrel Shifter and Byte Mask generation logic, automatically concatenating fragmented internal data streams into optimal burst chunks conforming to the AXI bus specification. Watermark Flow Control.

[0050] The designed DMA module passed module verification analysis. The specific functional verification results are shown below: Figure 2 For the process of writing a set of data to system memory, DMA starts a transfer transaction by pulling the mdc_start_xfer_o signal high. The address bus mdc_addr_o then transitions from the idle value 0000e6f8 to 0000e6fc and increments by four bytes to 0000e714. During the transfer, the burst length mdc_burst_count_o remains 07, and the read / write indicator mdc_rd_wrn_o remains low to execute the write operation. After the master write data valid signal mdc_wdata_val_o goes high, the slave ready signal mdc_wdata_rdy_i responds by going high, and the data bus mdc_wdata_o synchronously outputs the corresponding data 00000607, 0006b9c3, 00029300 to 00027170 sequentially. After the data transmission is complete, the handshake signal goes low, and mdc_xfer_done_i returns a high-level pulse to indicate the end of the transfer process.

[0051] Figure 3To initiate the transfer process of writing a set of descriptors into system memory, the transfer start signal rndc_start_xfer_o is pulled high for one clock cycle to trigger the operation. The address bus rndc_addr_o then jumps from 0000e468 to 000041dc, and the burst count rndc_burst_count_o synchronously changes from 07 to 00. The read / write control rndc_rd_wrn_o remains low to indicate write mode. The write data bus rndc_wdata_o loads the value 300000a0. The write data valid signal rndc_wdata_val_o is pulled high and held in wait. During this period, the fixed burst rndc_fixed_burst_o and the packet end signal rndc_eop_o remain low. The slave write ready signal rndc_wdata_rdy_i is pulled high after several cycles to respond to the handshake. The transmission complete signal rndc_xfer_done_i is pulled high synchronously. The write valid signal rndc_wdata_val_o is then pulled low to end the current transmission. The address bus is updated to the next address 000041e0, and the write data bus becomes 000000a0. No error signal rndc_err_i is generated throughout the process. Although the read channel displays the data c1000000, the read valid signal rndc_rdata_val_i remains low, indicating that this data set should be ignored.

[0052] Figure 4 The process shown involves reading data from system memory into the MTL data buffer. The transmission start signal `mdc_start_xfer_o` is pulled low from high to trigger a bus transaction, and the address bus `mdc_addr_o` immediately outputs the starting address 000092f8. The read / write control signal `mdc_rd_wrn_o` remains low to indicate a read operation, the burst transfer count `mdc_burst_count_o` is set to 07, and the master receive ready signal `mdc_rdata_rdy_o` remains high. After several cycles of waiting, the slave device pulls high the data valid signal `mdc_rdata_val_i`, and the data bus `mdc_rdata_i` then continuously outputs eight data words: 08090a0b, 59000607, 0026b9c3, bcbbfa00, c0bfbebd, c4c3c2c1, c8c7c6c5, and cccbcac9. During the transmission, the address bus synchronization increments from 000092f8 to 00009318 in 4-byte steps. As the last data word is transmitted, the data valid signal mdc_rdata_val_i returns to zero, and the transmission completion signal mdc_xfer_done_i goes high, marking the end of this burst data read process.

[0053] like Figure 5The process shown is the process of reading the descriptor from system memory. mdc_start_xfer_o is pulled high for one clock cycle to trigger the transfer transaction. The address bus mdc_addr_o starts at address 00000180. The read / write control line mdc_rd_wrn_o is kept high to indicate the read operation mode. The burst counter mdc_burst_count_o is initialized to 03 to set the four-times data transfer length. The master side continuously sets `mdc_rdata_rdy_o` high to indicate data reception readiness. The slave side, after a response delay, pulls `mdc_rdata_val_i` high to indicate valid data read. The first clock cycle, address 00000180, corresponds to the returned data 000092f8. In the next clock cycle, the address increments to 00000184, the counter decrements to 02, and the data bus is at 00009400. In the third clock cycle, the address changes to 00000188, the counter drops to 01, and the sampled data is 80000108. In the final clock cycle, when the address reaches 0000018c, the counter resets to zero, and the last bit of data, b0000108, is transmitted. After data transmission, `mdc_rdata_val_i` is pulled low, and the feedback signal `mdc_xfer_done_i` is immediately pulled high for one clock cycle to confirm the transaction's completion. The address bus is then updated to the next base address, 00000190, and the relevant control signals are reset.

[0054] Example 2 This embodiment provides a multi-level, multi-priority DMA arbitration method, including: Receive concurrent requests from multiple independent DMA channels and select the only authorized channel; Perform data and attribute selection for the authorized channel; The system distinguishes between request types, including descriptor requests and data requests, and dynamically adjusts the arbitration priority between different requests through priority inversion signals provided by the internal state machine and the external configuration state register interface. The internal arbitration result is mapped to the external system bus, maximizing bus bandwidth utilization through read / write separation and response demultiplexing mechanisms.

[0055] It should be noted that the specific implementation of the multi-level multi-priority DMA arbitration method in this embodiment of the invention is similar to the specific implementation of the multi-level multi-priority DMA arbitration circuit in this embodiment of the invention. Please refer to the description in the method section for details. In order to reduce redundancy, it will not be repeated here.

[0056] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multi-level, multi-priority DMA arbitration circuit, characterized in that, The circuit is decoupled into three orthogonal scheduling levels at both the logical and physical levels, including: The first-level arbitration layer includes a channel-level arbitrator and a channel-level multiplexer. The channel-level arbitrator is used to receive concurrent requests from multiple independent DMA channels and filter out the uniquely authorized channel. The channel-level multiplexer is used to perform data and attribute gating of the authorized channel. The second-level arbitration layer is connected to the output of the first-level arbitration layer and includes a functional-level arbitrator and a main arbitrator. The main arbitrator has sub-arbitrators deployed within it for mutual orthogonality between read and write channels. The functional-level arbitrator is used to distinguish request types, including descriptor requests and data requests, and dynamically adjusts the arbitration priority between different requests through priority inversion signals provided by the internal state machine and the external configuration state register interface. The third-level arbitration layer includes independent read crossbar switches, write crossbar switches, and response demultiplexers; it is connected to the output of the second-level arbitration layer and is used to map the internal arbitration results to the external system bus. Through read-write separation and response demultiplexing mechanisms, the bus bandwidth utilization is maximized; the third-level arbitration layer also includes an asynchronous clock domain adaptation module. The channel-level multiplexer in the first-level arbitration layer also includes a pipelined register slice, which is used to pass authorized data requests through a first-level D flip-flop for sizing before sending them to the second-level arbitration layer, in order to eliminate combinational logic glitches and strictly limit the timing critical path within the first-level arbitration layer. During state transition evaluation, the state machine of the functional arbitrator forcibly places the decision threshold for data requests before that for descriptor requests.

2. The multi-level, multi-priority DMA arbitration circuit as described in claim 1, characterized in that, In the first-level arbitration layer, the channel-level arbitrator is a priority encoder with an enable terminal implemented using a deeply cascaded pure combinational logic network. It fixes the channel priority order from high to low. When multiple channels pull up the request level in the same clock cycle, all signals lower than the current highest request bit are shielded by the logic gate array, and the authorization vector of a single hotspot is output.

3. The multi-level, multi-priority DMA arbitration circuit as described in claim 1, characterized in that, The first-level arbitration layer channel-level multiplexer internally constructs a high-dimensional AND-OR logic tree. Using a single hotspot authorization signal as a mask, it sets all unauthorized channel data to zero. Then, through a bitwise OR operation, it aggregates the data of the only selected channel onto the one-dimensional output bus.

4. The multi-level, multi-priority DMA arbitration circuit as described in claim 1, characterized in that, The second-level arbitration layer also includes two request generation engines: a descriptor engine and a data path engine. The descriptor engine is used to prefetch descriptor streams from system memory and parse out the physical storage location of data packets, packet length, and various offload flags. These requests are uniformly classified as descriptor requests by the system. The data path engine is used to move large amounts of Ethernet frame payload data, which are classified as data requests.

5. The multi-level, multi-priority DMA arbitration circuit as described in claim 1, characterized in that, When the receive priority configuration input signal is enabled, all write memory requests from the receive data path gain preemption rights over the transmit data path; when the read priority enable signal is enabled, memory requests to read the TSO header template are set as privileged operations, allowing them to seamlessly jump the queue between normal descriptor polling.

6. The multi-level, multi-priority DMA arbitration circuit as described in claim 1, characterized in that, The read crossbar switch and the write crossbar switch are two completely independent arbiters, which are used to map internal read requests to the read address channel of the system bus and internal write requests to the write address and write data channel of the system bus, respectively, so as to realize the concurrent utilization of bidirectional bandwidth for reading and writing. The response demultiplexer is used to: when a request is sent to the system bus via a crossbar switch, tag each transaction with a transaction tag containing the identity information of the request source; when the system bus returns data or a completion status, parse the transaction tag and route the returned payload to the corresponding request source through an internal demultiplexing matrix.

7. The multi-level, multi-priority DMA arbitration circuit as described in claim 1, characterized in that, The asynchronous clock domain adaptation module includes a first asynchronous FIFO and a second asynchronous FIFO, which serve to read data and write data respectively. Internally, it adopts a Gray code pointer FIFO structure and has built-in data shifter and byte mask generation logic to splice the internal data stream into the best burst block that conforms to the bus protocol.

8. A multi-level, multi-priority DMA arbitration method, characterized in that, A multi-level, multi-priority DMA arbitration circuit based on any one of claims 1-7 includes: Receive concurrent requests from multiple independent DMA channels and select the only authorized channel; Perform data and attribute selection for the authorized channel; The system distinguishes between request types, including descriptor requests and data requests, and dynamically adjusts the arbitration priority between different requests through priority inversion signals provided by the internal state machine and the external configuration state register interface. The internal arbitration result is mapped to the external system bus, maximizing bus bandwidth utilization through read / write separation and response demultiplexing mechanisms.

Citation Information

Patent Citations

  • Deterministic packet scheduling and DMA for time sensitive networking

    CN114285510A

  • Data transmission DMA controller

    CN121560792A