A data transfer method and device for a PCIe switch

By configuring the DMA engine in the virtual endpoint device of the PCIe switch, the number of accesses to registers is reduced, data transmission efficiency is improved, the problem of frequent access to registers by the DMA engine is solved, and a unified data handling method is realized.

CN118484422BActive Publication Date: 2025-07-25WUXI STARS MICRO SYSTEM TECHNOLOGIES CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410550908.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-06
Publication Date
2025-07-25
Estimated Expiration
2044-05-06

AI Technical Summary

Technical Problem

In the prior art, the DMA engine of PCIe switch needs to frequently access registers for data transfer, resulting in low data transmission efficiency between the host and EP equipment, and the DMA drivers of different EP equipment are not unified, so it is impossible to provide a unified interface.

Method used

Configure the DMA engine in the virtual endpoint device of the PCIe switch, submit the DMA data transfer task through the host access base address register space, and use the DMA engine to extract the descriptor from the task queue in the host memory for data transfer, and add the completion descriptor to the task completion queue to reduce the number of accesses to the register.

Benefits of technology

By reducing the number of accesses to PCIe switch registers, data handling performance is improved, more efficient data transmission is achieved, and unified data handling between different EP devices is supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118484422B_ABST
    Figure CN118484422B_ABST
Patent Text Reader

Abstract

The present application provides a data transfer method and apparatus for a PCIe switch. The method includes: configuring a DMA engine in a virtual endpoint device of the PCIe switch, accessing a base address register space of the virtual endpoint device through a host, and submitting a DMA data transfer task to the DMA engine; using the DMA engine to retrieve a task descriptor from a task submission queue in a host memory, obtaining a source address and a destination address in the task descriptor, and when reading data to be transferred from the source address and writing it to the destination address, adding a corresponding completion descriptor to a task completion queue in the host memory; and completing the DMA data transfer task according to the statuses of the task submission queue and the task completion queue. The technical solution of the present application reduces the number of accesses by the host to the registers of the PCIe switch while implementing DMA transfer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of high-speed interconnection, and particularly relates to a data transfer method and apparatus for a PCIe switch. Background Art

[0002] The PCIe link uses end-to-end data transfer, and only one device can be connected to one end of a PCIe link. To expand multiple devices, a PCIe switch is introduced, enabling an RC (Root complex) to connect to multiple endpoint devices EP (Endpoint) downstream. The typical topology of PCIe is as Figure 1 shown.

[0003] The logic inside the switch is as Figure 2 shown. A PCIe switch has an upstream port UP directly or indirectly connected to the RC and multiple downstream ports DP. The DP is connected to the EP or to the next-level switch. The inside of the switch consists of multiple virtual PCI-PCI bridges. Each UP and DP corresponds to a virtual PCI-PCI bridge and conforms to the standard definition of a PCI bridge, with its own configuration space. The switch also uses the mechanism of the PCI bridge to forward traffic. When the traffic between the RC and the EP passes through the switch for forwarding, the system software is not aware of it. The switch can forward the traffic between the CPU and the EP and the P2P traffic between the EP and the EP.

[0004] In the related art, the DMA engine can be implemented through the virtual EP inside the switch and is used for data transfer between different EPs or between the host and the EP. For example, the DMA in the SSD disk is responsible for reading data from the host memory and writing it into the SSD disk, or vice versa, writing the data in the SSD disk into the host memory. Two task queues are used to represent the submission task queue and the task completion queue, and the head pointer and the tail pointer are used to represent the current submission and completion positions of the queue respectively. The host obtains the head and tail pointers by accessing the registers of the EP inside the switch. For example, in the submission queue, the CPU submits a task to the DMA engine, and the DMA engine processes the submission task and needs to modify the head and tail pointers. Similarly, in the completion queue, the CPU obtains the task processed by the DMA by accessing the registers to get the tail and head pointers, realizing data transfer between different types of EPs or between the EP and the host. However, this method requires the CPU to frequently access the registers inside the switch.

[0005] In addition, due to the different hardware designs of each EP device, not all EP devices have DMA. For EP devices with DMA, their interfaces and implementations are also different. The kernel cannot provide a unified interface, resulting in different DMA drivers for different devices. The DMA of the same device can only handle its own data transfer. Summary of the Invention

[0006] The purpose of this application is to provide a data transfer method and device for PCIe switch, aiming to reduce the number of accesses by the host to the registers of the PCIe switch while implementing DMA transfer.

[0007] According to the first aspect of this application, a data transfer method for PCIe switch is provided, including:

[0008] Configure a DMA engine in the virtual endpoint device of the PCIe switch. Through the host accessing the base address register space of the virtual endpoint device, submit a DMA data transfer task to the DMA engine;

[0009] Use the DMA engine to retrieve a task descriptor from the task submission queue in the host memory, obtain the source address and destination address in the task descriptor. When reading the data to be transferred from the source address and writing it to the destination address, add the corresponding completion descriptor to the task completion queue in the host memory;

[0010] Complete the DMA data transfer task according to the status of the task submission queue and the task completion queue.

[0011] In a possible implementation, before retrieving the task descriptor from the task submission queue in the host memory, it further includes:

[0012] Allocate memory space for the task submission queue and the task completion queue in the host memory, write the memory address and size into the register mapped to the base address register space of the virtual endpoint device, and initialize the head pointer and tail pointer of the task submission queue and the task completion queue.

[0013] In a possible implementation, the step of submitting the DMA data transfer task to the DMA engine further includes:

[0014] Write the task descriptor unit into the task submission queue, update the tail pointer of the task submission queue, and write the updated tail pointer into the register mapped to the base address register space of the virtual endpoint device to complete the task submission.

[0015] In a possible implementation, the step of fetching a task descriptor from a task submission queue in the host memory by using the DMA engine further includes:

[0016] When the task descriptor unit is fetched and processed, update the head pointer until the values of the head pointer and the tail pointer are equal, that is, all task descriptor units are processed.

[0017] In a possible implementation, after adding a corresponding completion descriptor to a task completion queue in the host memory, it further includes:

[0018] Update the tail pointer of the task completion queue, write the updated tail pointer into a register mapped to the base address register space of the virtual endpoint device, and report an interrupt to notify the host at the same time.

[0019] According to a second aspect of the present application, there is provided a data transfer device for a PCIe switch, including:

[0020] A submission unit, configured to configure a DMA engine in a virtual endpoint device of the PCIe switch, and submit a DMA data transfer task to the DMA engine by accessing the base address register space of the virtual endpoint device through the host;

[0021] A transfer unit, configured to fetch a task descriptor from a task submission queue in the host memory by using the DMA engine, obtain a source address and a destination address in the task descriptor, and add a corresponding completion descriptor to a task completion queue in the host memory when reading data to be transferred from the source address and writing it to the destination address;

[0022] A completion unit, configured to complete the DMA data transfer task according to the statuses of the task submission queue and the task completion queue.

[0023] Compared with the related art, the technical solution of the present application has at least the following advantages:

[0024] It is not necessary for the host to access registers mapped to the BAR space of the vEP inside the PCIe switch, but only to access the memory on the host side, and to sense whether there is a new TCE according to the change of the completion mark of the next TCE in the memory, so as to obtain the transfer result of the DMA engine, reducing the number of times the host accesses the registers of the PCIe switch, thereby improving the data transfer performance.

[0025] Other features and advantages of the present application will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present application. The objectives and other advantages of the present application may be realized and attained by the structure and processes pointed out in the specification, claims and drawings. Description of the Drawings

[0026] To more clearly illustrate the technical solutions in the embodiments of the present application or in the related art, the following briefly introduces the drawings required for use in the description of the embodiments or the related art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0027] Figure 1 is a typical topology diagram of PCIe according to the related art.

[0028] Figure 2 is a logic diagram inside the switch according to the related art.

[0029] Figure 3 is a flowchart of a data transfer method for a PCIe switch according to an exemplary embodiment of the present application.

[0030] Figure 4 is a connection state diagram of a virtual EP (vEP) implemented inside a PCIe switch according to an exemplary embodiment of the present application.

[0031] Figure 5 is a flowchart of applying for queue memory space according to an exemplary embodiment of the present application.

[0032] Figure 6 is a flowchart of submitting a DMA data transfer task according to an exemplary embodiment of the present application.

[0033] Figure 7 is a structure diagram of TDQ and TCQ on the host side and registers on the chip side according to an exemplary embodiment of the present application.

[0034] Figure 8 is a structure diagram of a TDE descriptor according to an exemplary embodiment of the present application.

[0035] Figure 9 is a flowchart of retrieving a task descriptor according to an exemplary embodiment of the present application.

[0036] Figure 10 is a structure diagram of a TCE descriptor according to an exemplary embodiment of the present application.

[0037] Figure 11It is a flowchart according to another exemplary embodiment of the present application.

[0038] Figure 12 It is an example process diagram of DMA data transfer between EP and host according to an exemplary embodiment of the present application.

[0039] Figures 13 - 16 It is a schematic diagram of the change of the Finish flag during the use of the TCQ circular queue according to an exemplary embodiment of the present application. Detailed implementation manners

[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0041] Based on the above analysis, the detailed implementation manners of the present application provide a data transfer method and apparatus for a PCIe switch, implementing a virtual EP (vEP) inside the PCIe switch as a carrier of the DMA engine, supporting data transfer between any ports, including but not limited to data transfer between host and host, host and EP device, and various types of EP devices. The base address register (BAR) space of the vEP is mapped to the control register of the DMA engine, and the host submits DMA tasks by accessing the BAR space. After receiving the notification, the DMA engine reads the task execution data in the host memory to perform data transfer, writes the completed tasks into the host-side memory, and notifies the host. The host obtains the number of completed tasks by accessing the BAR space of the vEP, and then processes the completed tasks.

[0042] To reduce the interaction between the host and the vEP, the embodiments of the present application further set a flag bit while writing the completed tasks into the host-side memory. The host only needs to check the flag bit of the next element in the task completion queue. If there is a change, it can be determined that a new task is completed. The task completion queue exists in the host memory rather than on the switch. Therefore, when the host processes the task completion queue, it does not need to access the switch register every time, thereby improving the performance of processing the task completion queue.

[0043] Refer to Figure 3 the flowchart of, the data transfer method for a PCIe switch provided by the embodiments of the present application includes:

[0044] Step 101: Configure the DMA engine in the virtual endpoint device of the PCIe switch. Access the base address register space of the virtual endpoint device through the host, and submit a DMA data transfer task to the DMA engine.

[0045] Exemplarily, as Figure 4 shown, vEP is implemented inside the PCIe switch and is connected to the UP through the virtual DP (vDP). The host can normally enumerate the vDP and vEP. The BAR space of the vEP maps to the control register of the DMA. The host driver initializes the task queue, submits tasks, and processes the completed tasks by accessing the BAR space.

[0046] Step 102: Use the DMA engine to retrieve a task descriptor from the task submission queue in the host memory, obtain the source address and destination address in the task descriptor, and when reading the data to be transferred from the source address and writing it to the destination address, add the corresponding completion descriptor to the task completion queue in the host memory. Exemplarily, the host driver first creates a task submission queue (TDQ) and a task completion queue (TCQ) in the host memory. Each queue is a section of independent and continuous physical space. When the host driver issues a task to the DMA, the DMA first reads the task descriptor from the host TDQ, reads data from the host according to the task descriptor information, and then writes the data to the destination address; when the data transfer of a task descriptor is completed, the DMA writes the completion descriptor (task completion unit TCE) to the TCQ queue and sends an MSI interrupt to the host.

[0047] See Figure 5 , in a specific embodiment of the present application, in step 102, before retrieving the task descriptor from the task submission queue in the host memory, it further includes:

[0048] Allocate memory space for the task submission queue and the task completion queue in the host memory, write the memory address and size into the register mapped by the base address register space of the virtual endpoint device, and initialize the head pointer and tail pointer of the task submission queue and the task completion queue.

[0049] During the initialization process, the host driver allocates memory space for the TDQ in the host-side memory and writes the memory address and size into the chip-side register mapped by the BAR of the vEP. The head and tail pointers Head and Tail of the queue are respectively initialized to 0, that is, both point to the first element.

[0050] The host driver initializes the TCQ using the same process, and configures the memory address and size of the TCQ into the chip-side register.

[0051] See Figure 6, in a specific embodiment of the present application, the step of submitting a DMA data transfer task to the DMA engine further includes:

[0052] Writing a task descriptor unit to the task submission queue, updating the tail pointer of the task submission queue, and writing the updated tail pointer to a register mapped to the base address register space of the virtual endpoint device to complete the task submission.

[0053] In some alternative embodiments, both the TDQ and TCQ are circular queue structures for easy cyclic use. The host driver is responsible for submitting tasks and is the producer of the TDQ queue. Refer to Figure 7 , each time a task descriptor unit TDE is written to the TDQ queue, the tail pointer is incremented by 1, and the updated tail pointer is written to a register mapped to the BAR of the vEP to complete the task submission. Figure 8 It shows that the task descriptor unit TDE descriptor includes information such as source address, destination address, length, descriptor ID, etc. Optionally, the host driver can write multiple TDEs at once and then update the tail pointer to the chip-side register.

[0054] Refer to Figure 9 , in a specific embodiment of the present application, the step of using the DMA engine to retrieve a task descriptor from a task submission queue in the host memory further includes:

[0055] When the task descriptor unit is retrieved and processed, the head pointer is updated until the values of the head pointer and the tail pointer are equal, indicating that all task descriptor units have been processed.

[0056] The DMA engine in the chip is the consumer of the TDQ queue. The DMA engine determines whether there are new tasks to be processed based on the positions of the tail and head pointers. Each time a TDE is processed, the head pointer is incremented by 1 until the head and tail values are equal, indicating that all TDEs have been processed.

[0057] After the DMA finishes processing a TDE, a task completion unit TCE is written to the TCQ queue, and the host driver processes the TCE to end the previously submitted TDE task. Figure 10 Exemplarily, it shows that the task completion unit TCE descriptor includes information such as completion flag, status, TDE_ID, etc. The TCE is the basic unit of the TCQ.

[0058] Refer to Figure 11 , in a specific embodiment of the present application, after adding the corresponding completion descriptor to the task completion queue in the host memory, it further includes:

[0059] Update the tail pointer of the task completion queue, write the updated tail pointer into the register mapped in the base address register space of the virtual endpoint device, and report an interrupt to notify the host at the same time.

[0060] The DMA engine is the producer of the TCQ queue. After processing each TDE, it generates a TCE. The TDE_ID in the TCE indicates the corresponding processed TDE. The DMA engine writes the TCE into the TCQ queue, adds 1 to the tail pointer, and writes it into the TCQ register of the chip for the host software to query, and reports an MSI interrupt to notify the host.

[0061] Step 103: Complete the DMA data transfer task according to the status of the task submission queue and the task completion queue.

[0062] It can be understood that the host driver is the consumer of the TCQ queue. It reads the TCE produced by the DMA engine according to the tail and head pointers of the TCQ, and completes the task closing work according to the status of the TCE and the mapped TDE. Each time the host driver consumes a TCE, it adds 1 to the head pointer until the head and tail values are equal, indicating that all TCEs have been processed.

[0063] Exemplarily, see Figure 12 , the following takes the DMA data transfer between the EP and the host as an example to describe the entire data flow.

[0064] Step ①, the host driver issues a task to the DMA engine and writes the Tail pointer of the TDQ into the DMA register.

[0065] Step ②, the DMA engine determines that there is a new task in the TDQ, sends a read operation request to the host, and reads the TDE in the TDQ.

[0066] Step ③, return the read TDE to the DMA engine.

[0067] Step ④, the DMA engine parses the TDE, and according to information such as the source address and length, sends a Memory Read to the RAM of the EP to read the content.

[0068] Step ⑤, the EP returns the read data and saves it to the FIFO of the DMA.

[0069] Step ⑥, the DMA sends a Memory Write command according to the destination address in the TDE to write the data in the FIFO into the host memory.

[0070] Step ⑦, after the DMA completes the data transfer from the source address to the destination address, it writes the TCE into the host's TCQ queue.

[0071] Step ⑧: The DMA sends an MSI interrupt to notify the host that the DMA operation is completed. The host driver receives the MSI interrupt and writes the head pointer of the TCQ into the DMA register.

[0072] In step ⑧, after the host driver receives the MSI interrupt, it reads the TCE in two optional ways:

[0073] The first way is to read the tail and head values of the TCQ in the DMA register of the chip, then read all the TCEs between the head and the tail, and process each TCE. After completion, update the head to the DMA register.

[0074] The second way is to add a completion flag (Finish flag) to the TCE structure. When initializing the TCQ, the Finish flag of all TCEs = 0, as Figure 13 shown.

[0075] When the DMA engine writes the TCE for the first round, set the Finish flag of the new TCE to 1. Since the TCQ is a circular queue, when the TCQ queue returns from the tail to the head, flip the Finish flag. For example, when writing the TCE in the second round, set the Finish flag of the new TCE to 0, and so on, flipping the flag once every round, as Figure 14 shown.

[0076] The host driver reads the TCE at the current head position in the TCQ and determines whether the Finish flag has changed. If there is a change, it is determined that the TCE is a newly generated TCE by the DMA engine, reads and processes the TCE, adds 1 to the head value, and reads the TCE corresponding to the new head value. Until the Finish flag in the TCE does not change, the traversal ends and all the TCEs are consumed.

[0077] Since the TCQ is a circular queue and is used cyclically, the Finish flag is flipped as a whole once after each traversal of the queue. For example, after the TCQ queue is full in the first round, the Finish flag of all TCEs = 1, as Figure 15 shown. When the DMA engine generates a new TCE again, it returns to the head of the TCQ queue and writes 0 to the Finish flag of the TCE, as Figure 16 shown. When the host driver consumes the TCEs in the TCQ queue, it also flips the flag judgment value once after each traversal of the TCQ queue.

[0078] Those skilled in the art can understand that the DMA data transfer from EP to EP or from the host to EP, etc. is similar to the above process and will not be elaborated here.

[0079] In the related art, the host driver must send a Memory Read message to the vEP through the PCIe link, read the registers mapped in the BAR space of the vEP, obtain the head and tail pointers of the TCQ, and sense whether there is a new TCE based on the head and tail pointers to know the transfer result of the DMA engine. It can be seen that the data transfer method for the PCIe switch proposed in this application has the following advantages compared with the related art:

[0080] It only needs to access the memory on the host side, sense whether there is a new TCE based on the change of the completion flag of the next TCE in the memory to obtain the transfer result of the DMA engine, without the host accessing the registers mapped in the BAR space of the vEP inside the PCIe switch, thus reducing the number of times the host accesses the registers of the PCIe switch and improving the data transfer performance.

[0081] Correspondingly, a specific embodiment of this application also provides a data transfer device for a PCIe switch, including:

[0082] A submission unit, configured to configure a DMA engine in the virtual endpoint device of the PCIe switch, and submit a DMA data transfer task to the DMA engine by the host accessing the base address register space of the virtual endpoint device;

[0083] A transfer unit, configured to use the DMA engine to fetch a task descriptor from a task submission queue in the host memory, obtain the source address and destination address in the task descriptor, and when reading the data to be transferred from the source address and writing it to the destination address, add the corresponding completion descriptor to the task completion queue in the host memory;

[0084] A completion unit, configured to complete the DMA data transfer task according to the status of the task submission queue and the task completion queue.

[0085] The above device can be implemented by the data transfer method for the PCIe switch provided in the embodiments of the first aspect. The specific implementation manner can refer to the description in the embodiments of the first aspect and will not be elaborated here.

[0086] It can be understood that the structures, names, and parameters described in the above embodiments are only examples. Those skilled in the art can also easily combine and adjust the structural features of the above multiple embodiments according to the usage needs, and should not limit the concept of this application to the specific details of the above examples.

[0087] Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A data transfer method for a PCIe switch, characterized in that, Including: Configure a DMA engine in the virtual endpoint device of the PCIe switch, and through the host accessing the base address register space of the virtual endpoint device, submit a DMA data transfer task to the DMA engine; Use the DMA engine to fetch a task descriptor from the task submission queue in the host memory, obtain the source address and destination address in the task descriptor, and when reading the data to be transferred from the source address and writing it to the destination address, add the corresponding completion descriptor to the task completion queue in the host memory; Complete the DMA data transfer task according to the status of the task submission queue and the task completion queue; After adding the corresponding completion descriptor to the task completion queue in the host memory, further including: Update the tail pointer of the task completion queue, write the updated tail pointer into the register mapped to the base address register space of the virtual endpoint device, and at the same time report an interrupt to notify the host; After the host receives the interrupt, read the completion descriptor in one of the following ways: Read the values of the tail pointer and the head pointer of the task completion queue, then read all the completion descriptors between the head pointer and the tail pointer, process each completion descriptor, and after completion, update the head pointer to the DMA register; Or Add a completion flag to the completion descriptor structure, and when initializing the task completion queue, set the completion flags of all completion descriptors to 0.

2. The data transfer method for a PCIe switch according to claim 1, wherein Before fetching the task descriptor from the task submission queue in the host memory, further including: Apply for memory space for the task submission queue and the task completion queue in the host memory, write the memory address and size into the register mapped to the base address register space of the virtual endpoint device, and initialize the head pointer and the tail pointer of the task submission queue and the task completion queue.

3. The data transfer method for a PCIe switch according to claim 2, wherein The step of submitting the DMA data transfer task to the DMA engine further includes: Write the task descriptor unit into the task submission queue, update the tail pointer of the task submission queue, and write the updated tail pointer into the register mapped to the base address register space of the virtual endpoint device to complete the task submission.

4. The data transfer method for PCIe switch according to claim 3, wherein The step of using the DMA engine to fetch the task descriptor from the task submission queue in the host memory further includes: When the task descriptor unit is fetched and processed, update the head pointer until the values of the head pointer and the tail pointer are equal, that is, all the task descriptor units are processed.

5. A data transfer device for a PCIe switch, characterized in that, Including: A submission unit, configured to configure a DMA engine in the virtual endpoint device of the PCIe switch, and through the host accessing the base address register space of the virtual endpoint device, submit a DMA data transfer task to the DMA engine; A transfer unit, configured to use the DMA engine to fetch a task descriptor from the task submission queue in the host memory, obtain the source address and destination address in the task descriptor, and when reading the data to be transferred from the source address and writing it to the destination address, add the corresponding completion descriptor to the task completion queue in the host memory; A completion unit, configured to complete the DMA data transfer task according to the status of the task submission queue and the task completion queue; The transfer unit is further configured to: after adding the corresponding completion descriptor to the task completion queue in the host memory, update the tail pointer of the task completion queue, write the updated tail pointer into the register mapped to the base address register space of the virtual endpoint device, and report an interrupt to notify the host at the same time; After the host receives the interrupt, read the completion descriptor in one of the following ways: Read the values of the tail pointer and the head pointer of the task completion queue, then read all the completion descriptors between the head pointer and the tail pointer, process each completion descriptor, and update the head pointer to the DMA register after completion; Or Add a completion flag to the completion descriptor structure, and set the completion flags of all completion descriptors to 0 when initializing the task completion queue.

6. The data transfer device for a PCIe switch according to claim 5, wherein The transfer unit is further configured to: Before taking out the task descriptor from the task submission queue in the host memory, apply for memory space for the task submission queue and the task completion queue in the host memory, write the memory address and size into the register mapped to the base address register space of the virtual endpoint device, and initialize the head pointer and the tail pointer of the task submission queue and the task completion queue.

7. The data transfer device for a PCIe switch according to claim 6, wherein The submission unit is further configured to: Write the task descriptor unit into the task submission queue, update the tail pointer of the task submission queue, and write the updated tail pointer into the register mapped to the base address register space of the virtual endpoint device to complete the task submission.

8. The data transfer device for a PCIe switch according to claim 7, wherein The transfer unit is further configured to: When the task descriptor unit is taken out and processed, update the head pointer until the values of the head pointer and the tail pointer are equal, that is, all task descriptor units are processed.

Citation Information

Patent Citations

  • Intelligent network card rapid DMA design method and system, equipment and terminal

    CN113986791A