Data issuing method, data uploading method and related device

By adopting a single virtual queue management mechanism in FPGA heterogeneous accelerators, the problem of inter-chip communication complexity after the expansion of high-end FPGA chip resources is solved, efficient network and computing accelerated data transmission is achieved, design complexity is reduced and memory access efficiency is improved.

CN120602555AActive Publication Date: 2025-09-05INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202511088231.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-09-05
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

As high-end FPGA chip resources expand, when FPGA heterogeneous devices simultaneously have virtualized network cards and computing acceleration functions, the power consumption and latency overhead of inter-chip communication increase, leading to increased design complexity.

Method used

A single virtual queue management mechanism is adopted to achieve efficient transmission of data packets sent from the network and accelerated by computing through direct storage access requests and prefetch mechanisms, avoiding the performance loss of multiple context switches and virtual address translations, and supporting high-bandwidth and low-latency data transmission.

Benefits of technology

It achieves efficient data transmission of network functions and computing acceleration functions in heterogeneous accelerators, reduces design complexity, improves memory access efficiency and CPU load, and supports robust reliability of frequent two-way data interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602555A_ABST
    Figure CN120602555A_ABST
Patent Text Reader

Abstract

The invention discloses a data issuing method, a data uploading method and a related device, and relates to the technical field of data processing.The data issuing method comprises the steps that when a network issuing instruction is received, a first direct storage access request is initiated according to a first issuing descriptor, and a first issuing data packet is obtained from a storage area, transmitting the first issuing data packet to the first equipment through the first issuing virtual queue; and when a calculation acceleration issuing instruction is received, initiating a second direct storage access request according to a second issuing descriptor, acquiring a second issuing data packet from the storage area, storing the second issuing data packet into the cache, and transmitting the second issuing data packet to a target processing module for processing according to a protocol frame type of the second issuing data packet, and the response frame is generated and returned to the upper computer, so that the technical problem that a heterogeneous accelerator is complicated in design when multifunctional data transmission is realized in related technologies is solved, and data transmission of a network function and a calculation acceleration function under the management of the same virtual queue is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data downloading method, a data uploading method and related devices. Background Art

[0002] With the increasing availability of high-end FPGA chips, data center users often require heterogeneous FPGA devices to simultaneously support both virtualized network cards and compute acceleration. This places higher demands on the design of data transmission systems between FPGAs (heterogeneous accelerators) and host computers. If multiple FPGA chips are used, each implementing a single function, applications requiring both functions will face the power consumption and latency overhead of inter-chip communication, increasing the complexity of inter-chip and even inter-board synchronization. Summary of the Invention

[0003] The present application provides a data sending method, an uploading method and related devices to at least solve the problem of complex design of heterogeneous accelerators when implementing multifunctional data transmission.

[0004] This application provides a data delivery method, which is applied to a heterogeneous accelerator. The data delivery method includes: In response to receiving a first type of sending instruction sent by the host computer, the first type of sending instruction is a network sending instruction, the first type of sending instruction is parsed, a first sending descriptor is obtained, a first direct storage access request is initiated according to the first sending descriptor, a first sending data packet is obtained from the storage area, a first sending virtual queue is determined based on the first sending descriptor, and the first sending data packet is transmitted to the first device through the first sending virtual queue; in response to receiving a second type of sending instruction sent by the host computer, the second type of sending instruction is a computing acceleration sending instruction, the second type of sending instruction is parsed, a second sending descriptor is obtained, a second direct storage access request is initiated according to the second sending descriptor, a second sending data packet is obtained from the storage area, the second sending data packet is stored in a cache, a protocol frame is parsed for the second sending data packet, a protocol frame type is obtained, and the second sending data packet is transmitted to a target processing module of the second device for processing according to the protocol frame type; in response to completion of processing of the second sending data packet, a response frame is generated, and the response frame is returned to the host computer.

[0005] This application also provides a data uploading method, which is applied to a heterogeneous accelerator. The data uploading method includes: In response to receiving a first type of upload instruction sent by a first device, the first type of upload instruction is a network upload instruction, the first type of upload instruction carries a first upload data packet, a first upload descriptor is selected from a pre-cached upload descriptor set, and the first upload data packet is written to a first storage address corresponding to the first upload descriptor according to the first upload descriptor, so as to transmit the first upload data packet to the host computer based on the first storage address; in response to receiving a second type of upload instruction sent by a second device, the second type of upload instruction is a computing acceleration upload instruction, the second type of upload instruction carries a second upload data packet, a second upload descriptor is selected from a pre-cached upload descriptor set, and the second upload data packet is written to a second storage address corresponding to the second upload descriptor according to the second upload descriptor, so as to transmit the second upload data packet to the host computer based on the second storage address.

[0006] The present application also provides a data sending device, which is applied to a heterogeneous accelerator. The data sending device includes: The first data sending unit is configured to receive a first type of sending instruction sent by a host computer, the first type of sending instruction being a network sending instruction, parse the first type of sending instruction, obtain a first sending descriptor, initiate a first direct storage access request based on the first sending descriptor, obtain a first sending data packet from a storage area, determine a first sending virtual queue based on the first sending descriptor, and transmit the first sending data packet to the first device through the first sending virtual queue; and the second data sending unit is configured to receive a second type of sending instruction sent by the host computer, the second type of sending instruction being a computing acceleration sending instruction, parse the second type of sending instruction, obtain a second sending descriptor, initiate a second direct storage access request based on the second sending descriptor, obtain a second sending data packet from the storage area, store the second sending data packet in a cache, perform protocol frame parsing on the second sending data packet, obtain a protocol frame type, transmit the second sending data packet to a target processing module of the second device for processing according to the protocol frame type, generate a response frame in response to completion of processing of the second sending data packet, and return the response frame to the host computer.

[0007] The present application also provides a data upload device, which is applied to a heterogeneous accelerator. The data upload device includes: The first data upload unit is used to receive a first type of upload instruction sent by a first device, the first type of upload instruction is a network upload instruction, the first type of upload instruction carries a first upload data packet, selects a first upload descriptor from a pre-cached upload descriptor set, and writes the first upload data packet to a first storage address corresponding to the first upload descriptor according to the first upload descriptor, so as to transmit the first upload data packet to the host computer based on the first storage address; the second data upload unit is used to receive a second type of upload instruction sent by a second device, the second type of upload instruction is a computing acceleration upload instruction, the second type of upload instruction carries a second upload data packet, selects a second upload descriptor from a pre-cached upload descriptor set, writes the second upload data packet to a second storage address corresponding to the second upload descriptor according to the second upload descriptor, so as to transmit the second upload data packet to the host computer based on the second storage address.

[0008] The present application also provides a computer device, including: a memory for storing a computer program; a processor for implementing the steps of the data sending method in the following embodiments or the steps of the data uploading method in any of the following embodiments when executing the computer program.

[0009] The data sending method includes: in response to receiving a first type of sending instruction sent by a host computer, the first type of sending instruction is a network sending instruction, parsing the first type of sending instruction, obtaining a first sending descriptor, initiating a first direct storage access request according to the first sending descriptor, obtaining a first sending data packet from a storage area, determining a first sending virtual queue based on the first sending descriptor, and transmitting the first sending data packet to a first device through the first sending virtual queue; in response to receiving a second type of sending instruction sent by the host computer, the second type of sending instruction is a computing acceleration sending instruction, parsing the second type of sending instruction, obtaining a second sending descriptor, initiating a second direct storage access request according to the second sending descriptor, obtaining a second sending data packet from the storage area, storing the second sending data packet in a cache, performing protocol frame parsing on the second sending data packet, obtaining a protocol frame type, transmitting the second sending data packet to a target processing module of a second device for processing according to the protocol frame type, generating a response frame in response to completion of processing of the second sending data packet, and returning the response frame to the host computer.

[0010] The data upload method includes: in response to receiving a first type of upload instruction sent by a first device, the first type of upload instruction is a network upload instruction, the first type of upload instruction carries a first upload data packet, selects a first upload descriptor from a pre-cached upload descriptor set, writes the first upload data packet to a first storage address corresponding to the first upload descriptor according to the first upload descriptor, and transmits the first upload data packet to a host computer based on the first storage address; in response to receiving a second type of upload instruction sent by a second device, the second type of upload instruction is a computing acceleration upload instruction, the second type of upload instruction carries a second upload data packet, selects a second upload descriptor from a pre-cached upload descriptor set, writes the second upload data packet to a second storage address corresponding to the second upload descriptor according to the second upload descriptor, and transmits the second upload data packet to the host computer based on the second storage address.

[0011] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the data sending method in any of the following embodiments or the steps of the data uploading method in any of the following embodiments are implemented.

[0012] The data sending method includes: in response to receiving a first type of sending instruction sent by a host computer, the first type of sending instruction is a network sending instruction, parsing the first type of sending instruction, obtaining a first sending descriptor, initiating a first direct storage access request according to the first sending descriptor, obtaining a first sending data packet from a storage area, determining a first sending virtual queue based on the first sending descriptor, and transmitting the first sending data packet to a first device through the first sending virtual queue; in response to receiving a second type of sending instruction sent by the host computer, the second type of sending instruction is a computing acceleration sending instruction, parsing the second type of sending instruction, obtaining a second sending descriptor, initiating a second direct storage access request according to the second sending descriptor, obtaining a second sending data packet from the storage area, storing the second sending data packet in a cache, performing protocol frame parsing on the second sending data packet, obtaining a protocol frame type, transmitting the second sending data packet to a target processing module of a second device for processing according to the protocol frame type, generating a response frame in response to completion of processing of the second sending data packet, and returning the response frame to the host computer.

[0013] The data upload method includes: in response to receiving a first type of upload instruction sent by a first device, the first type of upload instruction is a network upload instruction, the first type of upload instruction carries a first upload data packet, selects a first upload descriptor from a pre-cached upload descriptor set, writes the first upload data packet to a first storage address corresponding to the first upload descriptor according to the first upload descriptor, and transmits the first upload data packet to a host computer based on the first storage address; in response to receiving a second type of upload instruction sent by a second device, the second type of upload instruction is a computing acceleration upload instruction, the second type of upload instruction carries a second upload data packet, selects a second upload descriptor from a pre-cached upload descriptor set, writes the second upload data packet to a second storage address corresponding to the second upload descriptor according to the second upload descriptor, and transmits the second upload data packet to the host computer based on the second storage address.

[0014] The data sending method, uploading method and related devices provided by the present application include: when receiving a network sending instruction, initiating a first direct storage access request according to a first sending descriptor, obtaining a first sending data packet from a storage area, and transmitting the first sending data packet to a first device through a first sending virtual queue; when receiving a computing acceleration sending instruction, initiating a second direct storage access request according to a second sending descriptor, obtaining a second sending data packet from a storage area, storing the second sending data packet in a cache, transmitting the second sending data packet to a target processing module for processing according to a protocol frame type of the second sending data packet, generating a response frame, and returning the response frame to a host computer, thereby solving the technical problem of complex design of heterogeneous accelerators when realizing multi-functional data transmission in related technologies, and realizing data transmission of network functions and computing acceleration functions under the management of the same virtual queue. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0016] Figure 1 A flowchart of a data delivery method according to an embodiment of the present invention; Figure 2 A flowchart of a data uploading method provided in an embodiment of the present application; Figure 3 A schematic diagram of a descriptor provided in an embodiment of the present application; Figure 4 A schematic diagram of the structure of a heterogeneous accelerator provided in one embodiment of the present application; Figure 5A schematic structural diagram of a heterogeneous accelerator provided in another embodiment of the present application; Figure 6 A schematic structural diagram of a heterogeneous accelerator provided in yet another embodiment of the present application; Figure 7 This is a diagram of the internal structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0017] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0018] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0019] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0020] In cloud services that rely on data center networks, FPGA virtualization technology, hosted on heterogeneous devices such as GPUs and FPGA accelerator cards, is widely used by cloud service providers due to its programmability and manageability. This technology centrally manages and abstracts FPGA logic resources, flexibly providing compute acceleration and network functions, either provided by heterogeneous device vendors or customized by users, to virtual machine users. This enables the export of FPGA computing power to the cloud and high-speed interconnection between virtual machines. FPGAs, due to their advantages in high-bandwidth parallel data processing and reconfigurability, are often preferred in the construction of network cards and reconfigurable compute acceleration devices.

[0021] Heterogeneous device virtualization can be achieved through VFs created using SRIOV (Single Root I / O Virtualization) technology. By designing a virtual queue-based DMA (Direct Memory Access) transfer mechanism, VF (Virtual Function) devices and drivers can achieve high-speed communication over PCIe links. This communication mechanism is user-friendly for virtual machines and containers, avoiding additional communication overhead between the host and the user. For network interface cards (NICs), their service characteristics are high-bandwidth bidirectional data transmission. Therefore, using relatively independent unidirectional transmit and receive queues for DMA design facilitates higher descriptor scheduling efficiency, thereby achieving high bandwidth. For computing accelerators, their primary function lies in the algorithm logic design. For data transmission, the transmission of parameters and calculation results requires high-bandwidth burst transmission, while the control of algorithm execution requires frequent bidirectional data exchange. Unidirectional queues are more optimal for the former, while bidirectional queues are more optimal for the latter.

[0022] This application proposes a data transmission method, data upload method, and related apparatus that can simultaneously meet the data transmission requirements of network virtual devices (such as virtual network cards) and computing acceleration virtual devices. Using a single virtual queue type—a unidirectional receive and transmit queue—this method, on the one hand, meets the high-bandwidth, low-latency requirements of network card virtual devices through a receive-direction prefetch mechanism (pre-caching upload descriptor sets); and on the other hand, by designing a computing acceleration data transmission protocol, robust and reliable frequent two-way data exchange with handshakes is achieved using unidirectional receive and transmit queues.

[0023] like Figure 1 As shown, an embodiment of the present application provides a data sending method, which includes: Step 101: In response to receiving a first type of sending instruction sent by a host computer, the first type of sending instruction is a network sending instruction, the first type of sending instruction is parsed, a first sending descriptor is obtained, a first direct storage access request is initiated according to the first sending descriptor, a first sending data packet is obtained from the storage area, a first sending virtual queue is determined based on the first sending descriptor, and the first sending data packet is transmitted to the first device through the first sending virtual queue.

[0024] Specifically, in response to obtaining the first downlink descriptor, a first direct storage access request is initiated, and the first downlink data packet is obtained from the first downlink data packet storage address based on the first direct storage access request, and the first downlink descriptor includes the first downlink data packet storage address and the first downlink data packet length; in response to obtaining the first downlink data packet, the first downlink data packet is converted into a first downlink data packet in a target format; based on the first downlink descriptor, a first downlink virtual queue is determined, and based on the first downlink virtual queue, the first downlink data packet in the target format is transmitted to the first device; in response to the first device receiving the first downlink data packet in the target format, the first downlink data packet in the target format is verified, fragmented, and encapsulated to obtain a processed first downlink data packet; and the processed first downlink data packet is sent to an external network via Ethernet.

[0025] The host computer of the present application may specifically include a virtual machine or a container. When the virtual machine receives a network-sent instruction, the host computer may parse the network-sent instruction and obtain network-sent data packet information (first sent data packet related information). The virtual network driver in the virtual machine obtains the physical address and length of the network-sent data packet information. The virtual network driver generates a network-sent descriptor (first sent descriptor) based on the physical address and length of the network-sent data packet.

[0026] The heterogeneous accelerator in this application may include a high-speed serial computer expansion bus hard core, an adapter module, a virtual queue management module, a virtual network card device, a network unit, a memory unit, and an Ethernet.

[0027] In response to the high-speed serial computer expansion bus hard core detecting the generation of a network delivery descriptor via a first data channel (the CQ channel of the PCIE hard core), the high-speed serial computer expansion bus hard core initiates a direct storage access request and, based on the direct storage access request, retrieves a network delivery data packet (a first delivery data packet) from a memory unit (storage area). In response to the adapter module retrieving the network delivery data packet, the adapter module parses the network delivery data packet into a target format. The virtual queue management module determines a target delivery network virtual queue (the first delivery virtual queue) and, based on the target delivery network virtual queue, sends the network delivery data packet to a virtual network card device (the first device). The virtual network card device transmits the network delivery data packet to the network unit (the first device). The network unit verifies, fragments, and encapsulates the network delivery data packet to obtain a target network delivery data packet and sends the target network delivery data packet to the Ethernet.

[0028] In this way, the high-speed serial computer expansion bus (such as PCIE) hard core actively monitors descriptors and triggers DMA requests (direct memory access requests), bypassing the overhead of multiple context switches in traditional virtualization and significantly reducing latency. The virtual network driver directly operates the physical address, avoiding the performance loss of virtual address conversion and improving memory access efficiency. DMA directly obtains data packets from the memory unit without the CPU participating in data transfer, reducing memory bandwidth usage and CPU load.

[0029] Step 102: In response to receiving a second type of sending instruction sent by the upper computer, the second type of sending instruction is a computing acceleration sending instruction, the second type of sending instruction is parsed, a second sending descriptor is obtained, a second direct storage access request is initiated according to the second sending descriptor, a second sending data packet is obtained from the storage area, the second sending data packet is stored in the cache, a protocol frame is parsed for the second sending data packet, a protocol frame type is obtained, and the second sending data packet is transmitted to the target processing module of the second device for processing according to the protocol frame type. In response to the completion of the processing of the second sending data packet, a response frame is generated and the response frame is returned to the upper computer.

[0030] Specifically, in response to obtaining the second downlink descriptor, a second direct storage access request is initiated, and the second downlink data packet is obtained from the second downlink data packet storage address based on the second direct storage access request, and the second downlink descriptor includes the second downlink data packet storage address and the second downlink data packet length; in response to obtaining the second downlink data packet, the second downlink data packet is stored in the cache; the second downlink data packet is parsed to obtain the protocol frame type, and the protocol frame type includes a memory write command frame, a memory read command frame, a status register read and write frame, and a control register read and write frame; the first processing module of the second device is used to process the second downlink data packet corresponding to the memory write command frame, the memory read command frame, and the control register read and write frame; the second processing module of the second device is used to process the second downlink data packet corresponding to the status register read and write frame; in response to the completion of the processing of the second downlink data packet, a response frame is generated, and the response frame is returned to the host computer.

[0031] When the virtual machine receives the acceleration delivery instruction, it parses it and obtains the acceleration delivery data packet (the second delivery data packet). The virtual acceleration driver in the virtual machine generates an acceleration protocol frame based on the parameter transfer relationship, register read / write, and memory read information in the acceleration delivery data packet. It also generates an acceleration delivery descriptor based on the host physical address of the acceleration protocol frame.

[0032] The heterogeneous accelerator in this application may also include a virtual computing acceleration device and an acceleration unit.

[0033] In response to the high-speed serial computer expansion bus hard core detecting the generation of a computing acceleration delivery descriptor (a second delivery descriptor) via the first data channel, the high-speed serial computer expansion bus hard core initiates a second direct storage access request and retrieves the computing acceleration delivery data packet from the memory unit based on the second storage access request. In response to receiving the accelerated delivery data packet, the high-speed serial computer expansion bus hard core stores the accelerated delivery data packet into a cache via the second data channel (the RC channel of the PCIE hard core).

[0034] In response to storing the accelerated downlink data packet in the cache, the protocol frame header information of the accelerated downlink data packet is parsed to obtain the protocol frame type, which includes a memory write command frame, a memory read command frame, an acceleration unit status register read / write frame, and an acceleration unit control register read / write frame. In response to the protocol frame type being a first protocol frame type, information related to the accelerated downlink data packet corresponding to the first protocol frame type is collected and transmitted to a first processing module (which may be an interface conversion module) of the second device for processing. The first protocol frame type includes at least one of a memory write command frame, a memory read command frame, and an acceleration unit control register read / write frame. In response to the protocol frame type being a second protocol frame type, which is an acceleration unit status register read / write frame, the accelerated downlink data packet corresponding to the second protocol frame type is routed to a second processing module (which may be an acceleration unit management module) of the second device for processing. In response to completion of processing of the accelerated downlink data packet, a response frame is created and sent to a target receive queue.

[0035] In this way, the virtual acceleration driver directly generates computing acceleration protocol frames, abstracting computing tasks into hardware-recognizable instruction streams and avoiding the software protocol stack parsing overhead. After the PCIE hard core receives data packets via the first channel, they are cached via the second channel, achieving complete hardware offload of data movement and computation without CPU intervention. This enables dual-channel DMA transfers. Dynamically distinguishing between memory operations, register control, and other types based on the protocol frame header supports the diverse interaction requirements of heterogeneous acceleration units (such as GPUs, FPGAs, and ASICs). It also supports parallel processing of data transmissions corresponding to protocol frame types. The first protocol frame type (memory / control register operations) is directly connected to the acceleration unit hardware through the interface conversion module, achieving low-latency data read and write. The second protocol frame type (status register operations) is uniformly handled by the acceleration unit management module, ensuring atomicity and consistency of state synchronization.

[0036] like Figure 2 As shown, an embodiment of the present application provides a data uploading method, which includes: Step 201: In response to receiving a first type of upload instruction sent by a first device, the first type of upload instruction is a network upload instruction, the first type of upload instruction carries a first upload data packet, a first upload descriptor is selected from a pre-cached upload descriptor set, and the first upload data packet is written to a first storage address corresponding to the first upload descriptor according to the first upload descriptor, so as to transmit the first upload data packet to a host computer based on the first storage address.

[0037] Among them, writing the first upload data packet into the first storage address corresponding to the first upload descriptor according to the first upload descriptor, so as to transmit the first upload data packet to the host computer based on the first storage address includes: in response to the completion of writing the first upload data packet into the first storage address corresponding to the first upload descriptor, triggering the first interrupt instruction, and sending the first interrupt instruction to the host computer; in response to the host computer receiving the first interrupt instruction, the host computer reads the first upload data packet based on the first storage address.

[0038] When the network unit receives a network upload instruction, which carries a network upload data packet, the virtual queue management module selects a network upload descriptor from a pre-cached set of upload descriptors. The network upload descriptor includes the memory address of the target network upload virtual machine. In response to the high-speed serial computer expansion bus hard core acquiring the network upload descriptor, the adapter module initiates and writes the network upload data packet to the memory address corresponding to the network upload descriptor based on a direct storage write request. In response to writing the network upload data packet to the memory address corresponding to the network upload descriptor, the adapter module triggers a network interrupt instruction and sends it to the virtual machine (host computer). In response to the virtual machine receiving the network interrupt instruction, the virtual machine reads the network upload data from the memory address corresponding to the network upload descriptor.

[0039] Step 202: In response to receiving a second type of upload instruction sent by the second device, the second type of upload instruction is a computing acceleration upload instruction, the second type of upload instruction carries a second upload data packet, a second upload descriptor is selected from a pre-cached upload descriptor set, and the second upload data packet is written to a second storage address corresponding to the second upload descriptor according to the second upload descriptor, so as to transmit the second upload data packet to the host computer based on the second storage address.

[0040] Among them, writing the second upload data packet into the second storage address corresponding to the second upload descriptor according to the second upload descriptor, so as to transmit the second upload data packet to the host computer based on the second storage address includes: in response to the completion of writing the second upload data packet into the second storage address corresponding to the second upload descriptor, triggering the second interrupt instruction, and sending the second interrupt instruction to the host computer; in response to the host computer receiving the second interrupt instruction, the host computer reads the second upload data packet based on the second storage address.

[0041] Specifically, when the acceleration unit receives an accelerated upload instruction, the accelerated upload instruction carries a computing accelerated upload data packet, and the virtual queue management module selects an accelerated upload descriptor from a pre-cached upload descriptor set, and the accelerated upload descriptor includes the target accelerated upload virtual machine memory address. In response to the high-speed serial computer expansion bus hard core obtaining the accelerated upload descriptor, it initiates and writes the computing accelerated upload data packet to the memory address corresponding to the accelerated upload descriptor based on a direct storage write request. In response to writing the computing accelerated upload data packet to the memory address corresponding to the accelerated upload descriptor, the adapter module triggers an accelerated interrupt instruction and sends the accelerated interrupt instruction to the virtual machine. In response to the virtual machine receiving the accelerated interrupt instruction, the virtual machine reads the accelerated upload data from the memory address corresponding to the accelerated upload descriptor.

[0042] To do this, a prefetch mechanism is implemented for descriptors in the receive virtual queues: a cache of a certain size is allocated for all receive queues within the heterogeneous accelerator, and the driver is used to prefill the available descriptor queues. This allows the receive descriptors to be cached in the heterogeneous accelerator in advance of incoming uploads. This allows upload packets from network devices or response packets from compute accelerators to be immediately uploaded and directly stored using the locally cached descriptors upon reaching the virtual queue management module.

[0043] In one embodiment, if Figure 3 Figure 2 shows the ring structure of virtual queues (network downlink and / or accelerated downlink virtual queues, network upload and / or accelerated upload virtual queues) and the structure of the descriptor table. Virtual queues use a ring structure, and the queue depth is the maximum number of descriptors the queue can accommodate. Queues are divided into available queues and used queues. The available queue is configured by the host computer, from which the FPGA retrieves and uses descriptors. After the FPGA uses a descriptor, it updates it to the used queue, where the host computer driver recognizes it and invalidates the corresponding descriptor in the descriptor table. The queue depth is the maximum number of descriptors that can be cached at any one time. It determines the efficiency of descriptor retrieval and, in turn, affects the queue's data transmission bandwidth.

[0044] In the virtual queue management module, virtual queues for both devices are handled identically within the FPGA. Software determines whether a queue belongs to a network device or a compute accelerator. The number of devices, the number of VFs that can be expanded via SRIOV, and the number of per-device queues for network devices and compute accelerator devices are all configurable.

[0045] In this application, each virtual machine and container user can be assigned one or more VF (virtual) devices. Virtual devices include virtual network devices (network VFs) and virtual computing acceleration devices (computing acceleration VFs). Each virtual network device and virtual computing acceleration device can be assigned one or more pairs of virtual queues. In this application, multiple pairs of virtual queues are assigned to each virtual network device to meet the communication needs of high-bandwidth heterogeneous acceleration systems for network devices. For example, four pairs of network virtual queues can be assigned to a virtual network device, and two pairs of network virtual queues can be assigned to a virtual computing acceleration device.

[0046] In one example, assuming that the FPGA implements a total of 15 network PFs (physical) and 1 computing acceleration PF (physical), each device can be expanded to 15 VFs (virtual) through SRIOV, each network device has 4 virtual queues, and each computing acceleration device has 2 virtual queues. At this time, there can be 992 virtual queues. In this application, once the above configuration parameters are determined, the queue number assigned to the device is determined. Therefore, the processing logic of the host computer driver and queue management module can determine all the queue numbers corresponding to the device based on the device number (BDF) and formulate a load distribution plan among multiple queues according to business needs; in the process of data transfer from FPGA to the host computer, the device number and device type can also be determined based on the queue number to complete the sideband signal information supplement of the upload data transmission request.

[0047] like Figure 4 As shown, an embodiment of the present application provides a heterogeneous acceleration system, which includes: The heterogeneous acceleration system includes a host computer and a heterogeneous accelerator. The host computer is provided with a virtual machine, and the heterogeneous accelerator is provided with a management module, a virtual network device, a virtual computing acceleration device, a network unit, an acceleration unit, and a memory unit. The first end of the management module is communicatively connected to the virtual machine, the second end of the management module is communicatively connected to the first end of the virtual network device and the first end of the virtual computing acceleration device, the second end of the virtual network device is communicatively connected to the first end of the network unit, the second end of the virtual computing acceleration device is communicatively connected to the second end of the acceleration unit, and the memory unit is communicatively connected to the management module.

[0048] The host computer of the heterogeneous acceleration system is also equipped with a resource management module (computing resource management and network resource management shown in the figure). The heterogeneous accelerator is also equipped with a physical network device, a physical computing acceleration device and an Ethernet network. The resource management module communicates with the physical network device (network PF) and the physical computing acceleration device (computing acceleration PF) respectively. The physical network device provides network service functions to the resource management module through a virtual queue. The physical computing acceleration device provides computing acceleration service functions to the resource management module through a virtual queue.

[0049] The host computer in this application includes a physical management layer (resource management module) and a virtualization management layer (virtual machines / containers), and the (virtual machines / containers) run in a virtualized environment on the host side.

[0050] Heterogeneous accelerators can be FPGA accelerators. FPGA accelerators leverage the parallel computing and reconfigurability of FPGAs (field programmable gate arrays) to provide hardware acceleration for specific tasks, such as deep learning and neural networks. FPGA accelerators achieve acceleration by converting certain computing tasks into hardware circuits, offering strong programmability, high flexibility, and excellent performance.

[0051] The virtual network device can be a virtual network card (VF), a virtual bridge, or other virtual network devices. This application uses the virtual network card as an example for description. The virtual computing acceleration device can be a virtual GPU, a virtual NPU, and so on.

[0052] This application incorporates a physical network card (NIC) and a network unit (NU) within the heterogeneous accelerator. The NU logically supports protocol processing, while the physical NIC converts data signals into electrical / optical signals. By decoupling the physical NIC and NU, the NU handles protocol logic while the NU handles signal conversion. The two are connected via standard interfaces (such as SGMII and XAUI), facilitating independent upgrades.

[0053] Physical computing accelerators support the actual execution of AI inference and other computing tasks at the physical execution layer. Accelerator units, on the control layer, convert command frames into accelerator control signals. This setup allows the accelerator unit to handle general control logic (such as memory address mapping), while the physical accelerator focuses on specific algorithms. The same set of control logic can be used to interface with different accelerators.

[0054] This application also includes PF (physical function) devices, which are physical compute accelerators and physical network interface cards (NICs) that also have virtual queues. Therefore, they can provide network or compute acceleration services to the host or virtual machine administrator, but not to virtual machine or container users. PF devices primarily implement compute or network resource management functions, such as supporting the expansion of the PCIe base address register space and the implementation of SRIOV.

[0055] The heterogeneous accelerator includes an Ethernet, a first end of the Ethernet is communicatively connected to the network unit, and a second end of the Ethernet is communicatively connected to the external network; in response to the Ethernet receiving a network downlink data packet transmitted by the network unit, the network downlink data packet is transmitted to the external network; in response to the Ethernet receiving a network uplink data packet transmitted by the external network, the network uplink data packet is transmitted to the network unit.

[0056] In this application, the FPGA hardware resources allocated to virtual machines, or container users on the host computer are presented in the form of VFs (virtual functions) and virtual queues. Virtual functions are categorized as network devices and compute accelerators. Each virtual machine or container user can be assigned one or more VFs, and each network device or compute accelerator can be assigned one or more pairs of virtual queues. The number of virtual queue pairs allocated determines the bandwidth of the network device and the performance of the compute accelerator. The FPGA hardware resources here can include logic resources such as programmable logic power supplies, on-chip memory, and DSPs; interface resources such as high-speed serial network ports (PCIE, Ethernet) and memory controllers; and resources such as routing channels and clock networks.

[0057] In this application, PCIE SR-IOV virtualization technology is used to abstract the physical queue of the FPGA into a virtual queue, which is directly managed by the virtual machine, bypassing the host overhead. Each virtual network device and virtual computing acceleration device in this application has at least one receive unidirectional queue (RXQ) and one transmit unidirectional queue (TXQ).

[0058] FPGA accelerators (heterogeneous accelerators), network devices or compute acceleration devices (VFs), and virtual machine VFs (VFs) logically communicate directly through virtual queues. The two devices share data paths during the virtual queue and descriptor stages. After data transfer in the virtual queues is completed, user data is arbitrated to the network and compute acceleration branches based on sideband signals. The network branch completes standard operations such as packet verification, packet padding, TSO fragmentation, and checksums in the network subsystem before connecting to the external network via an Ethernet port. The compute acceleration branch uses the compute acceleration data transmission protocol to control the compute acceleration unit and access the memory subsystem.

[0059] Specifically, in response to the virtual machine receiving a network and / or accelerated delivery instruction, a network and / or accelerated delivery descriptor is generated based on the network delivery and / or accelerated delivery data packet information, in response to the management module obtaining the network and / or accelerated delivery descriptor, a network delivery and / or accelerated delivery data packet is initiated and obtained from the memory unit based on a direct storage access request, the target network delivery and / or accelerated delivery virtual queue is determined, and the network delivery and / or accelerated delivery data packet is transmitted to the virtual network device and / or virtual computing acceleration device through the target network delivery and / or accelerated delivery virtual queue, and the virtual network device and / or virtual computing acceleration device transmits the network delivery and / or accelerated delivery data packet to the network unit and / or acceleration unit.

[0060] In response to the network unit and / or acceleration unit receiving a network and / or acceleration upload instruction, the network and / or acceleration upload instruction carries a network and / or computing acceleration upload data packet, the management module selects a network and / or acceleration upload descriptor from a pre-cached upload descriptor set, initiates and writes the network and / or computing acceleration upload data packet to the memory address corresponding to the network and / or acceleration upload descriptor based on a direct storage write request, and in response to writing the network and / or computing acceleration upload data packet to the memory address corresponding to the network and / or acceleration upload descriptor, triggers a network and / or acceleration interrupt instruction, and sends the network and / or acceleration interrupt instruction to the virtual machine. In response to the virtual machine receiving the network and / or acceleration interrupt instruction, the virtual machine reads the network and / or acceleration upload data from the memory address corresponding to the network and / or acceleration upload descriptor.

[0061] like Figure 5 As shown, the management module in the present application also includes a high-speed serial computer expansion bus hard core, an adapter module and a virtual queue management module; the first end of the high-speed serial computer expansion bus hard core is communicatively connected to the virtual machine, the second end of the high-speed serial computer expansion bus hard core is communicatively connected to the first end of the adapter module, the second end of the adapter module is communicatively connected to the first end of the virtual queue management module, and the second end of the virtual queue management module is communicatively connected to the first end of the virtual network device and the first end of the virtual computing acceleration device respectively, wherein the high-speed serial computer expansion bus hard core includes multiple data interfaces, and the multiple data interfaces correspond to multiple data channels.

[0062] In response to the high-speed serial computer expansion bus hard core obtaining the network and / or accelerated delivery descriptor, a network delivery and / or accelerated delivery data packet is initiated and obtained from the memory unit based on a direct storage access request. In response to the adaptation module obtaining the network and / or accelerated delivery data packet, the adaptation module parses the network and / or accelerated delivery data packet into a network and / or accelerated delivery data packet in a target format, and the virtual queue management module determines the target delivery network and / or acceleration virtual queue.

[0063] In response to the high-speed serial computer expansion bus hard core obtaining the network and / or acceleration upload descriptor, a network and / or computing acceleration upload data packet is initiated and written to the memory address corresponding to the network and / or acceleration upload descriptor based on a direct storage write request. In response to the adaptation module obtaining the network and / or computing acceleration upload data packet, the network and / or computing acceleration upload data packet is parsed into a network and / or computing acceleration upload data packet in a target format, and the virtual queue management module determines the target upload network and / or acceleration virtual queue.

[0064] The AMD platform's PCIE core (serial computer expansion bus core) provides a TLP data interface encapsulated with four channels: CC, CQ, RC, and RQ. The TLP frame is parsed and adapted by the adapter module. The TLP information here includes both the service payload and PCIE-related register configuration information. This module intercepts access information from the host computer's device configuration space and extended configuration space, and communicates the FPGA's device topology information to the PCIE bus, including the number of network devices and computing acceleration devices, SRIOV expansion capabilities, and the size of the base address register space. For service information from both devices, the PCIE adapter module performs protocol parsing, flag distribution, cache reorganization, and other operations, receiving or sending data packets using the AXIS protocol. The virtual queue management module caches and maintains the status information of all virtual queues.

[0065] The virtual queues used by the virtual network devices and virtual computing acceleration devices in this application are uniformly numbered, specifically using odd and even numbers to represent send queues (network downlink and / or accelerated downlink virtual queues) and receive queues (network upload and / or accelerated upload virtual queues), respectively. A queue is a circular queue composed of descriptors, which are the basic units for completing DMA data transfers and contain the address of the data payload, data size, descriptor type, and other additional information. This application uses unidirectional send and receive queues, meaning that only send descriptors exist in the send queue and only receive descriptors exist in the receive queue. The send descriptor points to the address and length information of the data packet from the host computer to the FPGA (heterogeneous accelerator). After obtaining this descriptor, the FPGA will issue a data read request to the address pointed to by this descriptor, completing the data transfer operation from the host computer memory to the FPGA device. The receive descriptor points to an empty memory block in the host computer. After obtaining this descriptor, the FPGA will initiate a data write request to the address pointed to by this descriptor, completing the data transfer operation from the FPGA to the host computer memory.

[0066] In the upload data path of the virtual queue management module, in order to simultaneously achieve the low latency requirements of the network device and the real-time handshake requirements of the computing acceleration device, this application adopts a pre-fetch mechanism for the descriptors of the receiving virtual queue: a certain size of cache is opened for all receiving queues inside the FPGA, and the driver is used to pre-fill the available descriptor queue, and the receiving descriptor is cached to the FPGA in advance before the upload service arrives. In this way, after the uploaded network data packet from the network device or the response packet from the computing acceleration device arrives at the virtual queue management module, the upload DMA transfer operation can be completed immediately using the locally cached descriptor. The larger the cache, the higher the ability to cope with burst traffic and the greater the resource overhead. In order to further improve the efficiency of descriptor use, chained and indirect descriptor technologies can also be used separately in the receiving queue.

[0067] See also Figure 6 The heterogeneous accelerator in this application is also provided with a protocol processing module, which includes a command parsing module, a response generation module, a data control module, an interface conversion module, an acceleration unit management module and an interrupt management module.

[0068] The first end of the command parsing module is communicatively connected to the second end of the virtual computing acceleration device, the second end of the command parsing module is communicatively connected to the first end of the data control module, the second end of the data control module is communicatively connected to the first end of the interface conversion module, the first end of the data controller module is communicatively connected to the first end of the response generation module, the second end of the data control module is communicatively connected to the first end of the acceleration management unit, and the second end of the acceleration management unit is communicatively connected to the first end of the interrupt management module.

[0069] Unlike network devices that only focus on the transmission bandwidth of a single queue and the independent operation of the transceiver queues without affecting each other, the computing acceleration device needs to respond to the host computer after executing parameter transfer, register reading and writing, or memory reading, and the transceiver queues need to cooperate. Therefore, this application designs the above-mentioned data transmission protocol, and completes the protocol parsing and processing, response frame generation, interrupt management and other functions after the virtual queue management module outputs the data load of the computing acceleration device. The command parsing module will parse the control information in the computing acceleration data transmission protocol frame header and identify whether the frame type is a memory write command frame, a memory read command frame, an acceleration unit status register read and write frame, or an acceleration unit control register read and write frame. After the identification is completed, the subsequent modules will perform the corresponding processing. The response generation module creates a response frame and specifies the receiving queue corresponding to the request to receive it. The data handling and unit control module completes the collection of memory read and write command information and acceleration unit control register read and write information and hands it over to the interface conversion module for processing. The acceleration unit status register read and write command is routed to the acceleration status unit management module. The acceleration status unit management module completes the computing acceleration device information maintenance, acceleration unit binding, and interrupt vector allocation. The interrupt management module is responsible for collecting interrupt requests from the acceleration unit, querying the interrupt vector information in the status management module and initiating DMA interrupt requests to the PCIE adapter module.

[0070] In this way, the acceleration unit status register read and write commands are distinguished from the memory write command frame, memory read command frame, and acceleration unit control register read and write frame and executed in different modules, separating the status register (read-only / low-frequency access) and the control register (write-only / critical configuration). This avoids read and write conflicts and prevents memory bandwidth-intensive operations from blocking each other with register control operations, which can significantly improve the performance, security, and maintainability of FPGA heterogeneous systems.

[0071] In one embodiment, the storage space of each acceleration unit can be mapped to a unified address space through a hardware-level interconnect protocol (such as PCIE / CXL / NoC), forming a virtualized shared memory pool. This allows the storage space of other computing acceleration units within a heterogeneous accelerator to be shared. After the computing acceleration device completes computation in the shared memory pool, it triggers the hardware message construction module, directly reads the computation parameters from the shared memory pool, fills the message header (source / destination MAC / IP address, port number) according to a predefined protocol format, such as RoCEv2 / UDP / TCP, and constructs a network message. The message construction module is directly connected to the Ethernet port of the network card device through a hardware queue, bypassing the operating system protocol stack. This enables the computation parameters of the computing acceleration device to be widely shared directly through the virtualized network device. This leverages hardware to implement data storage and forwarding, which was previously performed by software, further improving the computational efficiency of the virtualized FPGA device.

[0072] In this application, by using one-way receiving and sending queues to simultaneously meet the high bandwidth and low latency requirements of network card virtual devices and the burst bandwidth and reliable handshake communication requirements of computing acceleration devices, the two devices share the descriptor queue management logic and queue-level data paths, realizing an integrated design with low resource overhead for FPGA virtual devices; a computing acceleration data transmission protocol is used to complete a reliable handshake mechanism based on one-way queues to meet the burst bandwidth requirements and interactive communication requirements of computing acceleration devices. The virtual queue sharing allocation mechanism set up in this application allows the device quantity, expansion capabilities, queue resources, etc. of the two devices to be freely allocated during the design phase, and with the help of SRIOV, direct mapping of FPGA resources to virtual users can be achieved.

[0073] The present application also provides a data sending device, which is applied to a heterogeneous accelerator. The data sending device includes: a first data sending unit and a second data sending unit.

[0074] The first data sending unit is used to receive a first type of sending instruction sent by an upper computer, where the first type of sending instruction is a network sending instruction, parse the first type of sending instruction, obtain a first sending descriptor, initiate a first direct storage access request according to the first sending descriptor, obtain a first sending data packet from a storage area, determine a first sending virtual queue based on the first sending descriptor, and transmit the first sending data packet to the first device through the first sending virtual queue.

[0075] The second data sending unit is used to receive a second type of sending instruction sent by the upper computer, where the second type of sending instruction is a computing acceleration sending instruction, parse the second type of sending instruction, obtain a second sending descriptor, initiate a second direct storage access request based on the second sending descriptor, obtain a second sending data packet from the storage area, store the second sending data packet into the cache, perform protocol frame parsing on the second sending data packet, obtain the protocol frame type, transmit the second sending data packet to the target processing module of the second device for processing according to the protocol frame type, generate a response frame in response to the completion of the processing of the second sending data packet, and return the response frame to the upper computer.

[0076] The present application also provides a data upload device, which is applied to a heterogeneous accelerator. The data upload device includes: a first data upload unit and a second data upload unit.

[0077] The first data upload unit is used to receive a first type of upload instruction sent by a first device, where the first type of upload instruction is a network upload instruction and carries a first upload data packet, select a first upload descriptor from a pre-cached upload descriptor set, and write the first upload data packet to a first storage address corresponding to the first upload descriptor according to the first upload descriptor, so as to transmit the first upload data packet to a host computer based on the first storage address.

[0078] The second data upload unit is used to receive a second type of upload instruction sent by the second device, the second type of upload instruction is a computing acceleration upload instruction, the second type of upload instruction carries a second upload data packet, selects a second upload descriptor from a pre-cached upload descriptor set, and writes the second upload data packet to a second storage address corresponding to the second upload descriptor according to the second upload descriptor, so as to transmit the second upload data packet to the host computer based on the second storage address.

[0079] The description of the features in the embodiments corresponding to the data sending device and the data uploading device can be found in the relevant description of the embodiment corresponding to the data transmission method, and will not be repeated here.

[0080] The embodiment of the present application also provides a computer device, such as Figure 7 As shown, it includes a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the steps of any of the above-mentioned data sending methods or the steps of the data uploading method in any of the following embodiments.

[0081] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned data sending methods or the steps of the data uploading method in any of the following embodiments when running.

[0082] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0083] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0084] The above is a detailed introduction to a data transmission method provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method and core ideas of the present application. It should be noted that, for those skilled in the art, without departing from the principles of the present application, several improvements and modifications may be made to the present application, and such improvements and modifications also fall within the scope of protection of the present application.

Claims

1. A data delivery method, characterized in that: Applied to heterogeneous accelerators, the data delivery method includes: In response to receiving a first type of delivery instruction sent by a host computer, the first type of delivery instruction being a network delivery instruction, parsing the first type of delivery instruction to obtain a first delivery descriptor, initiating a first direct storage access request according to the first delivery descriptor, obtaining a first delivery data packet from a storage area, determining a first delivery virtual queue based on the first delivery descriptor, and transmitting the first delivery data packet to the first device through the first delivery virtual queue; In response to receiving a second type of sending instruction sent by the upper computer, the second type of sending instruction is a computing acceleration sending instruction, the second type of sending instruction is parsed, a second sending descriptor is obtained, a second direct storage access request is initiated according to the second sending descriptor, a second sending data packet is obtained from the storage area, the second sending data packet is stored in the cache, a protocol frame is parsed for the second sending data packet, a protocol frame type is obtained, and the second sending data packet is transmitted to the target processing module of the second device for processing according to the protocol frame type. In response to the completion of the processing of the second sending data packet, a response frame is generated and the response frame is returned to the upper computer.

2. The data sending method according to claim 1, characterized in that: The initiating a first direct storage access request according to the first delivery descriptor, obtaining a first delivery data packet from a storage area, determining a first delivery virtual queue based on the first delivery descriptor, and transmitting the first delivery data packet to the first device through the first delivery virtual queue includes: In response to obtaining the first delivery descriptor, initiating a first direct storage access request, and obtaining the first delivery data packet from the first delivery data packet storage address based on the first direct storage access request, wherein the first delivery descriptor includes the first delivery data packet storage address and the first delivery data packet length; In response to obtaining the first delivered data packet, converting the first delivered data packet into a first delivered data packet in a target format; determining a first delivery virtual queue based on the first delivery descriptor, and transmitting a first delivery data packet in a target format to a first device based on the first delivery virtual queue; In response to the first device receiving the first delivered data packet in the target format, verifying, fragmenting, and encapsulating the first delivered data packet in the target format to obtain a processed first delivered data packet; The processed first sent data packet is sent to an external network via Ethernet.

3. The data sending method according to claim 1, characterized in that: The initiating a second direct storage access request according to the second delivery descriptor, obtaining a second delivery data packet from the storage area, storing the second delivery data packet in a cache, performing protocol frame parsing on the second delivery data packet to obtain a protocol frame type, transmitting the second delivery data packet to a target processing module of the second device for processing according to the protocol frame type, generating a response frame in response to completion of processing the second delivery data packet, and returning the response frame to the host computer includes: In response to obtaining the second delivery descriptor, initiating a second direct storage access request, and obtaining the second delivery data packet from the second delivery data packet storage address based on the second direct storage access request, wherein the second delivery descriptor includes the second delivery data packet storage address and the second delivery data packet length; In response to obtaining the second delivered data packet, storing the second delivered data packet in a cache; Performing protocol frame parsing on the second delivered data packet to obtain a protocol frame type, where the protocol frame type includes a memory write command frame, a memory read command frame, a status register read / write frame, and a control register read / write frame; Using the first processing module of the second device to process the second sent data packet corresponding to the memory write command frame, the memory read command frame, and the control register read / write frame; using the second processing module of the second device to process the second sent data packet corresponding to the status register read / write frame; In response to the completion of processing the second transmitted data packet, a response frame is generated and returned to the host computer.

4. A data uploading method, characterized in that: Applied to heterogeneous accelerators, the data uploading method includes: In response to receiving a first-type upload instruction sent by a first device, the first-type upload instruction is a network upload instruction, the first-type upload instruction carries a first upload data packet, selecting a first upload descriptor from a pre-cached upload descriptor set, writing the first upload data packet to a first storage address corresponding to the first upload descriptor according to the first upload descriptor, and transmitting the first upload data packet to a host computer based on the first storage address; In response to receiving a second type of upload instruction sent by a second device, the second type of upload instruction is a computing acceleration upload instruction, the second type of upload instruction carries a second upload data packet, a second upload descriptor is selected from a pre-cached upload descriptor set, and the second upload data packet is written to a second storage address corresponding to the second upload descriptor according to the second upload descriptor, so as to transmit the second upload data packet to the host computer based on the second storage address.

5. The data uploading method according to claim 4, characterized in that: Writing the first upload data packet into a first storage address corresponding to the first upload descriptor according to the first upload descriptor, so as to transmit the first upload data packet to the host computer based on the first storage address, includes: In response to the completion of writing the first upload data packet into the first storage address corresponding to the first upload descriptor, a first interrupt instruction is triggered and the first interrupt instruction is sent to the host computer. In response to the host computer receiving the first interrupt instruction, the host computer reads the first upload data packet based on the first storage address.

6. The data uploading method according to claim 4, characterized in that: Writing the second upload data packet into a second storage address corresponding to the second upload descriptor according to the second upload descriptor, so as to transmit the second upload data packet to the host computer based on the second storage address, includes: In response to the completion of writing the second upload data packet into the second storage address corresponding to the second upload descriptor, a second interrupt instruction is triggered and sent to the host computer. In response to the host computer receiving the second interrupt instruction, the host computer reads the second upload data packet based on the second storage address.

7. A data sending device, characterized in that: Applied to a heterogeneous accelerator, the data sending device includes: a first data sending unit, configured to receive a first type of sending instruction sent by a host computer, the first type of sending instruction being a network sending instruction, parse the first type of sending instruction, obtain a first sending descriptor, initiate a first direct storage access request based on the first sending descriptor, obtain a first sending data packet from a storage area, determine a first sending virtual queue based on the first sending descriptor, and transmit the first sending data packet to the first device through the first sending virtual queue; The second data sending unit is used to receive a second type of sending instruction sent by the upper computer, where the second type of sending instruction is a computing acceleration sending instruction, parse the second type of sending instruction, obtain a second sending descriptor, initiate a second direct storage access request according to the second sending descriptor, obtain a second sending data packet from the storage area, store the second sending data packet into the cache, perform protocol frame parsing on the second sending data packet, obtain the protocol frame type, transmit the second sending data packet to the target processing module of the second device for processing according to the protocol frame type, generate a response frame in response to the completion of the processing of the second sending data packet, and return the response frame to the upper computer.

8. A data uploading device, characterized in that: Applied to a heterogeneous accelerator, the data uploading device includes: a first data upload unit, configured to receive a first-type upload instruction sent by a first device, where the first-type upload instruction is a network upload instruction and carries a first upload data packet; select a first upload descriptor from a pre-cached upload descriptor set; write the first upload data packet to a first storage address corresponding to the first upload descriptor according to the first upload descriptor, and transmit the first upload data packet to a host computer based on the first storage address; A second data upload unit is used to receive a second type of upload instruction sent by a second device, where the second type of upload instruction is a computing acceleration upload instruction, and the second type of upload instruction carries a second upload data packet. A second upload descriptor is selected from a pre-cached upload descriptor set, and the second upload data packet is written to a second storage address corresponding to the second upload descriptor according to the second upload descriptor, so as to transmit the second upload data packet to the host computer based on the second storage address.

9. A computer device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the data sending method according to any one of claims 1 to 3 or the steps of the data uploading method according to any one of claims 4 to 6 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the data sending method according to any one of claims 1 to 3 or the steps of the data uploading method according to any one of claims 4 to 6 are implemented.

Citation Information

Patent Citations

  • Ethernet frame issuing method, Ethernet frame uploading method and related devices

    CN114925012A

  • Configurable device interface

    CN115113973A

  • Accelerator processing method and device, storage medium and processor

    CN115373810A

  • Data transmission method and device, accelerator equipment, host and storage medium

    CN117312201A

  • Configurable device interface

    US20210232528A1

Cited By

  • Virtual network card and electronic equipment

    CN121210030A