Data transmission method and apparatus, accelerator device, and host
By receiving and storing application data packets in the accelerator device and writing their stored information into the descriptor, the problem of waste of resources in direct memory access transmission is solved, and efficient and real-time data transmission is achieved.
Patent Information
- Application Number
- PCT/CN2024/133312
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-11-20
- Publication Date
- 2025-06-05
AI Technical Summary
In application scenarios such as network transmission and image transmission, direct memory access transmission cannot predict the amount of data transmission in advance, resulting in the need to apply for a large enough memory space in advance, resulting in waste of resources, especially under embedded platforms.
By receiving and storing the application data packets in memory, recording their stored information, including the packet size and storage address, writing this information into the descriptor, causing the host to apply for memory based on the descriptor information, and starting direct memory access transmission.
It avoids waste of resources during direct memory access transmission, reduces the demand for CPU performance, and ensures real-time data transmission.
Smart Images

Figure CN2024133312_05062025_PF_FP_ABST
Abstract
Description
Data transmission method, device, accelerator device, and host
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on November 29, 2023, with application number 202311608717.X, and application name “A data transmission method, device and accelerator device, host and storage medium”, all contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of computer technology, and in particular, to a data transmission method, apparatus, accelerator device, host, and non-volatile readable storage medium. Background Art
[0004] DMA (Direct Memory Access) has become the most common method of large-scale data transmission due to its performance advantages such as high transmission bandwidth and no need for CPU (Central Processing Unit) participation. Normally, the host side (Host) applies for memory in advance and configures the DMA transfer descriptor, which contains information such as data transmission length, source address, and destination address. The DMA controller initiates data transmission based on the descriptor information without the participation of the CPU, freeing up the CPU. However, in application scenarios such as network transmission and image transmission, since the data transmission volume cannot be predicted in advance, it is necessary to apply for a sufficiently large memory space on the host side in advance to prevent data packet loss, resulting in a waste of resources. Especially in embedded platforms, due to the limited memory resources themselves, excessive application of memory space may cause abnormalities in other functions.
[0005] Therefore, how to avoid resource waste during direct memory access transmission is a technical problem that those skilled in the art need to solve. Summary of the Invention
[0006] The purpose of the present application is to provide a data transmission method, device, electronic device and non-volatile readable storage medium, which avoids the waste of resources in the direct memory access transmission process.
[0007] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present application provides a data transmission method, which is applied to an accelerator device, the accelerator device being connected to a host, and the method comprising:
[0008] receiving an application data packet and storing the application data packet in a memory;
[0009] Recording storage information of the application data packet; wherein the storage information includes the size of the application data packet and the storage address of the application data packet in the memory;
[0010] Write the storage information of the application data packet into the descriptor so that the host can apply for memory according to the storage information in the descriptor and fill the applied memory address into the descriptor to start direct memory access transmission;
[0011] The application data packets in the memory are transferred to the host via direct memory access transfer according to the memory address in the descriptor filled by the host.
[0012] The step of receiving an application data packet and storing the application data packet in a memory includes:
[0013] Receive application data packets;
[0014] Preprocessing the application data packet; wherein the preprocessing includes any one or a combination of analog-to-digital conversion, data filtering, and image decompression;
[0015] The pre-processed application data packets are stored in the memory.
[0016] Receiving an application data packet includes:
[0017] Receive application data packets through multiple data channels;
[0018] Accordingly, the application data packet is pre-processed, including:
[0019] The application data packets received by multiple data channels are preprocessed in parallel by multiple data preprocessing modules.
[0020] The storage information of the application data packet is written into the descriptor so that the host can apply for memory according to the storage information in the descriptor and fill the applied memory address into the descriptor to start direct memory access transmission, including:
[0021] The target descriptor is determined in the descriptors whose status is idle, the storage information of the application data packet is written into the target descriptor, and the status of the target descriptor is modified to occupied, so that the host can apply for memory according to the size of the application data packet in the target descriptor, and fill the applied memory address into the target descriptor to start direct memory access transmission.
[0022] After the application data packet transmission is completed, the host modifies the state of the target descriptor to idle.
[0023] After the state of the target descriptor is changed to occupied, the following steps are further included:
[0024] The descriptor count is updated and transmitted to the host; wherein the descriptor count is used to describe the number of descriptors in the occupied state.
[0025] The host cyclically queries the descriptor count, and if the descriptor count is not zero, reads the target descriptor and executes the step of applying for memory according to the size of the application data packet in the target descriptor;
[0026] After the host changes the state of the target descriptor to idle, it updates the descriptor count according to the number of descriptors that are currently occupied.
[0027] The storage information of the application data packet is written into the descriptor, including:
[0028] After each application data packet is received and each application data packet is stored in the memory, the storage information of the application data packet is written into the descriptor.
[0029] The memory is divided into multiple memory blocks, and application data packets are stored in the memory, including:
[0030] storing the application data packets sequentially into the storage blocks in the memory;
[0031] Accordingly, the storage information of the application data packet is written into the descriptor, including:
[0032] After each storage block is filled, the storage information of the application data packet in the filled storage block is written into the descriptor; wherein the storage information of each application data packet corresponds to one descriptor.
[0033] The accelerator device includes a field programmable logic gate array and a memory. The field programmable logic gate array includes a data cache module, a data transmission control module, a storage controller, a read / write module, and a direct memory access module.
[0034] The data cache module is configured to cache received application data packets;
[0035] The storage controller is configured to store the cached application data packets into the memory;
[0036] The data transmission control module is configured to record the storage information of the application data packet and write the storage information of the application data packet into the descriptor through the read-write module;
[0037] The read / write module is configured to transmit information to and from the host computer;
[0038] The direct memory access module is configured to transfer application data packets in the memory to the host according to the memory addresses in the descriptors filled in by the host.
[0039] The field programmable logic gate array further includes a data preprocessing module, which is configured to preprocess received application data packets.
[0040] The accelerator device further includes a data preprocessing module independent of the field programmable logic gate array. The data preprocessing module is connected to the field programmable logic gate array and is configured to preprocess received application data packets.
[0041] The memory is a double rate synchronous dynamic random access memory.
[0042] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a data transmission device, which is applied to an accelerator device connected to a host, and includes:
[0043] a storage module, configured to receive application data packets and store the application data packets in a memory;
[0044] a recording module configured to record storage information of the application data packet; wherein the storage information includes the size of the application data packet and the storage address of the application data packet in the memory;
[0045] a write module configured to write storage information of an application data packet into a descriptor so that the host can apply for memory according to the storage information in the descriptor and fill the applied memory address into the descriptor to start direct memory access transmission;
[0046] The transmission module is configured to transmit the application data packet in the memory to the host through direct memory access transmission according to the memory address in the descriptor filled by the host.
[0047] To achieve the above-mentioned purpose, a third aspect of the embodiments of the present application provides an accelerator device, including:
[0048] a memory arranged to store a computer program;
[0049] The processor is configured to implement the steps of the above-mentioned data transmission method when executing the computer program.
[0050] To achieve the above objectives, the present application provides a computer non-volatile readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above data transmission method are implemented.
[0051] To achieve the above-mentioned purpose, a fourth aspect of an embodiment of the present application provides a data transmission method, which is applied to a host, the host being connected to an accelerator device, and the method includes:
[0052] Obtaining a descriptor and applying for memory based on storage information in the descriptor; wherein the storage information includes the size of the application data packet to be transmitted and the storage address of the application data packet in the memory of the accelerator device;
[0053] Fill the requested memory address into the descriptor to start direct memory access transfer;
[0054] Receive application data packets sent by the accelerator device.
[0055] Before obtaining the descriptor, the following steps are also included:
[0056] The descriptor count is queried. If the descriptor count is not zero, the steps of obtaining the descriptor and applying for memory according to the storage information in the descriptor are executed.
[0057] Before obtaining the descriptor, the following steps are also included:
[0058] Request descriptor storage space through direct memory access consistent memory.
[0059] To achieve the above objectives, the present application provides a data transmission device, which is applied to a host computer, the host computer is connected to an accelerator device, and the device includes:
[0060] a first application module configured to obtain a descriptor and apply for memory according to storage information in the descriptor; wherein the storage information includes the size of the application data packet to be transmitted and the storage address of the application data packet in the memory of the accelerator device;
[0061] The startup module is configured to fill the requested memory address into the descriptor to start direct memory access transfer;
[0062] The receiving module is configured to receive application data packets sent by the accelerator device.
[0063] To achieve the above-mentioned purpose, a fifth aspect of an embodiment of the present application provides a host, including:
[0064] a memory arranged to store a computer program;
[0065] The processor is configured to implement the steps of the above-mentioned data transmission method when executing the computer program.
[0066] To achieve the above-mentioned purpose, the sixth aspect of the embodiment of the present application provides a computer non-volatile readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned data transmission method are implemented.
[0067] To achieve the above-mentioned purpose, a seventh aspect of an embodiment of the present application provides a data transmission system, including the above-mentioned accelerator device and the above-mentioned host, wherein the host is connected to the accelerator device.
[0068] From the above scheme, it can be seen that a data transmission method provided by the present application includes: an accelerator device receives an application data packet and stores the application data packet in a memory; records the storage information of the application data packet; wherein the storage information includes the size of the application data packet and the storage address of the application data packet in the memory; writes the storage information of the application data packet into a descriptor, so that the host applies for memory according to the storage information in the descriptor, and fills the applied memory address into the descriptor to start direct memory access transmission; transmits the application data packet in the memory to the host through direct memory access transmission according to the memory address in the descriptor filled by the host.
[0069] The data transmission method provided by the present application is that the accelerator device writes the storage information of the application data packet to be transmitted into the descriptor, and the DMA transmission is initiated by the host. The corresponding memory space is requested according to the storage information in the descriptor. There is no need to open up a large amount of memory space in advance, which reduces the demand for CPU performance and ensures the real-time performance of data transmission. It can be seen that the data transmission method provided by the present application avoids the waste of resources in the direct memory access transmission process. The present application also discloses a data transmission device and an accelerator device, a host and a computer non-volatile readable storage medium, which can also achieve the above technical effects.
[0070] It should be understood that the foregoing general description and the following detailed description are merely illustrative and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. The drawings are used to provide an understanding of the present disclosure and constitute part of the specification. Together with the following optional implementation methods, they are used to explain the present disclosure, but do not constitute a limitation of the present disclosure. In the drawings:
[0072] FIG1 is a flow chart showing a data transmission method according to an exemplary embodiment;
[0073] FIG2 is a schematic diagram showing a block structure of a memory according to an exemplary embodiment;
[0074] FIG3 is a structural diagram of a data transmission device according to an exemplary embodiment;
[0075] FIG4 is a structural diagram of an accelerator device according to an exemplary embodiment;
[0076] FIG5 is a structural diagram of another accelerator device according to an exemplary embodiment;
[0077] FIG6 is a structural diagram of another accelerator device according to an exemplary embodiment;
[0078] FIG7 is a flow chart showing another data transmission method according to an exemplary embodiment;
[0079] FIG8 is a structural diagram of a data transmission device according to an exemplary embodiment;
[0080] FIG9 is a structural diagram of a host according to an exemplary embodiment;
[0081] FIG10 is a structural diagram of a data transmission system according to an exemplary embodiment;
[0082] FIG11 is a structural diagram of another data transmission system according to an exemplary embodiment;
[0083] FIG12 is a flowchart of a data transmission method in an application embodiment provided by the present application. DETAILED DESCRIPTION
[0084] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. In addition, in the embodiments of the present application, "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0085] The embodiment of the present application discloses a data transmission method, which avoids resource waste in the direct memory access transmission process.
[0086] Referring to FIG1 , a flow chart of a data transmission method according to an exemplary embodiment is shown. As shown in FIG1 , the method includes:
[0087] S101: Receive an application data packet and store the application data packet in a memory;
[0088] The embodiments of the present application are executed by an accelerator device, such as an FPGA (Field Programmable Gate Array) accelerator, which is connected to a host computer. The embodiments of the present application can be applied to parallel application scenarios such as multi-channel analog-to-digital (A / D) conversion and multi-channel image input.
[0089] In actual implementation, the accelerator device receives the application data packet and stores it in a memory. The memory here can be DDR (Double Data Rate) or other types of memory, which are not limited here.
[0090] As an optional implementation, receiving an application data packet and storing the application data packet in a memory includes: receiving an application data packet; preprocessing the application data packet; wherein the preprocessing includes any one or a combination of any several of analog-to-digital conversion, data filtering, and image decompression; and storing the preprocessed application data packet in a memory.
[0091] In practice, after receiving application data packets, the accelerator device performs preprocessing, which can include analog-to-digital conversion, data filtering, image decompression, and other techniques. Preprocessing algorithms can be developed based on the specific application. The preprocessed application data packets are then stored in memory. This indicates that performing preprocessing of application data packets within the accelerator device can effectively reduce data processing pressure on the host CPU and improve the performance of the data transmission system.
[0092] As a feasible implementation, receiving an application data packet includes: receiving the application data packet via multiple data channels; and correspondingly, preprocessing the application data packet includes: performing parallel preprocessing on the application data packets received via the multiple data channels using multiple data preprocessing modules. In actual implementation, the accelerator device can receive the application data packets via the multiple data channels in parallel, and perform parallel preprocessing on the application data packets received via the multiple data channels using the multiple data preprocessing modules, thereby improving data processing efficiency.
[0093] As a feasible implementation, storing the application data packet in the memory includes: caching the application data packet and storing the cached application data packet in the memory. In actual implementation, the accelerator device first caches the pre-processed application data packet and then packages and writes it to the memory.
[0094] S102: Recording storage information of the application data packet; wherein the storage information includes the size of the application data packet and the storage address of the application data packet in the memory;
[0095] In actual implementation, the accelerator device records the storage information of the application data packet, including the size of the application data packet and its storage address in the memory.
[0096] S103: writing the storage information of the application data packet into the descriptor so that the host can apply for memory according to the storage information in the descriptor, and filling the applied memory address into the descriptor to start direct memory access transmission;
[0097] In actual implementation, the accelerator device writes the storage information for the application data packet into a descriptor prepared in advance by the host. A descriptor is a fixed-size block of host-side memory, prepared by the host before data is transferred. It can store up to 256 transfer messages. When the descriptor is empty, it is in an idle state; when filled with data, it is occupied. Each transfer occupies one block and can be reused. The host then requests memory based on the storage information in the descriptor and fills the requested memory address into the descriptor. The host then initiates the data transfer operation, essentially starting a direct memory access transfer.
[0098] As a feasible implementation, the storage information of the application data packet is written into a descriptor, so that the host can request memory based on the storage information in the descriptor and fill the requested memory address into the descriptor to initiate direct memory access transmission. This includes: determining a target descriptor in a descriptor with an idle status, writing the storage information of the application data packet into the target descriptor, and changing the status of the target descriptor to occupied. This allows the host to request memory based on the size of the application data packet in the target descriptor and fill the requested memory address into the target descriptor to initiate direct memory access transmission. After the application data packet transmission is completed, the host changes the status of the target descriptor to idle.
[0099] In actual implementation, the target descriptor is identified from the idle descriptors, the application data packet storage information is written to the target descriptor, and the target descriptor status is changed to occupied. The host then requests memory based on the application data packet size in the target descriptor and fills the requested memory address into the target descriptor to initiate a direct memory access transfer. After the application data packet transfer is complete, the host changes the target descriptor status to idle.
[0100] As a feasible implementation, writing the storage information of the application data packet into the descriptor includes: writing the storage information of each application data packet into the descriptor after each application data packet is stored in the memory. In actual implementation, for applications with high real-time requirements, the storage information of each application data packet is written into the descriptor immediately after it is stored in the memory.
[0101] As another feasible implementation, the memory is divided into multiple storage blocks, and the storage information of the application data packet is written into the descriptor, including: after each storage block is filled, the storage information of the application data packet in the filled storage block is written into the descriptor. In actual implementation, for applications with low real-time requirements, the memory is divided into blocks, and after each data block is filled, the storage information of the application data packet therein is filled into the descriptor. For example, as shown in Figure 2, the memory space is divided into 8 data blocks 0-7, and they are filled in sequence starting from data block 0. After each data block is filled, the storage information of the application data packet therein is filled into the descriptor, informing the host side that data reading can be performed. In application scenarios where real-time requirements are not high, the host side is avoided from frequently initiating DMA data reading operations, thereby ensuring the stability of the data transmission system.
[0102] As an optional embodiment, after writing the storage information of the application data packet into the descriptor, the method further includes: updating a descriptor count and transmitting the descriptor count to the host; wherein the descriptor count is used to describe the number of descriptors in the occupied state. The host cyclically queries the descriptor count. If the descriptor count is not zero, the host reads the target descriptor and executes the step of allocating memory based on the size of the application data packet in the target descriptor. After the host changes the status of the target descriptor to idle, the host updates the descriptor count based on the number of descriptors in the current occupied state.
[0103] In actual implementation, the accelerator device counts the stored application data packets. Each time an application data packet is stored, its storage information will occupy a descriptor. Each time a descriptor is filled, the descriptor count is increased by one. That is, the stored application data packets are counted by counting the number of descriptors in the occupied state, and the descriptor count is uploaded to the host side to inform the host side of the number of descriptors that can currently be processed. The host cyclically queries the descriptor count. If the descriptor count is not zero, it applies for memory according to the storage information in the description information. After the application data packet transmission is completed, the host updates the descriptor count according to the number of descriptors in the current occupied state, that is, the descriptor count is reduced by one. It should be noted that while updating the descriptor count, the host side and the accelerator device side also need to update the head pointer and tail pointer of the descriptor.
[0104] S104: Transmit the application data packet in the memory to the host through direct memory access transmission according to the memory address in the descriptor filled in by the host.
[0105] In actual implementation, the accelerator device transmits the application data packet to the host through DMA transmission according to the memory address requested by the host. The host waits to receive the transmission completion interrupt and ends the operation.
[0106] In the data transmission method provided by the embodiment of the present application, the accelerator device writes the storage information of the application data packet to be transmitted into the descriptor. The DMA transmission is initiated by the host, and the corresponding memory space is requested based on the storage information in the descriptor. This eliminates the need to allocate a large amount of memory space in advance, reduces the demand for CPU performance, and ensures the real-time performance of data transmission. Therefore, the data transmission method provided by the embodiment of the present application avoids the waste of resources in the direct memory access transmission process.
[0107] A data transmission device provided in an embodiment of the present application is introduced below. The data transmission device described below and the data transmission method described above can be referenced to each other.
[0108] Referring to FIG3 , a structural diagram of a data transmission device according to an exemplary embodiment is shown. As shown in FIG3 , the device includes:
[0109] The storage module 101 is configured to receive application data packets and store the application data packets in a memory;
[0110] The recording module 102 is configured to record storage information of the application data packet; wherein the storage information includes the size of the application data packet and the storage address of the application data packet in the memory;
[0111] The writing module 103 is configured to write the storage information of the application data packet into the descriptor so that the host can apply for memory according to the storage information in the descriptor and fill the applied memory address into the descriptor to start direct memory access transmission;
[0112] The transmission module 104 is configured to transmit the application data packet in the memory to the host through direct memory access transmission according to the memory address in the descriptor filled by the host.
[0113] In the embodiment of the present application, the accelerator device writes the storage information of the application data packet to be transmitted into the descriptor. The DMA transmission is initiated by the host, and the corresponding memory space is requested based on the storage information in the descriptor. This eliminates the need to allocate a large amount of memory space in advance, reduces the demand for CPU performance, and ensures the real-time nature of data transmission. As can be seen, the embodiment of the present application avoids the waste of resources in the direct memory access transmission process.
[0114] Based on the above embodiment, as an optional implementation, the storage module 101 is configured to: receive an application data packet; preprocess the application data packet; wherein the preprocessing includes any one or a combination of any several of analog-to-digital conversion, data filtering, and image decompression; and store the preprocessed application data packet in the memory.
[0115] Based on the above embodiment, as an optional implementation, the storage module 101 is configured to: receive application data packets through multiple data channels; and perform parallel preprocessing on the application data packets received through the multiple data channels through multiple data preprocessing modules.
[0116] Based on the above embodiment, as an optional implementation, the write module 103 is configured to: determine the target descriptor in the descriptor whose status is idle, write the storage information of the application data packet into the target descriptor, and modify the status of the target descriptor to occupied, so that the host applies for memory according to the size of the application data packet in the target descriptor, and fills the applied memory address into the target descriptor to start direct memory access transmission.
[0117] Based on the above embodiment, as an optional implementation manner, after the application data packet transmission is completed, the host changes the state of the target descriptor to idle.
[0118] Based on the above embodiment, as an optional implementation, the writing module 103 is configured to write the storage information of the application data packet into the descriptor after each application data packet is received and each application data packet is stored in the memory.
[0119] Based on the above embodiment, as an optional implementation, the memory is divided into multiple storage blocks, and the storage module 101 is configured to: store the application data packets in the storage blocks in the memory in sequence; the writing module 103 is configured to: after each storage block is filled, write the storage information of the application data packet in the filled storage block into the descriptor; wherein the storage information of each application data packet corresponds to a descriptor.
[0120] Based on the above embodiment, as an optional implementation manner, the following is further included:
[0121] The descriptor count is updated and transmitted to the host; wherein the descriptor count is used to describe the number of descriptors in the occupied state.
[0122] Based on the above embodiment, as an optional implementation method, the host cyclically queries the descriptor count. If the descriptor count is not zero, the target descriptor is read and the step of applying for memory according to the size of the application data packet in the target descriptor is executed; after the host changes the status of the target descriptor to idle, the descriptor count is updated according to the number of descriptors that are currently occupied.
[0123] Based on the above embodiment, as an optional implementation, the accelerator device includes a field programmable logic gate array and a memory, and the field programmable logic gate array includes a data cache module, a data transmission control module, a storage controller, a read / write module, and a direct memory access module;
[0124] The data cache module is configured to cache received application data packets;
[0125] The storage controller is configured to store the cached application data packets into the memory;
[0126] The data transmission control module is configured to record the storage information of the application data packet and write the storage information of the application data packet into the descriptor through the read-write module;
[0127] The read / write module is configured to transmit information to and from the host computer;
[0128] The direct memory access module is configured to transfer application data packets in the memory to the host according to the memory addresses in the descriptors filled in by the host.
[0129] Based on the above embodiment, as an optional implementation, the field programmable gate array further includes a data preprocessing module, which is configured to preprocess the received application data packets.
[0130] Based on the above embodiment, as an optional implementation, the accelerator device further includes a data preprocessing module independent of the field programmable gate array, the data preprocessing module is connected to the field programmable gate array, and the data preprocessing module is configured to preprocess the received application data packets.
[0131] Based on the above embodiment, as an optional implementation, the memory is a double data rate synchronous dynamic random access memory.
[0132] Regarding the apparatus in the above embodiment, the manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.
[0133] The present application discloses an accelerator device, comprising:
[0134] a memory arranged to store a computer program;
[0135] The processor is configured to implement the steps of the above-mentioned data transmission method when executing the computer program.
[0136] As an optional embodiment, as shown in Figure 4, the accelerator device includes a field programmable logic gate array and a memory, and the field programmable logic gate array includes a data cache module, a data transmission control module, a storage controller, a read-write module, and a direct memory access module; the data cache module is configured to cache received application data packets; the storage controller is configured to store the cached application data packets in the memory; the data transmission control module is configured to record the storage information of the application data packets and write the storage information of the application data packets into the descriptor; the read-write module is configured to transmit information with the host; and the direct memory access module is configured to transmit the application data packets in the memory to the host according to the memory address in the descriptor filled in by the host.
[0137] In actual implementation, the data cache module is configured to cache the pre-processed data packets first, and the storage controller is configured to package the cached data and write it into the memory. The transmission control module is configured to record the size of the application data packet and the storage address in the memory, and write the information into the descriptor prepared in advance by the host. The transmission control module is also configured to count the data packets and upload the descriptor count value to the host side through the read-write module to inform the host side of the number of descriptors that can currently be processed, where the descriptors are used in a circular manner. The read-write module is responsible for the information transmission between the read-write module and the accelerator device. For example, the accelerator device can inform the host of the number of currently operable descriptors, that is, the descriptor count, through the read-write module. The direct memory access module is configured to receive the descriptors that have been filled in by the accelerator device, and then the host side applies for memory, fills in the descriptors, and initiates data transmission operations.
[0138] The memory in the embodiment of the present application may be DDR or other types of memory, which is not limited here.
[0139] Optionally, the accelerator device also includes a data preprocessing module configured to preprocess received application data packets. This module performs preprocessing on data sent from parallel applications such as multi-channel ADC (Analog-to-Digital Converter) conversion and image transmission, including ADC conversion, data filtering, and image decompression. The processing algorithm can be developed based on the actual application. Performing this parallel data processing within the FPGA accelerator effectively reduces CPU processing pressure and improves system performance.
[0140] As a feasible implementation, the FPGA further includes a data preprocessing module configured to preprocess received application data packets. In actual implementation, as shown in FIG5 , the data preprocessing module is located in the FPGA.
[0141] As another feasible implementation, the accelerator device also includes a data preprocessing module independent of the field programmable gate array (FPGA). The data preprocessing module is connected to the field programmable gate array and is configured to preprocess the received application data packets. In actual implementation, as shown in FIG6 , the data preprocessing module is independent of the field programmable gate array (FPGA). A dedicated scene-oriented device is selected to build a data receiving preprocessing module, such as RK3399, which is specialized in image processing. The dedicated data preprocessing module transmits the processed results to the field programmable gate array via interfaces such as PCIe (Peripheral Component Interconnect Express, a high-speed serial computer expansion bus standard) or SDIO (Secure Digital Input and Output). The field programmable gate array only performs data caching and forwarding. The dedicated data preprocessing module can effectively improve the efficiency of data preprocessing, reduce the difficulty of FPGA development, and improve the overall efficiency of the system.
[0142] In the embodiment of the present application, the accelerator device writes the storage information of the application data packet to be transmitted into the descriptor. The DMA transmission is initiated by the host, and the corresponding memory space is requested based on the storage information in the descriptor. This eliminates the need to allocate a large amount of memory space in advance, reduces the demand for CPU performance, and ensures the real-time nature of data transmission. As can be seen, the embodiment of the present application avoids the waste of resources in the direct memory access transmission process.
[0143] The embodiment of the present application discloses a data transmission method, which avoids resource waste in the direct memory access transmission process.
[0144] Referring to FIG. 7 , a flow chart of another data transmission method according to an exemplary embodiment is shown. As shown in FIG. 7 , the method includes:
[0145] S201: Obtain a descriptor and apply for memory according to storage information in the descriptor; wherein the storage information includes the size of the application data packet to be transmitted and the storage address of the application data packet in the memory of the accelerator device;
[0146] The embodiment of the present application is executed by a host computer, which is connected to an accelerator device. The embodiment of the present application can be applied to parallel application scenarios such as multi-channel AD conversion and multi-channel image input.
[0147] As a feasible implementation, before obtaining the descriptor, the method further includes: applying for descriptor storage space via direct memory access (DMA) to the consistent memory. In actual implementation, the host first applies for descriptor storage space via DMA to the consistent memory and starts the system.
[0148] In actual implementation, the accelerator device receives the application data packet, stores it in the memory, records the storage information of the application data packet, and writes it into the descriptor. The host requests memory based on the size of the application data packet to be transmitted in the storage information in the descriptor.
[0149] As an optional implementation, before obtaining the descriptor, the method further includes: querying the descriptor count; if the descriptor count is not zero, obtaining the descriptor and applying for memory according to the storage information in the descriptor.
[0150] In practice, the accelerator device counts received application data packets using a descriptor count and uploads this descriptor count to the host, informing it of the number of descriptors it can currently process. The host then loops through the descriptor count. If the descriptor count is non-zero, it allocates memory based on the storage information in the descriptor information.
[0151] S202: Fill the requested memory address into the descriptor to start direct memory access transmission;
[0152] In actual implementation, the host fills the requested memory address into the descriptor, and the host initiates a data transfer operation, that is, starts a direct memory access transfer.
[0153] S203: Receive an application data packet sent by the accelerator device.
[0154] In actual implementation, the host waits for receiving the application data packet transmission completion interrupt to end the operation.
[0155] In the embodiment of the present application, the accelerator device writes the storage information of the application data packet to be transmitted into the descriptor. The DMA transmission is initiated by the host, and the corresponding memory space is requested based on the storage information in the descriptor. This eliminates the need to allocate a large amount of memory space in advance, reduces the demand for CPU performance, and ensures the real-time nature of data transmission. As can be seen, the embodiment of the present application avoids the waste of resources in the direct memory access transmission process.
[0156] A data transmission device provided in an embodiment of the present application is introduced below. The data transmission device described below and the data transmission method described above can be referenced to each other.
[0157] Referring to FIG8 , a structural diagram of a data transmission device according to an exemplary embodiment is shown. As shown in FIG8 , the device includes:
[0158] The first application module 201 is configured to obtain a descriptor and apply for memory according to storage information in the descriptor; wherein the storage information includes the size of the application data packet to be transmitted and the storage address of the application data packet in the memory of the accelerator device;
[0159] The starting module 202 is configured to fill the requested memory address into the descriptor to start direct memory access transmission;
[0160] The receiving module 203 is configured to receive an application data packet sent by the accelerator device.
[0161] In the embodiment of the present application, the accelerator device writes the storage information of the application data packet to be transmitted into the descriptor. The DMA transmission is initiated by the host, and the corresponding memory space is requested based on the storage information in the descriptor. This eliminates the need to allocate a large amount of memory space in advance, reduces the demand for CPU performance, and ensures the real-time nature of data transmission. As can be seen, the embodiment of the present application avoids the waste of resources in the direct memory access transmission process.
[0162] Based on the above embodiment, as an optional implementation manner, the following is further included:
[0163] The query module is configured to query the descriptor count, and if the descriptor count is not zero, the workflow of the first application module 201 is started.
[0164] Based on the above embodiment, as an optional implementation manner, the following is further included:
[0165] The second application module is configured to apply for a descriptor storage space via direct memory access consistent memory.
[0166] Regarding the apparatus in the above embodiment, the manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.
[0167] Based on the hardware implementation of the above program modules, and in order to implement the method of the embodiment of the present application, the embodiment of the present application further provides a host. FIG9 is a structural diagram of a host according to an exemplary embodiment. As shown in FIG9 , the host includes:
[0168] Communication interface 1, capable of exchanging information with other devices such as network devices;
[0169] The processor 2 is connected to the communication interface 1 to implement information exchange with other devices and is configured to execute the data transmission method provided by one or more of the above technical solutions when running a computer program. The computer program is stored in the memory 3.
[0170] Of course, in actual use, the various components in the host are coupled together via bus system 4. It will be appreciated that bus system 4 is configured to enable communication between these components. In addition to a data bus, bus system 4 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in FIG9 , all of these buses are labeled as bus system 4.
[0171] The memory 3 in the embodiment of the present application is configured to store various types of data to support the operation of the host, and examples of such data include any computer program for operating on the host.
[0172] When the processor 2 executes the program, the corresponding processes in each method of the embodiment of the present application are implemented. For the sake of brevity, they are not repeated here.
[0173] In an exemplary embodiment, the embodiment of the present application further provides a non-volatile readable storage medium, that is, a computer non-volatile readable storage medium, which can be a computer non-volatile readable storage medium, for example, including a memory storing a computer program, and the above-mentioned computer program can be executed by a processor to complete the steps of the data transmission method on the accelerator device side and the host side. The computer non-volatile readable storage medium can be FRAM (Ferroelectric Random Access Memory), ROM (Read-Only Memory), PROM (Programmable Read-Only Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), Flash Memory, magnetic surface memory, optical disk, CD-ROM (Compact Disc Read-Only Memory) and other memories.
[0174] The embodiment of the present application discloses a data transmission system, comprising an accelerator device as provided in the above embodiment and a host as provided in the above embodiment, wherein the host is connected to the accelerator device. The accelerator device can be connected to the host via a PCIe hard core module.
[0175] As a feasible implementation method, the data preprocessing module is located in a field programmable logic gate array, and the data transmission system is shown in FIG10 .
[0176] As another feasible implementation, the data preprocessing module is independent of the FPGA. The data transmission system is shown in Figure 11. A dedicated scenario-specific device, such as the RK3399, is selected to construct the data reception preprocessing module. The dedicated data preprocessing module transmits the processed results to the FPGA via interfaces such as PCIe or SDIO (Secure Digital Input and Output). The FPGA only performs data caching and forwarding. A dedicated data preprocessing module can effectively improve data preprocessing efficiency, reduce FPGA development difficulty, and improve overall system efficiency.
[0177] In the embodiment of the present application, the accelerator device writes the storage information of the application data packet to be transmitted into the descriptor. The DMA transmission is initiated by the host, and the corresponding memory space is requested based on the storage information in the descriptor. This eliminates the need to allocate a large amount of memory space in advance, reduces the demand for CPU performance, and ensures the real-time nature of data transmission. As can be seen, the embodiment of the present application avoids the waste of resources in the direct memory access transmission process.
[0178] The following describes an application embodiment provided by the present application, a data processing and transmission system based on an FPGA accelerator. The system is targeted at parallel application scenarios such as multi-channel ADC conversion and multi-channel image input. First, the cache module receives the data and then performs data preprocessing, such as filtering and image decompression. The data preprocessing module can effectively reduce the data processing pressure of the CPU. After the data is preprocessed, it is first temporarily stored in the FPGA accelerator DDR, and then the data storage information in the DDR is recorded by the data transmission control module. The storage information is then configured in the DMA descriptor, and finally the host is notified to perform data reading operations through DMA. Compared with the traditional processing method, the DMA startup of the data transmission is initiated by the host, which avoids the need to open up a large amount of memory space in advance, reduces the demand for CPU performance, and ensures the real-time performance of data transmission.
[0179] As shown in Figure 12, the following steps are included:
[0180] Step 1: The host first applies for descriptor storage space through DMA consistency memory and starts the system;
[0181] Step 2: The data preprocessing module in the FPGA accelerator receives parallel data such as multi-channel ADCs and images and performs preprocessing operations such as ADC conversion and image processing. Each application processing algorithm can be developed according to different applications.
[0182] Step 3: The data cache module caches and packages the pre-processed data and writes it into DDR;
[0183] Step 4: The data transmission module records and controls the transmission information. If there is no data, the data preprocessing module continuously monitors the data on the parallel interface. If there is data, it records the length of each data packet and the storage address of each data packet in the DDR.
[0184] Step 5: For applications with high real-time requirements, the data transmission module fills the length of each data packet and the storage address information in the DDR into the descriptor, and updates the count of available descriptors through the read / write module (the count increases by one each time a descriptor is filled). For applications with low real-time requirements, the DDR is divided into blocks, and the descriptor is filled after each DDR data block is filled, and the count of available descriptors is updated through the read / write module (the count increases by one each time a descriptor is filled).
[0185] Step 6: The host cyclically queries the descriptor counter. If the value is not zero, it means that there is data. Then the descriptor is read back, memory is requested according to the length information in the descriptor, and the requested memory address is filled into the descriptor, and then DMA transfer is started.
[0186] Step 7: The FPGA accelerator checks and records the data transmission status. The host waits to receive the transmission completion interrupt and ends the operation.
[0187] The embodiment of the present application designs a data processing and transmission system based on an FPGA accelerator, designs a data preprocessing module, and offloads work suitable for parallel computing to the data preprocessing module to reduce the workload of the CPU; designs a data reading mode actively initiated by the FPGA end to reduce the memory overhead of the host end; uses a data interaction channel based on the read-write module to ensure the correctness of data transmission and avoid data packet loss; designs different data storage and transmission strategies for different real-time application scenarios to improve system stability.
[0188] The above are merely optional implementations of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A data transmission method, characterized in that: Applied to an accelerator device, the accelerator device is connected to a host, and the method comprises: receiving an application data packet and storing the application data packet in a memory; Recording storage information of the application data packet; wherein the storage information includes the size of the application data packet and the storage address of the application data packet in the memory; Writing the storage information of the application data packet into the descriptor, so that the host applies for memory according to the storage information in the descriptor, and filling the applied memory address into the descriptor to start direct memory access transmission; The application data packet in the memory is transmitted to the host according to the memory address in the descriptor filled by the host through direct memory access transmission.
2. The data transmission method according to claim 1, characterized in that: The receiving of the application data packet and storing the application data packet in a memory comprises: Receive application data packets; Preprocessing the application data packet; wherein the preprocessing includes any one or a combination of any one of analog-to-digital conversion, data filtering, and image decompression; The preprocessed application data packets are stored in the memory.
3. The data transmission method according to claim 2, characterized in that: The receiving of the application data packet comprises: receiving application data packets via multiple data channels; Accordingly, the application data packet is preprocessed, including: The application data packets received by the multiple data channels are preprocessed in parallel by multiple data preprocessing modules.
4. The data transmission method according to claim 1, characterized in that: Writing the storage information of the application data packet into the descriptor so that the host can apply for memory according to the storage information in the descriptor, and filling the applied memory address into the descriptor to start direct memory access transmission, including: A target descriptor is determined in the descriptors whose status is idle, the storage information of the application data packet is written into the target descriptor, and the status of the target descriptor is modified to occupied, so that the host applies for memory according to the size of the application data packet in the target descriptor, and fills the applied memory address into the target descriptor to start direct memory access transmission.
5. The data transmission method according to claim 4, characterized in that: After the application data packet transmission is completed, the host modifies the state of the target descriptor to idle.
6. The data transmission method according to claim 5, characterized in that: After the state of the target descriptor is changed to occupied, the method further includes: A descriptor count is updated and transmitted to the host; wherein the descriptor count is used to describe the number of descriptors in an occupied state.
7. The data transmission method according to claim 6, characterized in that: The host cyclically queries the descriptor count, and if the descriptor count is not zero, reads the target descriptor, and executes the step of applying for memory according to the size of the application data packet in the target descriptor; After the host modifies the state of the target descriptor to idle, the host updates the descriptor count according to the number of descriptors whose current state is occupied.
8. The data transmission method according to claim 1, characterized in that: Writing the storage information of the application data packet into the descriptor includes: After each application data packet is received and each application data packet is stored in the memory, the storage information of the application data packet is written into the descriptor.
9. The data transmission method according to claim 1, characterized in that: The memory is divided into a plurality of storage blocks, and the application data packet is stored in the memory, including: storing the application data packets in sequence into storage blocks in a memory; Accordingly, the storage information of the application data packet is written into the descriptor, including: After each storage block is filled, the storage information of the application data packet in the filled storage block is written into the descriptor; wherein each storage information of the application data packet corresponds to a descriptor.
10. The data transmission method according to claim 1, characterized in that: The accelerator device includes a field programmable logic gate array and a memory, and the field programmable logic gate array includes a data cache module, a data transmission control module, a storage controller, a read-write module, and a direct memory access module; The data cache module is configured to cache received application data packets; The storage controller is configured to store the cached application data packets into the memory; The data transmission control module is configured to record the storage information of the application data packet, and write the storage information of the application data packet into the descriptor through the read-write module; The read / write module is configured to transmit information with the host; The direct memory access module is configured to transfer the application data packet in the memory to the host according to the memory address in the descriptor filled by the host.
11. The data transmission method according to claim 10, characterized in that: The field programmable logic gate array further comprises a data preprocessing module, and the data preprocessing module is configured to preprocess the received application data packets.
12. The data transmission method according to claim 10, characterized in that: The accelerator device further includes a data preprocessing module independent of the field programmable gate array, the data preprocessing module is connected to the field programmable gate array, and the data preprocessing module is configured to preprocess received application data packets.
13. A data transmission device, characterized in that: Applied to an accelerator device, the accelerator device is connected to a host, and the device comprises: A storage module, configured to receive an application data packet and store the application data packet in a memory; A recording module, configured to record storage information of the application data packet; wherein the storage information includes the size of the application data packet and the storage address of the application data packet in the memory; A writing module is configured to write the storage information of the application data packet into the descriptor, so that the host applies for memory according to the storage information in the descriptor, and fills the applied memory address into the descriptor to start direct memory access transmission; The transmission module is configured to transmit the application data packet in the memory to the host according to the memory address in the descriptor filled by the host through direct memory access transmission.
14. An accelerator device, characterized in that: include: a memory arranged to store a computer program; A processor, configured to implement the steps of the data transmission method according to any one of claims 1 to 12 when executing the computer program.
15. A data transmission method, characterized in that: Applied to a host, the host is connected to an accelerator device, the method comprises: Obtaining a descriptor, and applying for memory according to storage information in the descriptor; wherein the storage information includes the size of the application data packet to be transmitted and the storage address of the application data packet in the memory of the accelerator device; Fill the requested memory address into the descriptor to start direct memory access transmission; Receive the application data packet sent by the accelerator device.
16. The data transmission method according to claim 15, characterized in that: Before obtaining the descriptor, the following step is also included: The descriptor count is queried, and if the descriptor count is not zero, the steps of obtaining the descriptor and applying for memory according to the storage information in the descriptor are executed.
17. The data transmission method according to claim 15, characterized in that: Before obtaining the descriptor, the following step is also included: Request descriptor storage space through direct memory access consistent memory.
18. A host, characterized in that: include: a memory arranged to store a computer program; A processor, configured to implement the steps of the data transmission method according to any one of claims 15 to 17 when executing the computer program.
19. A computer non-volatile readable storage medium, characterized in that: The computer non-volatile readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the data transmission method according to any one of claims 1 to 12 or claims 15 to 17 are implemented.
20. A data transmission system, characterized in that: The invention comprises the accelerator device as claimed in claim 14 and the host as claimed in claim 18, wherein the host is connected to the accelerator device.
Citation Information
Patent Citations
RDMA-based communication method and device and storage medium
CN109426631A
A DMA transmission method and a DMA controller suitable for network transmission
CN109558344A
DMA device
CN115168257A
Accelerator processing method and device, storage medium and processor
CN115373810A
Data transmission method and device, accelerator equipment, host and storage medium
CN117312201A
Cited By
Data transmission system and method, electronic equipment and storage medium
CN120561050A