A Method for Implementing 10 Gigabit Ethernet Embedded Data Offloading Based on UDP Protocol through FPGA
The 10Gbps embedded data offload method based on UDP protocol is realized through FPGA, and the two-way handshake, IP sharding and dynamic flow control mechanism are used to solve the problem of insufficient transmission speed in the existing technology, and efficient 10Gbps data transmission is achieved.
Patent Information
- Application Number
- CN202510161616.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-14
AI Technical Summary
In the prior art, the transmission speed of Gigabit Ethernet and USB offloading methods cannot meet the requirements of large-capacity storage for data offloading time.
The 10Gigabit network embedded data offload method based on UDP protocol is realized through FPGA, and two-way handshake, IP sharding and dynamic flow control mechanism are adopted to solve the problem of inefficient transmission efficiency caused by frequent IP frame interruption at the host receiver.
It achieves data transmission speeds up to 10Gbps, improves the effectiveness and rate of data transmission, and completely solves the problem of insufficient transmission speed.
Smart Images

Figure CN119652839B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of embedded storage technologies, and more specifically, to a method for realizing 10 Gigabit Ethernet embedded data offloading based on the UDP protocol through an FPGA. Background Art
[0002] With the rapid development of information technology, high-speed and massive embedded data storage has been widely applied in many fields such as radar, communication, and remote sensing. For the current large-capacity and high-speed data storage, a convenient, fast, and reliable data offloading method has become increasingly important.
[0003] The existing technologies have the following deficiencies:
[0004] Currently, there are two main data offloading methods for embedded memories: one is to offload through Gigabit Ethernet, and the other is to offload through USB. Both offloading methods have the advantages of simple interfaces, generality, and convenient transmission, but their transmission speeds are limited. The transmission speed of Gigabit Ethernet is 1 Gbps, and the fastest transmission speed of USB 3.0 is only 5 Gbps. For some applications that require large-capacity storage (in the order of TB) and have requirements for data offloading time, the offloading speeds of Gigabit Ethernet and USB cannot meet the application requirements. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the existing technologies, an embodiment of the present invention provides a method for realizing 10 Gigabit Ethernet embedded data offloading based on the UDP protocol through an FPGA. By comprehensively using two-way handshake, IP fragmentation, and formulating a dynamic flow control mechanism at the FPGA hardware layer to solve the problems of the native UDP protocol without an acknowledgment response mechanism and the low transmission efficiency caused by frequent interruptions of IP frames at the host receiving end, the effectiveness and rate of data transmission are improved. The transmission speed of up to 10 Gbps is achieved through the 10G-UDP protocol, which completely solves the problems raised in the above-mentioned background art.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A method for realizing 10 Gigabit Ethernet embedded data offloading based on the UDP protocol through an FPGA includes the following steps:
[0008] Step S1: The device side offloads the stored data in the solid-state drive through the SATA or PCIE interface and packs the offloaded data into frame data;
[0009] Step S2: After the frame data is cached, it is packed into IP frames and transmitted to the host along with the transmission packet sequence number. After receiving the IP frames, the host side parses the transmission packet sequence number and generates response frames, and sends the response frames back to the receiving end two-way handshake protocol module through a 10 Gigabit Ethernet. The two-way handshake protocol module parses the received response frames to obtain the received packet sequence number, and performs verification with the transmission packet sequence number, and selects to retransmit the IP frames or confirm the IP frames according to the verification result;
[0010] Step S3: When sending IP frames, the adaptive IP fragmentation module fragments the IP frames to obtain multiple IP sub-frames. The IP sub-frames will be automatically assembled into IP frames at the host side driver layer and generate interrupts to notify the host to receive. At the same time, the transmission traffic of the IP sub-frames is controlled by inserting a delay count method;
[0011] Step S4: After using the IP fragmentation module to fragment the IP frames, the multiple IP sub-frames are converted into UDP data through the 10G-UDP protocol module for information interaction with the host side.
[0012] In a preferred embodiment, in step S1, the data offloading module includes four parts: a SATA / PCIE interface control module, a Raid control module, a FIFO cache module, and a data packing module, which are used to complete the data offloading and packing of the solid-state drive;
[0013] The SATA / PCIE interface control module is used to implement the parsing of the SATA / PCIE interface protocol, and send the parsed data to the Raid control module. The SATA / PCIE interface control module parses for different hard disk protocol types and supports interfaces such as SATA1.0 / 2.0 / 3.0 and PCIE1.0 / 2.0 / 3.0;
[0014] The Raid control module is used for disk array management to improve the performance of the storage system, supports custom configuration of the number of devices, and has a built-in DMA data transfer engine; after sorting out the received parsed data, the Raid control module sends it to the FIFO cache module for caching through DMA;
[0015] The FIFO cache module is used to cache the offloaded data;
[0016] The data packing module packs the cached data in the FIFO cache module, and the default data packing length can be customized in the data packing module.
[0017] In a preferred embodiment, in step S2, the two-way handshake protocol module realizes the two-way handshake between the device side and the host side through the packet sequence number feedback verification mechanism, and retransmits the current sending frame when the handshake fails or the response times out.
[0018] In a preferred embodiment, in step S2, the two-way handshake protocol module includes a sending packet sequence number module, a receiving packet sequence number module, a timeout module, an SRAM cache module, and a handshake protocol logic control module. The functions of each module are as follows:
[0019] The sending packet sequence number module is used to generate the sending packet sequence number required for the handshake. When the IP frame data is sent successfully and the correct response packet sequence number is received, the sending packet sequence number is automatically incremented by 1;
[0020] The receiving packet sequence number module is used to receive the response frame and output the received packet sequence number for verification and comparison with the sending packet sequence number;
[0021] The timeout module is used to generate a response timeout interrupt. When the IP frame data starts to be sent, the timeout module starts timing. If the response frame sent from the host side is not received within the specified time, the timeout module generates a timeout response interrupt to retransmit the IP frame; when the response frame is received, the timeout timer counter is cleared;
[0022] The SRAM cache module is used to cache the frame data. It adopts a dual-SRAM mechanism and uses the ping-pong operation method to read and write data;
[0023] The handshake protocol logic control module is used to complete the logic control of the handshake protocol.
[0024] In a preferred embodiment, in step S2, the handshake protocol logic control module completes the logic control of the handshake protocol, including the following steps:
[0025] Step 1: Perform the SRAM cache read and write mechanism for the frame data;
[0026] Step 2: Package the frame data and the sending packet sequence number and output the IP frame;
[0027] Step 3: Perform a handshake verification on the sending packet sequence number and the received packet sequence number. If the handshake verification is successful, the next IP frame is sent;
[0028] Step 4: If the handshake verification fails or the response times out, retransmit the current IP frame.
[0029] In a preferred embodiment, in step S2, the specific steps of the two-way handshake protocol process are as follows:
[0030] The device - side caches the received frame data using the ping - pong operation method, reads the transmission frame data from the SRAM cache, packs the transmission frame data and the transmission packet sequence number into an IP frame, sends the IP frame and enters the waiting - for - response - frame state. After receiving the IP frame, the host - side parses the corresponding packet sequence number, packs the extracted packet sequence number into a response frame and sends it back to the device - side. After the device - side receives the response frame or times out waiting for the response, it checks and compares the transmission packet sequence number and the received packet sequence number;
[0031] If the transmission packet sequence number is equal to the received packet sequence number, the transmission packet sequence number is automatically incremented by 1, the re - transmission counter is cleared to 0, and preparations are made to send the next frame. If the transmission packet sequence number is not equal to the received packet sequence number, the current frame is re - transmitted. If the number of re - transmissions is greater than the preset re - transmission threshold, the re - transmission is terminated and a network transmission failure is reported.
[0032] In a preferred embodiment, in step S3, the adaptive IP fragmentation module includes six parts: an automatic flow control module, a FIFO cache module, a fragmentation logic control module, a fragmentation identification module, a fragmentation flag module, and a fragmentation offset module. The functions of each module are as follows:
[0033] The automatic flow control module is used to complete the flow control of IP frame sending. After the IP frame sending is completed, by inserting a delay count, it controls the sending time between IP frames;
[0034] The FIFO cache module is used to cache IP frames;
[0035] The fragmentation logic control module is used to complete fragmentation logic control and data fragmentation;
[0036] The fragmentation identification module is used to generate an IP fragmentation identification;
[0037] The fragmentation flag module is used to generate an IP fragmentation flag;
[0038] The fragmentation offset module is used to generate an IP fragmentation offset coordinate, which is used to mark the relative position of the current fragment in the entire IP packet.
[0039] In a preferred embodiment, in step S3, the principle of automatic flow control in the automatic flow control module is as follows:
[0040] When the flow control module detects an IP frame retransmission interruption, it indicates that there may be congestion in the current transmission network, there is a frame loss phenomenon, the delay counting time is automatically incremented by 1, the transmission time between IP frames is extended, the transmission rate of IP frames is reduced, and at the same time the 1-second timer count is cleared. When no IP frame retransmission interruption is detected within 1 second, it indicates that the current transmission network is unobstructed, the delay counting time is automatically decremented by 1, the transmission time between IP frames is shortened, the transmission rate of IP frames is increased, and thus the data transmission efficiency is improved. When a network transmission error is detected, the delay counting time is cleared, and after the network is restored, flow control is restarted.
[0041] In a preferred embodiment, in step S4, the 10G-UDP protocol module is a network protocol implemented based on a 10 Gigabit Ethernet, and is used to complete the underlying logical interface protocol conversion.
[0042] The technical effects and advantages of the method for realizing a 10 Gigabit Ethernet embedded data offloading based on the UDP protocol through an FPGA according to the present invention:
[0043] The device end of the present invention unloads the stored data in the solid state drive through the SATA or PCIE interface, packs the unloaded data into frame data, caches the frame data and then packs it into an IP frame and attaches the transmission packet sequence number and transmits it to the host end. After receiving the IP frame, the host end parses the transmission packet sequence number and generates a response frame and sends it back to the device end. The device end parses the received packet sequence number from the received response frame and performs a check with the transmission packet sequence number. By performing a two-way handshake check on the sequence numbers of the transmitted packet and the received packet, it is ensured whether there is any loss in the previous data during the transmission process. According to the check result, an IP frame retransmission or an IP frame confirmation is selected. At the same time, when sending an IP frame, fragmentation is performed to obtain multiple IP sub-frames. The IP sub-frames will be automatically assembled into an IP frame at the host end driver layer and generate an interruption to notify the host to receive. Through the IP fragmentation method, the single data transmission volume of the IP frame is increased, the interruption reception frequency of the IP frame at the host end is effectively reduced. At the same time, when sending an IP sub-frame, the transmission flow of the IP sub-frame is controlled by inserting a delay count method, the retransmission times are effectively reduced, and the data transmission efficiency is improved. Finally, through the 10G-UDP protocol module, multiple IP sub-frames are converted into UDP data for information interaction with the host end. By comprehensively using two-way handshake, IP fragmentation and the FPGA hardware layer to formulate a dynamic flow control mechanism, the problems of the native UDP protocol without a response mechanism and the low transmission efficiency caused by the frequent interruption of the IP frame at the host receiving end are solved, and the effectiveness and rate of data transmission are improved. Description of the Drawings
[0044] Figure 1 It is a flowchart of the method for realizing a 10 Gigabit Ethernet embedded data offloading based on the UDP protocol through an FPGA according to the present invention.
[0045] Figure 2It is a logic flow block diagram of a method for implementing 10 Gigabit Ethernet embedded data offloading based on the UDP protocol through FPGA in the present invention.
[0046] Figure 3 It is a data offloading module diagram of a method for implementing 10 Gigabit Ethernet embedded data offloading based on the UDP protocol through FPGA in the present invention.
[0047] Figure 4 It is a two-way handshake protocol module diagram of a method for implementing 10 Gigabit Ethernet embedded data offloading based on the UDP protocol through FPGA in the present invention.
[0048] Figure 5 It is a two-way handshake protocol flow chart of a method for implementing 10 Gigabit Ethernet embedded data offloading based on the UDP protocol through FPGA in the present invention.
[0049] Figure 6 It is an adaptive IP fragmentation module diagram of a method for implementing 10 Gigabit Ethernet embedded data offloading based on the UDP protocol through FPGA in the present invention.
[0050] Figure 7 It is an automatic flow control schematic diagram of a method for implementing 10 Gigabit Ethernet embedded data offloading based on the UDP protocol through FPGA in the present invention.
[0051] Figure 8 It is an adaptive IP fragmentation flow chart of a method for implementing 10 Gigabit Ethernet embedded data offloading based on the UDP protocol through FPGA in the present invention.
[0052] Figure 9 It is a 10G-UDP protocol module diagram of a method for implementing 10 Gigabit Ethernet embedded data offloading based on the UDP protocol through FPGA in the present invention. Detailed implementation manners
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] The device side of the present invention unloads the stored data in the solid-state drive through the SATA or PCIE interface, packs the unloaded data into frame data, caches the frame data and then packs it into an IP frame with a transmission packet sequence number attached and sends it to the host side. After receiving the IP frame, the host side parses the transmission packet sequence number and generates a response frame to be sent back to the device side. The device side parses the received packet sequence number from the received response frame and performs a check with the transmission packet sequence number. According to the check result, it selects to retransmit the IP frame or confirm the IP frame. At the same time, when sending the IP frame, the IP frame is fragmented into multiple IP sub-frames to reduce the IP frame interruption reception frequency at the host side, controls the sending traffic of the IP sub-frames by inserting a delay count method, reduces the number of retransmissions, improves the data transmission efficiency, and finally converts the multiple IP sub-frames into UDP data through the 10G-UDP protocol module to interact with the host side for information.
[0055] Embodiment, a method for realizing 10 Gigabit Ethernet embedded data offloading based on the UDP protocol through FPGA, as Figure 1 shown, includes the following steps:
[0056] Step S1: The device side unloads the stored data in the solid-state drive through the SATA or PCIE interface and packs the unloaded data into frame data;
[0057] Step S2: After caching the frame data, it is packed into an IP frame with a transmission packet sequence number attached and sent to the host. After receiving the IP frame, the host side parses the transmission packet sequence number and generates a response frame, and sends the response frame back to the receiving end two-way handshake protocol module through the 10 Gigabit Ethernet. The two-way handshake protocol module parses the received packet sequence number from the received response frame and performs a check with the transmission packet sequence number, and selects to retransmit the IP frame or confirm the IP frame according to the check result;
[0058] Step S3: When sending the IP frame, the adaptive IP fragmentation module fragments the IP frame into multiple IP sub-frames. The IP sub-frames will be automatically assembled into an IP frame at the host side driver layer and generate an interrupt to notify the host to receive, and at the same time controls the sending traffic of the IP sub-frames by inserting a delay count method;
[0059] Step S4: After using the IP fragmentation module to fragment the IP frame, convert the multiple IP sub-frames into UDP data through the 10G-UDP protocol module to interact with the host side for information.
[0060] The specific implementation is as follows:
[0061] The method for realizing 10 Gigabit Ethernet embedded data offloading based on the UDP protocol through FPGA mainly includes four modules, as Figure 2 shown, which are the data offloading module, the two-way handshake protocol module, the adaptive IP fragmentation module, and the 10G-UDP protocol module.
[0062] The data offloading module is mainly used to complete the data offloading of the solid-state drive and pack the offloaded data into the corresponding frame format and send it to the two-way handshake protocol module.
[0063] The two-way handshake protocol module is mainly used to implement the protocol handshake and message retransmission for data transmission between the device side and the host side; to improve the reliability of data transmission, because the UDP protocol is an unreliable transmission protocol without an acknowledgment response confirmation mechanism, but has the advantages of fast transmission speed and simple implementation.
[0064] The adaptive IP fragmentation module is used to automatically complete IP fragmentation according to different frame lengths; it can effectively reduce the number of network receive interrupt responses on the host side and improve the transmission efficiency of network data.
[0065] The 10G-UDP protocol module is used to complete the parsing of network protocols such as UDP, ARP, and ICMP, and implement the sending and receiving of communication data.
[0066] It should be noted that the solid-state drive is a data storage device used to store various types of data in a computer. The acknowledgment response confirmation mechanism is a mechanism used to ensure reliable data transmission in network communication and data transmission. UDP, ARP, and ICMP are different network protocols. UDP is a transport layer protocol used to provide connectionless data transmission services. ARP and ICMP are network layer protocols. The ARP protocol is used to resolve IP addresses into MAC addresses so that devices within a local area network can communicate. The ICMP protocol is used to send error messages and network control information to help diagnose and manage network connections.
[0067] In step S1, the data offloading module includes four parts: the SATA / PCIE interface control module, the Raid control module, the FIFO cache module, and the data packing module. As Figure 3 shown, the data offloading module is used to complete the data offloading and packing of the solid-state drive.
[0068] The SATA / PCIE interface control module is used to implement the parsing of the SATA / PCIE interface protocol and send the parsed data to the Raid control module. The SATA / PCIE interface control module can parse different hard disk protocol types and support interfaces such as SATA1.0 / 2.0 / 3.0 and PCIE1.0 / 2.0 / 3.0.
[0069] The Raid control module is used for disk array management to improve the performance of the storage system, supports custom configuration of the number of devices, and has a built-in DMA data transmission engine; after organizing the received parsed data, the Raid control module sends it to the FIFO cache module for caching through DMA.
[0070] The FIFO cache module is used to cache the offloaded data;
[0071] The data packing module packs the cached data in the FIFO cache module. The default data packing length can be customized in the data packing module. For example, the default data packing length is set to 48KB. After the packing is completed, the framed data is sent to the two-way handshake protocol module for subsequent UDP protocol transmission. The framed data parameters include: the framed data length, marked as frame_len; the framed data, marked as frame_data; the framed data validity, marked as frame_data_val; and the framed data end flag, marked as frame_data_last.
[0072] It should be noted that SATA and PCIE are two common computer interface standards, which are used to connect storage devices and expansion cards respectively; the DMA data transfer engine is a hardware mechanism used to achieve high-speed data transfer in a computer system; since the maximum length of a network IP frame does not exceed 64KB, when customizing the default data packing length, ensure that its length does not exceed 64KB.
[0073] In step S2, the two-way handshake protocol module realizes the two-way handshake between the device side and the host side through the packet sequence number feedback verification mechanism. When the handshake fails or the response times out, the currently sent frame is retransmitted to make up for the transmission problem caused by the UDP transmission protocol without an acknowledgment mechanism, thereby improving the reliability of data transmission.
[0074] The two-way handshake protocol module includes five parts: a sending packet sequence number module, a receiving packet sequence number module, a timeout module, an SRAM cache module, and a handshake protocol logic control module. As Figure 4 shown, the functions of each module are as follows:
[0075] The sending packet sequence number module is used to generate the sending packet sequence number required for the handshake. When the IP frame data is sent successfully and the correct response packet sequence number is received, the sending packet sequence number is automatically incremented by 1;
[0076] The receiving packet sequence number module is used to receive the response frame and output the received packet sequence number for verification and comparison with the sending packet sequence number;
[0077] The timeout module is used to generate a timeout response interrupt. When the IP frame data starts to be sent, the timeout module starts timing. If the response frame sent by the host side is not received within the specified time, the timeout module generates a timeout response interrupt to retransmit the IP frame; when the response frame is received, the timeout timer counter is cleared;
[0078] The SRAM cache module is used to cache frame data. It adopts a dual-SRAM mechanism and uses the ping-pong operation method to read and write data. It should be noted that since the maximum length of an IP frame is 64KB, the SRAM cache capacity is 64KB.
[0079] The handshake protocol logic control module is mainly used to complete the logic control of the handshake protocol, including the following steps:
[0080] Step 1: Implement the SRAM cache read and write mechanism for frame data;
[0081] Step 2: Package the frame data and the sending packet sequence number and output an IP frame;
[0082] Step 3: Perform handshake verification on the sending packet sequence number and the receiving packet sequence number. If the handshake verification is successful, send the next IP frame;
[0083] Step 4: If the handshake verification fails or the response times out, retransmit the current IP frame.
[0084] The IP frame parameters output by the handshake protocol logic control module include: the IP frame length, including the frame data length and the sending packet sequence number length, marked as ip_len; the IP frame data, including the frame data and the packet sequence number, marked as ip_data; the IP frame data validity, marked as ip_data_val; the IP frame data end flag, marked as ip_data_last; the IP frame retransmission interruption, used for fragment flow control, marked as ip_retrans_int; the network error flag, marked as ethnet_error.
[0085] The two-way handshake protocol process is as Figure 5 shown, and the specific steps are as follows:
[0086] The device side caches the received frame data using the ping-pong operation method, reads the sending frame data from the SRAM cache, packages the sending frame data and the sending packet sequence number into an IP frame, sends the IP frame and enters the waiting for the response frame state. After the host side receives the IP frame, it parses the corresponding packet sequence number, and the host side packages the extracted packet sequence number into a response frame and sends it back to the device side. After the device side receives the response frame or receives a response timeout, it compares the sending packet sequence number and the receiving packet sequence number for verification.
[0087] If the sending packet sequence number and the receiving packet sequence number are equal, the sending packet sequence number is automatically incremented by 1, the retransmission counter is cleared to 0, and it is ready to send the next frame; if the sending packet sequence number and the receiving packet sequence number are not equal, the current frame is retransmitted. If the number of retransmissions is greater than the preset retransmission threshold, the retransmission is terminated and a network transmission failure is reported.
[0088] It should be noted that the dual-SRAM mechanism is a cache management strategy for embedded system and microcontroller designs, which is used to improve the efficiency and reliability of data access. The dual-SRAM mechanism realizes data storage and management by using two independent SRAM blocks, namely two independent static random access memories. The ping-pong operation mode is a common data transfer and processing strategy, which is widely used in computer systems, embedded systems, and signal processing fields. By alternately using two storage areas, the efficiency and speed of data processing are improved.
[0089] In step S3, the maximum packet length of the network IP frame is 64KB, but the maximum transmission unit allowed by the link layer is mostly 1500 bytes. When the length of the IP transmission packet is greater than the maximum transmission unit of the link layer, IP fragmentation is required. IP fragmentation is implemented through FPGA, which reduces the number of interrupt responses received by the host-side network and improves the transmission efficiency of network data.
[0090] The adaptive IP fragmentation module includes six parts: an automatic flow control module, a FIFO cache module, a fragmentation logic control module, a fragmentation identification module, a fragmentation flag module, and a fragmentation offset module. As Figure 6 shown, the adaptive IP fragmentation module is used to implement adaptive IP fragmentation.
[0091] The automatic flow control module is used to complete the flow control of IP frame sending. After the IP frame sending is completed, by inserting a delay count, the sending time between IP frames is controlled to achieve the flow control of IP frame sending. The delay count time is automatically adjusted by detecting the IP retransmission interrupt to achieve automatic flow control;
[0092] The FIFO cache module is used to cache IP frames;
[0093] The fragmentation logic control module is used to complete fragmentation logic control and data fragmentation. The fragmentation logic control module includes the following control and output parameters: the number of IP frames is marked as ip_seg_num, which is the quotient of the IP frame length divided by the IP frame length, and if there is a remainder, it is incremented by 1; the IP frame count is marked as ip_seg_cnt, which is automatically incremented by 1 after sending one IP frame; the IP frame sending completion flag is marked as ip_seg_tx_done; the IP frame sending completion flag is marked as ip_tx_done; the IP frame length is marked as ip_seg_len; the IP frame data is marked as ip_seg_data; the IP frame data validity flag is marked as ip_seg_data_valid; the start bit of the IP frame data is marked as ip_seg_data_sof; the end bit of the IP frame data is marked as ip_seg_data_eof;
[0094] The fragmentation identification module is used to generate IP fragmentation identifiers to distinguish different IP frames. For the fragments of the same IP packet, the same fragmentation identifier is used. When all IP fragments are sent, the IP fragmentation identifier is automatically incremented by 1. The fragmentation identification module sets the IP fragmentation identifier ip_seg_id by detecting the IP frame transmission completion flag ip_tx_done parameter. When ip_seg_id is equal to 1, that is, the IP frame is sent, it is automatically incremented by 1;
[0095] The fragmentation flag module is used to generate IP fragmentation flags. The fragmentation flag contains two flag bits, DF and MF; when DF is 1, fragmentation is prohibited, and when DF is 0, fragmentation is allowed; when MF is 1, it means there are more fragments, and when MF is 0, it means the current fragment is the last fragment. The fragmentation flag module sets the following parameters by detecting the IP fragmentation number ip_seg_num and IP fragmentation count ip_seg_cnt parameters:
[0096] Setting 1: Mark the DF bit of the IP fragmentation flag as ip_seg_df. When ip_seg_num is greater than 1, ip_seg_df is set to 0 to allow fragmentation; otherwise, ip_seg_df is set to 1 to prohibit fragmentation;
[0097] Setting 2: Mark the MF bit of the IP fragmentation flag as ip_seg_mf. When ip_seg_cnt is equal to ip_seg_num minus 1, ip_seg_mf is set to 0, indicating that the current IP fragment is the last fragment; otherwise, ip_seg_mf is set to 1, indicating that the current IP fragment is not the last fragment;
[0098] The fragmentation offset module is used to generate IP fragmentation offset coordinates to mark the relative position of the current fragment in the entire IP packet. Taking 8 bytes as a unit, the fragmentation offset module sets the IP fragmentation offset by detecting the IP fragmentation transmission completion flag ip_seg_tx_done parameter:
[0099] Mark the IP fragmentation offset as ip_seg_offset. When ip_seg_tx_done is set to 1, that is, the IP fragmentation is sent, it is automatically incremented by ip_seg_len divided by 8.
[0100] The automatic flow control principle in the automatic flow control module is as Figure 7 shown, and the specific steps are as follows:
[0101] When the flow control module detects an IP frame retransmission interruption, it indicates that there may be congestion in the current transmission network, there is a frame loss phenomenon, the delay count time is automatically incremented by 1, the transmission time between IP frames is extended, the transmission rate of IP frames is reduced, and at the same time the 1-second timer count is cleared. When no IP frame retransmission interruption is detected within 1 second, it indicates that the current transmission network is unobstructed, the delay count time is automatically decremented by 1, the transmission time between IP frames is shortened, the transmission rate of IP frames is increased, and thus the data transmission efficiency is improved. When a network transmission error is detected, the delay count time is cleared, and after the network is restored, flow control is restarted.
[0102] The working principle of the adaptive IP fragmentation module is as Figure 8 shown, and the specific steps are as follows:
[0103] After receiving a new IP frame, first cache the IP frame, and at the same time calculate the number of IP fragments ip_seg_num through the IP frame length ip_len and the IP fragment length ip_seg_len;
[0104] When ip_seg_num is greater than 1, the IP fragmentation flag ip_seg_df is set to 0, allowing fragmentation; otherwise ip_seg_df is set to 1, prohibiting fragmentation;
[0105] When the IP fragment count ip_seg_cnt is equal to the number of IP fragments ip_seg_num minus 1, the IP fragmentation flag ip_seg_mf is set to 0, and the current fragment is the last fragment; otherwise the IP fragmentation flag ip_seg_mf is set to 1, and the current fragment is not the last fragment;
[0106] After the IP fragmentation flag is configured, read the IP fragmentation data from the FIFO cache for transmission, the data length is the IP fragment length ip_seg_len, after the IP frame transmission is completed, the IP fragment count ip_seg_cnt is incremented by 1, and the IP fragment offset ip_seg_offset is incremented by ip_seg_len divided by 8;
[0107] After the flow control delay timing is completed, determine whether the IP fragment count ip_seg_cnt is equal to the number of IP fragments ip_seg_num. If they are equal, the IP frame transmission is completed, the IP fragmentation identifier ip_seg_id is incremented by 1, and wait for the next frame of IP frame data; if they are not equal, prepare to transmit the next IP frame data.
[0108] It should be noted that the delay count is a technique for measuring and controlling time delay. In this example, the delay count is used to implement automatic flow control.
[0109] In step S4, the 10G-UDP protocol module is a network protocol implemented based on the 10 Gigabit Ethernet, including five parts: a UDP sending module, a UPD receiving module, an ARP protocol module, an ICMP protocol module, and a 10G Ethernet PCS / PMA IP interface control module. As Figure 9 shown, the 10G-UDP protocol module supports multiple protocols and can implement two-way UDP communication.
[0110] The UDP sending module is used to implement the UDP protocol to send packets in a group and complete the packet sending of the fragmented frames;
[0111] The UDP receiving module is used to implement the UDP protocol parsing and complete the correct reception of the response packet sequence numbers;
[0112] The ARP protocol module is used to complete the active or passive address protocol parsing;
[0113] The ICMP protocol module is used to complete the response to the host PING command;
[0114] The 10G Ethernet PCS / PMA IP interface control module, as the PHY module of the 10G Ethernet, is directly implemented by calling the 10G Ethernet PCS / PMA IP inside the FPGA and is used to complete the conversion of the underlying logic interface protocol.
[0115] It should be noted that the PHY module, that is, the physical layer module, is used to implement the physical layer functions of data communication and ensure the transmission of data on the physical medium.
[0116] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product.
[0117] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application of the technical solution and the invention constraints. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0118] In addition, in each embodiment of the present application, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0119] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claimed rights.
[0120] Finally: The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall all be included within the protection scope of the present invention.
Claims
1. A method for implementing 10G network embedded data unloading based on UDP protocol through FPGA, characterized in that: The following steps are involved: Step S1: The device side unloads the stored data in the solid state drive through the SATA or PCIE interface, and packages the unloaded data into frame data; Step S2: After the frame data is cached, it is packaged into an IP frame and transmitted to the host with a sending packet sequence number. After receiving the IP frame, the host parses the sending packet sequence number and generates a response frame, and transmits the response frame back to the receiving end two-way handshake protocol module through the 10 Gigabit network. The two-way handshake protocol module parses the received response frame to obtain the receiving packet sequence number, and verifies it with the sending packet sequence number, and selects to retransmit the IP frame or confirm the IP frame according to the verification result; Step S3: When sending IP frames, the adaptive IP fragmentation module fragments the IP frames to obtain multiple IP frames. The IP frames are automatically assembled at the host driver layer to generate IP frames and an interrupt is generated to notify the host to receive the IP frames. At the same time, the sending flow of the IP frames is controlled by inserting a delay count. Step S4: After the IP frame is fragmented by the IP fragmentation module, the multiple IP frames are converted into UDP data through the 10G-UDP protocol module to exchange information with the host end; The host driver layer is a key software component in the computer system that is responsible for communicating and controlling hardware devices and is used to assemble IP frames.
2. The method for implementing 10G network embedded data unloading based on UDP protocol by FPGA according to claim 1, characterized in that: In step S1, the data unloading module includes four parts: SATA / PCIE interface control module, Raid control module, FIFO cache module, and data packaging module, which are used to complete data unloading and packaging of the solid state drive; The SATA / PCIE interface control module is used to implement SATA / PCIE interface protocol parsing and send the parsed data to the Raid control module. The SATA / PCIE interface control module parses different hard disk protocol types and supports SATA1.0 / 2.0 / 3.0 and PCIE1.0 / 2.0 / 3.0 interfaces; The Raid control module is used for disk array management, improving storage system performance, supporting custom configuration of device quantity, and built-in DMA data transmission engine; the Raid control module organizes the received parsed data and sends it to the FIFO cache module for caching through DMA; The FIFO buffer module is used to buffer the offload data; The data packing module packs the cache data in the FIFO cache module, and the default data packing length can be customized in the data packing module.
3. The method for implementing 10G network embedded data unloading based on UDP protocol by FPGA according to claim 1, characterized in that: In step S2, the two-way handshake protocol module implements a double handshake between the device and the host through a packet sequence number return verification mechanism, and retransmits the current sending frame when the handshake fails or the response times out.
4. The method for implementing 10G network embedded data unloading based on UDP protocol by FPGA according to claim 3, characterized in that: In step S2, the two-way handshake protocol module includes five parts: a sending packet sequence number module, a receiving packet sequence number module, a timeout module, an SRAM cache module, and a handshake protocol logic control module. The functions of each module are as follows: The sending packet sequence number module is used to generate the sending packet sequence number required for handshake. When the IP frame data is sent and the correct response packet sequence number is successfully received, the sending packet sequence number is automatically increased by 1; The receiving packet sequence number module is used to receive the response frame and output the receiving packet sequence number, which is used to check and compare the sending packet sequence number with the receiving packet sequence number; The timeout module is used to generate an over-response interrupt. When the IP frame data starts to be sent, the timeout module starts timing. If the response frame sent by the host is not received within the specified time, the timeout module generates a timeout response interrupt and retransmits the IP frame. When the response frame is received, the timeout timing counter is cleared. The SRAM cache module is used to cache frame data, adopts a dual SRAM mechanism, and uses a ping-pong operation method to read and write data; The handshake protocol logic control module is used to complete the logic control of the handshake protocol.
5. The method for implementing 10G network embedded data unloading based on UDP protocol by FPGA according to claim 4, characterized in that: In step S2, the handshake protocol logic control module completes the logic control of the handshake protocol, including the following steps: Step 1: Perform SRAM cache read and write mechanism on frame data; Step 2: Pack the frame data and the sending packet sequence number and output the IP frame; Step 3: Perform handshake verification on the sending packet sequence number and the receiving packet sequence number. If the handshake verification succeeds, the next IP frame is sent; Step 4: If the handshake verification fails or the response times out, the current IP frame is retransmitted.
6. The method for implementing 10G network embedded data unloading based on UDP protocol by FPGA according to claim 3, characterized in that: In step S2, the specific steps of the two-way handshake protocol process are as follows: The device caches the received frame data in a ping-pong operation mode, reads the sending frame data from the SRAM cache, packages the sending frame data and the sending packet sequence number into an IP frame, sends the IP frame and enters the waiting response frame state. After receiving the IP frame, the host parses the corresponding packet sequence number, and packages the extracted packet sequence number into a response frame and transmits it back to the device. After receiving the response frame or receiving the response timeout, the device verifies and compares the sending packet sequence number with the receiving packet sequence number. If the sending packet sequence number is equal to the receiving packet sequence number, the sending packet sequence number is automatically increased by 1, the retransmission counter is cleared to 0, and preparations are made to send the next frame; if the sending packet sequence number is not equal to the receiving packet sequence number, the current frame is retransmitted. If the number of retransmissions is greater than the preset retransmission threshold, the retransmission is terminated and the network transmission failure is reported.
7. The method for implementing 10G network embedded data unloading based on UDP protocol by FPGA according to claim 1, characterized in that: In step S3, the adaptive IP fragmentation module includes six parts: automatic flow control module, FIFO buffer module, fragmentation logic control module, fragmentation identification module, fragmentation mark module, and fragmentation offset module. The functions of each module are as follows: The automatic flow control module is used to complete the flow control of IP frame transmission. After the IP frame transmission is completed, the transmission time between IP frames is controlled by inserting a delay count; The FIFO buffer module is used to buffer IP frames; The sharding logic control module is used to complete sharding logic control and data sharding; The fragment identification module is used to generate IP fragment identification; The fragmentation mark module is used to generate IP fragmentation marks; The fragment offset module is used to generate IP fragment offset coordinates to mark the relative position of the current fragment in the entire IP message.
8. The method for implementing 10G network embedded data unloading based on UDP protocol by FPGA according to claim 7, characterized in that: In step S3, the automatic flow control principle in the automatic flow control module is as follows: When the flow control module detects an IP frame retransmission interruption, it indicates that the current transmission network is congested and frame loss occurs. The delay count time automatically increases by 1, extending the sending time between IP frames, reducing the sending rate of IP frames, and at the same time, the 1-second timer is reset to zero. When no IP frame retransmission interruption is detected within 1 second, it indicates that the current transmission network is unobstructed. The delay count time automatically decreases by 1, shortening the sending time between IP frames, increasing the sending rate of IP frames, and thus improving data transmission efficiency. When a network transmission error is detected, the delay count time is reset to zero, and flow control is performed again after the network is restored.
9. The method for implementing 10G network embedded data unloading based on UDP protocol by FPGA according to any one of claims 1 to 7, characterized in that: In step S4, the 10G-UDP protocol module is a network protocol implemented based on 10 Gigabit network and is used to complete the underlying logical interface protocol conversion.
Citation Information
Patent Citations
10-gigabit Ethernet TCP offload engine (TOE) system realized based on FPGA
CN105516191A
FPGA-based UDP / IP hardware protocol stack and implementation method
CN108462642A