Data input and acceleration processing system and method based on Rocket IO

Through a RocketIO-based data input and acceleration processing system, the PCIe bus is used to directly transmit data, which solves the problems of low data transmission efficiency, delay and inaccurate data in the prior art, and achieves efficient and stable data transmission and processing.

CN120336236APending Publication Date: 2025-07-18JIANGSU HUACHUANG MICROSYSTEM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510207858.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing data transmission methods are inefficient, delayed and inaccurate, especially when data transmission in PCIe links, resulting in cumbersome processes and data loss.

Method used

Using a RocketIO-based data input and acceleration processing system, the RocketIO data input unit and acceleration processing unit are used to directly transmit data using the PCIe bus to avoid intermediate forwarding, reduce the transmission path and data copying times, and improve transmission efficiency in combination with DMA.

Benefits of technology

It reduces data transmission delay, improves data transmission efficiency and processing efficiency, reduces resource consumption of control units, and realizes efficient and stable data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336236A_ABST
    Figure CN120336236A_ABST
Patent Text Reader

Abstract

The invention discloses a data input and acceleration processing system based on Rocket IO. The system comprises a Rocket IO data input unit, a main control unit and an acceleration processing unit, the invention further discloses a data input and acceleration processing method based on the Rocket IO, which comprises the following steps of: inputting data through the optical fiber interface, performing Rocket IO protocol conversion, sending the data into the acceleration processing unit in a PCIe data packet form, notifying the main control unit by using a message mechanism after data transmission is completed, and controlling the acceleration processing unit through interaction of a control command with the acceleration processing unit. The acceleration processing unit executes acceleration instruction processing. Data input based on the Rocket IO interconnection protocol is achieved, the input data are directly transmitted to the acceleration processing unit through the PCIe bus, intermediate forwarding of the data through the control unit is avoided, the transmission path and the data copying frequency are reduced, and therefore the data transmission delay is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of data transmission and high-performance computing, and particularly relates to a data input and acceleration processing system and method based on RocketIO. Background Art

[0002] RocketIO is a high-speed serial data transmission channel with a custom protocol. It uses two pairs of differential pairs to send and receive data, and can achieve two simplex or one pair of full-duplex data transmissions. Among them, RocketIO channel binding is to bind several RocketIO channels into a consistent parallel channel by adding P characters to the transmitted data stream, thereby improving the data throughput rate. The RocketIO channel has the advantages of flexible channel quantity and convenient network topology definition, and also has the advantages of high real-time performance, high bandwidth, high reliability, high stability, and resistance to electromagnetic interference. In addition, the high-speed interconnect bus PCIe is used as a local bus to connect the processor and other external devices, and the two-way data transmission bandwidth of the PCIe4.0 protocol can reach up to 64GB / s at most.

[0003] Currently, there is a design method for a storage system based on PCIE data transmission. The CPU PCIE2.0 interface of the server motherboard is led out as a data exchange channel, and every 16 PCIE channels are used as a group corresponding to a physical address port; then the signals of these 16 PCIE transmission channels are branched into four paths and respectively correspond to four virtual addresses to form a virtual PCIE address channel; then the virtual PCIE address channel is connected to the input end of the storage FLASH controller, and the storage FLASH controller directly decodes the PCIE protocol and directly distributes and stores the data into the FLASH particles; when performing a read operation, the storage FLASH controller encodes the data in the FLASH with PCIE and directly places the data into the virtual PCIE address channel.

[0004] However, the existing data transmission methods currently have at least the following three problems: 1) Low data transmission efficiency; when the existing data transmission methods perform data transmission, processes such as branching the signals of the PCIE transmission channels into four paths and establishing a virtual PCIE address channel are required, resulting in a cumbersome data transmission process and low efficiency; 2) Data transmission latency; during the data transmission process of the existing data transmission methods, time is required for the storage FLASH controller to establish a transmission path, decode the PCIE protocol, and read data, resulting in data transmission latency; 3) Inaccurate data transmission: In the existing data transmission method, data is directly placed into the virtual PCIE address channel for data transmission across the entire PCIE link, that is, multi-channel data transmission is carried out simultaneously, which may result in data loss, thus leading to inaccurate data transmission. Summary of the Invention

[0005] In view of the above three problems, the object of the present invention is to propose a data input and acceleration processing system and method based on RocketIO. Through the RocketIO data input unit, the main control unit, and the acceleration processing unit, problems such as data transmission delay and low data transmission efficiency existing in the existing data transmission process are overcome. Data can be directly transmitted to the acceleration processing unit through the PCIe bus, thereby reducing the data transmission path and lowering the data transmission delay.

[0006] It is achieved through the following technical solutions: In the first aspect, a data input and acceleration processing system based on RocketIO is proposed, including: a RocketIO data input unit, which is used to receive data input from an external optical fiber, send the data to the acceleration processing unit, and notify the main control unit through a message mechanism to control the next program processing; a main control unit, which is used to interact control commands and data with the acceleration processing unit through the PCIe bus, and load the code for large-scale computing into the acceleration processing unit; an acceleration processing unit, which is used to process PCIe data packets, cache the data obtained from the PCIe data packets, and execute the code and control commands for large-scale computing sent by the main control unit. The system of the present invention is based on the data input of the RocketIO interconnection protocol, and directly transmits the input data to the acceleration processing unit through the PCIe bus, avoiding intermediate forwarding of data through the control unit, reducing the transmission path and the number of data copies, thereby lowering the data transmission delay.

[0007] Preferably, the DMA method is used for data transmission between the RocketIO data input unit and the acceleration processing unit. By using the DMA method for data transmission between the RocketIO data input unit and the acceleration processing unit, the data transmission efficiency and the overall performance of the system can be improved.

[0008] Preferably, the RocketIO data input unit includes a RocketIO protocol conversion subunit, a data processing subunit, a transmission control subunit, a memory subunit, and a PCIe interface subunit; wherein, the RocketIO protocol conversion subunit is used to receive the input data; the data processing subunit is used to write the received data into a preset buffer; the transmission control subunit is used to configure and control the data transmission channel; the memory subunit is used to cache the data; the PCIe interface subunit is used to connect to other devices. By performing RocketIO protocol conversion on the input data through the RocketIO data input unit, the data can be transmitted to the acceleration processing unit efficiently and stably.

[0009] Preferably, the transmission control subunit is used to provide multiple data transmission channels, and each data transmission channel has an independent register set, at least including a control status register, an address register, and a transmission length register. By using the transmission control subunit to provide multiple data transmission channels, the data transmission efficiency is further improved.

[0010] Preferably, the main control unit includes a processor subunit and a memory subunit; wherein, the processor subunit is used to interact with the acceleration processing unit for control commands and data, and load the code for large-scale computing into the acceleration processing unit; the memory subunit is used to store the programs running on the processor. Through the main control unit, the code and control commands for computing can be sent to the acceleration processing unit, thereby improving the data processing rate.

[0011] Preferably, the acceleration processing unit includes an acceleration processing subunit, a data processing subunit, a memory subunit, and a PCIe interface subunit; wherein, the acceleration processing subunit is used to load and execute the code for large-scale computing and execute acceleration processing instructions; the data processing subunit is used to process the data; the memory subunit is used to store the processed data; the PCIe interface subunit is used to connect to other devices. By executing the code for large-scale computing and acceleration processing instructions through the acceleration processing unit, and integrating acceleration computing components, the data processing efficiency can be improved.

[0012] Second aspect, a data input and acceleration processing method based on RocketIO is also proposed. The method includes the following steps: S1. Input data into the RocketIO data input unit through an external optical fiber. After the data is subjected to RocketIO protocol conversion by the RocketIO protocol conversion subunit, the data is sent to the acceleration processing unit in the form of PCIe data packets. After the data transmission is completed, the main control unit is notified through the message mechanism; S2. After receiving the notification that the data input to the acceleration processing unit is completed, the main control unit sends the code for large-scale calculation and control commands to the acceleration processing unit through the PCIe bus; S3. After receiving the PCIe data packets sent by the RocketIO data input unit, the code for large-scale calculation and control commands sent by the main control unit, the acceleration processing unit first loads and executes the code for large-scale calculation, and then executes the acceleration processing instruction according to the control command, accelerates the processing of the PCIe data packets, and writes the data obtained therefrom into the memory subunit for caching. By directly transmitting the input data from the RocketIO data input unit to the acceleration processing unit, the method of the present invention can avoid the intermediate forwarding of data through the control unit, reduce the data transmission path and the number of data copies, thereby reducing the data transmission delay and improving the data transmission efficiency.

[0013] Preferably, in step S2, the main control unit is composed of multiple execution threads, and an execution thread is allocated for each data transmission channel. By allocating an execution thread for each data transmission channel, the system performance can be improved through parallel processing, and the data can be transmitted efficiently and stably.

[0014] The beneficial effects of the present invention compared with the prior art are: The technical solution of the present invention converts the data input by the external optical fiber through the RocketIO data input unit after RocketIO protocol conversion, and directly inputs the data to the acceleration processing unit through the PCIe bus. At the same time, the main control unit interacts with the acceleration processing unit through the PCIe bus to exchange control commands and data, and the acceleration processing unit executes the acceleration processing instruction according to the control command, successfully avoiding the intermediate forwarding of data through the control unit, reducing the data transmission path and the number of data copies, improving the overall performance of the system, thereby reducing the data transmission delay and the resource consumption of the control unit, and improving the data transmission efficiency and the data processing efficiency. Description of the Drawings

[0015] Figure 1 It is a schematic structural diagram of a data input and acceleration processing system based on RocketIO; Figure 2 It is a flowchart of a data input and acceleration processing method based on RocketIO; Figure 3Schematic diagram of data transmission between the RocketIO data input unit and the acceleration processing unit; Figure 4 Flow chart of data transmission between the RocketIO data input unit and the acceleration processing unit. Detailed implementation

[0016] The following will describe in detail the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention.

[0017] As Figure 1 shown, it is a schematic diagram of the structure of a data input and acceleration processing system based on RocketIO, which includes a RocketIO data input unit 101, a main control unit 102, and an acceleration processing unit 103; as Figure 2 shown, it is a flow chart of a data input and acceleration processing method based on RocketIO. Combining Figure 1 and Figure 2 , after the data is input through the optical fiber interface and undergoes RocketIO protocol conversion, the data is sent to the acceleration processing unit in the form of PCIe data packets; after the data transmission is completed, the main control unit is notified using the message mechanism, and by interacting with the acceleration processing unit to control commands, the acceleration processing unit executes acceleration instruction processing, thereby improving the data transmission efficiency and data processing efficiency.

[0018] The system specifically includes the following: The RocketIO data input unit is used to receive data input from the external optical fiber, send the data to the acceleration processing unit, and notify the main control unit through the message mechanism to control the next program processing. Among them, the RocketIO data input unit includes: a RocketIO protocol conversion sub-unit, including a RocketIO controller, which is used to receive the data input through the external optical fiber, perform RocketIO protocol conversion, and output the data to the data processing sub-unit; the data processing sub-unit is used to write the received data into a preset buffer; the transmission control sub-unit is used to configure and control the data transmission status, form a data transmission channel with the acceleration processing unit, and write data to the acceleration processing unit through the DMA method, and notify the main control unit after the data transmission is completed; the memory sub-unit is used to cache the data; the PCIe interface sub-unit is used to connect to other devices; thus, the data can be efficiently and stably transmitted to the acceleration processing unit through the RocketIO data input unit.

[0019] The main control unit is used to interact control commands and data with the acceleration processing unit through the PCIe bus, and load the code for large-scale computing into the acceleration processing unit. Among them, the main control unit includes: a memory sub-unit for storing programs running on the processor; a processor sub-unit for interacting control commands and data with the acceleration processing unit, and loading the code for large-scale computing into the acceleration processing unit; through the main control unit, the code and control commands for computing can be sent to the acceleration processing unit, thereby improving the data processing rate.

[0020] The acceleration processing unit is used to process PCIe data packets, cache the data obtained from the PCIe data packets, and execute the code and control commands for large-scale computing sent by the main control unit. Among them, the acceleration processing unit includes: an acceleration processing sub-unit responsible for loading and executing the code for large-scale computing, and integrating acceleration computing components; a data processing sub-unit responsible for writing the received data into a preset buffer; a memory sub-unit for caching data; a PCIe interface sub-unit for connecting to other devices; by executing the code for large-scale computing and acceleration processing instructions in the acceleration processing unit, and integrating acceleration computing components, the data processing efficiency can be improved.

[0021] The method specifically includes the following steps: S1. Input data into the RocketIO data input unit through an external optical fiber, and after the data is completed with RocketIO protocol conversion by the RocketIO protocol conversion sub-unit, output it to the data processing sub-unit. The data processing sub-unit caches the data into the memory sub-unit, and then configures and controls the data transmission status through the transmission control sub-unit to form a data transmission channel between the RocketIO data input unit and the acceleration processing unit. Then, use the DMA method to convert the input data into the form of PCIe data packets through the PCIe EP device, that is, divide the data into multiple TLP data packets, and transmit them to the acceleration processing unit through the data transmission channel. When the data transmission is completed, the RocketIO data input unit notifies the main control unit of the completion of data transmission through the message mechanism. Among them, DMA, the full name is Direct Memory Access, is a data transmission method used to significantly improve the data transmission efficiency when processing a large amount of data.

[0022] In this embodiment, the PCIe EP device, i.e., the PCIe Endpoint, is an endpoint device for data transmission via the PCIe bus; the PCIe EP device is a component of the PCIe architecture system. The driver program is started by the processor subunit to scan the PCIe bus, and a unique address is assigned to each PCIe EP device. The PCIe EP device obtains the addresses of other PCIe EP devices by reading the configuration register, and then uses the PCIe EP device to split the input data into multiple TLP data packets for transmission, thereby further ensuring the accuracy and reliability of data transmission.

[0023] In this embodiment, PCIe, whose full name is Peripheral Component Interconnect Express, is a high-speed serial computer expansion bus standard used to link the host unit, the RocketIO data input unit, and the acceleration processing unit for efficient data transmission. Among them, the PCIe architecture system includes: the RC component, i.e., the Root Complex, which is an important component of PCIe and the transmission medium between the processor subunit and the memory subunit, the PCIe switch, and the PCIe EP device. It is used to generate TLP data packets and send them to the RocketIO data input unit, as well as parse the received TLP data packets and transmit the parsed data to the processor subunit. Among them, the TLP data packet, i.e., the Transaction Layer Packet, is the basic unit for data communication in the PCIe bus; the PCIe switch is used to expand the PCIe bus and connect multiple PCIe EP devices; and the PCIe EP device.

[0024] As Figure 3 shown, it is a schematic diagram of the structure for data transmission between the RocketIO data input unit and the acceleration processing unit. The figure includes: the processor subunit, the memory subunit, the RC component, the PCIe switch, the TLP1 data packet and the TLP2 data packet, three endpoint devices, i.e., PCIe EP0, PCIe EP1, and PCIe EP2, and the DMA engine, etc.; as Figure 4 shown, it is a flowchart of data transmission between the RocketIO data input unit and the acceleration processing unit. Combining Figure 3 and Figure 4, data transmission between the RocketIO data input unit and the acceleration processing unit is carried out by initiating DMA mode. The specific steps of the transmission include: S1. Pre-allocate a section of physically contiguous memory space through the acceleration processing unit and lock this section of memory space to prevent other modules from accessing this section of memory space; S2. Configure the DMA controller in the DMA engine built into the RocketIO data input unit through the main control unit. The configured parameters mainly include: the destination address of data transmission, the source address of data transmission, and the size of data transmission, and initiate a DMA transmission request to the DMA controller through the main control unit; S3. After the DMA controller receives the DMA transmission request, the DMA engine automatically divides the data into multiple TLP data packets according to the PCIe protocol and sends the multiple TLP data packets to the acceleration processing unit. Moreover, the multiple TLP data packets point to the start address of the locked section of memory space in the acceleration processing unit, and the valid data on the TLP data packets will be automatically filled into the memory space where the start address is located; S4. After the DMA transmission is completed, the RocketIO data input unit notifies the main control unit of the completion of data transmission through the message mechanism. Among them, the message mechanism means that the RocketIO data input unit sends an interrupt signal to the main control unit to notify the main control unit of the completion of data transmission.

[0025] In this embodiment, the transmission control subunit is used to provide multiple data transmission channels, and each data transmission channel has an independent register set, at least including a control status register, an address register, and a transmission length register; among them, the transmission control subunit embeds multiple data transmission channels, and under the control and scheduling of the processor subunit driver program, the transmission control subunit completes the configuration of multiple data transmission channels through the DMA engine for data transmission. By using the transmission control subunit to provide multiple data transmission channels, the data transmission efficiency is further improved.

[0026] S2. After the main control unit receives the notification of the completion of data transmission to the acceleration processing unit, load the code for large-scale computing to the acceleration processing unit through the PCIe bus by controlling the processor subunit and interact with the acceleration processing unit to send control commands to the acceleration processing unit. Among them, the main control unit consists of multiple execution threads, and the multiple execution threads correspond to multiple data transmission channels and computing engines. An execution thread can be assigned to each data transmission channel, so as to improve the system performance through parallel processing and enable the data to be transmitted efficiently and stably.

[0027] S3. After the acceleration processing unit receives multiple TLP data packets sent by the RocketIO data input unit and the code and control commands for large-scale computing sent by the main control unit, it first integrates the acceleration computing components through the acceleration processing subunit to cooperate with the acceleration processing subunit to complete various high-performance computing tasks. Among them, the acceleration computing components include GPUs or FPGAs. Secondly, it then loads and executes the code for large-scale computing through the acceleration processing subunit, and executes the acceleration processing instructions according to the control commands, accelerates the data processing subunit to process multiple TLP data packets, and writes the data obtained from the multiple TLP data packets into the memory subunit for caching, thereby improving the data processing efficiency.

[0028] In summary, after the present invention converts the data input from the external optical fiber through the RocketIO data input unit using the RocketIO protocol, it directly inputs the data into the acceleration processing unit through the PCIe bus. At the same time, the main control unit interacts with the acceleration processing unit through the PCIe bus to exchange control commands and data, and the acceleration processing unit executes the acceleration processing instructions according to the control commands, successfully avoiding the intermediate forwarding of data through the control unit, reducing the data transmission path and the number of data copies, improving the overall performance of the system, thereby reducing the data transmission delay and the resource consumption of the control unit, improving the data transmission efficiency and the data processing efficiency, and having significant progressiveness.

[0029] The above embodiments are only used to illustrate the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution according to the technical idea proposed by the present invention fall within the protection scope of the present invention.

Claims

1. A data input and acceleration processing system based on RocketIO, characterized in that, Including: A RocketIO data input unit, configured to receive data input from an external optical fiber, send the data to an acceleration processing unit, and notify a main control unit to control the next program processing through a message mechanism; A main control unit, configured to interact control commands and data with the acceleration processing unit through a PCIe bus, and load codes for large-scale computing into the acceleration processing unit; An acceleration processing unit, configured to process PCIe data packets, cache the data obtained from the PCIe data packets, and execute the codes for large-scale computing and control commands sent by the main control unit.

2. The data input and acceleration processing system based on RocketIO according to claim 1, wherein Data transmission is performed between the RocketIO data input unit and the acceleration processing unit by using the DMA method.

3. A data input and acceleration processing system based on RocketIO according to claim 1, characterized in that The RocketIO data input unit includes a RocketIO protocol conversion subunit, a data processing subunit, a transmission control subunit, a memory subunit, and a PCIe interface subunit; wherein, the RocketIO protocol conversion subunit is configured to receive the input data; the data processing subunit is configured to write the received data into a preset buffer; the transmission control subunit is configured to configure and control a data transmission channel; the memory subunit is configured to cache the data; and the PCIe interface subunit is configured to connect to other devices.

4. The data input and acceleration processing system based on RocketIO according to claim 3, wherein, The transmission control subunit is configured to provide multiple data transmission channels, and each data transmission channel has an independent register group. Any register group includes at least a control status register, an address register, and a transmission length register.

5. A data input and acceleration processing system based on RocketIO according to claim 1, characterized in that, The main control unit includes a processor subunit and a memory subunit; wherein, the processor subunit is configured to interact control commands and data with the acceleration processing unit, and load codes for large-scale computing into the acceleration processing unit; and the memory subunit is configured to store programs running on the processor.

6. The data input and acceleration processing system based on RocketIO according to claim 1, characterized in that The acceleration processing unit includes an acceleration processing subunit, a data processing subunit, a memory subunit, and a PCIe interface subunit; wherein, the acceleration processing subunit is configured to load and execute codes for large-scale computing and execute acceleration processing instructions; the data processing subunit is configured to process data; the memory subunit is configured to store the processed data; and the PCIe interface subunit is configured to connect to other devices.

7. A data input and acceleration processing method based on RocketIO, characterized in that, Using the RocketIO-based data input and acceleration processing system according to any one of claims 1-6, the method includes the following steps: S1. Input data into the RocketIO data input unit through an external optical fiber, and after completing the RocketIO protocol conversion of the data through the RocketIO protocol conversion subunit, send the data to the acceleration processing unit in the form of a PCIe data packet. After the data transmission is completed, notify the main control unit through a message mechanism; S2. After receiving the notification that the data input to the acceleration processing unit is completed, the main control unit sends codes for large-scale computing and control commands to the acceleration processing unit through the PCIe bus; S3. After the acceleration processing unit receives the PCIe data packet sent by the RocketIO data input unit, the code for large-scale computing sent by the main control unit, and the control command, it first loads and executes the code for large-scale computing, then executes the acceleration processing instruction according to the control command, accelerates the processing of the PCIe data packet, and writes the data obtained therefrom into the memory sub-unit for caching.

8. A data input and acceleration processing method based on RocketIO according to claim 7, characterized in that, In step S2, the main control unit consists of multiple execution threads and is used to allocate an execution thread for each data transmission channel.