Self-adaptive PCIe bandwidth data processing system

By dividing data into sub-data and using task units to compete for processing identifier updates, the problem of insufficient bandwidth of PCIe switch interconnection links is solved, and efficient data transmission and processing is achieved.

CN120596426AActive Publication Date: 2025-09-05MUXI INTEGRATED CIRCUIT (WUHAN) CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510673684.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-05
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

In the prior art, in scenarios where multiple GPU chips are interconnected, the actual bandwidth of the communication link interconnected by PCIe switches is difficult to reach a given reference bandwidth, resulting in low data transmission and processing efficiency.

Method used

The target data is divided into N sub-data, and the task units of the first and second communication links compete for application processing. The atomic add operation is used to update the processing identifier on the idle link, and the data processing path is adjusted in real time to avoid pre-allocation of data volume and improve bandwidth utilization.

Benefits of technology

It effectively improves the data transmission and processing efficiency in multi-GPU chip interconnection scenarios and improves bandwidth utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596426A_ABST
    Figure CN120596426A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data transmission, in particular to an adaptive PCIe bandwidth data processing system which divides target data into N sub-data, and a first task unit corresponding to a first communication link and a second task unit corresponding to a second communication link compete for processing the sub-data, when any one of the first communication link and the second communication link is in an idle state, the corresponding task unit updates the processing identifier, and the task unit processes the sub-data corresponding to the processing identifier after the update succeeds, so that the amount of data to be processed by different communication links does not need to be allocated in advance, and the processing efficiency is improved. And the sub-data is processed according to the real-time state of the communication link, so that the bandwidth utilization rate of each communication link can be effectively improved, and the data transmission and processing efficiency in a multi-GPU chip interconnection scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data transmission, and in particular to a data processing system with adaptive PCIe bandwidth. Background Art

[0002] At present, the demand for computing power is exploding in many fields such as scientific research, artificial intelligence, and big data analysis, and a single GPU chip is increasingly unable to meet the growing computing power demand. Therefore, existing technologies propose to form a GPU cluster composed of multiple GPU chips to provide more powerful computing power to cope with increasingly complex and large-scale computing tasks.

[0003] However, there is a need for communication between GPU chips in a GPU cluster. In the existing technology, multiple GPU chips can usually be connected to a PCIe switch to achieve interconnection between multiple GPU chips, and multiple GPU chips can be interconnected through preset lines to achieve interconnection between multiple GPU chips. In the existing method, two interconnection methods can be applied simultaneously to improve the efficiency of data transmission. However, applying the two interconnection methods simultaneously requires allocating the amount of data transmitted and processed on different communication links, so as to maximize bandwidth utilization and thereby improve the efficiency of data transmission and processing.

[0004] In the prior art, the amount of data allocated to a communication link interconnected by a PCIe switch is typically determined based on the reference bandwidth specified by the PCIe version. However, due to factors such as line interference, voltage fluctuations, and hardware connections, the actual bandwidth of the communication link interconnected by the PCIe switch is difficult to reach the given reference bandwidth. This makes it difficult for the communication link interconnected by the PCIe switch to complete the transmission and processing of the allocated data within the expected time, reducing bandwidth utilization and the efficiency of data transmission and processing.

[0005] Therefore, how to improve the efficiency of data transmission and processing in multi-GPU chip interconnection scenarios has become an urgent problem to be solved. Summary of the Invention

[0006] In view of the above technical problems, the technical solution adopted by the present invention is:

[0007] A data processing system with adaptive PCIe bandwidth, the system comprising: a PCIe switch, M GPU chips, a processor, and a memory storing a computer program, wherein the M GPU chips are interconnected via the PCIe switch to form a first communication link, and the M GPU chips are interconnected via a preset connection relationship to form a second communication link, wherein the first communication link corresponds to a first task unit, and the second communication link corresponds to a second task unit, and M is a positive integer. When the computer program is executed by the processor, the following steps are implemented:

[0008] S101, the target data A to be processed is divided into N sub-data {a1, a2, ..., a n ,…,a N}, where a n is the nth sub-data, n is a n The corresponding sub-data identifier, n is an integer in the range [1, N].

[0009] S102, when the first communication link is in an idle state, the first task unit updates the processing identifier. If the first task unit successfully updates the processing identifier, the sub-data corresponding to the sub-data identifier that is the same as the updated processing identifier is processed through the first communication link. The initial value of the processing identifier is 0.

[0010] S103, when the second communication link is in an idle state, the second task unit updates the processing identifier. If the second task unit successfully updates the processing identifier, the sub-data corresponding to the sub-data identifier that is the same as the updated processing identifier is processed through the second communication link.

[0011] S104, whenever the processing identifier is successfully updated, the updated processing identifier is detected, and if the updated processing identifier meets a preset condition, the processing of the target data A is completed.

[0012] The present invention has significant advantages over the prior art. By utilizing the above technical solution, the present invention provides an adaptive PCIe bandwidth data processing system that achieves considerable technological advancement and practicality, and has wide industrial application value. It has at least the following advantages:

[0013] The present invention divides the target data into N sub-data, and the first task unit corresponding to the first communication link and the second task unit corresponding to the second communication link compete to apply for processing the sub-data. When either the first communication link or the second communication link is in an idle state, the corresponding task unit updates the processing identifier. After the update is successful, the task unit processes the sub-data corresponding to the processing identifier. There is no need to pre-allocate the amount of data to be processed by different communication links. Instead, the sub-data is processed according to the real-time status of the communication link, thereby effectively improving the bandwidth utilization of each communication link, and further improving the efficiency of data transmission and processing in the multi-GPU chip interconnection scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0015] Figure 1 A schematic diagram of a flow chart of a computer program executed by a processor in a data processing system with adaptive PCIe bandwidth provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0017] This embodiment provides a data processing system with adaptive PCIe bandwidth. Figure 1 , is a schematic diagram of a flow chart of a computer program executed by a processor in a data processing system with adaptive PCIe bandwidth provided by an embodiment of the present invention. The system includes: a PCIe switch, M GPU chips, a processor, and a memory storing the computer program. The M GPU chips are interconnected via the PCIe switch to form a first communication link. The M GPU chips are interconnected via a preset connection relationship to form a second communication link. The first communication link corresponds to a first task unit, and the second communication link corresponds to a second task unit. M is a positive integer. When the computer program is executed by the processor, the following steps are implemented:

[0018] S101, the target data A to be processed is divided into N sub-data {a1, a2, ..., a n ,…,a N}, where a n is the nth sub-data, n is a n The corresponding sub-data identifier, n is an integer in the range [1, N];

[0019] S102: When the first communication link is in an idle state, the first task unit updates the processing identifier. If the first task unit successfully updates the processing identifier, the sub-data corresponding to the sub-data identifier that is the same as the updated processing identifier is processed via the first communication link. The initial value of the processing identifier is 0.

[0020] S103: When the second communication link is in an idle state, the second task unit updates the processing identifier. If the second task unit successfully updates the processing identifier, the sub-data corresponding to the sub-data identifier identical to the updated processing identifier is processed via the second communication link.

[0021] S104, whenever the processing identifier is successfully updated, the updated processing identifier is detected, and if the updated processing identifier meets a preset condition, the processing of the target data A is completed.

[0022] The PCIe switch is connected to M GPU chips respectively to implement data routing and exchange between the M GPU chips. In this embodiment, a single PCIe switch is used as an example. The number M of GPU chips should be less than or equal to the number of ports of the PCIe switch.

[0023] There is a preset connection relationship between every two GPU chips to support data transmission between GPU chips. A first communication link interconnecting M GPU chips can be formed through the PCIe switch, and a second communication link interconnecting M GPU chips can be formed through the preset connection relationship.

[0024] The first task unit may be used to apply for sub-data to be processed in the first communication link, and the second task unit may be used to apply for sub-data to be processed in the second communication link.

[0025] Specifically, in this embodiment, the data amounts corresponding to the N sub-data obtained by segmenting the target data may be consistent.

[0026] The state of the communication link may include an idle state and a busy state, and the processing identifier may be a global variable.

[0027] In a specific implementation, when the first communication link is in an idle state, the first task unit updates the processing identifier, including:

[0028] When the first communication link is in an idle state, the first task unit performs an atomic addition operation on the processing identifier at the current moment.

[0029] The atomic add operation is an indivisible operation that cannot be interrupted by operations in other threads or processes during its execution. Specifically, when an atomic add operation is performed on a process identifier, it adds a specified increment to the current value of the process identifier. In this embodiment, the increment is 1. This process is atomic, meaning that it is either completely executed successfully or not executed at all, with no partial execution.

[0030] In a specific embodiment, when the second communication link is in an idle state, the second task unit updates the processing identifier, including:

[0031] When the second communication link is in an idle state, the second task unit performs an atomic addition operation on the processing identifier at the current moment.

[0032] Among them, the first task unit and the second task unit can independently update the processing identifier. Since the update of the processing identifier is an atomic addition operation, even if the first task unit and the second task unit update the processing identifier at the same time, only one of the first task unit and the second task unit can successfully update the processing identifier. Accordingly, the sub-data corresponding to the sub-data identifier identical to the updated processing identifier is processed by the communication link corresponding to the task unit that successfully updates the processing identifier.

[0033] In a specific implementation, when there is no sub-data to be processed in the first communication link, the first communication link is in an idle state.

[0034] When there is no sub-data to be processed in the first communication link, the state of the first communication link may be an idle state; when there is sub-data to be processed in the first communication link, the state of the first communication link may be a busy state.

[0035] In a specific implementation, when there is no sub-data to be processed in the second communication link, the second communication link is in an idle state.

[0036] When there is no sub-data to be processed in the second communication link, the state of the second communication link may be an idle state; when there is sub-data to be processed in the second communication link, the state of the second communication link may be a busy state.

[0037] In a specific embodiment, processing the sub-data corresponding to the sub-data identifier identical to the updated processing identifier through the first communication link includes:

[0038] Processing the sub-data corresponding to the sub-data identifier identical to the updated processing identifier via the first communication link using a collective communication algorithm;

[0039] The processing of the sub-data corresponding to the sub-data identifier identical to the updated processing identifier through the second communication link includes:

[0040] The sub-data corresponding to the sub-data identifier identical to the updated processing identifier is processed through the second communication link using the collective communication algorithm.

[0041] The collective communication algorithms may include an All-Reduce algorithm, an All-Gather algorithm, a Reduce-Scatter algorithm, an All-to-All algorithm, and the like.

[0042] In a specific implementation, the preset condition is: the updated processing identifier is greater than N.

[0043] When the updated processing identifier is greater than N, it indicates that all sub-data are respectively allocated to the first communication link or the second communication link for data processing, and the allocation of sub-data can be stopped at this time.

[0044] In a specific implementation, step S104 further includes:

[0045] If the updated processing identifier does not meet the preset conditions, the process returns to step S102 and step S103.

[0046] Among them, when the updated processing identifier is less than or equal to N, it means that there is sub-data that has not been assigned to the first communication link or the second communication link for data processing. At this time, it is necessary to return to execute steps S102 and S103, that is, when either the first communication link or the second communication link is in an idle state, the task unit corresponding to the idle communication link applies to update the processing identifier to assign the sub-data to the idle communication link for data processing. It should be noted that steps S102 and S103 are parallel processing and there is no order of precedence, that is, when the first communication link is in an idle state, the first task unit updates the processing identifier, and when the second communication link is in an idle state, the second task unit updates the processing identifier.

[0047] In this embodiment, the target data is divided into N sub-data, and the first task unit corresponding to the first communication link and the second task unit corresponding to the second communication link compete to apply for processing the sub-data. When either the first communication link or the second communication link is in an idle state, the corresponding task unit updates the processing identifier. After the update is successful, the task unit processes the sub-data corresponding to the processing identifier. There is no need to pre-allocate the amount of data to be processed by different communication links. Instead, the sub-data is processed according to the real-time status of the communication link, thereby effectively improving the bandwidth utilization of each communication link, and further improving the efficiency of data transmission and processing in the multi-GPU chip interconnection scenario.

[0048] Although some specific embodiments of the present invention have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present invention. It should also be understood by those skilled in the art that various modifications may be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.

Claims

1. A data processing system with adaptive PCIe bandwidth, characterized in that: The system includes: a PCIe switch, M GPU chips, a processor, and a memory storing a computer program. The M GPU chips are interconnected through the PCIe switch to form a first communication link. The M GPU chips are interconnected through a preset connection relationship to form a second communication link. The first communication link corresponds to a first task unit, and the second communication link corresponds to a second task unit. M is a positive integer. When the computer program is executed by the processor, the following steps are implemented: S101, the target data A to be processed is divided into N sub-data {a1, a2, ..., a n ,…,a N }, where a n is the nth sub-data, n is a n The corresponding sub-data identifier, n is an integer in the range [1, N]; S102: When the first communication link is in an idle state, the first task unit updates the processing identifier. If the first task unit successfully updates the processing identifier, the sub-data corresponding to the sub-data identifier that is the same as the updated processing identifier is processed via the first communication link. The initial value of the processing identifier is 0. S103: When the second communication link is in an idle state, the second task unit updates the processing identifier. If the second task unit successfully updates the processing identifier, the sub-data corresponding to the sub-data identifier identical to the updated processing identifier is processed via the second communication link. S104, whenever the processing identifier is successfully updated, the updated processing identifier is detected, and if the updated processing identifier meets a preset condition, the processing of the target data A is completed.

2. The data processing system with adaptive PCIe bandwidth according to claim 1, wherein: When the first communication link is in an idle state, the first task unit updates the processing identifier, including: When the first communication link is in an idle state, the first task unit performs an atomic addition operation on the processing identifier at the current moment.

3. The data processing system with adaptive PCIe bandwidth according to claim 1, wherein: When the second communication link is in an idle state, the second task unit updates the processing identifier, including: When the second communication link is in an idle state, the second task unit performs an atomic addition operation on the processing identifier at the current moment.

4. The data processing system with adaptive PCIe bandwidth according to claim 1, wherein: When there is no sub-data to be processed in the first communication link, the first communication link is in an idle state.

5. The data processing system with adaptive PCIe bandwidth according to claim 1, wherein: When there is no sub-data to be processed in the second communication link, the second communication link is in an idle state.

6. The data processing system with adaptive PCIe bandwidth according to claim 1, wherein: The processing of the sub-data corresponding to the sub-data identifier identical to the updated processing identifier through the first communication link includes: Processing the sub-data corresponding to the sub-data identifier identical to the updated processing identifier via the first communication link using a collective communication algorithm; The processing of the sub-data corresponding to the sub-data identifier identical to the updated processing identifier through the second communication link includes: The sub-data corresponding to the sub-data identifier identical to the updated processing identifier is processed through the second communication link using the collective communication algorithm.

7. The data processing system with adaptive PCIe bandwidth according to claim 1, wherein: The preset condition is: the updated processing identifier is greater than N.

8. The data processing system with adaptive PCIe bandwidth according to claim 1, wherein: Step S104 further includes: If the updated processing identifier does not meet the preset conditions, the process returns to step S102 and step S103.

Citation Information

Patent Citations

  • Dual controller data communication method, device, equipment and readable storage medium

    CN108924008A

  • Multiprocessor interconnection system

    CN114968902A

  • Distributed training set communication control method and device and medium

    CN119336451A

  • Request processing method and device, storage medium and program product

    CN119544797A

  • Data transmission method, electronic equipment, storage medium and program product

    CN119583467A