A data processing method, system, computer device and storage medium

CN122817152APending Publication Date: 2026-09-25BEIJING KINGSOFT CLOUD NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510353050.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

而且,每个步骤都会产生额外时延t,总体额外时延将呈线性增长,即(N-1)t

Benefits of technology

[0025]本公开实施例提供的技术方案与现有技术相比具有如下优点:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817152A_ABST
    Figure CN122817152A_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of computers, in particular to a data processing method and system, computer device and storage medium; the method comprises: constructing a bidirectional tree structure corresponding to a GPU cluster; performing global reduction processing of data according to the bidirectional tree structure; the present scheme performs data reduction based on the bidirectional tree structure, which not only reduces the time complexity and delay in the reduction process, but also improves the utilization rate of the overall bandwidth and the training efficiency of the artificial intelligence model, and provides basic support for the training of larger and more complex artificial intelligence models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a data processing method, system, computer device, and storage medium. Background Technology

[0002] The Reduce-Scatter algorithm is a collective communication operation widely used in distributed computing, combining the Reduce and Scatter steps. In Graphics Processing Unit (GPU) clusters, the NVIDIA Collective Communications Library (NCCL) can utilize its efficient communication mechanism to execute the Reduce-Scatter algorithm.

[0003] During the execution of the Reduce-Scatter algorithm, the GPU cluster adopts a ring topology, in which all GPUs in the cluster are connected end to end to form a ring network. Data is transmitted along the ring path, so that each GPU can eventually obtain data from other GPUs, thereby achieving efficient interaction among multiple machines and multiple GPUs.

[0004] In the data processing described above, data needs to be transmitted through N-1 steps in a ring network (where N is the total number of GPUs), with a time complexity of O(N). Furthermore, each step introduces an additional latency t, and the total additional latency increases linearly, i.e., (N-1)t. As the size of the GPU cluster increases, the communication overhead will increase significantly, leading to a decrease in overall computational efficiency. Summary of the Invention

[0005] To address the aforementioned technical problems, this disclosure provides a data processing method, system, computer device, and storage medium.

[0006] In a first aspect, this disclosure provides a data processing method, the method comprising:

[0007] Construct a bidirectional tree structure corresponding to the GPU cluster; perform global data reduction processing according to the bidirectional tree structure.

[0008] In some alternative implementations, a bidirectional tree structure corresponding to the GPU cluster is constructed, including:

[0009] Obtain at least two pre-divided first communication domains of the GPU cluster, wherein the connection method of the GPU nodes in the first communication domain is a preset connection method; connect the at least two first communication domains through the network to form a second communication domain; construct at least one of the first communication domains or the second communication domain into a bidirectional tree structure.

[0010] In some alternative implementations, at least two pre-divided first communication domains of the GPU cluster are obtained, including:

[0011] Obtain the connection method of each GPU node in the GPU cluster; divide the GPU nodes with preset connection methods to form at least two first communication domains.

[0012] In some optional implementations, when the first communication domain is constructed as a ring topology and the second communication domain is constructed as a bidirectional tree structure, global data reduction processing is performed according to the bidirectional tree structure, including:

[0013] In the first communication domain, local data reduction is performed according to the ring topology; in the second communication domain, global data reduction is performed according to the bidirectional tree structure corresponding to the second communication domain.

[0014] In some optional implementations, when the first communication domain is constructed as a bidirectional tree structure and the second communication domain is constructed as a ring topology, global data reduction processing is performed according to the bidirectional tree structure, including:

[0015] In the first communication domain, local data reduction is performed according to the bidirectional tree structure corresponding to the first communication domain; in the second communication domain, global data reduction is performed according to the ring topology structure.

[0016] In some optional implementations, when both the first and second communication domains are constructed as bidirectional tree structures, global data reduction processing is performed according to the bidirectional tree structure, including:

[0017] In the first communication domain, local data reduction is performed according to the bidirectional tree structure corresponding to the first communication domain; in the second communication domain, global data reduction is performed according to the bidirectional tree structure corresponding to the second communication domain.

[0018] In some alternative implementations, after performing global reduction of the data according to the bidirectional tree structure, the method further includes:

[0019] Data is distributed and processed using a bidirectional tree structure.

[0020] In a second aspect, this disclosure provides a data processing system for performing the methods described in the first aspect and any corresponding embodiment thereof, the system comprising:

[0021] It integrates a communication library and a GPU cluster; the communication library is used to build a bidirectional tree structure corresponding to the GPU cluster; the GPU cluster is used for global data reduction processing according to the bidirectional tree structure.

[0022] Thirdly, this disclosure provides a computer device, including:

[0023] The memory and the processor are interconnected and communicate with each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the data processing method described in the first aspect and any embodiment thereof.

[0024] Fourthly, this disclosure provides a computer-readable storage medium storing computer instructions for causing a computer to perform the data processing method described in the first aspect and any embodiment thereof.

[0025] The technical solution provided in this disclosure has the following advantages compared with the prior art:

[0026] The data processing method provided in this embodiment constructs a bidirectional tree structure corresponding to the GPU cluster; performs global data reduction processing according to the bidirectional tree structure; data reduction based on the bidirectional tree structure not only reduces the time complexity and latency in the reduction process, but also improves the overall bandwidth utilization and the training efficiency of artificial intelligence models, providing basic support for the training of larger-scale and more complex artificial intelligence models. Attached Figure Description

[0027] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0028] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 A schematic diagram illustrating the data reduction process for a GPU cluster using a ring topology.

[0030] Figure 2 This is a structural connection diagram of the data processing system provided in the embodiments of this disclosure;

[0031] Figure 3 A schematic flowchart illustrating the data processing method provided in this embodiment of the disclosure;

[0032] Figure 4(a) is a schematic diagram of the structure corresponding to the down tree provided in the embodiment of this disclosure;

[0033] Figure 4(b) is a schematic diagram of the structure corresponding to the up tree provided in the embodiment of this disclosure;

[0034] Figure 5A structural link diagram of a computer device provided in an embodiment of this disclosure. Detailed Implementation

[0035] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0036] Numerous specific details are set forth in the following description to provide a thorough understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of this disclosure, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.

[0037] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0038] Before elaborating on the specific real-time methods of this solution, a brief introduction to the technical background and processing methods of related technologies is provided below.

[0039] Training artificial intelligence models is typically achieved through the deployment of large-scale GPU clusters. During training, to address the communication issues between GPU nodes within the cluster, parallel computation across multiple GPU nodes relies on a high-performance collective communication library. Specifically, this library depends on the communication primitives and interfaces provided to enable data exchange and synchronization between GPUs.

[0040] This section will use the Reduce-Scatter primitive from the NVIDIA Collective Communications Library (NCCL) as an example to illustrate the data reduction process in a GPU cluster. Typically, NCCL uses a ring topology approach for collective communication primitives like Reduce-Scatter. The ring topology approach means that the GPU cluster uses a ring topology structure during the data reduction process; all GPU nodes are connected end-to-end to form a ring network, and data is transmitted along the ring path, ensuring that each GPU node ultimately receives data from other GPU nodes, achieving efficient multi-machine, multi-GPU interaction.

[0041] like Figure 1 As shown, the GPU cluster consists of 4 GPU nodes, each corresponding to a rank. Figure 1 The data is categorized into rank1, rank2, rank3, and rank4. Each rank is divided into four data blocks, such as a0, a1, a2, and a3 for rank1. Figure 1 In the reduction process shown, each rank goes through n-1 steps from its initial state to the final state corresponding to the end of this reduction, where n is the number of GPU nodes in the GPU cluster. That is, the time complexity of the reduction process is O(n). As the size of the GPU cluster increases, its complexity also increases. In addition, each step introduces an extra latency t, so the total extra latency of a single reduction operation is (N-1)t, increasing linearly. This also means that the larger the cluster size, the more significantly the communication overhead will increase. Moreover, as the number of GPUs increases, the ring topology cannot fully utilize all available network bandwidth, resulting in a waste of bandwidth resources.

[0042] Therefore, when using the ring topology method to implement collective communication primitives such as Reduce-Scatter, it has high time complexity, long latency and low bandwidth utilization, thus becoming a bottleneck that limits the improvement of artificial intelligence model training efficiency.

[0043] Based on this, the present disclosure provides a data processing method, system, computer device, and storage medium that performs data reduction based on a bidirectional tree structure. This not only reduces the time complexity and latency in the reduction process but also improves the overall bandwidth utilization and the training efficiency of artificial intelligence models, providing fundamental support for the training of larger-scale and more complex artificial intelligence models.

[0044] This embodiment provides a data processing system, such as Figure 2As shown, the data processing system includes a collection communication library 201 and a GPU cluster 202. The collection communication library 201 is used to construct the bidirectional tree structure corresponding to the GPU cluster 202; the GPU cluster 202 is used to perform global data reduction processing according to the constructed bidirectional tree structure when implementing collection communication primitives such as Reduce-Scatter.

[0045] According to an embodiment of the present invention, a data processing method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0046] This embodiment provides a data processing method, applied to Figure 2 The data processing system shown is illustrated. The data processing method provided in this embodiment is specifically designed for scenarios such as distributed training, high-performance computing, and parallel algorithms, and is used to implement data reduction using collective communication algorithms such as reduce-distribute. Figure 3 This is a flowchart of a data processing method according to an embodiment of the present invention, such as... Figure 3 As shown, the process includes the following steps:

[0047] S301 constructs a bidirectional tree structure corresponding to the GPU cluster.

[0048] Specifically, the aggregated communication library in the data processing system first senses the physical topology of the GPU cluster, obtaining the hardware information of each GPU node and the connection methods between GPU nodes. The hardware information includes, but is not limited to, GPU model and PCI bus address, and the connection methods between GPU nodes include, but are not limited to, NVLink and PCIe connections. Then, the aggregated communication library constructs a bidirectional tree structure corresponding to the GPU cluster based on the sensed hardware information and connection methods. This embodiment of the present disclosure, by constructing a bidirectional tree structure corresponding to the GPU cluster in the aggregated communication library, breaks the linear constraint of transmitting data node by node in a traditional ring topology. This not only reduces the time complexity and latency of inter-CPU communication but also improves the overall bandwidth utilization by introducing a multi-communication domain parallel communication mode, allowing data to be transmitted simultaneously along multiple channels.

[0049] The bidirectional tree structure includes a down tree as shown in Figure 4(a) and an up tree as shown in Figure 4(b). In the down tree, the data transmission direction is the data broadcast direction from the parent node to the child node, while in the up tree, the data transmission direction is the data sending direction from the child node to the parent node. During data reduction, each child node sends its data blocks (i.e., the data processing results generated by the child node's reduction operation) to the parent node according to the data sending direction, so that the parent node can perform reduction based on the received data. During data distribution, i.e., after data reduction is completed, the parent node broadcasts the reduced data to its subordinate child nodes and their recursive child nodes according to the data broadcast direction.

[0050] It should be noted that the bidirectional tree structure in this embodiment is not the actual physical structure of the GPU cluster, but a virtual data transmission structure used when implementing collective communication primitives such as reduce-distribute.

[0051] In some alternative embodiments, to ensure the efficiency of data transmission, after constructing the bidirectional tree structure corresponding to the GPU cluster, it is also necessary to calculate the path parameters of each path in the bidirectional tree structure, such as bandwidth and latency, and select the optimal communication path based on the calculated path parameters so that the optimal communication path is selected first for data transmission in subsequent communication processes.

[0052] S302 performs global data reduction processing according to a bidirectional tree structure.

[0053] Specifically, after the bidirectional tree structure corresponding to the GPU cluster is constructed in the collective communication library, the GPU cluster performs global data reduction processing according to the bidirectional tree structure to realize the data reduction part in collective communication primitives such as reduction-dispersion.

[0054] In some alternative implementations, after S302, the present disclosure further includes: performing distributed processing of data according to the bidirectional tree structure.

[0055] Specifically, when the data processing method corresponding to the embodiments of this disclosure is used to implement the reduction-distribution communication primitive, after performing global reduction processing of data according to the up-tree in the bidirectional tree structure, the data is then distributed according to the down-tree in the bidirectional tree structure. Distribution processing means that the root node broadcasts the data reduction result to each of its child nodes and each recursive child node.

[0056] The data processing method provided in this embodiment constructs a bidirectional tree structure corresponding to the GPU cluster; performs global data reduction processing according to the bidirectional tree structure; data reduction based on the bidirectional tree structure not only reduces the time complexity and latency in the reduction process, but also improves the overall bandwidth utilization and the training efficiency of artificial intelligence models, providing basic support for the training of larger-scale and more complex artificial intelligence models.

[0057] In related technologies, since the aggregated communication library only perceives the topology of physical nodes, it often uses physical nodes as the basis for communication layering and partitioning. However, in the era of large-scale models, due to the iteration of multi-machine NVLink communication technology, this node-based communication layering method uses physical topology as the topology-aware basis for communication optimization, rather than using differences in communication capabilities as the basis for topology partitioning, resulting in low bandwidth utilization. Therefore, in some optional implementations, to solve the above problems, S301 has been refined, and S301 specifically includes:

[0058] S3011, acquire at least two first communication domains pre-divided by the GPU cluster.

[0059] In the first communication domain, the GPU nodes are connected in a preset way. This can also be understood as the first communication domain consisting of at least two GPU nodes, and the GPU nodes are connected through a preset connection method. Therefore, the first communication domain corresponds to the preset connection method.

[0060] Specifically, first, the connection method of each GPU node in the GPU cluster is obtained; then, the GPU nodes with preset connection methods are divided to form at least two first communication domains.

[0061] For example, in this example, the preset connection method is NVLink, and the first communication domain is the NVLink communication domain. The process of the collection communication library dividing the first communication domain is as follows: the collection communication library senses the NVLink connection method between GPU nodes in the GPU cluster and obtains all GPU nodes with NVLink connection methods from the GPU cluster. Then, all GPU nodes with NVLink connection methods obtained from the GPU cluster can be divided into at least two categories, and the GPU nodes included in each category form an NVLink communication domain. The division method can be based on the NVLink connection method between GPU nodes (such as direct connection or indirect connection), etc.

[0062] S3012, connect at least two first communication domains through a network to form a second communication domain.

[0063] Specifically, after obtaining the pre-divided first communication domain in the GPU cluster, each first communication domain is treated as a whole, and a second communication domain is formed by connecting at least two first communication domains through a network.

[0064] S3013, construct at least one of the first communication domain or the second communication domain into a bidirectional tree structure.

[0065] Specifically, for the first communication domain, the aggregated communication library first obtains at least one first communication domain by dividing the GPU cluster based on the perception of the preset connection methods; then, the aggregated communication library constructs the topology corresponding to each first communication domain. For the second communication domain, the topology corresponding to the second communication domain is implemented through the network connection methods between the first communication domains. Therefore, the construction of the second communication domain and the construction of the corresponding topology of the second communication domain are completed simultaneously. In this embodiment, at least one of the first or second communication domains needs to be constructed as a bidirectional tree structure. That is, the first communication domain can be constructed as a ring topology, and the second communication domain as a bidirectional tree structure. Alternatively, the first communication domain can be constructed as a bidirectional tree structure, and the second communication domain as a ring topology. Alternatively, both the first and second communication domains can be constructed as bidirectional tree structures. This embodiment achieves layering in the data processing process by constructing the first and second communication domains. By constructing at least one of the first and second communication domains as a bidirectional tree structure, it no longer relies on simple ring data transmission, but adopts a divide-and-conquer approach, decomposing the data communication process into more efficient recursive subproblems, thereby significantly reducing the time complexity of the communication process.

[0066] In some alternative implementations, when the first communication domain is constructed as a ring topology and the second communication domain is constructed as a bidirectional tree structure, the global reduction processing of data according to the bidirectional tree structure is as follows: in the first communication domain, the local reduction processing of data is performed according to the ring topology; in the second communication domain, the global reduction processing of data is performed according to the bidirectional tree structure corresponding to the second communication domain.

[0067] Specifically, when the first communication domain is constructed as a ring topology and the second communication domain is constructed as a bidirectional tree structure, during the reduction process, firstly, according to the ring topology corresponding to each first communication domain, local data reduction is performed in parallel in at least two first communication domains. After the local reduction operations in all first communication domains are completed, then in the second communication domain, based on the results of the local reduction in each first communication domain, global data reduction is performed according to the bidirectional tree structure corresponding to the second communication domain.

[0068] In some alternative implementations, when the first communication domain is constructed as a bidirectional tree structure and the second communication domain is constructed as a ring topology, the global reduction processing of data according to the bidirectional tree structure is as follows: in the first communication domain, the local reduction processing of data is performed according to the bidirectional tree structure corresponding to the first communication domain; in the second communication domain, the global reduction processing of data is performed according to the ring topology.

[0069] Specifically, similar to the previous implementation, when the first communication domain is constructed as a bidirectional tree structure and the second communication domain is constructed as a ring topology, during the reduction process, firstly, according to the bidirectional tree structure corresponding to each first communication domain, local data reduction is performed in parallel in at least two first communication domains. After the local reduction operations in all first communication domains are completed, then in the second communication domain, based on the results of the local reduction in each first communication domain, global data reduction processing is performed according to the ring topology corresponding to the second communication domain.

[0070] In some optional implementations, when both the first communication domain and the second communication domain are constructed as bidirectional tree structures, the global reduction processing of data according to the bidirectional tree structure is performed as follows: in the first communication domain, local reduction processing of data is performed according to the bidirectional tree structure corresponding to the first communication domain; in the second communication domain, global reduction processing of data is performed according to the bidirectional tree structure corresponding to the second communication domain.

[0071] Specifically, when both the first and second communication domains are constructed as bidirectional tree structures, during the reduction process, local data reduction is first performed in parallel in at least two first communication domains according to the bidirectional tree structure corresponding to each first communication domain. After the local reduction operations in all first communication domains are completed, global data reduction is then performed in the second communication domain according to the bidirectional tree structure corresponding to the second communication domain, based on the results of the local reduction in each first communication domain.

[0072] The data processing method provided in this disclosure is hierarchical according to communication domains, and at least one of the first and second communication domains is constructed into a bidirectional tree structure. The hierarchical approach enables parallel reduction of multiple first communication domains. The combination of parallel approach and bidirectional tree structure significantly reduces the communication overhead of the data processing system, reducing the communication time complexity from O(n) to O(log n). Moreover, based on communication domain awareness, the bandwidth of each communication domain is fully utilized, solving the performance limitation problem in the traditional ring topology method. It has good scalability, and its performance advantage becomes more significant as the system scale increases.

[0073] This invention also provides a computer device; please refer to [link / reference]. Figure 5 .like Figure 5As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 5 Take a processor 10 as an example.

[0074] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0075] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.

[0076] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0077] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0078] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.

[0079] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0080] In addition to the computer devices and computer-readable storage media described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps of the sound source localization method provided in any embodiment of this application.

[0081] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0082] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data processing method, characterized in that, The method includes: Construct a bidirectional tree structure corresponding to the GPU cluster; Global data reduction is performed according to the bidirectional tree structure described above.

2. The method according to claim 1, characterized in that, The bidirectional tree structure corresponding to the GPU cluster includes: Obtain at least two pre-divided first communication domains of the GPU cluster, wherein the connection method of the GPU nodes in the first communication domain is a preset connection method; At least two of the first communication domains are connected via a network to form a second communication domain; At least one of the first communication domain or the second communication domain is constructed as a bidirectional tree structure.

3. The method according to claim 1, characterized in that, The step of obtaining at least two pre-divided first communication domains of the GPU cluster includes: Obtain the connection method of each GPU node in the GPU cluster; GPU nodes with preset connection methods are divided into at least two first communication domains.

4. The method according to claim 2, characterized in that, When the first communication domain is constructed as a ring topology and the second communication domain is constructed as a bidirectional tree structure, the global data reduction processing according to the bidirectional tree structure includes: In the first communication domain, local data reduction processing is performed according to the ring topology; In the second communication domain, global data reduction is performed according to the bidirectional tree structure corresponding to the second communication domain.

5. The method according to claim 2, characterized in that, When the first communication domain is constructed as a bidirectional tree structure and the second communication domain is constructed as a ring topology, the global data reduction processing according to the bidirectional tree structure includes: In the first communication domain, local data reduction processing is performed according to the bidirectional tree structure corresponding to the first communication domain. In the second communication domain, global data reduction processing is performed according to the ring topology.

6. The method according to claim 2, characterized in that, When both the first communication domain and the second communication domain are constructed as bidirectional tree structures, the global data reduction processing according to the bidirectional tree structure includes: In the first communication domain, local data reduction processing is performed according to the bidirectional tree structure corresponding to the first communication domain. In the second communication domain, global data reduction is performed according to the bidirectional tree structure corresponding to the second communication domain.

7. The method according to claim 1, characterized in that, After performing global data reduction according to the bidirectional tree structure, the method further includes: Data is distributed according to the bidirectional tree structure.

8. A data processing system, characterized in that, The system is used to execute the method of any one of claims 1 to 7, and the system includes: a collection communication library and a GPU cluster; The collection communication library is used to construct a bidirectional tree structure corresponding to the GPU cluster; The GPU cluster is used to perform global data reduction processing according to the bidirectional tree structure.

9. A computer device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the data processing method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the data processing method according to any one of claims 1 to 7.