Ensemble communication method, CXL switching equipment and computing equipment
The CXL switching device directly obtains and fuses gradient data in the distributed training system, solving the problem of repeated transmission of gradient data and achieving efficient distributed training.
Patent Information
- Application Number
- CN202510423441.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-08-19
AI Technical Summary
During the distributed training of neural network models, the repeated transmission of gradient data leads to high communication costs, excessive communication delay and overhead, which affects training efficiency.
The CXL switching device is used to directly obtain gradient data from the local memory of each computing node through the CXL protocol, calculate the fusion gradient and store it in shared memory. Each computing node establishes a communication connection with the shared memory through the CXL switching device, and directly obtains the fusion gradient update model parameters from the shared memory.
The communication rounds during model training are reduced, communication delay and overhead are reduced, and distributed training efficiency is improved.
Smart Images

Figure CN120508410A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a collective communication method, a CXL switching device, and a computing device. Background Art
[0002] In recent years, the field of artificial intelligence has rapidly developed, and neural network models have received increasing attention. As the parameters of deep learning-based network models grow larger and larger, and as the training datasets required to train these models grow larger and larger, a single computing node can no longer meet the computational requirements of most network model training processes.
[0003] Using a computing cluster containing multiple computing nodes to perform distributed training of deep learning-based target models is becoming the key to effectively solving the computing requirements of the target model training process.
[0004] However, using a computing cluster containing multiple computing nodes to perform distributed training of network models places high demands on the memory capacity of the computing nodes and the communication overhead between the computing nodes.
[0005] Currently, the distributed training model approach has the problem of repeated transmission of gradient data and high communication costs. Summary of the Invention
[0006] The embodiments of the present application provide a collective communication method, a CXL switching device, and a computing device, which reduce the communication rounds required for collective communication during model training, significantly reduce communication delays, and improve the efficiency of distributed training.
[0007] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:
[0008] In a first aspect, a collective communication method is provided, which is applied to a model training system. Multiple computing nodes in the model training system are used to perform distributed training on a target model. The model training system also includes at least one Compute Express Link (CXL) switch device. The CXL switch device is communicatively connected to the local memory of each of the multiple computing nodes via the CXL protocol. Each computing node is communicatively connected to the shared memory via the CXL switch device. The method is performed by the at least one CXL switch device and includes:
[0009] Obtain the gradient data in the local memory of each computing node based on the CXL protocol; the gradient data is obtained by the computing node training the locally deployed target model;
[0010] The fused gradient is calculated based on the gradient data of each computing node and stored in the shared memory; the fused gradient is used by each computing node to update the parameters of the locally deployed target model.
[0011] The CXL switch directly retrieves gradient data from each compute node's local memory via the CXL protocol, calculates the fused gradient, and stores it in shared memory. Because each compute node has a communication connection to the shared memory via the CXL switch, it can directly retrieve the fused gradient from the shared memory and subsequently update its own model parameters.
[0012] As can be seen, there is no need to repeatedly transmit gradient data, reducing the number of collective communication rounds required during model training. Furthermore, communication based on the CXL protocol eliminates the need to transmit data over the network, significantly reducing communication latency and improving the efficiency of distributed training.
[0013] In addition, since gradient fusion is performed directly on the CXL switching device, there is no need to send the aggregated gradient data of each computing node to a single service node. Compared with the method that requires gradient fusion at the service node, the communication overhead is further reduced.
[0014] In one possible implementation, the local memory of each computing node is a CXL memory managed by a CXL controller, and the CXL switch device is in communication with the CXL controller;
[0015] Obtain the gradient data in the local memory of each computing node based on the CXL protocol, including:
[0016] The CXL controller of each computing node receives gradient data sent in response to a first instruction. The first instruction is sent by the sending entity in response to each computing node being in a first state. The first state indicates that the computing node completes local calculation of the gradient data and stores the calculated gradient data in the local memory.
[0017] As can be seen, further configuration of the sending agent is possible, such as configuring the scheduling node as the sending agent. The sending agent can summarize the status of each compute node and promptly notify the CXL controller to send gradient data to the CXL switch. This coordinates communication between the CXL switch and the CXL controllers of each compute node, further reducing communication latency and improving the efficiency of collective communication, thereby improving the efficiency of distributed training.
[0018] In one possible implementation, calculating the fused gradient based on the gradient data of each computing node includes:
[0019] In response to a second instruction, a fused gradient is calculated based on the gradient data of each computing node. The second instruction is sent by the sending entity in response to the CXL controllers of each computing node being in a second state. The second state indicates that each CXL controller has completed sending the gradient data.
[0020] As can be seen, the sending entity can aggregate the status of the CXL controllers of each compute node. Once it confirms that all CXL controllers have sent gradient data, it can promptly notify the CXL switch to calculate the fused gradients. Further coordinating communication between the CXL switch and the CXL controllers of each compute node minimizes communication latency, improves the efficiency of collective communication, and ultimately enhances the efficiency of distributed training.
[0021] In a possible implementation, after storing the fused gradient in the shared memory, the following steps are further included:
[0022] A ready message is sent to the sending entity so that the sending entity instructs each computing node to obtain the fused gradient in the shared memory through the CXL switch device. The ready message is used to indicate that the CXL switch device has completed the calculation of the fused gradient.
[0023] The CXL switch promptly notifies the compute nodes of the completed fused gradient calculations, allowing the sending entity to promptly notify each compute node to obtain the fused gradients, thereby completing the model parameter update. Further coordinating communications between the CXL switch and the compute nodes minimizes communication latency, improves collective communication efficiency, and contributes to the efficiency of distributed training.
[0024] In one possible implementation, a model training system includes multiple CXL switching devices, which form a topology with N levels. N is a positive integer greater than or equal to 2. A CXL switching device at level 1 has a downstream port connected to the local memory of at least one computing node, and an upstream port connected to a CXL switching device at level 2. A CXL switching device at level i has a downstream port connected to at least one CXL switching device at level (i-1), and an upstream port connected to a CXL switching device at level (i+1). A CXL switching device at level N has a downstream port connected to at least one CXL switching device at level (N-1), and an upstream port connected to at least one computing node. i is a positive integer greater than 1 and less than N.
[0025] Obtain the gradient data in the local memory of each computing node based on the CXL protocol, including:
[0026] Read the gradient data in the local memory of each connected computing node through the first-level CXL switch device;
[0027] Calculate the fused gradient based on the gradient data of each computing node, including:
[0028] Calculate the first-layer fused gradient based on the read gradient data through the first-layer CXL switching device;
[0029] The CXL switching device at the i-th layer reads the fusion gradient of the i-1th layer calculated by the CXL switching device at the i-1th layer, and calculates the fusion gradient of the i-th layer;
[0030] The CXL switching device at the Nth layer reads the N-1th layer fusion gradient calculated by the CXL switching device at the N-1th layer, and calculates the fusion gradient.
[0031] As can be seen, multiple CXL switches can be used to form a multi-level topology. Through data exchange between CXL switches at different levels, the fused gradient is ultimately calculated. This is suitable for situations with a large number of computing nodes, minimizing the number of gradient data communication rounds and improving the efficiency of distributed training models.
[0032] In a second aspect, a collective communication method is provided, which is applied to a model training system. Multiple computing nodes in the model training system are used to perform distributed training on a target model. The model training system also includes at least one Compute Express Link (CXL) switch device. The CXL switch device is communicatively connected to the local memory of each of the multiple computing nodes via the CXL protocol. Each computing node is communicatively connected to the shared memory via the CXL switch device. The method is executed by a scheduling node and includes:
[0033] Sending a first instruction to each computing node, the first instruction is used to instruct the computing node to send gradient data in the local memory to the CXL switch device, where the gradient data is obtained by the computing node training the locally deployed target model;
[0034] When each computing node has sent gradient data, a second instruction is sent to the CXL switch device, where the second instruction is used to instruct the CXL switch device to calculate a fused gradient based on the gradient data of each computing node and store the fused gradient in the shared memory;
[0035] When the CXL switching device has completed the calculation of the fused gradient, a third instruction is sent to each computing node. The third instruction is used to instruct each computing node to read the fused gradient from the shared memory and use the fused gradient to update the parameters of the locally deployed target model.
[0036] As can be seen, there is no need to repeatedly transmit gradient data, reducing the number of collective communication rounds required during model training. Furthermore, communication based on the CXL protocol eliminates the need to transmit data over the network, significantly reducing communication latency and improving the efficiency of distributed training.
[0037] In addition, since gradient fusion is performed directly on the CXL switching device, there is no need to send the aggregated gradient data of each computing node to a single service node. Compared with the method that requires gradient fusion at the service node, the communication overhead is further reduced.
[0038] In a third aspect, a collective communication method is provided, which is applied to a model training system. Multiple computing nodes in the model training system are used to perform distributed training on a target model. The model training system also includes at least one Compute Express Link (CXL) switch device. The CXL switch device is communicatively connected to the local memory of each of the multiple computing nodes via the CXL protocol. Each computing node is communicatively connected to the shared memory via the CXL switch device. The method is executed by the computing node and includes:
[0039] Train the locally deployed target model to obtain gradient data;
[0040] Store gradient data in its own local memory;
[0041] The fused gradients in the shared memory are accessed through the CXL switch device, and the parameters of the locally deployed target model are updated based on the fused gradients. The fused gradients are calculated based on the gradient data read by the CXL switch device from the local memory of each computing node and stored in the shared memory.
[0042] As can be seen, each compute node only needs to train the locally deployed target model, store the gradient data obtained from the training in its own local memory, and then wait for the CXL switch to calculate the fused gradient. After confirming that the fused gradient has been stored in shared memory, the CXL switch can access the fused gradient in shared memory to obtain the fused gradient.
[0043] The steps required of compute nodes in obtaining fused gradients are relatively simple, reducing the requirements for compute nodes. The entire gradient fusion process is performed directly on the CXL switch, eliminating the need to send aggregated gradient data from each compute node to a single compute node or service node. This further reduces communication overhead compared to methods that require gradient fusion to be performed locally on the compute node or service node.
[0044] In a fourth aspect, a collective communication device is provided, which is applied to a CXL switching device in a model training system. Multiple computing nodes in the model training system are used to perform distributed training on a target model. The model training system also includes at least one computing fast link CXL switching device, which is connected to the local memory of each computing node in the multiple computing nodes via the CXL protocol; each computing node is connected to the shared memory via the CXL switching device; the device includes: a functional unit for executing any one of the methods provided in the first aspect, the actions performed by each functional unit being implemented by hardware or by hardware executing corresponding software implementations. For example, the collective communication device may include: an acquisition unit for acquiring gradient data in the local memory of each computing node based on the CXL protocol; the gradient data is obtained by the computing node training a locally deployed target model; a computing unit for calculating a fused gradient based on the gradient data of each computing node, and storing the fused gradient in the shared memory; the fused gradient is used by each computing node to update the parameters of the locally deployed target model.
[0045] In a fifth aspect, a CXL switching device is provided, comprising: a controller and a memory; the controller is coupled to the memory; the memory is used for computer program instructions; and the controller is used to call the computer program instructions in the memory to execute any one of the methods provided in the first aspect.
[0046] In the sixth aspect, a computing device is provided, comprising: a controller and a memory; the controller is coupled to the memory; the memory is used for computer program instructions; and the controller is used to call the computer program instructions in the memory to execute any one of the methods provided in the second or third aspect above.
[0047] In a seventh aspect, a model training system is provided, comprising: a plurality of computing nodes and at least one CXL switch; the CXL switch being communicatively connected to a local memory of each of the plurality of computing nodes via a CXL protocol; and each computing node being communicatively connected to a shared memory via the CXL switch;
[0048] Any one of the multiple computing nodes is configured to train a locally deployed target model to obtain gradient data; and the gradient data is stored in its own local memory;
[0049] The CXL switch obtains the gradient data in the local memory of each computing node based on the CXL protocol; calculates the fused gradient based on the gradient data of each computing node and stores the fused gradient in the shared memory;
[0050] Any one of the multiple computing nodes is further configured to access the fused gradient in the shared memory through the CXL switch device, and update the parameters of the locally deployed target model based on the fused gradient.
[0051] In an eighth aspect, a computer-readable storage medium is provided, which stores computer execution instructions. When the computer execution instructions are executed on a computing device, the computing device executes any one of the methods provided in the first aspect above.
[0052] In a ninth aspect, a computer program product is provided, comprising: computer execution instructions, which, when executed on a computing device, cause the computing device to execute any one of the methods provided in the first aspect.
[0053] Among them, the technical effects brought about by any implementation method in the second to ninth aspects can refer to the technical effects brought about by different implementation methods in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 A schematic diagram of distributed training provided in an embodiment of the present application;
[0055] Figure 2 A schematic diagram of the structure of the model training system provided in an embodiment of the present application;
[0056] Figure 3 A schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0057] Figure 4 A schematic diagram of a flow chart of a collective communication method provided in an embodiment of the present application;
[0058] Figure 5 A signaling interaction diagram of a collective communication method provided in an embodiment of the present application;
[0059] Figure 6 A schematic diagram of the communication rounds of the collective communication method provided in an embodiment of the present application;
[0060] Figure 7 A schematic diagram of a collective communication method provided in an embodiment of the present application;
[0061] Figure 8 A schematic diagram of a flow chart of a collective communication method applied to a scheduling node provided in an embodiment of the present application;
[0062] Figure 9 A schematic diagram of a flow chart of a collective communication method applied to computing nodes provided in an embodiment of the present application;
[0063] Figure 10 A schematic diagram of the structure of a collective communication device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0064] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0065] In the description of this application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship, for example, A / B can represent A or B; "and / or" in this application is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural.
[0066] Furthermore, in the description of this application, unless otherwise specified, "plurality" means two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.
[0067] In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit differences. At the same time, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way for easy understanding.
[0068] The following is a brief introduction to the relevant terms involved in the embodiments of this application.
[0069] CXL, short for Compute Express Link, is a new high-speed interconnect technology for high-bandwidth, low-latency device interconnection. It aims to provide higher data throughput and lower latency to meet the needs of modern computing and storage systems. It includes three protocols: CXL.io, CXL.memory, and CXL.cache. The CXL.io protocol is an enumeration configuration protocol primarily used for device discovery, enumeration, and error reporting. The CXL.memory protocol allows computing devices to access the memory of CXL-supported memory devices as if they were their own local memory. In the CXL.memory protocol, the CPU, called the master, is responsible for sending requests, while the CXL memory device, acting as a slave, responds to the master's requests. Requests are categorized as either data-carrying or data-free, and responses are categorized as either data-carrying or data-free. The CXL.cache protocol enables devices to resolve memory consistency issues.
[0070] Collective communication: Common collective communication operations include gather, scatter, allgather, reduce, and allreduce. For example, the gather operation gathers data from multiple nodes to a single node; the scatter operation sends data from a single source node to multiple destination nodes; the allgather operation gathers data from all nodes and distributes it to each node; the reduce operation performs an operation (such as sum, product, or maximum value) on the elements of a dataset at a single node, reducing the dataset to a single value; and the allreduce operation executes the reduce operation on all nodes and distributes the result to all nodes.
[0071] Large model: A large model, also known as a foundation model, refers to a machine learning model with a large number of parameters and a complex structure. It can process massive amounts of data and complete various complex tasks, such as natural language processing, computer vision, and speech recognition.
[0072] The memory wall refers to the mismatch between memory performance and processor performance, resulting in memory becoming a performance bottleneck. As neural network models grow in size and complexity, memory requirements increase significantly. However, memory speeds have not kept pace with processor speeds, forcing the processor to constantly wait for memory data transfers, limiting its full computing power.
[0073] Communication barriers: Distributed training is often used in large model training due to the large number of model parameters and the high computational complexity. Frequent exchange of data and parameters between nodes is limited by communication overhead and can lead to communication delays. For large model training, this can result in significant performance losses.
[0074] As mentioned earlier, using a computing cluster consisting of multiple computing nodes to perform distributed training of deep learning-based target models is becoming the key to effectively solving the computing requirements of the target model training process.
[0075] However, distributed training approaches face the aforementioned memory and communication barriers. While optimization algorithms have been introduced to mitigate these issues, the repeated transmission of gradient data and the high communication costs still remain.
[0076] In view of the above situation, an embodiment of the present application provides a collective communication method, which is applied to a model training system, in which multiple computing nodes of the model training system are used to perform distributed training on a target model. The model training system also includes at least one CXL switching device. The CXL switching device is connected to the local memory of each computing node in the multiple computing nodes through the CXL protocol; each computing node is connected to the shared memory through the CXL switching device; the CXL switching device obtains the gradient data in the local memory of each computing node based on the CXL protocol; the gradient data is obtained by the computing node training the locally deployed target model; the fused gradient is calculated based on the gradient data of each computing node, and the fused gradient is stored in the shared memory; the fused gradient is used for each of the computing nodes to update the parameters of the locally deployed target model.
[0077] The CXL switch directly retrieves gradient data from each compute node's local memory via the CXL protocol, calculates the fused gradient, and stores it in shared memory. Because each compute node has a communication connection to the shared memory via the CXL switch, it can directly retrieve the fused gradient from the shared memory and subsequently update its own model parameters.
[0078] As can be seen, there is no need to repeatedly transmit gradient data, reducing the number of collective communication rounds required during model training. Furthermore, communication based on the CXL protocol eliminates the need to transmit data over the network, significantly reducing communication latency and improving the efficiency of distributed training.
[0079] In addition, since gradient fusion is performed directly on the CXL switching device, there is no need to send the aggregated gradient data of each computing node to a single service node. Compared with the method that requires gradient fusion at the service node, the communication overhead is further reduced.
[0080] The following is an illustrative introduction to the application scenarios of the embodiments of the present application.
[0081] During distributed model training, gradient data is transmitted through collective communication, with the goal of allowing each computing node to ultimately obtain the fused gradient and subsequently update the local model. To reduce the number of communication rounds in this process, an embodiment of the present application provides a collective communication method for application in a model training system.
[0082] The following combination Figure 1 , a brief introduction to distributed training is given below. The basic process includes: (1) Divide the entire training dataset into multiple small batches and distribute these small batches to different devices. Each device has a complete copy of the model and processes the assigned data independently. During training, the device performs forward propagation, loss calculation, backpropagation and other operations. (2) Aggregate the gradient data calculated by each device, calculate the average gradient, and use it to update the network parameters. The above process is iterated until the training is completed.
[0083] The model training system provided in the embodiment of the present application includes multiple computing nodes and at least one CXL switching device. Figure 2 , Figure 2 In the embodiment shown, three computing nodes and one CXL switch device are included. Figure 2 The local memory of each computing node is also shown, and the CXL switch device communicates with the local memory of each computing node through the CXL protocol. Figure 2 The figure also shows a shared memory, and each computing node is connected to the shared memory through a CXL switch device.
[0084] It should be noted that in the embodiment of the present application, the local memory of the computing node refers to the memory directly connected to the computing node, for example, a memory stick connected to the memory slot of the motherboard of the computing node.
[0085] Figure 3 A schematic diagram of the structure of a computing device provided in an embodiment of the present application.
[0086] Need to explain, Figure 3 The system architecture shown is merely an example and does not constitute a limitation on the system architecture of the computing device provided in the embodiments of the present application.
[0087] In the embodiments of the present application, the computing device may specifically be a network device. The network device may include a server, etc. The server may be a single physical server, or two or more physical servers sharing different responsibilities and cooperating to implement various server functions.
[0088] For example, the server may be a blade server, a high-density server, a rack server, or a tower server, etc. The terminal device may include a personal digital assistant (PDA), an ultra-mobile personal computer (UMPC), a notebook computer, a netbook, a desktop computer, an all-in-one computer, etc.
[0089] The hardware of the computing device includes a processor, a basic input / output system (BIOS) chip, an out-of-band controller, and memory, while the software mainly includes the BIOS, an out-of-band management module, and an operating system (OS). Figure 3 shown.
[0090] A processor may include a central processing unit (CPU), which includes one or more CPU cores. The CPU's data processing operations are all performed by the CPU cores. The more CPU cores a CPU includes, the faster it can process data.
[0091] The BIOS chip is a chip installed on the motherboard that initializes and detects various hardware components during the computer's startup process. The BIOS chip includes a flash memory area.
[0092] The out-of-band management module is located in the out-of-band controller, and the operating system is located in the processor.
[0093] An out-of-band management module can be a management unit for non-business modules. For example, an out-of-band management module can remotely maintain and manage a computing device through a dedicated data channel. This out-of-band management module is completely independent of the computing device's operating system and can communicate with the BIOS and operating system through the computing device's out-of-band management interface.
[0094] Exemplarily, the out-of-band management module may include a management unit for the computing device's operating status, a management system in a management chip, a baseboard management controller (BMC) for the computing device, a system management module (SMM), etc. It should be noted that the embodiments of the present application do not limit the specific form of the out-of-band management module, and the above description is merely exemplary.
[0095] The operating system (OS) is a computer program that manages and controls the hardware and software resources of a computing device. Any other software must be supported by the OS to run. After a computing device is powered on, the BIOS first performs a series of operations, including self-tests and initialization, and then boots the OS, allowing the user to use the computing device normally.
[0096] BIOS is a set of programs embedded in the BIOS chip on the motherboard of a computing device. The main function of BIOS is to provide the lowest-level and most direct hardware settings and control for the computing device.
[0097] Memory, also known as internal storage or main memory, is installed in memory slots on the motherboard of a computing device.
[0098] It should be noted that the system architecture and application scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0099] For ease of understanding, the collective communication method provided in the embodiment of the present application is exemplarily introduced below in combination with the above system architecture and accompanying drawings.
[0100] See also Figure 4 , is a flow chart of a collective communication method provided in an embodiment of the present application, the method being executed by at least one CXL switching device, and the method comprising:
[0101] S401: Obtain the gradient data in the local memory of each computing node based on the CXL protocol; the gradient data is obtained by the computing node training the locally deployed target model.
[0102] During the model training process, each computing node trains the locally deployed target model based on local training data, obtains gradient data, and stores the gradient data in its own local memory.
[0103] Exemplarily, the computational node training process is as follows: the computational node inputs sample data from the training data into a locally deployed target model to obtain prediction data corresponding to the sample data. The node then uses a preset loss function to calculate the prediction loss based on the difference between the prediction data corresponding to the sample data and the labeled data in the training data. The node then calculates the partial derivatives of the prediction loss with respect to each model parameter to be trained in the target model to obtain the gradient data for each model parameter. The preset loss function can be of any type, including, but not limited to, a cross-entropy loss function, a mean squared error loss function, and the like.
[0104] Subsequently, the CXL switch device obtains the gradient data in the local memory of each computing node through the CXL protocol. Specifically, the local memory of each computing node can be pre-configured as a CXL memory device.
[0105] CXL memory devices include the CXL controller and the storage resources controlled by the CXL controller. The local memory of each compute node can be considered the storage resources controlled by the CXL controller. That is, the local memory of each compute node is managed by the CXL controller, and the CXL switch device communicates with the CXL controller.
[0106] The CXL controller can be a CXL memory expansion card, also known as a CXL control chip, a CXL memory controller (CXL memory expander controller), or a CXL memory pooling chip. The CXL controller can receive memory requests from external devices and provide the external devices with memory resources provided by the memory it controls.
[0107] In the embodiment of the present application, the CXL switch device is connected to the local memory of each computing node through the CXL controller, and thus can obtain the gradient data in the local memory of each computing node through the CXL controller.
[0108] Specifically, gradient data can be obtained based on the CXL.memory protocol mentioned above. Through the CXL.memory protocol, the CXL switch can access the memory resources controlled by the CXL controller (i.e., the local memory of each compute node) as if it were accessing its own local memory.
[0109] S402: Calculate the fused gradient based on the gradient data of each computing node, and store the fused gradient in the shared memory; the fused gradient is used by each computing node to update the parameters of the locally deployed target model.
[0110] In the embodiment of the present application, the CXL switch device may be programmable and support expansion of computing power and storage capacity. The CXL switch device may be pre-expanded to a device with stronger computing power and storage capacity, and gradient data fusion may be directly implemented on the CXL switch device.
[0111] A CXL switching device, also known as a CXL switch, generally includes a CXL control switch chip, a first port and a second port. The first port is an upstream switch port (USP), and the second port is a downstream switch port (DSP).
[0112] Exemplarily, the computing node is connected to an upstream port of a CXL switch device; and the local memory of the computing node is connected to a downstream port of the CXL switch device via a CXL controller.
[0113] For easier understanding, see Figure 2 , Figure 2 Figure 3 shows three compute nodes: compute node 1, compute node 2, and compute node 3. a1, b1, and c1 represent the gradient data calculated by compute node 1 and stored in local memory 1 of compute node 1. a2, b2, and c2 represent the gradient data calculated by compute node 2 and stored in local memory 2 of compute node 2. a3, b3, and c3 represent the gradient data calculated by compute node 3 and stored in local memory 3 of compute node 3.
[0114] Among them, a1, a2 and a3 are the gradient data corresponding to the model parameter A in the target model, b1, b2 and b3 are the gradient data corresponding to the model parameter B in the target model, and c1, c2 and c3 are the gradient data corresponding to the model parameter C in the target model.
[0115] As mentioned above, the CXL switch device is connected to the local memory of each computing node through the CXL controller, and can directly obtain data in the local memory of each computing node through the CXL protocol.
[0116] Figure 2 In the illustrated embodiment, the CXL switch device obtains gradient data from the local memory of the computing nodes host1, host2, and host3 respectively through the CXL protocol, and directly calculates the fused gradient on the CXL switch device. Figure 2 In the illustrated embodiment, a1+a2+a3 may represent the fused gradient calculated for model parameter A, b1+b2+b3 may represent the fused gradient calculated for model parameter B, and c1+c2+c3 may represent the fused gradient calculated for model parameter C.
[0117] It should be noted that Figure 2 This is for illustrative purposes only. Other calculation methods may also be used to calculate the fused gradient, such as summing and averaging the gradients of the model parameters for each computing node to obtain the fused gradient for the model parameter. The present embodiment does not limit the calculation algorithm for the fused gradient.
[0118] The CXL switch then stores the calculated fused gradients in shared memory.
[0119] In the embodiment of the present application, the shared memory may include any one of the following: a local memory of a CXL switch device, an extended memory of a CXL switch device, and a CXL memory device; the CXL memory device refers to a memory device managed by a CXL controller.
[0120] The local memory of a CXL switch refers to the memory directly connected to the switch, such as a memory module directly connected to the switch. Extended memory refers to additional memory in addition to the local memory, such as a memory module connected to the switch via an expansion slot.
[0121] CXL memory devices are memory devices managed by the CXL controller, such as the local memory of any of the aforementioned compute nodes. However, the local memory of a single compute node typically has limited storage capacity and cannot be used as shared memory.
[0122] For example, the model training system provided in embodiments of the present application may include a compute node with a large local memory, such as compute node host x. The local memory of compute node host x can be used as shared memory, that is, a CXL switch device stores fused gradients in the local memory of host x based on the CXL protocol.
[0123] In the embodiment of the present application, each computing node can obtain fused gradient data from the shared memory through the CXL switch device, and update the parameters of the locally deployed target model according to the fused gradient.
[0124] The CXL architecture allows processors to communicate directly with different types of memory and storage devices over high-speed serial peripheral component interconnect express (PCIe) channels.
[0125] As a possible implementation of the present invention, the standard PCIe address space can be used for addressing. Specifically, although CXL extends the functionality of PCIe, basic memory addressing can still rely on the PCIe address space. In the PCIe configuration space, each device has a unique address, which can usually be determined by the bus number, device number, and function number.
[0126] In an embodiment of the present application, the PCIe address of the shared memory can be notified to each computing node in advance. Therefore, after the computing node confirms that the fused gradient has been stored in the shared memory, it accesses the PCIe address of the shared memory through the CXL protocol to obtain the fused gradient.
[0127] As a possible implementation of the present invention, shared memory can also be accessed based on the memory addressing mode newly introduced by CXL. Specifically, CXL introduces a new memory addressing mode, especially for direct memory access of devices such as persistent memory and field programmable gate arrays (FPGAs).
[0128] In an embodiment of the present application, the extended function of CXL can be used to pre-configure the shared memory as a persistent memory area so that it can be directly accessed by the processor of each computing node without going through a traditional memory controller.
[0129] It should be noted that the above is only an exemplary description of computing nodes accessing shared memory, and the embodiments of the present application are not limited to this.
[0130] In the embodiment of the present application, since there are usually multiple parameters to be trained in the model and each parameter corresponds to a fused gradient, the fused gradient calculated by the CXL switching device is actually a fused gradient set, that is, a fused gradient containing multiple parameters to be trained.
[0131] When each computing node updates the parameters of the local target model, it updates the corresponding parameters to be trained according to the fusion gradient of each parameter to be trained.
[0132] The CXL switch uses the CXL protocol to retrieve gradient data from each compute node's local memory, perform fusion calculations, and store it in shared memory. Each compute node then accesses the fused gradients from shared memory through the CXL controller. This allows each compute node to obtain the fused gradients with just one round of CXL communication.
[0133] It can be seen that by applying the collective communication method provided in the embodiment of the present application, the CXL switching device is connected to the local memory of each computing node in the multiple computing nodes through the CXL protocol; each computing node is connected to the shared memory through the CXL switching device; the CXL switching device obtains the gradient data in the local memory of each computing node based on the CXL protocol; the gradient data is obtained by the computing node training the locally deployed target model; the fused gradient is calculated based on the gradient data of each computing node, and the fused gradient is stored in the shared memory; the fused gradient is used by each computing node to update the parameters of the locally deployed target model.
[0134] There's no need to repeatedly transmit gradient data, reducing the number of collective communication rounds required during model training. Furthermore, communication based on the CXL protocol eliminates the need to transmit data over the network, significantly reducing communication latency and improving the efficiency of distributed training.
[0135] In addition, since gradient fusion is performed directly on the CXL switching device, there is no need to send the aggregated gradient data of each computing node to a single service node. Compared with the method that requires gradient fusion at the service node, the communication overhead is further reduced.
[0136] In addition, the local memory and shared memory of each computing node exist independently, and there is no need to pool the local memory and shared memory of each computing node. Compared with the solution of pooling multiple memory devices, there is no need for additional memory pool management.
[0137] In the embodiment of the present application, a scheduling node may also be provided in the model training system, and the scheduling node is responsible for the overall scheduling of collective communications.
[0138] As an example, a CXL switching device obtains gradient data in the local memory of each computing node based on the CXL protocol, specifically including: receiving gradient data sent by the CXL controller of each computing node in response to a first instruction; the first instruction is sent by the sending entity in response to each computing node being in a first state; the first state indicates that the computing node completes local calculation of the gradient data and stores the calculated gradient data in the local memory.
[0139] The first instruction may be sent by a scheduling node. For example, after completing local calculations of gradient data, each compute node notifies the scheduling node. Upon receiving the notification, the scheduling node determines that the compute node is in the first state. Once the scheduling node determines that all compute nodes are in the first state, it sends a first instruction to the CXL controller of each compute node, instructing the CXL controller of each compute node to send the gradient data to the CXL switch.
[0140] It can be seen that in the embodiment of the present application, a scheduling node can be configured. The scheduling node can summarize the status of each computing node and then promptly notify the CXL controller to send gradient data to the CXL switching device, thereby coordinating the communication between the CXL switching device and the CXL controller of each computing node, further reducing communication delays, improving the efficiency of collective communication, and helping to improve the efficiency of distributed training.
[0141] In the embodiment of the present application, calculating the fused gradient based on the gradient data of each computing node may specifically include: calculating the fused gradient based on the gradient data of each computing node in response to a second instruction; the second instruction is sent by the sending entity in response to the CXL controllers of each computing node being in the second state; the second state indicates that each CXL controller has completed sending the gradient data.
[0142] Specifically, after each compute node's CXL controller finishes sending gradient data to the CXL switch, it can notify the scheduling node. Upon receiving the notification, the scheduling node determines that the compute node's CXL controller is in the second state. Once the scheduling node determines that all compute node CXL controllers are in the second state, it sends a second instruction to the CXL switch, instructing the CXL switch to begin calculating the fused gradient.
[0143] As can be seen, in this embodiment of the present application, the scheduling node aggregates the status of the CXL controllers of each compute node. Once it confirms that each CXL controller has sent gradient data, it promptly notifies the CXL switch to calculate the fused gradient. This further coordinates communication between the CXL switch and the CXL controllers of each compute node, minimizing communication latency and improving the efficiency of collective communication, thereby enhancing the efficiency of distributed training.
[0144] In this embodiment of the present application, after storing the fused gradient in the shared memory, the CXL switch further includes: sending a ready message to a sending entity, so that the sending entity instructs each compute node to obtain the fused gradient from the shared memory through the CXL switch. The ready message indicates that the CXL switch has completed calculation of the fused gradient.
[0145] Specifically, to ensure that compute nodes obtain the fused gradients as quickly as possible, the CXL switch stores the fused gradients in shared memory and sends a ready message to the scheduling node. Upon receiving the ready message, the scheduling node instructs each compute node to obtain the fused gradients.
[0146] Exemplarily, the scheduling node may send a third instruction to each computing node. After receiving the third instruction, each computing node obtains the fused gradient in the shared memory through the CXL switching device, and then completes the parameter update of the target model deployed by itself based on the fused gradient.
[0147] As can be seen, the CXL switch can promptly notify the completion of the fused gradient calculation, and the scheduling node can promptly notify each compute node to obtain the fused gradient, thereby completing the model parameter update. Further coordinating communication between the CXL switch and each compute node can minimize communication latency, improve the efficiency of collective communication, and help improve the efficiency of distributed training.
[0148] The following is combined with Figure 5 , taking three computing nodes as an example, the collective communication method provided in the embodiment of the present application is further explained.
[0149] like Figure 5 As shown, the following steps are included:
[0150] S501: The scheduling node notifies each computing node to send the gradient data in the local memory to the CXL switching device.
[0151] For example, after each computing node obtains local gradient data, it notifies the scheduling node. After the scheduling node confirms that each computing node has completed gradient calculation, it notifies each computing node to send the gradient data in the local memory to the CXL switch device.
[0152] S502: Each computing node sends the gradient data in the local memory to the CXL switch device and notifies the scheduling node.
[0153] Exemplarily, each computing node sends the gradient data in the local memory to the CXL switch device through the CXL controller.
[0154] S503: The scheduling node notifies the CXL switching device to calculate the fusion gradient.
[0155] After completing the transmission of gradient data, each computing node can notify the scheduling node. After the scheduling node confirms that each computing node has completed the transmission of gradient data, it notifies the CXL switching device to calculate the fusion gradient.
[0156] S504: The CXL switch calculates the fused gradient, writes the fused gradient into the shared memory, and notifies the scheduling node that the fused gradient is ready.
[0157] S505: The scheduling node notifies each computing node that the fused gradient is ready, and instructs each computing node to obtain the fused gradient and update the model parameters of the local target model.
[0158] For example, the scheduling node determines that the fused gradient is ready and can publish a gradient calculation completion event. Each computing node can pre-subscribe to the above event, that is, the scheduling node and the computing node pre-establish a publish-subscribe message mode.
[0159] The computing node determines that the fused gradient is ready, starts to obtain the fused gradient in the shared memory through the CXL switch, and updates the model parameters based on the fused gradient.
[0160] See also Figure 6 This diagram illustrates the communication rounds of the collective communication method provided in an embodiment of the present application. A CXL switch first uses the CXL protocol to retrieve data from each compute node's local memory and integrate the fused gradients. Each compute node then retrieves the fused gradients through the CXL switch. The entire gradient data transfer and calculation process requires only a single CXL communication round, reducing the number of communication rounds. Furthermore, communication based on the CXL protocol significantly reduces communication latency and improves the efficiency of distributed training.
[0161] In an embodiment of the present application, the model training system may include multiple CXL switching devices, and these CXL switching devices may form a topology structure with a hierarchy of N, where N is a positive integer greater than or equal to 2.
[0162] The downstream port of the CXL switch device at the first level is connected to the local memory of at least one computing node, and the upstream port is connected to the CXL switch device at the second level; the downstream port of the CXL switch device at the i-th level is connected to at least one CXL switch device at the i-1th level, and the upstream port is connected to the CXL switch device at the i+1th level; the downstream port of the CXL switch device at the Nth level is connected to at least one CXL switch device at the N-1th level, and the upstream port is connected to at least one computing node.
[0163] In such a topology, obtaining the gradient data in the local memory of each computing node based on CXL may specifically include: reading the gradient data in the local memory of each connected computing node through the first-level CXL switch device.
[0164] Calculate the fused gradient based on the gradient data of each computing node, including:
[0165] The CXL switching device at the first layer calculates the first-layer fusion gradient based on the read gradient data; the CXL switching device at the i-th layer reads the i-1th layer fusion gradient calculated by the CXL switching device at the i-1th layer, and calculates the i-th layer fusion gradient; the CXL switching device at the N-th layer reads the N-1th layer fusion gradient calculated by the CXL switching device at the N-1th layer, and calculates the fusion gradient.
[0166] Specifically, there can be multiple CXL switching devices at the first level, each of which is connected to the local memory of at least one computing node. In the process of calculating the fusion gradient,
[0167] The first-level CXL switch device obtains gradient data from the local memory of at least one computing node connected to it through the CXL protocol, and performs fusion calculation based on the obtained gradient data to obtain the first-level fused gradient.
[0168] Because there can be multiple CXL switches at the first layer, there can also be multiple calculated converged gradients for the first layer. The CXL switches at the second layer read the converged gradients calculated by each CXL switch at the first layer using the CXL protocol and calculate the converged gradient for the second layer. This continues in this manner until the CXL switches at the Nth layer read the converged gradients for the N-1th layer calculated by the CXL switches at the N-1th layer and calculate the converged gradient for the second layer.
[0169] See also Figure 7 , Figure 7 A schematic diagram of a collective communication method provided in an embodiment of the present application is provided. Figure 7In the illustrated embodiment, two layers of CXL switches are included. The first layer of CXL switches includes CXL switch 1 and CXL switch 2, and the second layer of CXL switches includes CXL switch 3. CXL switch 1 connects to the local memories of compute nodes 1, 2, and 3, namely, local memory 1, local memory 2, and local memory 3. CXL switch 2 connects to the local memories of compute nodes 4, 5, and 6, namely, local memory 4, local memory 5, and local memory 6.
[0170] CXL switch 1 calculates the first-layer fused gradient based on the read gradient data, which can be expressed as a1+a2+a3, b1+b2+b3, and c1+c2+c3. CXL switch 2 calculates another first-layer fused gradient based on the read gradient data, which can be expressed as a4+a5+a6, b4+b5+b6, and c4+c5+c5. CXL switch 3 reads the fused gradients calculated by each first-layer CXL switch and calculates the second-layer fused gradient, which can be expressed as a1+a2+a3+a4+a5+a6, b1+b2+b3+b4+b5+b6, and c1+c2+c3+c4+c5+c6.
[0171] As can be seen, multiple CXL switches can be used to form a multi-level topology. Through data exchange between CXL switches at different levels, the fused gradient is ultimately calculated. This is suitable for situations with a large number of computing nodes, minimizing the number of gradient data communication rounds and improving the efficiency of distributed training models.
[0172] See also Figure 8 An embodiment of the present application further provides a collective communication method executed by a scheduling node, the method being applied to a model training system, wherein multiple computing nodes of the model training system are used to perform distributed training on a target model, the model training system further comprising at least one CXL switch device, the CXL switch device being communicatively connected to the local memory of each of the multiple computing nodes via the CXL protocol; each computing node being communicatively connected to the shared memory via the CXL switch device, the method comprising the following steps:
[0173] S801: Send a first instruction to each computing node. The first instruction is used to instruct the computing node to send gradient data in the local memory to the CXL switching device. The gradient data is obtained by the computing node training the locally deployed target model.
[0174] S802: After each computing node has sent gradient data, a second instruction is sent to the CXL switch device. The second instruction is used to instruct the CXL switch device to calculate a fused gradient based on the gradient data of each computing node and store the fused gradient in the shared memory.
[0175] S803: When the CXL switching device has completed the calculation of the fused gradient, a third instruction is sent to each computing node. The third instruction is used to instruct each computing node to read the fused gradient from the shared memory and use the fused gradient to update the parameters of the locally deployed target model.
[0176] As can be seen, there is no need to repeatedly transmit gradient data, reducing the number of collective communication rounds required during model training. Furthermore, communication based on the CXL protocol eliminates the need to transmit data over the network, significantly reducing communication latency and improving the efficiency of distributed training.
[0177] In addition, since gradient fusion is performed directly on the CXL switching device, there is no need to send the aggregated gradient data of each computing node to a single service node. Compared with the method that requires gradient fusion at the service node, the communication overhead is further reduced.
[0178] See also Figure 9 An embodiment of the present application further provides a collective communication method performed by a computing node. The method is applied to a model training system. Multiple computing nodes of the model training system are used to perform distributed training on a target model. The model training system also includes at least one CXL switch device. The CXL switch device is connected to the local memory of each computing node in the multiple computing nodes via the CXL protocol. Each computing node is connected to the shared memory via the CXL switch device. The method includes the following steps:
[0179] S901: Train the locally deployed target model to obtain gradient data.
[0180] S902: Storing the gradient data in its own local memory.
[0181] S903: Access the fused gradients in the shared memory through the CXL switch device, and update the parameters of the locally deployed target model based on the fused gradients. The fused gradients are calculated based on the gradient data read by the CXL switch device from the local memory of each computing node and stored in the shared memory.
[0182] Each compute node simply trains the locally deployed target model, stores the trained gradient data in its local memory, and then waits for the CXL switch to calculate the fused gradient. After confirming that the fused gradient has been stored in shared memory, the CXL switch accesses the fused gradient in shared memory to obtain the fused gradient.
[0183] As can be seen, the steps required of compute nodes in the process of obtaining fused gradients are relatively simple, reducing the requirements for compute nodes. The entire gradient fusion process is performed directly on the CXL switch, eliminating the need to send the aggregated gradient data from each compute node to a single compute node or service node. This further reduces communication overhead compared to methods that require gradient fusion to be performed on compute nodes or service nodes.
[0184] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the method. In order to realize the above functions, the training device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0185] In the embodiment of the present application, the functional modules of the collective communication device can be divided according to the above method. For example, the collective communication device can include functional modules corresponding to the functional divisions, or two or more functions can be integrated into one processing module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. In actual implementation, there may be other division methods.
[0186] For example, Figure 10A possible schematic diagram of the collective communication device involved in the above-mentioned embodiment is shown, which is applied to a CXL switching device in a model training system. Multiple computing nodes of the model training system are used to perform distributed training on a target model. The model training system also includes at least one computing fast link CXL switching device. The CXL switching device is connected to the local memory of each computing node in the multiple computing nodes through the CXL protocol; each computing node is connected to the shared memory through the CXL switching device; the device includes: a functional unit for executing any one of the methods provided in the first aspect, and the actions performed by each functional unit are implemented by hardware or by hardware executing corresponding software implementations. For example, the collective communication device 1000 may include: an acquisition unit 1001, which is used to obtain gradient data in the local memory of each computing node based on the CXL protocol; the gradient data is obtained by the computing node training the locally deployed target model; a computing unit 1002, which is used to calculate a fused gradient based on the gradient data of each computing node and store the fused gradient in the shared memory; the fused gradient is used by each computing node to update the parameters of the locally deployed target model.
[0187] Optionally, the local memory of each computing node is a CXL memory managed by a CXL controller, and the CXL switch device is communicatively connected to the CXL controller. The acquiring unit 1001 is specifically configured to receive the gradient data sent by the CXL controller of each computing node in response to the first instruction.
[0188] Optionally, the computing unit 1002 is specifically configured to: calculate the fused gradient based on the gradient data of each computing node in response to a second instruction; the second instruction is sent by the sending entity after the CXL controller of each computing node has sent the gradient data.
[0189] Optionally, the apparatus further includes: a sending unit, configured to send a ready message to a sending subject of the second instruction after storing the fused gradient in the shared memory, wherein the ready message is used to indicate that the CXL switch device has completed calculation of the fused gradient.
[0190] Optionally, the model training system includes a plurality of CXL switching devices, which form a topology with N levels; N is a positive integer greater than or equal to 2; wherein a downstream port of a CXL switching device at the first level is connected to a local memory of at least one computing node, and an upstream port is connected to a CXL switching device at the second level; a downstream port of a CXL switching device at the i-th level is connected to at least one CXL switching device at the i-1th level, and an upstream port is connected to a CXL switching device at the i+1th level; a downstream port of a CXL switching device at the Nth level is connected to at least one CXL switching device at the N-1th level, and an upstream port is connected to at least one computing node; wherein i is a positive integer greater than 1 and less than N;
[0191] The acquisition unit 1001 is specifically configured to: read the gradient data in the local memory of each of the computing nodes connected to the acquisition unit 1001 through the first-level CXL switch device;
[0192] The calculation unit 1002 is specifically configured to calculate the first-layer fusion gradient based on the read gradient data through the first-layer CXL switching device;
[0193] The CXL switching device at the i-th layer reads the fusion gradient of the i-1th layer calculated by the CXL switching device at the i-1th layer, and calculates the fusion gradient of the i-th layer;
[0194] The CXL switching device at the Nth level reads the N-1th layer fusion gradient calculated by the CXL switching device at the N-1th level, and calculates the fusion gradient.
[0195] Optionally, the shared memory includes any one of the following:
[0196] The local memory of the CXL switch device, the extended memory of the CXL switch device, and the CXL memory device; the CLX memory device refers to a memory device managed by a CXL controller.
[0197] An embodiment of the present application also provides a computing device, comprising: a controller and a memory; the controller is coupled to the memory; the memory is used for computer program instructions; and the controller is used to call the computer program instructions in the memory to execute any one of the methods in the above embodiments.
[0198] An embodiment of the present application further provides a CXL switching device, comprising: a controller and a memory; the controller is coupled to the memory; the memory is used for computer program instructions; and the controller is used to call the computer program instructions in the memory to execute any one of the methods in the above embodiments.
[0199] An embodiment of the present application further provides a computer-readable storage medium storing computer execution instructions. When the computer execution instructions are executed on a computing device, the computing device executes any one of the methods in the above embodiments.
[0200] For explanations of the relevant contents and descriptions of the beneficial effects of any of the computer-readable storage media provided above, reference may be made to the corresponding embodiments described above, and no further details will be given here.
[0201] The embodiment of the present application also provides a computer program product comprising instructions, which, when run on a computing device, causes the computing device to perform any one of the methods in the above embodiments. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. It should be noted that the above-mentioned devices for storing computer instructions or computer programs provided in the embodiment of the present application, such as but not limited to the above-mentioned memory, computer-readable storage medium and communication chip, etc., are all non-transitory.
[0202] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using a software program, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.
[0203] Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. A computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. Available media can be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
[0204] Although the present application has been described with reference to specific features and embodiments thereof, it is apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are merely illustrative of the present application as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art may make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, the present application is intended to include such modifications and variations as fall within the scope of the claims of the present application and their equivalents.
Claims
1. A collective communication method, characterized in that: Applied to a model training system, wherein the plurality of computing nodes of the model training system are used to perform distributed training on a target model, the model training system further comprising at least one computing express link (CXL) switch device, the CXL switch device being communicatively connected to the local memory of each of the plurality of computing nodes via the CXL protocol; Each of the computing nodes is communicatively connected to the shared memory via the CXL switch device; The method is performed by the at least one CXL switching device, and includes: Acquiring gradient data in the local memory of each computing node based on the CXL protocol; the gradient data is obtained by the computing node training the locally deployed target model; A fused gradient is calculated based on the gradient data of each computing node, and the fused gradient is stored in the shared memory; the fused gradient is used by each computing node to update the parameters of the locally deployed target model.
2. The method according to claim 1, characterized in that The local memory of each computing node is a CXL memory managed by a CXL controller, and the CXL switch device is communicatively connected to the CXL controller; The acquiring of the gradient data in the local memory of each computing node based on the CXL protocol includes: receiving the gradient data sent by the CXL controller of each of the computing nodes in response to a first instruction; The first instruction is sent by the sending entity in response to each of the computing nodes being in the first state; the first state indicates that the computing node completes local calculation of gradient data and stores the calculated gradient data in the local memory.
3. The method according to claim 2, characterized in that The calculating of the fused gradient based on the gradient data of each computing node includes: In response to a second instruction, the fused gradient is calculated based on the gradient data of each computing node. The second instruction is sent by the sending entity in response to the CXL controllers of each computing node being in a second state. The second state indicates that each CXL controller has completed sending the gradient data.
4. The method according to claim 3, characterized in that After storing the fused gradient in the shared memory, the method further includes: A ready message is sent to the sending entity, so that the sending entity instructs each computing node to obtain the fused gradient in the shared memory through the CXL switch device; the ready message is used to indicate that the CXL switch device has completed the calculation of the fused gradient.
5. The method according to claim 1, wherein The model training system includes a plurality of CXL switching devices, which form a topology with N levels; N is a positive integer greater than or equal to 2; wherein a downstream port of a CXL switching device at the first level is connected to a local memory of at least one computing node, and an upstream port is connected to a CXL switching device at the second level; a downstream port of a CXL switching device at the i-th level is connected to at least one CXL switching device at the i-1th level, and an upstream port is connected to a CXL switching device at the i+1th level; a downstream port of a CXL switching device at the Nth level is connected to at least one CXL switching device at the N-1th level, and an upstream port is connected to at least one computing node; wherein i is a positive integer greater than 1 and less than N; The acquiring of the gradient data in the local memory of each computing node based on the CXL protocol includes: Reading gradient data in the local memory of each of the computing nodes connected to the first-level CXL switch device through the first-level CXL switch device; The calculating of the fused gradient based on the gradient data of each computing node includes: Calculating a first-layer fusion gradient based on the read gradient data by the first-layer CXL switching device; The CXL switching device at the i-th layer reads the fusion gradient of the i-1th layer calculated by the CXL switching device at the i-1th layer, and calculates the fusion gradient of the i-th layer; The CXL switching device at the Nth level reads the N-1th layer fusion gradient calculated by the CXL switching device at the N-1th level, and calculates the fusion gradient.
6. A collective communication method, characterized in that: Applied to a model training system, wherein the plurality of computing nodes of the model training system are used to perform distributed training on a target model, the model training system further comprising at least one computing express link (CXL) switch device, the CXL switch device being communicatively connected to the local memory of each of the plurality of computing nodes via the CXL protocol; Each of the computing nodes is communicatively connected to the shared memory via the CXL switch device; The method is executed by a scheduling node and includes: Sending a first instruction to each of the computing nodes, wherein the first instruction is used to instruct the computing node to send gradient data in a local memory to the CXL switch device, where the gradient data is obtained by the computing node training a locally deployed target model; When each computing node has sent the gradient data, sending a second instruction to the CXL switch device, wherein the second instruction is used to instruct the CXL switch device to calculate a fused gradient based on the gradient data of each computing node and store the fused gradient in the shared memory; When the CXL switch has completed the calculation of the fused gradient, a third instruction is sent to each computing node. The third instruction is used to instruct each computing node to read the fused gradient from the shared memory and update the parameters of the locally deployed target model using the fused gradient.
7. A collective communication method, characterized in that: Applied to a model training system, wherein the plurality of computing nodes of the model training system are used to perform distributed training on a target model, the model training system further comprising at least one computing express link (CXL) switch device, the CXL switch device being communicatively connected to the local memory of each of the plurality of computing nodes via the CXL protocol; Each of the computing nodes is communicatively connected to the shared memory via the CXL switch device; The method is performed by the computing node and includes: Train the locally deployed target model to obtain gradient data; Storing the gradient data in its own local memory; The fused gradient in the shared memory is accessed through the CXL switch device, and the parameters of the locally deployed target model are updated based on the fused gradient. The fused gradient is obtained by the CXL switch device reading the gradient data in the local memory of each computing node, calculating the gradient data based on the gradient data, and storing the calculated gradient in the shared memory.
8. A CXL switching device, characterized in that: include: controller and memory; The controller is coupled to the memory; The memory is used for computer program instructions; The controller is configured to call the computer program instructions in the memory to execute the method according to any one of claims 1 to 5.
9. A computing device, characterized in that include: controller and memory; The controller is coupled to the memory; The memory is used for computer program instructions; The controller is configured to call the computer program instructions in the memory to execute the method according to claim 6 or 7.
10. A model training system, characterized in that: The model training system includes: multiple computing nodes and at least one CXL switch device; the CXL switch device is connected to the local memory of each of the multiple computing nodes through the CXL protocol; each computing node is connected to the shared memory through the CXL switch device; Any one of the plurality of computing nodes is configured to train a locally deployed target model to obtain gradient data; and store the gradient data in its own local memory; The CXL switch device acquires the gradient data in the local memory of each computing node based on the CXL protocol; Calculating a fused gradient based on the gradient data of each computing node, and storing the fused gradient in the shared memory; Any one of the plurality of computing nodes is further configured to access the fused gradient in the shared memory through the CXL switch device, and update parameters of the locally deployed target model based on the fused gradient.