Artificial intelligence large model training method in heterogeneous multi-machine multi-card environment

By building a unified interface and heterogeneous communication protocol in a heterogeneous multi-machine and multi-card environment, the problems of hardware compatibility and low data transmission efficiency are solved, achieving efficient load balancing and dynamic resource management, and improving training efficiency and fault tolerance.

CN120909794APending Publication Date: 2025-11-07SICHUAN HUIXIN INTELLIGENT COMPUTING TECHNOLOGY CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202511084375.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-07

Smart Images

  • Figure CN120909794A_ABST
    Figure CN120909794A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence large model training method in a heterogeneous multi-machine and multi-card environment, and belongs to the technical field of artificial intelligence large model training. Load balancing of heterogeneous equipment is realized by constructing a uniform interface, and the communication efficiency is optimized by adopting hierarchical pipeline aggregation and dynamic quantization compression; the node dynamic adjustment is realized in combination with the elastic topological structure, the problems of poor equipment compatibility, high communication delay and rigid topological structure in the prior art are effectively solved, and the method has the remarkable advantages of improving the utilization rate of heterogeneous computing resources, reducing the cross-node communication overhead and enhancing the fault-tolerant capability of the system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence large model training, and in particular to an artificial intelligence large model training method in a heterogeneous multi-machine multi-card environment. BACKGROUND

[0002] Multi-machine multi-card distributed training is the mainstream solution for large-scale deep learning model training at present. Its core value lies in that multiple servers equipped with multiple GPUs work cooperatively to distribute massive training data and complex computing tasks to different computing nodes for parallel execution. This architecture can theoretically significantly improve model training efficiency and break through the single-machine memory capacity limit.

[0003] However, the prior art has exposed many systematic defects in practical application: First, in terms of hardware adaptation, the compatibility of traditional solutions for heterogeneous computing devices (such as mixed deployment of NVIDIA GPUs and Ascend AI processors) is poor, and there is a lack of unified resource scheduling interface, resulting in low coordination efficiency between different architecture accelerators.

[0004] In addition, the existing distributed training framework has obvious shortcomings in long sequence parallel processing capability. When facing super large-scale language models, its static topology structure is difficult to adapt to the needs of dynamically increasing and decreasing nodes, and the fault recovery mechanism often needs to restart the whole cluster, causing serious waste of computing resources.

[0005] In addition, the configuration complexity of open source frameworks is high, and ordinary developers need to manually adjust a large number of underlying parameters to achieve basic functions, and the utilization rate of heterogeneous hardware resources is generally low, and the cross-node communication delay becomes the main bottleneck restricting the improvement of training efficiency.

[0006] The superimposed effect of these problems makes it difficult for the current distributed training system to fully exert the theoretical performance advantages of the multi-machine multi-card cluster. SUMMARY

[0007] To solve the problems raised in the background art, the present application provides an artificial intelligence large model training method in a heterogeneous multi-machine multi-card environment to solve the problems of poor compatibility of traditional solutions for heterogeneous computing devices, low cross-node data transmission efficiency, and obvious shortcomings in long sequence parallel processing capability.

[0008] To achieve the above purpose, the present application provides the following technical solutions: An artificial intelligence large model training method in a heterogeneous multi-machine multi-card environment, comprising the following steps: S1: using a large language model as a processing engine, processing original data through a data processing script to output training data in a standardized format; S2: constructing a unified interface, distributing the training data output by S1 to each node of the heterogeneous multi-machine multi-card cluster based on the unified interface, achieving load balancing and stateful computing; S3: constructing a communication architecture including zero-copy memory management, hierarchical pipeline aggregation, and dynamic perception scheduling based on a heterogeneous communication protocol, and configuring a double buffering mechanism, the communication architecture being used to support the transmission of training data, model parameters, and gradients between cluster nodes; S4: performing multi-machine multi-card distributed training on the training data distributed by S2 based on a multi-dimensional parallel architecture, including forward propagation, backward propagation, and gradient aggregation; Specifically, the forward propagation is processed by each node in parallel, the generated gradients in the backward propagation are transmitted through the communication architecture of S3, and the gradient aggregation is achieved by real-time merging of local gradients of adjacent nodes and global synchronization through a sub-graph level micro-pipeline. During the gradient transmission process, the model parameters and gradients are dynamically quantized to 8-4bit and compressed through cross-node cooperation, the compressed data is transmitted through the communication architecture of S3, the receiving end decompresses and dequantizes the data for gradient aggregation. In addition, based on the preset elastic topology structure, the state of the cluster nodes is monitored, and when the nodes are dynamically added or reduced, the communication link is reconstructed through a near-end recovery strategy to restore the distributed training process.

[0009] Preferably, in S1, the large language model includes DeepSeek, and the standardized format is json or jsonl format.

[0010] Preferably, in S2, the unified interface can express both task-based parallel computing and actor-based parallel computing, and load balancing is achieved by dynamically allocating the training data processing amount of each node.

[0011] Preferably, in S3, the heterogeneous communication protocol is CrossBridge protocol, and the communication architecture includes: a zero-copy memory management layer; the zero-copy memory management layer is used to establish a shared virtual memory pool between a host and an accelerator, store the training data, model parameters, and gradients of S4, and eliminate the serialization / deserialization and memory copy overhead of data transmission; a hierarchical pipeline aggregation layer; the hierarchical pipeline aggregation layer is used to decompose the gradient aggregation task of S4 into a sub-graph level micro-pipeline, merge the local gradients of adjacent nodes in the backward propagation through a dedicated aggregation unit, and multicast synchronization; a dynamic perception scheduler; the dynamic perception scheduler is used to split the super-tensor of S4 into multiple sub-packets, and transmit them in parallel through PCIe / CXL and RDMA double channels.

[0012] Preferably, in S3, the double buffering mechanism is implemented by reconstructing the DMA controller, which adopts a three-level architecture: the hardware execution layer contains a double buffering storage area and a cyclic transmission engine, the driver abstraction layer contains a dynamic address mapping engine and a hardware trigger unit, and the protocol control layer runs an adaptive buffer algorithm to dynamically adjust the buffer depth.

[0013] Preferably, in S4, the multi-dimensional parallel architecture is a 5D parallel architecture, including data parallelism, model parallelism, pipeline parallelism, sequence blocking, and gradient remapping, and each GPU node in the forward propagation temporarily stores the calculation result in a shared virtual memory pool, and the global synchronization is only performed after every 3 micro-pipeline segments to perform a lightweight global consistency check.

[0014] Preferably, in S4, the compression adopts an LZ4-SLIM hybrid algorithm, the compressed data packet is transmitted in parallel through the RDMA / TCP dual-channel, and the receiving end decompresses and dequantizes the FP32 gradient for gradient aggregation in S4.

[0015] Preferably, in S4, the elastic topology is a (5x7)+3-15 logical topology, which is specifically expressed as: 5 groups of 7-node computing units + 3 redundant management nodes - 15 cross-group long connections. The near-end recovery strategy includes reconstructing the short connection of the group where the failed node is located, synchronizing the model shards of S4 by allocating a new instance from the redundant management node, and reconstructing the missing gradient by using the gradient history cached by the adjacent nodes to recover the training.

[0016] Preferably, the artificial intelligence large model training method further includes a distributed caching step: A training-aware heterogeneous storage engine is deployed to cache the training data of S1 and the model parameters of S4, and based on the model operator, the topology relationship between the video memory and the secondary storage is dynamically mapped to realize zero-wait switching between the video memory and the secondary storage.

[0017] Preferably, in S4, the distributed parallel training further includes sequence acceleration: The multi-head grouping fusion technology based on hardware topology perception reconstructs the attention calculation path to improve the hardware utilization rate and throughput of the forward propagation and the backward propagation in S4.

[0018] Compared with the prior art, the present application has the following advantages: The application provides an artificial intelligence large model training method in a heterogeneous multi-machine multi-card environment, which realizes load balancing of heterogeneous devices through construction of a unified interface, adopts hierarchical pipeline aggregation and dynamic quantization compression to optimize communication efficiency, and realizes dynamic adjustment of nodes in combination with an elastic topology structure, effectively solving the problems of poor device compatibility, high communication delay and rigid topology structure in the prior art, and having the significant advantages of improving utilization of heterogeneous computing resources, reducing cross-node communication overhead and enhancing system fault tolerance. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 The flowchart of the application is shown. DETAILED DESCRIPTION

[0020] In order to enable those skilled in the art to understand the technical content of the application, the application will be further described in detail below in combination with the drawings and specific examples. It should be understood that the specific examples described herein are only used to explain the application and do not limit the application.

[0021] Embodiment 1 An artificial intelligence large model training method in a heterogeneous multi-machine multi-card environment, as shown in Figure 1 The method comprises the following steps: S1: using a large language model as a processing engine, processing original data through a data processing script, and outputting training data in a standardized format; S2: constructing a unified interface, distributing the training data output by S1 to each node of a heterogeneous multi-machine multi-card cluster based on the unified interface, realizing load balancing and stateful computing; S3: constructing a communication architecture including zero-copy memory management, hierarchical pipeline aggregation and dynamic sensing scheduling based on a heterogeneous communication protocol, and configuring a double buffering mechanism, the communication architecture being used to support transmission of training data, model parameters and gradients between cluster nodes; S4: performing multi-machine multi-card distributed training on the training data distributed by S2 based on a multi-dimensional parallel architecture, including forward propagation, backward propagation and gradient aggregation; Specifically, the forward propagation is processed by each node in parallel, the gradients generated by the backward propagation are transmitted through the communication architecture of S3, the gradient aggregation realizes real-time merging of local gradients of adjacent nodes and global synchronization through a subgraph-level micro-pipeline, and in the gradient transmission process, the model parameters and the gradients are dynamically quantized to 8-4bit and compressed through cross-node cooperation, the compressed data is transmitted through the communication architecture of S3, the receiving end decompresses and dequantizes the data for gradient aggregation, and in addition, the state of the cluster nodes is monitored based on a preset elastic topology structure, and when the nodes are dynamically added or reduced, the communication link is reconstructed through a near-end recovery strategy to recover the distributed training process.

[0022] In this embodiment, in the data processing stage, the large language model cleans and formats the original data, and outputs standardized data to adapt to different hardware nodes. The unified interface dynamically allocates data volume according to the computing capacity of each node to achieve load balancing. The communication architecture eliminates serialization overhead through a shared memory pool, and the hierarchical pipeline aggregation decomposes gradient synchronization into multiple subgraph-level operations, and only performs lightweight global verification at key nodes. During distributed training, model parameters and gradients are dynamically quantized to low-bit-width data, compressed through double-channel transmission, and restored to the receiving end for gradient aggregation after restoring the data precision. When the number of nodes is monitored to increase or decrease, the elastic topology structure triggers the near-end recovery strategy, and uses redundant nodes and cached gradients to reconstruct the training state.

[0023] Compared with the prior art, the traditional method relies on fixed topology structure and single communication protocol, which cannot adapt to the mixed deployment scene of heterogeneous hardware, and gradient synchronization needs frequent global communication. This scheme converts global communication into local micro-pipeline operation through standardized interface and hierarchical aggregation mechanism, and reduces the amount of transmission data by combining dynamic quantization compression. In addition, the elastic topology structure breaks through the limitation of traditional static node relationship, supports dynamic expansion of hardware resources while ensuring the continuity of training, and effectively solves the problems of low data processing efficiency, insufficient resource utilization, high communication delay and weak fault tolerance in heterogeneous environment. Standardized data format and unified interface simplify the collaborative training process of heterogeneous hardware, zero-copy memory management and hierarchical pipeline aggregation reduce serialization and communication overhead, dynamic quantization compression technology balances transmission efficiency and calculation accuracy, and elastic topology structure guarantees the high availability of distributed training.

[0024] Embodiment 2 The difference between this embodiment and embodiment 1 is that in S1, the large language model includes DeepSeek, and the standardized format is json or jsonl format, to solve the problem of low data processing efficiency in heterogeneous hardware environment and the compatibility problem of distributed training caused by insufficient standardization, to realize zero-conversion analysis of cross-platform data processing, and to reduce the resource consumption of data preprocessing in heterogeneous clusters.

[0025] Embodiment 3 The difference between this embodiment and embodiment 1 is that in S2, the unified interface can express task-based parallel computing and actor-based parallel computing at the same time, and realizes load balancing by dynamically allocating the training data processing amount of each node.

[0026] Embodiment 4 The difference between this embodiment and embodiment 1 is that in S3, the heterogeneous communication protocol is CrossBridge protocol, and the communication architecture includes: Zero-Copy Memory Management Layer (Zero-Copy Memory Manager); Establish a shared virtual memory pool with unified physical addresses between the host and each accelerator (GPU / FPGA) to eliminate the serialization / deserialization and memory copying overhead of traditional frameworks (such as Horovod Elastic) when transmitting gradients and model parameters between devices; Hierarchical Pipeline Aggregation; The gradient aggregation task, such as ResNet-50, is decomposed into a subgraph-level micropipeline. During the backpropagation process, a dedicated aggregation unit merges the local gradients of adjacent GPUs and synchronizes them to the entire cluster via multicast, thus avoiding the centralized bottleneck of global reduction. Dynamic Awareness Scheduler; It monitors device load and link status in real time, automatically segments ultra-large tensors (such as deep convolution weights) into multiple sub-packets, and transmits them in parallel through PCIe / CXL and RDMA dual channels to maximize the use of heterogeneous bandwidth.

[0027] In this embodiment, the data processing flow is as follows: Forward propagation; each GPU processes batches of input data in parallel, and the calculation results are temporarily stored in a local shared memory pool; Backpropagation; 1. When the backpropagation of a single layer is completed, the local gradient is written directly to the shared memory pool; 2. The aggregation unit sniffs the memory pool of adjacent nodes in real time and triggers the subgraph-level gradient merging (such as the gradients of Layer 2~4 being aggregated between GPU 0-1 first); 3. After merging, the gradient is broadcast to other devices through the multicast engine, and at the same time, the dynamic scheduler allocates the sub-packet transmission path according to the link status (such as cutting the FP32 gradient into blocks, 50% going through low-latency PCIe, and 50% going through high-throughput RDMA).

[0028] After each compute node completes backpropagation, the quantization engine analyzes the gradient distribution in real time and dynamically selects 4-8 bit precision to perform non-uniform quantization on the FP32 gradients to generate low-precision tensors. The compression engine immediately applies the LZ4-SLIM hybrid algorithm (sparse mapping + pattern coding) for compression, generating data packets with a 90% reduction in payload. Based on real-time MTU detection results, the network adapter intelligently segments large packets / aggregates small packets to form optimized transmission frames and injects metadata through the frame header. During transmission, the frames are dynamically allocated to RDMA / TCP dual-channel parallel transmission according to the frame type. The receiving end reassembles the frames in order, decompresses them, and dequantizes them into FP32 gradients for aggregation. This process reduces the single-round gradient synchronization time from 1.2ms to 0.7ms, achieving a 100-epoch training time of 13.3 minutes in an Ascend-NVIDIA 8-card hybrid cluster, and improving bandwidth utilization from 58% to 85%.

[0029] By dynamically framing 4-8 bit low-precision gradients according to the TCP / IP payload and combining this with 90% traffic compression technology, the gradient synchronization time per round was reduced from 1.2ms to 0.7ms. Real-world testing shows that in an 8-GPU Ascend-NVIDIA hybrid cluster, the training time for 100 epochs was reduced to 13.3 minutes, and bandwidth utilization increased from 58% with Horovod Elastic to 85%.

[0030] Example 5 The difference between this embodiment and embodiment 4 is that in S3, the double buffering mechanism is implemented by refactoring the Ascend DMA controller to process the gradient calculation and transmission in S4 in parallel. The controller adopts a three-level architecture: Hardware execution layer: includes a double buffered storage area (BufferZone0 / 1) and a cyclic transfer engine (CyclicTransferEngine), which is directly mounted to the Ascend memory bus through the physical address mapping register group (PA_MAPRegisters); Driver Abstraction Layer: Embedded Dynamic Address Mapping Engine and Hardware Trigger Unit (HWTriggerUnit). The former maintains the bidirectional binding between the logical address of the buffer and the PA_MAP register in real time, while the latter generates a hardware interrupt signal INT_0→1 when the Buffer0 transfer is completed. Protocol control layer: Runs the Adaptive Buffer Scheduler algorithm and sends depth adjustment instructions D(t)=8⋅MTU / (B(t)⋅T_compute) to the driver layer through the Penetrative Control Bus, where T_compute is the Ascend calculation cycle and MTU is the maximum transmission unit.

[0031] In this embodiment, by nesting asynchronous gradient aggregation and double buffering mechanism, when the Ascending device performs the forward calculation and back propagation of the Nth round of model, the Asynchronous Aggregation Engine writes the gradient generated in the Nth round to the background buffer (such as BufferB) of the double-buffered gradient pool (Buffer Pool A / B) in real time; at the same time, the network transmission thread directly extracts the pre-compressed gradient of the N-1th round from the current ready front buffer pool (BufferA) to initiate transmission, realizing the transmission of the N-1th round of compressed gradient while calculating the Nth round. When the Nth round of reverse calculation is completed, the buffer pool master-slave state is instantaneously switched (BufferB→front), and the network thread immediately grabs the fresh gradient for compression and transmission, while the Ascending device starts the N+1th round of calculation without blocking. This nested scheduling completely hides 92% of the communication latency in the calculation period, compared with the fixed pipeline waiting mechanism of Horovod Elastic based on RingAllReduce, the model parameter update frequency is increased from 0.83 times / ms to 1.75 times / ms, with a performance gain of 2.1 times, realizing "transmission of the Nth round of compressed gradient while calculating the N+1th round of gradient" of the Ascending, hiding 92% of the communication latency. Compared with the fixed pipeline of Horovod Elastic based on Ring AllReduce, the model parameter update frequency is increased by 2.1 times. The key connection relationship is: Storage path: gradient data is written from the Ascending computing core to the double-buffered storage area through the memory bus in parallel (parallel writing); Hardware linkage path: when the loop transmission engine completes Buffer0 transmission, it immediately triggers an INT_0→1 signal to drive the layer to instantaneously switch the PA_MAP register to point to Buffer1 (hardware-level switching), and at the same time, trigger the CPU to process Buffer0 data (asynchronous release); Control feedback path: the protocol layer dynamically adjusts the depth allocation of the double-buffered area through the control bus by monitoring the network bandwidth B(t) and the calculation period T_compute, and feeds back the state to the protocol layer through the interrupt reporting channel for closed-loop optimization. This structure makes the communication-computation overlap rate jump to 89%, and the bandwidth stability reaches 85%±3%.

[0032] Through "hardware abstraction + software coordination" architecture reconstruction of the Ascending DMA transmission process, the data throughput is increased by 3.8 times; Dynamic mapping of the driving layer: learn from the ping-pong mechanism of STM32 DMA double buffering, realize the hardware-level switching of the buffer address in the Ascend driver: when the DMA controller completes the gradient transmission of buffer 0, automatically map the target address to buffer 1, and trigger the CPU to process the data of buffer 0. This mechanism changes the traditional single-buffer "transmission-processing" serial process into a parallel pipeline, and the transmission delay is reduced from 5.6 μs to 3.5 μs.

[0033] Embodiment 6 The difference between this embodiment and embodiment 5 is that in S4, the multi-dimensional parallel architecture is a 5D parallel architecture, including data parallelism, model parallelism, pipeline parallelism, sequence blocking, and gradient remapping, each GPU node temporarily stores the calculation result in the shared virtual memory pool in the forward propagation, and the global synchronization only performs a light global consistency check after every 3 micro-pipeline segments, avoiding the waiting delay of layer-by-layer global reduction.

[0034] Embodiment 7 The difference between this embodiment and embodiment 4 is that in S4, the compression adopts a LZ4-SLIM hybrid algorithm, the compressed data packet is transmitted in parallel through the RDMA / TCP double-channel, and the receiving end decompresses and dequantizes it into FP32 gradient for gradient aggregation of S4.

[0035] Embodiment 8 The difference between this embodiment and embodiment 1 is that in S4, the elastic topology structure is (5x7)+3-15 logical topology, which is specifically expressed as: 5 groups of 7-node computing units + 3 redundant management nodes - 15 cross-group long connections; The proximal recovery strategy includes locally reconstructing the short connection of the group where the failed node is located, allocating a new instance from the redundant management node to synchronize the model shard of S4, and using the gradient history cached by the adjacent node to reconstruct the missing gradient to recover the training.

[0036] In this embodiment, the adaptive process of the elastic topology is as follows: when a node dynamic increase / decrease event (such as device offline / expansion) is monitored, the Topology Manager immediately freezes the training state, starts (Proximal Recovery) based on the current (5x7)+3-15 logical topology structure: local reconstruction: only release the short connection in the computing unit group where the failed node is located (such as GPU0-1-2 group), and keep the other 34 healthy links active. Hot replacement activation: allocate a new instance from the 3 redundant management nodes, inject the fault position, and synchronize the latest model shard in the group; Gradient compensation: use the gradient history cached by the adjacent node in the previous 2 rounds to reconstruct the missing gradient of the failed node through a weighted interpolation algorithm; When nodes are dynamically added or removed, the (5×7)+3-15 elastic topology uses a near-end recovery strategy to keep the training interruption time within 15 seconds. In contrast, Horovod Elastic requires reinitializing the communication ring after node changes, resulting in a 12% loss of training time.

[0037] Example 9 The difference between this embodiment and Embodiment 1 is that the large-scale artificial intelligence model training method also includes a distributed caching step: Deploy a training-aware heterogeneous storage engine to cache training data from S1 and model parameters from S4. Based on the model operator access heatmap, dynamically map the topological relationship between GPU memory and secondary storage to achieve zero-wait switching between GPU memory and secondary storage.

[0038] In this embodiment, the training-aware heterogeneous storage engine refers to a storage management system capable of sensing data access patterns during training. Specifically, it can be implemented using a dynamic caching strategy based on hardware performance monitoring. By analyzing the bandwidth and latency characteristics of different storage media in real time, a cross-device cache hierarchy is constructed. The model operator access heatmap is a statistical graph recording the frequency of storage resource access by each computational unit of the model. Specifically, it can be generated using a sliding window-based statistical method to identify high-frequency data access areas and optimize storage locations. The topology relationship between GPU memory and secondary storage refers to the data distribution mapping rules between the graphics processor's memory and external storage devices. Specifically, it can be implemented using virtualization technology based on address space remapping. By dynamically adjusting the distribution ratio of data blocks across storage layers, cross-layer access overhead is reduced. The zero-wait switching mechanism eliminates the latency generated during storage layer switching. Specifically, it can be implemented using a hardware prefetching and cache coherency protocol collaborative control method, completing the readiness state preparation of the target storage medium before the data call request arrives.

[0039] Example 10 The difference between this embodiment and embodiment 1 is that, in S4, distributed parallel training also includes sequence acceleration: By reconstructing the attention computation path through hardware topology-aware multi-head grouping fusion technology, the hardware utilization and throughput of forward and backward propagation in S4 are improved.

[0040] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1.A method for training an artificial intelligence large model in a heterogeneous multi-machine multi-card environment, the method comprising: receiving a first model parameter from a first machine; receiving a second model parameter from a second machine; and updating the first model parameter based on the second model parameter. The method comprises the following steps: S1: using a large language model as a processing engine, processing raw data through a data processing script to output training data in a standardized format; S2: building a unified interface, distributing the training data output by S1 to each node of a heterogeneous multi-machine multi-card cluster based on the unified interface, achieving load balancing and stateful computing; S3: based on a heterogeneous communication protocol, a communication architecture is built, which includes zero-copy memory management, hierarchical pipeline aggregation, and dynamic perception scheduling, and a double buffering mechanism is configured, the communication architecture is used to support the transmission of training data, model parameters, and gradients between cluster nodes; S4: based on a multi-dimensional parallel architecture, the training data allocated by S2 is executed for multi-machine multi-card distributed training, including forward propagation, backward propagation, and gradient aggregation; Specifically, the forward propagation is processed by each node in parallel, the generated gradient in the backward propagation is transmitted through the communication architecture of S3, and the gradient aggregation is realized by the real-time merging and global synchronization of local gradients of adjacent nodes through the sub-graph level micro-pipeline. During the gradient transmission process, the model parameters and the gradient are dynamically quantized to 8-4bit and compressed through cross-node cooperation, the compressed data is transmitted through the communication architecture of S3, the receiving end decompresses and dequantizes for gradient aggregation. In addition, based on the preset elastic topology structure, the cluster node state is monitored, and when the nodes are dynamically added or reduced, the communication link is reconstructed through the near-end recovery strategy to restore the distributed training process. 2.The AI large model training method of claim 1, wherein, In S1, the large language model includes DeepSeek, and the standardized format is json or jsonl format. 3.The AI large model training method of claim 1, wherein, In S2, the unified interface can express task-based parallel computing and actor-based parallel computing at the same time, and load balancing is achieved by dynamically allocating the training data processing amount of each node. 4.The AI large model training method of claim 1, wherein, In S3, the heterogeneous communication protocol is CrossBridge protocol, and the communication architecture includes: Zero-copy memory management layer; the zero-copy memory management layer is used to establish a shared virtual memory pool between the host and the accelerator, store the training data, model parameters and gradients of S4, and eliminate the serialization / deserialization and memory copy overhead of data transmission; Hierarchical pipeline aggregation layer; the hierarchical pipeline aggregation layer is used to decompose the gradient aggregation task of S4 into a sub-graph level micro-pipeline, and through a dedicated aggregation unit, the local gradients of adjacent nodes are merged and multicast in the backward propagation; Dynamic perception scheduler; the dynamic perception scheduler is used to split the super tensor of S4 into multiple sub-packets, and transmit them in parallel through the PCIe / CXL and RDMA double channels. 5.The AI large model training method of claim 4, wherein, In S3, the double buffering mechanism is realized by reconstructing the DMA controller, which is used to parallel process the gradient calculation and transmission of S4. The controller adopts a three-level architecture: the hardware execution layer includes a double buffering storage area and a cyclic transmission engine, the driver abstraction layer includes a dynamic address mapping engine and a hardware trigger unit, and the protocol control layer runs an adaptive buffer algorithm to dynamically adjust the buffer depth. 6.The AI large model training method of claim 5, wherein, In S4, the multi-dimensional parallel architecture is a 5D parallel architecture, including data parallelism, model parallelism, pipeline parallelism, sequence blocking, and gradient remapping, each GPU node in the forward propagation temporarily stores the calculation result in a shared virtual memory pool, and a global synchronization is performed after every 3 micro-pipeline segments to perform a lightweight global consistency check. 7.The AI large model training method of claim 4, wherein, In S4, compression is performed by using an LZ4-SLIM hybrid algorithm, the compressed data packet is transmitted in parallel through a RDMA / TCP double-channel, and the receiving end decompresses and dequantizes the gradient into an FP32 gradient for gradient aggregation of S4. 8.The AI large model training method of claim 1, wherein, In S4, the elastic topology is a (5x7)+3-15 logical topology, specifically expressed as: 5 groups of 7-node computing units + 3 redundant management nodes - 15 cross-group long connections; the near-end recovery strategy includes reconstructing the short connection of the group in which the failed node is located, synchronizing the model shard of S4 by allocating a new instance from the redundant management node, and reconstructing the missing gradient by using the gradient history cached by the adjacent node to recover the training. 9.The AI large model training method of claim 1, wherein, The artificial intelligence large model training method further includes a distributed caching step: A training perception type heterogeneous storage engine is deployed to cache the training data of S1 and the model parameters of S4, and based on the model operator, the topology relationship between the video memory and the secondary storage is dynamically mapped to the heat map to realize zero-wait switching between the video memory and the secondary storage. 10.The AI large model training method of claim 1, wherein, In S4, the distributed parallel training further includes sequence acceleration: A multi-head grouping fusion technology based on hardware topology perception is used to reconstruct the attention calculation path, thereby improving the hardware utilization rate and throughput of the forward propagation and the backward propagation in S4.

Citation Information

Cited By

  • Communication resource allocation method, device, equipment, medium and product

    CN121301028A

  • Communication resource allocation method, apparatus, device, medium, and product

    CN121301028B

  • Multi-machine multi-card distributed training optimization system and method oriented to credential heterogeneous environment

    CN121603382A

  • A multi-machine multi-card distributed training optimization system and method for a signal creation heterogeneous environment

    CN121603382B

  • CFD simulation agent model construction method and system based on storage and calculation separation architecture

    CN121659857A