Large model cross-domain training scheduling method, system, device, equipment, medium and product
By obtaining the resource requirements and status of cross-domain training jobs for large models through the scheduling layer, determining the target scheduling scheme, and constructing data transmission and task communication paths, the problems of high training cost and high scheduling complexity in existing technologies are solved, and efficient cross-domain training execution is achieved.
Patent Information
- Application Number
- CN202511138098.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-12-02
AI Technical Summary
Existing cross-domain training scheduling methods for large models suffer from high training costs and scheduling complexity, especially when the computing centers belong to different entities, making it impossible to plan and construct dedicated fiber optic networks in a unified manner.
The scheduling layer obtains the computational, storage, and network resource requirements of large-scale cross-domain training jobs, determines the target scheduling scheme based on the resource status, and constructs data transmission paths and task communication paths through the scheduling layer and the control layer, so that cross-domain training can be carried out without the need to build a dedicated fiber optic network.
It reduces training costs and scheduling complexity, and enables efficient execution of cross-domain data transmission and task communication.
Smart Images

Figure CN121050880A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a scheduling method, system, device, equipment, medium, and product for cross-domain training of large models. Background Technology
[0002] Current industry practices for cross-domain, multi-computing-center training of large-scale models only consider computing power scheduling. The storage and network within each computing center, as well as the networks between computing centers, are planned and constructed in advance by the respective owners of the computing centers, requiring no scheduling and thus eliminating collaboration among computing power, storage, and networks. However, when the computing centers are owned by different entities, unified planning and construction become impossible. To enable cross-domain training across multiple computing centers, dedicated fiber optic networks must be pre-built, undoubtedly increasing training costs and scheduling complexity. Summary of the Invention
[0003] This invention provides a scheduling method, system, device, equipment, medium, and product for cross-domain training of large models, in order to solve the problems of high training cost and scheduling complexity in existing cross-domain training scheduling methods for large models.
[0004] According to one aspect of the present invention, a scheduling method for cross-domain training of large models is provided, applicable to a scheduling system including a scheduling layer and a control layer, the method comprising:
[0005] The scheduling layer obtains the job requirements and resource status of large-scale cross-domain training jobs; the job requirements include computational resource requirements, storage resource requirements, and network resource requirements.
[0006] The scheduling layer determines the corresponding target scheduling scheme based on job requirements and resource status.
[0007] Based on the target scheduling scheme, data transmission paths and task communication paths are constructed through the scheduling layer control and management layer, and cross-domain training of large models is started based on the data transmission paths and task communication paths.
[0008] According to another aspect of the present invention, a scheduling system for cross-domain training of large models is provided, the system comprising:
[0009] The scheduling layer includes task schedulers, data schedulers, and traffic schedulers deployed on global scheduling nodes;
[0010] The control layer includes computing power controllers and storage controllers deployed within each computing power center, as well as network controllers deployed on each wide-area deterministic network control plane;
[0011] Each computing power controller is connected to the task scheduler, each storage controller is connected to the data scheduler, and each network controller is connected to the traffic scheduler.
[0012] According to another aspect of the present invention, a scheduling device for cross-domain training of large models is provided, applied to a scheduling system including a scheduling layer and a control layer, the device comprising:
[0013] The acquisition module is used to obtain the job requirements and resource status of large-scale cross-domain training jobs through the scheduling layer; the job requirements include computing resource requirements, storage resource requirements, and network resource requirements.
[0014] The scheduling scheme generation module is used to determine the corresponding target scheduling scheme based on job requirements and resource status through the scheduling layer;
[0015] The scheduling scheme deployment module is used to construct data transmission paths and task communication paths through the scheduling layer control and management layer based on the target scheduling scheme, and to start cross-domain training of large models based on the data transmission paths and task communication paths.
[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to execute the scheduling method for cross-domain training of large models as described in any embodiment of the present invention.
[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the scheduling method for cross-domain training of large models as described in any embodiment of the present invention.
[0021] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the scheduling method for cross-domain training of large models as described in any embodiment of the present invention.
[0022] The scheduling method for large-scale cross-domain training provided in this invention obtains the job requirements and resource status of the large-scale cross-domain training job through a scheduling layer. The job requirements include computational resource requirements, storage resource requirements, and network resource requirements. The scheduling layer determines a corresponding target scheduling scheme based on the job requirements and resource status. Based on the target scheduling scheme, the scheduling layer controls the management layer to construct data transmission paths and task communication paths, and then initiates the large-scale cross-domain training based on these paths. This technical solution, through the collaboration of the scheduling layer and the management layer, satisfies the computational, storage, and network resource requirements of the large-scale cross-domain training job. Furthermore, by constructing data transmission paths and task communication paths, it enables cross-domain data transmission and task communication without the need for a dedicated fiber optic network, effectively reducing training costs and scheduling complexity.
[0023] The scheduling system, apparatus, electronic device, computer-readable storage medium, and computer program product for cross-domain training of large models provided in this embodiment of the invention also have the above-mentioned technical effects.
[0024] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a schematic diagram of the structure of a scheduling system for cross-domain training of a large model according to Embodiment 1 of the present invention;
[0027] Figure 2 This is a flowchart of a scheduling method for cross-domain training of a large model according to Embodiment 1 of the present invention;
[0028] Figure 3 This is a schematic diagram of the structure of another scheduling system for cross-domain training of large models according to Embodiment 1 of the present invention;
[0029] Figure 4 This is a flowchart of a scheduling method for cross-domain training of a large model according to Embodiment 2 of the present invention;
[0030] Figure 5 This is a schematic diagram of the structure of a scheduling device for cross-domain training of a large model according to Embodiment 3 of the present invention;
[0031] Figure 6 This is a schematic diagram of the structure of an electronic device that implements the scheduling method for cross-domain training of large models according to embodiments of the present invention. Detailed Implementation
[0032] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0034] Example 1
[0035] Figure 1 This is a schematic diagram of the scheduling system for cross-domain training of a large model provided in Embodiment 1 of the present invention. Figure 1 As shown, the scheduling system includes a scheduling layer 11 and a control layer 12, which are communicatively connected. The scheduling layer 11 serves as the global coordination and decision-making center for large-scale cross-domain training of models. Its main functions include: acquiring the computational, storage, and network resource requirements of large-scale cross-domain training jobs, and determining the optimal target scheduling scheme based on the resource status reported by the control layer 12; issuing tasks such as path construction, resource allocation, and task deployment to the control layer 12, and receiving execution status feedback to achieve global coordination of cross-domain training.
[0036] The control layer 12 serves as the carrier for local resource execution and control. Its main functions include: receiving tasks issued by the scheduling layer 11, allocating computing resources, configuring storage resources and synchronizing data, and constructing data transmission paths and task communication paths; and feeding back the local resource status to the scheduling layer 11 to support the scheduling layer's dynamic decision-making and adjustments, ensuring efficient execution of cross-domain training.
[0037] The scheduling system proposed in this invention achieves cross-domain resource collaborative decision-making through the scheduling layer and achieves localized efficient execution through the control layer. The collaboration between the two ensures the efficient execution of large-scale cross-domain training jobs while reducing the complexity of scheduling.
[0038] Based on the above scheduling system Figure 2 This is a flowchart of a scheduling method for cross-domain training of a large model according to Embodiment 1 of the present invention. This embodiment is applicable to situations where large-scale cross-domain training is collaboratively scheduled while simultaneously considering computational, storage, and network resource requirements. This method can be executed by a scheduling device for large-scale cross-domain training, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 2 As shown in the figure, the scheduling method for cross-domain training of large models provided in this embodiment includes the following steps:
[0039] S110. Obtain the job requirements and resource status of large model cross-domain training jobs through the scheduling layer; job requirements include computing resource requirements, storage resource requirements and network resource requirements.
[0040] In this context, "large-scale cross-domain training job" refers to a large-scale model training task that needs to be collaboratively executed across multiple geographically dispersed computing centers in a distributed training scenario spanning wide area networks. This task may include information such as training objectives, model information, and resource requirements. Job requirements refer to the specific resource requirements for completing the training in a large-scale cross-domain training job. This is a core input for the scheduling layer to formulate scheduling strategies and can include specific quantitative requirements for computing, storage, and network resources. Computing resource requirements describe the requirements of large-scale model training for accelerated computing power (GPU / NPU / TPU, etc.), including but not limited to the type, quantity, and computing power metrics of accelerator cards. Storage resource requirements describe the requirements of large-scale model training for data storage, including but not limited to storage capacity, read / write performance, and storage medium type. Network resource requirements describe the requirements of large-scale model training for data transmission and task communication, including but not limited to network latency, network bandwidth, and packet loss rate. Resource status refers to the current status of various resources reported from the management layer to the scheduling layer, specifically including computing resource status, storage resource status, and network resource status. This is also an important input for the scheduling layer to formulate scheduling strategies.
[0041] In this embodiment of the invention, the user can submit a corresponding large-scale cross-domain training job to the scheduling layer of the scheduling system based on information such as the model characteristics and resource requirements of the model to be trained. After receiving the large-scale cross-domain training job, the scheduling layer will extract the job requirements containing three types of core resource requirements (computation, storage, and network). In addition, the scheduling layer can also obtain various resource statuses (computation, storage, and network) fed back by the control layer, and will subsequently make scheduling decisions based on the job requirements and resource status.
[0042] S120. The scheduling layer determines the corresponding target scheduling scheme based on job requirements and resource status.
[0043] The target scheduling scheme can refer to the optimal scheduling scheme ultimately decided based on job requirements and resource status. It includes information such as the target computing center participating in training, storage node allocation method, data transmission and task communication path planning, and computing resource deployment scheme. It is the core basis for starting cross-domain training.
[0044] In this embodiment of the invention, the scheduling layer can, based on the acquired job requirements (computing resource requirements, storage resource requirements, and network resource requirements) and resource status (computing resource status, storage resource status, and network resource status), match the job requirements and resource status one by one to filter out a set of candidate schemes that simultaneously meet the computing resource requirements, storage resource requirements, and network resource requirements. Then, the scheduling layer or the user selects the final target scheduling scheme from the set of candidate schemes. In one embodiment, the computing resource requirements and computing resource status can be matched first to filter out candidate schemes that meet the computing resource requirements (i.e., schemes that include suitable combinations of computing centers and computing node allocation methods) to form a first scheduling scheme set; then, the storage resource requirements and storage resource status can be matched to remove schemes that do not meet the storage resource requirements from the first scheduling scheme set to form a second scheduling scheme set; then, the network resource requirements and network resource status can be matched to remove schemes that do not meet the network resource requirements from the second scheduling scheme set to form a third scheduling scheme set; finally, a preset scheduling scheme selection strategy (such as comprehensively considering training efficiency, resource utilization, cross-domain costs, etc.) can be invoked to quantitatively score each scheme in the third scheduling scheme set, and the scheme with the highest score can be selected as the target scheduling scheme.
[0045] S130. Based on the target scheduling scheme, the data transmission path and task communication path are constructed through the scheduling layer control and management layer, and the cross-domain training of the large model is started based on the data transmission path and task communication path.
[0046] The data transmission path can refer to a logical channel constructed by the network controller for transmitting data such as training datasets, checkpoint data, and model parameter files. It is characterized by large transmission volumes and high stability requirements. The task communication path can refer to a logical channel constructed by the network controller for cross-domain communication between tasks (such as gradient synchronization and parameter updates). It is characterized by high real-time performance (low latency and low jitter).
[0047] In this embodiment of the invention, the scheduling layer can determine the target computing power center and the wide area network involved in the cross-domain training based on the target scheduling scheme, and plan the bandwidth and transmission stability requirements that the data transmission path must meet, as well as the real-time indicators such as latency and jitter that the task communication path must meet. Then, the data transmission path and the task communication path are opened sequentially on the specified wide area network, and the training data is migrated and synchronized between multiple computing power centers based on the opened data transmission path. Finally, the scheduling layer can issue the deployment task of large-scale model cross-domain training to the management and control layer, so that the management and control layer can start and execute large-scale model cross-domain training based on the opened data transmission path and the task communication path.
[0048] The scheduling method for large-scale cross-domain training provided in this invention obtains the job requirements and resource status of the large-scale cross-domain training job through a scheduling layer. The job requirements include computational resource requirements, storage resource requirements, and network resource requirements. The scheduling layer determines a corresponding target scheduling scheme based on the job requirements and resource status. Based on the target scheduling scheme, the scheduling layer controls the management layer to construct data transmission paths and task communication paths, and then initiates the large-scale cross-domain training based on these paths. This technical solution, through the collaboration of the scheduling layer and the management layer, satisfies the computational, storage, and network resource requirements of the large-scale cross-domain training job. Furthermore, by constructing data transmission paths and task communication paths, it enables cross-domain data transmission and task communication without the need for a dedicated fiber optic network, effectively reducing training costs and scheduling complexity.
[0049] Figure 3 This is a schematic diagram of another scheduling system for cross-domain training of large models provided in Embodiment 1 of the present invention. Figure 3 As shown, this scheduling system is... Figure 1 The scheduling system is further refined in this way. The scheduling layer 11 includes a task scheduler 111, a data scheduler 112, and a traffic scheduler 113 deployed on the global scheduling node. The control layer 12 includes a computing power controller 121 and a storage controller 122 deployed within each computing power center, and a network controller 123 deployed on each wide-area deterministic network control plane. Each computing power controller 121 is communicatively connected to the task scheduler 111, each storage controller 122 is communicatively connected to the data scheduler 112, and each network controller 123 is communicatively connected to the traffic scheduler 113.
[0050] Among them, the global scheduling node can refer to the physical or virtual node that deploys the scheduling layer, has a global view of cross-domain resources (can obtain the resource status of each computing center and wide area network in real time), and coordinates resource decisions of multiple computing centers. For example, it can be the control node of the cloud platform or a dedicated scheduling server.
[0051] The task scheduler can be the core component of the scheduling layer. It is responsible for receiving and parsing large-scale cross-domain training jobs, coordinating the matching of computing, storage, and network resource needs, generating target scheduling schemes, and coordinating the execution of other schedulers and controllers. It is the initiator of the "decision-execution" link.
[0052] A data scheduler can refer to a component in the scheduling layer that focuses on storage resource management. It is responsible for allocating appropriate storage resources for the data involved in large-scale cross-domain training jobs and, by collaborating with the task scheduler, completing data synchronization operations between various computing centers to ensure the consistency of cross-domain training data.
[0053] A traffic scheduler can be a component in the scheduling layer that focuses on network resource management. It is responsible for allocating appropriate network resources for data transmission and task communication involved in cross-domain training of large models. By cooperating with the task scheduler and data scheduler, it opens deterministic network paths for data transmission and task communication of training jobs, ensuring the quality of service for cross-domain data transmission and task communication.
[0054] A computing center can refer to a physical or logical cluster with large-scale computing resources (such as a GPU data center or a supercomputing center). It is the actual execution site for large model training tasks and internally deploys computing power controllers and storage controllers.
[0055] A computing power controller can refer to a computing resource management component deployed inside a computing power center. It is responsible for monitoring and allocating local computing resources and can provide the task scheduler with the status of computing resources in its computing power center as a basis for task scheduling.
[0056] A storage controller can refer to a storage resource management component deployed inside a computing center. It is responsible for monitoring and allocating local storage resources and can provide the data scheduler with the storage resource status of its computing center as a basis for data scheduling.
[0057] Wide area deterministic networks (hereinafter referred to as WANs) can refer to networks that cover multiple geographical areas, have predictable latency and bandwidth (such as time-sensitive networks, flexible Ethernet, etc.), and support cross-domain data transmission and task communication.
[0058] Wide area deterministic network control plane can refer to the control node (such as software-defined network controller) responsible for managing the wide area network. It has functions such as path planning, resource allocation, and status monitoring. The network controller is deployed here to realize logical control of physical links.
[0059] A network controller can refer to a network resource management component deployed on the control plane of a wide area network (WAN). It is responsible for monitoring and allocating local network resources, planning and opening data transmission paths and task communication paths, and providing the traffic scheduler with the network resource status of its WAN as a basis for traffic scheduling.
[0060] Furthermore, based on Figure 3 The scheduling system shown, S110 specifically includes the following steps:
[0061] S1101. Extract large model cross-domain training jobs from the preset job queue through the task scheduler; large model cross-domain training jobs are generated from the model features and resource requirements description of the model to be trained;
[0062] S1102. The task scheduler parses the computational resource requirements, storage resource requirements, and network resource requirements in the cross-domain training job of the large model.
[0063] The preset job queue can refer to a queue pre-configured in the task scheduler for storing large model cross-domain training jobs to be executed. It can support scheduling by priority or time to ensure the orderliness of job scheduling.
[0064] The model to be trained can refer to a large artificial intelligence model whose parameters are optimized through distributed training across a wide area network. For example, the model to be trained can be a model such as GPT or Llama.
[0065] Model features can refer to the inherent attribute information of the model to be trained, which reflects the model's potential demand for resources. These features may include, but are not limited to: model type (such as GPT3 / 4, llama2 / 3 / 4, etc.), parameter size (such as 10B, 32B, 561B, etc.), number of model layers, parallel mode (such as data parallelism, tensor parallelism, pipelined parallelism, etc.), training method (such as pre-training, incremental training, partial fine-tuning, etc.), and training dataset (such as the location, size, and type of the dataset).
[0066] Resource requirement descriptions can refer to specific descriptions of the resources required for training, generated based on model features. They are resource requirements predefined by users for cross-domain training of large models and are a core component of the training job.
[0067] In this embodiment of the invention, the task scheduler can maintain a preset job queue for centrally storing large-scale cross-domain training jobs to be executed. This queue can be sorted according to job priority (e.g., urgency, training task importance) or submission time to ensure orderly scheduling. The task scheduler can select large-scale cross-domain training jobs to be executed from the preset job queue according to preset rules (e.g., prioritizing high-priority jobs, sequential extraction, etc.), and parse them to obtain their computational resource requirements, storage resource requirements, and network resource requirements, which serve as the core input for subsequently determining the target scheduling scheme. The large-scale cross-domain training job can be generated from the model features and resource requirement description of the model to be trained, and submitted to the task scheduler by the user.
[0068] Furthermore, based on Figure 3 The scheduling system shown, S120 specifically includes the following steps:
[0069] S1201. Obtain the corresponding computing resource status, storage resource status and network resource status through the task scheduler, data scheduler and traffic scheduler respectively.
[0070] S1202. The task scheduler determines the first scheduling scheme set according to the computing resource requirements and computing power resource status, and sends the first scheduling scheme set and storage resource requirements to the data scheduler.
[0071] S1203. The data scheduler selects the second scheduling scheme set from the first scheduling scheme set according to the storage resource requirements and storage resource status, and sends the second scheduling scheme set to the task scheduler.
[0072] S1204. Send the second scheduling scheme set and network resource requirements to the traffic scheduler through the task scheduler;
[0073] S1205. The traffic scheduler selects the third scheduling scheme set from the second scheduling scheme set according to the network resource demand and network resource status, and sends the third scheduling scheme set to the task scheduler.
[0074] S1206. The target scheduling scheme is determined in the third scheduling scheme set by calling the preset scheduling scheme selection strategy through the task scheduler.
[0075] Among them, computing resource status can refer to the real-time information of computing resources of each computing center, including but not limited to the computing power usage of accelerator cards, memory utilization, health status, etc.
[0076] Storage resource status can refer to real-time information on storage resources in each computing center, including but not limited to: available capacity, read / write bandwidth, checkpoint write rate, and storage medium type.
[0077] Network resource status can refer to real-time link quality information of various wide area networks, including but not limited to bandwidth, end-to-end latency, packet loss rate, path availability, and load status of cross-domain backbone networks.
[0078] The first scheduling scheme set can refer to the set of candidate scheduling schemes that, after being filtered by the task scheduler, only meet the computing resource requirements.
[0079] The second scheduling scheme set can refer to the set of candidate scheduling schemes that, after being filtered by the data scheduler, simultaneously meet the requirements for computing resources and storage resources.
[0080] The second scheduling scheme set can refer to the set of candidate scheduling schemes that, after being filtered by the traffic scheduler, simultaneously meet the requirements of computing, storage, and network resources.
[0081] The preset scheduling scheme selection strategy can refer to the strategy used to select the optimal target scheduling scheme from the third scheduling scheme set. For example, the target scheduling scheme can be determined by comprehensively considering indicators such as total training time, resource utilization, cross-domain communication cost, and fault tolerance, and by scoring or ranking.
[0082] In this embodiment of the invention, the process for generating the target scheduling scheme specifically includes:
[0083] ① Obtain the computing resource status and storage resource status of each computing center through the task scheduler and data scheduler respectively, and obtain the network resource status of each wide area network through the traffic scheduler.
[0084] ② The task scheduler uses computing resource requirements as a filtering condition to match the acquired computing resource status and selects a set of schemes that meet the computing resource requirements from all possible scheduling schemes, namely the first scheduling scheme set; then, the task scheduler synchronizes the first scheduling scheme set and storage resource requirements to the data scheduler.
[0085] ③ The data scheduler uses storage resource requirements as a filtering condition and combines the obtained storage resource status to further filter the first scheduling scheme set, eliminating schemes that do not meet storage requirements, and obtains the second scheduling scheme set; then, the data scheduler feeds back the second scheduling scheme set to the task scheduler.
[0086] ④ The task scheduler synchronizes the second scheduling scheme set and network resource requirements to the traffic scheduler.
[0087] ⑤ The traffic scheduler uses network resource requirements as a filtering condition and combines the obtained network resource status to perform a final filtering of the second scheduling scheme set, eliminating schemes that do not meet network requirements, and obtaining the third scheduling scheme set; then, the traffic scheduler feeds back the third scheduling scheme set to the task scheduler.
[0088] ⑥ The task scheduler calls the preset scheduling scheme selection strategy, such as comprehensively considering training efficiency, resource utilization, cost, etc., to quantitatively evaluate and prioritize all schemes in the third scheduling scheme set, and selects the optimal scheme as the target scheduling scheme.
[0089] It should be understood that the above embodiment uses a hierarchical filtering logic of "computation → storage → network". In practical applications, hierarchical filtering logics such as "computation → network → storage", "storage → computation → network", "storage → network → computation", "network → computation → storage", and "network → storage → computation" can also be used. That is, this embodiment does not impose specific restrictions on the order of filtering computation, storage, and network.
[0090] Furthermore, based on Figure 3 The scheduling system shown, S130 specifically includes the following steps:
[0091] S1301. The task scheduler determines the target computing power centers and target wide-area deterministic networks participating in cross-domain training according to the target scheduling scheme, and generates data transmission network scheme, data storage scheme, task communication network scheme and training task deployment scheme based on each target computing power center and target wide-area deterministic network.
[0092] S1302. The data transmission network scheme and the task communication network scheme are sent to the traffic scheduler through the task scheduler, and the data storage scheme is sent to the data scheduler.
[0093] S1303. The traffic scheduler generates data transmission path activation tasks corresponding to each target wide area deterministic network according to the data transmission network scheme, and sends each data transmission path activation task to the network controller of the corresponding target wide area deterministic network.
[0094] S1304. Each network controller opens the corresponding data transmission path according to its own data transmission path opening task.
[0095] S1305. The data scheduler generates data storage tasks corresponding to each target computing center according to the data storage scheme, and sends each data storage task to the storage controller of the corresponding target computing center.
[0096] S1306. Each storage controller performs the corresponding storage resource allocation operation according to its own data storage task, and completes the cross-computing center data synchronization operation based on the opened data transmission paths.
[0097] S1307. The traffic scheduler generates task communication path activation tasks corresponding to each target wide area deterministic network according to the task communication network scheme, and sends each task communication path activation task to the network controller of the corresponding target wide area deterministic network.
[0098] S1308. Each network controller opens the corresponding task communication path according to its own task communication path.
[0099] S1309. The task scheduler generates training task deployment tasks corresponding to each target computing center according to the training task deployment plan, and sends each training task deployment task to the computing power controller of the corresponding target computing center.
[0100] S13010: Each computing power controller deploys tasks according to their respective training tasks, binds the corresponding task communication paths, and simultaneously starts cross-domain training of large models.
[0101] Among them, the data transmission network scheme can refer to the network configuration scheme generated based on the target scheduling scheme for planning cross-domain training data transmission. The scheme defines the starting point and ending point of data transmission, the target wide area network through which the path passes, bandwidth and latency requirements, transmission protocols, and other information, which is the basis for the traffic scheduler to generate data transmission path opening tasks.
[0102] A data storage scheme can refer to a configuration scheme generated based on a target scheduling scheme for planning storage resources for training data (dataset / model parameters / checkpoints). The scheme defines the storage capacity to be allocated to each target computing center, data storage location, data read and write permissions (such as access priority of training nodes), data consistency maintenance rules (such as verification cycle, version management strategy), and other information, which is the basis for the data scheduler to generate data storage tasks.
[0103] A task communication network scheme can refer to a network configuration scheme generated based on a target scheduling scheme and used to plan real-time interaction between task nodes during the training process. The scheme defines the participating nodes in task communication (such as training nodes in each computing center), the target wide area network through which the path passes, real-time indicators (such as end-to-end latency and packet loss rate), communication protocols, path redundancy design (such as backup link planning), etc., and serves as the basis for the traffic scheduler to generate task communication path opening tasks.
[0104] A training task deployment scheme can refer to a configuration scheme generated based on a target scheduling scheme to plan the training division of labor among various computing centers. The scheme defines the training task types (such as sharding computation in model parallelism and sample allocation in data parallelism), the computing resources to be activated (such as the number of GPUs and memory configuration), training environment parameters (such as framework version and precision mode), local iteration counts and global synchronization nodes, and fault recovery strategies (such as task migration rules after a node goes offline), etc., which are the basis for the task scheduler to generate training task deployment tasks.
[0105] A data transmission path activation task can refer to a specific operation instruction generated by the traffic scheduler based on the data transmission network scheme and issued to the network controller. This instruction may include information such as the start and end nodes of the data transmission path to be activated, the reserved bandwidth resources, the path identifier, and the transmission priority configuration. It is used to guide the network controller to activate a dedicated channel in the target wide area network that meets the data transmission requirements.
[0106] Data storage tasks can refer to specific operation instructions generated by the data scheduler based on the data storage scheme and issued to the storage controller. These instructions may include information such as the amount of storage resources to be allocated, the data storage location, the triggering conditions for data synchronization (such as timed synchronization or incremental synchronization), the verification algorithm (such as SHA-256), and the mounting method of the storage node (such as object storage protocol). These instructions are used to guide the storage controller in completing the storage and consistency maintenance of cross-domain data.
[0107] The task communication path opening task can refer to the specific operation instructions generated by the traffic scheduler according to the task communication network scheme and issued to the network controller. It may include the node pairs of the task communication path to be opened (such as training node A and parameter server B), path monitoring indicators (such as latency jitter threshold), fault switching trigger mechanism, path bandwidth guarantee rules and other information, which are used to guide the network controller to build a dedicated communication link that meets the low latency interaction requirements.
[0108] Training task deployment tasks can refer to specific operation instructions generated by the task scheduler according to the training task deployment plan and issued to the computing power controller. These instructions may include details of the computing resources to be allocated (such as GPU model and quantity), training task startup parameters (such as batch size and learning rate), model shard loading paths, synchronization node addresses with other computing power centers, training log upload rules, and other information, which are used to guide the computing power controller to complete the initialization of the training environment and the start of cross-domain training.
[0109] In this embodiment of the invention, the deployment process of the target scheduling scheme specifically includes:
[0110] ① The task scheduler can identify all target computing centers (such as GPU cluster centers distributed in different regions) and target wide area networks (such as dedicated fiber optic networks and backbone networks connecting various computing centers) participating in this large-scale cross-domain model training, based on the determined target scheduling scheme. Based on these target participants, it further refines and generates four sub-schemes: data transmission network scheme, data storage scheme, task communication network scheme, and training task deployment scheme. The data transmission network scheme specifies the network path planning, bandwidth allocation, and transmission priority for training data transmission between target computing centers; the data storage scheme specifies the amount of storage resources to be allocated to each target computing center, data sharding rules, and storage node locations; the task communication network scheme plans the network path, latency control, and fault tolerance mechanisms for real-time interaction (such as gradient parameter synchronization) between nodes in each computing center during training; and the training task deployment scheme determines the division of training tasks among the target computing centers (such as model sharding allocation, number of computing nodes, and training progress connection rules).
[0111] ②The task scheduler sends the data transmission network scheme and the task communication network scheme to the flow scheduler, and sends the data storage scheme to the data scheduler.
[0112] ③ The traffic scheduler generates corresponding data transmission path activation tasks for each target WAN based on the received data transmission network scheme, and distributes these tasks to the network controller (such as the control node of a regional backbone network) corresponding to that network.
[0113] ④ After receiving the data transmission path opening task, the network controller of each target WAN performs path opening operations within the network it manages (such as configuring routing tables, reserving bandwidth resources, and enabling data verification mechanisms), and finally constructs a dedicated path that meets the data transmission requirements.
[0114] ⑤ Based on the received data storage scheme, the data scheduler generates corresponding data storage tasks for each target computing center (including the storage capacity to be allocated, data read and write permissions, shard storage locations, etc.) and distributes these tasks to the storage controller (such as the distributed storage management node within the computing center) corresponding to that computing center.
[0115] ⑥ After receiving the data storage task, the storage controller of each target computing center first allocates sufficient storage resources locally (such as mounting disks and dividing storage pools), and then transfers the training data (such as training datasets, model parameter files, etc.) from the data source node to the local storage node through the data transmission path opened in step S1304. It then completes the cross-computing center data synchronization operation through operations such as hash verification and version alignment.
[0116] ⑦ Based on the received task communication network scheme, the traffic scheduler generates corresponding task communication path activation tasks for each target WAN (including real-time requirements such as path delay limit, jitter tolerance, and packet loss rate threshold), and distributes these tasks to the network controller corresponding to the network.
[0117] ⑧ After receiving the task communication path opening task, the network controller of each target WAN constructs a low-latency, high-reliability dedicated communication path, namely the task communication path, within the network range it manages, to ensure the real-time interaction needs between task nodes during the training process.
[0118] ⑨ The task scheduler generates corresponding training task deployment tasks (including the number of GPUs to be activated, training framework version, number of local iterations, etc.) for each target computing center according to the training task deployment plan, and distributes these tasks to the computing power controller corresponding to the computing power center.
[0119] ⑩ After receiving the training task deployment task, the computing power controller of each target computing power center first binds the local training node to the task communication path opened in step S1308 (to ensure that the real-time interaction channel is available), and then completes the initial parameter alignment of each node (such as model shard parameter synchronization) through the path. Finally, under the coordination of the task scheduler, all computing power centers synchronously start large model cross-domain training.
[0120] Furthermore, based on the above embodiments, the scheduling method provided in this embodiment further includes:
[0121] Each computing power controller reports the computing power resource status of its respective computing power center to the task scheduler every first preset period or when a preset computing resource change event is detected.
[0122] Each storage controller reports the storage resource status of its respective computing center to the data scheduler every second preset period or when a preset storage resource change event is detected.
[0123] Each network controller reports the network resource status of its wide area deterministic network to the traffic scheduler every third preset period or when a preset network resource change event is detected.
[0124] The first preset period can refer to the time interval at which the computing power controller actively reports the computing resource status to the task scheduler. This can be pre-configured by the system based on the sensitivity of the large model training to changes in computing power, balancing real-time performance and resource consumption. The second preset period can refer to the time interval at which the storage controller actively reports the storage resource status to the data scheduler. This can be pre-configured by the system based on the update frequency of the stored data. The third preset period can refer to the time interval at which the network controller actively reports the network resource status to the traffic scheduler. This can be pre-configured by the system based on the stability of the network link.
[0125] Preset computing resource change events can refer to pre-configured events that trigger the immediate reporting of computing resource status, such as including but not limited to: GPU node failure, addition / decommissioning of computing nodes, computing load exceeding a preset threshold, etc.
[0126] Preset storage resource change events can refer to pre-configured events that trigger the immediate reporting of storage resource status, such as including but not limited to: available storage capacity falling below a preset threshold, storage volume expansion / shrinkage, disk failure, etc.
[0127] Preset network resource change events can refer to pre-configured events that trigger the immediate reporting of network resource status, such as including but not limited to: link interruption, link bandwidth / latency exceeding a preset threshold, etc.
[0128] In this embodiment of the invention, a dual mechanism of "periodic triggering + event triggering" can be used to enable the management layer to report resource status to the scheduling layer in real time and accurately. The specific process is as follows:
[0129] ①The computing power controller of each computing power center continuously monitors the local computing resources. When any of the following conditions are met, it immediately reports the latest computing power resource status to the task scheduler: (1) the first preset period is reached; (2) a preset computing resource change event is detected.
[0130] ② The storage controller of each computing center continuously monitors the local storage system. When any of the following conditions are met, it immediately reports the latest storage resource status to the data scheduler: (1) the second preset cycle is reached; (2) a preset storage resource change event is detected.
[0131] ③ The network controllers of each WAN continuously monitor the backbone network link quality (such as bandwidth, latency, packet loss rate, etc.) and when any of the following conditions are met: (1) the third preset period is reached; (2) a preset network resource change event is detected.
[0132] Example 2
[0133] Figure 4 This is a flowchart illustrating a scheduling method for cross-domain training of a large model, as provided in Embodiment 2 of the present invention. Figure 4 As shown, the essence of large-scale cross-domain training scheduling is to select and allocate appropriate computing, storage, and network resources for large-scale model jobs. Therefore, this embodiment describes the scheduling process of large-scale cross-domain training from the aspects of scheduling input / output and working principle. The scheduling input includes a description of the characteristics and resource requirements of the large model itself, as well as a description of cross-domain computing, network, and storage resources; the scheduling output is the correspondence and amount of computing, storage, and network resources allocated to the large-scale model job. The working principle of scheduling is the process of generating output based on input.
[0134] 1. Model characteristics and resource requirements for large model training tasks
[0135] Large-scale models possess extremely large parameter counts (tens of billions to trillions) and deep network structures (hundreds to thousands of layers). They employ distributed training strategies such as data / model / pipeline parallelism and rely on high-performance computing resources to process massive amounts of data, thereby achieving high-precision modeling for complex tasks (such as natural language understanding and multimodal generation). Therefore, the model features and resource requirements of large-scale model training jobs can serve as inputs for cross-domain scheduling.
[0136] ① The specific features of the model include:
[0137] Model type: This section describes the type and version of the large model itself, such as GPT3 / 4, llama2 / 3 / 4, etc.
[0138] Parameter size: This section describes the parameter size of large model training jobs, such as 10B, 32B, 561B, etc.
[0139] Model layer count: This section describes the number of neural network layers in a large model training job, such as 96 layers, 256 layers, etc.
[0140] Parallel mode: This section describes the parallelism and degree of parallelism of large model training jobs, such as data parallelism DP=16, tensor parallelism TP=8, pipeline parallelism PP=32, etc.
[0141] Training methods: This section describes the training methods for large model training jobs, including pre-training, incremental training, and partial fine-tuning.
[0142] Training dataset: This section describes the name and location of the large model dataset, as well as the size and type of the data, such as hundreds of billions of token text data.
[0143] ② Resource requirements specifically include:
[0144] Computing resource requirements are mainly used to describe the needs of large model training jobs for accelerated computing power (GPU / NPU / TPU, etc.), including GPU type and quantity;
[0145] Storage resource requirements are mainly used to describe the storage resource requirements of pre-training data, checkpoint data, model parameter files, etc. for large model training jobs, including storage capacity, read and write performance, etc.
[0146] Network resource requirements are mainly used to describe the network resource requirements for the transmission of pre-training data, checkpoint data, and model parameter files in large model training jobs, as well as the communication between various tasks in large model training jobs, including network latency and bandwidth.
[0147] 2. Computing, storage, and network resource structure across wide area networks
[0148] The state information of resources covered by cross-domain scheduling is also an important component of scheduling input. In scenarios where computing and storage resources are distributed across domains, the resource state information that needs to be considered when scheduling large model training jobs across wide area networks includes computing, storage, and network resources.
[0149] ① Computing resource requirements: This mainly includes the computing power usage, memory utilization, and fault status of heterogeneous accelerator cards (such as GPUs / TPUs), specifically including indicators such as floating-point operations per second (FLOPS), remaining memory capacity, and CUDA core utilization.
[0150] ② Storage resource requirements: It is necessary to obtain the data distribution topology, access latency and throughput fluctuations of the storage system, including key information such as the location of the training dataset, the storage medium of the parameter server (NVMe / HDD), and the checkpoint write rate (Input / Output Operations Per Second, IOPS).
[0151] ③ Network Resource Requirements: Real-time collection of WAN link quality data is required, mainly including cross-domain backbone network bandwidth, end-to-end latency, packet loss rate, and other indicators. Furthermore, the WAN topology is also a component of the scheduling input. A global view of the topology needs to model the physical connections and logical architecture coupling relationships of cross-domain resources, including the multi-layered BGP (Interior Gateway Protocol) routing topology of the WAN, the rack-level layout of computing and storage nodes, and the PFC (Priority Flow Control) flow control status of RoCEv2. During 3D parallel training, a virtual communication domain needs to be constructed based on the topology-aware algorithm.
[0152] like Figure 3 As shown, the scheduling layer consists of a task scheduler, a data scheduler, and a traffic scheduler, while the management and control layer consists of a computing power controller, a storage controller, and a network controller. The functions of each component are described below:
[0153] ① Task Scheduler: This function handles cross-domain scheduling requests for large model training jobs and serves as the entry point for cross-domain scheduling. It has the capability to process the GPU computing power requirements of large model training jobs and allocate appropriate computing resources to them. It can also split the GPU computing power requirements of large model training jobs across multiple computing centers, enabling the deployment of large model training jobs across multiple computing centers.
[0154] ② Data Scheduler: Capable of handling the data storage needs of large model training jobs, allocating appropriate storage resources for training datasets, checkpoints, model parameter files, and other data from large model training jobs. It can collaborate with the task scheduler to synchronize the training datasets, checkpoints, and model parameter files of large model training jobs to multiple suitable computing centers.
[0155] ③ Traffic Scheduler: Capable of handling the communication and transmission traffic requirements of large model training jobs, allocating appropriate network resources for inter-task communication and data transmission in large model training jobs. It can collaborate with the task scheduler and data scheduler to open deterministic network paths for inter-task communication and data transmission in large model training jobs, ensuring the quality of service for communication and transmission.
[0156] ④ Computing Power Controller: Used to manage computing resources within the computing power center, it can handle computing power requests from the task scheduler. The computing power controller provides the task scheduler with the status of computing resources in its computing power center as the basis for task scheduling.
[0157] ⑤ Storage controller: This unit manages storage resources within the computing center and can handle storage requests from the data scheduler. The storage controller provides the data scheduler with information on the storage resource status of its computing center, serving as the basis for data scheduling.
[0158] ⑥ The network controller manages the network resources of its WAN and can handle network requests from the traffic scheduler. The network controller provides the traffic scheduler with the status of the network resources of its WAN as a basis for traffic scheduling.
[0159] like Figure 4 As shown, the scheduling process for cross-domain training of large models includes a scheme generation stage and a scheme deployment stage, specifically including the following steps:
[0160] S1. Submit large-scale cross-domain training jobs. Users can generate and submit large-scale model training jobs based on the model features and resource requirements of the model to be trained. The task scheduler takes over the large-scale model training jobs, parses the computing resource requirements, storage resource requirements, and network resource requirements, and guides the entire process of cross-domain scheduling of large-scale model training jobs. At the same time, the corresponding computing power resource status, storage resource status, and network resource status are obtained through the task scheduler, data scheduler, and traffic scheduler, respectively.
[0161] S2. Perform task scheduling. The task scheduler selects large model cross-domain training jobs to be scheduled from the large model job queue, then performs task scheduling, filters out the first set of scheduling schemes that meet the computing resource requirements of the large model training jobs based on the obtained computing resource status, and continues to execute S3. If there is no scheduling scheme that meets the computing resource requirements, it returns to execute S1.
[0162] S3. Request Data Scheduling. The task scheduler sends the first scheduling scheme set and storage resource requirements to the data scheduler.
[0163] S4. Execute data scheduling. After receiving a task scheduling request, the data scheduler iterates through the first scheduling scheme set, determining whether the computing center storage resource status of each scheme in the set meets the storage resource requirements of the large model training job. If not, the scheme is removed from the first scheduling scheme set, and a new scheduling result set, i.e., the second scheduling scheme set, is formed after the iteration. If the second scheduling scheme set is empty, return to execute S2; if the second scheduling scheme set is not empty, continue to execute S5.
[0164] S5. Return the data scheduling result. The data scheduler feeds back the second scheduling scheme set to the task scheduler.
[0165] S6. Request traffic scheduling. The task scheduler sends the second scheduling scheme set and network resource requirements to the traffic scheduler.
[0166] S7. Execute traffic scheduling. After receiving a task scheduling request, the traffic scheduler iterates through the second scheduling scheme set, determining whether the network resource status of the interconnected wide area network between computing centers for each scheme in the set meets the network resource requirements for inter-task communication and data transmission between computing centers for large model training jobs. If not, the scheme is removed from the second scheduling scheme set, and a new scheduling result set, i.e., the third scheduling scheme set, is formed after the iteration. If the third scheduling scheme set is empty, return to execute S2; if the third scheduling scheme set is not empty, continue to execute S8.
[0167] S8. Return the traffic scheduling result. The traffic scheduler feeds back the third scheduling scheme set to the task scheduler.
[0168] S9. Generate a scheduling scheme. The task scheduler calls the preset scheduling scheme selection strategy to determine the target scheduling scheme from the third scheduling scheme set.
[0169] S10. Synchronizing Data Transmission Network Scheme. The task scheduler generates a data transmission network scheme for transmitting training data based on the target scheduling scheme, and synchronizes this data transmission network scheme to the flow scheduler.
[0170] S11. Issue data transmission path activation task. After receiving the data transmission network plan, the traffic scheduler generates a data transmission path activation task according to the plan requirements and issues it to the network controller.
[0171] S12. Data transmission path activation. After receiving the data transmission path activation task, the network controller completes the activation of the data transmission path.
[0172] S13. Synchronous Storage Scheme. The task scheduler generates a data storage scheme for storing the dataset based on the target scheduling scheme, and synchronizes the data storage scheme to the data scheduler.
[0173] S14. Distribute data storage tasks. The data scheduler receives the data storage plan, generates data storage tasks for each computing center according to the plan's requirements, and distributes them to the storage controllers of the corresponding computing centers.
[0174] S15. Training dataset is stored across computing centers. After receiving the data storage task from the data scheduler, the storage controller of each computing center completes the allocation of storage resources for the corresponding computing center, and completes the data transfer and synchronization between the computing centers based on the established data transmission paths.
[0175] S16. Synchronizing the task communication network scheme. The task scheduler generates a task communication network scheme for cross-area communication of training tasks based on the target scheduling scheme, and synchronizes the task communication network scheme to the traffic scheduler.
[0176] S17. Issue the task communication path activation task. After receiving the task communication network plan, the traffic scheduler generates a task communication path activation task according to the plan requirements and issues it to the network controller.
[0177] S18. Task communication path activation. After receiving the task communication path activation task, the network controller completes the activation of the task communication path.
[0178] S19. Distribute large model training tasks. The task scheduler generates a training task deployment plan based on the target scheduling scheme, and generates training task deployment tasks for each computing center based on the deployment plan, and then distributes them to the computing power controller of the corresponding computing power center.
[0179] S20. Start large model training. After receiving the training task deployment task issued by the task scheduler, each computing power controller synchronously starts cross-domain training of the large model.
[0180] The scheduling method for cross-domain training of large models provided in this invention, through the collaboration of the scheduling layer and the control layer, not only meets the computational, storage and network resource requirements of cross-domain training of large models, but also realizes cross-domain data transmission and task communication by constructing data transmission paths and task communication paths, replacing high-cost dedicated lines, and effectively reducing training costs and scheduling complexity.
[0181] Example 3
[0182] Figure 5 This is a schematic diagram of a scheduling device for cross-domain training of a large model, provided in Embodiment 3 of the present invention. Figure 5As shown, this device is applied to a scheduling system that includes a scheduling layer and a control layer, and includes:
[0183] The acquisition module 21 is used to acquire the job requirements and resource status of large model cross-domain training jobs through the scheduling layer; the job requirements include computing resource requirements, storage resource requirements and network resource requirements;
[0184] The scheduling scheme generation module 22 is used to determine the corresponding target scheduling scheme based on job requirements and resource status through the scheduling layer;
[0185] The scheduling scheme deployment module 23 is used to construct data transmission paths and task communication paths through the scheduling layer control and management layer based on the target scheduling scheme, and to start cross-domain training of large models based on the data transmission paths and task communication paths.
[0186] Furthermore, based on the above embodiments of the invention, the scheduling layer includes a task scheduler, a data scheduler, and a traffic scheduler deployed on the global scheduling node; the management and control layer includes a computing power controller and a storage controller deployed within each computing power center, as well as a network controller deployed on each wide-area deterministic network control plane.
[0187] Furthermore, based on the above embodiments of the invention, the acquisition module 21 includes:
[0188] The job extraction unit is used to extract large model cross-domain training jobs from a preset job queue through the task scheduler; large model cross-domain training jobs are generated from the model features and resource requirements description of the model to be trained;
[0189] The job parsing unit is used to parse the computational resource requirements, storage resource requirements, and network resource requirements in large model cross-domain training jobs through the task scheduler.
[0190] Furthermore, based on the above embodiments of the invention, the scheduling scheme generation module 22 includes:
[0191] The resource status acquisition unit is used to acquire the corresponding computing resource status, storage resource status and network resource status through the task scheduler, data scheduler and traffic scheduler respectively.
[0192] The first screening unit is used to determine the first scheduling scheme set according to the computing resource requirements and computing power resource status through the task scheduler, and send the first scheduling scheme set and storage resource requirements to the data scheduler;
[0193] The second screening unit is used to select a second scheduling scheme set from the first scheduling scheme set by the data scheduler according to the storage resource requirements and storage resource status, and send the second scheduling scheme set to the task scheduler.
[0194] The sending unit is used to send the second scheduling scheme set and network resource requirements to the traffic scheduler through the task scheduler;
[0195] The third screening unit is used to select the third scheduling scheme set from the second scheduling scheme set by the traffic scheduler according to the network resource demand and network resource status, and send the third scheduling scheme set to the task scheduler.
[0196] The target scheduling scheme determination unit is used to determine the target scheduling scheme from the third scheduling scheme set by calling the preset scheduling scheme selection strategy through the task scheduler.
[0197] Furthermore, based on the above embodiments of the invention, the scheduling scheme deployment module 23 includes:
[0198] The scheme generation unit is used to determine the target computing power centers and target wide-area deterministic networks participating in cross-domain training according to the target scheduling scheme through the task scheduler, and to generate data transmission network scheme, data storage scheme, task communication network scheme and training task deployment scheme based on each target computing power center and target wide-area deterministic network.
[0199] The scheme sending unit is used to send the data transmission network scheme and the task communication network scheme to the traffic scheduler through the task scheduler, and to send the data storage scheme to the data scheduler.
[0200] The data transmission path activation task distribution unit is used to generate data transmission path activation tasks corresponding to each target wide area deterministic network according to the data transmission network scheme through the traffic scheduler, and distribute each data transmission path activation task to the network controller of the corresponding target wide area deterministic network.
[0201] The data transmission path activation unit is used to activate the corresponding data transmission path according to the respective data transmission path activation tasks of each network controller.
[0202] The data storage task distribution unit is used to generate data storage tasks corresponding to each target computing center according to the data storage scheme through the data scheduler, and distribute each data storage task to the storage controller of the corresponding target computing center.
[0203] The data storage task execution unit is used to execute the corresponding storage resource allocation operation according to the respective data storage tasks through each storage controller, and to complete the cross-computing center data synchronization operation based on the opened data transmission paths;
[0204] The task communication path activation task distribution unit is used to generate task communication path activation tasks corresponding to each target wide area deterministic network according to the task communication network scheme through the traffic scheduler, and to distribute each task communication path activation task to the network controller of the corresponding target wide area deterministic network.
[0205] The task communication path opening unit is used to open the corresponding task communication path for each task according to the respective task communication path of each network controller.
[0206] The training task deployment task distribution unit is used to generate training task deployment tasks corresponding to each target computing power center according to the training task deployment plan through the task scheduler, and distribute each training task deployment task to the computing power controller of the corresponding target computing power center.
[0207] The cross-domain training initiation unit is used to deploy tasks according to their respective training tasks and bind the corresponding task communication paths through each computing power controller, and to synchronously start cross-domain training of large models.
[0208] Furthermore, based on the above embodiments of the invention, the scheduling device further includes:
[0209] The computing power resource status reporting module is used to report the computing power resource status of the computing power center to the task scheduler at each computing power controller every first preset period or when a preset computing resource change event is detected.
[0210] The storage resource status reporting module is used to report the storage resource status of the computing center to the data scheduler every second preset period or when a preset storage resource change event is detected by each storage controller.
[0211] The network resource status reporting module is used to report the network resource status of the wide area deterministic network to the traffic scheduler every third preset period or when a preset network resource change event is detected.
[0212] The scheduling device for cross-domain training of large models provided in the embodiments of the present invention can execute the scheduling method for cross-domain training of large models provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0213] Example 4
[0214] Figure 6A schematic diagram of an electronic device 30 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0215] like Figure 6 As shown, the electronic device 30 includes at least one processor 31 and a memory, such as a read-only memory (ROM) 32 or a random access memory (RAM) 33, communicatively connected to the at least one processor 31. The memory stores computer programs executable by the at least one processor. The processor 31 can perform various appropriate actions and processes based on the computer program stored in the ROM 32 or loaded from storage unit 38 into the RAM 33. The RAM 33 can also store various programs and data required for the operation of the electronic device 30. The processor 31, ROM 32, and RAM 33 are interconnected via a bus 34. An input / output (I / O) interface 35 is also connected to the bus 34.
[0216] Multiple components in electronic device 30 are connected to I / O interface 35, including: input unit 36, such as keyboard, mouse, etc.; output unit 37, such as various types of monitors, speakers, etc.; storage unit 38, such as disk, optical disk, etc.; and communication unit 39, such as network card, modem, wireless transceiver, etc. Communication unit 39 allows electronic device 30 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0217] Processor 31 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 31 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 31 performs the various methods and processes described above, such as scheduling methods for cross-domain training of large models.
[0218] In some embodiments, the scheduling method for large model cross-domain training can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 38. In some embodiments, part or all of the computer program can be loaded and / or mounted on electronic device 30 via ROM 32 and / or communication unit 39. When the computer program is loaded into RAM 33 and executed by processor 31, one or more steps of the scheduling method for large model cross-domain training described above can be performed. Alternatively, in other embodiments, processor 31 can be configured to perform the scheduling method for large model cross-domain training by any other suitable means (e.g., by means of firmware).
[0219] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0220] In some embodiments, the scheduling method for cross-domain training of large models can be implemented as a computer program, which is implicitly included in a computer program product. When executed by a processor, the computer program implements the scheduling method for cross-domain training of large models of the present invention. The computer program product can be understood as a software product that primarily implements its solution through a computer program. The computer program used to implement the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a standalone software package, or entirely on a remote machine or server.
[0221] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0222] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0223] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0224] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0225] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0226] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A scheduling method for cross-domain training of large models, characterized in that, The method, applied to a scheduling system comprising a scheduling layer and a control layer, includes: The scheduling layer obtains the job requirements and resource status of large-scale cross-domain training jobs; the job requirements include computing resource requirements, storage resource requirements, and network resource requirements. The scheduling layer determines the corresponding target scheduling scheme based on the job requirements and the resource status. Based on the target scheduling scheme, the scheduling layer controls the management layer to construct data transmission paths and task communication paths, and starts cross-domain training of large models based on the data transmission paths and task communication paths.
2. The method according to claim 1, characterized in that, The scheduling layer includes a task scheduler, a data scheduler, and a traffic scheduler deployed on global scheduling nodes; the management and control layer includes a computing power controller and a storage controller deployed within each computing power center, as well as a network controller deployed on each wide-area deterministic network control plane.
3. The method according to claim 2, characterized in that, The process of obtaining the job requirements and resource status of large-scale cross-domain training jobs through the scheduling layer includes: The task scheduler extracts the large model cross-domain training job from the preset job queue; the large model cross-domain training job is generated by the model features and resource requirement description of the model to be trained; The task scheduler parses the computational resource requirements, storage resource requirements, and network resource requirements in the large model cross-domain training job.
4. The method according to claim 2, characterized in that, The step of determining the corresponding target scheduling scheme through the scheduling layer based on the job requirements and the resource status includes: The corresponding computing resource status, storage resource status, and network resource status are obtained through the task scheduler, the data scheduler, and the traffic scheduler, respectively. The task scheduler determines a first scheduling scheme set according to the computing resource requirements and the computing power resource status, and sends the first scheduling scheme set and the storage resource requirements to the data scheduler. The data scheduler selects a second scheduling scheme set from the first scheduling scheme set according to the storage resource requirements and the storage resource status, and sends the second scheduling scheme set to the task scheduler. The task scheduler sends the second scheduling scheme set and the network resource requirements to the traffic scheduler. The traffic scheduler selects a third scheduling scheme set from the second scheduling scheme set according to the network resource requirements and the network resource status, and sends the third scheduling scheme set to the task scheduler. The target scheduling scheme is determined in the third scheduling scheme set by invoking a preset scheduling scheme selection strategy through the task scheduler.
5. The method according to claim 2, characterized in that, The step of constructing data transmission paths and task communication paths by controlling the management layer through the scheduling layer based on the target scheduling scheme, and initiating large-scale cross-domain training of the model based on the data transmission paths and task communication paths, includes: The task scheduler determines each target computing power center and each target wide-area deterministic network participating in cross-domain training according to the target scheduling scheme, and generates a data transmission network scheme, a data storage scheme, a task communication network scheme, and a training task deployment scheme based on each target computing power center and each target wide-area deterministic network. The task scheduler sends the data transmission network scheme and the task communication network scheme to the traffic scheduler, and the data storage scheme to the data scheduler. The traffic scheduler generates data transmission path activation tasks corresponding to each of the target wide area deterministic networks according to the data transmission network scheme, and sends each of the data transmission path activation tasks to the network controller corresponding to the target wide area deterministic network. Each of the network controllers activates the corresponding data transmission path according to its respective data transmission path activation task. The data scheduler generates data storage tasks corresponding to each target computing center according to the data storage scheme, and sends each data storage task to the storage controller corresponding to the target computing center. Each of the aforementioned storage controllers performs the corresponding storage resource allocation operation according to its respective data storage task, and completes the cross-computing center data synchronization operation based on the opened data transmission paths. The traffic scheduler generates task communication path activation tasks corresponding to each of the target wide area deterministic networks according to the task communication network scheme, and sends each of the task communication path activation tasks to the network controller corresponding to the target wide area deterministic network. Each of the network controllers activates the corresponding task communication path according to its respective task communication path. The task scheduler generates training task deployment tasks corresponding to each target computing power center according to the training task deployment scheme, and sends each training task deployment task to the computing power controller corresponding to the target computing power center. Each computing power controller deploys tasks according to its respective training task, binds the corresponding task communication path, and synchronously starts cross-domain training of large models.
6. The method according to claim 2, characterized in that, Also includes: Each computing power controller reports the computing power resource status of its respective computing power center to the task scheduler at the first preset interval or when a preset computing resource change event is detected. Each of the aforementioned storage controllers reports the storage resource status of its respective computing center to the data scheduler every second preset period or when a preset storage resource change event is detected; Each of the aforementioned network controllers reports the network resource status of its respective wide-area deterministic network to the traffic scheduler every third preset period or when a preset network resource change event is detected.
7. A scheduling system for cross-domain training of large models, characterized in that, The scheduling method applied to cross-domain training of large models according to any one of claims 1-6, the system comprising: The scheduling layer includes task schedulers, data schedulers, and traffic schedulers deployed on global scheduling nodes; The control layer includes computing power controllers and storage controllers deployed within each computing power center, as well as network controllers deployed on each wide-area deterministic network control plane; Each of the computing power controllers is communicatively connected to the task scheduler, each of the storage controllers is communicatively connected to the data scheduler, and each of the network controllers is communicatively connected to the traffic scheduler.
8. A scheduling device for cross-domain training of large models, characterized in that, The apparatus, applied to a scheduling system comprising a scheduling layer and a control layer, includes: The acquisition module is used to acquire the job requirements and resource status of large model cross-domain training jobs through the scheduling layer; the job requirements include computing resource requirements, storage resource requirements and network resource requirements; The scheduling scheme generation module is used to determine the corresponding target scheduling scheme through the scheduling layer based on the job requirements and the resource status; The scheduling scheme deployment module is used to construct data transmission paths and task communication paths by controlling the management and control layer through the scheduling layer based on the target scheduling scheme, and to start cross-domain training of large models based on the data transmission paths and task communication paths.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the scheduling method for cross-domain training of large models according to any one of claims 1-6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the scheduling method for cross-domain training of a large model as described in any one of claims 1-6.
11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the scheduling method for cross-domain training of large models as described in any one of claims 1-6.
Citation Information
Cited By
Multi-scene adaptive cloud machine collaboration method
CN121309584A