A model training control method, program product, electronic device, and medium
By expanding the physical network interface card (NIC) of the training server to store intermediate information in parallel using virtual NICs, the problem of low efficiency in saving intermediate information during model training is solved, and more efficient intermediate information saving is achieved.
Patent Information
- Application Number
- CN202511311807.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-09-15
AI Technical Summary
The low efficiency of saving intermediate information during model training has become a performance bottleneck.
By expanding the training server's physical network interface card (NIC) into multiple virtual NICs, and segmenting intermediate information based on the virtual NIC information, the intermediate information is stored in parallel to the storage server using multiple virtual NICs, thereby improving the efficiency of intermediate information storage.
It makes more efficient use of the network transmission resources of the training server and improves the efficiency of saving intermediate information during model training.
Smart Images

Figure CN120803755B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of model training technology, and in particular to a control method, program product, electronic device and medium for model training. Background Technology
[0002] With the rapid development of deep learning technology, especially the widespread application of large models in natural language processing, computer vision, and multimodal fusion, the demand for computing and storage resources during model training has exploded. As the number of model parameters increases, the amount of intermediate information that needs to be stored during training also grows larger, and storing this intermediate information has gradually become a significant performance bottleneck during training. Summary of the Invention
[0003] This application provides a control method, program product, electronic device, and medium for model training, to at least solve the problem of low efficiency in saving intermediate information during model training in related technologies.
[0004] This application provides a model training control method, comprising: acquiring intermediate information of a target model training task executed on a training server, wherein the training server is used to execute multiple model training tasks, the multiple model training tasks including a target model training task, the training server is connected to a storage server through multiple training physical network cards, the training physical network cards are expanded into multiple virtual network cards, the model training task is assigned virtual network cards on at least two training physical network cards, and the intermediate information is used to indicate the training status of the target model training task during execution; segmenting the intermediate information according to the virtual network card information of the target virtual network card assigned to the target model training task to obtain multiple sets of corresponding target virtual network cards and intermediate sub-information, wherein the virtual network card information is used to indicate the network status of the information transmission network provided by the target virtual network card; and storing the corresponding intermediate sub-information in parallel to the storage server through the target virtual network card.
[0005] This application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of any of the above-described model training control methods.
[0006] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the control method for training any of the above-described models.
[0007] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of any of the above-described model training control methods.
[0008] This application obtains intermediate information of the target model training task executed on the training server. This intermediate information indicates the training status of the target model training task during execution. The intermediate information is segmented based on the network conditions of the information transmission network provided by the target virtual network interface card (NIC) allocated to the target NIC, resulting in multiple sets of corresponding target NICs and intermediate sub-information. The corresponding intermediate sub-information is then stored in parallel to the storage server via the target NICs. The training physical NICs on the training server are expanded into multiple virtual NICs. The model training task executed on the training server, including the target model training task, is allocated virtual NICs on at least two training physical NICs. This allows the target model training task to utilize the network resources of at least two training physical NICs to store intermediate information in parallel, thus making more efficient use of the network transmission resources on the training server. Therefore, this application solves the technical problem of low efficiency in saving intermediate information during model training in related technologies, achieving the technical effect of improving the efficiency of saving intermediate information during model training. Attached Figure Description
[0009] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a hardware structure block diagram of the model training control method according to an embodiment of this application;
[0011] Figure 2 This is a flowchart of a control method for model training according to an embodiment of this application;
[0012] Figure 3 This is an implementation architecture diagram of a model training control method according to an embodiment of this application;
[0013] Figure 4 This is a schematic diagram of a storage service mounting method according to an embodiment of this application. Figure 1 ;
[0014] Figure 5 This is a schematic diagram of a storage service mounting method according to an embodiment of this application. Figure 2 ;
[0015] Figure 6 This is a schematic diagram of an intermediate information storage method according to an embodiment of this application;
[0016] Figure 7 This is a schematic diagram of an intermediate information recovery method according to an embodiment of this application;
[0017] Figure 8 This is a structural block diagram of a model training control device according to an embodiment of this application. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0019] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0020] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0021] The specific application environment architecture or specific hardware architecture on which the control method for model training depends is described here.
[0022] The methods and embodiments provided in this application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a hardware structure block diagram of the model training control method according to an embodiment of this application. For example... Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown in the image. A processor 102 (which may include, but is not limited to, a central processing unit (CPU), microprocessor (MCU), or programmable logic device (FPGA), etc.) and a memory 104 for storing data are also shown. The server device may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the server equipment described above. For example, the server equipment may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0023] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the model training control method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the aforementioned method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to server devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0024] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the server device. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0025] The embodiments of this application provide a control method for model training, and the method is described in detail in conjunction with the execution flow of the control method for model training.
[0026] This embodiment provides a method for controlling model training. Figure 2 This is a flowchart of a control method for model training according to an embodiment of this application, such as... Figure 2 As shown, the method includes the following steps:
[0027] Step S202: Obtain intermediate information of the target model training task executed on the training server. The training server is used to execute multiple model training tasks, including the target model training task. The training server is connected to the storage server through multiple training physical network cards. The training physical network cards are expanded into multiple virtual network cards. The model training task is assigned virtual network cards on at least two training physical network cards. The intermediate information is used to indicate the training status of the target model training task during execution.
[0028] Step S204: Based on the virtual network card information of the target virtual network card assigned in the target model training task, the intermediate information is segmented to obtain multiple sets of target virtual network cards and intermediate sub-information with corresponding relationships. Among them, the virtual network card information is used to indicate the network status of the information transmission network provided by the target virtual network card.
[0029] Step S206: Store the corresponding intermediate sub-information in parallel to the storage server through the target virtual network card.
[0030] Through the above steps, intermediate information of the target model training task executed on the training server is obtained. This intermediate information indicates the training status of the target model training task during execution. The intermediate information is segmented based on the network conditions of the information transmission network provided by the target virtual network card, as indicated by the virtual network card information of the target virtual network card allocated to the target model training task. This results in multiple sets of corresponding target virtual network cards and intermediate sub-information. The corresponding intermediate sub-information is then stored in parallel to the storage server via the target virtual network card. The training physical network card on the training server is expanded into multiple virtual network cards. The model training task executed on the training server, including the target model training task, is allocated virtual network cards on at least two training physical network cards. This allows the target model training task to utilize the network resources of at least two training physical network cards to store intermediate information in parallel, thus making more efficient use of the network transmission resources on the training server. Therefore, this solves the technical problem of low efficiency in saving intermediate information during model training in related technologies, achieving the technical effect of improving the efficiency of saving intermediate information during model training.
[0031] In the embodiment provided in step S202, the training server may be, but is not limited to, a server used for training models, and the training server may, but is not limited to, train one or more models by executing multiple model training tasks.
[0032] Optionally, in this embodiment, a model to be trained may correspond to one or more model training tasks, but is not limited to. When the complexity of the model to be trained is high (e.g., when the complexity parameter used to indicate complexity is greater than a complexity parameter threshold), the model training may be divided into multiple model training tasks to be executed, and multiple model training tasks may be executed using multiple accelerator cards on the training server. Alternatively, when the complexity of the model to be trained is low (e.g., when the complexity parameter used to indicate complexity is less than or equal to a complexity parameter threshold), the model training may be treated as a single model training task, and this model training task may be executed using one accelerator card on the training server. In this case, other accelerator cards on the training server may be used, but are not limited to, to execute model training tasks related to other models to be trained.
[0033] Optionally, in this embodiment, the training server may, but is not limited to, connect to the storage server through multiple training physical network cards, and may, but is not limited to, store the intermediate information generated on the training server on the storage server through multiple training physical network cards.
[0034] Optionally, in this embodiment, the connection between the training server and the storage server can be achieved through multiple training physical network cards on the training server and multiple storage physical network cards on the storage server, but is not limited to.
[0035] Optionally, in this embodiment, the number of training physical network cards on the training server may be, but is not limited to, greater than, equal to, or less than the number of multiple model training tasks executed on the training server.
[0036] Optionally, in this embodiment, intermediate information may be used, but is not limited to, to indicate the training state of the target model training task during execution. Intermediate information may be, but is not limited to, data generated at various stages of training, used for subsequent recovery of the training state or evaluation of the model training process. Intermediate information may include, but is not limited to, checkpoint data. More specifically, intermediate information may include, but is not limited to, model weight parameters, optimizer state, training iteration count, and loss function value, etc.
[0037] Optionally, in this embodiment, the training physical network interface card (NIC) on the training server can be, but is not limited to, expanded into multiple virtual NICs. NIC services for the target model training task can be provided through virtual NICs on multiple training physical NICs, but is not limited to.
[0038] Optionally, in this embodiment, the virtual network interface card (NIC) may be created by the virtualization layer, and may, but is not limited to, mimicking the functions of a physical NIC in software, including MAC address (Media Access Control address), IP address (Internet Protocol address), and data packet transmission and reception. Each virtual NIC can be configured with an independent IP address, subnet mask, and network interface name, thereby logically isolating the network traffic of different virtual machines or containers. The bandwidth, flow control, and other resources of the physical NIC may, but are not limited to, be shared by multiple virtual NICs, or the resource usage of each virtual NIC may be restricted by configuration to ensure overall network performance.
[0039] Optionally, in this embodiment, when a training physical network interface card (NIC) is expanded into multiple virtual NICs, bandwidth quotas can be set for each virtual NIC, but are not limited to. When setting bandwidth quotas, each virtual NIC can be allocated a full bandwidth quota, ensuring that the network interface card resources of the training physical NIC are fully utilized. For example, for a training physical NIC A with bandwidth 'a', it can be expanded into three virtual NICs, A1, A2, and A3. A bandwidth quota of 'a' can be allocated to A1, A2, and A3, respectively, ensuring sufficient bandwidth for each NIC when used individually. Alternatively, a larger bandwidth quota can be configured for each virtual NIC, reserving some bandwidth to handle situations where multiple model training tasks simultaneously trigger the saving of intermediate information. For example, A1 can be allocated a bandwidth quota of 4a / 5, A2 can be allocated a bandwidth quota of 4a / 5, and A3 can be allocated a bandwidth quota of 4a / 5.
[0040] Optionally, in this embodiment, at least two virtual network cards on the training physical network cards may be allocated to each model training task, but not limited to. This can also be understood as at least two training physical network cards being allocated to each model training task, so that each model training task can use multiple information transmission channels when saving intermediate information, thereby accelerating the process of saving intermediate information.
[0041] Optionally, in this embodiment, at least two virtual network cards on the physical network cards of each training model may be randomly assigned to each training model training task. However, in order to avoid wasting the physical network cards of the training model, the virtual network card on each physical network card may be assigned to a model training task.
[0042] Optionally, in this embodiment, to better utilize the various training physical network cards on the training server, the processes of expanding virtual network cards and allocating virtual network cards can be combined, but are not limited to. Before obtaining intermediate information of the target model training task executed on the training server, a corresponding model training task can be allocated to each training physical network card, resulting in multiple training physical network cards and a main model training task with corresponding relationships. Each training physical network card is expanded into a main virtual network card and at least one secondary virtual network card, and bandwidth quota information is configured for the main and secondary virtual network cards. The bandwidth quota file is used to indicate the full bandwidth quota allocated to the main or secondary virtual network card when it is used independently. When primary and secondary virtual network interface cards (NICs) are used simultaneously, a first bandwidth quota is allocated to the primary NIC, and a second bandwidth quota is allocated to the secondary NIC. The first quota is greater than the second quota, and the sum of the first and second quotas is less than the full quota. This bandwidth quota represents the allowed bandwidth. The primary NIC in each training physical NIC is assigned to the corresponding primary model training task, and at least one secondary NIC is assigned to each model training task. The primary and secondary NICs assigned to each model training task are obtained by extending different training physical NICs. It is possible, but not limited to, establishing a correspondence between the virtual NICs and storage services on the storage server after allocating virtual NICs to each model training task, and then assigning each storage service to the model training task corresponding to the virtual NIC of that storage service, thus obtaining multiple information transmission channels for the model training tasks. This approach fully utilizes the bandwidth resources of the training physical NICs and avoids bandwidth shortages caused by simultaneous use. Furthermore, by assigning each training physical NIC to a specific model training task, the primary purpose of each physical NIC is clearly defined, avoiding resource idleness and waste. Each physical network interface card (NIC) has a clearly defined task objective, making network resource allocation more targeted. Furthermore, by assigning the primary virtual NIC to the primary model training task, and then allocating secondary virtual NICs to each model training task, the priority of tasks on each training physical NIC is clearly defined. This allocation method ensures that each primary model training task can preferentially obtain sufficient network resources from its corresponding training physical NIC. From an overall perspective, this guarantees that intermediate information from model training tasks on the training server will not be lost due to network resource contention.
[0043] Optionally, in this embodiment, since the model training tasks executed by the training server may, but are not limited to, simultaneously trigger the saving operation of intermediate information for the model training tasks, the network transmission resources of other training physical network cards on the training server are actually wasted when using the training physical network card corresponding to the model training task to store intermediate information. It is possible, but not limited to, to expand each training physical network card into multiple virtual network cards using the method of this application, with each model training task allocated virtual network cards on at least two training physical network cards. This allows for more efficient use of the training server's training network card resources when saving the intermediate information of a model training task, thereby improving the storage efficiency of intermediate information.
[0044] Optionally, in this embodiment, when the target model training task is detected to have entered the intermediate information saving state, various data in the target model training task can be extracted to obtain the intermediate information of the target model training task.
[0045] Optionally, in this embodiment, if an abnormal state of the training server is detected, various data in the target model training task can be extracted to obtain intermediate information of the target model training task.
[0046] In the embodiment provided in step S204, intermediate information can be segmented according to the virtual network card information of the target virtual network card assigned to the target model training task, but not limited to this step.
[0047] Optionally, in this embodiment, the virtual network interface card (NIC) information may be used, but is not limited to, to indicate the network conditions of the information transmission network provided by the target virtual NIC. For example, it may be used, but is not limited to, to indicate the network bandwidth provided by the target virtual NIC.
[0048] Optionally, in this embodiment, the intermediate information can be segmented according to the data size, or according to the bandwidth ratio between the target virtual network cards allocated to the target model training task, and the segmented intermediate sub-information can be allocated to the corresponding target virtual network cards according to the bandwidth ratio.
[0049] Optionally, in this embodiment, intermediate information can be segmented according to the target virtual network interface card (NIC) level and the target virtual NIC bandwidth quota. The virtual NIC information can include, but is not limited to, the aforementioned bandwidth quota. Specifically, target sub-information can be segmented from the intermediate information according to the bandwidth quota of the primary virtual NIC of the target model training task, obtaining the remaining sub-information of the remaining part of the intermediate information and the primary virtual NIC and target sub-information with corresponding relationships; the remaining sub-information can be segmented according to the quota ratio between the bandwidth quotas of each secondary virtual NIC of the target model training task; the remaining intermediate sub-information obtained after segmenting the remaining sub-information can be allocated to the corresponding secondary virtual NICs according to the quota ratio, resulting in one or more sets of secondary virtual NICs and remaining intermediate sub-information with corresponding relationships. Among them, the multiple sets of target virtual NICs and intermediate sub-information with corresponding relationships include primary virtual NICs and target sub-information with corresponding relationships, and one or more sets of secondary virtual NICs and remaining intermediate sub-information with corresponding relationships. By using the methods described above, we can make the most of virtual network interface cards that prioritize the target model training task, thereby reducing the impact of other model training tasks on the intermediate information storage process of the target model training task.
[0050] In the embodiment provided in step S206, intermediate sub-information corresponding to the target virtual network card can be stored in the storage server through the target virtual network card and the storage physical network card on the storage server, but is not limited to.
[0051] Optionally, in this embodiment, the target virtual network interface card assigned to the target model training task may, but is not limited to, correspond one-to-one with a portion or all of the multiple storage physical network interfaces on the storage server.
[0052] Optionally, in this embodiment, in order to further improve the parallel storage rate, multiple independent remote storage services corresponding one-to-one with multiple physical network cards can be set up on the storage server. Each remote storage service can independently manage the storage process of each target virtual network card-physical network card.
[0053] Optionally, in this embodiment, intermediate sub-information can be stored in parallel to the storage server according to the importance of the components in each intermediate sub-information through the target virtual network interface card (NIC). For example, for intermediate sub-information 1, which includes two parts 1-1 and 1-2, and intermediate sub-information 2, which includes two parts 2-1 and 2-2, 1-1 can be stored to the storage server first, and then 1-2 can be stored to the storage server, independently but synchronously, through the target virtual NIC 1; and 2-1 can be stored to the storage server first, and then 2-2 can be stored to the storage server, through the target virtual NIC 2. In this case, the importance of 1-1 is higher than that of 1-2, and the importance of 2-1 is higher than that of 2-2.
[0054] Optionally, in this embodiment, but not limited to, the intermediate sub-information can be stored in parallel to the storage server according to the generation time of the components in each intermediate sub-information through the target virtual network card.
[0055] Optionally, in this embodiment, to ensure the reasonable segmentation of various intermediate information, a corresponding target virtual network interface card (VNIC) may be assigned to the target model training task before segmenting the intermediate information according to the VNIC information of the target VNIC allocated to the target model training task. The operation of assigning a corresponding target VNIC to the target model training task may occur before, but is not limited to, obtaining the intermediate information of the target model training task, or it may occur after obtaining the intermediate information of the target model training task. It is also possible, but is not limited to, assigning a corresponding target VNIC to the target model training task while simultaneously assigning corresponding VNICs to other model training tasks.
[0056] Optionally, in this embodiment, to ensure the smooth parallel storage of intermediate sub-information and reduce interference between the storage of various intermediate sub-information, multiple information transmission channels, including corresponding storage directories and virtual network cards, can be created before, but are not limited to, the corresponding intermediate sub-information is stored in parallel to the storage server via the target virtual network card. The creation of information transmission channels can occur before, but is not limited to, the splitting of intermediate information, or it can occur after, but is not limited to, the splitting of intermediate information. The creation of information transmission channels can be performed after, but is not limited to, allocating the corresponding target virtual network card for the target model training task, thus centrally completing the preparation work for the parallel transmission of intermediate information.
[0057] Optionally, in this embodiment, when storing the corresponding intermediate sub-information in parallel to the storage server through the target virtual network card, the storage status of each intermediate sub-information may be recorded, including but not limited to the relationship between intermediate sub-information, the virtual network card that transmits the intermediate sub-information, and the final storage location of the intermediate sub-information, etc.
[0058] Optionally, in this embodiment, after storing the corresponding intermediate sub-information in parallel to the storage server through the target virtual network card, when it is necessary to restore the historical training state of the model in the future, the intermediate sub-information can be retrieved from the storage server to the training server using the recorded storage status of the intermediate sub-information, and then restored as intermediate information.
[0059] Optionally, in this embodiment, the training process of the model can be controlled by combining a Kubernetes (container orchestration system) environment and SR-IOV (Single Root I / O Virtualization) multi-NIC technology. Kubernetes can be applied to the training server, which acts as a Kubernetes node, and each model training task is executed in a group of training containers within the Kubernetes node. Multiple SR-IOV VFs (virtual function NICs) can be configured in the training container group, and these NICs can be connected to multi-NIC storage servers with independent storage services, laying the foundation for parallel splitting and sharding of checkpoints (intermediate information). The model training task can, but is not limited to, automatically parse the tensor structure when saving checkpoints, and can, but is not limited to, shard according to the number of VFs, and use a dedicated checkpoint migration service to write the sharded data at high speed and in parallel to remote storage such as NFS (Network File System) within the corresponding network segment. In handling the aforementioned storage process, when the training task restarts, the migration service can also pull shards in parallel from each VF and assemble them to restore the complete checkpoint, thereby accelerating loading. Compared to the traditional single-channel write method, the method in this application does not rely on the complex RDMA (Remote Direct Memory Access) protocol stack configuration. It fully utilizes the low latency, high bandwidth of SR-IOV and the resource isolation capabilities in the Kubernetes environment, significantly improving the concurrency of checkpoint saving and loading and the overall I / O efficiency of the training task. It has significant advantages such as flexible deployment, strong scalability, and high engineering feasibility.
[0060] Specifically, multiple physical network interface cards (NICs) can be configured on the storage server, with each NIC bound to an independent remote storage service (such as NFS), and each service residing in an independent subnet. Multiple SR-IOV Virtual File Exchanges (VFs) can be configured within the training container group, ensuring each VF is on the same network segment as its corresponding storage service, achieving high-speed point-to-point connections and thus constructing multi-channel data transmission paths. When saving checkpoints, the model training task parses the tensor data in the checkpoint based on the number of VF NICs in the training container group and intelligently splits it, dividing large checkpoints into multiple independent data shards (i.e., intermediate sub-information) to adapt to multi-channel parallel storage. During the saving phase, the migration service in the training container group can, but is not limited to, transfer each shard of data to multiple remote storage services via its corresponding VF, based on the pre-parsed tensor shard information, constructing a distributed checkpoint data structure. This process can, but is not limited to, be executed in parallel with the main training task to avoid blocking training progress. During the model recovery phase, the migration service concurrently pulls the corresponding data shards through multiple VF network cards and reassembles them into a complete model checkpoint locally on the training server. This supports the rapid recovery of large models in multi-node or multi-stage training and significantly reduces the reload time required for training interruptions.
[0061] Figure 3 This is an implementation architecture diagram of a model training control method according to an embodiment of this application. Figure 3 As shown, the training server is a Kubernetes node, which may include, but is not limited to, three training container groups: container group 1, container group 2, and container group 3. Each training container group may be used to execute different model training tasks. The training server may have, but is not limited to, three physical network interfaces (NICs), and may, but is not limited to, each physical NIC being expanded into three virtual NICs, with one virtual NIC from each physical NIC assigned to each training container group, as shown below. Figure 3 As shown, each training container group is allocated three virtual network interface cards (NICs). The storage server may, but is not limited to, have three physical storage NICs and corresponding storage services. Container group 1 may, but is not limited to, transmit intermediate sub-information in parallel through virtual LAN transmission channels formed by the allocated virtual NICs and the corresponding physical storage NICs, thereby accelerating the storage speed of intermediate information in container group 1. Container groups 2 and 3 are similar to container group 1 and will not be described further.
[0062] As an optional implementation, intermediate information of the target model training task executed on the training server can be obtained, but is not limited to, through the following methods: detecting instructions called by the training container through a migration container, wherein the training container group includes a training container and a migration container, the training container group is used to execute the target model training task, the training container is used to execute the training process of the target model training task, and the migration container is used to control intermediate information; when a save instruction is detected by the training container, intermediate information is obtained from the target storage location in the training container group through the migration container, wherein the target storage location is used to provide storage space for the training container.
[0063] Optionally, in this embodiment, each training container group may include, but is not limited to, a training container and a transfer container. The training container performs the model training process, and the transfer container performs the intermediate information control process.
[0064] By setting up the training container and the transfer container as described above, the training process and the intermediate information storage process can be separated, reducing the impact of the intermediate information storage process on the training process and improving the execution efficiency of the model training task.
[0065] As an optional implementation, before segmenting the intermediate information based on the virtual network card information of the target virtual network card allocated to the target model training task, virtual network cards may be allocated to the target model training task in the following ways, but not limited to: expanding each training physical network card into multiple virtual network cards according to the number of multiple model training tasks executed on the training server; allocating virtual network cards on at least two training physical network cards to the target model training task according to the task information of the target model training task, wherein the task information is used to indicate the complexity of the target model training task.
[0066] Optionally, in this embodiment, in order to ensure that each model training task has the possibility of using each training physical network card, each training physical network card can be expanded to a number of virtual network cards equal to the number of tasks.
[0067] Optionally, in this embodiment, the first training physical network interface card (NIC) can be expanded into a number of virtual NICs corresponding to the number of tasks, and the second training physical NIC can be expanded into two virtual NICs. The bandwidth of the first training physical NIC is greater than a bandwidth threshold, and the bandwidth of the second training physical NIC is either less than or greater than the bandwidth threshold. Through these operations, excessively small bandwidth can be avoided from participating in too much intermediate information transmission, and excessively small virtual bandwidth allocation can be avoided.
[0068] Based on the above, at least two virtual network cards on the physical network cards of the target model are allocated for the training task, which makes it possible to use multiple physical network cards of the training model to transmit intermediate information.
[0069] As an optional implementation, each training physical network interface card (NIC) can be expanded into multiple virtual NICs based on the number of model training tasks executed on the training server, but not limited to: the number of detection tasks; expanding each training physical NIC into multiple virtual NICs, wherein the number of virtual NICs obtained by expanding each training physical NIC is equal to the number of tasks; determining the address information of the multiple virtual NICs obtained by expanding each training physical NIC based on the address information of each storage physical NIC on the storage server, wherein the address information of the multiple virtual NICs obtained by expanding each training physical NIC corresponds one-to-one with the address information of each storage physical NIC, and the address information is used to indicate the subnet where the NIC device is located. The NIC device includes multiple storage physical NICs and multiple virtual NICs obtained by expanding multiple training physical NICs, and the multiple storage physical NICs are in different subnets.
[0070] Optionally, in this embodiment, address information can be determined for multiple virtual network interfaces extended from a training physical network interface card (NIC) based on, but is not limited to, the address information of multiple physical network interfaces on the storage server. Specifically, IP addresses in the same subnet as the IP addresses of each physical network interface card can be determined for the multiple virtual network interfaces extended from a training physical network interface card based on, but is not limited to, the IP addresses of each physical network interface card. For example, for training physical network interface card A, virtual network interface card A1 is in the same subnet 1 as physical network interface card A, virtual network interface card A2 is in the same subnet 2 as physical network interface card B, and virtual network interface card A3 is in the same subnet 3 as physical network interface card C.
[0071] Optionally, in this embodiment, multiple high-performance physical network interface cards (e.g., 10 Gigabit Ethernet cards) can be pre-configured on the remote storage server (i.e., the storage server) to carry multiple independent storage service traffic streams. Each storage physical network interface card can be bound to an independent IP address, and each can be placed in a different subnet (e.g., 192.168.10.0 / 24, 192.168.11.0 / 24) to achieve network-level isolation. An independent remote file system service (i.e., storage service), such as NFS service instances nfs-a, nfs-b, etc., can be deployed on the network interface card of each subnet, allowing it to listen to the corresponding subnet. The SR-IOV function of the physical server (i.e., the training server) can be enabled, and the SR-IOV CNI (Single Root I / O Virtualization Container Network Interface) plugin can be deployed on the Kubernetes node (i.e., the training server). Multiple SR-IOV virtual network interfaces (VFs) can be declared in the training container group specification, but are not limited to. Each VF can be assigned to the container runtime environment via an SR-IOV Device Plugin, but is not limited to. Each VF can be assigned an IP address consistent with the target storage service subnet, and static routes can be added to ensure that the VF can only access storage services within its corresponding subnet.
[0072] Based on the above, each training physical network card is expanded into a virtual network card in the same subnet as each storage physical network card. This provides a basis for simultaneously transmitting different parts of intermediate information through independent transmission channels, avoiding mutual interference when transmitting different parts of intermediate information.
[0073] As an optional implementation, virtual network interfaces on at least two training physical network interfaces can be allocated to the target model training task based on the task information of the target model training task in the following manner: matching the target task parameters with the corresponding number of target network interfaces from the corresponding task parameters and network interface numbers, wherein the target task parameters are the number of model parameters to be trained in the target model training task, and the task information includes the task parameters; selecting the target number of training physical network interfaces from multiple training physical network interfaces to obtain multiple target training network interfaces; and allocating one virtual network interface from each target training network interface to the target model training task to obtain multiple target virtual network interfaces corresponding to the target model training task.
[0074] Optionally, in this embodiment, the correspondence between task parameters and the number of network cards may be recorded in the training server, and the number of virtual network cards to be allocated for the target model training task may be determined based on the corresponding task parameters and the number of network cards.
[0075] Based on the above, different training physical network cards are assigned to model training tasks of different complexity, avoiding over-allocation and waste of training physical network cards.
[0076] As an optional implementation, before matching corresponding information transmission channels for each intermediate sub-information based on the correspondence between each target virtual network interface and each intermediate sub-information, multiple information transmission channels can be created, but are not limited to, in the following ways: matching corresponding storage services for each target virtual network interface based on the address information of each target virtual network interface, resulting in multiple sets of target virtual network interfaces and storage services with corresponding relationships, wherein each storage service is used to listen to the subnet where multiple storage physical network interfaces on the storage server are located; mounting each storage service to multiple storage directories of the target model training task based on the multiple sets of target virtual network interfaces and storage services with corresponding relationships, resulting in multiple information transmission channels, wherein the information transmission channels include target virtual network interfaces and storage directories with corresponding relationships.
[0077] Optionally, in this embodiment, in the aforementioned training server combined with Kubernetes, multiple information transmission channels can be created after allocating a virtual network card for the target model training task, to centrally complete the preparation work for parallel storage of intermediate information.
[0078] Optionally, in this embodiment, the mounted VF CIDR (Classless Inter-Domain Routing) information (i.e., address information) can be parsed, and the corresponding remote storage service (i.e., storage service) can be matched based on the CIDR information. For example:
[0079] VF-0 CIDR→nfs-a;
[0080] VF-1 CIDR→nfs-b;
[0081] VF-2 CIDR→nfs-c.
[0082] Optionally, in this embodiment, by creating multiple information transmission channels, a basis is provided for the parallel transmission of multiple intermediate sub-information, avoiding mutual interference in the transmission and storage processes of multiple intermediate sub-information, and ensuring the storage quality of intermediate information.
[0083] As an optional implementation, but not limited to, multiple information transmission channels can be obtained by mounting various storage services to multiple storage directories of the target model training task based on multiple sets of corresponding target virtual network interface cards (NICs) and storage services in the following manner: detecting the start command of the training container group, wherein the training container group is used to execute the target model training task, and the start command is used to start the training container group; upon detecting the start command, mounting each storage service to each storage directory, thereby obtaining multiple sets of corresponding storage services and storage directories, wherein the storage directories are used to store data generated during the operation of the training container group; and determining the storage directory corresponding to each target virtual network interface card based on the multiple sets of corresponding target virtual network interface cards and storage services, and the multiple sets of corresponding storage services and storage directories, thereby obtaining multiple information transmission channels.
[0084] Optionally, in this embodiment, each storage service may be mounted to multiple storage directories configured for the training container group before the training container group is started, but not limited to this.
[0085] Optionally, in this embodiment, Figure 4 This is a schematic diagram of a storage service mounting method according to an embodiment of this application. Figure 1 ,like Figure 4 As shown, a pre-mount command operation can be injected into the startup command of the initial container group through the Kubernetes webhook mechanism, but not limited to. This operation is responsible for mounting remote storage services. Multiple remote storage paths can be mounted to different mount points within the training container group (such as / mnt / a, / mnt / b, etc.) through mounting remote storage services. Figure 5 This is a schematic diagram of a storage service mounting method according to an embodiment of this application. Figure 2 ,like Figure 5 As shown, it is possible, but not limited to, to traverse each target virtual network interface, determine the storage service corresponding to each target virtual network interface, mount the storage service to a storage directory, and if the mounting is successful, proceed to mount the storage service corresponding to the next target virtual network interface. If the mounting fails, end the mounting operation and issue an alarm to wait for further processing.
[0086] Optionally, in this embodiment, it is possible, but not limited to, to configure each virtual network card in the training container group to establish a point-to-point mapping connection with a specific subnet / storage service on the remote storage server (i.e., the storage server) to ensure the determinism and exclusivity of the data path (i.e., the information transmission channel).
[0087] As an optional implementation, intermediate information can be segmented based on the virtual network interface card (NIC) information of the target NIC assigned to the target model training task in the following ways: extracting multiple intermediate arrays included in the intermediate information and determining the array parameters of the intermediate arrays, wherein the array parameters are used to indicate the data size of the intermediate arrays; segmenting or aggregating multiple intermediate arrays according to the bandwidth parameters of each target NIC and each array parameter to obtain multiple sets of target NICs and intermediate sub-information with corresponding relationships, wherein the bandwidth parameters are used to indicate the bandwidth allowed to be used by the target NIC, and the virtual NIC information includes the bandwidth parameters.
[0088] Optionally, in this embodiment, the intermediate information may include, but is not limited to, multiple intermediate arrays, and the intermediate arrays may be, but are not limited to, tensors.
[0089] Optionally, in this embodiment, to achieve parallel storage, intermediate information may be segmented, but is not limited to this. Since the intermediate information includes one or more intermediate arrays, the segmentation of the intermediate information may, but is not limited to, the segmentation or aggregation of intermediate arrays. Large intermediate arrays may be divided into several parts, and multiple smaller intermediate arrays may be aggregated into a larger part to better adapt to the bandwidth allowed by the virtual network interface card, ultimately improving the speed of parallel storage.
[0090] As an optional implementation, multiple intermediate arrays can be segmented or aggregated based on the bandwidth parameters and array parameters of each target virtual network interface card (NIC) to obtain multiple sets of corresponding target virtual NICs and intermediate sub-information: Calculating the proportion of the bandwidth parameter of each target virtual NIC to the total bandwidth parameter to obtain the target proportion of each target virtual NIC, wherein the total bandwidth parameter indicates the total allowed bandwidth of the multiple target virtual NICs; calculating the parameters to be allocated for each target virtual NIC based on the target proportion of each target virtual NIC and the total array parameters of the multiple intermediate arrays, wherein the total array parameters indicate the total data size of the multiple intermediate arrays, and the parameters to be allocated indicate the amount of data to be allocated to the target virtual NIC; segmenting or aggregating multiple intermediate arrays based on the parameters to be allocated and the array parameters to obtain multiple sets of corresponding target virtual NICs and intermediate sub-information.
[0091] Optionally, in this embodiment, the amount of intermediate sub-information to be transmitted to each target virtual network interface card (NIC) can be calculated based on, but is not limited to, the ratio between the bandwidth parameters of each virtual NIC and the total amount of intermediate information to be transmitted. This includes, but is not limited to, further dividing or aggregating multiple intermediate arrays based on the data size of each parameter to be allocated and each intermediate array.
[0092] Based on the above, by segmenting or aggregating intermediate arrays, intermediate information is divided into intermediate sub-information corresponding to the bandwidth parameters of each target virtual network card. This allows each intermediate sub-information to start and end transmission at the same time, maximizing the utilization of network resources and reducing the storage time of intermediate information.
[0093] As an optional implementation, multiple intermediate arrays can be split or aggregated according to each parameter to be assigned and each array parameter in the following manner: polling each intermediate array in descending order of array parameters; assigning each intermediate array to multiple target virtual network interfaces and updating the parameters to be assigned of the target virtual network interfaces.
[0094] Optionally, in this embodiment, but not limited to, polling each intermediate array in descending order of array parameters, and matching the corresponding intermediate sub-information for the target virtual network card in descending order of the initial parameters to be assigned. This approach effectively reduces the splitting operation of the intermediate arrays, facilitating the recovery of intermediate information later.
[0095] Optionally, in this embodiment, the parameters to be allocated for the target virtual network interface can be updated after each allocation of the intermediate array's split / aggregate to the target virtual network interface, so that subsequent allocations of the intermediate array will not result in errors.
[0096] As an optional implementation, the intermediate arrays can be assigned to multiple target virtual network interfaces in the following ways, but not limited to: when the array parameter of the reference intermediate array is greater than the assignment parameter of the reference virtual network interface, the reference intermediate array is split to obtain a first subarray and a second subarray, wherein the array parameter of the first subarray is equal to the assignment parameter of the reference virtual network interface, the multiple intermediate arrays include the reference intermediate array, and the multiple target virtual network interfaces include the reference virtual network interface; the first subarray is assigned to the reference virtual network interface, and the second subarray is assigned to other target virtual network interfaces besides the reference virtual network interface; when a candidate virtual network interface has been assigned information from at least two intermediate arrays, the information assigned to the candidate virtual network interface is aggregated to obtain intermediate sub-information corresponding to the candidate virtual network interface, wherein the multiple target virtual network interfaces include the candidate virtual network interface.
[0097] Optionally, in this embodiment, in the aforementioned training server integrated with Kubernetes, after the training container group starts, the Checkpoint migration service can, but is not limited to, determine the number of currently allocated VFs (i.e., target virtual network interfaces) and their corresponding IP addresses by reading the network status injected by the SR-IOV CNI plugin, such as ` / sys / class / net / ` (a directory for representing and managing network devices and their related attributes), `ip link` (a subcommand for managing and configuring link layer attributes of network devices such as network interface cards), or ` / sys / class / net / ` (a directory for representing and managing network devices and their related attributes), `ip link` (a subcommand for managing and configuring link layer attributes of network devices such as network interface cards), or the network status injected by the SR-IOV CNI plugin. It is important to note that during this process, primary service network interfaces such as `lo` (Loopback Interface) and `eth0` (Ethernet Interface) need to be filtered out, retaining only the VFs allocated by SR-IOV as candidate storage transmission channels. The Checkpoint file structure of the training framework can, but is not limited to, be parsed to extract metadata such as tensor name, shape, data type, and size. All tensor entries (i.e., intermediate arrays) can, but is not limited to, be sorted in descending order of data size and aggregated and split to distribute the total data as evenly as possible across each virtual network interface. Specifically, for cases where a single tensor is large, it is possible, but not limited to, splitting the tensor body by dimension; for cases where tensors are small, multiple tensors can be merged to form a data shard. After this, the following mapping relationships can be constructed, but are not limited to:
[0098] Fragment (i.e., intermediate sub-information) A→VF (i.e., target virtual network interface card)-0→NFS (i.e., storage service)-A;
[0099] Fragment B→VF-1→NFS-B;
[0100] Fragmentation C→VF-2→NFS-C.
[0101] It records the offset, path, and target storage for each fragment so that subsequent migration services (i.e., migration containers) can schedule it. Each fragment can be saved as a raw binary file or packaged using a uniform container format (such as zip (a compressed file format), npz (a file format for storing arrays), or tar (an archive file format)) to unify the data format.
[0102] As an optional implementation, the corresponding intermediate sub-information can be stored in parallel to the storage server through the target virtual network interface card (NIC) in the following manner: according to the correspondence between each target virtual NIC and each intermediate sub-information, a corresponding information transmission channel is matched for each intermediate sub-information, wherein the information transmission channel includes the target virtual NIC with the corresponding relationship and the storage directory of the target model training task; the corresponding intermediate sub-information is transmitted in parallel to the storage server through each information transmission channel.
[0103] Optionally, in this embodiment, since the information transmission channel includes target virtual network cards and storage directories with corresponding relationships, it is possible, but not limited to, matching the corresponding target virtual network card for each intermediate sub-information from the target virtual network cards and intermediate sub-information with corresponding relationships, then determining the information transmission channel where the target virtual network card is located as the information transmission channel corresponding to the intermediate sub-information, and transmitting the corresponding intermediate sub-information to the storage server through the information transmission channel.
[0104] Optionally, in this embodiment, to ensure the normal operation of parallel transmission, multiple information transmission channels can be created using the aforementioned method before matching the corresponding information transmission channel for each intermediate sub-information based on the correspondence between each target virtual network interface and each intermediate sub-information. The creation of the information transmission channel can be performed, but is not limited to, as described above, together with the operation of allocating target virtual network interfaces for the target model training task, i.e., between the segmentation of intermediate information, or it can be performed, but is not limited to, after the segmentation of intermediate information and before matching the corresponding information transmission channel for each intermediate sub-information.
[0105] Optionally, in this embodiment, in the aforementioned training server combined with Kubernetes, the Checkpoint migration service (i.e., migration container) can, but is not limited to, bind each tensor shard (i.e., intermediate sub-information) to a specific storage path and network card channel (i.e., information transmission channel) according to the data sharding logic, to achieve channel-level concurrent writing and reading.
[0106] Through the above methods, the parallel transmission of various intermediate sub-information effectively accelerates the process of saving intermediate information and saves the time required to save intermediate information.
[0107] As an optional implementation, after matching the corresponding information transmission channel for each intermediate sub-information based on the correspondence between each target virtual network interface and each intermediate sub-information, a global index file can be generated in the following ways, but not limited to: determining the transmission attributes of each intermediate sub-information based on the information transmission channels configured for multiple intermediate sub-information; recording the security attributes, content attributes, and transmission attributes of each intermediate sub-information to obtain a sub-index file for each intermediate sub-information, wherein the security attribute is used to indicate the integrity of the intermediate sub-information, the content attribute is used to indicate the intermediate array corresponding to the intermediate sub-information, and the intermediate information includes multiple intermediate arrays; summarizing the sub-index files of each intermediate sub-information and recording the concatenation attribute of the intermediate information to obtain a global index file for the training state, wherein the concatenation attribute is used to indicate the order relationship between multiple intermediate arrays in the intermediate information.
[0108] Optionally, in this embodiment, the transmission attribute may be used, but is not limited to, to indicate the information transmission channel corresponding to the intermediate sub-information.
[0109] Optionally, in this embodiment, after segmenting the intermediate information into multiple intermediate sub-information and matching corresponding information transmission channels for the intermediate sub-information, a global index file for the intermediate information can be recorded. The global index file can be used, but is not limited to, to guide the parallel storage of each intermediate sub-information and to guide the recovery operation of the intermediate information.
[0110] Optionally, in this embodiment, the content attribute may, but is not limited to, indicate the intermediate arrays included in each intermediate sub-information. The concatenation attribute may, but is not limited to, indicate the order relationship between multiple intermediate arrays in the intermediate information. Since some models have strict requirements on the order of intermediate arrays, it may, but is not limited to, need to record the concatenation attribute to ensure that the original intermediate information can be recovered to meet the needs of model training.
[0111] Optionally, in this embodiment, in the aforementioned training server integrated with Kubernetes, a global index file checkpoint.manifest.json can be generated, but is not limited to. This global index file can include, but is not limited to, the content hash (i.e., security attribute) of each shard, the corresponding list of tensors (i.e., content attribute), the target path, and the transmission channel (i.e., transmission attribute). For example:
[0112] {"tensor_name (name of the intermediate array) (i.e., content attribute)":"model.layer.5.weight",
[0113] "shard_index (the number in the intermediate array) (i.e., the content attribute)": 3,
[0114] "file_name (intermediate sub-information file name):"layer.5.weight.3.part",
[0115] "file_size (intermediate sub-information file size)": 64_000_000,
[0116] "hash (hash value) (i.e., security attribute)":"XXXXX",
[0117] "target_mount (mount directory) (i.e., transfer attribute)":" / mnt / b",
[0118] "vf_interface (target virtual network interface identifier) (i.e., transport attributes)":"vf1",
[0119] "expected_path (full path, including mount directory and intermediate sub-file name) (i.e., transfer attributes)":" / mnt / b / checkpoints / layer.5.weight.3.part"}
[0120] Optionally, in this embodiment, the parallel storage of intermediate sub-information can be completed according to the guidance of the global index file, but is not limited to. For example, in the aforementioned training server combined with Kubernetes, the migration service (i.e., the migration container) can be triggered when the training task (i.e., the training container) calls save_checkpoint (i.e., the save command). It can, but is not limited to, first reading the aforementioned generated global index file (such as checkpoint.manifest.json) to identify the path, size, target mount point, and target virtual network interface of each tensor fragment (i.e., intermediate sub-information). A mapping table from the target virtual network interface to the storage path can be established, but is not limited to...
[0121] For example: {"vf0":" / mnt / a",
[0122] "vf1":" / mnt / b",
[0123] "vf2":" / mntc"}
[0124] Ensure that each target virtual network interface card only accesses the storage service corresponding to its subnet.
[0125] It is possible, but not limited to, organizing all shards to be written (i.e., intermediate sub-information) into a task queue, recording the shard data path, target storage mount path, and corresponding VF interface identifier (i.e., information transmission channel) for each task. Then, start a parallel write thread or coroutine pool, binding a write thread or coroutine to each VF channel (i.e., information transmission channel). Each worker (write thread) independently processes its own shard queue, performing the following operations:
[0126] Open the file descriptor;
[0127] Perform asynchronous / direct write;
[0128] Once completed, mark the fragment write status (success / failure).
[0129] It can, but is not limited to, update the status of each successfully written shard and generate a corresponding .done file (a file used to mark that a task or operation has been completed) or a status identifier (used for subsequent verification / retry mechanisms). After all shards have been written, the migration service updates or generates the checkpoint.manifest.done file (a file used to mark that the saving operation of intermediate information has been completed), recording integrity verification information (such as the hash of each shard, write path, timestamp, etc.).
[0130] Optionally, in this embodiment, the intermediate information saving process can be decoupled from the main training process, but not limited to this. The migration service (i.e., the migration container) can run as an asynchronous thread or a sidecar process, but not limited to this. The main training process executed by the training container can continue to the next round of training after the save command is triggered, without being blocked by IO (Input / Output) operations. The main training process can check the migration service status before the next iteration through event polling, callback hooks, or checkpoint file locking mechanisms to ensure data consistency.
[0131] Figure 6 This is a schematic diagram of an intermediate information storage method according to an embodiment of this application, such as... Figure 6 As shown, the intermediate information, i.e., the initial checkpoint, can be divided into three intermediate sub-information, namely checkpoint_1, checkpoint_2, and checkpoint_3, in the manner described above, but not limited to the method described above. Then, through the storage network card (i.e., the target virtual network card) allocated to the training container group and the corresponding storage service, each intermediate sub-information is transmitted in parallel to the storage server via RDMA (Remote Direct Memory Access) network.
[0132] As an optional implementation, after aggregating the sub-index files of various intermediate sub-information and recording the splicing attributes of the intermediate information, and after transmitting the corresponding intermediate sub-information to the storage server in parallel through various information transmission channels, the intermediate information can be recovered in the following ways, but not limited to: receiving a recovery instruction, wherein the recovery instruction is used to instruct the target model training task to be restored to the training state; responding to the recovery instruction, searching for the global index file corresponding to the training state in the task state and index files with corresponding relationships; extracting multiple intermediate sub-information from the storage server according to the global index file; and recovering the intermediate information according to the content attributes, splicing attributes, and multiple intermediate sub-information.
[0133] Optionally, in this embodiment, the recovery instruction can be received and the recovery process of the intermediate information can be started after the intermediate information is stored on the storage server in the form of multiple intermediate sub-information, but not limited to this.
[0134] Optionally, in this embodiment, the corresponding global index file can be found based on the training state to be restored, and intermediate information can be restored with the help of the global index file.
[0135] Based on the above, the intermediate information can still be recovered normally after being stored in multiple intermediate sub-information formats, ensuring the normal use of the intermediate information and reducing the adverse impact on the functionality of the intermediate information caused by changes in the storage control method.
[0136] As an optional implementation, multiple intermediate sub-information can be extracted from the storage server based on the global index file in the following ways, but not limited to: traversing each sub-index file included in the global index file to determine the multiple intermediate sub-information to be extracted; and extracting each intermediate sub-information from the storage server based on transmission attributes and security attributes.
[0137] Optionally, in this embodiment, it is possible, but not limited to, first determining all the intermediate sub-information to be extracted, and then extracting each intermediate sub-information from the storage server according to the transmission attributes and security attributes; or, but not limited to, extracting each intermediate sub-information as soon as it is determined, repeating the above operation until all the intermediate sub-information to be extracted is obtained.
[0138] As an optional implementation, intermediate sub-information can be extracted from the storage server based on transmission attributes and security attributes in the following ways, but not limited to: extracting intermediate sub-information from the storage directory corresponding to each target virtual network interface card (NIC) through the target NIC in the information transmission channel configured for each intermediate sub-information; verifying each extracted intermediate sub-information based on each security attribute; storing the intermediate sub-information on the training server if the intermediate sub-information verification is successful; and re-extracting the intermediate sub-information if the intermediate sub-information verification fails.
[0139] Optionally, in this embodiment, it is possible, but not limited to, extracting only the intermediate sub-information that has passed the security attribute verification. For intermediate sub-information that fails verification for the first time or within the limit number of times, it is extracted again. For intermediate sub-information that fails verification after reaching or exceeding the limit number of times, the extraction of the intermediate sub-information is skipped and an alarm is triggered.
[0140] Through the above, by verifying security attributes, we avoid extracting abnormal intermediate sub-information, avoid using abnormal intermediate sub-information to recover abnormal intermediate information, and thus avoid the adverse effects of abnormal intermediate information on the target model training task.
[0141] As an optional implementation, intermediate information can be recovered based on content attributes, splicing attributes, and multiple intermediate sub-information in the following ways: recovering each intermediate array based on each content attribute and each intermediate sub-information; recovering intermediate information based on splicing attributes and each intermediate array.
[0142] Optionally, in this embodiment, in the aforementioned Kubernetes-integrated training server, when the migration service (i.e., the migration container) starts the recovery process, it may, but is not limited to, first parse the checkpoint.manifest.json file (i.e., the global index file) generated during the saving phase, extracting all tensor shard paths, corresponding storage paths and mount points, mapping relationships of corresponding VF channels (i.e., transfer attributes), and shard hashes and size verification information (i.e., security attributes). It may, but is not limited to, establish an access routing table with the following structure for subsequent download scheduling:
[0143] {" / mnt / a":"vf0",
[0144] " / mnt / b":"vf1",
[0145] " / mnt / c":"vf2"}
[0146] It is possible, but not limited to, constructing a fetch task queue, organizing all tensor fragments to be loaded into a task queue, grouping them by mount path, and recording: fragment filename, channel, target local cache path, and expected verification hash value. Bind an independent download thread or coroutine pool to each VF channel to implement concurrent fetch logic for fragments: using mechanisms such as open() (open a file) + read() (read data from a file), or aio_read (a system call for asynchronous file reading), sendfile (a file transfer system call); if it is a remote mount path (such as NFS), concurrent IO or fadvise (a system call that provides suggestions to the operating system about file access modes) can be used to optimize caching behavior; it is possible, but not limited to, writing fragments (i.e., intermediate sub-information) to a local temporary directory and marking them as complete upon completion. It is important to note that the migration service performs hash verification on each completed fragment; if the verification fails, a retry is automatically triggered, with a configurable maximum number of retries (e.g., 3 times). It is possible, but not limited to, merging all the retrieved fragment files into a specified directory according to their index order (i.e., concatenation attributes and content attributes) to construct a complete Checkpoint file (i.e., intermediate information) that can be recognized by the training framework. In other words, it constructs a Checkpoint file in a container format that can be recognized by each training framework. The main training process loads the reconstructed Checkpoint data through a Hook or explicit call (such as load_checkpoint()) to complete the restoration of the model training task state.
[0147] Figure 7 This is a schematic diagram of an intermediate information recovery method according to an embodiment of this application, such as... Figure 7 As shown, but not limited to, the three intermediate sub-information, namely checkpoint_1, checkpoint_2 and checkpoint_3, can be extracted into the training container group by using the various storage services and the allocated storage network card (i.e. target virtual network card) through the RDMA network, and then merged to obtain the initial checkpoint (i.e. intermediate information).
[0148] The aforementioned method significantly improves the efficiency of saving and loading checkpoints in large-scale model training tasks by introducing a parallel transmission mechanism based on SR-IOV multi-NICs and a checkpoint fragmentation and migration strategy. Compared with traditional single-channel serial storage solutions, this application can fully utilize multiple independent network channels between training nodes and storage servers to achieve high-bandwidth, low-latency concurrent data transmission. Simultaneously, by decoupling the main training process through tensor-level intelligent fragmentation and migration services, it effectively reduces blocking time during model saving and recovery, improving the stability and throughput of training tasks. Without relying on high-cost technologies such as RDMA, this application solves the storage bottleneck problem in large-scale model training in a more easily deployable manner, demonstrating strong versatility and engineering application value.
[0149] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0150] Embodiments of this application also provide a control device for model training. Figure 8 This is a structural block diagram of a model training control device according to an embodiment of this application, such as... Figure 8 As shown, the device includes:
[0151] The acquisition module 802 is used to acquire intermediate information of the target model training task executed on the training server. The training server is used to execute multiple model training tasks, including the target model training task. The training server is connected to the storage server through multiple training physical network cards. The training physical network cards are expanded into multiple virtual network cards. The model training task is assigned virtual network cards on at least two training physical network cards. The intermediate information is used to indicate the training status of the target model training task during the execution process.
[0152] The segmentation module 804 is used to segment intermediate information based on the virtual network card information of the target virtual network card assigned to the target model training task, and obtain multiple sets of target virtual network cards and intermediate sub-information with corresponding relationships. Among them, the virtual network card information is used to indicate the network status of the information transmission network provided by the target virtual network card.
[0153] Storage module 806 is used to store the corresponding intermediate sub-information in parallel to the storage server through the target virtual network card.
[0154] The above device acquires intermediate information of the target model training task executed on the training server. This intermediate information indicates the training status of the target model training task during execution. The intermediate information is segmented based on the network conditions of the information transmission network provided by the target virtual network card, as indicated by the virtual network card information of the target virtual network card allocated to the target model training task. This results in multiple sets of corresponding target virtual network cards and intermediate sub-information. The corresponding intermediate sub-information is then stored in parallel to the storage server via the target virtual network card. The training physical network card on the training server is expanded into multiple virtual network cards. The model training task executed on the training server, including the target model training task, is allocated virtual network cards on at least two training physical network cards. This allows the target model training task to utilize the network resources of at least two training physical network cards to store intermediate information in parallel, thus making more efficient use of the network transmission resources on the training server. Therefore, this addresses the technical problem of low efficiency in saving intermediate information during model training in related technologies, achieving the technical effect of improving the efficiency of saving intermediate information during model training.
[0155] In some embodiments, the segmentation module includes: a first extraction unit, configured to extract multiple intermediate arrays included in the intermediate information from the intermediate information and determine the array parameters of the intermediate arrays, wherein the array parameters are used to indicate the data size of the intermediate arrays; and a segmentation unit, configured to segment or aggregate multiple intermediate arrays according to the bandwidth parameters of each target virtual network interface and each array parameter to obtain multiple sets of target virtual network interfaces and intermediate sub-information with corresponding relationships, wherein the bandwidth parameters are used to indicate the bandwidth allowed to be used by the target virtual network interface, and the virtual network interface information includes the bandwidth parameters.
[0156] In some embodiments, the segmentation unit is further configured to: calculate the proportion of the bandwidth parameter of each target virtual network interface to the total bandwidth parameter, to obtain the target proportion of each target virtual network interface, wherein the total bandwidth parameter is used to indicate the total allowed bandwidth of multiple target virtual network interfaces; calculate the parameters to be allocated for each target virtual network interface based on the target proportion of each target virtual network interface and the total array parameter of multiple intermediate arrays, wherein the total array parameter is used to indicate the total data size of multiple intermediate arrays, and the parameters to be allocated are used to indicate the amount of data to be allocated to the target virtual network interface; and segment or aggregate multiple intermediate arrays based on each parameter to be allocated and each array parameter to obtain multiple sets of target virtual network interfaces and intermediate sub-information with corresponding relationships.
[0157] In some embodiments, the segmentation unit is further configured to: poll each intermediate array in descending order of array parameters; allocate each intermediate array to multiple target virtual network interfaces, and update the allocation parameters of the target virtual network interfaces.
[0158] In some embodiments, the segmentation unit is further configured to: split the reference intermediate array to obtain a first subarray and a second subarray when the array parameter of the reference intermediate array is greater than the allocation parameter of the reference virtual network interface card, wherein the array parameter of the first subarray is equal to the allocation parameter of the reference virtual network interface card, the multiple intermediate arrays include the reference intermediate array, and the multiple target virtual network interfaces include the reference virtual network interface card; allocate the first subarray to the reference virtual network interface card, and allocate the second subarray to other target virtual network interfaces besides the reference virtual network interface card; and aggregate the information allocated to the candidate virtual network interface card to obtain intermediate sub-information corresponding to the candidate virtual network interface card when the candidate virtual network interface card has been allocated information from at least two intermediate arrays, wherein the multiple target virtual network interfaces include the candidate virtual network interface card.
[0159] In some embodiments, the aforementioned apparatus further includes: an expansion module, configured to expand each training physical network interface card (NIC) into multiple virtual NICs based on the number of tasks of multiple model training tasks executed on the training server; and an allocation module, configured to allocate virtual NICs on at least two training physical NICs to a target model training task based on task information of the target model training task, wherein the task information is used to indicate the complexity of the target model training task.
[0160] In some embodiments, the expansion module includes: a first detection unit for detecting the number of tasks; an expansion unit for expanding each training physical network interface card (NIC) into multiple virtual NICs, wherein the number of virtual NICs obtained by expanding each training physical NIC is equal to the number of tasks; and a determination unit for determining the address information of the multiple virtual NICs obtained by expanding each training physical NIC based on the address information of each storage physical NIC on the storage server, wherein the address information of the multiple virtual NICs obtained by expanding each training physical NIC corresponds one-to-one with the address information of each storage physical NIC, and the address information is used to indicate the subnet where the NIC device is located. The NIC device includes multiple storage physical NICs and multiple virtual NICs obtained by expanding multiple training physical NICs, and the multiple storage physical NICs are located in different subnets.
[0161] In some embodiments, the allocation module includes: a first matching unit, configured to match a corresponding number of target network cards to the target task parameters from the corresponding task parameters and the number of network cards, wherein the target task parameters are the number of model parameters to be trained in the target model training task, and the task information includes the task parameters; a filtering unit, configured to filter out the target number of training physical network cards from multiple training physical network cards to obtain multiple target training network cards; and an allocation unit, configured to allocate one virtual network card from each target training network card to the target model training task to obtain multiple target virtual network cards corresponding to the target model training task.
[0162] In some embodiments, the storage module includes: a second matching unit, configured to match a corresponding information transmission channel for each of the intermediate sub-information based on the correspondence between each of the target virtual network interface cards and each of the intermediate sub-information, wherein the information transmission channel includes the target virtual network interface cards with a corresponding relationship and the storage directory of the target model training task; and a transmission unit, configured to transmit the corresponding intermediate sub-information to the storage server in parallel through each of the information transmission channels.
[0163] In some embodiments, the storage module further includes: a mounting unit, configured to: before matching the corresponding information transmission channel for each of the intermediate sub-information based on the correspondence between each of the target virtual network cards and each of the intermediate sub-information, match the corresponding storage service for each of the target virtual network cards based on the address information of each of the target virtual network cards, thereby obtaining multiple sets of target virtual network cards and storage services with corresponding relationships, wherein each storage service is used to listen to the subnet where multiple storage physical network cards on the storage server are located; and mount each storage service to multiple storage directories of the target model training task based on the multiple sets of target virtual network cards and storage services with corresponding relationships, thereby obtaining multiple information transmission channels.
[0164] In some embodiments, the mounting unit is further configured to: detect a startup command for the training container group, wherein the training container group is used to execute a target model training task, and the startup command is used to start the training container group; upon detecting a startup command, mount each storage service to each storage directory to obtain multiple sets of corresponding storage services and storage directories, wherein the storage directories are used to store data generated during the operation of the training container group; and determine the storage directory corresponding to each target virtual network interface card based on the multiple sets of corresponding target virtual network interface cards and storage services, and the multiple sets of corresponding storage services and storage directories, thereby obtaining multiple information transmission channels.
[0165] In some embodiments, the storage module further includes: a configuration unit, configured to, after matching the corresponding information transmission channel for each of the intermediate sub-information according to the correspondence between each of the target virtual network interface cards and each of the intermediate sub-information, determine the transmission attributes of each intermediate sub-information according to the information transmission channels configured for the multiple intermediate sub-information; record the security attributes, content attributes, and transmission attributes of each intermediate sub-information to obtain a sub-index file for each intermediate sub-information, wherein the security attributes are used to indicate the integrity of the intermediate sub-information, the content attributes are used to indicate the intermediate array corresponding to the intermediate sub-information, and the intermediate information includes multiple intermediate arrays; summarize the sub-index files of each intermediate sub-information and record the concatenation attributes of the intermediate information to obtain a global index file for the training state, wherein the concatenation attributes are used to indicate the order relationship between the multiple intermediate arrays in the intermediate information.
[0166] In some embodiments, the storage module further includes: a receiving unit, configured to receive a recovery instruction after summarizing the sub-index files of various intermediate sub-information, recording the splicing attributes of the intermediate information, and after transmitting the corresponding intermediate sub-information in parallel to the storage server through the various information transmission channels, wherein the recovery instruction is used to instruct the target model training task to be restored to the training state; a searching unit, configured to respond to the recovery instruction and search for the global index file corresponding to the training state in the task state and index files with corresponding relationships; a second extraction unit, configured to extract multiple intermediate sub-information from the storage server according to the global index file; and a recovery unit, configured to recover the intermediate information according to the content attributes, splicing attributes, and multiple intermediate sub-information.
[0167] In some embodiments, the second extraction unit is further configured to: traverse the various sub-index files included in the global index file to determine multiple intermediate sub-information to be extracted; and extract each intermediate sub-information from the storage server according to the transmission attributes and security attributes.
[0168] In some embodiments, the second extraction unit is further configured to: extract intermediate sub-information from the storage directory corresponding to each target virtual network card through the target virtual network card in the information transmission channel configured for each intermediate sub-information; verify each extracted intermediate sub-information according to each security attribute; store the intermediate sub-information on the training server if the intermediate sub-information verification is successful; and re-extract the intermediate sub-information if the intermediate sub-information verification fails.
[0169] In some embodiments, the recovery unit is further configured to: recover each intermediate array according to each content attribute and each intermediate sub-information; and recover intermediate information according to the splicing attribute and each intermediate array.
[0170] In some embodiments, the acquisition module includes: a second detection unit, configured to detect instructions called by the training container through a transfer container, wherein the training container group includes a training container and a transfer container, the training container group is used to execute a target model training task, the training container is used to execute the training process of the target model training task, and the transfer container is used to control intermediate information; and an acquisition unit, configured to acquire intermediate information from a target storage location in the training container group through a transfer container when a save instruction is detected by the training container, wherein the target storage location is used to provide storage space for the training container.
[0171] For a description of the features in the embodiment corresponding to the control device for model training, please refer to the relevant description in the embodiment corresponding to the control method for model training, which will not be repeated here.
[0172] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-described control method embodiments for model training.
[0173] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described control method embodiments for model training at runtime.
[0174] In some embodiments, the computer-readable storage medium described above may be, but is not limited to, a non-volatile computer-readable storage medium.
[0175] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0176] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described control method embodiments for model training.
[0177] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described control method embodiments for model training.
[0178] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0179] The foregoing has provided a detailed description of the control method, program product, electronic device, and medium for model training provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to aid in understanding the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A control method for model training, characterized in that, include: The intermediate information of the target model training task executed on the training server is obtained. The training server is used to execute multiple model training tasks, including the target model training task. The training server is connected to the storage server through multiple training physical network cards. The training physical network cards are expanded into multiple virtual network cards. The model training task is assigned to at least two of the virtual network cards on the training physical network cards. The intermediate information is used to indicate the training status of the target model training task during execution. The intermediate information is segmented based on the virtual network card information of the target virtual network card assigned in the target model training task to obtain multiple sets of target virtual network cards and intermediate sub-information with corresponding relationships. The virtual network card information is used to indicate the network status of the information transmission network provided by the target virtual network card. The corresponding intermediate sub-information is stored in parallel to the storage server through the target virtual network interface card; The step of segmenting the intermediate information based on the virtual network card information of the target virtual network card assigned by the target model training task to obtain multiple sets of corresponding target virtual network cards and intermediate sub-information includes: extracting multiple intermediate arrays included in the intermediate information from the intermediate information and determining the array parameters of the intermediate arrays, wherein the array parameters are used to indicate the data size of the intermediate arrays; segmenting or aggregating multiple intermediate arrays based on the bandwidth parameters of each target virtual network card and each array parameter to obtain multiple sets of corresponding target virtual network cards and intermediate sub-information, wherein the bandwidth parameters are used to indicate the bandwidth allowed to be used by the target virtual network card, and the virtual network card information includes the bandwidth parameters; The step of dividing or aggregating multiple intermediate arrays based on the bandwidth parameters and array parameters of each target virtual network interface card (NIC) to obtain multiple sets of corresponding target virtual NICs and intermediate sub-information includes: calculating the proportion of the bandwidth parameter of each target virtual NIC to the total bandwidth parameter to obtain the target proportion of each target virtual NIC, wherein the total bandwidth parameter is used to indicate the total allowed bandwidth of the multiple target virtual NICs; calculating the unallocated parameters of each target virtual NIC based on the target proportion of each target virtual NIC and the total array parameter of the multiple intermediate arrays, wherein the total array parameter is used to indicate the total data size of the multiple intermediate arrays, and the unallocated parameters are used to indicate the amount of data to be allocated to the target virtual NIC; and dividing or aggregating multiple intermediate arrays based on the unallocated parameters and array parameters to obtain multiple sets of corresponding target virtual NICs and intermediate sub-information.
2. The method according to claim 1, characterized in that, The step of splitting or aggregating multiple intermediate arrays according to each of the parameters to be assigned and each of the array parameters includes: The intermediate arrays are iterated in descending order of the array parameters. Each of the intermediate arrays is assigned to one of the target virtual network interfaces, and the parameters to be assigned of the target virtual network interfaces are updated.
3. The method according to claim 2, characterized in that, The step of assigning each of the intermediate arrays to the multiple target virtual network interfaces includes: If the array parameter of the reference intermediate array is greater than the parameter to be assigned of the reference virtual network interface card (NIC), the reference intermediate array is split to obtain a first subarray and a second subarray, wherein the array parameter of the first subarray is equal to the parameter to be assigned of the reference virtual network interface card (NIC), a plurality of intermediate arrays include the reference intermediate array, and a plurality of target virtual network interfaces include the reference virtual network interface card (NIC); the first subarray is assigned to the reference virtual network interface card (NIC), and the second subarray is assigned to other target virtual network interfaces (NICs) other than the reference virtual network interface card (NIC). When a candidate virtual network interface card (NIC) is assigned information from at least two of the intermediate arrays, the information assigned to the candidate virtual NIC is aggregated to obtain the intermediate sub-information corresponding to the candidate virtual NIC, wherein the plurality of target virtual NICs include the candidate virtual NIC.
4. The method according to claim 1, characterized in that, Before segmenting the intermediate information based on the virtual network interface card information of the target virtual network interface card assigned by the target model training task, the method further includes: The training physical network interface card is expanded into multiple virtual network interface cards based on the number of tasks of the multiple model training tasks executed on the training server. Based on the task information of the target model training task, at least two virtual network cards on the training physical network cards are allocated to the target model training task, wherein the task information is used to indicate the complexity of the target model training task.
5. The method according to claim 4, characterized in that, Expanding each training physical network interface card (NIC) into multiple virtual NICs based on the number of training tasks executed on the training server includes: Detect the number of tasks; Each of the training physical network cards is expanded into a plurality of virtual network cards, wherein the number of the plurality of virtual network cards obtained by expanding each of the training physical network cards is equal to the number of tasks; The address information of the multiple virtual network cards obtained by expanding each of the training physical network cards is determined based on the address information of each physical network card on the storage server. The address information of the multiple virtual network cards obtained by expanding each physical network card corresponds one-to-one with the address information of each physical network card. The address information is used to indicate the subnet where the network card device is located. The network card device includes multiple physical network cards and multiple virtual network cards obtained by expanding multiple physical network cards. The multiple physical network cards are in different subnets.
6. The method according to claim 4, characterized in that, The step of allocating at least two virtual network cards on the training physical network cards to the target model training task based on the task information of the target model training task includes: Match the target network card quantity to the target task parameter from the corresponding task parameters and network card quantity, wherein the target task parameter is the number of model parameters to be trained in the target model training task, and the task information includes the task parameter; The target number of training physical network cards is selected from the multiple training physical network cards to obtain multiple target training network cards; One of the virtual network cards in each of the target training network cards is assigned to the target model training task, thereby obtaining multiple target virtual network cards corresponding to the target model training task.
7. The method according to claim 1, characterized in that, The step of storing the corresponding intermediate sub-information in parallel to the storage server through the target virtual network interface card includes: Based on the correspondence between each target virtual network interface and each intermediate sub-information, a corresponding information transmission channel is matched for each intermediate sub-information, wherein the information transmission channel includes the target virtual network interface with the corresponding relationship and the storage directory of the target model training task; The corresponding intermediate sub-information is transmitted in parallel to the storage server through each of the information transmission channels.
8. The method according to claim 7, characterized in that, Before matching the corresponding information transmission channel for each of the intermediate sub-informations based on the correspondence between each of the target virtual network interface cards and each of the intermediate sub-informations, the method further includes: Based on the address information of each target virtual network card, a corresponding storage service is matched for each target virtual network card to obtain multiple sets of target virtual network cards and storage services with corresponding relationships. Each storage service is used to monitor the subnet where multiple storage physical network cards on the storage server are located. Based on multiple sets of corresponding target virtual network cards and the storage services, each of the storage services is mounted to multiple storage directories of the target model training task to obtain multiple information transmission channels.
9. The method according to claim 8, characterized in that, The step of mounting each of the storage services to multiple storage directories of the target model training task based on multiple sets of corresponding target virtual network interface cards and the storage services, thereby obtaining multiple information transmission channels, including: A startup command for a training container group is detected, wherein the training container group is used to execute the target model training task, and the startup command is used to start the training container group; Upon detecting the start command, each storage service is mounted to each storage directory, resulting in multiple sets of corresponding storage services and storage directories, wherein the storage directories are used to store data generated during the operation of the training container group; Based on multiple sets of corresponding target virtual network interface cards and storage services, and multiple sets of corresponding storage services and storage directories, the storage directory corresponding to each target virtual network interface card is determined, thereby obtaining multiple information transmission channels.
10. The method according to claim 7, characterized in that, After matching the corresponding information transmission channel for each of the intermediate sub-informations based on the correspondence between each of the target virtual network interface cards and each of the intermediate sub-informations, the method further includes: The transmission attributes of each of the intermediate sub-informations are determined based on the information transmission channels configured for the multiple intermediate sub-informations; Record the security attributes, content attributes, and transmission attributes of each intermediate sub-information to obtain a sub-index file for each intermediate sub-information, wherein the security attributes are used to indicate the integrity of the intermediate sub-information, the content attributes are used to indicate the intermediate array corresponding to the intermediate sub-information, and the intermediate information includes multiple intermediate arrays; The sub-index files of each intermediate sub-information are aggregated, and the concatenation attributes of the intermediate information are recorded to obtain the global index file of the training state, wherein the concatenation attributes are used to indicate the order relationship between multiple intermediate arrays in the intermediate information.
11. The method according to claim 10, characterized in that, After summarizing the sub-index file of each of the intermediate sub-information and recording the concatenation attributes of the intermediate information, and after transmitting the corresponding intermediate sub-information to the storage server in parallel through each of the information transmission channels, the method further includes: Receive a recovery instruction, wherein the recovery instruction is used to instruct the target model training task to be restored to the training state; In response to the recovery command, the global index file corresponding to the training state is searched in the corresponding task states and index files; Based on the global index file, extract multiple intermediate sub-information from the storage server; The intermediate information is restored based on the content attributes, the splicing attributes, and multiple intermediate sub-information.
12. The method according to claim 11, characterized in that, The step of extracting multiple intermediate sub-information from the storage server based on the global index file includes: Traverse each of the sub-index files included in the global index file to determine the multiple intermediate sub-information to be extracted; Each of the intermediate sub-information is extracted from the storage server based on the transmission attribute and the security attribute.
13. The method according to claim 12, characterized in that, The step of extracting each of the intermediate sub-information from the storage server based on the transmission attribute and the security attribute includes: The intermediate sub-information is extracted from the storage directory corresponding to each of the target virtual network cards in the information transmission channel configured by each of the intermediate sub-information. Each of the intermediate sub-information extracted based on the verification of each of the aforementioned security attributes; If the intermediate sub-information is successfully verified, the intermediate sub-information is stored on the training server. If the intermediate sub-information verification fails, the intermediate sub-information is extracted again.
14. The method according to claim 11, characterized in that, The step of restoring the intermediate information based on the content attributes, the splicing attributes, and multiple intermediate sub-information includes: Restore each intermediate array based on each of the content attributes and each of the intermediate sub-information; The intermediate information is recovered based on the splicing attributes and each of the intermediate arrays.
15. The method according to claim 1, characterized in that, The process of obtaining intermediate information about the target model training task executed on the training server includes: The instructions called by the training container are detected by the transfer container. The training container group includes the training container and the transfer container. The training container group is used to execute the target model training task. The training container is used to execute the training process of the target model training task. The transfer container is used to control the intermediate information. If a save instruction is detected in the training container, the intermediate information is obtained from the target storage location in the training container group through the transfer container, wherein the target storage location is used to provide storage space for the training container.
16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the control method for model training as described in any one of claims 1 to 15.
17. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the control method for model training as described in any one of claims 1 to 15.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the control method for model training as described in any one of claims 1 to 15.
Citation Information
Patent Citations
Container cloud distributed training data communication method based on optimal scheduling
CN110308986A
Task training method and device, equipment and storage medium
CN114647488A