Information processing method, device, equipment, storage medium and program product

By obtaining multiple threads and training sets of target tasks in deep learning training and matching threads for the acceleration chip, the alternating execution of logical operations is solved, and the problem of not being able to support elastic training while ensuring model accuracy is achieved in the existing technology, and the consistency of model quality with fixed resources is achieved.

CN114661474BActive Publication Date: 2025-05-16ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210334467.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-30
Publication Date
2025-05-16
Estimated Expiration
2042-03-30

AI Technical Summary

Technical Problem

The prior art cannot guarantee the accuracy of deep learning models while supporting elastic training.

Method used

By obtaining multiple threads and training sets of the target task, determine the acceleration chip that can provide elastic resources, and match threads for each acceleration chip to perform alternate logic operations until all threads have completed execution. The model parameters are updated.

Benefits of technology

In deep learning training tasks that support elastic resources, it is realized to maintain the consistency of model accuracy and model quality training when fixed resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114661474B_ABST
    Figure CN114661474B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides an information processing method, device, equipment, storage medium and program product, and the method includes: obtaining multiple threads and multiple training sets corresponding to a target task, the target task is executed by multiple threads, and the target task supports elastic training; determining at least one acceleration chip that provides elastic resources for the target task, and matching at least one thread for each acceleration chip according to multiple threads; for each training set, performing the following steps: for each acceleration chip, assigning a target thread to the acceleration chip, so that the current target thread performs a logical operation according to the training set, and after the current target thread completes the logical operation, switching the next target thread for the acceleration chip, until the multiple threads are all executed in at least one acceleration chip, to update the model parameters corresponding to the target task. It is possible to achieve the accuracy of the model while supporting elastic training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular, to an information processing method, apparatus, device, storage medium, and program product. Background Art

[0002] Deep neural networks have been used in many widely deployed systems across multiple fields, including computer vision, natural language processing, speech recognition, as well as recommendations and advertising. Therefore, deep learning has become a vital part of product data flow. In order to support such large-scale deep learning applications, large-scale shared acceleration chip clusters are generally built to perform multiple deep learning tasks.

[0003] However, it has been observed in these shared acceleration chip clusters that in fact, the acceleration chip resources are often still in a relatively low utilization state. At the same time, some tasks are still waiting in line, and the throughput of the entire cluster is not very high. In addition, resource sharing can also lead to task preemption. In order to solve the problem of long queue delays and failures caused by preemption in the above cluster tasks, adapting training tasks to elastic resources is a very direct method. After supporting elasticity, training tasks can start processing as soon as possible using available resources, eliminating the mandatory waiting caused by cluster scheduling, and using the remaining resources to continue training when resources are preempted, thereby improving cluster utilization and reducing task completion time. At present, elastic training methods often achieve model quality similar to synchronous distributed training with fixed resources by modifying synchronization methods, adjusting hyperparameters, and adjusting small batch sizes. However, it is impossible to completely achieve consistency with the accuracy when elastic resources are not used, that is, the model accuracy cannot be guaranteed.

[0004] Therefore, in the prior art, it is impossible to support elastic training while ensuring the accuracy of the model. Summary of the invention

[0005] The embodiments of the present application provide an information processing method, apparatus, device, storage medium and program product to solve the problem that the prior art cannot ensure the accuracy of the model while supporting elastic training.

[0006] In a first aspect, an embodiment of the present application provides an information processing method, the method comprising:

[0007] Acquire multiple threads and multiple training sets corresponding to a target task, wherein the target task is executed by multiple threads and the target task supports flexible training;

[0008] Determine at least one acceleration chip that provides elastic resources for the target task, and match at least one thread to each acceleration chip according to the multiple threads;

[0009] For each of the training sets, the following steps are performed: for each of the acceleration chips, a target thread is assigned to the acceleration chip so that the current target thread performs logical operations according to the training set, and after the current target thread completes the logical operations, the next target thread is switched to the acceleration chip until the multiple threads are all executed on the at least one acceleration chip, so as to update the model parameters corresponding to the target task.

[0010] Optionally, the acquiring of multiple threads corresponding to the target task includes:

[0011] Determining the number of processes used by the target task in distributed data parallelism;

[0012] Performing thread abstraction on the target task and determining the multiple threads, so that each acceleration chip supports the execution of any number of the threads during runtime;

[0013] The number of the multiple threads is the same as the number of the processes.

[0014] Optionally, matching at least one thread for each acceleration chip according to the multiple threads includes:

[0015] Determine the number of threads allocated to each acceleration chip according to the configuration information of each acceleration chip;

[0016] Based on the number of threads allocated to each acceleration chip, the multiple threads are allocated to corresponding acceleration chips, so that each acceleration chip provides resources for at least one thread;

[0017] Each of the acceleration chips supports one thread execution logic at the same time.

[0018] Optionally, after the current target thread completes executing the logic operation, switching the next target thread for the acceleration chip until all the multiple threads are executed in the at least one acceleration chip to update the model parameters corresponding to the target task includes:

[0019] After the current target thread completes executing logic, a random state and a corresponding gradient value of the current target thread are generated, the random state of the current target thread is stored as the initial state of the current target thread, and the gradient value corresponding to the current target thread is stored; wherein the random state is used to generate a random number sequence;

[0020] Obtaining an initial state of a next target thread, determining a random number and a random state of the next target thread corresponding to the random number according to the initial state of the next target thread, and storing the random state of the next target thread as the initial state of the next target thread; the random number is used to support the next target thread to perform a logic operation on the acceleration chip;

[0021] When all of the multiple threads are executed in the at least one acceleration chip, the model parameters corresponding to the target task are updated according to the gradient value corresponding to each of the threads.

[0022] Optionally, determining the random state of the next target thread according to the initial state of the next target thread includes:

[0023] Generate a random number and a random state of the next target thread corresponding to the random number through a random number generator according to the initial state of the next target thread;

[0024] The random state includes: a state of a deep learning framework for representing a Python interface, a state of an open source scientific computing library for representing Python, and a state of a random number generator.

[0025] Optionally, storing the random state of the current target thread as the initial state of the current target thread includes:

[0026] The initial state of the current target thread is stored in a preset shared pool of the at least one acceleration chip to support obtaining the initial state of the corresponding thread when elastic resources occur;

[0027] Correspondingly, storing the gradient value corresponding to the current target thread includes:

[0028] The gradient value corresponding to the current target thread is unloaded to the memory to support the fusion calculation of the gradient value corresponding to the next target thread.

[0029] Optionally, the determining of at least one acceleration chip that provides elastic resources for the target task includes:

[0030] When the elastic resource is triggered, a target resource that currently supports providing training for the target task is obtained, where the target resource is an acceleration chip, and the number of the acceleration chips is at least one.

[0031] In a second aspect, an embodiment of the present application provides an information processing device, the device comprising:

[0032] An acquisition module, used to acquire multiple threads and multiple training sets corresponding to a target task, wherein the target task is executed by multiple threads and the target task supports flexible training;

[0033] A first processing module is used to determine at least one acceleration chip that provides elastic resources for the target task, and match at least one thread to each acceleration chip according to the multiple threads;

[0034] The second processing module is used to perform the following steps for each of the training sets: for each of the acceleration chips, a target thread is allocated to the acceleration chip so that the current target thread performs logical operations according to the training set, and after the current target thread completes the logical operations, the next target thread is switched to the acceleration chip until the multiple threads are all executed on the at least one acceleration chip, so as to update the model parameters corresponding to the target task.

[0035] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;

[0036] The memory stores computer-executable instructions;

[0037] The processor executes the computer-executable instructions stored in the memory to implement the method as described in any one of the first aspects.

[0038] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the method described in any one of the first aspects is implemented.

[0039] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method described in any one of the first aspects.

[0040] The information processing method, device, equipment, storage medium and program product provided by the embodiment of the present application, the method performs elastic training tasks by acquiring target tasks and multiple training sets participating in training. When resource elasticity occurs, first determine the acceleration chip that can provide elastic resources, and allocate threads to the acceleration chip based on the training task, so that each acceleration chip only supports one thread execution logic at the same time. After the execution is completed, switch to the next thread execution logic to achieve alternating execution. After all threads are executed, update the model parameters once to achieve a round of training, and then repeat the above steps to perform the next round of training. Therefore, by alternating execution, elastic training can be achieved without adjusting or modifying the synchronization method, hyperparameters, small batch size, etc. At the same time, since there is no modification of data, the quality of the model is consistent with that of the existing fixed resources, that is, the model accuracy is guaranteed. Therefore, while supporting deep learning training tasks for elastic resources, the present application can improve the model accuracy of elastic training supported by the prior art, achieve consistency with the model quality trained with fixed resources, and ensure model accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0042] Figure 1 A schematic diagram of a scenario of an information processing method provided in an embodiment of the present application;

[0043] Figure 2 A flowchart of an information processing method provided in an embodiment of the present application;

[0044] Figure 3A Schematic diagram of application using GPU resources in existing technology Figure 1 ;

[0045] Figure 3B Schematic diagram of application using GPU resources in existing technology Figure 2 ;

[0046] Figure 3C A schematic diagram of an application using GPU resources provided in an embodiment of the present application;

[0047] Figure 4 A flowchart of an information processing method provided in yet another embodiment of the present application;

[0048] Figure 5 A schematic diagram of the structure of an information processing device provided in an embodiment of the present application;

[0049] Figure 6 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0050] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0051] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can also include other sequential instances in addition to those illustrated or described. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0052] At present, elastic training methods often achieve model quality similar to that of synchronous distributed training with fixed resources by modifying synchronization methods, adjusting hyperparameters, adjusting small batch sizes, etc. However, it is impossible to completely achieve consistency in accuracy with the case where elastic resources are not used, that is, the model accuracy cannot be guaranteed. Therefore, in the existing technology, it is impossible to improve the accuracy of the model supporting elastic training while supporting elastic training, and thus it is impossible to achieve consistency in model quality with that of the model trained with fixed resources, and the model accuracy cannot be guaranteed.

[0053] In order to solve the above problems, the invention concept of the present application is as follows: when elastic resources are triggered, first determine the acceleration chip that can provide resources, and based on the training task, allocate threads to the acceleration chip, so that each acceleration chip only supports one thread execution logic at the same time. When the execution is completed, switch to the next thread execution logic. When all threads are executed, update the model parameters once to implement one round of training, and then repeat the above steps to execute the next round of training, thereby realizing elastic training. Since the elastic training process does not require modification of the parameters when using fixed resources for training, the model accuracy of the existing technology supporting elastic training is improved, the model quality consistency with that of the model trained with fixed resources is achieved, and the model accuracy is guaranteed.

[0054] Figure 1 Scenario diagram of the information processing method provided for the embodiment of the present application. The user writes the program through the terminal device using the user programming interface, that is, the training is programmed in a class inherited from. The programming here is different from the prior art. It decouples the execution on the hardware from the high-level application programming interface, that is, there is no need to program the execution logic of the thread and the hardware (such as the acceleration chip, etc.). For example, thread 1 is assigned to execute on acceleration chip 1, thread 2 is assigned to execute on acceleration chip 2, etc.

[0055] At runtime, the server performs training operations based on the target task and training sample set input by the user. When elastic resources occur, the server abstracts the process corresponding to the target task into EasyScaleThread (i.e., a thread that supports elastic heterogeneity and lossless precision deep learning training framework, hereinafter referred to as thread) according to the available acceleration chip resources, and allocates threads to the available acceleration chips, and ensures that for each acceleration chip, at the same time, only one EasyScaleThread logic execution is arranged at runtime, and other executions are frozen. When the execution is completed, the next thread execution logic is switched to by context switching. When all threads are executed, the model parameters are updated once to achieve a round of training, and then the above steps are repeated to perform the next round of training until the training stops, and the training results are fed back to the terminal device. This elastic training process can achieve elastic training without modifying the synchronization method, hyperparameters, small batches (size, etc.). At the same time, since there is no modification of data, the model quality is consistent with that of the existing fixed resources, that is, the model accuracy is guaranteed. Therefore, this application can ensure model accuracy while supporting deep learning training tasks with elastic resources.

[0056] Among them, small batch is used to indicate: in deep learning scenarios, when executing the gradient descent algorithm on the training set, the entire data set is divided into several small training sets, and a small subset is trained each time. On the one hand, it avoids the huge amount of computation caused by the participation of the entire data set in training at one time; on the other hand, the gradient direction of the subset will not be too different from that of the entire data set, thus ensuring the correctness of the training.

[0057] The technical solution of the present application is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0058] Figure 2 This is a flow chart of an information processing method provided in an embodiment of the present application. The method of this embodiment can be executed by a server, and a cluster scheduler can be deployed in the server to implement resource scheduling. Figure 2 As shown, the method of this embodiment may include:

[0059] S201: Acquire multiple threads and multiple training sets corresponding to a target task.

[0060] The target task is executed by multiple threads, and the target task supports flexible training.

[0061] In this embodiment, when writing a target task, the user can consider the application logic (for example, the learning rate hyperparameter) based on the number of theoretical workers (that is, the number of EasyScaleThreads, which is consistent with the number of workers used in the native distributed data parallel method (Distributed Data Parallel, DDP) of PyTorch (deep learning computing framework). Therefore, the obtained target task can be executed by multiple threads.

[0062] S202: Determine at least one acceleration chip that provides elastic resources for the target task, and match at least one thread to each acceleration chip according to the multiple threads.

[0063] The acceleration chip here can be a GPU (Graphics Processing Unit, GPU) chip or an AI chip. The following embodiments all take GPU as an example to explain the information processing method in detail.

[0064] For example, at least one GPU that provides elastic resources for the target task is determined, and at least one thread is matched to each of the GPUs according to the multiple threads.

[0065] Elastic resources or resource elasticity are used to indicate that training tasks can use resources that change dynamically in quantity. For example, the number of available GPUs fluctuates between 1 and 100.

[0066] Optionally, determining at least one acceleration chip that provides elastic resources for the target task may include: when the elastic resources are triggered, obtaining the target resources that currently support providing training for the target task, the target resources are acceleration chips, and the number of the acceleration chips is at least one.

[0067] The triggering of resource elasticity here is in the scenario where the computing cluster is in a multi-tenant scenario, that is, multiple users submit multiple tasks on the same GPU cluster. When a user submits a task, computing resources will be occupied, causing changes in the number of resources.

[0068] In this embodiment, threads (i.e., EasyScaleThread) of a deep learning training framework that supports elastic heterogeneity and lossless precision can be allocated between heterogeneous GPUs to maximize GPU utilization. Here, EasyScaleThread: Here, users program applications and algorithm logic to execute in one thread, which is in line with the current programming model in the field of deep learning.

[0069] During operation, each acceleration chip can execute any number of EasyScaleThreads to achieve flexible training. Through such abstraction, training tasks (such as DLT (Deep Learning Training, DLT) jobs, in this embodiment, focusing on distributed training tasks) are no longer tightly coupled with the fixed parallelism of the GPU.

[0070] S203: For each of the training sets, perform the following steps: for each of the acceleration chips, assign a target thread to the acceleration chip so that the current target thread performs logical operations according to the training set, and after the current target thread completes the logical operations, switch the next target thread to the acceleration chip until all the multiple threads are executed on at least one acceleration chip, so as to update the model parameters corresponding to the target task.

[0071] In this embodiment, when elastic resources occur, threads are allocated to available GPUs so that the threads perform separate execution logic in an alternating manner. That is, each GPU only supports one thread to execute logic at the same time. When the execution is completed, the next thread is switched to execute logic. When all threads are executed, the model parameters are updated once to implement a round of training, and then the above steps are repeated to perform the next round of training until the training stops.

[0072] Exemplary, combined Figure 3A and Figure 3B The application schematic diagram of using GPU resources in the prior art is shown (wherein, Figure 3A A diagram showing the current GPU utilization method for fixed-parallelism data parallel training jobs for deep learning; Figure 3B A diagram showing an intuitive way to support resource elasticity), and Figure 3C The application schematic diagram of using GPU resources provided by the embodiment of the present application is shown, and the difference between the prior art and the embodiment of the present application can be determined.

[0073] Specifically, take a 4-worker (i.e., job, one job corresponds to one process) PyTorch (i.e., a deep learning framework) as an example. Each worker uses a specific GPU (a total of 4) for model training and gradient synchronization (see Figure 3AAs shown in the figure, 4-worker on fixed 4GPUs means that 4 workers use 4 GPUs. That is, PyTorchworker0 is assigned to execute on GPU0, PyTorch worker1 is assigned to execute on GPU1, PyTorch worker2 is assigned to execute on GPU2, and PyTorch worker3 is assigned to execute on GPU3.

[0074] Ideally, when resource elasticity occurs, for example, the number of available GPUs is reduced to 2, assuming that the GPUs have unlimited video memory and computing power, 4 PyTorch workers can still do lossless synchronous training, with each GPU taking on two processes (see Figure 3B As shown in the figure, 4-worker executing on 2ideal GPUs means that 4 workers are executed using 2 GPUs (i.e. 2ideal GPUs) under ideal circumstances. That is, PyTorch worker0 and PyTorch worker1 are assigned to execute on idealGPU0; PyTorch worker2 and PyTorch worker3 are assigned to execute on idealGPU1). In this case, the working set (the working set here refers to the working status of all workers, mainly including the process status corresponding to each worker) is consistent with when 4 GPUs are used, because the status of each process is consistent. However, the video memory and computing power of a GPU are not infinite. When running a DLT job on a GPU, at least the model weights and optimizer status must be saved on the GPU. By concurrently executing multiple workers of a DLT job on a single GPU, during their forward propagation process, concurrent video memory consumption can cause video memory exhaustion or generate a lot of overhead in swapping in and out of video memory and main memory. Even without considering the worker's working set, the memory overhead of each worker's CUDA (Compute Unified Device Architecture) context is still not negligible. For example, executing 32 workers on a V100 GPU requires 23GB of video memory for the CUDA context alone (about 750MB for each CUDA context), while the V100 has only 32GB of video memory in total. Therefore, it is possible to achieve model quality similar to that of synchronous distributed training with fixed resources (i.e., the loss of the training set / validation set) by modifying synchronization methods, adjusting hyperparameters, and adjusting the size of small batches. However, it is impossible to achieve complete bitwise consistency with the case where elastic resources are not used.

[0075] In order to improve the low accuracy problem caused by the existing elastic training, through the abstract concept EasyScaleThread, users program the application and algorithm logic to execute in one thread. When programming, users only need to consider the application logic based on the number of theoretical workers. At runtime, each GPU can execute any number of EasyScaleThreads, thus achieving elastic training. Through such abstraction, DLT jobs are no longer tightly coupled with the fixed parallelism of the GPU (see Figure 3C , PY can represent Python, and Python language is used for programming. PyTorch worker0 includes two threads EST_0 and EST_1, which are assigned to GPU0 and executed in sequence; PyTorch worker1 includes two threads EST_2 and EST_3, which are assigned to GPU1 and executed in sequence).

[0076] Specifically, the abstraction of EasyScaleThread decouples execution on the hardware from the high-level application programming interface. In addition, it minimizes the user's effort in the migration work and avoids introducing additional constraints that limit the user's programming capabilities. For example, users can still adjust hyperparameters and use third-party libraries. Users can program training in a class that inherits from EasyScaleThread. Therefore, it allows the runtime system to replicate it according to the number of workers assigned to the GPU, hook the execution flow, and access the state of the model and optimizer to minimize the context switching overhead, and through a series of logic, provide consistent accuracy results. In addition, the iteration steps of iterative training should be annotated in the user code, which is the key to the solution to perform consistent accuracy context switching at the boundary of mini-batches.

[0077] The information processing method provided by the present application performs elastic training tasks by acquiring target tasks and multiple training sets participating in training. When resource elasticity occurs, first determine the acceleration chip that can provide elastic resources, and allocate threads to the acceleration chip based on the training task, so that each acceleration chip only supports one thread execution logic at the same time. After the execution is completed, switch to the next thread execution logic to achieve alternating execution. After all threads are executed, update the model parameters once to achieve a round of training, and then repeat the above steps to perform the next round of training. Therefore, by alternating execution, elastic training can be achieved without adjusting or modifying the synchronization method, hyperparameters, mini-batch size, etc. At the same time, since there is no modification to the data, the quality of the model trained is consistent with that of the existing fixed resources, that is, the model accuracy is guaranteed. Therefore, while supporting deep learning training tasks with elastic resources, the present application can improve the model accuracy of elastic training supported by the prior art, achieve consistency with the model quality trained with fixed resources, and ensure model accuracy.

[0078] Optionally, to obtain multiple threads corresponding to the target task, you can do the following:

[0079] Step a1: Determine the number of processes used by the target task in distributed data parallelism.

[0080] Step a2: perform thread abstraction on the target task and determine the multiple threads so that each acceleration chip supports the execution of any number of threads during runtime.

[0081] The number of the multiple threads is the same as the number of the processes.

[0082] In this embodiment, through thread abstraction processing, when programming, the user only needs to consider the logic of the application based on the number of theoretical workers (i.e., the number of EasyScaleThreads, which is consistent with the number of workers (i.e., processes) used in DDP). Therefore, when training with fixed resources, for example, there are 4 processes, which will be abstracted as 4 threads here. Each GPU can support the execution of any number of threads, but each GPU allows one thread to execute logic at the same time. After the current thread is executed, the state of the next thread is restored by context switching.

[0083] Optionally, matching at least one thread for each acceleration chip according to the multiple threads may be achieved by the following steps:

[0084] Step b1: Determine the number of threads allocated to each acceleration chip according to the configuration information of each acceleration chip.

[0085] Step b2: Based on the number of threads allocated to each acceleration chip, allocate the multiple threads to corresponding acceleration chips, so that each acceleration chip provides resources for at least one thread.

[0086] Each of the acceleration chips supports one thread execution logic at the same time.

[0087] The configuration information here may include GPU type, GPU throughput and computing power, GPU performance, etc.

[0088] In this embodiment, the computing power of each type of GPU of the target task under the current configuration can be determined according to the throughput of each type of GPU of the target task under the current configuration. That is, the throughput of each type of GPU of the target task under the current configuration is proportionally calculated to obtain the performance ratio of each type of GPU as the computing power of the GPU. Then, the allocation information is determined according to the computing power of each type of GPU of the target task under the current configuration, the GPU type, the number of each type of GPU, the number of threads allocated on each type of GPU, and the maximum number of threads corresponding to the acquired target task.

[0089] For example, there are two GPUs (GUP0 and GUP1) and three threads. If GPU1 has strong computing power and more remaining resources, two threads are allocated to it, and GPU0 is allocated the remaining thread. The total number of threads is guaranteed to be consistent in the allocation.

[0090] Therefore, in order to allocate resources reasonably, ensure high availability of resources and improve training efficiency, resources can be provided for multiple threads based on GPU configuration information, such as type, throughput, available resources, performance and other information.

[0091] Optionally, after the current target thread completes executing the logic operation, switching the next target thread for the acceleration chip until all the multiple threads are executed in the at least one acceleration chip to update the model parameters corresponding to the target task can be achieved by the following steps:

[0092] Step c1: after the current target thread completes its execution logic, generate a random state and a corresponding gradient value of the current target thread, store the random state of the current target thread as the initial state of the current target thread, and store the gradient value corresponding to the current target thread; wherein the random state is used to generate a random number sequence.

[0093] Step c2, obtaining the initial state of the next target thread, determining a random number and the random state of the next target thread corresponding to the random number according to the initial state of the next target thread, and storing the random state of the next target thread as the initial state of the next target thread; the random number is used to support the next target thread to perform logical operations on the acceleration chip.

[0094] Optionally, determining the random state of the next target thread according to the initial state of the next target thread may be achieved by the following steps:

[0095] Generate a random number and a random state of the next target thread corresponding to the random number through a random number generator according to the initial state of the next target thread;

[0096] The random state includes: a state of a deep learning framework for representing a Python interface, a state of an open source scientific computing library for representing Python, and a state of a random number generator.

[0097] In this embodiment, in order to save training overhead, the context switching method adopted is not to load all process-related parameters, but to use the key influencing factors affecting the accuracy of the model, that is, the random state (the state of factors that can support the generation of random numbers, such as PyTorch, NumPy (Numerical Python, an open source scientific computing library of Python) and the state of the random number generator of Python's random).

[0098] In addition, storing the random state of the current target thread as the initial state of the current target thread may include:

[0099] The initial state of the current target thread is stored in a preset shared pool of the at least one acceleration chip to support obtaining the initial state of the corresponding thread when elastic resources occur.

[0100] Accordingly, storing the gradient value corresponding to the current target thread may include:

[0101] The gradient value corresponding to the current target thread is unloaded to the memory to support the fusion calculation of the gradient value corresponding to the next target thread.

[0102] Step c3: when all the multiple threads are executed in the at least one acceleration chip, the model parameters corresponding to the target task are updated according to the gradient value corresponding to each of the threads.

[0103] In this embodiment, after completing the gradient calculation of an EasyScaleThread, the state of the next thread is restored by context switching. This is because each thread is treated as an independent worker that completes execution on a GPU, and its state needs to be obtained to implement logical execution. The state here is the random state of the thread, that is, the state of the part that generates the random number, and the random number is used to support the thread execution logic.

[0104] For example, see Figure 4 As shown, Figure 4 A flowchart of an information processing method provided for another embodiment of the present application. Taking a GPU and 4 threads (EST_0, EST_1, EST_2, EST_3) as an example, in a GPU, there is only one GPU computing process. Therefore, only one CUDA context can start the GPU kernel. At the same time, the runtime only arranges the logical execution of one EasyScaleThread, and freezes the execution of other threads. Each EasyScaleThread starts by loading the current small batch of data input, then performs data enhancement, forward-reverse calculation, and then unloads the gradient to the host DRAM (ie, CPU memory). Since deep learning training is performed on a DAG (Directed Acyclic Graph) computational graph, and the output gradient is for subsequent distributed synchronization. The worker's gradient is propagated to the CPU memory (ie, the data bucket for network communication), which can naturally overlap with its reverse calculation and the forward calculation of the next EasyScaleThread. After completing the gradient calculation of an EasyScaleThread, a context switch is used to restore the state of the next thread. Among them, Figure 4 In the above code, tx stands for thread x, which indicates thread (x represents 0, 1, 2, or 3). comp stands for computation, which indicates the computation process. grad stands for gradient, which indicates the gradient value.

[0105] Among them, unloading the gradient value to the memory can reduce the resource usage of the GPU. Figure 4 In it, Context switching for accuracy-consistent means context switching to ensure accuracy; Schedule EasyScaleThread computation on GPU means scheduling EasyScaleThread computation on GPU; Copy gradient means copy gradient; Sync gradient means synchronized gradient; GPU computing process means GPU computing process.

[0106] Specifically, the state can be optimized in the following ways: 1. Locate the non-deterministic sources that affect the final accuracy and minimize the state that needs to be recorded. 2. Take advantage of the data parallel characteristics of DL (deep learning) to minimize the working set of data switching. The working set of EasyScaleThread in GPU memory can be divided into temporary tensors (i.e., a multi-dimensional array of any number of primitive values, a data unit for storing data in deep learning frameworks), activations, model parameters and optimizer states, and gradients. They can be treated differently: for temporary tensors and activations, they are created in the forward calculation and destroyed after the gradient generation is completed in the backward calculation. Therefore, by constraining the minimum gap of context switching to at least one mini-batch, they will be automatically released as expected. For model parameters and optimization states, each data parallel worker will keep a copy during the training process and update it after a mini-batch is processed. Therefore, the model can be reused when switching EasyScaleThread. Gradients are calculated based on different data inputs, which are different in different EasyScaleThreads.

[0107] Since the gradients are generated during the backward computation and are only used in the distributed gradient synchronization after the mini-batch is finished, the gradients are migrated to the host DRAM during the context switch and overlapped with the computation of the next EasyScaleThread. In this way, all EasyScaleThreads are executed alternately until all computations are completed. After that, the distributed synchronization is triggered and the model update is performed once to complete the computation of the mini-batch.

[0108] Among them, the process of locating the source of non-determinism that affects the final accuracy can be as follows: The root cause of non-determinism is scattered throughout the software stack of almost the entire training process.

[0109] First, although deep learning training is composed of operators in a DAG graph (e.g., convolution, batch normalization), in a DAG graph, some operators implicitly rely on states other than the output of their predecessor operators. For example, Dropout generates results based on the state of the random generator in the GPU, BatchNorm (an algorithm often used in deep networks to accelerate neural network training, convergence speed and stability) tracks the mean and variance of tensors based on the worker number considered, and data loaders and data enhancers rely on random states such as random number generators in PyTorch, NumPy, and Python.

[0110] Second, the predictable nature of deep learning is used during DL runtimes to select the most appropriate algorithms for deep learning frameworks, compilers, accelerator algorithm libraries, etc. They often collect performance statistics by applying different algorithms on different mini-batches or function calls, which often differ in implementation and may therefore produce subtly different outputs.

[0111] Then, when resource elasticity occurs, distributed communication may introduce non-deterministic Allreduce (communication operation for distributed deep learning) results. Gradients are organized into gradient communication buckets to optimize the performance of distributed communication, and Allreduce is started once the bucket is full. However, resource elasticity naturally forces the reconstruction of the communication channel. However, these buckets can be different during the reconstruction process because they are actually determined by the order in which the tensors get gradients in the backpropagation. This is because the forward-backward propagation is performed on a DAG graph, and the gradients of concurrent nodes may be generated in different orders. Due to the implementation of the ring Allreduce, non-deterministic order of floating point aggregation is finally introduced. Finally, the GPU kernel implementation of DL operations may be hardware-dependent. For example, some kernel implementations are based on the number of stream processor units, hardware-specific low-precision units, and so on. In order to use elastic heterogeneous GPUs (for example: using V100 and P100 at the same time), it is necessary to select kernel algorithms that are not related to the hardware when executing jobs. Among them, different workloads may suffer additional overhead.

[0112] This embodiment tracks these root causes of uncertainty. The state of data loaders, enhancers, DL operators, etc. (such as random states) is recorded in the context of EasyScaleThread. The context is restored before starting the small batch calculation execution of EasyScaleThread and saved after the execution ends. Hardware factors and gradient synchronization order are recorded in the checkpoint so that they are consistent when resource elasticity occurs.

[0113] Therefore, this application performs thread abstraction on computing tasks in the deep learning framework, allowing multiple distributed data parallel training processes to be executed on a single GPU, decoupling the execution of deep learning from specific hardware, and making full use of the characteristics of deep learning to perform efficient context switching. The state of the thread (such as the random state) is completely saved in the context, and there is no interference between threads, so the accuracy is completely intact.

[0114] Based on the same idea, the present application also provides a device corresponding to the above method, such as Figure 5 As shown, Figure 5 A schematic diagram of the structure of an information processing device provided in an embodiment of the present application. The information processing device may include:

[0115] An acquisition module 501 is used to acquire multiple threads and multiple training sets corresponding to a target task, wherein the target task is executed by multiple threads and the target task supports flexible training;

[0116] A first processing module 502 is used to determine at least one acceleration chip that provides elastic resources for the target task, and match at least one thread to each acceleration chip according to the multiple threads;

[0117] The second processing module 503 is used to perform the following steps for each of the training sets: for each of the acceleration chips, a target thread is allocated to the acceleration chip so that the current target thread performs logical operations according to the training set, and after the current target thread completes the logical operations, the next target thread is switched to the acceleration chip until all the multiple threads are executed in at least one acceleration chip, so as to update the model parameters corresponding to the target task.

[0118] In this embodiment, by setting an acquisition module 501, a first processing module 502 and a second processing module 503, elastic training tasks are performed by acquiring target tasks and multiple training sets participating in training. When resource elasticity occurs, first determine the acceleration chip that can provide elastic resources, and allocate threads to the acceleration chip based on the training task, so that each acceleration chip only supports one thread execution logic at the same time. After the execution is completed, switch to the next thread execution logic to achieve alternating execution. After all threads are executed, update the model parameters once to achieve a round of training, and then repeat the above steps to perform the next round of training. Therefore, by alternating execution, elastic training can be achieved without adjusting or modifying the synchronization method, hyperparameters, mini-batch size, etc. At the same time, since there is no modification of data, the quality of the model trained is consistent with that of the existing fixed resources, that is, the model accuracy is guaranteed. Therefore, while supporting deep learning training tasks for elastic resources, this application can improve the model accuracy of elastic training supported by the prior art, achieve consistency with the quality of the model trained with fixed resources, and ensure model accuracy.

[0119] Optional, get module, specifically used for:

[0120] Determining the number of processes used by the target task in distributed data parallelism;

[0121] Performing thread abstraction on the target task and determining the multiple threads, so that each acceleration chip supports the execution of any number of the threads during runtime;

[0122] The number of the multiple threads is the same as the number of the processes.

[0123] Optionally, the first processing module is specifically used to:

[0124] Determine the number of threads allocated to each acceleration chip according to the configuration information of each acceleration chip;

[0125] Based on the number of threads allocated to each acceleration chip, the multiple threads are allocated to corresponding acceleration chips, so that each acceleration chip provides resources for at least one thread;

[0126] Each of the acceleration chips supports one thread execution logic at the same time.

[0127] Optionally, the second processing module includes a first processing unit, a second processing unit and a third processing unit;

[0128] A first processing unit is used to generate a random state and a corresponding gradient value of the current target thread after the current target thread completes executing logic, store the random state of the current target thread as the initial state of the current target thread, and store the gradient value corresponding to the current target thread; wherein the random state is used to generate a random number sequence;

[0129] a second processing unit, configured to obtain an initial state of a next target thread, determine a random number and a random state of the next target thread corresponding to the random number according to the initial state of the next target thread, and store the random state of the next target thread as the initial state of the next target thread; the random number is used to support the next target thread to perform a logic operation on the acceleration chip;

[0130] The third processing unit is configured to update the model parameters corresponding to the target task according to the gradient value corresponding to each of the threads when all the multiple threads are executed in the at least one acceleration chip.

[0131] Optionally, the second processing unit is specifically configured to:

[0132] Generate a random number and a random state of the next target thread corresponding to the random number through a random number generator according to the initial state of the next target thread;

[0133] The random state includes: a state of a deep learning framework for representing a Python interface, a state of an open source scientific computing library for representing Python, and a state of a random number generator.

[0134] Optionally, the first processing unit is specifically configured to:

[0135] The initial state of the current target thread is stored in a preset shared pool of the at least one acceleration chip to support obtaining the initial state of the corresponding thread when elastic resources occur.

[0136] Optionally, the first processing unit is further specifically configured to:

[0137] The gradient value corresponding to the current target thread is unloaded to the memory to support the fusion calculation of the gradient value corresponding to the next target thread.

[0138] Optionally, the first processing module is further specifically configured to:

[0139] When the elastic resource is triggered, a target resource that currently supports providing training for the target task is obtained, where the target resource is an acceleration chip, and the number of the acceleration chips is at least one.

[0140] The device provided in the embodiment of the present application can realize the above-mentioned Figure 1-4 The implementation principle and technical effect of the method in the embodiment shown are similar and will not be described in detail here.

[0141] Figure 6 The hardware structure diagram of the electronic device provided in the embodiment of the present application is shown in FIG. Figure 6 As shown, the device 600 provided in this embodiment includes: a processor 601 and a memory connected to the processor in communication. The processor 601 and the memory 602 are connected via a bus 603 .

[0142] In a specific implementation process, the processor 601 executes the computer execution instructions stored in the memory 602, so that the processor 601 executes the method in the above method embodiment.

[0143] The specific implementation process of the processor 601 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.

[0144] In the above Figure 6 In the illustrated embodiment, it should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the invention may be directly implemented as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.

[0145] The memory may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk storage.

[0146] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.

[0147] An embodiment of the present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the information processing method of the above method embodiment is implemented.

[0148] An embodiment of the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the information processing method as described above is implemented.

[0149] The computer-readable storage medium mentioned above can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.

[0150] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, referred to as: ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0151] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.

[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. An information processing method, characterized in that: The method comprises: Acquire multiple threads and multiple training sets corresponding to a target task, wherein the target task is executed by multiple threads and the target task supports flexible training; Determine at least one acceleration chip that provides elastic resources for the target task, and match at least one thread to each acceleration chip according to the multiple threads; For each of the training sets, the following steps are performed: for each of the acceleration chips, a target thread is allocated to the acceleration chip, so that the current target thread performs a logic operation according to the training set, and after the current target thread completes the logic operation, the next target thread is switched for the acceleration chip, until the multiple threads are all executed in the at least one acceleration chip, so as to update the model parameters corresponding to the target task; Each of the acceleration chips supports one thread execution logic at the same time.

2. The method according to claim 1, characterized in that: The obtaining of multiple threads corresponding to the target task includes: Determining the number of processes used by the target task in distributed data parallelism; Performing thread abstraction on the target task and determining the multiple threads, so that each acceleration chip supports the execution of any number of the threads during runtime; The number of the multiple threads is the same as the number of the processes.

3. The method according to claim 1 or 2, characterized in that: The matching at least one thread to each acceleration chip according to the multiple threads includes: Determine the number of threads allocated to each acceleration chip according to the configuration information of each acceleration chip; Based on the number of threads allocated to each acceleration chip, the multiple threads are allocated to the corresponding acceleration chips, so that each acceleration chip provides resources for at least one thread.

4. The method according to claim 1 or 2, characterized in that: After the current target thread completes executing the logic operation, switching the next target thread for the acceleration chip until the multiple threads are all executed in the at least one acceleration chip to update the model parameters corresponding to the target task, including: After the current target thread completes executing logic, a random state and a corresponding gradient value of the current target thread are generated, the random state of the current target thread is stored as the initial state of the current target thread, and the gradient value corresponding to the current target thread is stored; wherein the random state is used to generate a random number sequence; Obtaining an initial state of a next target thread, determining a random number and a random state of the next target thread corresponding to the random number according to the initial state of the next target thread, and storing the random state of the next target thread as the initial state of the next target thread; the random number is used to support the next target thread to perform a logic operation on the acceleration chip; When all of the multiple threads are executed in the at least one acceleration chip, the model parameters corresponding to the target task are updated according to the gradient value corresponding to each of the threads.

5. The method according to claim 4, characterized in that Determining the random state of the next target thread according to the initial state of the next target thread includes: Generate a random number and a random state of the next target thread corresponding to the random number through a random number generator according to the initial state of the next target thread; The random state includes: a state of a deep learning framework for representing a Python interface, a state of an open source scientific computing library for representing Python, and a state of a random number generator.

6. The method according to claim 4, characterized in that The storing the random state of the current target thread as the initial state of the current target thread includes: The initial state of the current target thread is stored in a preset shared pool of the at least one acceleration chip to support obtaining the initial state of the corresponding thread when elastic resources occur; Correspondingly, storing the gradient value corresponding to the current target thread includes: The gradient value corresponding to the current target thread is unloaded to the memory to support the fusion calculation of the gradient value corresponding to the next target thread.

7. The method according to claim 1 or 2, characterized in that: The determining of at least one acceleration chip that provides elastic resources for the target task includes: When the elastic resource is triggered, a target resource that currently supports providing training for the target task is obtained, where the target resource is an acceleration chip, and the number of the acceleration chips is at least one.

8. An information processing device, characterized in that: The device comprises: An acquisition module, used to acquire multiple threads and multiple training sets corresponding to a target task, wherein the target task is executed by multiple threads and the target task supports flexible training; A first processing module is used to determine at least one acceleration chip that provides elastic resources for the target task, and match at least one thread to each acceleration chip according to the multiple threads; The second processing module is used to perform the following steps for each of the training sets: for each of the acceleration chips, assign a target thread to the acceleration chip, so that the current target thread performs a logic operation according to the training set, and after the current target thread completes the logic operation, switch the next target thread for the acceleration chip until the multiple threads are all executed in the at least one acceleration chip, so as to update the model parameters corresponding to the target task; Each of the acceleration chips supports one thread execution logic at the same time.

9. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the information processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the processor executes the computer-executable instructions, the information processing method according to any one of claims 1 to 7 is implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the information processing method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Image concurrent processing method, device and system based on multi-GPU card

    CN109388496A

  • Distributed parallel training method and device, and readable medium

    CN111381966A