Method for constructing virtual environment for training model and method for training model
By building a virtual environment and generating soft links of unified paths, the path consistency problem of distributed deep learning model training when limited resources are found in Kubernetes clusters and Horovod and Spark is solved, achieving larger-scale training and more efficient resource utilization.
Patent Information
- Application Number
- CN202211180643.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-09-26
AI Technical Summary
The existing distributed deep learning model training scheme limits the training scale and resource utilization when resources are limited in Kubernetes clusters, and the combination of Horovod and Spark has problems with path consistency and toolkit management difficulty.
By building a virtual environment, installing a distributed processing system and a deep learning framework, and generating soft links of unified paths in the virtual environment, ensuring the consistent operation of the distributed deep learning framework in the cluster, and then distributing it to each node of the cluster for training.
It supports the expansion of model training scale, shortening training time, and improving resource utilization in Kubernetes clusters, while simplifying the management and upgrading process of toolkits.
Smart Images

Figure CN115587623B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a method for constructing a virtual environment for training a deep learning model, a method for training a deep learning model, a computing device, and a readable storage medium. Background Art
[0002] In recent years, deep learning has been widely applied in many fields. The excellent performance of deep learning models usually depends on the participation of large-scale data, making the models very large, which continuously challenges the training methods and training speed of the models. Therefore, in the field of deep learning, there is a wide demand for distributed training of models because it can improve the training speed of the models.
[0003] Distributed training usually uses multiple GPU / CPU servers, and forms a distributed deep learning computing mode by constructing a high-performance communication network for data distribution and model synchronization. Distributed training usually has two methods: data parallelism and model parallelism. Among them, data parallelism is suitable for accelerating the training of most deep learning models. Common deep learning frameworks or platforms, including TensorFlow, PyTorch, and Apache MXNet, provide built-in methods to support distributed training with multiple GPUs and multiple worker nodes. In addition, distributed training of models can also be achieved by using a distributed deep learning framework (e.g., Horovod).
[0004] There are two existing distributed model training solutions. One solution is to install Horovod in a Docker container and schedule it through Kubernetes. This solution is suitable for independent Kubernetes clusters. However, this solution is not conducive to Kubernetes using large-scale Spark cluster resources. When Kubernetes cluster resources are limited, the training scale of the job is limited, and it is not conducive to using more cluster resources. Another solution is to install Horovod on all nodes of the Spark cluster and enable support for Horovod on Spark. The main disadvantages of this solution include: 1. It is difficult to keep the installation environment of Python and Horovod toolkits consistent on all nodes of the cluster. For example, if the GCC version is inconsistent, the compilation of Gloo, which Horovod depends on, may fail, resulting in the failure of Horovod installation. Moreover, it is difficult to install any toolkit on a large-scale cluster. 2. After Python and Horovod are installed, it is very difficult to expand and upgrade various toolkits. It is difficult to upgrade the installed toolkits, expand other toolkits, and install customized toolkits for each job. 3. Horovod's own characteristics require that the job scripts and the tokens for accessing HDFS need to be distributed to all nodes in the cluster in the same path. However, each node in a Spark cluster usually has multiple data disks. Therefore, the jobs and virtual environments distributed to each node will most likely be temporarily stored in different data disks, resulting in different storage directories. This leads to conflicts in the combination of Horovod and Spark.
[0005] To this end, the present invention provides a solution for constructing a virtual environment for training a deep learning model and a solution for training the deep learning model to solve the problems existing in the prior art. Summary of the invention
[0006] To this end, the present invention provides a method for constructing a virtual environment for training a deep learning model, a training method for a deep learning model, a computing device, and a readable storage medium to solve or at least alleviate the above problems.
[0007] According to a first aspect of the present invention, a method for constructing a virtual environment for training a deep learning model is provided, the method comprising: constructing a virtual environment that runs a predetermined computer language; installing a distributed processing system and a distributed deep learning framework in the virtual environment; in the virtual environment, generating unified paths for token files corresponding to an executor that runs a predetermined computer language and a distributed deep learning framework, respectively; and packaging the virtual environment and distributing it to a cluster based on the distributed processing system.
[0008] Optionally, in the method for constructing a virtual environment for training a deep learning model according to the present invention, generating a unified path for an executor running a predetermined computer language includes: in a first predetermined script and a second predetermined script of a distributed deep learning framework, determining whether there is a first unified path corresponding to the executor of the predetermined computer language; if not, creating a first soft link for the executor of the predetermined computer language, where the first soft link points to the first unified path.
[0009] Optionally, in the method for constructing a virtual environment for training a deep learning model according to the present invention, generating a first unified path for an executor running a predetermined computer language includes: in a third predetermined script of a distributed deep learning framework, setting the path of the executor of the predetermined computer language as the first unified path.
[0010] Optionally, in the method for constructing a virtual environment for training a deep learning model according to the present invention, generating a unified path for a token file corresponding to a distributed deep learning framework includes: in a first predetermined script and a second predetermined script of a distributed deep learning framework, determining whether there is a second unified path corresponding to the token file; if not, creating a second soft link for the token file, where the second soft link points to the second unified path.
[0011] Optionally, in the method for constructing a virtual environment for training a deep learning model according to the present invention, generating a unified path for a token file corresponding to a distributed deep learning framework includes: in a third predetermined script of a distributed deep learning framework, setting the path of the token file as the second unified path.
[0012] Optionally, in the method for constructing a virtual environment for training a deep learning model according to the present invention, after packaging the virtual environment, distributing it to a cluster based on a distributed processing system includes: packaging the virtual environment; submitting the packaged virtual environment to the cluster and distributing the packaged virtual environment to each node of the cluster; setting the environment variables of the cluster.
[0013] Optionally, in the method for constructing a virtual environment for training a deep learning model according to the present invention, the first predetermined script includes a driver-related program of a distributed processing system, and the second predetermined script includes a task-related program of a distributed system.
[0014] Optionally, in the method for constructing a virtual environment for training a deep learning model according to the present invention, the third predetermined script includes a collective communication library-related program.
[0015] Optionally, in the method for constructing a virtual environment for training a deep learning model according to the present invention, the distributed processing system includes spark.
[0016] Optionally, in the method for constructing a virtual environment for training a deep learning model according to the present invention, the distributed deep learning framework includes Horovod.
[0017] According to a second aspect of the present invention, there is provided a method for training a deep learning model, the method including: packing and distributing the constructed virtual environment to a cluster based on a distributed processing system by the method as described above; splitting the training task of the deep learning model into multiple sub-training tasks; and distributing the multiple sub-training tasks to each node of the cluster, so that each node performs distributed training on the deep learning model through a distributed deep learning framework based on an executor and a token file of a predetermined computer language under a unified path.
[0018] According to a third aspect of the present invention, there is provided a computing device, including: at least one processor; a memory storing program instructions, wherein the program instructions are configured to be executed by the at least one processor as described above, and the program instructions include instructions for executing the method as described above.
[0019] According to a fourth aspect of the present invention, there is provided a readable storage medium storing program instructions, which when read and executed by a computing device, cause the computing device to execute the method as described above.
[0020] In the method for constructing a virtual environment for training a deep learning model of the present invention, by generating a unified path for an executor running a predetermined computer language and a token file corresponding to a distributed deep learning framework in the constructed virtual environment, it is feasible to run the distributed deep learning framework in a cluster of a distributed processing system through the virtual environment. When applied in actual production, it can not only support expanding the training scale of the model, shortening the training time of the model, but also improve resource utilization.
[0021] In the training method of deep learning of the present invention, a virtual environment is constructed by the method for constructing a virtual environment for training a deep learning model and distributed to a cluster based on a distributed processing system, so as to realize distributed training of the deep learning model. It can also support expanding the training scale of the model, shortening the training time of the model, and improving resource utilization.
[0022] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other objects, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. Brief Description of the Drawings
[0023] To achieve the above and related purposes, certain illustrative aspects are described herein in conjunction with the following description and drawings, which indicate various ways in which the principles disclosed herein can be practiced, and all aspects and their equivalent aspects are intended to fall within the scope of the claimed subject matter. The above and other purposes, features, and advantages of the present disclosure will become more apparent by reading the following detailed description in conjunction with the drawings. Throughout the present disclosure, like reference numerals generally refer to like components or elements.
[0024] Figure 1 A block diagram showing the physical components of a computing device 100 is presented;
[0025] Figure 2 A flowchart showing a method 200 for constructing a virtual environment for training a deep learning model according to an embodiment of the present invention is presented. Detailed Description of Specific Embodiments
[0026] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Instead, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0027] Figure 1 A block diagram showing the physical components (i.e., hardware) of a computing device 100 is presented. In a basic configuration, the computing device 100 includes at least one processing unit 102 and a system memory 104. According to one aspect, depending on the configuration and type of the computing device, the processing unit 102 can be implemented as a processor. The system memory 104 includes, but is not limited to, volatile storage (e.g., random access memory), non-volatile storage (e.g., read-only memory), flash memory, or any combination of such memories. According to one aspect, the system memory 104 includes an operating system 105 and program modules 106, and the program modules 106 include program instructions 120 for executing the method for constructing a virtual environment for training a deep learning model and the method for training a deep learning model of the present invention.
[0028] According to one aspect, the operating system 105 is suitable for controlling the operation of the computing device 100, for example. In addition, the examples are practiced in conjunction with a graphics library, other operating systems, or any other application, and are not limited to any specific application or system. In Figure 1The basic configuration is shown by those components within the dashed line 108. According to one aspect, the computing device 100 has additional features or functions. For example, according to one aspect, the computing device 100 includes additional data storage devices (removable and / or non-removable), such as magnetic disks, optical disks, or magnetic tapes. Such additional storage is Figure 1 shown by the removable storage device 109 and the non-removable storage device 110 in
[0029] As stated above, according to one aspect, program modules 106 are stored in the system memory 104. According to one aspect, the program modules 106 may include one or more application programs, and the present invention does not limit the types of application programs. For example, the application programs may include: email and contact applications, word processing applications, spreadsheet applications, database applications, slide show applications, painting or computer-aided applications, web browser applications, etc.
[0030] According to one aspect, the examples may be practiced in a circuit including discrete electronic components, a packaged or integrated electronic chip containing logic gates, a circuit utilizing a microprocessor, or on a single chip containing electronic components or a microprocessor. For example, the examples may be practiced via a system-on-chip (SOC) in which each or many of the components shown in Figure 1 may be integrated on a single integrated circuit. According to one aspect, such an SOC device may include one or more processing units, graphics units, communication units, system virtualization units, and various application functions, all of which are integrated (or "burned") onto a chip substrate as a single integrated circuit. When operating via the SOC, the functions described herein may be operated via dedicated logic integrated with other components of the computing device 100 on a single integrated circuit (chip). Embodiments of the present invention may also be practiced using other technologies capable of performing logical operations (such as AND, OR, and NOT), including but not limited to mechanical, optical, fluidic, and quantum technologies. Additionally, embodiments of the present invention may be practiced within a general-purpose computer or in any other circuit or system.
[0031] According to one aspect, the computing device 100 may also have one or more input devices 112, such as a keyboard, mouse, pen, voice input device, touch input device, etc. It may also include output devices 114, such as a display, speaker, printer, etc. The foregoing devices are examples and other devices may also be used. The computing device 100 may include one or more communication connections 116 that allow communication with other computing devices 118. Examples of suitable communication connections 116 include but are not limited to: RF transmitter, receiver, and / or transceiver circuits; Universal Serial Bus (USB), parallel, and / or serial ports.
[0032] As used herein, the term computer-readable medium includes computer storage media. Computer storage media can include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (e.g., computer-readable instructions, data structures, or program modules). System memory 104, removable storage device 109, and non-removable storage device 110 are all examples of computer storage media (i.e., memory storage). Computer storage media can include random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other article that can be used to store information and can be accessed by computer device 100. According to one aspect, any such computer storage media can be part of computing device 100. Computer storage media does not include carrier waves or other propagated data signals.
[0033] According to one aspect, communication media is implemented by computer-readable instructions, data structures, program modules, or other data in a modulated data signal (e.g., a carrier wave or other transmission mechanism), and includes any information delivery medium. According to one aspect, the term "modulated data signal" describes a signal having one or more sets of characteristics or a signal that has been altered in such a way as to encode information in the signal. By way of example and not limitation, communication media includes wired media such as a wired network or direct wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.
[0034] In one embodiment of the present invention, computing device 100 includes one or more processors and one or more readable storage media storing program instructions. When the program instructions are configured to be executed by one or more processors, the computing device is caused to execute the method for constructing a virtual environment for training a deep learning model and the method for training a deep learning model in the embodiments of the present invention.
[0035] Some terms related to the present invention are described below:
[0036] Horovod, an open-source distributed deep learning framework developed by Uber, can be used in conjunction with popular deep learning toolkits such as TensorFlow, Keras, PyTorch, and Apache MXNet. Horovod uses the AllReduce algorithm to replace the previous parameter server method for fast distributed training, and also provides various optimization methods such as tensor fusion, gradient compression, and support for NCCL communication to further accelerate the execution speed of distributed training. The goal of the All Reduce algorithm is to efficiently integrate (reduce) the data in different machines and then distribute the results to each machine.
[0037] Spark (full name: Apache Spark), a distributed open-source processing system for big data workloads. It uses in-memory caching and optimized query execution methods to enable fast analytical queries on data of any scale. It provides development APIs in Java, Scala, Python, and R languages, supporting code reuse across multiple workloads, batch processing, interactive queries, real-time analytics, machine learning, and graph processing, etc.
[0038] TensorFlow, an end-to-end open-source machine learning platform that allows you to easily build models and train and deploy models in the cloud, on-premises, in the browser, or on devices.
[0039] MPI is a message-passing application programming interface that includes protocols and semantic descriptions that specify how it behaves in various implementations. The goal of MPI is high performance, large scale, and portability. MPI remains the main model for high-performance computing today.
[0040] Gloo, a collective communication library with many algorithms useful for machine learning applications, including barriers, broadcasts, and AllReduce.
[0041] Docker, an open-source application container engine that allows developers to package their applications and dependencies into a lightweight, portable container and then publish it to any popular Linux machine, and can also achieve virtualization.
[0042] Kubernetes, an open-source container orchestration platform for running distributed applications and services at scale.
[0043] Figure 2 A flowchart of a method 200 for constructing a virtual environment for training a deep learning model according to an embodiment of the present invention is shown. The method 200 is adapted to be executed in a computing device (such as the aforementioned computing device 100). As Figure 2 shown, the method 200 begins at 210.
[0044] 210. Build a virtual environment for running a predetermined computer language.
[0045] According to an embodiment of the present invention, the predetermined computer language may be the Python language or other computer languages. Optionally, build a virtual environment for running the Python language. First, create a Python virtual environment. Specifically, create a directory for storing Python-related data, and use an environment manager (e.g., conda) to install Python-related data of the environment into the specified directory. Then, activate the Python virtual environment.
[0046] The following is exemplary code for building a virtual environment for running Python:
[0047] mkdir -p / data / python3.6
[0048] conda create -p / data / python3.6 python=3.6
[0049] source activate / data / python3.6
[0050] In the exemplary code, the version of the Python language is 3.6, but it is not limited thereto. Of course, it can also be other versions of the predetermined computer language.
[0051] 220. Install a distributed processing system and a distributed deep learning framework in the virtual environment.
[0052] Among them, the distributed processing system may be Spark or other distributed processing systems, such as Flink, Storm, etc. The distributed deep learning framework may be Horovod, for example, or other distributed deep learning frameworks, such as PipeDream, GPipe, etc.
[0053] According to an embodiment of the present invention, install the toolkits of Spark and Horovod in a virtual environment. Optionally, other toolkits for training deep learning models can also be installed in the virtual environment, such as the toolkit of TensorFlow, etc. Optionally, install various tools on which the distributed processing system and the distributed deep learning framework depend in the virtual environment. For example, install Gloo on which Horovod depends in the virtual environment. When Gloo is installed in the virtual system, its version needs to be above GCC5. Since Gloo is first compiled into a static library and then linked to the Horovod binary execution program, Gloo can be compiled on a machine with a high version of GCC during compilation, and the GCC version of the running machine can be either a high version or a low version. Therefore, a certain limit is imposed on its version when Gloo is installed.
[0054] The following is exemplary code for installing toolkits in a virtual environment:
[0055] python -m pip install --upgrade pip
[0056] pip install tensorflow==2.2.1
[0057] pip install pyspark==2.4.6
[0058] HOROVOD_WITH_MPI=1
[0059] HOROVOD_WITH_GLOO=1
[0060] pip install horovod[tensorflow,spark]
[0061] In the above exemplary code, first install the pip tool and upgrade it, then install TensorFlow with version 2.2.1 and PySpark with version 2.4.6 through the pip tool. PySpark is a Python library provided by Spark. Then set the MPI and Gloo tools on which Horovod depends, and finally install Horovod and the frameworks it needs.
[0062] The technical solution of the present invention installs a distributed processing system, a distributed deep learning framework, and other toolkits in a virtual environment, which can not only make the installation simple and easy, but also upgrade, expand, and customize the distributed processing system, the distributed deep learning framework, and other toolkits more conveniently and easily. It improves the flexibility of managing the distributed processing system and the distributed deep learning framework, and can also easily improve its performance.
[0063] 230. In the virtual environment, generate unified paths for the executor running a predetermined computer language and the token file corresponding to the distributed deep learning framework respectively.
[0064] According to an embodiment of the present invention, in the first predetermined script and the second predetermined script of the distributed deep learning framework, determine whether there is a first unified path corresponding to the path of the executor of the predetermined computer language. If not, create a first soft link of the executor of the predetermined computer language, and the first soft link points to the first unified path. Wherein, the first predetermined script includes the driver-related program of the distributed processing system, and the second predetermined script includes the task-related program of the distributed system.
[0065] Specifically, modify the source code of the distributed deep learning framework to achieve that before the distributed processing system distributes jobs to each node of the cluster and starts running, and before the subprocess of the distributed deep learning framework that executes the jobs starts, by generating a soft link to the unified path of the virtual environment on the cluster, the requirement of the distributed deep learning framework for the unified path of the executor of the predetermined computer language is met. The first predetermined script is the horovod / spark / driver / driver_service.py script, that is, the driver_service.py script under the horovod / spark / driver / path. The second predetermined script is the horovod / spark / task / task_service.py script, that is, the task_service.py script under the horovod / spark / task / path. Add code for determining whether there is a first unified path corresponding to the executor of the predetermined computer language in the SparkDriverClient.__init__ function of the first predetermined script and the SparkTaskClient.__init__ function of the second predetermined script. If not, create a first soft link of the executor of the predetermined computer language, and the first soft link points to the first unified path.
[0066] According to an embodiment of the present invention, in the first predetermined script and the second predetermined script of the distributed deep learning framework, it is determined whether there is a second unified path corresponding to the token file. If not, a second soft link of the token file is created, where the second soft link points to the second unified path.
[0067] Specifically, by generating a soft link of the unified path for the token files stored in the same directory of each distribution job, the requirement of the distributed deep learning framework for unifying the token file paths is met. The first predetermined script is the horovod / spark / driver / driver_service.py script, that is, the driver_service.py script in the horovod / spark / driver / path. The second predetermined script is the horovod / spark / task / task_service.py script, that is, the task_service.py script in the horovod / spark / task / path. Code for determining whether there is a second unified path corresponding to the token file is added in the SparkDriverClient.__init__ function of the first predetermined script and the SparkTaskClient.__init__ function of the second predetermined script. If not, a second soft link of the token file is created, and the second soft link points to the second unified path.
[0068] The following is the exemplary code added in the first predetermined script and the second predetermined script in the distributed deep learning framework:
[0069] cur_path = os.getcwd() # Return the working directory of the current process
[0070] print('cur_path:', cur_path)
[0071] appid = cur_path.split(" / ")[-2] # Get the second-to-last content separated by " / " in the working directory of the current process
[0072] user = cur_path.split(" / ")[-4] # Get the fourth-to-last content separated by " / " in the working directory of the current process
[0073] python_path = ' / tmp / horovod_spark_' + user + ' / ' + appid # Concatenate to form the python_path
[0074] os.system('mkdir -p'+ python_path) # Ensure the python_path exists, create it if it doesn't
[0075] if not os.path.exists(python_path + " / container_tokens"):
[0076] os.system('cp -rf. / container_tokens'+ python_path) # If the path pythonpath / container tokens (representing the second unified path) doesn't exist, forcefully overwrite files in the specified directory and create a soft link to the token files for a unified path
[0077] if not os.path.exists(python_path + " / python3.6"):
[0078] os.system('ln -s'+ cur_path + ' / horovodtest / python3.6'+ python_path + ' / python3.6') # If the path python path / python 3.6 (representing the first unified path) doesn't exist, create a soft link to the first unified path
[0079] According to an embodiment of the present invention, in the third predetermined script of the distributed deep learning framework, the path of the executor of the predetermined computer language is set to the first unified path. Wherein, the third predetermined script includes programs related to the collective communication library.
[0080] Specifically, modify the source code of the distributed deep learning framework, and specify the unified Python executor path in the command for the driver to distribute tasks in the source code of the distributed deep learning framework. The third predetermined script is the gloo_run.py script in the horovod / runner / path, that is, the gloo_run.py script under the horovod / runner / path. Add code for setting the path of the executor of the predetermined computer language to the first unified path at the beginning of the launch_gloo function in the third predetermined script.
[0081] The following is the exemplary code added in the third predetermined script of the distributed deep learning framework:
[0082] cur_path = os.getcwd() # Return the working directory of the current process
[0083] appid = cur_path.split(" / ")[-2]
[0084] user = cur_path.split(" / ")[-4]
[0085] python_path = ' / tmp / horovod_spark_' + user + ' / ' + appid
[0086] sys.executable = python_path + ' / python3.6 / bin / python' # Set the unified path of the Python interpreter.
[0087] According to an embodiment of the present invention, in the third predetermined script of the distributed deep learning framework, the path of the token file is set to the second unified path. Wherein, the third predetermined script includes programs related to the collective communication library.
[0088] Specifically, modify the source code of the distributed deep learning framework, and specify the unified Python executor path in the command for the driver of the distributed deep learning framework source code to distribute tasks. The third predetermined script is the gloo_run.py script in the horovod / runner / , that is, the gloo_run.py script under the horovod / runner / path. Add code for setting the path of the executor of the predetermined computer language to the first unified path at the beginning of the launch_gloo function in the third predetermined script.
[0089] The following is the exemplary code added in the third predetermined script of the distributed deep learning framework:
[0090] cur_path = os.getcwd()
[0091] appid = cur_path.split(" / ")[-2]
[0092] user = cur_path.split(" / ")[-4]
[0093] python_path = ' / tmp / horovod_spark_' + user + ' / ' + appid
[0094] token_path = python_path + ' / container_tokens'
[0095] env['HADOOP_TOKEN_FILE_LOCATION'] = token_path
[0096] According to the technical solution of the present invention, the executor running a predetermined computer language in the virtual environment and the token file corresponding to the distributed deep learning framework are unified respectively. By means of soft links, the executors of the predetermined computer language on the Driver on the cluster and all taskers and the tokens for accessing the distributed file system have the same absolute path, ensuring the consistency of the paths.
[0097] 240. Package the virtual environment and distribute it to the cluster based on the distributed processing system.
[0098] According to an embodiment of the present invention, the constructed virtual environment is packaged, and then the packaged virtual environment is submitted to the cluster based on the distributed processing system, the packaged virtual environment is distributed to each node of the cluster, and the environment variables of the cluster are set.
[0099] The following shows exemplary code for packaging the virtual environment and distributing it to the cluster based on the distributed processing system:
[0100] tar - cvf python3.6.tar python3.6 / * / / Package the constructed virtual environment
[0101] spark - submit -- master yarn -- deploy - mode cluster
[0102] -- conf spark.yarn.dist.archives=python3.6.tar#horovod
[0103] -- conf spark.yarn.appMasterEnv.PYSPARK_PYTHON=. / horovod / python3.6 / bin / python3
[0104] -- conf spark.yarn.appMasterEnv.PYSPARK_DRIVER_PYTHON=. / horovod / python3.6 / bin / python3
[0105] -- conf spark.executorEnv.PYSPARK_PYTHON=. / horovod / python3.6 / bin / python3
[0106] --conf spark.executorEnv.PYSPARK_DRIVER_PYTHON=. / horovod / python3.6 / bin / python3
[0107] run.py
[0108] Submit the virtual environment to the cluster through spark-submit, distribute the virtual environment, and set relevant environment variables, including: PYSPARK_PYTHON, PYSPARK_DRIVER_PYTHON, PYSPARK_PYTHON, and PYSPARK_DRIVER_PYTHON.
[0109] The construction solution of the virtual environment for training deep learning models in the present invention solves the problem in the prior art that after Python and Horovod virtual environment packages are distributed to each computing node through Spark, the submitted job script and virtual environment package are placed in a certain directory on a certain data disk. Since each node in the Spark cluster generally has multiple data disks, the jobs and virtual environments distributed to each node are probably temporarily stored on different data disks, resulting in different storage directories, which cannot meet the requirement of Horovod that the Python executors on the Driver and all taskers and the tokens for accessing HDFS use the same absolute path. The solution of the present invention can make the Python executors on the Driver and all taskers and the tokens for accessing HDFS have the same absolute path by constructing soft links, ensuring the consistency of the paths.
[0110] According to an embodiment of the present invention, the construction method of the virtual environment of the present invention can be applied to training deep learning models. Method 200 can be applied to various scenarios such as images, speech, video, machine translation, etc. For example, in the image scenario, the corresponding deep learning model can be an image classification model, an object detection model, etc.; in the machine translation scenario, the corresponding deep learning model can be a neural machine translation model. Among them, the neural machine translation model can be a sequence-to-sequence model, having an encoder made by a gated recurrent unit, an encoder made by a gated recurrent unit, and an attention mechanism.
[0111] The present invention also provides a method for training a deep learning model. According to an embodiment of the present invention, the constructed virtual environment is packaged by the method 200 of the present invention and then distributed to a cluster based on a distributed processing system. The training task of the deep learning model is split into multiple sub-training tasks. The multiple sub-training tasks are distributed to each node of the cluster so that each node can perform distributed training on the deep learning model through an executor and a token file of a predetermined computer language under a unified path based on a distributed deep learning framework.
[0112] Among them, the method for training a deep learning model can be applied to various scenarios such as images, speech, video, machine translation, etc. For example, in the image scenario, the corresponding deep learning model can be an image classification model, an object detection model, etc.; in the machine translation scenario, the corresponding deep learning model can be a neural network machine translation model.
[0113] Optionally, the parts of the deep learning model that can be split for training are split out to form multiple sub-training models, and the sub-training models are distributed to each node of the cluster for synchronous training.
[0114] Optionally, the training data of the deep learning model is split into multiple sub-training data, and the sub-training data is distributed to each node of the cluster for training the model. Before performing distributed training, each node in the cluster will obtain the deep learning model to be trained, and the initial values of the model parameters of the deep learning model have been preset. The types of training data can be: image samples, speech samples, natural language processing samples.
[0115] The method for constructing a virtual environment for training a deep learning model according to the present invention generates a unified path for an executor that runs a predetermined computer language and a token file corresponding to a distributed deep learning framework in the constructed virtual environment, making it feasible to run the distributed deep learning framework in a cluster of a distributed processing system. In practical production applications, it can not only support expanding the training scale of the model, shortening the training time of the model, but also improve resource utilization.
[0116] The training method of deep learning according to the present invention constructs a virtual environment through the method for constructing a virtual environment for training a deep learning model and distributes it to a cluster based on a distributed processing system, realizing distributed training of the deep learning model. It can also support expanding the training scale of the model, shortening the training time of the model, and improving resource utilization.
[0117] Furthermore, through the application of the virtual environment, it is convenient to upgrade, expand, or customize the distributed processing system, the distributed deep learning framework, and various toolkits.
[0118] A8. The method according to A3 or A5, wherein the third predetermined script includes a program related to a collective communication library. A9. The method according to any one of A1 to A8, wherein the distributed processing system includes spark. A10. The method according to any one of A1 to A9, wherein the distributed deep learning framework includes Horovod.
[0119] The various techniques described herein can be implemented in combination with hardware or software, or a combination thereof. Thus, the methods and devices of the present invention, or certain aspects or portions of the methods and devices of the present invention, may take the form of program code (i.e., instructions) embedded in a tangible medium, such as a removable hard disk, a USB flash drive, a floppy disk, a CD-ROM, or any other machine-readable storage medium, wherein when the program is loaded into a machine such as a computer and executed by the machine, the machine becomes a device for practicing the present invention.
[0120] In the case where the program code is executed on a programmable computer, the mobile terminal generally includes a processor, a processor-readable storage medium (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device. Among them, the memory is configured to store the program code; the processor is configured to execute the method for constructing a virtual environment for training a deep learning model and the method for training a deep learning model of the present invention according to the instructions in the program code stored in the memory.
[0121] By way of example and not limitation, the readable medium includes a readable storage medium and a communication medium. The readable storage medium stores information such as computer-readable instructions, data structures, program modules, or other data. The communication medium generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and includes any information delivery medium. A combination of any of the above is also included within the scope of the readable medium.
[0122] In the specification provided herein, the algorithms and displays are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the examples of the present invention. Based on the above description, the structure required to construct such a system is obvious. In addition, the present invention is not directed to any particular programming language. It should be understood that the content of the present invention described herein can be implemented using various programming languages, and the description of a particular language above is for the purpose of disclosing the best mode of the present invention.
[0123] In the description provided herein, numerous specific details are set forth. It will be understood, however, that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description.
[0124] Similarly, it should be understood that in order to streamline this disclosure and help understand one or more of the various inventive aspects, in the foregoing description of the exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the invention.
[0125] Those skilled in the art will appreciate that the modules or units or components of the devices in the examples disclosed herein may be arranged in the devices as described in that embodiment, or alternatively may be located in one or more devices different from those of the example. The modules in the foregoing examples may be combined into one module or further divided into multiple sub-modules.
[0126] Those skilled in the art can understand that the modules in the devices of the embodiments can be adaptively changed and disposed in one or more devices different from those of the embodiment. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be adopted to combine all the features disclosed in this specification (including the accompanying claims, abstract and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract and drawings) can be replaced by an alternative feature that provides the same, equivalent or similar purpose.
[0127] In addition, those skilled in the art will be able to understand that although some of the embodiments described herein include certain features included in other embodiments but not other features, the combination of features of different embodiments means that it is within the scope of the invention and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0128] In addition, some of the embodiments described herein are described as methods or combinations of method elements that can be implemented by a processor of a computer system or by other devices performing the functions. Therefore, a processor having the necessary instructions for implementing the method or method element forms an apparatus for implementing the method or method element. In addition, the elements described herein of the apparatus embodiments are examples of apparatuses for performing the functions performed by the elements for the purpose of implementing the invention.
[0129] As used herein, unless otherwise specified, the use of ordinal numbers "first", "second", "third", etc. to describe ordinary objects merely indicates different instances of similar objects and is not intended to imply that the objects so described must have a given order in time, space, ranking, or in any other manner.
[0130] Although the invention has been described in terms of a limited number of embodiments, those skilled in the art of the present technology will appreciate that other embodiments can be contemplated within the scope of the invention as thus described. In addition, it should be noted that the language used in this specification has been principally selected for readability and instructional purposes and not for the purpose of explaining or limiting the subject matter of the invention. Accordingly, many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the appended claims. For the scope of the invention, the disclosure of the invention is illustrative and not restrictive, and the scope of the invention is defined by the appended claims.
Claims
1. A method for constructing a virtual environment for training a deep learning model, the method comprising: Constructing a virtual environment that runs a predetermined computer language; Installing a distributed processing system and a distributed deep learning framework in the virtual environment; In the virtual environment, generating unified paths for an executor that runs the predetermined computer language and a token file corresponding to the distributed deep learning framework, including in a first predetermined script and a second predetermined script of the distributed deep learning framework, determining whether there is a second unified path corresponding to the token file, and if not, creating a second soft link to the token file, where the second soft link points to the second unified path; Packaging the virtual environment and distributing it to a cluster based on the distributed processing system.
2. The method according to claim 1, wherein, Generating a unified path for an executor that runs the predetermined computer language, including: In the first predetermined script and the second predetermined script of the distributed deep learning framework, determining whether there is a first unified path corresponding to the executor of the predetermined computer language; If not, creating a first soft link to the executor of the predetermined computer language, where the first soft link points to the first unified path.
3. The method according to claim 1 or 2, wherein, Generating a first unified path for an executor that runs the predetermined computer language, including: In a third predetermined script of the distributed deep learning framework, setting the path of the executor of the predetermined computer language to the first unified path.
4. The method according to claim 1, wherein Generating a unified path for a token file corresponding to the distributed deep learning framework, including: In a third predetermined script of the distributed deep learning framework, setting the path of the token file to the second unified path.
5. The method according to claim 1, wherein, The step of packaging the virtual environment and distributing it to a cluster based on the distributed processing system includes: Packaging the virtual environment; Submitting the packaged virtual environment to the cluster and distributing the packaged virtual environment to each node of the cluster; Setting the environment variables of the cluster.
6. The method according to claim 1 or 2, wherein, The first predetermined script includes a driver-related program of the distributed processing system, and the second predetermined script includes a task-related program of the distributed processing system.
7. The method according to claim 3, wherein The third predetermined script includes a collective communication library-related program.
8. The method according to claim 1, wherein The distributed processing system includes spark.
9. The method according to claim 1, wherein, The distributed deep learning framework includes Horovod.
10. A method for training a deep learning model, the method comprising: Packaging the constructed virtual environment by the method according to any one of claims 1 to 9 and distributing it to a cluster based on a distributed processing system; Splitting the training task of the deep learning model into multiple sub-training tasks; Distributing the multiple sub-training tasks to each node of the cluster so that each node performs distributed training on the deep learning model through the distributed deep learning framework based on the executor and the token file of the predetermined computer language under a unified path.
11. A computing device, comprising: At least one processor; And A memory stores program instructions, wherein the program instructions are configured to be executed by the at least one processor, and the program instructions include instructions for executing the method according to any one of claims 1 to 10.
12. A readable storage medium storing program instructions, which, when read and executed by a computing device, cause the computing device to execute the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Distributed development method and device, storage medium and computer equipment
CN112394944A