Cloud-based deep learning training platform and method, medium and equipment

By migrating deep learning model training to the cloud and using cloud resources and platform technology, the problems of traditional local training are solved, efficient and stable training management and resource scheduling are achieved, and the work efficiency and product quality of algorithm students are improved.

CN120372268APending Publication Date: 2025-07-25MOMENTA (SUZHOU) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410095252.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Traditional deep learning model training is inefficient and poorly stable under local resource limitations, making it difficult to meet efficient and stable training needs.

Method used

Migrate the deep learning model training platform from local to the cloud, and utilize the cloud computing resources and platforms to achieve efficient management and resource scheduling of training tasks through technical means such as load balancing, GPU resource management, data preprocessing and training task management.

Benefits of technology

It improves training speed and efficiency, ensures the stability and data security of training tasks, provides a unified training platform and visual interface, and supports algorithm students to better manage and optimize the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372268A_ABST
    Figure CN120372268A_ABST
Patent Text Reader

Abstract

The invention provides a cloud-based deep learning training platform, method, medium and device, the training platform comprises a platform front end located at the local of a user and a platform rear end deployed at the cloud, the platform front end is used for configuring a deep learning training task, and the platform rear end is used for configuring the deep learning training task. The training task progress and / or result fed back by the platform rear end are / is monitored in real time; the platform rear end is used for preprocessing source data to obtain a data set and presetting a plurality of model architectures, and is also used for selecting the corresponding data set and model architecture for at least one training task submitted by the platform front end, managing the training task by using an experiment and operation two-dimensional unit, and regularly storing training products, the progress and / or result of the training tasks are / is fed back according to the training products, load balancing is carried out on the multiple training tasks, and the GPU resources of the multiple training tasks are correspondingly distributed. According to the scheme, the training efficiency and stability can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning technology, and more particularly to a cloud-based deep learning training platform, method, medium, and device. Background Art

[0002] With the continuous development of artificial intelligence technology, deep learning has been widely applied in various fields, such as image recognition, speech recognition, natural language processing, etc. In the application of deep learning, the training mode is a very important link. However, the traditional deep learning model training is usually carried out in the local environment. Due to the limitations of local resources, there are technical problems such as low efficiency and poor stability. Summary of the Invention

[0003] In view of this, this application provides a cloud-based deep learning training platform, method, medium, and electronic device, mainly aiming to improve the stability and efficiency of deep learning model training.

[0004] According to one aspect of this application, a cloud-based deep learning training platform is provided. The training platform includes a platform front-end located at the user's local end and a platform back-end deployed in the cloud. Among them,

[0005] The platform front-end is used to configure deep learning training tasks and monitor the progress and / or results of the training tasks fed back by the platform back-end in real time;

[0006] The platform back-end is used to preprocess source data to obtain a data set, and pre-configure multiple model architectures. It is also used to select corresponding data sets and model architectures for at least one training task submitted by the platform front-end, manage the training tasks in terms of both experiment and operation dimensions, and regularly store training products, and feedback the progress and / or results of the training tasks according to the training products. Among them, load balancing is performed for multiple training tasks and corresponding graphics processing unit (GPU) resources are allocated respectively.

[0007] In one implementation,

[0008] The platform front-end is specifically used to submit a new training task based on the original data and label data to be trained, and set training task parameters on the task configuration page. The training task parameters include at least one of learning rate, batch size, optimizer type, and loss function.

[0009] In one implementation,

[0010] The platform front-end is also used to adjust the training task parameters according to the progress and / or results of the training tasks.

[0011] In one implementation,

[0012] The platform backend includes: a load balancing layer, a model training process management layer, and a storage layer. Among them,

[0013] The load balancing layer is used to set and provide a load balancing policy;

[0014] The model training process management layer includes multiple training process management servers. Among them, according to the load balancing policy provided by the load balancing layer, load balancing is performed on multiple training tasks among the training process management servers, and, according to the progress of each training task and / or the resource occupancy statistical results, graphics processing unit (GPU) resources are allocated to each training task;

[0015] The storage layer is used to store at least one of source data, data sets, model architecture data, and training product data.

[0016] In one implementation,

[0017] The training process management server further includes: a data preprocessing module, a training task determination module, a training production line call module, a round evaluation module, and a GPU resource management call module. Among them,

[0018] The data preprocessing module is used to preprocess the collected source data and segment the preprocessed data to obtain multiple data sets;

[0019] The training task determination module is used to call the corresponding data set and model architecture from the storage layer for the current training task;

[0020] The training production line call module is used to call the perception production line tool to make the perception production line tool run each processing process of the training task. Among them, the training task is managed in units of experiment and operation dimensions. Each experiment includes multiple operations, and each operation includes multiple training rounds;

[0021] The round evaluation module is used to evaluate the results of the training round when the training round is completed, summarize the evaluation results, and feedback the summarized evaluation results to the platform front end;

[0022] The GPU resource management call module is used to intelligently schedule the GPU cluster and allocate GPU resources according to the progress of each training task and / or the resource occupancy results.

[0023] In one implementation,

[0024] The training process management server further includes a model laboratory call module and a perception evaluation call module. Among them,

[0025] The model laboratory calling module is used to call the model laboratory tool, export the optimal training model result obtained from training to the model laboratory tool, and start the quantization evaluation process;

[0026] The perception evaluation calling module is used to call the perception evaluation tool to evaluate the performance of the model exported to the model laboratory, which is used as the basis for deploying the model.

[0027] In one implementation,

[0028] The training process management server further includes:

[0029] The training task visualization module is used to call the visualization tool to display the training task progress and / or result to the platform front end in the form of a visualization curve, where the visualization curve supports filtering, zooming, and adjustment.

[0030] According to one aspect of the present application, a cloud-based deep learning training method is provided, characterized in that the method is used for the platform to execute the deep learning training process.

[0031] According to one aspect of the present application, a storage medium is provided, in which a computer program is stored, where the computer program is set to execute the above-mentioned cloud-based deep learning training method when running.

[0032] According to one aspect of the present application, an electronic device is provided, including a memory and a processor, where a computer program is stored in the memory, and the processor is set to run the computer program to execute the above-mentioned cloud-based deep learning training method.

[0033] By means of the above technical solution, a cloud-based deep learning training platform, method, medium, and device provided by the present application transfer the traditional deep learning model training from the local to the cloud, greatly improving the work efficiency and product quality of algorithm engineers. Compared with the traditional deep learning model training, the cloud training platform of the present invention has the following advantages:

[0034] 1. High efficiency and stability: In the cloud cluster environment, computing resources can be fully utilized to improve the training speed and efficiency;

[0035] 2. Unity: Unify the training task creation method, enabling algorithm engineers in each business line to use the same training platform, improving work efficiency;

[0036] 3. Systematization: Systematic training process management and tracking can help algorithm workers better understand the training results and improve product quality;

[0037] 4. Reliability: The cloud environment has high reliability, which can ensure the stability of training tasks and data security;

[0038] 5. Scalability: The cloud environment has high scalability, which can dynamically expand computing resources according to needs to meet the requirements of different business scales;

[0039] 6. Visualization: The platform provides a visual interface, enabling algorithm engineers to more intuitively understand the training task situation and improve work efficiency;

[0040] 7. Intelligence: The platform adopts the MLOps concept, applies artificial intelligence technology to the training process, improves the training effect through means such as automation and monitoring, and can be adjusted and optimized according to the training results.

[0041] In summary, the cloud deep learning training platform of the embodiments of the present invention has many advantages such as high efficiency, stability, unity, systematicness, reliability, scalability, visualization, and intelligence, and is applicable to deep learning applications in various fields.

[0042] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the following specifically gives the specific implementation manners of this application. Brief Description of the Drawings

[0043] The drawings described herein are used to provide a further understanding of this application, and constitute a part of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0044] Figure 1 Shows a schematic diagram of a cloud-based deep learning training platform provided by an embodiment of this application;

[0045] Figure 2 Shows a schematic diagram of the platform backend structure of a cloud-based deep learning training platform provided by an embodiment of this application;

[0046] Figure 3 Shows a schematic diagram of the platform backend instance structure of a cloud-based deep learning training platform provided by an embodiment of this application;

[0047] Figure 4 Shows a flowchart of a cloud-based deep learning training method provided by an embodiment of this application. Detailed Description of the Embodiments

[0048] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application. It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.

[0049] The embodiment of the present invention provides a cloud-based deep learning training platform, aiming to provide efficient, stable, and scalable cloud services, provide systematic training services for algorithm personnel in each business line, and optimize the training resource scheduling and management capabilities.

[0050] See Figure 1 , which shows a schematic diagram of a cloud-based deep learning training platform provided by an embodiment of this application. The training platform may include a platform front-end located on the user's local side and a platform back-end located in the cloud (network side). Among them, the platform front-end can be understood as a client running on a terminal (such as a computer, mobile phone, etc.) for user interaction, including user configuration of training tasks, viewing training progress, adjusting training parameters, etc.; the platform back-end refers to the server side that provides data support to the client, and realizes functions including data storage, data processing and calculation, resource allocation, process management, tool invocation, etc.

[0051] The cloud-based deep learning training platform provided by the embodiment of this application has a training platform including a platform front-end located on the user's local side and a platform back-end deployed in the cloud.

[0052] The platform front-end is used to configure deep learning training tasks and monitor the progress and / or results of the training tasks fed back by the platform back-end in real time. For example, the platform front-end is specifically used to submit a new training task based on the original data and label data to be trained, and set training task parameters on the task configuration page. The training task parameters include at least one of the learning rate, batch size, optimizer type, and loss function. In addition, the platform front-end can also be used to adjust the training task parameters according to the progress and / or results of the training tasks.

[0053] The platform back-end is used to preprocess the source data to obtain a data set, and pre-sets multiple model architectures. It is also used to select the corresponding data set and model architecture for at least one training task submitted by the platform front-end, manage the training tasks in units of experiment and operation dimensions, and regularly store the training products, and feedback the progress and / or results of the training tasks according to the training products, where load balancing is performed on multiple training tasks and the corresponding GPU (graphics processing unit) resources are allocated to each.

[0054] For example, the platform backend implements functions such as data storage, training task process management, and load balancing. In one implementation, the platform backend includes: a load balancing layer, a model training process management layer, and a storage layer.

[0055] The load balancing layer is used to set and provide load balancing policies. Load Balancing is a technology that processes requests and distributes them to multiple servers, which is used to improve the ability to handle a large number of user accesses. Its goal is to achieve the best utilization of resources, maximize throughput, minimize response time, and avoid overloading any single server. Load balancing can be applied at different levels, such as the DNS level, the hardware level (e.g., using a load balancer), or the software level (e.g., using a proxy server). There are many load balancing algorithms, including Round Robin, Weighted Round Robin, Least Connection, and Source Address Hashing, etc. Selecting an appropriate load balancing policy according to requirements and scenarios can effectively improve system performance and reliability. In the embodiments of the present application, the model training process management layer includes multiple training process management servers. Therefore, according to the load balancing policy set by the load balancing layer, load balancing is performed among multiple training process management servers for multiple training tasks, so as to meet the requirements of multiple training tasks being performed simultaneously.

[0056] The model training process management layer includes multiple training process management servers, where, according to the load balancing policy provided by the load balancing layer, load balancing is performed among multiple training tasks among the respective training process management servers, and, according to the progress and / or resource occupancy statistical results of each training task, GPU resources are allocated to each training task. Machine learning requires a large amount of computing resources to process data and perform model training. GPUs are designed specifically for executing parallel computing tasks. GPUs have a large number of processing units (CUDA cores), which can process multiple computing tasks simultaneously. Therefore, GPUs are more excellent in parallel computing, which enables GPUs to execute machine learning tasks faster, and the advantages of GPUs are more obvious when dealing with large-scale data sets and complex models. In the embodiments of the present application, in order to meet the parallel execution of multiple training tasks, it is necessary to allocate GPU resources for each training task. Specifically, GPU resources can be allocated and adjusted for each training task according to the progress of each training task and the resource occupancy situation.

[0057] A storage layer for storing at least one of source data, data sets, model architecture data, and training product data. In specific implementation, a relational database (such as Mysql) or object storage can be used to store data.

[0058] It can be seen that the cloud-based deep learning training platform provided by the embodiments of the present application can make full use of the advantages of rich cloud network resources, schedule and allocate resources for multiple training tasks at the same time. In the cloud cluster environment, computing resources can be fully utilized to improve the training speed and efficiency. In addition, the unified training task creation method enables algorithm staff in each business line to use the same training platform, improving work efficiency. Moreover, due to the high reliability of the cloud environment, the stability of training tasks and data security can be guaranteed. In addition, the cloud environment has high scalability and can dynamically expand computing resources according to needs to meet the requirements of different business volumes.

[0059] See Figure 2 , which shows a schematic diagram of the platform backend structure of a cloud-based deep learning training platform provided by the embodiments of the present application.

[0060] In each training process management server, it may further include: a data preprocessing module, a training task determination module, a training production line call module, a round evaluation module, and a GPU resource management call module.

[0061] The data preprocessing module is used to preprocess the collected source data and segment the preprocessed data to obtain multiple data sets. In actual operation, first, source data needs to be collected. For example, a large number of high-resolution road images are collected from different sources (such as cameras, lidar, etc.). Among them, it is necessary to ensure that the data has diversity, covering different weather conditions (sunny, rainy, snowy), lighting conditions (daytime, night), locations (city, countryside), and traffic environments (congested, unobstructed). Then, preprocessing operations are performed on the data, including the following processing steps: (1) Cropping: Crop the image area that is meaningful for lane line detection according to requirements; (2) Scaling: Unify the image size to the input size required by the model; (3) Grayscale / Color: Select grayscale or color images as input according to the model requirements; (4) Normalization: Normalize the pixel value range in the image to between 0-1 or -1 to 1; (5) Data augmentation: Apply methods such as rotation, translation, and flipping to generate new training samples to improve the generalization ability of the model; (6) Segment the data set: Divide the entire data set into a training set, a validation set, and a test set according to a certain proportion.

[0062] The training task determination module is used to call the corresponding dataset and model architecture from the storage layer for the current training task. Before creating a training task, lane line annotation can be performed on each image to be processed, and a binary image can be created to represent whether each pixel belongs to a lane line. For example, semi-automatic or fully automatic methods can be used for annotation, such as semantic segmentation models like DeepLab. Then, the label data (binary image) and the original image are uploaded to the cloud together. When creating a training task, the user can log in to the front end of the training platform and submit a new training task through the front-end interface or API. Among them, relevant parameters such as learning rate, batch size, optimizer type, loss function, etc. are set on the task configuration page. Then, the cloud matches a suitable preprocessed dataset, associates it with the label data, and matches a model architecture suitable for the lane line detection task, such as U-Net.

[0063] The training production line call module is used to call the perception production line tool (an external tool not shown in the figure), so that the perception production line tool runs each processing flow of the training task. Among them, the training task is managed in units of experiments and runs in two dimensions. Each experiment includes multiple runs, and each run includes multiple training epochs. In the embodiments of the present application, the training task is managed in units of experiments (Experiment) and runs (Run) in multiple dimensions. Each experiment contains multiple runs, and each run contains multiple epochs. Here, an epoch refers to the number of times the training set is iteratively trained by the entire network. For example, epoch = 10 means that the entire dataset is trained 10 times. That is, during the training process for a certain training task, the most optimal training results can be analyzed and determined by increasing the number of training epochs, thereby improving the training effect. Figure 2 The epoch evaluation module is used to evaluate the results of the training epoch when the training epoch is completed, summarize the evaluation results, and feedback the summarized evaluation results to the platform front end. After each epoch ends, the platform will automatically conduct an evaluation and summarize the evaluation results with the previous evaluation results. The high-density evaluation and statistical analysis at the epoch level can submit an evaluation once every epoch is completed. The evaluation can gather all the epoch evaluation results of an experiment and conduct aggregated statistical analysis. The platform provides an aggregated statistical analysis function to help users understand the trends and problems during the training process.

[0064]

[0065] ​The GPU resource management and invocation module is used to intelligently schedule the GPU cluster and allocate GPU resources according to the progress and / or resource occupancy results of each training task. In a specific implementation, a cluster management and job scheduling system can be used to invoke the GPU cluster, where the GPU cluster includes multiple GPU servers. A GPU server is a high-performance computer specifically designed to handle graphics and compute-intensive tasks. It is equipped with one or more powerful graphics processing units (GPUs), enabling it to quickly process large amounts of data and complex operations in parallel. GPU servers are applied in artificial intelligence and deep learning (neural network models require a large amount of computing resources for training and inference, and GPUs provide high parallelism to accelerate these processes), graphics rendering and animation (in 3D modeling, rendering, and animation production, GPU servers can efficiently process large amounts of graphic data), and scientific simulation and data analysis (for large-scale data analysis and simulation in fields such as experimental physics, bioinformatics, and meteorology, GPU servers have superior computing performance). Using GPU servers can significantly improve the performance of compute-intensive tasks. For example, the platform automatically allocates GPU resources to ensure that each training task can be carried out efficiently. The platform intelligently schedules resources according to the task progress and resource occupancy to improve the training speed and efficiency.

[0066] In one implementation, the training process management server may further include a model laboratory invocation module and a perception evaluation invocation module.

[0067] The model laboratory invocation module is used to invoke the model laboratory tool and export the optimal training model result obtained from training to the model laboratory tool to initiate the quantization evaluation process. For example, export the best model obtained from training to ModelLab (the model laboratory tool) to initiate the quantization, inference, and evaluation processes.

[0068] The perception evaluation invocation module is used to invoke the perception evaluation tool to evaluate the performance of the model exported to the model laboratory, which serves as the basis for deploying the model. For example, interface with the perception evaluation tool to achieve an end-to-end fully automated training evaluation process. After verifying that the model performance meets the standards, deploy the model to the autonomous driving system for lane line detection.

[0069] In one implementation, the training process management server may further include a training task visualization module, which is used to invoke the visualization tool to display the training task progress and / or results to the platform front end in the form of a visualization curve, where the visualization curve supports filtering, zooming, and adjustment. For example, use the platform's visualization tool to view the key task metrics in real time, such as the Loss curve, Accuracy curve, etc. Users can filter, zoom, and adjust the curve according to their needs.

[0070] See Figure 3, which shows the schematic structural diagram of the platform backend instance of a cloud-based deep learning training platform provided by an embodiment of the present application.

[0071] In this example, Loadbalancer corresponds to the load balancing layer, each training process management server in the model training process management layer is implemented by MLFlow server, and Storage Layer corresponds to the storage layer. Among them, the model training process management layer calls the GPU training cluster through a cluster management and job scheduling system (Slurm Proxy). The GPU training cluster includes multiple GPU servers. Moreover, the model training process management layer produces a toolset to implement functions such as training, evaluation, and visualization. For example, the toolset includes tools such as PPL (Perception Production Line Tool), PEP (Perception Evaluation Tool), and MoedelLab (Model Laboratory).

[0072] For example, PPL (Perception Production Line), as a perception production line tool, has the following functions: managing computing power and storage resources, automatically running various processing processes required for deep learning, providing solutions for data storage, providing a maintenance framework for deep learning processes, implementing process calling and publishing capabilities, standardizing the review mechanism for ground truth data and processing processes, and promoting teamwork and organization in deep learning tasks. PEP (Perception Evaluation Platform) is a set of evaluation frameworks provided to better implement training. Through abstraction and standardization, PEP unifies the evaluation processes of various algorithms, and at the same time provides a unified computing and visualization framework, with the ability to submit and execute automatically with one key, effectively supporting the delivery and use of various algorithms.

[0073] MLFlow is an open-source machine learning platform that can cover the entire process of machine learning (from data preparation to model training). It can be understood that MLFlow is a tool for managing machine learning workflows, and its core functions include: managing experiment parameters, code, and results in the form of APIs and comparing them in the form of a UI; packaging code; deploying models, managing models, and registering models.

[0074] See Figure 4 , which shows the flowchart of a cloud-based deep learning training method provided by an embodiment of the present application. This method is described by taking the lane line detection task of the cloud-based deep learning training platform as an example. The following are the specific steps of data processing and deep learning training:

[0075] 1. Data collection:

[0076] Collect a large number of high - resolution road images from different sources (such as cameras, lidar, etc.). Ensure that the data is diverse, covering different weather conditions (sunny, rainy, snowy), lighting conditions (daytime, nighttime), locations (urban, rural), and traffic environments (congested, unobstructed).

[0077] 2. Data pre - processing:

[0078] Cropping: Crop out the image regions that are meaningful for lane line detection according to requirements.

[0079] Scaling: Unify the image size to the input size required by the model.

[0080] Grayscale / Color: Select grayscale or color images as inputs according to the model requirements.

[0081] Normalization: Normalize the pixel value range in the image to between 0 - 1 or - 1 to 1.

[0082] Data augmentation: Apply methods such as rotation, translation, and flipping to generate new training samples to improve the generalization ability of the model.

[0083] Split the dataset: Divide the entire dataset into training set, validation set, and test set according to a certain proportion.

[0084] 3. Annotation:

[0085] Perform lane line annotation on each image, creating a binary image to indicate whether each pixel belongs to a lane line. Semi - automatic or fully automatic methods can be used for annotation, such as semantic segmentation models like DeepLab.

[0086] Upload the label data (binary image) and the original image to the cloud platform together.

[0087] 4. Create a training task:

[0088] Log in to the cloud deep learning training platform.

[0089] Submit a new training task through the front - end interface or API.

[0090] Set relevant parameters on the task configuration page, such as learning rate, batch size, optimizer type, loss function, etc.

[0091] Select the appropriate pre - processed dataset and associate it with the label data.

[0092] Select a model architecture suitable for the lane line detection task, such as U - Net.

[0093] 5. Training process management and tracking:

[0094] The platform manages training tasks in units of Experiment and Run. Each experiment contains multiple runs, and each run contains multiple epochs.

[0095] The platform will automatically save the model weights, logs, and other product information regularly. Users can also save this information manually.

[0096] The platform provides a monitoring function to view the training task progress, performance, and resource occupancy in real time. Users can adjust the tasks based on the monitoring information.

[0097] 6. Visualization of training tasks:

[0098] Use the platform's visualization tools to view the key metrics of the task in real time, such as the Loss curve, Accuracy curve, etc.

[0099] Users can filter, zoom, and adjust the curves according to their needs.

[0100] 7. Evaluation and statistical analysis at the epoch level:

[0101] After each epoch ends, the platform will automatically conduct an evaluation and summarize the evaluation results with the previous ones.

[0102] The platform provides an aggregation statistical analysis function to help users understand the trends and problems during the training process.

[0103] 8. Model management and deployment:

[0104] Export the best model obtained from training to ModelLab and start the quantization, inference, and evaluation processes.

[0105] Connect to tools such as PPL and PEP to achieve an end-to-end fully automated training and evaluation process.

[0106] After verifying that the model performance meets the standards, deploy the model to the autonomous driving system for lane line detection.

[0107] 9. Resource scheduling and optimization:

[0108] The platform automatically allocates GPU resources to ensure that each training task can be carried out efficiently.

[0109] The platform intelligently schedules resources according to the task progress and resource occupancy to improve the training speed and efficiency.

[0110] It can be seen that the cloud-based deep learning training platform, method, medium, and device proposed in the embodiments of the present invention transfer traditional deep learning model training from local to the cloud, greatly improving the work efficiency and product quality of algorithm engineers. Compared with traditional deep learning model training, the cloud training platform of the present invention has the following advantages:

[0111] 1. High efficiency and stability: In the cloud cluster environment, computing resources can be fully utilized to improve the training speed and efficiency;

[0112] 2. Unity: Unify the training task creation method, enabling algorithm engineers in each business line to use the same training platform, improving work efficiency;

[0113] 3. Systematization: Systematic training process management and tracking can help algorithm staff better understand the training results and improve product quality;

[0114] 4. Reliability: The cloud environment has high reliability, which can ensure the stability of training tasks and data security;

[0115] 5. Scalability: The cloud environment has high scalability, and computing resources can be dynamically expanded according to needs to meet the requirements of different business scales;

[0116] 6. Visualization: The platform provides a visualization interface, enabling algorithm engineers to more intuitively understand the training task situation and improve work efficiency;

[0117] 7. Intelligence: The platform adopts the MLOps concept, applies artificial intelligence technology to the training process, improves the training effect through means such as automation and monitoring, and can be adjusted and optimized according to the training results.

[0118] In summary, the cloud deep learning training platform of the embodiments of the present invention has many advantages such as high efficiency and stability, unity, systematization, reliability, scalability, visualization, and intelligence, and is applicable to deep learning applications in various fields.

[0119] An embodiment of the present application also provides a storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0120] Optionally, in this embodiment, the above storage medium may be configured to store a computer program for executing the following steps:

[0121] Configure a deep learning training task and monitor the progress and / or results of the training task fed back by the backend of the platform in real time;

[0122] Preprocess the source data to obtain a dataset, and pre-set multiple model architectures. It is also used to select the corresponding dataset and model architecture for at least one training task submitted by the front end of the platform, manage the training tasks in terms of both experiment and operation dimensions, and regularly store the training products, and feedback the progress and / or results of the training tasks according to the training products. Among them, load balancing is performed for multiple training tasks and the corresponding graphics processing unit (GPU) resources are allocated respectively.

[0123] Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store computer programs.

[0124] An embodiment of the present application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0125] Optionally, the above electronic device may further include a transmission device and input / output devices. Among them, the transmission device is connected to the above processor, and the input / output devices are connected to the above processor.

[0126] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:

[0127] Configure deep learning training tasks, and monitor the progress and / or results of the training tasks fed back by the back end of the platform in real time;

[0128] Preprocess the source data to obtain a dataset, and pre-set multiple model architectures. It is also used to select the corresponding dataset and model architecture for at least one training task submitted by the front end of the platform, manage the training tasks in terms of both experiment and operation dimensions, and regularly store the training products, and feedback the progress and / or results of the training tasks according to the training products. Among them, load balancing is performed for multiple training tasks and the corresponding graphics processing unit (GPU) resources are allocated respectively.

[0129] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.

[0130] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0131] In the above embodiments of the present application, the descriptions of the respective embodiments each have their own emphasis. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0132] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0133] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0134] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0135] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks or optical discs and other various media that can store program codes.

[0136] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A cloud-based deep learning training platform, characterized in that, The training platform includes a platform front-end located on the user's local side and a platform back-end deployed in the cloud. Among them, the platform front-end is used to configure deep learning training tasks and monitor the progress and / or results of the training tasks fed back by the platform back-end in real time; the platform back-end is used to preprocess source data to obtain a data set, and pre-configures multiple model architectures. It is also used to select corresponding data sets and model architectures for at least one training task submitted by the platform front-end, manage the training tasks in terms of both experiment and operation dimensions, and regularly store training products and feedback the progress and / or results of the training tasks according to the training products. Among them, load balancing is performed on multiple training tasks and corresponding graphics processing unit (GPU) resources are allocated respectively.

2. The platform according to claim 1, wherein the platform front-end is specifically used to submit a new training task based on the raw data and label data to be trained, and set training task parameters on the task configuration page. The training task parameters include at least one of learning rate, batch size, optimizer type, and loss function.

3. The platform according to claim 2, wherein the platform front-end is further used to adjust the training task parameters according to the progress and / or results of the training tasks.

4. The platform according to claim 1, characterized in that The platform back-end includes: a load balancing layer, a model training process management layer, and a storage layer. Among them, the load balancing layer is used to set and provide a load balancing strategy; the model training process management layer includes multiple training process management servers. Among them, according to the load balancing strategy provided by the load balancing layer, load balancing is performed on multiple training tasks among the training process management servers, and graphics processing unit (GPU) resources are allocated to each training task according to the progress and / or resource occupancy statistical results of each training task; the storage layer is used to store at least one of source data, data sets, model architecture data, and training product data.

5. The platform according to claim 4, characterized in that The training process management server further includes: a data preprocessing module, a training task determination module, a training production line call module, a round evaluation module, and a GPU resource management call module. Among them, the data preprocessing module is used to preprocess the collected source data and split the preprocessed data to obtain multiple data sets; the training task determination module is used to call the corresponding data set and model architecture from the storage layer for the current training task; the training production line call module is used to call the perception production line tool to make the perception production line tool run each processing process of the training task. Among them, the training tasks are managed in terms of both experiment and operation dimensions. Each experiment includes multiple runs, and each run includes multiple training rounds; the round evaluation module is used to evaluate the results of the training round when the training round is completed, summarize the evaluation results, and feedback the summarized evaluation results to the platform front-end; the GPU resource management call module is used to intelligently schedule the GPU cluster and allocate GPU resources according to the progress and / or resource occupancy results of each training task.

6. The platform according to claim 5, characterized in that The training process management server further includes a model laboratory calling module and a perception evaluation calling module, wherein, the model laboratory calling module is configured to call model laboratory tools, export the optimal training model result obtained from training to the model laboratory tools, so as to initiate a quantization evaluation process; the perception evaluation calling module is configured to call perception evaluation tools to evaluate the performance of the model exported to the model laboratory, so as to be used as a basis for deploying the model.

7. The platform according to claim 6, characterized in that, The training process management server further includes: a training task visualization module, configured to call visualization tools to display the training task progress and / or results to the front end of the platform in the form of a visualization curve, wherein the visualization curve supports filtering, zooming, and adjustment.

8. A cloud-based deep learning training method, characterized in that, The method is applied to the platform according to any one of claims 1-7 and is used to execute the deep learning training process.

9. A storage medium, characterized in that, A computer program is stored in the storage medium, wherein the computer program is configured to execute the method according to claim 8 when running.

10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method according to claim 8.