Method for automatically optimizing running resources of deep learning image classification model
By automatically optimizing the resource allocation of deep learning image classification model, the training efficiency and resource waste caused by unreasonable allocation of computing resources are solved, and more efficient and stable model training and higher resource utilization are achieved.
Patent Information
- Application Number
- CN202411943280.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-06-03
AI Technical Summary
In deep learning image classification tasks, unreasonable allocation of computing resources leads to inefficient training efficiency, waste or insufficient resources, especially in the case of tight computing resources, uneven resource allocation may occur.
A method for automatically optimizing the operation resources of the deep learning image classification model is proposed. It judges whether to initialize parameters when the system is started, obtains the number of resources running in the system, and automatically initializes the optimization strategy based on the number of resources obtained. During multiple runs, the number of resources is dynamically adjusted according to the changes in operation speed, and the adaptive algorithm ensures maximum resource utilization efficiency.
By automatically optimizing resource allocation, the efficiency and stability of deep learning model training are significantly improved, the problems of resource waste and insufficient resources are reduced, resource utilization is improved, and the threshold for users to use the system is lowered.
Smart Images

Figure CN120088621A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer information technology, and mainly relates to a method for automatically optimizing the operating resources of a deep learning image classification model. Background Art
[0002] In the deep learning image classification task, the training process of the model is often restricted by computing resources (such as CPU, GPU, memory, etc.). For example, when the available computing resources are insufficient, the training may become very slow, or when there are too many resources, it may lead to resource waste.
[0003] Deep learning models often require a large amount of memory during training, especially for tasks such as image classification where the dataset may be very large. Without reasonable memory management, it may lead to memory overflow or system lag, especially in an environment with limited memory.
[0004] If there is no reasonable resource allocation during the model training process, resource waste may occur (such as using too much CPU computing instead of GPU), thus affecting the overall training efficiency.
[0005] In some cases where computing resources are relatively tight, the phenomenon of uneven resource allocation may occur. For example, the CPU or GPU may be overloaded in some stages while idle in other stages, resulting in underutilization of computing resources.
[0006] Without automated resource adjustment, the training process may be inefficient due to inappropriate manually set parameters. For example, choosing an inappropriate batch size or data loading method may lead to slower training speed or memory overflow. Summary of the Invention
[0007] In view of the above deficiencies, the present invention proposes a method for automatically optimizing the operating resources of a deep learning image classification model. The method for automatically optimizing the operating resources of a deep learning image classification model can save cluster computing resources, improve the utilization rate of cluster computing resources, and at the same time reduce the threshold for users to use the system. The specific steps are as follows:
[0008] S1. When the system starts, determine whether to initialize the parameters. If not, obtain the number of resources used by the system.
[0009] S2. In the case where the system is not initialized, the system automatically initializes the optimization strategy according to the number of resources obtained in step S1.
[0010] S3. When the deep learning image classification model runs for the first time, set the operating resources of the model to the number of resources available during the first run in step S2, and run the model.
[0011] S4. The deep learning image classification model runs for the second time. Define the optimization method for this run as increasing resources. Add the resource increment in the optimization strategy to the resource quantity in the first run to obtain the resource quantity for this run of the model, and then run the model;
[0012] S5. The deep learning image classification model runs for the (2 + N)th time, where N is greater than or equal to 1. If increasing the resource quantity promotes speed improvement, it is considered that the optimization is effective, indicating that there is room for further optimization. Continue to increase the running resources to improve the running speed. Otherwise, reduce the resources. If the model running speed remains unchanged when reducing the resource quantity, it indicates that there is surplus running resources and the running resource quantity needs to be further reduced. If the running speed decreases, it indicates that the resources are insufficient and the resource quantity needs to be increased.
[0013] Furthermore, the specific configurations initialized in the automatic initialization optimization strategy include:
[0014] S2.1. Whether to turn on the switch of the optimization strategy;
[0015] S2.2. Configure the available resource quantity for the first run of the model;
[0016] S2.3. Configure the threshold for whether the model running speed changes. Optimization is only considered effective when the time of two runs of the model exceeds the threshold;
[0017] S2.4. Configure the resource increment or decrement for each optimization;
[0018] S2.5. Configure the maximum and minimum resource quantities that can be allocated for the model to run;
[0019] S2.6. Set to optimize only the selected models, and record the numbers of the selected models as an array.
[0020] Furthermore, the step S5 specifically further includes:
[0021] S5.1. If the previous model run is to increase the resource quantity and the running speed increases, it indicates that the optimization method of increasing the running resources is effective, and the resource quantity for this run of the model continues to increase;
[0022] S5.2. If the previous model run is to increase the resource quantity and the running speed slows down or remains unchanged, it indicates that the optimization method of increasing the running resources is ineffective, and the resource quantity for this run is reduced;
[0023] S5.3. If the previous model run is to reduce the resources and the running speed increases or remains unchanged, it indicates that the optimization method of reducing the running resources is effective, and the resource quantity for this run continues to be reduced;
[0024] S5.4. If the previous model run fails, regardless of the existing resource quantity, the resource quantity for this run will be increased.
[0025] Furthermore, when the system determines whether to initialize parameters during startup, the system also initializes the monitoring function for resource utilization, including monitoring the CPU utilization rate, memory usage, disk I / O, and network bandwidth of the current system, to ensure dynamic scheduling in different environments.
[0026] Furthermore, the resource quantity is dynamically adjusted according to the change in the model running speed. Among them, an adaptive algorithm is introduced to dynamically adjust resource allocation to ensure the stability of the system under different loads.
[0027] Furthermore, in the case where the previous model run fails, the system rolls back the model to a more stable configuration through an adaptive fault tolerance mechanism and continues with the optimization strategy and model run.
[0028] Furthermore, during the model running and training process, a deep reinforcement learning algorithm is introduced to enable the optimization strategy to perform autonomous learning by an intelligent agent; combined with the training process of the deep learning model, the hyperparameter scheduling during the model running is optimized, and the computing resources are dynamically adjusted based on the current training state of the model.
[0029] According to the second aspect of the present invention, a computer program product is provided, on which one or more computer programs are stored. When the one or more computer programs are executed by a computer processor, the above method is implemented.
[0030] One or more of the above technical solutions in the embodiments of the present application have at least one of the following technical effects:
[0031] 1) The resource utilization rate is improved: The training strategy is dynamically adjusted according to the available computing resources, avoiding resource waste or insufficiency.
[0032] 2) The memory and computing efficiency are improved: By optimizing the data loading method and batch size, the memory management and computing efficiency during the training process are enhanced.
[0033] 3) The training speed and stability are optimized: Dynamically adjusting the resource configuration can avoid overload problems and improve the training speed and stability of the model.
[0034] 4) Automatic adaptation: There is no need to manually adjust the resource configuration. The system automatically optimizes according to the actual situation, improving the convenience of the training process.
[0035] In summary, through more intelligent resource scheduling and adaptive optimization, the present invention significantly improves the efficiency, stability, and resource utilization rate of deep learning model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated in and constitute a part of this specification. The drawings illustrate the embodiments and, together with the description, are used to explain the principles of the invention. Other embodiments and many of the intended advantages of the embodiments will be readily apparent as they become better understood by reference to the following detailed description. The elements of the drawings are not necessarily to scale with each other. Like reference numerals refer to corresponding like parts.
[0037] Figure 1 The flowchart shows a method for automatically optimizing the operating resources of a deep learning image classification model according to an embodiment of the present invention.
[0038] Figure 2 It is a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed Embodiments
[0039] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the related invention and not for limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the invention are shown in the drawings.
[0040] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.
[0041] Figure 1 The flowchart shows a method for automatically optimizing the operating resources of a deep learning image classification model according to an embodiment of the present invention, as Figure 1 shown below:
[0042] In this embodiment, during the model training of a specific deep learning image classification model, the resource allocation of the training process is automatically adjusted according to system resources (such as CPU, GPU, memory), so as to optimize the training speed. Specifically, it includes:
[0043] S1. When the system starts, determine whether to initialize the parameters. If not, obtain the number of resources running in the system. The core code implementation is as follows:
[0044]
[0045]
[0046] In addition to obtaining the number of system resources, a monitoring function for resource usage can be added. For example, monitor the CPU usage rate, memory usage, disk I / O, network bandwidth, etc. of the current system to ensure dynamic scheduling in different environments.
[0047] S2. When the system is not initialized and configured, the system automatically initializes the policy according to the number of resources obtained in S1.
[0048] S2.1. The initial configuration includes a switch for enabling the optimization policy.
[0049] S2.2. Configure the number of resources available when the configuration model runs for the first time.
[0050] S2.3. Set the threshold for whether the running speed of the configuration model changes. Only when the time between two runs of the model exceeds the threshold is the optimization considered effective. When the optimization policy determines that the model needs to increase resources, there needs to be a judgment criterion. Only when the predetermined criterion is met is it considered that optimization is required. This configuration is used as the judgment criterion.
[0051] S2.4. Configure the number of resources increased or decreased during each optimization.
[0052] S2.5. Configure the maximum and minimum number of resources that can be allocated during the operation of the configuration model.
[0053] S2.6. Set to optimize only specific models. This setting is an array that records the numbers of specific models.
[0054] In this embodiment, the GPU quantity is used to select the initial batch size, or the data loading method is determined according to the memory size. The core implementation code is as follows:
[0055]
[0056]
[0057] S3. When the deep learning image classification model runs for the first time, set the running resources of the model to the number of resources available during the first run in S2, and run the model.
[0058] Not only the initial resource configuration can be set, but also the management of the resource pool can be realized. That is, different types of computing resources (such as CPU, GPU, memory, etc.) are regarded as independent resource pools, and the resources in the resource pool are flexibly scheduled according to the needs of tasks, rather than just a fixed total number of resources.
[0059] Before the initial run, the evaluation can evaluate the preliminary requirements of the model through some preprocessing stages (such as a quick benchmark test), and adjust the resource configuration during the initial run according to these results to avoid incorrect resource allocation too early.
[0060] S4. The deep learning image classification model runs for the second time. Define the optimization method for this run as increasing the number of resources. Based on the resources used in the first run, add the number of increased resources set in the optimization strategy to obtain the resources for this run of the model, and then run the model. The core code implementation is as follows:
[0061]
[0062]
[0063] Dynamically adjust the resource quantity according to the change in the model running speed to ensure the maximization of resource utilization efficiency. At the same time, introduce an adaptive algorithm, such as a PID controller, to dynamically adjust resource allocation to ensure the stability of the system under different loads.
[0064] In addition to simply increasing resources, an incremental adjustment strategy can be used. Based on the resource utilization rate and running duration of the model in the first run, infer the optimization increment. For example, after the first run, the system can decide which resource should be increased according to the bottleneck situation of the CPU or memory, and the amount of resources increased each time can be a dynamically adjusted parameter.
[0065] With the help of prediction algorithms (such as a load prediction model or an adaptive optimization algorithm), estimate the resource requirements before each run and dynamically adjust resource allocation accordingly. For example, use time series analysis methods to predict the load changes in the future for a period of time.
[0066] S5. The deep learning image classification model runs for the (2 + N)th time, where N is greater than or equal to 1. If increasing resources promotes an increase in the running speed, it is considered that the optimization is effective and there is still room for further optimization. Continue to increase the running resources to see if the model running speed can be further improved. Otherwise, reduce the resources. If reducing the resource quantity does not change the model running speed, it means that there is a surplus of running resources and the running resources can be further reduced. If the running speed drops, it indicates that the resources are insufficient and the resource quantity needs to be increased.
[0067] During the model running process, the task may fail. Since the model running failure may be caused by insufficient resources, in the case of the previous model running failure, this optimization is processed by increasing resources this time. If the resource quantity reaches the upper limit or the lower limit, the model running resource quantity will no longer be increased or decreased. The specific core implementation code is as follows:
[0068]
[0069]
[0070] S5.1. If the previous model run was to increase the resource quantity and the running speed increased, it indicates that the optimization method of increasing the running resources is effective. This run continues to increase the running resource quantity.
[0071] S5.2. If the previous model run was to increase resources and the running speed became slower or remained unchanged, it indicates that the optimization method of increasing running resources is ineffective, and the number of running resources for this run is reduced;
[0072] S5.3. If the previous model run was to reduce resources and the running speed became faster or remained unchanged, it indicates that the optimization method of reducing running resources is effective, and the number of running resources for this run continues to be reduced;
[0073] S5.4. If the previous model run was to reduce resources and the running speed became slower, it indicates that the optimization method of reducing running resources is ineffective, and the number of running resources for this run is increased;
[0074] S5.5. If the previous model run failed, regardless of whether resources were increased or decreased, considering that the failure might be caused by insufficient resources, the number of running resources for this run is increased in case of model run failure. If resource adjustment causes the model to error or fail, the system can roll back to the previous more stable configuration and re-evaluate the current optimization strategy.
[0075] The deep reinforcement learning algorithm is introduced, making the optimization strategy not only depend on existing rules but also be able to perform autonomous learning through intelligent agents. The system can adjust the resource allocation strategy through continuous attempts and feedback, so as to achieve more efficient resource management and utilization. At the same time, in combination with the training process of the deep learning model, its hyperparameter scheduling is optimized, and the computing resources are dynamically adjusted based on the training state of the current model (such as learning rate, gradient, etc.).
[0076] In summary, the present invention can not only perform resource allocation and optimization more intelligently, but also improve the robustness, flexibility and performance of the system, ensuring efficient operation in various situations.
[0077] Next, refer to Figure 2 , which shows a schematic structural diagram of a computer system 200 of an electronic device suitable for implementing the embodiments of the present application. Figure 2 The shown electronic device is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present application.
[0078] As Figure 2As shown, computer system 200 includes a central processing unit (CPU) 201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 202 or a program loaded from a storage section 208 into a random access memory (RAM) 203. In the RAM 203, various programs and data required for the operation of the system 200 are also stored. The CPU 201, ROM 202, and RAM 203 are connected to each other via a bus 204. An input / output (I / O) interface 205 is also connected to the bus 204.
[0079] The following components are connected to the I / O interface 205: an input section 206 including a keyboard, a mouse, etc.; an output section 207 including, for example, a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 208 including a hard disk, etc.; and a communication section 209 including a network interface card such as a LAN card, a modem, etc. The communication section 209 performs communication processing via a network such as the Internet. A drive 210 is also connected to the I / O interface 205 as needed. A removable medium 211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 210 as needed so that a computer program read therefrom is installed into the storage section 208 as needed.
[0080] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 209 and / or installed from the removable medium 211. When the computer program is executed by the central processing unit (CPU) 201, the above-described functions defined in the method of the present application are performed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0081] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0082] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0083] The modules described in the embodiments of this application can be implemented in software or in hardware.
[0084] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: obtain the number of resources for system operation if the parameters are not initialized when the system starts; the system automatically initializes the optimization strategy according to the obtained number of resources; when the model runs for the first time, set the number of operating resources of the model and run the model; when the model runs for the second time, define the optimization method for this run as increasing resources, add the number of resource increases to the number of resources in the first run to obtain the number of resources for running the model this time, and run the model; when the model runs for the 2+Nth time, if increasing the number of resources promotes the speed increase, it is considered that the optimization is effective, continue to increase the operating resources to improve the running speed, otherwise reduce the resources; if the number of resources is reduced and the running speed of the model remains unchanged, the number of operating resources needs to be further reduced; if the running speed decreases, the number of resources needs to be increased.
[0085] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principle. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present application.
Claims
1. A method for automatically optimizing the operating resources of a deep learning image classification model, characterized in that: include: S1. When the system starts, determine whether to initialize parameters. If not, obtain the number of resources running the system; S2. In the case where the system has no initial configuration, the system automatically initializes the optimization strategy according to the number of resources obtained in step S1; S3, the deep learning image classification model is run for the first time, the running resources of the model are set to the number of resources available during the first run in step S2, and the model is run; S4, the deep learning image classification model is run for the second time, and the optimization method for this operation is defined as increasing resources. The number of resources increased in the optimization strategy is added to the number of resources for the first operation to obtain the number of resources for this operation of the model, and the model is run; S5. The deep learning image classification model runs for the 2+Nth time, where N is greater than or equal to 1. If increasing the number of resources leads to a speed increase, the optimization is considered effective, indicating that there is room for further optimization. Continuing to increase the running resources will increase the running speed, otherwise reducing the resources will increase the running speed. If the model running speed remains unchanged when the number of resources is reduced, it indicates that there are surplus running resources and the number of running resources needs to be further reduced. If the running speed decreases, it indicates that the resources are insufficient and the number of resources needs to be increased.
2. The method according to claim 1, characterized in that The configuration initialized in the automatic initialization optimization strategy specifically includes: S2.1, whether to turn on the optimization strategy switch; S2.2, the number of resources available when the configuration model is run for the first time; S2.
3. Configure the threshold for whether the model running speed changes. The optimization is considered effective only when the model runs twice more than the threshold. S2.4, configure the number of resources to be increased or decreased during each optimization; S2.
5. Configure the maximum and minimum number of resources that can be allocated for model operation; S2.
6. Set to optimize only the selected model, and the setting records the number of the selected model as an array.
3. The method according to claim 1, characterized in that The step S5 specifically includes: S5.
1. If the previous model run increased the number of resources and the running speed was faster, it means that the optimization method of increasing the running resources is effective, and the number of running resources will continue to increase in this model run; S5.
2. If the previous model run increased the number of resources and the running speed slowed down or remained unchanged, it indicates that the optimization method of increasing the running resources is invalid, and the number of running resources is reduced this time; S5.
3. If the previous model run reduced resources and the running speed increased or remained unchanged, it indicates that the optimization method of reducing running resources is effective, and the number of resources will continue to be reduced in this run; S5.
4. If the previous model run failed, the existing resource quantity will not be considered and the current run will increase the number of running resources.
4. The method according to claim 1, characterized in that: When the system starts up and determines whether to initialize parameters, the system also initializes the monitoring function of resource utilization, including monitoring the CPU utilization, memory usage, disk IO and network bandwidth of the current system to ensure dynamic scheduling in different environments.
5. The method according to claim 1, characterized in that The number of resources is dynamically adjusted according to the changes in the model's running speed. An adaptive algorithm is introduced to dynamically adjust resource allocation to ensure the stability of the system under different loads.
6. The method according to claim 3, characterized in that: In the event that the previous model operation fails, the system rolls back the model to a previous more stable configuration through an adaptive fault-tolerance mechanism and continues to optimize the strategy and model operation.
7. The method according to claim 1, characterized in that During the model operation and training process, a deep reinforcement learning algorithm is introduced to enable the optimization strategy to be autonomously learned by the intelligent agent.
8. The method according to claim 7, characterized in that During the operation and training of the model, combined with the training process of the deep learning image classification model, the hyperparameter scheduling in the operation of the model is optimized, and the computing resources are dynamically adjusted based on the training status of the current model.
9. A computer program product, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
10. A computing system, characterized in that: The method comprises a processor and a memory, wherein the processor is configured to execute the method according to any one of claims 1 to 8.