Hyper-parameter adaptive multi-objective optimization method

Through the improved non-dominant sorting genetic algorithm NSGA-III and dynamic weight allocation, combined with the Katib hyperparameter optimization framework, multi-objective optimization of hyperparameters is achieved, solving the problem of low efficiency of hyperparameter tuning in the existing technology, and achieving efficient and adaptive hyperparameter optimization.

CN119987939AActive Publication Date: 2025-05-13SHANDONG INSPUR SCI RES INST CO LTD

Patent Information

Application Number
CN202510013218.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-13
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

The existing hyperparameter optimization methods are difficult to adapt to changes in different tasks and environments, and cannot achieve efficient multi-objective optimization, resulting in low efficiency of hyperparameter tuning in deep learning model training.

Method used

The improved non-dominant sorting genetic algorithm NSGA-III is adopted, combining dynamic weight allocation and adaptive search strategies to achieve multi-objective optimization of hyperparameters, and integrate it with the distributed training platform through the Katib hyperparameter optimization framework to dynamically adjust the hyperparameter configuration.

Benefits of technology

It realizes efficient hyperparameter optimization on the AI ​​training platform, improves optimization efficiency and resource utilization, can adapt to changes in the training environment, and is suitable for training large-scale deep learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987939A_ABST
    Figure CN119987939A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a hyper-parameter adaptive multi-objective optimization method. According to the hyper-parameter adaptive multi-objective optimization method, a user creates a user-defined Expertion experiment resource and verifies the resource; the method comprises the following steps: creating a Suggestation suggestion resource by an Expertion experiment controller, and executing improved NSGA-III (Non-dominated Sorting Genetic Algorithm-III) hyper-parameter optimization; the Trial test controller creates a training task and starts training; and collecting and storing a target index, and if an end condition is met, outputting an optimal hyper-parameter solution set for a decision maker to select. According to the hyper-parameter adaptive multi-objective optimization method, efficient hyper-parameter optimization for an AI training platform is realized, the optimization efficiency is improved, hyper-parameter configuration can be adaptively and dynamically adjusted according to the change of a training environment, and an intelligent and efficient solution is provided for large-scale deep learning model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a hyperparameter adaptive multi-objective optimization method. Background Art

[0002] With the rapid development of artificial intelligence technology, deep learning models are widely used in various fields, such as computer vision, natural language processing, speech recognition, etc. These models usually contain a large number of hyperparameters, such as learning rate, batch size, regularization coefficient, etc. The selection of these hyperparameters has an important impact on model performance. However, manually adjusting hyperparameters is a time-consuming and inefficient process that requires a lot of attempts and experience accumulation.

[0003] In order to improve the efficiency of hyperparameter tuning, researchers have proposed a variety of automated hyperparameter optimization methods, such as grid search, random search, Bayesian optimization, etc.

[0004] Among them, Katib hyperparameter optimization framework is an open source hyperparameter optimization framework based on Kubernetes, which supports multiple optimization algorithms and is widely used in artificial intelligence (AI) training platforms. However, the optimization algorithm of Katib hyperparameter optimization framework is usually based on fixed strategies, which makes it difficult to adapt to changes in different tasks and environments, and cannot achieve adaptive and efficient optimization.

[0005] In order to reduce the time overhead of distributed training, improve resource utilization efficiency, and provide an efficient and intelligent solution for the training of large-scale deep learning models, the present invention proposes a hyperparameter adaptive multi-objective optimization method. Summary of the invention

[0006] In order to overcome the defects of the prior art, the present invention provides a simple and efficient hyperparameter adaptive multi-objective optimization method.

[0007] The present invention is achieved through the following technical solutions:

[0008] A hyperparameter adaptive multi-objective optimization method, characterized in that it comprises the following steps:

[0009] Step S1: The user creates a custom Experiment resource and verifies it;

[0010] User-created custom Experiment resources include the following information:

[0011] Hyperparameter search space: Define the range and type of hyperparameters that need to be optimized, including learning rate, batch size, number of hidden layers, and number of neurons in each layer; each hyperparameter can be customized to continuous, discrete, or categorical values;

[0012] Target indicators: define the target according to the requirements, including accuracy, training time and resource consumption, and whether to maximize or minimize the target indicator;

[0013] Optimization algorithm: The improved non-dominated sorting genetic algorithm NSGA-III is used to search for hyperparameters;

[0014] Parallelism configuration: defines the number of trials that can be run simultaneously and controls the parallelism of training tasks;

[0015] Training task template: defines how each training task runs, including Pod configuration, image address, and startup command;

[0016] Submit the defined Experiment configuration information to the Experiment controller through the distributed training platform's API Server, and verify the correctness and completeness of custom resources based on the Experiment Webhook.

[0017] Step S2, the Experiment controller creates a Suggestion resource;

[0018] If the correctness and completeness of the custom resource are verified, the Experiment controller will create a Suggestion resource to generate a hyperparameter solution set;

[0019] Step S3, performing improved NSGA-III hyper-parameter optimization;

[0020] SuggestionThe controller checks whether the service resources of the improved non-dominated sorting genetic algorithm NSGA-III are ready;

[0021] If the service status of the improved non-dominated sorting genetic algorithm NSGA-III indicates that the service can be provided, the Suggestion controller generates a new hyperparameter solution set based on the non-dominated sorting genetic algorithm NSGA-III and writes it into the suggestion status status.suggestions field of the Suggestion resource;

[0022] Step S4: The Trial controller generates a training task and starts training;

[0023] The Trial controller creates an actual training task, i.e., a Kubernetes Job or Pod instance, for each Trial based on the training task template in the custom Experiment resource, and submits the training task to the Kubernetes cluster for execution.

[0024] Step S5: The Metrics Collector collects and stores target metrics.

[0025] Metrics Collector collects metrics and stores them in the backend database of the Katib hyperparameter optimization framework;

[0026] Step S6: If the end condition is met, the optimal hyperparameter solution set is output;

[0027] When the training task is completed, the trial controller updates the trial resource status of the training task;

[0028] When a trial resource is used up, the Suggestion controller updates the search strategy based on the current experimental results to generate a new set of hyperparameter solutions, and writes the updated set of hyperparameter solutions into the status.suggestions field of the Suggestion resource for use in the next round of trials.

[0029] The tuning process is repeated for multiple rounds of iterations to generate new trial experiment tasks and run them until the pre-defined termination conditions are met;

[0030] When the Experiment resource meets the end condition, the run ends and the optimal hyperparameter solution set is written to the Experiment resource status.paretoOptimalTrials field.

[0031] In step S1, the custom Experiment resource contains relevant configuration information of the training task, and the system triggers the Experiment Webhook to verify the submitted Experiment resource:

[0032] If the Experiment resource is complete and the configuration is legal and valid, the verification is passed, and the Experiment resource will be officially accepted and included in the system management;

[0033] Otherwise, the verification fails and the hook Webhook will return an error message, requiring the user to modify the custom Experiment resource and resubmit it.

[0034] In step S3, the improved non-dominated sorting genetic algorithm NSGA-III is implemented as follows:

[0035] Step S3.1, the improved non-dominated sorting genetic algorithm NSGA-III adopts a dynamic weight allocation method to adaptively adjust the weight of each optimization objective during the search process, find high-quality hyperparameter configurations in multi-objective optimization problems, and thus optimize the performance of the distributed training model;

[0036] The optimization objectives are obtained through actual training in the form of real-time feedback, including accuracy, inference speed, and resource consumption;

[0037] S3.2. Deeply integrate the improved non-dominated sorting genetic algorithm NSGA-III with the Katib hyperparameter optimization framework, so that the Katib hyperparameter optimization framework can adaptively adjust the hyperparameter configuration dynamically according to the changes in the training environment, thereby improving the optimization efficiency;

[0038] S3.3. Make full use of distributed computing resources to achieve parallel hyperparameter optimization;

[0039] S3.4, the improved non-dominated sorting genetic algorithm NSGA-III adopts an adaptive search strategy, which can dynamically adjust the algorithm parameters according to the feedback information in the search process to improve the convergence speed and solution quality;

[0040] S3.5. Combined with intelligent resource allocation and scheduling technology, appropriate computing resources are dynamically allocated according to the characteristics of the training task, and resources are intelligently scheduled to further improve training efficiency and resource utilization.

[0041] In step S3.1, the improved non-dominated sorting genetic algorithm NSGA-III adaptively adjusts the weights of each optimization objective. The specific process is as follows:

[0042] Step S3.1.1. First, obtain the hyperparameter configuration of the training model, including the learning rate lr∈[0.0001,0.1], the batch size bs∈[16,128], the number of hidden layers hn∈[1,5] and the number of neurons in each layer nn∈[32,512]; randomly generate individuals x from the candidate hyperparameter combinations in a normal distribution. i , individual x i It is expressed as:

[0043] x i =random(lr,bs,hn,nn)

[0044] Among them, random represents the normal distribution function;

[0045] Then, several individuals x i The initial population p 0 , the initial population p 0 For each individual x in iBoth contain a set of hyperparameters that need to be optimized, including learning rate, batch size, number of hidden layers, and number of neurons in each layer;

[0046] Step S3.1.2: Train and evaluate the model performance, construct a multi-objective fitness function F to calculate the individual x i The fitness value of

[0047] The multi-objective fitness evaluation function F is expressed as:

[0048] F=ω 1 ·f 1 (x i )-ω 2 ·f 2 (x i )-ω 3 ·f 3 (x i )

[0049] Among them, ω 1 is the initial weight of accuracy, ω 2 is the initial weight for inference speed, ω 3 is the initial weight of resource consumption, f 1 (x i ) is the accuracy, f 2 (x i ) is the training time, f 3 (x i ) is resource consumption;

[0050] Step S3.1.3: Dynamically adjust the weight of each target based on the performance of individuals in the current population:

[0051] The weight is dynamically adjusted based on the fitness value of the target, and the new weight ω 、 k The calculation is as follows:

[0052]

[0053] Among them, 1≦k≦3, ω 、 1 is the accuracy weight, ω 、 2 is the inference speed weight, ω 、 3 is the resource consumption weight;

[0054] Step S3.1.4, selection and evolution process:

[0055] The individuals in the population are sorted and ranked non-dominatedly. The process is as follows:

[0056] First, calculate the individual x iThe dominance relationship is calculated as follows:

[0057] When individual x j The accuracy is higher, the reasoning speed and resource consumption are lower, and it is considered that individual x j The performance on the three optimization objectives is better than that of individual x i ;

[0058] When individual x j The performance on the three optimization objectives is not inferior to that of individual x i , and at least on one optimization goal m, individual x j outperforms individual x i , m∈k, then it is considered that individual x j Dominant individual x i ;

[0059] If individual x j is not dominated by any other individual, then individual x j is called a non-dominated solution;

[0060] According to the dominance relationship, all individuals are divided into different hierarchical fronts;

[0061] Step S3.1.5: For each non-dominated level calculated in step S3.1.4, calculate the crowding distance of the individual to evaluate the distribution of the individual in the target space, so as to maintain the diversity of the population during the evolution process and avoid the concentration of solutions in certain specific areas;

[0062] Individual x i The crowding distance d between its neighboring individuals i Calculated by the following formula:

[0063]

[0064] Among them, f k max and f k min are the maximum and minimum values ​​of the optimization target in the population;

[0065] Step S3.1.6: Customize and select individuals based on non-dominated level and crowding distance to form a new population p 1-tmp ;

[0066] Step S3.1.7, performing crossover and mutation operations on the selected individuals to generate new individuals;

[0067] Among them, the crossover operation is realized by single-point crossover or multi-point crossover, and the mutation operation is realized by randomly changing the hyperparameter value;

[0068] Step S3.1.8: Add the newly generated individuals to the current population p1-tmp Merge to form a new temporary population p ` , and for the population p ` Perform non-dominated sorting and crowding distance calculation to select a new population p 1 ;

[0069] Step S3.1.9, repeat steps S3.1.3 to S3.1.7 until the custom termination condition is met;

[0070] Step S3.1.10: Output the solution set on the Pareto front and provide a set of uniformly distributed hyperparameter configurations for decision makers to choose from.

[0071] In step S3.1.4, all individuals are divided into fronts of different levels according to the dominance relationship. The division method is as follows:

[0072] The first level fronts contains all non-dominated solutions;

[0073] The second level fronts contain individuals that are dominated only by non-dominated solutions in the first level fronts;

[0074] The third level fronts contain individuals that are dominated only by individuals in the second level fronts;

[0075] And so on, divide into several levels of fronts until all individuals are divided;

[0076] In step S3.1.6, when selecting individuals based on the non-dominated level and the crowding distance, individuals with high non-dominated levels (close to the first level) are preferentially selected;

[0077] When the number of individuals in the same non-dominated level is greater than the population requirement, further screening is performed through crowding distance, and individuals with more dispersed distribution in the target space are preferentially selected, that is, the crowding distance d between them and their adjacent individuals i Large individual.

[0078] In step S3, the Experiment controller monitors the update of the Suggestion resource in real time. If it is found that the Suggestion resource has been updated, a Trial resource is generated for each new hyperparameter set.

[0079] A hyperparameter adaptive multi-objective optimization system, comprising:

[0080] The Experiment creation module is responsible for helping users create custom Experiment resources, submitting the defined Experiment configuration information to the Experiment controller through the API Server of the distributed training platform, and verifying the correctness and completeness of custom resources based on the Experiment Webhook.

[0081] The Experiment controller is responsible for creating a Suggestion resource after receiving the Experiment resource created by the user, which is used to generate a hyperparameter solution set. At the same time, it monitors the update of the Suggestion resource in real time. If it is found that the Suggestion resource has been updated, it will generate a Trial resource for each new hyperparameter set.

[0082] Suggestion controller is responsible for checking whether the service resources of the improved non-dominated sorting genetic algorithm NSGA-III are ready;

[0083] If the service status of the improved non-dominated sorting genetic algorithm NSGA-III indicates that the service can be provided, the Suggestion controller generates a new hyperparameter solution set based on the non-dominated sorting genetic algorithm NSGA-III and writes it into the suggestion status status.suggestions field of the Suggestion resource;

[0084] At the same time, the Suggestion controller is responsible for generating a new set of hyperparameter solutions based on the current indicator updates when the Trial resource execution is completed;

[0085] The Trial controller is responsible for creating actual training tasks for each Trial according to the template of the training tasks in the custom Experiment resource, i.e., Kubernetes Job or Pod instance, and submitting the training task to the Kubernetes cluster for running. When the training task is completed, the Trial resource status of the training task is updated. When the Experiment task meets the termination condition, the running is terminated and the optimal hyperparameter solution set is written to the Experiment resource status.paretoOptimalTrials field.

[0086] Metrics Collector is responsible for collecting target metrics and storing them in the backend database of the Katib hyperparameter optimization framework.

[0087] A hyperparameter adaptive multi-objective optimization device, characterized by comprising:

[0088] One or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the above methods.

[0089] A readable storage medium, characterized in that: a computer program is stored on the readable storage medium, and the computer program implements the above method when executed by a processor.

[0090] The beneficial effect of the present invention is that the hyperparameter adaptive multi-objective optimization method realizes efficient hyperparameter optimization for AI training platforms through adaptive search strategies and dynamic multi-objective optimization, which not only improves the optimization efficiency, but also can adaptively adjust the hyperparameter configuration dynamically according to changes in the training environment, providing an intelligent and efficient solution for large-scale deep learning model training. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0092] Attached Figure 1 Schematic diagram of the implementation architecture of the hyperparameter adaptive multi-objective optimization system of the present invention.

[0093] Attached Figure 2 The figure is a schematic diagram of the implementation process of the hyperparameter adaptive multi-objective optimization method of the present invention.

[0094] Attached Figure 3 Schematic diagram of the execution flow of the hyperparameter adaptive multi-objective optimization method of the present invention.

[0095] Attached Figure 4 Schematic diagram of an example configuration file for the Suggestion resource of the present invention DETAILED DESCRIPTION

[0096] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0097] The hyperparameter adaptive multi-objective optimization method comprises the following steps:

[0098] Step S1: The user creates a custom Experiment resource and verifies it;

[0099] User-created custom Experiment resources include the following information:

[0100] Hyperparameter search space: Define the range and type of hyperparameters that need to be optimized, including learning rate, batch size, number of hidden layers, and number of neurons in each layer; each hyperparameter can be customized to continuous, discrete, or categorical values;

[0101] Target indicators: define the target according to the requirements, including accuracy, training time and resource consumption, and whether to maximize or minimize the target indicator;

[0102] Optimization algorithm: The improved non-dominated sorting genetic algorithm NSGA-III is used to search for hyperparameters;

[0103] Parallelism configuration: defines the number of trials that can be run simultaneously and controls the parallelism of training tasks;

[0104] Training task template: defines how each training task runs, including Pod configuration, image address, and startup command;

[0105] Submit the defined Experiment configuration information to the Experiment controller through the distributed training platform's API Server, and verify the correctness and completeness of custom resources based on the Experiment Webhook.

[0106] Step S2, the Experiment controller creates a Suggestion resource;

[0107] If the correctness and completeness of the custom resource are verified, the Experiment controller will create a Suggestion resource to generate a hyperparameter solution set;

[0108] Step S3, performing improved NSGA-III hyper-parameter optimization;

[0109] SuggestionThe controller checks whether the service resources of the improved non-dominated sorting genetic algorithm NSGA-III are ready;

[0110] If the service status of the improved non-dominated sorting genetic algorithm NSGA-III indicates that the service can be provided, the Suggestion controller generates a new hyperparameter solution set based on the non-dominated sorting genetic algorithm NSGA-III and writes it into the suggestion status status.suggestions field of the Suggestion resource;

[0111] Step S4: The Trial controller generates a training task and starts training;

[0112] The Trial controller creates an actual training task, i.e., a Kubernetes Job or Pod instance, for each Trial based on the training task template in the custom Experiment resource, and submits the training task to the Kubernetes cluster for execution.

[0113] Step S5: The Metrics Collector collects and stores target metrics.

[0114] Metrics Collector collects metrics and stores them in the backend database of the Katib hyperparameter optimization framework;

[0115] Step S6: If the end condition is met, the optimal hyperparameter solution set is output;

[0116] When the training task is completed, the trial controller updates the trial resource status of the training task;

[0117] When a trial resource is used up, the Suggestion controller updates the search strategy based on the current experimental results to generate a new set of hyperparameter solutions, and writes the updated set of hyperparameter solutions into the status.suggestions field of the Suggestion resource for use in the next round of trials.

[0118] The tuning process is repeated for multiple rounds of iterations to generate new trial experiment tasks and run them until the pre-defined termination conditions are met;

[0119] When the Experiment resource meets the end condition, the run ends and the optimal hyperparameter solution set is written to the Experiment resource status.paretoOptimalTrials field.

[0120] In step S1, the custom Experiment resource contains relevant configuration information of the training task, and the system triggers the Experiment Webhook to verify the submitted Experiment resource:

[0121] If the Experiment resource is complete and the configuration is legal and valid, the verification is passed, and the Experiment resource will be officially accepted and included in the system management;

[0122] Otherwise, the verification fails and the hook Webhook will return an error message, requiring the user to modify the custom Experiment resource and resubmit it.

[0123] In step S3, the improved non-dominated sorting genetic algorithm NSGA-III is implemented as follows:

[0124] Step S3.1, the improved non-dominated sorting genetic algorithm NSGA-III adopts a dynamic weight allocation method to adaptively adjust the weight of each optimization objective during the search process, find high-quality hyperparameter configurations in multi-objective optimization problems, and thus optimize the performance of the distributed training model;

[0125] The optimization objectives are obtained through actual training in the form of real-time feedback, including accuracy, inference speed, and resource consumption;

[0126] S3.2. Deeply integrate the improved non-dominated sorting genetic algorithm NSGA-III with the Katib hyperparameter optimization framework, so that the Katib hyperparameter optimization framework can adaptively adjust the hyperparameter configuration dynamically according to the changes in the training environment, thereby improving the optimization efficiency;

[0127] S3.3. Make full use of distributed computing resources to achieve parallel hyperparameter optimization;

[0128] S3.4, the improved non-dominated sorting genetic algorithm NSGA-III adopts an adaptive search strategy, which can dynamically adjust the algorithm parameters according to the feedback information in the search process to improve the convergence speed and solution quality;

[0129] S3.5. Combined with intelligent resource allocation and scheduling technology, appropriate computing resources are dynamically allocated according to the characteristics of the training task, and resources are intelligently scheduled to further improve training efficiency and resource utilization.

[0130] In step S3.1, the improved non-dominated sorting genetic algorithm NSGA-III adaptively adjusts the weights of each optimization objective. The specific process is as follows:

[0131] Step S3.1.1. First, obtain the hyperparameter configuration of the training model, including the learning rate lr∈[0.0001,0.1], the batch size bs∈[16,128], the number of hidden layers hn∈[1,5] and the number of neurons in each layer nn∈[32,512]; randomly generate individuals x from the candidate hyperparameter combinations in a normal distribution. i , individual x i It is expressed as:

[0132] x i =random(lr,bs,hn,nn)

[0133] Among them, random represents the normal distribution function;

[0134] Then, several individuals x i The initial population p 0 , the initial population p 0 For each individual x in i Both contain a set of hyperparameters that need to be optimized, including learning rate, batch size, number of hidden layers, and number of neurons in each layer;

[0135] Step S3.1.2: Train and evaluate the model performance, construct a multi-objective fitness function F to calculate the individual x i The fitness value of

[0136] The multi-objective fitness evaluation function F is expressed as:

[0137] F=ω 1 ·f 1 (x i )-ω 2 ·f 2 (x i )-ω 3 ·f 3 (x i )

[0138] Among them, ω 1 is the initial weight of accuracy, ω 2 is the initial weight for inference speed, ω 3 is the initial weight of resource consumption, f 1 (x i ) is the accuracy, f 2 (x i ) is the training time, f 3 (x i ) is resource consumption;

[0139] Step S3.1.3: Dynamically adjust the weight of each target based on the performance of individuals in the current population:

[0140] The weight is dynamically adjusted based on the fitness value of the target, and the new weight ω 、 k The calculation is as follows:

[0141]

[0142] Among them, 1≦k≦3, ω 、 1 is the accuracy weight, ω 、 2 is the inference speed weight, ω 、 3 is the resource consumption weight;

[0143] Step S3.1.4, selection and evolution process:

[0144] The individuals in the population are sorted and ranked non-dominatedly. The process is as follows:

[0145] First, calculate the individual x i The dominance relationship is calculated as follows:

[0146] When individual x j The accuracy is higher, the reasoning speed and resource consumption are lower, and it is considered that individual x j The performance on the three optimization objectives is better than that of individual x i ;

[0147] When individual x j The performance on the three optimization objectives is not inferior to that of individual x i , and at least on one optimization goal m, individual x j outperforms individual x i , m∈k, then it is considered that individual x j Dominant individual x i ;

[0148] If individual x j is not dominated by any other individual, then individual x j is called a non-dominated solution;

[0149] According to the dominance relationship, all individuals are divided into different hierarchical fronts;

[0150] Step S3.1.5: For each non-dominated level calculated in step S3.1.4, calculate the crowding distance of the individual to evaluate the distribution of the individual in the target space, so as to maintain the diversity of the population during the evolution process and avoid the concentration of solutions in certain specific areas;

[0151] Individual x i The crowding distance d between its neighboring individuals i Calculated by the following formula:

[0152]

[0153] Among them, f k max and f k min are the maximum and minimum values ​​of the optimization target in the population;

[0154] Step S3.1.6: Customize and select individuals based on non-dominated level and crowding distance to form a new population p 1-tmp ;

[0155] Step S3.1.7, performing crossover and mutation operations on the selected individuals to generate new individuals;

[0156] Among them, the crossover operation is realized by single-point crossover or multi-point crossover, and the mutation operation is realized by randomly changing the hyperparameter value;

[0157] Step S3.1.8: Add the newly generated individuals to the current population p 1-tmp Merge to form a new temporary population p ` , and for the population p ` Perform non-dominated sorting and crowding distance calculation to select a new population p 1 ;

[0158] Step S3.1.9, repeat steps S3.1.3 to S3.1.7 until the custom termination condition is met;

[0159] Step S3.1.10: Output the solution set on the Pareto front and provide a set of uniformly distributed hyperparameter configurations for decision makers to choose from.

[0160] In step S3.1.4, all individuals are divided into fronts of different levels according to the dominance relationship. The division method is as follows:

[0161] The first level fronts contains all non-dominated solutions;

[0162] The second level fronts contain individuals that are dominated only by non-dominated solutions in the first level fronts;

[0163] The third level fronts contain individuals that are dominated only by individuals in the second level fronts;

[0164] And so on, divide into several levels of fronts until all individuals are divided;

[0165] In step S3.1.6, when selecting individuals based on the non-dominated level and the crowding distance, individuals with high non-dominated levels (close to the first level) are preferentially selected;

[0166] When the number of individuals in the same non-dominated level is greater than the population requirement, further screening is performed through crowding distance, and individuals with more dispersed distribution in the target space are preferentially selected, that is, the crowding distance d between them and their adjacent individuals i Large individual.

[0167] In step S3, the Experiment controller monitors the update of the Suggestion resource in real time. If it is found that the Suggestion resource has been updated, a Trial resource is generated for each new hyperparameter set.

[0168] A hyperparameter adaptive multi-objective optimization system, comprising:

[0169] The Experiment creation module is responsible for helping users create custom Experiment resources, submitting the defined Experiment configuration information to the Experiment controller through the API Server of the distributed training platform, and verifying the correctness and completeness of custom resources based on the Experiment Webhook.

[0170] The Experiment controller is responsible for creating a Suggestion resource after receiving the Experiment resource created by the user, which is used to generate a hyperparameter solution set. At the same time, it monitors the update of the Suggestion resource in real time. If it is found that the Suggestion resource has been updated, it will generate a Trial resource for each new hyperparameter set.

[0171] Suggestion controller is responsible for checking whether the service resources of the improved non-dominated sorting genetic algorithm NSGA-III are ready;

[0172] If the service status of the improved non-dominated sorting genetic algorithm NSGA-III indicates that the service can be provided, the Suggestion controller generates a new hyperparameter solution set based on the non-dominated sorting genetic algorithm NSGA-III and writes it into the suggestion status status.suggestions field of the Suggestion resource;

[0173] At the same time, the Suggestion controller is responsible for generating a new set of hyperparameter solutions based on the current indicator updates when the Trial resource execution is completed;

[0174] The Trial controller is responsible for creating actual training tasks for each Trial according to the template of the training tasks in the custom Experiment resource, i.e., Kubernetes Job or Pod instance, and submitting the training task to the Kubernetes cluster for running. When the training task is completed, the Trial resource status of the training task is updated. When the Experiment task meets the termination condition, the running is terminated and the optimal hyperparameter solution set is written to the Experiment resource status.paretoOptimalTrials field.

[0175] Metrics Collector is responsible for collecting target metrics and storing them in the backend database of the Katib hyperparameter optimization framework.

[0176] The hyperparameter adaptive multi-objective optimization device includes:

[0177] One or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the above methods.

[0178] The readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.

[0179] Compared with the existing technology, this hyperparameter adaptive multi-objective optimization method has the following characteristics:

[0180] (1) The algorithm of Katib hyperparameter optimization framework adopts optimized evolutionary algorithm to achieve intelligent adaptive hyperparameter tuning;

[0181] (2) The optimal multi-objective optimization scheme is learned through interactive training of the improved non-dominated sorting genetic algorithm NSGA-III, which enables the system to automatically search and adjust the hyperparameter configuration of complex models during the training process;

[0182] (3) By innovatively adopting the improved non-dominated sorting genetic algorithm NSGA-III, combined with an adaptive search strategy and a dynamic acquisition function, efficient multi-objective hyperparameter optimization is achieved, which solves the common parameter adjustment problems in the distributed training process and significantly improves the training speed and resource utilization efficiency.

[0183] The embodiment described above is only one specific implementation of the present invention. Common changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included in the protection scope of the present invention.

Claims

1. A hyperparameter adaptive multi-objective optimization method, characterized by: The following steps are involved: Step S1: The user creates a custom Experiment resource and verifies it; User-created custom Experiment resources include the following information: Hyperparameter search space: Define the range and type of hyperparameters that need to be optimized, including learning rate, batch size, number of hidden layers, and number of neurons in each layer; each hyperparameter can be customized to continuous, discrete, or categorical values; Target indicators: define the target according to the requirements, including accuracy, training time and resource consumption, and whether to maximize or minimize the target indicator; Optimization algorithm: The improved non-dominated sorting genetic algorithm NSGA-III is used to search for hyperparameters; Parallelism configuration: defines the number of trials that can be run simultaneously and controls the parallelism of training tasks; Training task template: defines how each training task runs, including Pod configuration, image address, and startup command; Submit the defined Experiment configuration information to the Experiment controller through the distributed training platform's API Server, and verify the correctness and completeness of custom resources based on the Experiment Webhook. Step S2, the Experiment controller creates a Suggestion resource; If the correctness and completeness of the custom resource are verified, the Experiment controller will create a Suggestion resource to generate a hyperparameter solution set; Step S3, performing improved NSGA-III hyper-parameter optimization; SuggestionThe controller checks whether the service resources of the improved non-dominated sorting genetic algorithm NSGA-III are ready; If the service status of the improved non-dominated sorting genetic algorithm NSGA-III indicates that the service can be provided, the Suggestion controller generates a new hyperparameter solution set based on the non-dominated sorting genetic algorithm NSGA-III and writes it into the suggestion status status.suggestions field of the Suggestion resource; Step S4: The Trial controller generates a training task and starts training; The Trial controller creates an actual training task, i.e., a Kubernetes Job or Pod instance, for each Trial based on the training task template in the custom Experiment resource, and submits the training task to the Kubernetes cluster for execution. Step S5: The Metrics Collector collects and stores target metrics. Metrics Collector collects metrics and stores them in the backend database of the Katib hyperparameter optimization framework; Step S6: If the end condition is met, the optimal hyperparameter solution set is output; When the training task is completed, the trial controller updates the trial resource status of the training task; When a trial resource is used up, the Suggestion controller updates the search strategy based on the current experimental results to generate a new set of hyperparameter solutions, and writes the updated set of hyperparameter solutions into the status.suggestions field of the Suggestion resource for use in the next round of trial experiments. The tuning process is repeated for multiple rounds of iterations to generate new trial experiment tasks and run them until the pre-defined termination conditions are met; When the Experiment resource meets the end condition, the run ends and the optimal hyperparameter solution set is written to the Experiment resource status.paretoOptimalTrials field.

2. The hyperparameter adaptive multi-objective optimization method according to claim 1, characterized in that: In step S1, the custom Experiment resource contains relevant configuration information of the training task, and the system triggers the ExperimentWebhook to verify the submitted Experiment resource: If the Experiment resource is complete and the configuration is legal and valid, the verification is passed, and the Experiment resource will be officially accepted and included in the system management; Otherwise, the verification fails and the hook Webhook will return an error message, requiring the user to modify the custom Experiment resource and resubmit it.

3. The hyperparameter adaptive multi-objective optimization method according to claim 1, characterized in that: In step S3, the improved non-dominated sorting genetic algorithm NSGA-III is implemented as follows: Step S3.1, the improved non-dominated sorting genetic algorithm NSGA-III adopts a dynamic weight allocation method to adaptively adjust the weight of each optimization objective during the search process, find high-quality hyperparameter configurations in multi-objective optimization problems, and thus optimize the performance of the distributed training model; The optimization objectives are obtained through actual training in the form of real-time feedback, including accuracy, inference speed, and resource consumption; S3.

2. Deeply integrate the improved non-dominated sorting genetic algorithm NSGA-III with the Katib hyperparameter optimization framework, so that the Katib hyperparameter optimization framework can adaptively adjust the hyperparameter configuration dynamically according to the changes in the training environment, thereby improving the optimization efficiency; S3.

3. Make full use of distributed computing resources to achieve parallel hyperparameter optimization; S3.4, the improved non-dominated sorting genetic algorithm NSGA-III adopts an adaptive search strategy, which can dynamically adjust the algorithm parameters according to the feedback information in the search process to improve the convergence speed and solution quality; S3.

5. Combined with intelligent resource allocation and scheduling technology, appropriate computing resources are dynamically allocated according to the characteristics of the training task, and resources are intelligently scheduled to further improve training efficiency and resource utilization.

4. The hyperparameter adaptive multi-objective optimization method according to claim 3, characterized in that: In step S3.1, the improved non-dominated sorting genetic algorithm NSGA-III adaptively adjusts the weights of each optimization objective. The specific process is as follows: Step S3.1.

1. First, obtain the hyperparameter configuration of the training model, including the learning rate lr∈[0.0001,0.1], the batch size bs∈[16,128], the number of hidden layers hn∈[1,5] and the number of neurons in each layer nn∈[32,512]; randomly generate individuals x from the candidate hyperparameter combinations in a normal distribution. i , individual x i It is expressed as: x i =random(lr,bs,hn,nn) Among them, random represents the normal distribution function; Then, several individuals x i Constitute the initial population p0, each individual x in the initial population p0 i Both contain a set of hyperparameters that need to be optimized, including learning rate, batch size, number of hidden layers, and number of neurons in each layer; Step S3.1.2: Train and evaluate the model performance, construct a multi-objective fitness function F to calculate the individual x i The fitness value of The multi-objective fitness evaluation function F is expressed as: F=ω1·f1(x i )-ω2·f2(x i )-ω3·f3(x i ) Among them, ω1 is the initial weight of accuracy, ω2 is the initial weight of inference speed, ω3 is the initial weight of resource consumption, and f1(x i ) is the accuracy, f2(x i ) is the training time, f3(x i ) is resource consumption; Step S3.1.3: Dynamically adjust the weight of each target based on the performance of individuals in the current population: The weight is dynamically adjusted based on the fitness value of the target, and the new weight ω 、 k The calculation is as follows: Among them, 1≦k≦3, ω 、 1 is the accuracy weight, ω 、 2 is the inference speed weight, ω 、 3 is the resource consumption weight; Step S3.1.4, selection and evolution process: The individuals in the population are sorted and ranked non-dominatedly. The process is as follows: First, calculate the individual x i The dominance relationship is calculated as follows: When individual x j The accuracy is higher, the reasoning speed and resource consumption are lower, and it is considered that individual x j The performance on the three optimization objectives is better than that of individual x i ; When individual x j The performance on the three optimization objectives is not inferior to that of individual x i , and at least on one optimization goal m, individual x j outperforms individual x i , m∈k, then it is considered that individual x j Dominant individual x i ; If individual x j is not dominated by any other individual, then individual x j is called a non-dominated solution; According to the dominance relationship, all individuals are divided into different hierarchical fronts; Step S3.1.5: For each non-dominated level calculated in step S3.1.4, calculate the crowding distance of the individual to evaluate the distribution of the individual in the target space, so as to maintain the diversity of the population during the evolution process and avoid the concentration of solutions in certain specific areas; Individual x i The crowding distance d between its neighboring individuals i Calculated by the following formula: Among them, f k max and f k min are the maximum and minimum values ​​of the optimization target in the population; Step S3.1.6: Customize and select individuals based on non-dominated level and crowding distance to form a new population p 1-tmp ; Step S3.1.7, performing crossover and mutation operations on the selected individuals to generate new individuals; Among them, the crossover operation is realized by single-point crossover or multi-point crossover, and the mutation operation is realized by randomly changing the hyperparameter value; Step S3.1.8: Add the newly generated individuals to the current population p 1-tmp Merge to form a new temporary population p ` , and for the population p ` Perform non-dominated sorting and crowding distance calculation to select a new population p1; Step S3.1.9, repeat steps S3.1.3 to S3.1.7 until the custom termination condition is met; Step S3.1.10: Output the solution set on the Pareto front and provide a set of uniformly distributed hyperparameter configurations for decision makers to choose from.

5. The hyperparameter adaptive multi-objective optimization method according to claim 4, characterized in that: In step S3.1.4, all individuals are divided into fronts of different levels according to the dominance relationship. The division method is as follows: The first level fronts contains all non-dominated solutions; The second level fronts contain individuals that are dominated only by non-dominated solutions in the first level fronts; The third level fronts contain individuals that are dominated only by individuals in the second level fronts; And so on, several levels of fronts are divided until all individuals are divided.

6. The hyperparameter adaptive multi-objective optimization method according to claim 4, characterized in that: In step S3.1.6, when selecting individuals based on the non-dominated level and the crowding distance, individuals with high non-dominated levels are preferentially selected; When the number of individuals in the same non-dominated level is greater than the population requirement, further screening is performed through crowding distance, and individuals with more dispersed distribution in the target space are preferentially selected, that is, the crowding distance d between them and their adjacent individuals i Large individual.

7. The hyperparameter adaptive multi-objective optimization method according to claim 1, characterized in that: In step S3, the Experiment controller monitors the update status of the Suggestion resource in real time. If it is found that the Suggestion resource has been updated, a Trial resource is generated for each new hyperparameter set.

8. A hyperparameter adaptive multi-objective optimization system, characterized by: include: The Experiment creation module is responsible for helping users create custom Experiment resources, submitting the defined Experiment configuration information to the Experiment controller through the API Server of the distributed training platform, and verifying the correctness and completeness of custom resources based on the Experiment Webhook. The Experiment controller is responsible for creating a Suggestion resource after receiving the Experiment resource created by the user, which is used to generate a hyperparameter solution set. At the same time, it monitors the update of the Suggestion resource in real time. If it is found that the Suggestion resource has been updated, it will generate a Trial resource for each new hyperparameter set. Suggestion controller is responsible for checking whether the service resources of the improved non-dominated sorting genetic algorithm NSGA-III are ready; If the service status of the improved non-dominated sorting genetic algorithm NSGA-III indicates that the service can be provided, the Suggestion controller generates a new hyperparameter solution set based on the non-dominated sorting genetic algorithm NSGA-III and writes it into the suggestion status status.suggestions field of the Suggestion resource; At the same time, the Suggestion controller is responsible for generating a new set of hyperparameter solutions based on the current indicator updates when the Trial resource execution is completed; The Trial controller is responsible for creating actual training tasks for each Trial according to the template of the training task in the custom Experiment resource, i.e., Kubernetes Job or Pod instance, and submitting the training task to the Kubernetes cluster for running. When the training task is completed, the Trial resource status of the training task is updated. When the Experiment task meets the termination condition, the running is terminated and the optimal hyperparameter solution set is written to the Experiment resource status.paretoOptimalTrials field. Metrics Collector is responsible for collecting target metrics and storing them in the backend database of the Katib hyperparameter optimization framework.

9. A hyperparameter adaptive multi-objective optimization device, characterized in that: include: One or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the methods according to claims 1 to 7.

10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Universal deep learning hyper-parameter optimization method and device

    CN111401567A

  • System resource and model hyper-parameter collaborative optimization method in deep learning training

    CN112836796A

  • Automatic machine learning method and system based on kubernetes

    CN113011559A

  • Method and system for determining technological parameters of selective laser melting technology

    CN116502455A

  • Resource allocation for tuning hyperparameters of large-scale deep learning workloads

    US20220035672A1

Cited By

  • Particle accelerator system loop parameter online intelligent regulation and control method and device based on large language model

    CN120583579A