Multi-objective optimization method based on meta-learning
By adopting the meta-learning method of internal and external double-layer loop structure in the multi-objective optimization algorithm, neural networking genetic operators are solved to achieve adaptive parameter adjustment and selection, and the existing multi-objective optimization algorithm is insufficiently efficient and flexible, and the algorithm's adaptability and generalization ability are improved.
Patent Information
- Application Number
- CN202510226656.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
AI Technical Summary
Existing multi-objective optimization algorithms require manual design of operators and parameters, and the learnable single-objective optimization algorithm cannot be directly applied to multi-objective optimization problems, resulting in inefficiency and insufficient flexibility.
The multi-objective optimization method based on meta-learning is adopted to neural network the genetic operators of the meta-learning multi-objective genetic algorithm through the internal and external double-layer circular structure, and the multi-layer perceptron network in the neural network is used for adaptive parameter adjustment and selection to realize the training of the multi-objective optimization model.
It improves the flexibility and adaptability of genetic operators, can better solve multi-objective optimization problems, enhances the generalization ability of algorithms, and avoids the tedious process of manual adjustment.
Smart Images

Figure CN120163050A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-objective optimization, and particularly relates to a multi-objective optimization method based on meta-learning. Background Art
[0002] In practical applications, multi-objective optimization problems have received extensive attention due to their universality. Multi-objective optimization problems are prevalent in various aspects such as industrial manufacturing, transportation, and economic finance. Multi-objective optimization aims to solve problems with multiple objectives, which often conflict with each other, and there is no single solution that can optimize all objectives. For example, when conducting industrial production, the profit obtained from production and production losses are two conflicting objectives, and multi-objective optimization is required to obtain the optimal trade-off. Multi-objective evolutionary algorithms perform well in solving such problems, especially in the absence of gradient information. However, due to the diversity of multi-objective optimization problems and the limitations of traditional multi-objective evolutionary algorithms, technicians usually need to manually design the operators and related parameters of multi-objective evolutionary algorithms for different multi-objective optimization problems. Some learnable algorithms have achieved success in single-objective optimization problems, but there is still a lack of research results for learnable multi-objective optimization.
[0003] Since multi-objective optimization problems require optimizing multiple objective problems simultaneously, it is difficult to obtain problem information such as gradients. Therefore, multi-objective evolutionary algorithms are often used to solve multi-objective optimization problems. Multi-objective evolutionary algorithms are divided into three categories. One category is multi-objective evolutionary algorithms based on dominance relationships, such as non-dominated sorting genetic algorithms (NSGA-II), multi-objective adaptive genetic algorithms (AGE-MOEA), etc. This type of algorithm usually selects the next-generation population based on dominance relationships during evolution. The second category is multi-objective evolutionary algorithms based on indicators, such as multi-objective evolutionary algorithms based on hypervolume indicators (SMS-EMOA). This type of algorithm judges the quality of the population during evolution through some indicators, such as hypervolume, inverted generational distance, etc., and uses this to guide the evolution direction. The third category is multi-objective algorithms based on decomposition (MOEA / D). This type of algorithm decomposes multi-objective optimization problems into multiple single-objective optimization sub-problems and optimizes multi-objective problems by optimizing these single-objective sub-problems.
[0004] The disadvantages of the prior art are as follows:
[0005] There is a wide variety of current multi-objective optimization problems. For example, in the body design of a car, safety, lightweight, and cost control need to be considered. After mathematical modeling, these can form multi-objective optimization problems. These multi-objective optimization problems vary greatly in mathematical form. Most of the existing multi-objective evolutionary optimization algorithms require researchers to manually design and adjust appropriate evolutionary operators and corresponding parameters according to the specific multi-objective optimization problems they face. This requires researchers to have rich prior information and sometimes even requires almost exhaustive adjustments, which is very time-consuming and energy-consuming.
[0006] Existing learnable evolutionary optimization algorithms can only solve single-objective optimization problems. However, in the real world, various problems are generally multi-objective optimization problems with multiple conflicting objectives. Moreover, due to the essential differences between single-objective optimization problems and multi-objective optimization problems, especially the fact that gradient information cannot be directly obtained for multi-objective problems, it is difficult to train learnable evolutionary operators. Therefore, existing learnable single-objective evolutionary algorithms cannot be directly applied to multi-objective optimization problems. Summary of the Invention
[0007] To solve the above problems existing in the prior art, the present invention provides a multi-objective optimization method based on meta-learning. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0008] The present invention provides a multi-objective optimization method based on meta-learning, and the method includes:
[0009] Obtain the data to be optimized corresponding to the multi-objective optimization task;
[0010] Process the data to be optimized by using a pre-trained multi-objective optimization model based on meta-learning to obtain an optimization result; wherein,
[0011] The multi-objective optimization model based on meta-learning adopts an inner and outer double-loop structure. In the inner loop, the genetic operators of the meta-learning multi-objective genetic algorithm are neural networked; the network parameters of the neural networked genetic operators are used as the optimization tasks of the outer loop, and the evaluation index of the inner loop is used as the objective function value of the outer loop to realize the training of the multi-objective optimization model based on meta-learning.
[0012] In an embodiment of the present invention, the training process of the multi-objective optimization model based on meta-learning includes:
[0013] S01, taking the parameters of the multi-objective optimization model based on meta-learning as the individuals of the outer loop population, and randomly initializing the individuals of the outer loop population;
[0014] S02. Take the objectives of the multi-objective optimization task as the individuals of the inner-loop population, initialize and evaluate the inner-loop population to obtain the initial parameters of the inner loop.
[0015] S03. Take the individuals of the initialized outer-loop population as the neural network weight parameters of the genetic operator, load them into the genetic operator of the meta-learning multi-objective genetic algorithm, and use the loaded genetic operator to perform several inner-loop processes on the initial parameters of the inner loop until the number of loops reaches the preset number of loops, and output the iterated inner-loop population.
[0016] S04. Perform hypervolume evaluation on the meta-learning-based multi-objective optimization model corresponding to the iterated inner-loop population to obtain a hypervolume score. If the hypervolume score meets the preset conditions, use the current meta-learning-based multi-objective optimization model as the pre-trained meta-learning-based multi-objective optimization model; if the hypervolume score does not meet the preset conditions, return to step S02, optimize the parameters of the meta-learning-based multi-objective optimization model, replace the individuals of the outer-loop population with the optimized parameters of the meta-learning-based multi-objective optimization model, and repeat steps S02 - S04 until the obtained hypervolume score meets the preset conditions or the number of repeated executions reaches the preset number of executions to obtain the pre-trained meta-learning-based multi-objective optimization model.
[0017] In one embodiment of the present invention, the genetic operator includes:
[0018] A sampling operator, a mutation operator, and a selection operator; where
[0019] Both the sampling operator and the mutation operator are constituted based on the self-attention mechanism;
[0020] The selection operator is constituted based on a multi-layer perceptron.
[0021] In one embodiment of the present invention, the multi-head self-attention mechanism network in the sampling operator and the mutation operator includes:
[0022] An input linear layer, a scaled dot-product attention mechanism, a fusion splicing module, and an output linear layer connected in sequence; where
[0023] The input linear layer is composed of three linear layers, and the scaled dot-product attention mechanism is composed of h attention heads; h represents a hyperparameter, h = 2 n , n≥1.
[0024] In one embodiment of the present invention, the multi-layer perceptron network in the selection operator includes:
[0025] An input layer, a hidden layer, and an output layer connected in sequence.
[0026] In one embodiment of the present invention, the initial parameters of the inner loop include:
[0027] A population vector X of size N P , the objective function values of m objective functions The non-dominated rank r of the population P and the initial mutation rate σ P .
[0028] In one embodiment of the present invention, the process of performing one inner loop in a number of inner loop processes on the initial parameters of the inner loop by using the loaded genetic operators includes:
[0029] Sampling the initial parameters by using the loaded sampling operator to obtain a parental population vector Parental objective function values Parental non-dominated rank and parental mutation rate
[0030] Processing the parental objective function values Parental non-dominated rank and parental mutation rate by using the loaded mutation operator to obtain a new mutation rate σ P '; Performing Gaussian mutation on the parental population vector P by using the new mutation rate σ ' to obtain a mutated offspring population vector X P '; Evaluating the mutated offspring population vector X P ' to obtain the objective function values of the mutated offspring
[0031] Processing the objective function values of the mutated offspring by using the loaded selection operator to obtain the population vector X C , objective function values the non-dominated rank r of the population C and the population mutation rate σ C .
[0032] In one embodiment of the present invention, processing the objective function values of the mutated offspring by using the loaded selection operator to obtain the population vector X C , objective function values the non-dominated rank r of the population C and the mutation rate σ C , includes:
[0033] Taking the objective function values of the mutated offspring Input the crowding distance of individuals in the mutated offspring population into the loaded selection operator to obtain the population vector X of the next generation C and the objective function value
[0034]
[0035] Perform non - dominated sorting on the population vector of the next generation to obtain the non - dominated rank r of the population of the next generation C ;
[0036] The loaded selection operator obtains the mutation rate σ of the population of the next generation from the parental mutation rate C .
[0037] Advantages of the present invention:
[0038] In the solution provided by the present invention, the proposed multi - objective optimization model based on meta - learning adopts an inner - outer double - loop structure. In the inner loop, the genetic operators of the meta - learning multi - objective genetic algorithm are neural network - ized, enabling the genetic operators to adaptively adjust parameters, endowing the genetic operators with the ability to learn, better adapting to multi - objective optimization problems, and enhancing the flexibility of the genetic operators. Compared with traditional multi - objective evolutionary optimization algorithms, the selection of traditional selection operators mostly only focuses on the quality of the current population and cannot make selections based on deeper information. By neural network - izing the genetic operators of the meta - learning multi - objective genetic algorithm, the multi - layer perceptron network in the neural network can select a more potential next - generation population based on deeper levels, which can better solve multi - objective optimization problems. By taking the network parameters of the neural network - ized genetic operators as the optimization task of the outer loop and the evaluation index of the inner loop as the objective function value of the outer loop, the purpose of training the network parameters in the multi - objective optimization model based on meta - learning without gradient information can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a schematic diagram of the steps of a multi - objective optimization method based on meta - learning provided by an embodiment of the present invention;
[0040] Figure 2 is a schematic diagram of the steps of the training process of a multi - objective optimization model based on meta - learning provided by an embodiment of the present invention;
[0041] Figure 3 is a schematic diagram of the structure of a multi - objective optimization model based on meta - learning provided by an embodiment of the present invention;
[0042] Figure 4 is a schematic diagram of the structure of the multi - head self - attention mechanism network in the multi - objective optimization model based on meta - learning provided by an embodiment of the present invention;
[0043] Figure 5 It is a schematic structural diagram of a multi-layer perceptron network in a multi-objective optimization model based on meta-learning provided by an embodiment of the present invention. Specific embodiments
[0044] The following further describes the present invention in detail with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0045] To solve the problem that in the prior art, multi-objective evolutionary optimization algorithms still require manual design of operators and have poor generalization ability, a multi-objective optimization method based on meta-learning provided by an embodiment of the present invention, as Figure 1 shown, may include the following steps:
[0046] S1. Obtain the data to be optimized corresponding to the multi-objective optimization task.
[0047] Specifically, S1 may include:
[0048] First, perform mathematical modeling according to the multi-objective optimization task to be solved. For example, when facing a space design task, it is necessary to perform mathematical modeling on the lift, drag, and structural strength of the wing; when facing a path planning task, it is necessary to perform mathematical modeling on transportation costs and time efficiency; when facing a financial investment problem, perform mathematical modeling on input cost risks and expected returns. The multi-objective optimization problem to be solved is obtained from the obtained mathematical modeling, so as to obtain the data to be optimized corresponding to the multi-objective optimization task.
[0049] S2. Use a pre-trained multi-objective optimization model based on meta-learning to process the data to be optimized to obtain an optimization result; wherein,
[0050] The multi-objective optimization model based on meta-learning adopts an inner and outer double-loop structure. In the inner loop, the genetic operators of the meta-learning multi-objective genetic algorithm are neural networked; the network parameters of the neural networked genetic operators are used as the optimization tasks of the outer loop, and the evaluation index of the inner loop is used as the objective function value of the outer loop to realize the training of the multi-objective optimization model based on meta-learning.
[0051] The training process of the multi-objective optimization model based on meta-learning, as Figure 2 shown, may include:
[0052] S01. Use the parameters of the multi-objective optimization model based on meta-learning as the individuals of the outer loop population, and randomly initialize the individuals of the outer loop population.
[0053] S02. Use the objectives of the multi-objective optimization task as the individuals of the inner loop population, initialize and evaluate the inner loop population to obtain the initial parameters of the inner loop.
[0054] For the structural schematic diagram of the multi-objective optimization model based on meta-learning, please refer to Figure 3 As shown, it can be seen that by taking the objectives of the multi-objective optimization task as the individuals of the inner-loop population and iteratively optimizing the individuals of this population in subsequent inner-loop steps, the optimization of the individuals of this population can be achieved. It can be understood that for each outer loop, it is necessary to initialize and evaluate the individuals of the inner-loop population to obtain the corresponding initial parameters of the inner loop.
[0055] S03, take the individuals of the initialized outer-loop population as the neural network weight parameters of the genetic operator, load them into the genetic operator of the meta-learning multi-objective genetic algorithm, and use the loaded genetic operator to perform several inner-loop processes on the initial parameters of the inner loop until the number of loops reaches the preset number of loops, and output the iterated inner-loop population.
[0056] Specifically, the genetic operator may include:
[0057] Sampling operator, mutation operator, and selection operator; among them,
[0058] Both the sampling operator and the mutation operator are constructed based on the self-attention mechanism;
[0059] The selection operator is constructed based on a multi-layer perceptron.
[0060] The structure of the inner loop is as shown in Figure 3 the lower half part. Take the parameters of the multi-objective optimization model based on meta-learning as the neural network weight parameters of the genetic operator, load them into the genetic operator of the meta-learning multi-objective genetic algorithm, and successively use the learnable sampling operator based on the self-attention mechanism, the learnable mutation operator based on the self-attention mechanism, and the learnable selection operator based on the multi-layer perceptron for processing respectively. After several inner loops, output the iterated inner-loop population.
[0061] Specifically, the initial parameters of the inner loop may include:
[0062] Population vector X of size N P and the objective function values of m objective functions Non-dominated rank r of the population P and initial mutation rate σ P .
[0063] The process of performing one inner-loop process on the initial parameters of the inner loop by using the loaded genetic operator may include:
[0064] S001, sample the initial parameters by using the loaded sampling operator to obtain the parental population vector Parent objective function value Parent non - dominated rank and the parent mutation rate
[0065] Specifically, the multi - head self - attention mechanism network in the sampling operator and the mutation operator, as Figure 4 shown, may include:
[0066] An input linear layer, a scaled dot - product attention mechanism, a fusion splicing module, and an output linear layer connected in sequence; where
[0067] The input linear layer is composed of three linear layers, and the scaled dot - product attention mechanism is composed of h attention heads; h represents a hyperparameter, h = 2 n , n≥1. The size of the network can be controlled by setting the number of attention heads h.
[0068] Both the sampling operator and the mutation operator proposed in the embodiments of the present invention utilize the multi - head self - attention mechanism network. The multi - head self - attention mechanism network is an important part of the Transformer architecture and performs very well in processing data. The embodiments of the present invention adopt the basic structure of the multi - head self - attention mechanism network. During the processing of this multi - head self - attention mechanism network, a linear transformation is performed on the input of the network to obtain three vectors Q (query), K (key), and V (value). After the above - mentioned linear transformation, the scaled dot - product attention mechanism is used to perform scaled dot - product attention calculation on Q (query), K (key), and V (value) to obtain attention scores.
[0069] Specifically, a similarity matrix is generated by performing the matrix product of Q and K. Divide this similarity matrix by (the dimension of K), use the softmax function to normalize the result of the division to obtain a normalized matrix, and multiply the normalized matrix by V to obtain the corresponding attention scores; where the expression of the attention scores is as follows:
[0070]
[0071] After obtaining the attention scores, aggregate the attention scores corresponding to the scaled dot - product attention on h attention heads, and through linear transformation, obtain the output of the multi - head self - attention mechanism network.
[0072] For the sampling operator, its input can include the comparison result between the objective function values in the current generation population and those in the previous generation population data, as well as the domination rank and objective function values in the current generation population. By using the multi-head self-attention mechanism network to learn from the above information and applying the acquired knowledge to the sampling probability. For the comparison result, since there are multiple objective values, the domination relationship is selected in the sampling operator to determine whether the objective function value result in the current generation population is better than that in the previous generation population. For the objective function values in the current generation population, the arctangent normalization is used to reduce the numerical magnitude differences between different problems and objective functions. In the sampling operator, the current population is sampled according to the obtained sampling probability, and at the same time, the objective function values, non-domination ranks, and mutation rates of the current population are also sampled.
[0073] S002. Use the loaded mutation operator to process the objective function values of the parent Parent non-domination rank and the parent mutation rate to obtain a new mutation rate σ P '; Use the new mutation rate σ P ' to perform Gaussian mutation on the parent population vector to obtain the mutated offspring population vector X P '; Evaluate the mutated offspring population vector X P ' to obtain the objective function values of the mutated offspring
[0074] Specifically, for the mutation operator, its input consists of the objective function values of the parent obtained by sampling and the comparison result with the objective function values of the previous generation population , the parent non-domination rank and the parent mutation rate . In the mutation operator, the domination relationship is used to determine whether the current objective function value is better than that of the previous generation. After learning the characteristics of the above population information, the mutation operator uses the multi-head self-attention mechanism network to dynamically adjust the mutation rate.
[0075] It can be understood that in Figure 3 and Figure 4 , the same color is used to represent the same structure for easy understanding of the specific composition of the sampling operator and the mutation operator.
[0076] S003. Use the loaded selection operator to process the objective function values of the mutated offspring to obtain the population vector X C of the next generation, the objective function value the non-domination rank r of the population Cand the population mutation rate σ C , may include:
[0077] Input the objective function value of the mutated offspring and the crowding distance of the individuals in the mutated offspring population into the loaded selection operator to obtain the population vector X of the next generation and the objective function value C
[0078]
[0079] Perform non - dominated sorting on the population vector of the next generation to obtain the non - dominated rank r of the population of the next generation C ;
[0080] The loaded selection operator obtains the mutation rate σ of the population of the next generation from the parental mutation rate C .
[0081] Specifically, the multi - layer perceptron network in the selection operator, as shown in Figure 5 , may include:
[0082] An input layer, a hidden layer, and an output layer connected in sequence.
[0083] The multi - layer perceptron network (MLP) is a feed - forward neural network, including an input layer, multiple hidden layers, and an output layer. Each layer includes many neurons with full interconnection. In the embodiments of the present invention, the multi - layer perceptron network is incorporated into the multi - objective genetic algorithm, and its learning ability is used to extract feature information from the population, so that the multi - layer perceptron network plays a crucial role in selecting the population of the next generation. Specifically, the multi - layer perceptron network can be used to perform selection, which is equivalent to a binary classification task. This task can divide the population into two groups: "tending to be selected" and "tending not to be selected", with the aim of maximizing the inclusion of individuals "tending to be selected" to generate a favorable population of the next generation.
[0084] Specifically, taking a multi - objective optimization problem with two - objective problems as an example, in the input layer, the objective function value of the parent population before mutation and the objective function value after mutation are used as the input of the multi - layer perceptron network. In addition, the crowding distance can also be selected as an additional input of the multi - layer perceptron network. The multi - layer perceptron network can learn from the objective function values and crowding distances received by itself, extract feature information and complex relationships, and use the learned knowledge to help select the population of the next generation. Multiple layers are set in the hidden layer of the selection operator to increase the capacity of the network so that it can learn more complex features. In the hidden layer, the ReLU function is selected as the activation function. Compared with the Sigmoid function, this activation function is more in line with the biological neural activation mechanism, has higher computational efficiency, and is not prone to overfitting.
[0085] At the output layer of the selection operator, the embodiments of the present invention process this process in a manner similar to a binary classification task, using the softmax function to convert the output of the hidden layer into two probabilities p1 and p2. Among them, p1 represents the probability that an individual is selected, and p2 represents the probability that an individual is not selected. Sort the probabilities p1 corresponding to the population, and select the top N individuals as the next-generation population. The multi-layer perceptron network can learn deeper information about the current population, enabling the learning-based selection to select more suitable and promising individuals, thereby improving the future performance of the algorithm. Moreover, using the crowding distance as an additional input to the multi-layer perceptron network can enable the selection operator to better handle the current performance and achieve better population diversity, so that the multi-objective optimization model based on meta-learning has higher performance.
[0086] S04. Perform a hypervolume evaluation on the meta-learning-based multi-objective optimization model corresponding to the population after the inner loop iteration to obtain a hypervolume score. If the hypervolume score meets the preset conditions, use the current meta-learning-based multi-objective optimization model as the pre-trained meta-learning-based multi-objective optimization model; if the hypervolume score does not meet the preset conditions, return to step S02 to optimize the parameters of the meta-learning-based multi-objective optimization model, replace the individuals in the outer loop population with the optimized parameters of the meta-learning-based multi-objective optimization model, and repeat steps S02 - S04 until the obtained hypervolume score meets the preset conditions or the number of repeated executions reaches the preset number of executions to obtain the pre-trained meta-learning-based multi-objective optimization model.
[0087] After performing the inner loop processing that meets the preset number of loops, perform a hypervolume evaluation on the meta-learning-based multi-objective optimization model corresponding to the population after the inner loop iteration to obtain a hypervolume score; compared with other metrics in the field of multi-objective optimization (such as generational distance, inverted generational distance, etc.), the hypervolume score does not require prior knowledge of the actual optimal solution set (Pareto front) of the current multi-objective optimization problem.
[0088] The multi-objective optimization model based on meta-learning proposed by the embodiments of the present invention adopts an inner and outer double-loop structure. In the inner loop, the genetic operator of the meta-learning multi-objective genetic algorithm is neural networked, enabling the genetic operator to adaptively adjust parameters, endowing the genetic operator with the ability to learn, better adapting to multi-objective optimization problems, and enhancing the flexibility of the genetic operator. Compared with traditional multi-objective evolutionary optimization algorithms, the selection of traditional selection operators mostly only focuses on the advantages and disadvantages of the current population and cannot make selections based on deeper information. By neural networking the genetic operator of the meta-learning multi-objective genetic algorithm, the multi-layer perceptron network in the neural network can select a more promising next-generation population based on deeper information, which can better solve multi-objective optimization problems. By taking the network parameters of the neural networked genetic operator as the optimization task of the outer loop and the evaluation index of the inner loop as the objective function value of the outer loop, the purpose of training the network parameters in the multi-objective optimization model based on meta-learning without gradient information can be achieved.
[0089] It should be noted that in the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0090] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A multi-objective optimization method based on meta-learning, characterized in that: include: Obtain the data to be optimized corresponding to the multi-objective optimization task; The data to be optimized is processed using a pre-trained meta-learning-based multi-objective optimization model to obtain an optimization result; wherein, The meta-learning-based multi-objective optimization model adopts an inner and outer double-layer loop structure. In the inner loop, the genetic operator of the meta-learning multi-objective genetic algorithm is neuralized; the network parameters of the neural networked genetic operator are used as the optimization task of the outer loop, and the evaluation index of the inner loop is used as the objective function value of the outer loop, so as to realize the training of the meta-learning-based multi-objective optimization model.
2. The multi-objective optimization method based on meta-learning according to claim 1, characterized in that: The training process of the multi-objective optimization model based on meta-learning includes: S01, taking the parameters of the multi-objective optimization model based on meta-learning as individuals of the outer loop population, and randomly initializing the individuals of the outer loop population; S02, taking the objective of the multi-objective optimization task as an individual of the inner loop population, initializing and evaluating the inner loop population to obtain the initial parameters of the inner loop; S03, using the individuals of the initialized outer loop population as the neural network weight parameters of the genetic operator, loading them into the genetic operator of the meta-learning multi-objective genetic algorithm, using the loaded genetic operator to perform several inner loop processing on the initial parameters of the inner loop until the number of loops reaches a preset number of loops, and outputting the iterated inner loop population; S04, perform hypervolume evaluation on the meta-learning-based multi-objective optimization model corresponding to the iterated inner loop population to obtain a hypervolume score. If the hypervolume score meets the preset conditions, use the current meta-learning-based multi-objective optimization model as a pre-trained meta-learning-based multi-objective optimization model. If the hypervolume score does not meet the preset conditions, return to step S02, optimize the parameters of the meta-learning-based multi-objective optimization model, replace the individuals of the outer loop population with the optimized parameters of the meta-learning-based multi-objective optimization model, and repeat steps S02-S04 until the obtained hypervolume score meets the preset conditions or the number of repeated executions reaches the preset number of executions, and obtain the pre-trained meta-learning-based multi-objective optimization model.
3. The multi-objective optimization method based on meta-learning according to claim 2, characterized in that: The genetic operator comprises: Sampling operator, mutation operator and selection operator; among them, The sampling operator and mutation operator are both based on the self-attention mechanism; The selection operator is based on a multi-layer perceptron.
4. The multi-objective optimization method based on meta-learning according to claim 3, characterized in that: The multi-head self-attention mechanism network in the sampling operator and the mutation operator includes: The input linear layer, scaled dot product attention mechanism, fusion concatenation module and output linear layer are connected in sequence; among them, The input linear layer consists of three linear layers, and the scaled dot product attention mechanism consists of h attention heads; h represents a hyperparameter, h = 2 n , n≥1.
5. The multi-objective optimization method based on meta-learning according to claim 3, characterized in that: The multilayer perceptron network in the selection operator includes: The input layer, hidden layer, and output layer are connected in sequence.
6. The multi-objective optimization method based on meta-learning according to claim 3, characterized in that: The initial parameters of the inner loop include: The population vector X of size N P , the objective function value of m objective functions The non-dominated level r of the population P and the initial mutation rate σ P .
7. The multi-objective optimization method based on meta-learning according to claim 6, characterized in that: The process of performing one inner loop among several inner loop processes on the initial parameters of the inner loop using the loaded genetic operator includes: The loaded sampling operator is used to sample the initial parameters to obtain the parent population vector Parent objective function value Parental non-dominant rank and parental mutation rate Use the loaded mutation operator to modify the parent objective function value Parental non-dominant rank and parental mutation rate Processing, get the new mutation rate σ P ′; Using the new mutation rate σ P ′ for the parent population vector Perform Gaussian mutation to obtain the mutated offspring population vector X P ′; for the mutated offspring population vector X P ′ is evaluated to obtain the target function value of the offspring after mutation Use the loaded selection operator to change the mutated offspring objective function value Processing is performed to obtain the next generation population vector X C , objective function value The non-dominated level r of the population C and the population mutation rate σ C .
8. The multi-objective optimization method based on meta-learning according to claim 7, characterized in that: The loaded selection operator is used to mutate the target function value of the offspring Processing is performed to obtain the next generation population vector X C , objective function value The non-dominated level r of the population C and mutation rate σ C ,include: The mutated offspring objective function value The crowding distance of the individuals in the offspring population after mutation is input into the loaded selection operator to obtain the population vector X of the next generation C and the objective function value Perform non-dominated sorting on the population vector of the next generation to obtain the non-dominated level r of the population of the next generation C ; The loaded selection operator is derived from the parent mutation rate The mutation rate σ of the next generation population is obtained from C .