Multi-target neural architecture search method based on decomposition dual-archive guidance

By adopting a decomposition-based dual archive guidance method in multi-objective neural architecture search, combining clustering and reference vector strategies, the problems of efficiency and mass balance in high-dimensional discrete search space are solved, and a more comprehensive multi-objective optimization effect is achieved.

CN120147675APending Publication Date: 2025-06-13YANSHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510208898.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to balance search efficiency and quality in high-dimensional discrete search spaces, and multi-objective optimization methods have challenges in convergence and diversity balance.

Method used

A multi-objective neural architecture search method based on decomposition is adopted, combined with the clustering method and reference vector strategy, by initializing the population, updating the convergence and diversity archives, cyclically optimized until the termination condition is reached, and the optimal solution set is output.

Benefits of technology

It effectively avoids small model traps and early convergence problems in traditional methods, improves the global optimization ability of search efficiency and results, and enhances the diversity of balanced solutions in multi-objective optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147675A_ABST
    Figure CN120147675A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-target neural architecture search method based on decomposition dual-archive guidance, and relates to the field of multi-target neural architecture search, which comprises the following steps: initializing a population to generate a random population; according to the random population and a preset archive size, preliminarily updating the convergence archive; preliminarily updating the diversity archive according to the random population and preset problem parameters; and respectively and circularly executing optimization updating on the initially updated convergence archive and diversity archive until a termination condition is reached, and outputting an optimal solution set. According to the method, the adaptive differential evolution algorithm is adopted, discrete architecture coding is directly operated, the problem of low efficiency of a traditional gradient-based optimization method in a discrete decision space is effectively solved, and by comprehensively exploring the discrete search space, a local optimal trap is avoided, and the search efficiency of the algorithm and the global optimization capability of a result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multi-objective neural architecture search. Specifically, it relates to a multi-objective neural architecture search method based on decomposition and dual-archive guidance. Background Art

[0002] With the rapid development of computer vision technology, neural networks have achieved remarkable results in tasks such as image classification. However, traditional neural network design methods rely on expert experience and a large number of experiments, and optimize performance by manually adjusting the network structure. This method is inefficient and limited by the knowledge and experience of designers, and it is impossible to explore the globally optimal network architecture in a short time. Therefore, how to automatically discover the optimal neural network architecture has become an important topic in deep learning research.

[0003] The emergence of Neural Architecture Search (NAS) technology provides an innovative solution to this problem. By using search algorithms, NAS can automatically find the optimal network architecture in a huge search space. Traditional NAS methods usually rely on three main components: the search space, the search strategy, and the performance evaluation strategy. The search space defines all possible combinations of neural architectures, the search strategy determines how to efficiently select the optimal architecture from this space, and the performance evaluation strategy is used to evaluate the performance of candidate architectures on specific tasks. However, since NAS is essentially a discrete optimization problem with a huge search space and extremely high computational complexity, traditional single-objective optimization methods often struggle to meet the multi-objective optimization requirements.

[0004] To address this computational complexity problem, researchers have proposed various optimization methods. For example, DARTS (Differentiable Architecture Search) transforms the architecture search problem into a differentiable continuous optimization problem and uses gradient descent to accelerate the search process, effectively reducing the computational overhead. Although DARTS has improved in computational efficiency, it is still a single-objective optimization-based method that usually focuses on optimizing the accuracy of the network while ignoring other key factors such as model size and inference latency.

[0005] With the rise of multi-objective optimization methods, NAS has started to focus not only on a single network performance objective but also on the trade-off between multiple optimization objectives. Researchers have proposed multi-objective optimization-based neural architecture search methods such as NSGA-Net and NAT, etc. These methods can find a balance between multiple objectives. For example, NSGA-Net optimizes the floating-point operation count and classification error rate on the CIFAR-10 dataset through the non-dominated sorting genetic algorithm (NSGA-II), considering both the accuracy of the network and the consumption of computing resources. NAT adopts the NSGA-III algorithm to optimize multiple objectives such as model accuracy, number of parameters, and number of multiply-accumulate operations on the ImageNet dataset. These multi-objective optimization methods provide a more comprehensive performance evaluation for NAS, can find a balance between multiple objectives, and avoid the limitations of relying solely on a single metric.

[0006] However, the above multi-objective optimization methods still face some challenges in the high-dimensional discrete search space, especially in balancing search efficiency and optimization quality. Specifically, multi-objective optimization methods need to simultaneously focus on the balance between convergence and diversity. Therefore, there is an urgent need to design a multi-objective neural architecture search method that can effectively improve search efficiency in a complex search space while maintaining the diversity and convergence of multi-objective optimization. Summary of the Invention

[0007] (1) Technical problems to be solved

[0008] Aiming at the deficiencies of the prior art, the present invention provides a decomposition-based dual-archive-guided multi-objective neural architecture search method, which has the advantages of combining clustering methods and reference vector strategies to effectively balance convergence and diversity in the discrete search space, thereby solving the problems that the prior art cannot balance search efficiency and quality and the balance between convergence and diversity of multi-objective optimization in the high-dimensional discrete search space.

[0009] (2) Technical solutions

[0010] To achieve the above advantages of combining clustering methods and reference vector strategies to effectively balance convergence and diversity in the discrete search space, the specific technical solutions adopted by the present invention are as follows:

[0011] A decomposition-based dual-archive-guided multi-objective neural architecture search method, the method includes:

[0012] Initialize the population to generate a random population;

[0013] Compare the number of solutions in the convergence archive with the preset archive size according to the random population and the preset archive size, and preliminarily update the convergence archive based on the comparison result;

[0014] According to the random population and the preset problem parameters, the diversity archive is initially updated using the clustering algorithm and the adaptive differential evolution algorithm;

[0015] The optimization update is repeatedly performed on the initially updated convergence archive and diversity archive until the termination condition is reached and then stopped, and the optimal solution set is output.

[0016] Preferably, according to the random population and the preset archive size, the number of solutions in the convergence archive is compared with the preset archive size. Based on the comparison result, the initial update of the convergence archive includes:

[0017] The solutions in the convergence archive and the newly generated solutions are combined to form a first candidate solution set;

[0018] The first candidate solution set is non-dominated sorted, and the solutions located on the non-dominated front are retained to ensure that only non-dominated solutions are included in the convergence archive;

[0019] The number of solutions in the convergence archive is compared with the preset maximum archive. If the number of solutions in the convergence archive is less than or equal to the preset maximum archive, the initially updated convergence archive is output; otherwise, the non-dominated solutions in the convergence archive are screened using the ISDE+ metric to control the archive size.

[0020] Preferably, screening the non-dominated solutions in the convergence archive using the ISDE+ metric includes:

[0021] The objective values in the convergence archive are standardized;

[0022] Calculate the objective value differences between each pair of solutions in the convergence archive, and calculate the ISDE+ metric value for each solution based on the objective value differences;

[0023] Gradually screen the solutions with higher ISDE+ values in the convergence archive according to the ISDE+ metric values, and use the finally retained solution set as the initially updated convergence archive.

[0024] Preferably, according to the random population and the preset problem parameters, the initial update of the diversity archive using the clustering algorithm and the adaptive differential evolution algorithm includes:

[0025] The solutions in the diversity archive and the newly generated solutions are combined to form a second candidate solution set, and the second candidate solution set is non-dominated sorted to ensure that only non-dominated solutions are included in the diversity archive;

[0026] Calculate the ideal point in the diversity archive, and dynamically allocate the individuals in the diversity archive to the corresponding sub-problems;

[0027] Normalize the target values in the convergence archive, cluster the target values in the diversity archive using a clustering algorithm to obtain a clustering result, and generate an assignment relationship of reference vectors according to the clustering result;

[0028] Generate a local mating pool and a global mating pool based on the assignment relationship of reference vectors, and generate parents through the local mating pool and the global mating pool;

[0029] Generate offspring using an adaptive differential evolution algorithm, calculate the fitness of the offspring, and update the parameters of the adaptive differential evolution algorithm according to the fitness of the offspring to obtain the diversity archive after the initial update.

[0030] Preferably, cluster the target values in the diversity archive using a clustering algorithm, and the obtained clustering result includes:

[0031] Calculate the similarity matrix between the solutions in the diversity archive, and initialize the parameters of the clustering algorithm according to the similarity matrix;

[0032] Iteratively calculate and update the responsibility value and support degree value of each solution to the reference vector until the clustering algorithm converges and stops, and select the clustering center according to the maximum responsibility value and support degree value;

[0033] Calculate the cosine similarity between the clustering center and the reference vector, and assign the solution to the nearest reference vector to realize the association between the solution and the reference vector.

[0034] Preferably, generate offspring using an adaptive differential evolution algorithm, calculate the fitness of the offspring, and update the parameters of the adaptive differential evolution algorithm according to the fitness of the offspring, including:

[0035] Extract the decision variables of the parents, and initialize the boundaries of the decision variables according to the preset upper and lower boundaries;

[0036] Perform the crossover operation of differential evolution, randomly select positions and control the probability of crossover with the crossover rate, and generate offspring based on the differences between the parents at the selected positions;

[0037] Perform polynomial mutation, control the decision variables to be mutated with the mutation probability, and control the amplitude of mutation according to the distribution index to ensure that the generated offspring variables are within the specified upper and lower boundaries;

[0038] Calculate the fitness of the offspring, and update the parameters of the adaptive differential evolution algorithm according to the fitness of the offspring.

[0039] Preferably, updating the parameters of the adaptive differential evolution algorithm includes:

[0040] Randomly select the historical crossover rate and historical scaling factor, and perform Gaussian distribution and tangent transformation processing in turn to generate new crossover rate and scaling factor;

[0041] Calculate the fitness difference between the parent generation and the offspring generation, record the new crossover rate and scaling factor corresponding to the parent generation with a fitness greater than that of the offspring generation, and update the values of the crossover rate and scaling factor in the adaptive differential evolution algorithm by means of weighted summation.

[0042] Preferably, perform optimization updates on the convergence archive and the diversity archive after preliminary update in a loop until the termination condition is reached and then stop. The output of the optimal solution set includes:

[0043] Select the convergence archive parent generation and the diversity parent generation from the convergence archive and the diversity archive after preliminary update;

[0044] Generate the convergence offspring generation and the diversity offspring generation from the convergence archive parent generation and the diversity parent generation using the genetic parameter set in the genetic algorithm;

[0045] Update the convergence archive using the convergence offspring generation, and update the diversity archive using the diversity offspring generation. Iteratively perform the optimization updates on the convergence archive and the diversity archive until the termination condition is reached and then stop, and output the optimal solution set.

[0046] Preferably, the calculation formula for the ISDE+ index value is:

[0047]

[0048] In the formula, ISDE represents the ISDE+ index value of the solution, I(i,:) represents the row of the difference matrix between the i-th solution and all other solutions, C(i) represents the maximum difference value between the i-th solution and other solutions, and N represents the size of the solution set.

[0049] Preferably, the calculation formula for the responsibility value is:

[0050]

[0051] The calculation formula for the support value is:

[0052]

[0053] 式中, r(i,k) 表示解i对参考向量k的责任值, a(i,k) 表示解i对参考向量k的支持度值, s(i,k) 表示解i对参考向量k的相似度矩阵, a(i,k′) 表示解i对参考向量 k′ 的支持度值, s(i,k′) 表示解i对参考向量 k′ 的相似度, r(i′,k) 表示解 i′ 对参考向量k的责任值。

[0054] (III) Beneficial effects

[0055] Compared with the prior art, the present invention provides a multi-objective neural architecture search method based on decomposition and dual-archive guidance, having the following beneficial effects:

[0056] (1) Aiming at the small model trap and early convergence problems in neural architecture search, the present invention provides a multi-objective neural architecture search method based on dual-archive guidance, thereby effectively avoiding the premature convergence phenomenon caused by small models in traditional genetic algorithms, ensuring the balance between convergence and diversity during the search process, and further exploring potential optimal architectures more comprehensively.

[0057] (2) When facing multi-objective optimization problems, through the multi-objective neural architecture search method based on dual-archive guidance provided by the present invention, the target problem is transformed into multiple simple sub-problems for solution, thereby effectively avoiding the dominance resistance phenomenon and solution space aggregation problem in the high-dimensional sparse target space of traditional MOEA algorithms, improving the uniformity and diversity of solutions, and further enhancing the overall optimization effect.

[0058] (3) The present invention adopts an adaptive differential evolution algorithm to directly operate on the discrete architecture encoding, effectively solving the low efficiency problem of traditional gradient-based optimization methods in the discrete decision space, and avoiding local optimal traps by comprehensively exploring the discrete search space, improving the search efficiency of the algorithm and the global optimization ability of the results.

[0059] (4) Through the optimization strategy guided by dual archives, the present invention can improve the balance of neural network architectures on multiple optimization objectives (such as performance, number of parameters, inference latency, etc.), solve the defects of existing multi-objective optimization algorithms in high-dimensional discrete search spaces, and is widely applied to fields such as generative artificial intelligence, automated design and optimization of deep learning models. Especially in the case of limited computing resources, it improves the efficiency and effect of neural network architecture search. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0061] Figure 1 is a flowchart of the multi-objective neural architecture search method based on decomposition and dual-archive guidance according to an embodiment of the present invention;

[0062] Figure 2 is a flowchart of the loop optimization part in the multi-objective neural architecture search method based on decomposition and dual-archive guidance according to an embodiment of the present invention;

[0063] Figure 3 It is a framework diagram of the update strategy for the convergence archive and the diversity archive in the decomposition-based dual-archive-guided multi-objective neural architecture search method according to an embodiment of the present invention;

[0064] Figure 4 It is a schematic diagram of the AP clustering algorithm in the decomposition-based dual-archive-guided multi-objective neural architecture search method according to an embodiment of the present invention;

[0065] Figure 5 It is one of the schematic diagrams of screening solutions by the ISDE+ metric in the convergence archive of the decomposition-based dual-archive-guided multi-objective neural architecture search method according to an embodiment of the present invention;

[0066] Figure 6 It is the second schematic diagram of screening solutions by the ISDE+ metric in the convergence archive of the decomposition-based dual-archive-guided multi-objective neural architecture search method according to an embodiment of the present invention;

[0067] Figure 7 It is a schematic diagram of dynamically decomposing problems by reference vectors in the diversity archive of the decomposition-based dual-archive-guided multi-objective neural architecture search method according to an embodiment of the present invention;

[0068] Figure 8 It is one of the convergence curve graphs of the average HV values of six algorithms on the C10MOP test problem in the decomposition-based dual-archive-guided multi-objective neural architecture search method according to an embodiment of the present invention;

[0069] Figure 9 It is the second convergence curve graph of the average HV values of six algorithms on the C10MOP test problem in the decomposition-based dual-archive-guided multi-objective neural architecture search method according to an embodiment of the present invention;

[0070] Figure 10 It is the third convergence curve graph of the average HV values of six algorithms on the C10MOP test problem in the decomposition-based dual-archive-guided multi-objective neural architecture search method according to an embodiment of the present invention;

[0071] Figure 11 It is a comparison graph of the HV values of eight algorithms on the C10MOP test problem in the decomposition-based dual-archive-guided multi-objective neural architecture search method according to an embodiment of the present invention;

[0072] Figure 12 It is a comparison graph of the HV values of eight algorithms on the IN1KMOP test problem in the decomposition-based dual-archive-guided multi-objective neural architecture search method according to an embodiment of the present invention;

[0073] Figure 13It is a comparison chart of the HV values of 8 algorithms in the CitySegMOP test problem for the multi-objective neural architecture search method based on decomposition-based dual-archive guidance according to an embodiment of the present invention. Detailed implementation manners

[0074] To further illustrate each embodiment, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. They are mainly used to illustrate the embodiments and can be combined with the relevant descriptions in the specification to explain the operating principle of the embodiments. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention.

[0075] According to an embodiment of the present invention, a multi-objective neural architecture search method based on decomposition-based dual-archive guidance is provided.

[0076] Specifically, the multi-objective neural architecture search (MONAS) problem can be formulated as:

[0077]

[0078] In the formula, Ω represents the search space, x and w(x) represent the encoded network architecture (i.e., the decision vector) and the corresponding weights of the decoded network, and w * (x) represents the optimal weights when x is trained on the training set D train to minimize the loss function; f e represents the objective function of the model prediction error, and f c and f H respectively represent the objective functions of the performance related to the model complexity and the hardware device.

[0079] Among them, Algorithm 1: The overall framework of the decomposition-based dual-archive guidance multi-objective neural architecture search method (DANAS) includes:

[0080] 1. Generate a random population Population.

[0081] 2. Initialize and update the convergence archive CA according to the generated population Population and the set archive size CAsize.

[0082] 3. Initialize and update the diversity archive DA according to the population Population and the problem parameters Problem.

[0083] 4. When the algorithm has not terminated, repeat the following steps:

[0084] Select two parent individuals ParentA and ParentB from the convergence archive CA and the diversity archive DA for the next operation according to the population size N;

[0085] Using a genetic algorithm (GA), offspring Offspring are generated respectively based on ParentA and ParentB using different sets of genetic parameters (e.g., {1, 20, 0, 0} and {0, 0, 1, 20}).

[0086] Update the convergence archive CA according to the new offspring Offspring.

[0087] Update the diversity archive DA according to the new offspring Offspring and the problem parameter Problem.

[0088] 5. When the algorithm terminates, output the diversity archive DA.

[0089] The present invention will be further described in conjunction with the accompanying drawings and specific embodiments, as Figure 1 shown. The multi-objective neural architecture search method based on decomposition and dual-archive guidance according to an embodiment of the present invention includes:[[]]

[0090] S1. Initialize the population to generate a random population.

[0091] It should be noted that the initial population is a randomly generated initialization candidate neural network architecture.

[0092] S2. Compare the number of solutions in the convergence archive with the preset archive size according to the random population and the preset archive size, and preliminarily update the convergence archive based on the comparison result.

[0093] Among them, comparing the number of solutions in the convergence archive with the preset archive size according to the random population and the preset archive size, and preliminarily updating the convergence archive based on the comparison result includes:[[]]

[0094] Merge the solutions in the convergence archive with the newly generated solutions to form a first candidate solution set.

[0095] It should be noted that before the archive is updated in a loop, the input to the convergence archive is the generated random population, and there are no newly generated solutions at this time. When the archive is updated in a loop, the newly generated solutions generated each time will be merged with the original solutions in the convergence archive when the convergence archive is updated.

[0096] Perform non-dominated sorting on the first candidate solution set, and retain the solutions located on the non-dominated front to ensure that only non-dominated solutions are included in the convergence archive;

[0097] Compare the number of solutions in the convergence archive with the preset maximum archive. If the number of solutions in the convergence archive is less than or equal to the preset maximum archive, output the preliminarily updated convergence archive; otherwise, screen the non-dominated solutions in the convergence archive through the ISDE+ index to control the archive size.

[0098] Among them, screening the non-dominated solutions in the convergence archive through the ISDE+ index includes:

[0099] Normalize the objective values in the convergence archive;

[0100] Calculate the difference in objective values between each pair of solutions in the convergence archive, and calculate the ISDE+ index value of each solution based on the difference in objective values.

[0101] Gradually screen the solutions with higher ISDE+ values in the convergence archive according to the ISDE+ index values, and use the finally retained solution set as the initially updated convergence archive.

[0102] To facilitate the understanding of the above technical solution of the present invention, the following will detail the initially updated convergence archive in the actual process of the present invention:

[0103] During the process of updating the convergence archive, in order to balance the convergence and diversity of the solutions in the archive, first merge the solutions in the current archive with the newly generated solutions to form a candidate solution set. Then, perform non-dominated sorting on the merged solution set, and only retain the solutions located on the first non-dominated front, so as to ensure that the archive only contains non-dominated solutions. If the number of solutions in the archive at this time does not exceed the set maximum archive size, directly output the updated archive; otherwise, continue to screen the solutions through the ISDE+ index to control the archive size.

[0104] To calculate the ISDE+ index, first normalize the objective values of the candidate solutions, adjust each objective value to the range of [0,1] to eliminate the influence of the objective scale. Then calculate the relative difference between every two solutions, which is represented by the maximum objective difference matrix I, where I(i,j) represents the maximum difference in objective values between solution i and solution j. Then calculate the maximum absolute value C of each column of the I matrix as the normalization factor to standardize the difference values between different solutions. Based on the normalized I and C, calculate the ISDE+ value of each solution, and this value quantifies the relative contribution of the solution through an exponential decay function. The smaller the value, the smaller the contribution of the solution to the overall archive.

[0105] During the archive reduction process, each time remove the solution with the smallest ISDE+ value, and dynamically update the ISDE+ values of the remaining solutions. Specifically, calculate its influence on other solutions through the I values of the removed solution and other solutions, and adjust the ISDE+ values of the remaining solutions with an exponential decay function to re-evaluate the importance of the solutions. This process is as Figures 5 - 6 shown. This process continues until the size of the archive meets the set maximum archive limit. The finally output archive contains a solution set that has been strictly screened in terms of both convergence and diversity.

[0106] The process of calculating the ISDE+ index is as follows:

[0107] Normalize the objective space and normalize the objective values in the remaining non-dominated solution set so that the objective values are within the range of [0, 1]. The normalization process is as follows:

[0108]

[0109] In the formula, CAObj represents the objective vectors of all solutions in the solution set. After normalization, the objective values of each column of CAObj will be mapped to the range of [0, 1].

[0110] Calculate the differences between each pair of solutions in the solution set, calculate the differences in objective values between every two solutions in the solution set, and the differences are calculated as follows:

[0111] I(i, j) = max(CAObj(i, :) - CAObj(j, :));

[0112] In the formula, I(i, j) represents the difference between the i-th solution and the j-th solution, which is used to measure the diversity of the solution set. CAObj represents the objective vectors of all solutions in the solution set; CAObj(i, :) represents the normalized values of the i-th solution on all objectives (i.e., the i-th row of the matrix); CAObj(j, :) represents the normalized values of the j-th solution on all objectives (i.e., the j-th row of the matrix).

[0113] Calculate the ISDE+ index value. According to the calculated difference matrix I, use the following formula to calculate the ISDE+ index value of each solution:

[0114]

[0115] In the formula, ISDE represents the ISDE+ index value of the solution; I(i, :) represents the row of the difference matrix between the i-th solution and all other solutions, C(i) represents the maximum difference value between the i-th solution and other solutions, N represents the size of the solution set, that is, the contributions of all solutions are accumulated, and 0.05 is a constant used to control the attenuation.

[0116] Filter solutions according to the ISDE+ index value. According to the calculated ISDE+ index value, gradually filter the solutions in the solution set and retain the solutions with higher ISDE+ values. The specific steps are as follows: Initialize a Choose vector to represent the current candidate solution set. In the solution set, gradually eliminate the solution with the smallest ISDE+ value until the size of the remaining solution set does not exceed the predetermined maximum capacity MaxSize. Each time a solution is eliminated, update the ISDE+ value:

[0117]

[0118] In the formula, Choose(x) represents the solution selected for elimination, and ISDE represents the ISDE+ value of the current solution.

[0119] Update the solution set. After screening, the finally retained solution set is the updated solution set, and its size does not exceed the maximum capacity MaxSize.

[0120] S3. According to the random population and the preset problem parameters, use the clustering algorithm and the adaptive differential evolution algorithm to preliminarily update the diversity archive.

[0121] Among them, using the clustering algorithm and the adaptive differential evolution algorithm to preliminarily update the diversity archive according to the random population and the preset problem parameters includes:

[0122] Merge the solutions in the diversity archive with the newly generated solutions to form a second candidate solution set, and perform non-dominated sorting on the second candidate solution set to ensure that only non-dominated solutions are included in the diversity archive;

[0123] Calculate the ideal point in the diversity archive, and dynamically allocate the individuals in the diversity archive to the corresponding sub-problems;

[0124] Standardize the objective values in the convergence archive, use the clustering algorithm to cluster the objective values in the diversity archive to obtain the clustering result, and generate the allocation relationship of the reference vectors according to the clustering result.

[0125] Among them, using the clustering algorithm to cluster the objective values in the diversity archive to obtain the clustering result includes:

[0126] Calculate the similarity matrix between the solutions in the diversity archive, and initialize the parameters of the clustering algorithm according to the similarity matrix;

[0127] Iteratively calculate and update the responsibility value and support value of each solution to the reference vector until the clustering algorithm converges and stops, and select the clustering center according to the maximum responsibility value and support value.

[0128] Among them, the calculation formula of the responsibility value is:

[0129]

[0130] The calculation formula of the support value is:

[0131]

[0132] 式中, r(i,k) 表示解i对参考向量k的责任值, a(i,k) 表示解i对参考向量k的支持度值, s(i,k) 表示解i对参考向量k的相似度矩阵, a(i,k′) 表示解i对参考向量 k′ 的支持度值, s(i,k′) 表示解i对参考向量 k′ 的相似度, r(i′,k) 表示解i′ 对参考向量k的责任值。

[0133] Calculate the cosine similarity between the clustering center and the reference vector, and assign the solution to the nearest reference vector to associate the solution with the reference vector.

[0134] Generate a local mating pool and a global mating pool based on the assignment relationship of the reference vectors, and generate parents through the local mating pool and the global mating pool;

[0135] Generate offspring using the adaptive differential evolution algorithm, calculate the fitness of the offspring, and update the parameters of the adaptive differential evolution algorithm according to the fitness of the offspring to obtain the initially updated diversity archive.

[0136] Among them, generating offspring using the adaptive differential evolution algorithm, calculating the fitness of the offspring, and updating the parameters of the adaptive differential evolution algorithm according to the fitness of the offspring include:

[0137] Extract the decision variables of the parents, and initialize the boundaries of the decision variables according to the preset upper and lower boundaries;

[0138] Perform the crossover operation of differential evolution. By randomly selecting positions and controlling the crossover probability with the crossover rate, generate offspring based on the differences between the parents at the selected positions;

[0139] Perform polynomial mutation. Control the decision variables to be mutated with the mutation probability, and control the mutation amplitude according to the distribution index to ensure that the generated offspring variables are within the specified upper and lower boundaries;

[0140] Calculate the fitness of the offspring, and update the parameters of the adaptive differential evolution algorithm according to the fitness of the offspring.

[0141] Among them, updating the parameters of the adaptive differential evolution algorithm includes:

[0142] Randomly select the historical crossover rate and historical scaling factor, and perform Gaussian distribution and tangent transformation processing on them in turn to generate new crossover rate and scaling factor;

[0143] Calculate the fitness difference between the parents and the offspring, record the new crossover rate and scaling factor corresponding to the parents with fitness greater than that of the offspring, and update the values of the crossover rate and scaling factor in the adaptive differential evolution algorithm by weighted summation.

[0144] To facilitate the understanding of the above technical solutions of the present invention, the following will give a detailed description of the initially updated diversity archive in the actual process of the present invention:

[0145] The diversity archive update strategy is shown in Algorithm 2 and Figure 3As shown. Merge the current diversity archive with the newly generated individuals, perform non-dominated sorting on the merged individuals, and only retain the non-dominated solutions; calculate the ideal point, and use Algorithm 3 to dynamically allocate the individuals in the DA archive to the corresponding sub-problems, where the ideal point vector z * represents the theoretical optimal value of all objective values in the objective space. For each objective function f i , the i-th dimensional coordinate of the ideal point is the minimum value of all current solutions on the i-th objective:

[0146] z * =(min f 1 (x), min f 2 (x),..., min f m (x));

[0147] In the formula, z * represents the ideal point vector, and f 1 (x), f 2 (x),..., f m (x) represent multiple objective functions.

[0148] Algorithm 2: Update the diversity archive using APClustering (Affinity Propagation Clustering Algorithm) and Adaptive Differential Evolution:

[0149] Input: Diversity archive DA, new solution set New, problem parameters Problem, memory value sf of the scaling factor, memory value cr of the crossover rate, memory value pos of the position.

[0150] Output: Updated diversity archive DA.

[0151] 1. Merge the new solution set New with the current diversity archive DA.

[0152] 2. Perform non-dominated sorting on the solutions in DA, and only retain the first-layer non-dominated solutions.

[0153] 3. Calculate the ideal point in the archive.

[0154] 4. Associate the solutions in the archive with the ideal point.

[0155] 5. Perform AP clustering on the objective values in DA to obtain the clustering results (including clusters, cluster centers, and indices).

[0156] 6. Determine the number of clusters K.

[0157] 7. For each reference vector, calculate its cosine similarity with all cluster centers and assign the nearest reference vector to each cluster.

[0158] 8. Perform the following operations on each cluster:

[0159] Determine the set of individuals conT corresponding to the reference vector closest to the cluster.

[0160] Count the number of solutions numT in the cluster.

[0161] If numT < S, randomly select individuals from the global population to fill.

[0162] If numT = S, directly select the solutions in conT.

[0163] If numT > S, randomly select S solutions from conT.

[0164] 9. Generate offspring according to the selected population.

[0165] 10. Calculate the fitness of the offspring.

[0166] 11. Adaptively update the scaling factor sf, crossover rate cr, and position pos according to the fitness of the offspring.

[0167] 12. Store the current cluster center.

[0168] Algorithm 3: Assignment of Solutions to Sub - problems:

[0169] Input: M: Number of objectives; N: Population size; V: Main reference vector; Va: Auxiliary reference vector; S: Size of each sub - problem; Z: Ideal point.

[0170] Output: Updated Population: Updated Population;

[0171] 1. Generate the reference vector V according to the number of objectives M;

[0172] 2. Generate the auxiliary reference vector Va according to the population size N;

[0173] 3. For each individual in the population:

[0174] Calculate the cosine similarity between each individual and each vector in V (main reference vector);

[0175] Calculate the cosine similarity between each individual and each vector in Va (auxiliary reference vector); (The calculation method of cosine similarity is the same as that in Algorithm 2)

[0176] 4. For each vector v in V:

[0177] Calculate the Euclidean distance to each vector in V a The calculation formula is as follows:

[0178]

[0179] where x i represents the auxiliary reference vector, y i represents the main reference vector, and M represents the dimension (number of targets) of the vector.

[0180] Find the vector nearest in Va Va (v).

[0181] 5. For each sub-problem i (from 1 to K):

[0182] Initialize the current solution set current as an empty set;

[0183] For each solution in the population:

[0184] If the cosine similarity of this solution with V(i) exceeds the threshold, add this solution to current;

[0185] If the size of current is less than S:

[0186] From the solutions associated with V a (i), find the solution closest to nearest Va (v);

[0187] For each of these solutions:

[0188] If this solution is not yet in current, add it;

[0189] If the size of current has reached S, terminate;

[0190] If the size of current is still less than S, randomly select solutions from the population to supplement it to size S;

[0191] If the size of current is greater than S:

[0192] Perform non-dominated sorting and crowding distance calculation on the solutions in current;

[0193] Retain the top S solutions with the highest crowding distance;

[0194] Assign current to partition(i);

[0195] 6. Return the updated population.

[0196] During the update process, if Figure 4As shown in the figure, first, the APC clustering algorithm is used for the target value matrix in the archive, the solutions in the archive are divided into several clusters, and the center points of each cluster are extracted. AP clustering determines the positions of cluster centers by constructing a similarity matrix between data points and using a message passing mechanism. The similarity matrix is obtained through the Euclidean distance, and the specific definition is as follows:

[0197] S(i,j) = -Px i -x j P 2 ;

[0198] In the formula, S(i,j) represents the similarity matrix between data points i and j, x i and x j respectively represent the feature vectors of data points i and j, and Px i -x j P 2 represents the square of their Euclidean distance.

[0199] Message passing includes two key steps: responsibility value calculation, which reflects the suitability (relative competitiveness) of data point k to become the cluster center of point i; and support degree calculation, which represents the support strength of other points for point i to select k as the cluster center. Calculating the responsibility value r(i,k) and support degree value a(i,k) of each solution includes:

[0200] The calculation formula for the responsibility value is:

[0201]

[0202] The calculation formula for the support degree value is:

[0203]

[0204] In the formula, r(i,k) represents the responsibility value of solution i to reference vector k, a(i,k) represents the support degree value of solution i to reference vector k; s(i,k) represents the similarity matrix of solution i to reference vector k, a(i,k′) represents the support degree value of solution i to reference vector k′; s(i,k′) represents the similarity of solution i to reference vector k′, and r(i′,k) represents the responsibility value of solution i′ to reference vector k.

[0205] As Figure 7 shown in the figure, after multiple iterations, the algorithm will adaptively select the cluster centers and assign each point to the nearest cluster. These cluster centers represent the local characteristics of the solution distribution. Subsequently, by calculating the cosine similarity between each cluster center and the reference vector, the cluster centers are assigned to the closest reference vector, thereby realizing the association between the solutions and the reference vectors. The cosine similarity calculation formula is as follows:

[0206]

[0207] In the formula, represents the dot product of vectors x and y, n represents the dimension of the vector space, that is, the number of elements of vectors x and y, x i and x j respectively represent the feature vectors of data points i and j, Px P and Py P respectively represent the magnitudes of the vectors, where:

[0208]

[0209] On this basis, according to the allocation relationship of each reference vector, a local mating pool is generated for each reference vector, and a global mating pool is generated at the same time to ensure that there are enough individuals in each cluster for offspring generation. If the number of individuals in some clusters is insufficient, the local mating pool size will be maintained by randomly supplementing individuals from the global mating pool.

[0210] Next, in each generation, 4 parent individuals are randomly selected from the mating pool, and the offspring are generated using Algorithm 4, the adaptive differential evolution (DE) algorithm. The generation of offspring is affected by the dynamically adjusted differential scaling factor sf and the crossover probability cr, so as to enhance the search ability of the algorithm. The generated offspring obtain fitness values through objective value calculation and constraint handling to evaluate their quality.

[0211] Among them, the objective value is calculated through the objective function. If there is a constraint function in the problem, constraint handling is required, that is, calculating the constraint penalty score. If there is a constraint function g j (x) and h k (x) in the problem, the calculation formula is as follows:

[0212] For the inequality constraint g j (x) ≤ 0, the expression of the penalty score P j of the inequality constraint is:

[0213] P j = max(0, g j (x));

[0214] In the formula, if g j (x) ≤ 0, it means the constraint is satisfied and the penalty score is 0; if g j (x) > 0, it means the constraint is violated and the penalty score is g j (x).

[0215] For the equality constraint h k (x) = 0, the expression of the penalty score P k of the equality constraint is:

[0216] P k = |h k (x)|;

[0217] The penalty score for the equality constraint is the absolute value of the deviation, and the total penalty score is the sum of the penalty scores for all constraints.

[0218] Finally, the memory parameters sf and cr of differential evolution are updated according to the performance of the new generation of individuals, enabling the parameters to dynamically adapt to the search requirements, thereby balancing global exploration and local exploitation.

[0219] The adaptive differential evolution algorithm generates offspring: First, the decision variables of four parents are extracted, and the boundaries of the decision variables are initialized according to the given lower and upper bounds. Then, the crossover operation of differential evolution begins. By randomly selecting positions and controlling the probability of crossover with the crossover rate, offspring are generated based on the differences between parents at the selected positions. Next, polynomial mutation is performed. The mutation operation controls which decision variables mutate with the mutation probability proM and controls the magnitude of mutation according to the distribution index. The mutation operation is divided into two cases: when the mutation intensity mu is less than or equal to 0.5, the offspring variables approach the lower bound; when mu is greater than 0.5, the offspring variables approach the upper bound. Finally, the algorithm ensures that the generated offspring variables are within the specified upper and lower boundaries and returns the generated offspring solutions.

[0220] As described in Algorithm 4, the adaptive differential evolution algorithm includes the following steps:

[0221] In each generation, memory_pos, memory_sf, and memory_cr are randomly initialized according to the population size pop_size, where memory_sf and memory_cr represent the memories of the scaling factor and crossover rate respectively.

[0222] The crossover rate is generated through a normal distribution, and it is ensured that the value of the crossover rate is restricted between 0 and 1. If the crossover rate is -1, it is set to 0.

[0223] The scaling factor sf is generated by offsetting the mean value and a random value, and it is ensured that all generated scaling factors are greater than 0 and do not exceed 1.

[0224] Calculate the fitness difference dif between the current population and the offspring population. Determine the superiority and inferiority of the parents and offspring based on the fitness difference, and mark the successful crossover rates and scaling factors.

[0225] The successful crossover rates and scaling factors are updated with weights, and the corresponding memories in memory_sf and memory_cr are updated. The update rule is based on the fitness difference, and successful parents have a greater impact on the memory.

[0226] After each update, memory_pos is updated cyclically to ensure that the memory always remains within the fixed size memory_size.

[0227] Algorithm 4: Adaptive Differential Evolution Algorithm:

[0228] Input: Parent1, Parent2, Parent3, Parent4: Selected parent individuals; Lower: Lower bound of decision variables; Upper: Upper bound of decision variables; pos: Memory location; sf: Memory parameter of scaling factor; cr: Memory parameter of crossover probability; proM: Mutation probability; D: Number of dimensions of decision variables; N: Population size.

[0229] Output: Offspring: Generated offspring individuals.

[0230] 1. Set Parent1, Parent2, Parent3, Parent4 as the decision variables of the parent individuals;

[0231] 2. Initialize Offspring = Parent1;

[0232] 3. For each position pos, randomly generate a number rand(N, D);

[0233] If, perform the following steps:

[0234] Randomly generate a position for Site

[0235] Adjust Offspring(Site) according to the following formula:

[0236] Offspring(Site) = Offspring(Site) + sf(pos) × (Parent2(Site) - Parent1(Site)) +;

[0237] sf(pos) × (Parent3(Site) - Parent4(Site))

[0238] 4. Generate a random number mu = rand(N, D) again;

[0239] 5. If mu ≤ 0.5:

[0240] Set temp = Site & mu;

[0241] Adjust Offspring: Offspring = Offspring(temp) + Mutation adjustment towards the lower bound;

[0242] 6. Otherwise:

[0243] Set temp = Site & mu;

[0244] Adjust Offspring: Offspring = Offspring(temp) + Mutation adjustment towards the upper bound;

[0245] 7. Finally, limit the value of Offspring between the upper and lower bounds:

[0246] Offspring = max(min(Offspring, Upper), Lower);

[0247] Specifically, this adaptive differential evolution algorithm adjusts the differential evolution parameters (crossover rate CR and scaling factor F) through the adaptive function. First, a certain number of historical parameters (memory_sf and memory_cr) are randomly selected, and the new crossover rate cr and scaling factor sf are generated through Gaussian distribution and tangent transformation. The values of the new crossover rate and the newly generated scaling factor should be within [0, 1]. Calculate the fitness differences between the parent and offspring, and determine which parents are fitter than the offspring and which offspring are better. For the fitter parents, record the corresponding crossover rate and scaling factor. On this basis, update the values of the crossover rate and scaling factor in memory by weighted summation.

[0248] Specifically, when updating the crossover rate, if its value is 0 or the initial state is -1, maintain it as -1, otherwise update it weighted according to the fitness difference. When updating the scaling factor, update it by weighted summation based on the fitness difference and the square of the scaling factor. Finally, update the memory pointer memory_pos to ensure that it loops within a certain range. In short, this function adjusts the crossover rate and scaling factor through feedback on fitness changes, gradually optimizing the search strategy of the differential evolution algorithm and enhancing the adaptive ability of the algorithm.

[0249] Every time the set number of function evaluations FE is reached, store the last num cluster centers, and calculate the cosine similarity between each cluster center and the reference vector. For each reference vector, use the mean of the corresponding cluster centers as the new reference vector, thereby achieving the update of the reference vector. Combine the current archive with the newly generated offspring individuals, perform association assignment again, and update the ideal point according to the objective value to ensure that the individuals in the archive have good distribution and diversity. When the archive update is completed and meets the preset conditions, the function returns the updated diverse archive.

[0250] S4. Perform optimization updates on the preliminary updated convergence archive and diversity archive respectively in a loop until the termination condition is reached and then stop, and output the optimal solution set.

[0251] Among them, performing optimization updates on the preliminary updated convergence archive and diversity archive respectively in a loop until the termination condition is reached and then stop, and outputting the optimal solution set includes:

[0252] Select the convergence archive parent and the diversity parent from the initially updated convergence archive and diversity archive.

[0253] Among them, selecting the convergence archive parent and the diversity parent from the initially updated convergence archive and diversity archive includes:

[0254] Randomly select half of the parent solutions from the convergence archive CA.

[0255] Then select the other half of the parent solutions from CA.

[0256] Compare the quality of the parent solutions according to the objective value, and select the solutions with a stronger dominance relationship.

[0257] Combine the excellent parents selected from the convergence archive with the solutions randomly selected from the diversity archive DA to form the final set of parent solutions.

[0258] Randomly select the remaining number of parent solutions from the diversity archive DA to enhance the diversity of the solution set.

[0259] Use the set of genetic parameters in the genetic algorithm to generate the convergence offspring and the diversity offspring from the convergence archive parent and the diversity archive parent;

[0260] Use the convergence offspring to update the convergence archive, and use the diversity offspring to update the diversity archive. Iteratively perform the optimization update of the convergence archive and the diversity archive until the termination condition is reached and then stop, and output the optimal solution set.

[0261] To facilitate the understanding of the above technical solution of the present invention, the following will perform a detailed description of the optimization update on the initially updated convergence archive and diversity archive respectively in the actual process of the present invention until the termination condition is reached and then stop, and output the optimal solution set:

[0262] As Figure 2As shown in the figure, the process of selecting parents from the convergence archive and the diversity archive: Appropriate parent individuals are selected from the convergence archive (CA) and the diversity archive (DA), and these individuals prepare for generating offspring individuals in subsequent genetic operations. ParentA: This set of parents consists of two parts. One part is the parents selected from CA through the dominance relationship, and the other part is the parents randomly selected from DA. First, two groups of parents, CAParent1 and CAParent2, each with a size of N / 2, are randomly drawn from CA. Compare the objective values of these two groups of parents, and select the better part of the parents using the dominance relationship. Then, randomly select N / 2 parents from DA and combine them to form ParentA with a population size of N. Randomly select N parents from CA to form ParentB. Among them, ParentA is used to generate offspring individuals with convergence, and ParentB is used to generate offspring individuals with diversity. These two groups of parents, one for generating convergent offspring and the other for generating diverse offspring, and then the two groups of offspring are merged into a new offspring population. ParentA only performs crossover operations, and ParentB only performs mutation operations.

[0263] The following combines specific embodiments to specifically describe the above-mentioned decomposition-based double-archive-guided multi-objective neural architecture search method.

[0264] The benchmark test suite used in this embodiment is based on the EvoXBench platform (a benchmark platform for evaluating and comparing the performance of evolutionary computing algorithms), which is used to generate a neural network architecture search benchmark test suite for EMO algorithms. Eight state-of-the-art multi-objective algorithms are selected for comparison with DANAS: RVEA, HEA, CMOSMA, NSGAIII, TS-NSGA-II, RPV-NSGA-II, MOEAD, and LOMONAS. The test suites are C10 / MOP (cifar10 dataset image classification test suite), IN1K / MOP (imagenet dataset image classification test suite), and CitySeg / MOP (cityspace dataset image segmentation test suite). Since it is difficult to obtain the true frontiers of some test problems, the main evaluation metric of the test suite generated by the EvoXBench platform is the HV value. By using the hypervolume metric HV as the performance evaluation metric to comprehensively judge the diversity and convergence of the population, and the performance metrics are used to describe DANAS and the comparison algorithms. Finally, the experimental results of the algorithms are analyzed.

[0265] Specifically, HV mainly evaluates the convergence performance of the population. The larger the HV value, the closer the non-dominated solutions in the objective space are to the true frontiers, and the better the convergence of the algorithm. The calculation formula of HV is as follows:

[0266]

[0267] In the formula, v(x, p) represents the hypervolume of the space formed between the solution x in the non-dominated solution set X and the reference point P corresponding to the true Pareto front.

[0268] Among the 8 comparison algorithms, RVEA and HEA are reference vector-based multi-objective optimization algorithms, NSGAIII, TS-NSGA-II, and RPV-NSGA-II are all based on the Pareto dominance relationship, CMOSMA and MOEAD are both based on decomposition. The CMOSMA algorithm is a multi-objective optimization algorithm focusing on neural network architecture search, which provides an efficient, stable, and high-quality solution neural network architecture search algorithm through the combination of clustering and multi-objective optimization, while LOMONAS pays more attention to the Pareto local search ability.

[0269] To evaluate the performance of these algorithms, all algorithms are executed 31 independent runs, and the maximum number of evaluations is 10,000. All algorithms are tested on matlab through the platemo platform, and this computer is equipped with an Intel(R) Core(TM) i7-13700K CPU@3.40GHz and an NVIDIA GeForce RTX 4080 SUPER GPU. The Wilcoxon rank-sum test is performed on all comparison algorithms at a 5% significance level. The symbols "+", "-", and "=" indicate that the performance of the compared algorithm is better than, worse than, and similar to DANAS, respectively. In the last row of the result data table, the sum results of the Wilcoxon rank-sum test for all test instances are calculated.

[0270] The population size N of all comparison algorithms is set to 100. In the diversity archive update strategy, when the adaptive differential evolution algorithm generates offspring, the storage parameters are memory_size = 5, scaling factor sf = 0.5, crossover probability cr = 0.9, and memory_pos = 1. During cyclic optimization, for the two parents obtained from the convergence archive, the first parent undergoes crossover and the second parent undergoes mutation. The parameters for crossover and mutation are set as P c = 1, n c = 20 (P c is the crossover probability, n c is the crossover parameter); P m = 1, n m = 20 (P m is the mutation probability, n m is the mutation parameter).

[0271] Specifically, the statistical results of the HV values and standard deviations (in parentheses after the mean) obtained by nine algorithms on three test suites are summarized in Table 1-3. In addition, a Wilcoxon rank-sum test was conducted on the experimental results at a 95% confidence level. ≈, +, and - indicate that the algorithm is statistically similar to, superior to, or inferior to the comparison algorithm, respectively. The bold values in the table are the optimal values of each test function.

[0272] Table 1: Performance Comparison of Different Algorithms on the C10MOP Problem

[0273]

[0274]

[0275] Table 2: Performance Comparison of Different Algorithms on the IN1KMOP Problem

[0276]

[0277]

[0278] Table 3: Performance Comparison of Different Algorithms on the CitySegMOP Problem

[0279]

[0280]

[0281] It can be seen that in Table 1, on the conventional test instances with 2-8 objectives, DANAS shows the best comprehensive performance among the eight comparison algorithms. Among the nine test instances, DANAS shows the best HV value in five instances, while RVEA (Preference Vector Guided Multi-Objective Optimization Algorithm), HEA (Hyperdominance-based Multi-Objective Optimization), CMOSMA (Self-Organizing Map-based Multi-Objective Optimization Algorithm), NSGAIII (Reference Point-based Non-Dominated Sorting Multi-Objective Optimization Algorithm), TS-NSGA-II (A Two-Stage Evolutionary Algorithm Balancing Convergence and Diversity), RPV-NSGA-II (Reference Vector-based Pareto Dominance Relationship Multi-Objective Optimization Algorithm), MOEAD (A Decomposition-based Multi-Objective Evolutionary Algorithm), and LOMONAS (Lightweight Multi-Objective Evolutionary Neural Architecture Search Algorithm with Low-Cost Surrogate Metrics) show the best performance in 0, 0, 3, 0, 0, 0, 0, and 1 instance(s), respectively. Among them, the HV average convergence curves of six algorithms on the C10MOP test problem are as Figures 8 - 10As shown in the figure, where "Number of FES" represents "Function Evaluations", that is, the number of function evaluations; "Hypervolume (HV)", namely the hypervolume. Table 2 shows the statistical results of the HV values and standard deviations (in parentheses after the mean) obtained by 9 algorithms on the IN-1K / MOP test suite. Table 3 shows the statistical results of the HV values and standard deviations (in parentheses after the mean) obtained by 9 algorithms on the CitySegMOP test suite.

[0282] Taking the results obtained on the C-10 / MOP test suite as an example, the HV convergence curve graph is drawn. Among them, the convergence effects of the TS-NSGA-II and RPV-NSGA-II algorithms are too poor, so they are not shown in the figure. And the LOMONAS algorithm does not open source its matlab code, so it is not compared in the figure either. Subtract the average HV value of the 8 comparison algorithms from the average HV value of DANAS, multiply the difference by 100,000, and then take the logarithm, as Figures 10 - 13 shown in the figure, where "HV Difference" represents the hypervolume difference. It can be seen from the figure that the HV value of DANAS is higher than that of most comparison algorithms, and its search ability is the strongest.

[0283] In a relatively large search space, DANAS has a good search effect, which is related to the reference vector allocation, local matching pool and global matching pool in the algorithm. By dynamically allocating reference vectors, clustering according to similarity, and generating offspring through local and global mating pools, most areas of the search space are explored. This leads the algorithm to explore more different domains and find better architectures.

[0284] In summary, by means of the above technical solutions of the present invention, the present invention provides a multi-objective neural architecture search method based on dual archive guidance for the small model trap and early convergence problems in neural architecture search, thereby effectively avoiding the premature convergence phenomenon caused by small models in traditional genetic algorithms, ensuring the balance between convergence and diversity during the search process, and further exploring potential optimal architectures more comprehensively; when facing multi-objective optimization problems, through the multi-objective neural architecture search method based on dual archive guidance provided by the present invention, the target problem is transformed into multiple simple sub-problems for solution, thereby effectively avoiding the dominance resistance phenomenon and solution space aggregation problem of traditional MOEA algorithms in high-dimensional sparse target spaces, improving the uniformity and diversity of solutions, and further enhancing the overall optimization effect; the present invention adopts an adaptive differential evolution algorithm to directly operate on the discrete architecture encoding, effectively solving the low efficiency problem of traditional gradient-based optimization methods in discrete decision spaces, and avoiding local optimal traps by comprehensively exploring the discrete search space, improving the search efficiency of the algorithm and the global optimization ability of the results; through the optimization strategy guided by dual archives, the present invention can improve the balance of neural network architectures in multiple optimization objectives (such as performance, number of parameters, inference latency, etc.), solve the deficiencies of existing multi-objective optimization algorithms in high-dimensional discrete search spaces, and is widely applied to fields such as generative artificial intelligence, automated design and optimization of deep learning models, etc. Especially in the case of limited computing resources, the efficiency and effect of neural network architecture search are improved.

[0285] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A decomposition-based dual-archive guided multi-objective neural architecture search method, characterized in that The method includes: Initialize the population and generate a random population; According to the random population and the preset archive size, the number of solutions in the convergence archive is compared with the preset archive size, and based on the comparison result, the convergence archive is preliminarily updated; Based on the random population and the preset problem parameters, the diversity archive is initially updated using the clustering algorithm and the adaptive differential evolution algorithm; The convergence archive and diversity archive after preliminary update are optimized and updated cyclically respectively until the termination condition is reached, and the optimal solution set is output.

2. The decomposition-based dual-archive guided multi-objective neural architecture search method according to claim 1, characterized in that: The step of comparing the number of solutions in the convergence archive with the preset archive size according to the random population and the preset archive size, and preliminarily updating the convergence archive based on the comparison result includes: Merge the convergence archived solution and the newly generated solution to form the first candidate solution set; Perform non-dominated sorting on the first candidate solution set and retain the solutions on the non-dominated front to ensure that the convergence archive contains only non-dominated solutions; The number of solutions in the convergence archive is compared with the preset maximum archive. If the number of solutions in the convergence archive is less than or equal to the preset maximum archive, the convergence archive after preliminary update is output; otherwise, the non-dominated solutions in the convergence archive are screened by the ISDE+ indicator to control the archive size.

3. The decomposition-based dual-archive guided multi-objective neural architecture search method according to claim 2, characterized in that: The screening of non-dominated solutions in the convergence archive by the ISDE+ indicator includes: Normalize the target values ​​in the convergence archive; Calculate the objective value difference between each pair of solutions in the convergence archive, and calculate the ISDE+ indicator value of each solution based on the objective value difference; According to the ISDE+ index value, the solutions with higher ISDE+ values ​​in the convergence archive are gradually screened, and the final retained solution set is used as the convergence archive after preliminary update.

4. The decomposition-based dual-archive guided multi-objective neural architecture search method according to claim 3, characterized in that: The method of preliminarily updating the diversity archive based on the random population and the preset problem parameters using the clustering algorithm and the adaptive differential evolution algorithm includes: The solutions in the diversity archive are merged with the newly generated solutions to form a second candidate solution set, and the second candidate solution set is non-dominated sorted to ensure that the diversity archive contains only non-dominated solutions; Calculate the ideal points in the diversity archive and dynamically assign individuals in the diversity archive to the corresponding subproblems; Standardize the target values ​​in the convergence archive, cluster the target values ​​in the diversity archive using a clustering algorithm to obtain clustering results, and generate a distribution relationship of reference vectors based on the clustering results; Generate a local mating pool and a global mating pool based on the allocation relationship of the reference vector, and generate a parent generation through the local mating pool and the global mating pool; The adaptive differential evolution algorithm is used to generate offspring, and the fitness of the offspring is calculated. The adaptive differential evolution algorithm parameters are updated according to the fitness of the offspring to obtain the diversity archive after the initial update.

5. The decomposition-based dual-archive guided multi-objective neural architecture search method according to claim 4, characterized in that: The clustering algorithm is used to cluster the target values ​​in the diversity archive to obtain the clustering results including: Calculate the similarity matrix between solutions in the diversity archive and initialize the clustering algorithm parameters based on the similarity matrix; Iteratively calculate and update the responsibility value and support value of each solution to the reference vector until the clustering algorithm converges, and select the cluster center based on the maximum responsibility value and support value; The cosine similarity between the cluster center and the reference vector is calculated, and the solution is assigned to the nearest reference vector to achieve the association between the solution and the reference vector.

6. The decomposition-based dual-archive guided multi-objective neural architecture search method according to claim 5, characterized in that: The method of generating offspring using the adaptive differential evolution algorithm, calculating the fitness of the offspring, and updating the adaptive differential evolution algorithm parameters according to the fitness of the offspring includes: Extract the decision variables of the parent generation and initialize the boundaries of the decision variables according to the preset upper and lower boundaries; Perform the crossover operation of differential evolution, by randomly selecting positions and controlling the probability of crossover with the crossover rate, and generating offspring at the selected positions according to the differences between the parents; Perform polynomial mutation, control the decision variables that need to be mutated by the mutation probability, and control the magnitude of the mutation according to the distribution index to ensure that the generated offspring variables are within the specified upper and lower bounds; Calculate the fitness of the offspring and update the adaptive differential evolution algorithm parameters according to the fitness of the offspring.

7. The decomposition-based dual-archive guided multi-objective neural architecture search method according to claim 6, characterized in that: The updating of the adaptive differential evolution algorithm parameters includes: Randomly select historical crossover rates and historical scaling factors, and perform Gaussian distribution and tangent transformation processing in turn to generate new crossover rates and scaling factors; Calculate the fitness difference between the parent and offspring, record the new crossover rate and scaling factor corresponding to the parent whose fitness is greater than that of the offspring, and update the crossover rate and scaling factor values ​​in the adaptive differential evolution algorithm using a weighted summation method.

8. The decomposition-based dual-archive guided multi-objective neural architecture search method according to claim 7, characterized in that: The optimization update is performed cyclically on the convergence archive and the diversity archive after the preliminary update respectively until the termination condition is reached, and the output optimal solution set includes: Select a convergence archive parent and a diversity parent from the convergence archive and the diversity archive after preliminary update; Genetic parameter sets in genetic algorithms are used to convert convergence archived parents and diversity parents into convergence offspring and diversity offspring; The convergence archive is updated using the convergence offspring, and the diversity archive is updated iteratively using the diversity offspring. The optimization update of the convergence archive and the diversity archive is iteratively performed until the termination condition is reached, and the optimal solution set is output.

9. The decomposition-based dual-archive guided multi-objective neural architecture search method according to claim 8, characterized in that: The calculation formula of the ISDE+ index value is: Where ISDE represents the ISDE+index value of the solution, I(i,:) represents the row of the difference matrix between the i-th solution and all other solutions, C(i) represents the maximum difference between the i-th solution and other solutions, and N represents the size of the solution set.

10. The decomposition-based dual-archive guided multi-objective neural architecture search method according to claim 9, characterized in that: The calculation formula of the responsibility value is: The calculation formula of the support value is: 式中, r(i,k) 表示解i对参考向量k的责任值, a(i,k) 表示解i对参考向量k的支持度值, s(i,k) 表示解i对参考向量k的相似度矩阵, a(i,k′) 表示解i对参考向量 k′ 的支持度值, s(i,k′) 表示解i对参考向量 k′ 的相似度, r(i′,k) 表示解 i′ 对参考向量k的责任值。