Neural network architecture search optimization method and device, equipment and storage medium

By modeling the search problem of neural network architecture as a multi-objective optimization problem, and adopting Chebishev polynomial decomposition and adaptive selection sub-problem strategies, the problem of slow convergence of neural network architecture search process is solved, and a more efficient search process of neural network architecture is achieved.

CN120069012AInactive Publication Date: 2025-05-30SHENZHEN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510555733.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the search process of neural network architectures converges slowly, making it difficult to find high-performance network architectures.

Method used

By randomly initializing the neural network architecture, multi-task populations are generated, neural network architecture search problems are modeled into multi-objective optimization problems, Chebishev polynomials are used to decompose the multi-objective optimization problems into multiple sub-problems, and based on the adaptive selection of sub-problems strategy of sub-problems improvement rate, select from multiple sub-problems, and each selected sub-problem is performed to evolve, produce offspring individuals, and update the multi-task population until it is satisfied.

Benefits of technology

Through the integration of adaptive selection sub-problem strategies and adaptive migration evolution strategies, the waste of computing resources is reduced, the convergence speed of populations is accelerated, and the performance of neural network architecture search is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069012A_ABST
    Figure CN120069012A_ABST
Patent Text Reader

Abstract

The invention is applicable to the technical field of neural network architecture search, and provides a neural network architecture search optimization method, device and equipment and a storage medium, and the method comprises the steps: randomly initializing a neural network architecture, generating a multi-task population, modeling a neural network architecture search problem into a multi-target optimization problem, and optimizing the multi-target optimization problem; the method comprises the following steps: decomposing a multi-objective optimization problem into a plurality of sub-problems by adopting a Chebyshev polynomial, selecting from the plurality of sub-problems based on a self-adaptive sub-problem selection strategy of a sub-problem improvement rate, performing evolution operation on each selected sub-problem, generating a filial generation individual, updating a multi-task population according to the filial generation individual, and performing multi-time iteration to obtain a multi-objective optimization result. And the optimal architecture of the neural network is output, so that the waste of computing resources is reduced, and the convergence speed of the population is accelerated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of neural network architecture search, and particularly relates to a method, device, equipment and storage medium for optimizing neural network architecture search. Background Art

[0002] Neural network architecture search aims to automatically design high-performance neural network structures to replace traditional manual design. It constructs a search space (including the topological structure of the neural network and candidate network operations, etc.), and uses a certain search strategy to automatically explore and evaluate different network structures, so as to find the optimal or near-optimal architecture. This method can alleviate the cumbersome process of manually designing networks to a certain extent and is expected to improve the model performance and adaptability. At present, some multi-task optimization algorithms have been introduced into neural network architecture search. The idea is to simultaneously consider the requirements of multiple classification tasks (such as CIFAR dataset, ImageNet dataset, etc.), and through sharing search resources or knowledge transfer, it is expected to find a general architecture suitable for all tasks or perform adaptive optimization for each task in a multi-task scenario. However, multi-task optimization algorithms need to simultaneously search for neural network architectures suitable for multiple tasks. The search space not only increases in the dimension of the network structure, but also needs to consider the compatibility between different tasks. This significantly increases the search difficulty and computational complexity, resulting in slow convergence or only finding sub-optimal network architectures during the neural network architecture search process. Summary of the Invention

[0003] The purpose of the present invention is to provide a method, device, equipment and storage medium for optimizing neural network architecture search, aiming to solve the problem of slow convergence during the neural network architecture search process in the prior art.

[0004] On the one hand, the present invention provides a method for optimizing neural network architecture search, and the method includes the following steps: Randomly initialize the neural network architecture to generate a multi-task population; Model the neural network architecture search problem as a multi-objective optimization problem; Use Chebyshev polynomials to decompose the multi-objective optimization problem into multiple sub-problems; Select from multiple sub-problems based on an adaptive sub-problem selection strategy based on the sub-problem improvement rate; Perform an evolutionary operation on each selected sub-problem to generate offspring individuals; Update the multi-task population according to the offspring individuals; Judge whether the iteration termination condition is satisfied. If not, jump to the step of selecting from multiple sub-problems based on the adaptive sub-problem selection strategy based on the sub-problem improvement rate. If satisfied, output the optimal architecture of the neural network.

[0005] Preferably, each individual in the multi-task population represents a neural network architecture.

[0006] Preferably, the multi-objective optimization problem includes a two-objective minimization problem of the error rate of the neural network on the dataset and the complexity of the neural network architecture.

[0007] Preferably, the step of selecting sub-problems based on the adaptive selection sub-problem strategy of the sub-problem improvement rate from multiple sub-problems includes: Calculate the improvement rate of each individual on each sub-problem and sort them in descending order; Select the first N sub-problems corresponding to the improvement rate, N where is a natural number greater than 0.

[0008] Preferably, the calculation formula of the improvement rate is as follows: where, x j i,old represents the i th parent individual of the j th task, x j i,new represents the offspring individual generated by the i th parent of the j th task, λ j i represents the i th sub-problem of the j th task, z i represents the preset ideal point of the i th task, FIR i j represents the improvement rate of the performance of the i th sub-problem of the j th task, g tch ( ) represents the Chebyshev value of the multi-task population individual, x represents the decision variable of an individual in the multi-task population, f j ( x ) represents the j th objective value of the individual.

[0009] Preferably, the evolutionary strategies used when performing evolutionary operations on each selected sub-problem include a migration evolutionary strategy and an intra-task evolutionary strategy, and the evolutionary strategies of each selected sub-problem are adaptively adjusted using a preset migration intensity.

[0010] Preferably, the evolutionary strategy for each selected sub - problem uses steps of adaptively adjusting with a preset migration intensity, including: When the randomly generated probability is less than the migration intensity, it is adjusted to the migration evolutionary strategy; When the randomly generated probability is not less than the migration intensity, it is adjusted to the in - task evolutionary strategy.

[0011] Preferably, the migration evolutionary strategy is an adaptive migration evolutionary strategy based on a sliding window and operator value evaluation, and the adaptive migration evolutionary strategy includes: Initialize the operator pool and the sliding window, where the operator pool includes at least two of a uniform crossover operator, a single - point crossover operator, a geometric crossover operator, a simulated binary crossover operator, and a differential evolution operator; Record the operator through the sliding window and the improvement rate of generating offspring using this operator; Conduct operator value evaluation according to the improvement rate and usage frequency corresponding to each operator, and dynamically select the optimal operator according to the evaluation results.

[0012] Preferably, the calculation formula for operator value evaluation is as follows: Where, op best represents the current optimal operator, l represents the sliding window size, Q ( op i ) represents the improvement rate of the operator in the i th sliding window, fre represents the corresponding usage frequency, N ( op j ) represents the usage times of the operator in the j th sliding window.

[0013] Preferably, the steps of updating the multi - task population according to the offspring individuals include: Compare the performance of the offspring individuals and the parent individuals of the corresponding sub - problems; If the performance of the offspring individuals is better than that of the parent individuals of the corresponding sub - problems, use the offspring individuals to replace the parent individuals of the corresponding sub - problems.

[0014] Preferably, if the Chebyshev value of the offspring individual is smaller than that of the parent individual of the corresponding sub - problem, it is determined that the performance of the offspring individual is better than that of the parent individual of the corresponding sub - problem.

[0015] On the other hand, the present invention provides a neural network architecture search and optimization device, and the device includes: An initialization unit for randomly initializing a neural network architecture to generate a multi-task population; A modeling unit for modeling the neural network architecture search problem as a multi-objective optimization problem; A sub-problem decomposition unit for decomposing the multi-objective optimization problem into multiple sub-problems by using Chebyshev polynomials; A sub-problem selection unit for selecting from multiple sub-problems based on an adaptive sub-problem selection strategy of sub-problem improvement rate; An evolution unit for performing evolution operations on each selected sub-problem; An update unit for updating the multi-task population according to the offspring individuals; and A judgment and output unit for judging whether the iteration termination condition is satisfied. If not, it triggers the sub-problem selection unit to execute the step of selecting from multiple sub-problems based on the adaptive sub-problem selection strategy of sub-problem improvement rate. If satisfied, it outputs the optimal architecture of the neural network.

[0016] Preferably, the sub-problem selection unit further includes: An improvement rate calculation unit for calculating the improvement rate of each individual on each sub-problem and sorting them in descending order; and A selection sub-unit for selecting the sub-problems corresponding to the top N improvement rates, N where is a natural number greater than 0.

[0017] Preferably, the formula for calculating the improvement rate is as follows: where, x j i,old represents the i th task and the j th parent individual, x j i,new represents the offspring individual generated by the i th task and the j th parent, and λ j i represents the i th task and the j th sub-problem, z i represents the i th task's preset ideal point, FIR i j represents the i th task and the j th sub-problem performance improvement rate, gtch () represents the Chebyshev value of an individual in the multi-task population, x represents the decision variable of an individual in the multi-task population, f j ( x ) represents the j th objective value of the individual.

[0018] On the other hand, the present invention also provides a neural network architecture search optimization device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned method are implemented.

[0019] On the other hand, the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method are implemented.

[0020] The present invention randomly initializes a neural network architecture, generates a multi-task population, models the neural network architecture search problem as a multi-objective optimization problem, uses Chebyshev polynomials to decompose the multi-objective optimization problem into multiple sub-problems, selects from multiple sub-problems based on an adaptive sub-problem selection strategy of the sub-problem promotion rate, performs evolutionary operations on each selected sub-problem to generate offspring individuals, updates the multi-task population according to the offspring individuals, and through multiple iterations, outputs the optimal architecture of the neural network. Thus, the differences between sub-problems are quantified through the adaptive sub-problem selection strategy, computing resources are allocated to more potential sub-problems, waste of computing resources is reduced, and the convergence speed of the population is accelerated; the present invention also improves the migration effect by fusing the adaptive sub-problem selection strategy and the adaptive migration evolution strategy, and further improves the performance of neural network architecture search. Description of the Drawings

[0021] Figure 1A is a flowchart of the implementation of the neural network architecture search optimization method provided in Embodiment 1 of the present invention; Figure 1B is a flowchart example diagram of the neural network architecture search optimization method provided in Embodiment 1 of the present invention; Figure 2 is a schematic structural diagram of the neural network architecture search optimization device provided in Embodiment 2 of the present invention; Figure 3 is a schematic structural diagram of the electronic device provided in Embodiment 3 of the present invention. Detailed Embodiments

[0022] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0023] The following describes the specific implementation of the present invention in detail with reference to specific embodiments: Example 1: Figure 1A and Figure 1B respectively show the implementation processes of the neural network architecture search optimization method provided in the first embodiment of the present invention. For the sake of convenience of description, only the parts related to the embodiments of the present invention are shown and are described in detail as follows: In step S101, a neural network architecture is randomly initialized to generate a multi-task population.

[0024] The embodiments of the present invention are applicable to the field of neural network search. In the embodiments of the present invention, a neural network architecture, that is, a chromosome, is randomly initialized to generate multiple initial task populations. As an example, as Figure 1B shown, Figure 1B the generated populations include population 1, population 1... population K, where K is a natural number greater than 1. Among them, each individual in the multi-task population represents a neural network architecture, and the decision variables of the chromosome are used to represent the topological structure of the neural network and the candidate network operations.

[0025] In step S102, the neural network architecture search problem is modeled as a multi-objective optimization problem.

[0026] In the embodiments of the present invention, the neural network architecture search problem can be modeled as a multi-objective optimization problem based on multiple task requirements. Preferably, the multi-objective optimization problem includes a two-objective minimization problem of the error rate of the neural network on the data set and the complexity of the neural network architecture. The multi-objective optimization problem can be specifically expressed as follows: Among them, x represents all the decision variables of each neural network architecture, that is, the chromosome decision variables of each individual, f 1 ( x ) represents the error generated on the data set when x is constructed into a neural network, f 2 ( x ) represents the physical data directly extracted from x , such as the number of parameters, computational complexity, etc., and is used to measure the complexity of the neural network.

[0027] In step S103, the multi-objective optimization problem is decomposed into multiple sub-problems by using Chebyshev polynomials.

[0028] In the embodiment of the present invention, the multi-objective optimization problem is decomposed into multiple single-objective optimization sub-problems by using Chebyshev polynomials. As an example, as Figure 1B shown, Figure 1B the sub-problems 1, sub-problem 2... sub-problem j are obtained by decomposition. Decomposing the multi-objective optimization problem into multiple single-objective optimization sub-problems by using Chebyshev polynomials belongs to the prior art and will not be elaborated in this embodiment.

[0029] In step S104, selection is performed from multiple sub-problems based on the adaptive sub-problem selection strategy of the sub-problem improvement rate.

[0030] In the embodiment of the present invention, when selection is performed from multiple sub-problems based on the adaptive sub-problem selection strategy of the sub-problem improvement rate, preferably, the improvement rate of each individual on each sub-problem is calculated and sorted in descending order, and the first N sub-problems corresponding to the improvement rates are selected. N is a natural number greater than 0, thus quantifying the differences between sub-problems and facilitating the selection of several sub-problems with high potential for optimization. As an example, as Figure 1B shown, Figure 1B the selected sub-problems in it are sub-problem 1 and sub-problem j .

[0031] More preferably, the calculation formula of the improvement rate is as follows: where, x j i,old represents the i th task's j th parental individual, x j i,new represents the i th task's j th offspring individual generated by the parental individual, λ j i represents the i th task's j th sub-problem, z i represents the i th task's preset ideal point, FIR i j represents the i th task's jThe improvement rate of the sub - problem performance g tch () represents the Chebyshev value of an individual in the multi - task population x represents the decision variable of an individual in the multi - task population f j ( x ) represents the j th objective value of the individual

[0032] In specific implementation, the objective value of each individual may include the above - mentioned error rate and complexity. For the objective value of each individual, before step S103, the corresponding data set can be input into each neural network architecture according to the task requirements for training, and its performance can be evaluated on the data set. Then, according to the modeled multi - objective optimization problem, the objective value of each individual in all task populations is calculated, including the above - mentioned error rate and complexity. Taking image classification as an example, it is modeled as a two - objective minimization problem of network accuracy and complexity. This algorithm inputs the image classification data sets CIFAR - 10 and CIFAR - 100 as two different tasks. Subsequently, the algorithm trains and evaluates the classification performance of different network architectures in the super - network according to the input data sets and calculates the size of their parameters. Then, according to the modeled two - objective optimization problem, the objective value of each individual in all task populations is calculated

[0033] In step S105, an evolutionary operation is performed on each selected sub - problem to generate offspring individuals

[0034] In the embodiment of the present invention, an evolutionary operation is performed on each selected sub - problem, so as to allocate more computing resources to potential sub - problems, reduce the waste of computing resources, and accelerate the convergence speed of the population. Among them, the evolutionary strategies used when performing an evolutionary operation on each selected sub - problem include a migration evolutionary strategy and an intra - task evolutionary strategy, and the evolutionary strategy of each selected sub - problem is adaptively adjusted using a preset migration intensity

[0035] When the evolutionary strategy of each selected sub - problem is adaptively adjusted using a preset migration intensity, preferably, when the randomly generated probability R is less than the migration intensity T P , it is adjusted to the migration evolutionary strategy. When the randomly generated probability R is not less than the migration intensity, it is adjusted to the intra - task evolutionary strategy, so as to achieve the adaptive adjustment of the evolutionary strategy through the migration intensity

[0036] Preferably, the migration evolutionary strategy is an adaptive migration evolutionary strategy based on a sliding window and operator value evaluation, so as to improve the migration effect by adaptively selecting migration operators. As an example, such as Figure 1BAs shown in the figure, the operator 2 is adaptively selected from the operator 1, operator 2, operator 3... to perform migration evolution on the parent 1 and parent 2.

[0037] Further preferably, the adaptive migration evolution strategy includes: initializing the operator pool and the sliding window, recording the operator and the promotion rate of generating offspring using this operator through the sliding window, evaluating the operator value according to the promotion rate and usage frequency corresponding to each operator, and dynamically selecting the optimal operator according to the evaluation result. Thus, through adaptive migration evolution, the exploration and exploitation of the operator are balanced, and the migration effect is further improved.

[0038] Among them, the operator pool may include at least two of the uniform crossover operator, single-point crossover operator, geometric crossover operator, simulated binary crossover operator, and differential evolution operator. The following will explain each type of operator respectively.

[0039] Uniform crossover operator: Among them, r i represents a random number uniformly distributed between [0, 1], i represents the position of the gene, c i represents the i -th gene position of the offspring individual, p 1i represents the i -th gene position of the first parent individual, p 2i represents the i-th gene position of the second parent individual.

[0040] Single-point crossover operator: Among them, u represents a random number of 0 or 1, i represents the position of the gene, c j represents the j-th gene position of the offspring individual, p j 1 represents the j -th gene position of the first parent individual, p j 2 represents the j -th gene position of the second parent individual. c 1 represents the first offspring individual, c 2 represents the second offspring individual.

[0041] Geometric crossover operator: c j 1 Represents the j th gene position of the first offspring individual, c j 2 Represents the j th gene position of the second offspring individual, p j 1 Represents the j th gene position of the first parent individual, p j 2 Represents the j th gene position of the second parent individual.

[0042] Simulated binary crossover operator: Wherein, u Represents a random number between [0, 1], n Represents the distribution index, usually 2 <= n <= 5.

[0043] Differential evolution operator: Wherein, rand Represents a random number between [0, 1], that is, the randomly generated probability above, CR and F are two preset control parameters, sub represents the sub-problem number, x j n1 and x j n2 are two parent individuals.

[0044] Preferably, the calculation formula for operator value evaluation is as follows: Wherein, op best Represents the current optimal operator, l Represents the sliding window size, Q ( op i ) Represents the improvement rate of the operator in the i th sliding window, fre Represents the corresponding usage frequency, N ( op j ) Represents thej The number of times an operator of a sliding window is used N ( op i ) represents the number of times an operator of the i th sliding window is used.

[0045] Optionally, the in-task evolution strategy includes: randomly selecting two individuals from the parental population, and using the differential evolution operator on the two obtained parental individuals to generate offspring individuals to achieve in-task evolution.

[0046] In step S106, update the multi-task population according to the offspring individuals.

[0047] In the embodiment of the present invention, update the multi-task population according to the offspring individuals obtained in step S105. When updating the multi-task population according to the offspring individuals, preferably, compare the performance of the offspring individuals with that of the parental individuals corresponding to the sub-problems. If the performance of the offspring individuals is better than that of the parental individuals corresponding to the sub-problems, use the offspring individuals to replace the parental individuals corresponding to the sub-problems to generate a new generation of population, thereby realizing the update of the multi-task population. Further preferably, if the Chebyshev value of the offspring individuals is smaller than the Chebyshev value of the parental individuals corresponding to the sub-problems, it is determined that the performance of the offspring individuals is better than that of the parental individuals corresponding to the sub-problems, and the offspring individuals are used to replace the parental individuals corresponding to the sub-problems, so as to realize performance comparison through the comparison results of the Chebyshev values.

[0048] In step S107, determine whether the iteration termination condition is satisfied. If not, jump to step S104. If satisfied, execute step S108.

[0049] In the embodiment of the present invention, the iteration termination condition may be reaching a preset maximum number of iterations. Of course, the iteration termination condition may also be that the change in the objective function value is less than a preset threshold, or reaching the maximum running time, which is not specifically limited in this embodiment.

[0050] In step S108, output the optimal architecture of the neural network.

[0051] In the embodiment of the present invention, output multiple populations, and the obtained optimal architecture is a set of final data for representing the neural network architecture, that is, chromosome information. Further, construct a target neural network according to the optimal architecture, thereby completing the search for the neural network architecture.

[0052] In the embodiment of the present invention, the neural network architecture is randomly initialized to generate a multi-task population. The neural network architecture search problem is modeled as a multi-objective optimization problem. The Chebyshev polynomial is used to decompose the multi-objective optimization problem into multiple sub-problems. An adaptive sub-problem selection strategy based on the sub-problem improvement rate is used to select from multiple sub-problems. Evolution operations are performed on each selected sub-problem to generate offspring individuals. The multi-task population is updated according to the offspring individuals. Through multiple iterations, the optimal architecture of the neural network is output. Thus, the differences between sub-problems are quantified by the adaptive sub-problem selection strategy, and computing resources are allocated to more promising sub-problems, reducing the waste of computing resources and accelerating the convergence speed of the population. The present invention also improves the migration effect by integrating the adaptive sub-problem selection strategy and the adaptive migration evolution strategy, thereby improving the performance of neural network architecture search.

[0053] Example 2: Figure 2 FIG. shows the structure of the neural network architecture search optimization device provided in the second embodiment of the present invention. For ease of description, only the parts related to the embodiment of the present invention are shown, including: An initialization unit 21, configured to randomly initialize a neural network architecture and generate a multi-task population; A modeling unit 22, configured to model the neural network architecture search problem as a multi-objective optimization problem; A sub-problem decomposition unit 23, configured to decompose the multi-objective optimization problem into multiple sub-problems by using the Chebyshev polynomial; A sub-problem selection unit 24, configured to select from multiple sub-problems according to an adaptive sub-problem selection strategy based on the sub-problem improvement rate; An evolution unit 25, configured to perform evolution operations on each selected sub-problem; An update unit 26, configured to update the multi-task population according to the offspring individuals; and A determination and output unit 27, configured to determine whether the iteration termination condition is satisfied. If not, it triggers the sub-problem selection unit 24 to execute the step of selecting from multiple sub-problems according to the adaptive sub-problem selection strategy based on the sub-problem improvement rate. If so, it outputs the optimal architecture of the neural network.

[0054] Preferably, each individual in the multi-task population represents a neural network architecture, and the multi-objective optimization problem includes a two-objective minimization problem of the error rate of the neural network on the data set and the complexity of the neural network architecture.

[0055] Preferably, the sub-problem selection unit further includes: An improvement rate calculation unit, configured to calculate the improvement rate of each individual on each sub-problem and sort them in descending order; and A selection sub-unit, configured to select the topN sub - problems corresponding to an improvement rate, N where \(n\) is a natural number greater than 0; Preferably, the calculation formula of the improvement rate is as follows: where, x j i,old represents the \(i\) - th i parent individual of the \(j\) - th j task, x j i,new represents the offspring individual generated by the \(j\) - th i parent of the \(k\) - th j task, \(\lambda\) j i represents the \(m\) - th i sub - problem of the \(j\) - th j task, z i represents the preset ideal point of the \(j\) - th i task, FIR i j represents the improvement rate of the performance of the \(m\) - th i sub - problem of the \(j\) - th j task, g tch \((\cdot)\) represents the Chebyshev value of the multi - task population individual, x \(\mathbf{x}\) represents the decision variable of an individual in the multi - task population, f j \((\) x ) represents the \(l\) - th j objective value of the individual.

[0056] Preferably, the evolutionary strategies used for evolving each selected sub - problem include a migration evolutionary strategy and an intra - task evolutionary strategy, and the evolutionary strategies of each selected sub - problem are adaptively adjusted using a preset migration intensity.

[0057] More preferably, the evolutionary unit is further configured to: When the randomly generated probability is less than the migration intensity, adjust to the migration evolutionary strategy; When the randomly generated probability is not less than the migration intensity, adjust to the intra - task evolutionary strategy.

[0058] Preferably, the migration evolutionary strategy is an adaptive migration evolutionary strategy based on a sliding window and operator value evaluation.

[0059] More preferably, the evolutionary unit further includes: An initialization subunit, configured to initialize an operator pool and a sliding window, where the operator pool includes at least two of a uniform crossover operator, a single-point crossover operator, a geometric crossover operator, a simulated binary crossover operator, and a differential evolution operator; A recording unit, configured to record an operator through the sliding window and the improvement rate of generating offspring by using the operator; and An operator value evaluation unit, configured to evaluate the operator value according to the improvement rate and the usage frequency corresponding to each operator, and dynamically select an optimal operator according to the evaluation result; The calculation formula for operator value evaluation is as follows: Wherein, op best represents the current optimal operator, l represents the sliding window size, Q ( op i ) represents the improvement rate of the operator in the i th sliding window, fre represents the corresponding usage frequency, N ( op j ) represents the usage times of the operator in the j th sliding window.

[0060] Preferably, the update unit further includes: A performance comparison unit, configured to compare the performance of an offspring individual and a parent individual of a corresponding sub-problem; and A performance determination unit, configured to use the offspring individual to replace the parent individual of the corresponding sub-problem if the performance of the offspring individual is better than that of the parent individual of the corresponding sub-problem; More preferably, the performance comparison unit is further configured to determine that the performance of the offspring individual is better than that of the parent individual of the corresponding sub-problem if the Chebyshev value of the offspring individual is smaller than the Chebyshev value of the parent individual of the corresponding sub-problem.

[0061] In the embodiments of the present invention, each unit of the neural network architecture search optimization device can be implemented by corresponding hardware or software units. Each unit can be an independent software or hardware unit, or can be integrated into a software or hardware unit. The present invention is not limited thereto. For the specific implementation manners of each unit of the neural network architecture search optimization device, reference can be made to the relevant descriptions in the foregoing method embodiments, and details are not described herein again.

[0062] Example 3: Figure 3 FIG. shows the structure of the electronic device provided in Embodiment 3 of the present invention. For the sake of convenience of description, only the parts related to the embodiments of the present invention are shown.

[0063] The electronic device 3 in the embodiments of the present invention includes a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30. When the processor 30 executes the computer program 32, it implements the steps in the embodiments of the above-mentioned various neural network architecture search optimization methods, such as Figure 1A the steps S101 to S108 shown. Alternatively, when the processor 30 executes the computer program 32, it implements the functions of each unit in the above-mentioned device embodiments, such as Figure 2 the functions of the units 21 to 27 shown.

[0064] In the embodiments of the present invention, a neural network architecture is randomly initialized to generate a multi-task population. The neural network architecture search problem is modeled as a multi-objective optimization problem. The Chebyshev polynomial is used to decompose the multi-objective optimization problem into multiple sub-problems. An adaptive sub-problem selection strategy based on the sub-problem improvement rate is used to select from multiple sub-problems. An evolutionary operation is performed on each selected sub-problem to generate offspring individuals. The multi-task population is updated according to the offspring individuals. Through multiple iterations, the optimal architecture of the neural network is output. Thus, the differences between sub-problems are quantified by the adaptive sub-problem selection strategy, and computing resources are allocated to more potential sub-problems, reducing the waste of computing resources and accelerating the convergence speed of the population. The present invention also improves the migration effect by integrating the adaptive sub-problem selection strategy and the adaptive migration evolution strategy, thereby improving the performance of neural network architecture search.

[0065] The electronic device in the embodiments of the present invention can be a neural network architecture search optimization device. The steps implemented when the processor 30 in the electronic device 3 executes the computer program 32 to implement the neural network architecture search optimization method can refer to the description of the foregoing method embodiments and will not be repeated here.

[0066] Example 4: In the embodiments of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the steps in the embodiments of the above-mentioned neural network architecture search optimization method, such as, Figure 1A the steps S101 to S108 shown. Alternatively, when the computer program is executed by a processor, it implements the functions of each unit in the above-mentioned device embodiments, such as Figure 2 the functions of the units 21 to 27 shown.

[0067] In an embodiment of the present invention, a neural network architecture is randomly initialized to generate a multi-task population. The neural network architecture search problem is modeled as a multi-objective optimization problem. The Chebyshev polynomial is used to decompose the multi-objective optimization problem into multiple sub-problems. An adaptive sub-problem selection strategy based on the sub-problem improvement rate is used to select from multiple sub-problems. Evolution operations are performed on each selected sub-problem to generate offspring individuals. The multi-task population is updated according to the offspring individuals. Through multiple iterations, the optimal architecture of the neural network is output. Thus, the differences between sub-problems are quantified by the adaptive sub-problem selection strategy, and computing resources are allocated to more promising sub-problems, reducing the waste of computing resources and accelerating the convergence speed of the population. The present invention also improves the migration effect and further improves the performance of neural network architecture search by integrating the adaptive sub-problem selection strategy and the adaptive migration evolution strategy.

[0068] The computer-readable storage medium of the embodiment of the present invention may include any entity or device, recording medium that can carry computer program code, for example, memories such as ROM / RAM, magnetic disks, optical discs, flash memories, etc.

[0069] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A neural network architecture search optimization method, characterized in that: The method comprises the following steps: Randomly initialize the neural network architecture to generate a multi-task population; Model the neural network architecture search problem as a multi-objective optimization problem; Decomposing the multi-objective optimization problem into a plurality of sub-problems by using Chebyshev polynomials; Selecting from the plurality of sub-problems using an adaptive sub-problem selection strategy based on sub-problem improvement rates; Perform evolutionary operations on each selected sub-problem to generate offspring individuals; Updating the multi-task population according to the offspring individuals; Determine whether the iteration termination condition is met. If not, jump to the step of selecting from the multiple sub-problems using an adaptive sub-problem selection strategy based on the sub-problem improvement rate. If it is met, output the optimal architecture of the neural network.

2. The method according to claim 1, characterized in that Each individual in the multi-task population represents a neural network architecture, and the multi-objective optimization problem includes a two-objective minimization problem of the error rate of the neural network on the data set and the complexity of the neural network architecture.

3. The method according to claim 1, characterized in that The step of selecting from the plurality of sub-problems using an adaptive sub-problem selection strategy based on the sub-problem lifting rate comprises: Calculate the improvement rate of each individual on each sub-problem and sort them in descending order; Before selection N The sub-problem corresponding to the improvement rate is N is a natural number greater than 0; The calculation formula of the improvement rate is as follows: in, x j i,old Indicates i Task No. j Parent individuals, x j i,new Indicates i Task No. j The offspring individuals produced by the parent generation, λ j i Indicates i Task No. j A sub-question, z i Indicates i The preset ideal point for each task, FIR i j Indicates i Task No. j The improvement rate of the performance of each sub-problem, g tch () represents the Chebyshev value of the multi-task population individual, x represents the decision variable of an individual in the multi-task population, f j ( x ) represents the individual j target value.

4. The method according to claim 1, characterized in that The evolutionary strategies used when performing evolutionary operations on each selected sub-problem include migration evolutionary strategies and intra-task evolutionary strategies. The evolutionary strategy of each selected sub-problem uses a preset migration intensity for adaptive adjustment. The evolutionary strategy for each selected sub-problem uses the preset migration strength adaptive adjustment steps, including: When the probability of random generation is less than the migration intensity, the migration evolution strategy is adjusted; When the probability of random generation is not less than the migration intensity, the intra-task evolution strategy is adjusted.

5. The method according to claim 4, characterized in that The migration evolution strategy is an adaptive migration evolution strategy based on sliding window and operator value evaluation, and the adaptive migration evolution strategy includes: Initializing an operator pool and a sliding window, wherein the operator pool includes at least two of a uniform crossover operator, a single-point crossover operator, a geometric crossover operator, a simulated binary crossover operator, and a differential evolution operator; Recording the operator and the promotion rate of the offspring generated by the operator through the sliding window; Operator value evaluation is performed based on the improvement rate and usage frequency of each operator, and the optimal operator is dynamically selected based on the evaluation results; The calculation formula for operator value assessment is as follows: in, op best represents the current optimal operator, l represents the sliding window size, Q ( op i ) indicates the i The improvement rate of the sliding window operator, fre Indicates the corresponding frequency of use, N ( op j ) indicates the j The number of times the operator of a sliding window is used, N ( op i ) indicates the i The number of times the operator of a sliding window is used.

6. The method according to claim 1, characterized in that The step of updating the multi-task population according to the offspring individuals comprises: Comparing the performance of the offspring individuals with the parent individuals of the corresponding subproblem; If the performance of the offspring individual is better than the parent individual of the corresponding sub-problem, the offspring individual is used to replace the parent individual of the corresponding sub-problem; If the Chebyshev value of the offspring individual is smaller than the Chebyshev value of the parent individual of the corresponding subproblem, it is determined that the performance of the offspring individual is better than the parent individual of the corresponding subproblem.

7. A neural network architecture search optimization device, characterized in that: The device comprises: Initialization unit, used to randomly initialize the neural network architecture and generate multi-task population; A modeling unit for modeling the neural network architecture search problem as a multi-objective optimization problem; A sub-problem decomposition unit, used to decompose the multi-objective optimization problem into multiple sub-problems using Chebyshev polynomials; A sub-problem selection unit, configured to select from the plurality of sub-problems based on an adaptive sub-problem selection strategy based on a sub-problem improvement rate; Evolution unit, used to perform evolution operations on each selected sub-problem; an updating unit, configured to update the multi-task population according to the offspring individuals; and The judgment and output unit is used to judge whether the iteration termination condition is met. If not, the sub-problem selection unit is triggered to execute the step of selecting from the multiple sub-problems based on the adaptive sub-problem selection strategy based on the sub-problem improvement rate. If it is met, the optimal architecture of the neural network is output.

8. The device according to claim 7, characterized in that The sub-question selection unit also includes: An improvement rate calculation unit, used to calculate the improvement rate of each individual on each sub-problem and sort them in descending order; and Select subunit, used to select the front N The sub-problem corresponding to the improvement rate is N is a natural number greater than 0; The calculation formula of the improvement rate is as follows: in, x j i,old Indicates i Task No. j Parent individuals, x j i,new Indicates i Task No. j The offspring individuals produced by the parent generation, λ j i Indicates i Task No. j A sub-question, z i Indicates i The preset ideal point for each task, FIR i j Indicates i Task No. j The improvement rate of the performance of each sub-problem, g tch () represents the Chebyshev value of the multi-task population individual, x represents the decision variable of an individual in the multi-task population, f j ( x ) represents the individual j target value.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Active power distribution network day-ahead high-dimensional target optimization scheduling method based on improved MOEA / D

    CN110956324A