Optimization method for realizing multi-performance architecture of neural network at low cost
Through the PASAEA method, the PA-Dropout proxy model and DDIMMS strategy are used to solve the problem of multi-performance architecture optimization in deep convolutional neural networks at low cost, and the efficient hyperparameter configuration and the balance of multi-performance indicators is achieved, which significantly reduces training costs and time.
Patent Information
- Application Number
- CN202510323644.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-13
Smart Images

Figure CN120146110A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and relates to neural network hyperparameter configuration and recommendation systems. Specifically, it is an optimization method for low-cost implementation of a multi-performance architecture of a neural network. Background Art
[0002] With the wide penetration of deep learning in various fields, the scale and complexity of deep convolutional neural network models have increased exponentially. In the field of image recognition, in order to accurately identify tiny targets in complex scenes, the model needs to have deeper network layers and more parameters; in speech recognition tasks, in order to adapt to different accents, speech rates, and complex environmental noises, the model structure is continuously refined. Although this development trend has significantly improved the model performance, it has also caused the computing resources required for the training process to increase explosively. Training a large-scale deep convolutional neural network often requires a large amount of GPU computing power and a long time. Such high costs have made many research institutions and enterprises face severe resource constraints and cost pressures when promoting related projects, and encounter many obstacles in the project planning and implementation process.
[0003] At the same time, the requirements for the performance of deep neural networks in actual application scenarios are becoming more and more diverse. In addition to the traditional focus on accuracy, in scenarios with high real-time requirements, such as target detection in autonomous driving, the inference speed of the model is crucial; on resource-constrained devices, such as image classification applications on mobile terminals, the memory occupancy of the model must be strictly controlled. Facing such diverse and complex performance requirements, the existing hyperparameter configuration methods have shown significant limitations. On the one hand, traditional single-performance optimization methods cannot simultaneously meet multiple performance indicators, resulting in a serious decline in other performances when improving a certain performance. On the other hand, even though some methods attempt multi-performance optimization, due to the extreme complexity of the hyperparameter space and the high computing cost, it is difficult to implement in actual applications.
[0004] Under this background, it is urgent to build a recommendation method for a multi-performance architecture solution that can address the expensive training problem of neural networks. This system can not only efficiently search in the vast hyperparameter space to accurately match the optimal hyperparameter combination for different tasks and datasets, but also comprehensively consider multiple performances such as the accuracy, inference speed, and memory occupancy of the model, so as to optimize the multi-performance architecture of the neural network at low cost, thereby strongly promoting the wide application of artificial intelligence technology in classification tasks in various complex scenarios. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide an optimization method for low-cost implementation of a multi-performance architecture of a neural network.
[0006] The present invention is implemented by the following technical solutions:
[0007] An optimization method for implementing a multi-performance architecture of a neural network at low cost, including four major modules, namely a data set input module, a parameter setting module, a multi-performance architecture hyperparameter search module, and a multi-performance neural network architecture recommendation module. Taking the error rate, computational complexity, and test time as performance goals, regarding the hyperparameter combinations in the deep convolutional neural network as population individuals, by designing PASAEA (including the PA-Dropout surrogate model and DDIMMS), inputting the corresponding data set, classification network, and hyperparameter search range, the corresponding multi-performance architecture solution can be efficiently and automatically searched for.
[0008] The data set input module: responsible for reading the input data set, used to perform classification tasks, so as to evaluate the actual performance of the entire architecture solution.
[0009] The parameter setting module: covers the hyperparameter range of the classification network and the parameter setting content of the PASAEA algorithm. This module will perform corresponding data conversion operations according to the given hyperparameter range, and then transmit the converted hyperparameters to the multi-performance architecture hyperparameter search module.
[0010] The multi-performance architecture hyperparameter search module: obtains the characteristic parameter information of the parameter setting module and the data set in the data set input module, and combines them with the input classification network. This module searches for the hyperparameters in the network, that is, for the established network, it truly calculates the error rate, computational complexity, and test time, constructs a PA-Dropout surrogate model with the samples of the true calculation as the training set, and then calculates the non-dominated solution set in the population. In this process, the target evaluation does not require expensive true calculations, and only relies on the low-cost PA-Dropout to obtain the performance estimation of the classification network; then through DDIMMS, select individuals that balance convergence and diversity for the true evaluation of the classification network and update the training set. Continuously iterate the PASAEA optimization process, and finally obtain a batch of optimal hyperparameter combinations and the corresponding multiple performance indicators, and save these contents to the optimization result path. Subsequently, read the multiple performance indicator data in the.mat file from the optimization result path, and draw a multi-performance target distribution map based on this. On this distribution map, select points A, B, and C. Among them, the error rate is the smallest at point A, the network calculation amount is the smallest at point B, and the network has the shortest test time at point C. Through the three scheme buttons of A, B, and C, the hyperparameters of these three schemes obtained by the search can be displayed on the hyperparameter combination configuration scheme interface.
[0011] The multi-performance neural network architecture recommendation module: can read the corresponding hyperparameter scheme from the hyperparameter combination configuration scheme according to the user's performance requirements. Substitute this set of schemes into the classification network respectively to implement the recommendation of a multi-performance architecture scheme that meets the user's performance requirements.
[0012] The present invention has the following beneficial effects:
[0013] First, the present invention proposes a multi-performance architecture scheme recommendation method for solving the expensive training problem of neural networks. With application performance objectives such as error rate, computational complexity, and perception function, high-dimensional hyperparameter combinations are used as population individuals, and by designing PASAEA (progressive accumulation neural network based surrogate-assisted evolution algorithm, PASAEA), including PA-Dropout and DDIMMS, the mapping relationship between hyperparameter configurations and performance objectives is fully fitted, and finally, the high-dimensional hyperparameters of the deep convolutional neural network are efficiently configured at a low computational cost.
[0014] Second, the present invention proposes a PA-Dropout neural network surrogate model. By designing a two-layer Dropout lightweight network and a weight screening strategy, the contribution degree of hyperparameter combinations to the multi-objective performance of the deep convolutional neural network is dynamically measured, and the effective calculation of discriminant information is iteratively screened to improve the fitting accuracy and efficiency of PA-Dropout for the multi-objective performance of the deep convolutional neural network.
[0015] Third, the present invention proposes DDIMMS. According to the high-dimensional hyperparameter space distribution state, a dynamic adaptive weighted calculation strategy that takes into account both convergence and diversity is designed. While reducing the influence of subjective human design of parameters, the reliability of the true evaluation data set for hyperparameter configuration is improved, and further, the update efficiency and fitting accuracy of PA-Dropout are improved, realizing the efficient configuration of hyperparameters in the multi-performance architecture of the final deep convolutional neural network.
[0016] Fourth, the present invention experimentally compares PASAEA and its four variants with six representative surrogate-assisted evolution algorithms, showing its excellent performance in the expensive training problem. For the hyperparameter configuration tasks of classification networks on the weld defect data set TIDIP and the CIFAR-10 data set, the accuracy of the classification model optimized by PASAEA reaches 93.52% and 84.8% on the weld data set TIDIP and the CIFAR-10 data set respectively, and the training time is increased by 29.87% and 14.15% on average compared with other surrogate-assisted evolution algorithms, and is even shortened by 50.6% and 48.2% compared with the multi-objective optimization algorithm, verifying the effectiveness and generalization of the proposed algorithm in the hyperparameter configuration of the deep convolutional neural network.
[0017] The present invention is reasonably designed and can be widely applied to the construction and research in the fields of automation, signal processing, etc., significantly reducing the human and time costs, and having extremely high engineering application value. Description of the Drawings
[0018] Figure 1 Shows the overall flow chart of the present invention.
[0019] Figure 2 Shows the flow chart of the hyperparameter configuration of the PASAEA-assisted deep convolutional neural network.
[0020] Figure 3 Represents the dataset input module.
[0021] Figure 4 Represents the parameter setting module.
[0022] Figure 5 Represents the multi-performance architecture hyperparameter search module.
[0023] Figure 6 Represents the multi-performance neural network architecture recommendation module.
[0024] Figure 7 Shows the schematic diagram of the model training of PA-Dropout.
[0025] Figure 8 Shows the IGD comparison convergence graph of PASAEA and 6 proxy-assisted evolutionary algorithms on DTLZ1.
[0026] Figure 9 Shows the IGD comparison convergence graph of PASAEA and 6 proxy-assisted evolutionary algorithms on DTLZ2.
[0027] Figure 10 Shows the task objective distribution graph of the hyperparameter configuration of the classification network by PASAEA in TIDIP.
[0028] Figure 11 Shows the task objective distribution graph of the hyperparameter configuration of the classification network by PASAEA in the CIFAR-10 dataset.
[0029] Figure 12 Shows the comparison graph of the optimization time of NoPA-Dropout without proxy model of PASAEA and 4 latest multi-objective optimization algorithms. Detailed Implementation Manner
[0030] The technical solution of the present invention will be further described in detail below in conjunction with the drawings and embodiments. These are some embodiments of the present invention, not all of them. It should be noted that, without conflict, the embodiments and features in the embodiments of the present invention can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0031] An optimization method for implementing a multi-performance architecture of a neural network at low cost, as Figure 1 shown, includes four major modules, namely a dataset input module, a parameter setting module, a multi-performance architecture hyperparameter search module, and a multi-performance neural network architecture recommendation module. This method takes the error rate, computational complexity, and test time as performance goals, regards the hyperparameter combinations in the deep convolutional neural network as population individuals, and by designing PASAEA (including the PA-Dropout surrogate model and DDIMMS), inputting the corresponding dataset, classification network, and hyperparameter search range, can efficiently and automatically search for a multi-performance architecture solution corresponding to the expensive training problem of the neural network.
[0032] (1) Dataset input module (as Figure 3 shown): Responsible for reading the input dataset, which is used to perform classification tasks, so as to evaluate the actual performance of the entire architecture solution.
[0033] (2) Parameter setting module (as Figure 4 shown): Covers the hyperparameter range of the classification network and the parameter setting content of the PASAEA algorithm. This module will perform corresponding data conversion operations according to the given hyperparameter range, and then transmit the converted hyperparameters to the multi-performance architecture hyperparameter search module.
[0034] (3) Multi-performance architecture hyperparameter search module (as Figure 5 shown): Obtains the characteristic parameter information of the parameter setting module and the dataset in the dataset input module, and combines them with the input classification network. This module searches for the hyperparameters in the network, that is, for the established network, it truly calculates the error rate, computational complexity, and test time, constructs a PA-Dropout surrogate model with the real calculation samples as the training set, and then calculates the non-dominated solution set in the population. In this process, the target evaluation does not require expensive real calculations, and only relies on the low-cost PA-Dropout to obtain the performance evaluation of the classification network; then through DDIMMS, it selects individuals that balance convergence and diversity for the real evaluation of the classification network and updates the training set. Continuously iterate the PASAEA optimization process, and finally obtain a batch of optimal hyperparameter combinations and the corresponding multiple performance indicators, and save these contents to the optimization result path. Subsequently, read the multiple performance indicator data in the.mat file from the optimization result path, and draw a multi-performance target distribution map based on this. On this distribution map, select points A, B, and C. At point A, the error rate is the smallest. At point B, the network calculation amount is the smallest. At point C, the network has the shortest test time. Through the three scheme buttons of A, B, and C, the hyperparameters of these three schemes obtained by the search can be displayed on the hyperparameter combination configuration scheme interface.
[0035] (4) Multi-performance neural network architecture recommendation module (asFigure 6 As shown in the figure: It can read the corresponding hyperparameter solution from the hyperparameter combination configuration solution according to the performance requirements of the user. Substitute this set of solutions into the classification network respectively to implement a multi-performance architecture solution that meets the user's performance requirements for the user.
[0036] First, in order to meet the user's needs and take into account the multi-performance requirements in the actual task, the present invention embodiment sets application objectives such as error rate, computational complexity, and test time as the optimization objectives of PASAEA, and carries out hyperparameter configuration work on the deep convolutional neural network. The specific configuration process is as follows, as Figure 2 shown:
[0037] First, input the data set, then preset the network architecture of the deep convolutional neural network, and determine the optimization parameter combination. Among them, the hyperparameter types are the number of convolutional channels, activation function, gradient descent function, learning rate, and batch size. Then initialize the population. Each individual in the population is regarded as a set of hyperparameter configurations. The dimension of the individual is the number of hyperparameters, and the decision variable is the corresponding hyperparameter, and substitute it into the neural network.
[0038] In various different application tasks, computational complexity and time are common performance objectives. Regarding the measurement of computational complexity, the floating point operation count FLOPs (floating point operations, abbreviated as FLOPs) of the model is used to reflect the computational efficiency of the model when processing data. FLOPs is calculated by the following formula:
[0039]
[0040] where A is the number of convolutional layers, B is the number of fully connected layers, C i in is the number of input channels of the i-th convolutional layer, C i out is the number of output channels of the i-th convolutional layer, K i is the side length of the convolutional kernel of the i-th convolutional layer, H i out is the height of the output feature map of the i-th convolutional layer, W i out is the width of the output feature map of the i-th convolutional layer; N j in is the number of input neurons of the j-th fully connected layer; N j out is the number of output neurons of the j-th fully connected layer.
[0041] Regarding the test time, it is defined as the time for the network to execute the application task on a single data set, and its calculation formula is:
[0042] T = Tend -T start
[0043] where T start is the start time of executing this task, and T end is the end time of executing this task.
[0044] For different application tasks, there are also specific performance goals:
[0045] (1) In the classification task, the classification error rate is used as the fitness function, denoted as F 分类任务 , and it is calculated according to the following formula, where t(i) represents the true label of each sample, p(i) is the corresponding predicted label, and N represents the total number of labels.
[0046]
[0047] (2) In the super-resolution task, the perception function PF is used as the fitness function, denoted as F 超分辨率 , and v lpips (i) represents the LPIPS value between the i-th image output by the generator and the real image, and v pI (i) represents the PI value between the i-th image output by the generator and the real image. batch is the batch size. When F 超分辨率 reaches the minimum value, the reconstructed image has better perceptual quality, and the specific calculation formula is as follows:
[0048]
[0049] (3) In the object detection task, in order to have a higher mean average precision, the sum of 10% weight mAP 0.5 and 90% weight mAP 0.5-0.95 is used as the fitness function, denoted as F 目标检测 , and the following calculation formula is used:
[0050] F 目标检测 = 0.1 × mAP 0.5 + 0.9 × mAP 0.5-0.95
[0051] In the formula, mAP 0.5 , mAP 0.5-0.95 respectively represent the average precision of the uniform distribution with the threshold IOU = 0.5 and the IOU step size 0.05 from 0.5 to 0.95.
[0052] During the PASAEA optimization process, the common performance and specific performance of different tasks are used as multi-performance objectives for real calculations. The samples of real calculations are used to construct the PA-Dropout surrogate model for the training set. Then, the non-dominated solution set in the population is calculated. In this process, the objective evaluation does not require expensive real calculations and only relies on the low-cost PA-Dropout to obtain the performance estimation of the deep convolutional neural network. Then, through DDIMMS, individuals that balance convergence and diversity are selected for the real evaluation of the deep convolutional neural network, and the training set is updated. The PASAEA optimization process is continuously iterated, and finally a batch of optimal hyperparameter configurations are obtained, realizing the balanced optimization of multi-performance objectives.
[0053] Second, in order to effectively improve the fitting accuracy and efficiency of the high-dimensional hyperparameter combination in the deep convolutional neural network and the multi-performance objectives when performing application tasks, the PA-Dropout surrogate model is constructed as follows:
[0054] First, initialize the weight matrix W i and the bias matrix B i , i = 1, 2, 3, to form the initial network Net(0). Divide the training data into a training set (X d , Y d ) composed of samples with a batch size of d (consistent with the decision variable size), ensuring that the input is in square matrix form to avoid partial irreversible operations in high-dimensional spaces. The training data passes through two-layer feedforward neural networks. The specific process is as shown in the Net update mechanism part in Figure 7 : In the initial stage, Net(0) and {X {1} , Y {1}} are used as the starting inputs, where X {1} is defined as the input data x o . For the input data x o , the predicted output y o is obtained through forward propagation. The specific calculation is as follows:
[0055] y h1 = f R ((W 1 oM)x o + B 1 )
[0056] y h2 = f R ((W 2 oM)y h1 + B 2 )
[0057] y o = W 3 oy h2 + B 3
[0058] Among them, y h1 , y h2 respectively represent the outputs of the first and second layer neurons. W 1 , W 2 , W 3 and B 1 , B 2 , B 3 represent the weight matrices and bias vectors required for solving each layer, and f R represents the Relu function activation method. The execution process of W○M is as follows Figure 7 Weight screening strategy part: First, expand the weight matrix into a one-dimensional vector, u is the dimension of this one-dimensional vector, and then sort the one-dimensional vector in ascending order so that W i >W i+1 , and retain the original index. Then create a one-dimensional vector M of the same type as the expanded weight vector. The first 1 / 2 of M is M 1 , and the remaining is M 2 , where the relationship between M and [M 1 , M 2 is as follows:
[0059] M = [M 1 , M 2
[0060] Then ○ represents the Hadamard product, that is, set the discarded weights to zero, and scale the retained weights by a factor of 1 / (1 - h) to ensure that the expected output of the network including the Dropout operation is the same as the expected output of the network without Dropout. W○M = W○[M 1 , M 2 , that is, all the weights in the M 1 part are retained, while for the weights in the M 2 part, where M 2 is a probability vector following the Bernoulli distribution:
[0061] M 2 : Bernourlli(h)
[0062] where the probability value h ∈ [0.5, 1), and the weights in the M 2 part are processed by Dropout according to the probability vector of the Bernoulli distribution. Since the output layer neurons correspond one-to-one with the specific prediction or classification results, if W 3 applies Dropout, the Net prediction result will be unstable, resulting in inconsistent outputs during training and testing, affecting performance evaluation. Therefore, W 3 does not perform weight screening.
[0063] After the weight screening, the weight vector is restored to the weight matrix according to the original index, and the updated W can be obtained. According to the previously calculated y o , where y o is the predicted value of the forward propagation data, and Y {1} is used as the data input value Y o . Then, calculate the error between the network input and output, and use the error as the loss function. The loss update formula is as follows:
[0064] Loss = y o - Y o
[0065] Combine the weights and biases updated by the gradient and the chain rule for backpropagation to obtain Net(1). Then, according to the set number of iterations run, continuously execute the above training process to obtain Net(run). Net(run) is the final surrogate model Net. After the surrogate model training is completed, the solution of multiple performance objectives of the deep convolutional neural network can be approximately equivalent to the target value by simply performing 1 fast calculation using the trained PA-Dropout. During testing, the closed Dropout neural network is directly fitted without additional parameter adjustment and complex calculations, and the target value y obj is quickly obtained, which not only improves the operation efficiency but also reduces the consumption of computing resources, strongly supporting the rapid deployment and efficient operation of the deep convolutional neural network in practice.
[0066] Through the above steps, the contribution degree of the dynamic metric hyperparameter combination of the PA-Dropout surrogate model to the multi-objective performance of the deep convolutional neural network is measured, and the effective calculation of discriminant information is iteratively screened to improve the fitting accuracy and efficiency of PA-Dropout for the multi-objective performance of the deep convolutional neural network.
[0067] Third, the embodiment of the present invention adaptively balances the convergence and diversity of the algorithm to achieve the scientific selection of real computing individuals, thereby improving the update efficiency and fitting efficiency of PA-Dropout, and finally achieving the efficient recommendation of the multi-performance architecture scheme in the deep convolutional neural network. A dual drive interactive model management strategy (DDIMMS) is proposed to provide reliable candidate solutions for real evaluation.
[0068] First, perform normalization processing on the initial population P and the population P c optimized by PA-Dropout. Then, calculate the minimum distance sum from the individuals in P and P c to their nearest neighbor individuals respectively. The minimum distance sum is denoted as Crowd, and the calculation formula of Crowd is as follows:
[0069] Crowd = ∑ x∈G min{d(x, y) = ||x - y|| 2 , |y ∈ G, y ≠ x}
[0070] In the above formula, G represents the population, x and y are individuals in the set G, and d(x, y) represents the distance between individual x and individual y. Through this formula, the result obtained is the value of Crowd. Calculate CrowdOld based on the population P, and calculate CrowdNew based on the population P c to reflect the tightness and diversity of the individual distribution in the population before and after optimization. Subsequently, taking the Euclidean distance as the measurement standard, calculate the spatial distance between individuals in P using the following formula c .
[0071]
[0072] Suppose x and y are the i-th and j-th individuals in P respectively c , and k represents the k-th decision variable. Dist_D(i, j) is the diversity evaluation of the individuals between i and j. To comprehensively evaluate the performance of the individuals in P c , the present invention introduces the calculation of the convergence index. This index measures the closeness between the non-dominated solution set in P c and the ideal point z*, aiming to capture the ability of the individuals to approach the optimal solution in the objective space. The calculation formula is as follows:
[0073]
[0074] where x represents the i-th individual, and k represents the k-th decision variable. Dist_C(i) is the convergence evaluation of the i-th individual. To avoid calculation errors caused by the measurement exceeding the data range, it is necessary to perform a normalization operation on Dist_D(i, j) and Dist_C(i) to obtain the normalized Dist_D 1 (i, j) and Dist_C 1 (i). Then compare the values of CrowdOld and CrowdNew. If CrowdOld ≤ CrowdNew, it indicates that the individual diversity of the optimized population is better. At this time, the convergence should be emphasized, with convergence as the main structure and diversity as the auxiliary structure. First, calculate the population diversity attention weight. The D matrix is an N×N matrix, and D ij represents the diversity attention weight assigned by the i-th individual to the j-th individual. D ij is calculated according to the following formula:
[0075]
[0076] Calculate the convergence of the i-th individual as the main comprehensive index and update it through the following formula:
[0077]
[0078] in Taking the inverse of Dist_D(i,j), the underlying logic is: the convergence index seeks to be minimized, the smaller it is, the closer it is to the optimal solution; while the diversity seeks to be maximized, the larger it is, the more areas can be explored. Through this transformation, the two work together on C i Calculate, reflect and optimize convergence and diversity at the same time, and then select C i The individuals with the smallest value are truly evaluated. These individuals have achieved a good balance between convergence and diversity. If CrowdOld>CrowdNew, it means that the population diversity is insufficient. At this time, diversity should be set as the main structure and convergence as the auxiliary structure. First calculate the population convergence attention weight and create an N×1 C matrix, C i1 It represents the convergent attention weight of the i-th individual and is calculated as follows.
[0079]
[0080] The diversity of the i-th individual is calculated as the main comprehensive index through the following formula:
[0081]
[0082] in Echoing the previous logic, Dist_C 1 (i) Take the reciprocal and finally choose D i The largest individual makes a true evaluation.
[0083] Fourth, in order to verify the effectiveness of the present invention, the embodiment of the present invention performs simulation comparisons with other agent-assisted evolutionary algorithms on the DTLZ and WFG test functions, and tests the hyperparameter configuration tasks of deep convolutional neural networks on the weld defect datasets TIDIP and CIFAR-10 datasets.
[0084] To deeply verify the performance of the PASAEA algorithm, it was comprehensively compared with six representative surrogate-assisted evolutionary algorithms (PCSAEA, MOL2SMEA, EDN-ARMOEA, ParEGO, K-RVEA, MOEA / D-EGO). In the experiment, all algorithms were initialized with real sampling to generate a training dataset of 11d - 1, and the maximum number of real evaluations was set to 11d + 119 to evaluate the effectiveness of the algorithms under limited evaluation resources. The research focused on investigating the influence of the decision space dimension and the number of objectives on the algorithm performance. In particular, given that the ParEGO algorithm was originally designed to focus on solving multi-objective optimization problems with no more than 4 objectives, the ParEGO algorithm was only compared and tested on problem instances with 3 objectives.
[0085] Table 1 shows the average IGD results of PASAEA and the six algorithms on the DTLZ benchmark when the fixed objective m = 3 and the decision variables d = 20, 40, 60, 100. It can be seen from Table 1 that PASAEA obtained 19, 11, 20, 23, 22, and 20 better results than other algorithms in 28 evaluations respectively, demonstrating excellent high-dimensional decision space solving ability. Among them, PASAEA performs well on DTLZ1 because it can handle linear problems and uniformly distributed Pareto fronts, and at the same time, its strong global search ability avoids the problem of falling into local optimal solutions existing in DTLZ3, making PASAEA perform better on DTLZ3.
[0086] Table 1 Average IGD values obtained by PASAEA and six comparison algorithms on the 3-m DTLZ test problem
[0087]
[0088] Table 2 shows the IGD results of PASAEA and the six algorithms on the DTLZ benchmark with different m values when d = 40. PASAEA obtained 16, 11, 15, 19, and 14 better IGDs than other algorithms in 28 test functions respectively. Even in the 3-objective optimization of ParEGO, among the total 7 problems, PASAEA performed better in 6 problems, highlighting its wide applicability and competitiveness. At the same time, as the number of objectives increases, the search space of the DTLZ7 problem becomes complex, requiring more computing resources and time, and the limited real evaluations lead to unsatisfactory performance on DTLZ7.
[0089] Table 2 Average IGD values obtained by PASAEA and six comparison algorithms on the 40-d DTLZ test problem
[0090]
[0091] Meanwhile, on the DTLZ1 and DTLZ2 test problemsFigure 8 Shows the IGD comparison convergence graph of PASAEA and six surrogate-assisted evolutionary algorithms on DTLZ1, Figure 9 shows the IGD comparison convergence graph of PASAEA and six surrogate-assisted evolutionary algorithms on DTLZ2. In Figure 8 , the initial evaluation number of PASAEA is 40, that of MOL2SMEA is 50, and the others are 219. As the number of evaluations increases, compared with the maximum true evaluation times at 254, it decreases by 25.07%. PASAEA and ParEGO obtain smaller IGD values and tend to be stable, and have better performance at the maximum evaluation times. In Figure 9 , the initial evaluation number is the same as that in Figure 8 . Before 254, MOL2SMEA has better performance because it is good at dealing with DTLZ2 containing non-linearity. While for PASAEA, as the number of evaluations increases, the fitting accuracy of PA-Dropout improves, and it can better handle non-linear relationships. Finally, the convergence result is better than that of similar algorithms. Generally speaking, this algorithm can quickly find a solution set better than the comparison algorithms.
[0092] In order to fully verify the effectiveness of the method of the present invention in solving practical classification problems, the embodiments of the present invention use the high-performance PASAEA verified by experiments to configure the hyperparameters of the deep convolutional neural network for classification. The datasets used are the weld defect dataset TIDIP and the CIFAR-10 dataset. Among them, TIDIP is cooperated with Shanxi Provincial Institute of Mechanical and Electrical Design and Research Co., Ltd., and it contains 5 categories: crack, porosity, slag inclusion, lack of fusion, and lack of penetration. Select the time-frequency domain samples obtained by converting the samples through Short-Time Fourier Transform. Finally, the dataset contains 10,880 pieces, including 2,420 cracks, 2,120 porosities, 2,120 slag inclusions, 2,100 lack of fusions, and 212 lack of penetrations. The dataset and the training set are divided according to 4:1. The CIFAR-10 dataset contains 60,000 pictures of 10 real animals and objects in real life, such as airplanes, cars, birds, deer, etc., and the training set and the dataset are divided according to the traditional 5:1.
[0093] Apply the PASAEA algorithm to the weld defect dataset TIDIP and the CIFAR-10 dataset to configure the hyperparameters of the deep convolutional neural network, and then construct the TIDIP recognition model and the CIFAR-10 recognition model. When PASAEA completes the iteration, a batch of hyperparameter combinations that can best balance the three objectives is obtained. For the performance of these combinations on the three key objectives, the objective distribution Figure 10 and Figure 11, the error rate in the figure is divided into two levels. 0 - 0.5 is the low error rate, and 0.5 - 1 is the high error rate. Under the condition of ensuring a relatively low recognition error rate, three points A, B, and C are selected from the points with low error rates. Among them, the error rate is the smallest at point A, the network calculation amount is the smallest at point B, and the network has the shortest test time at point C.
[0094] To further verify the acceleration effect of PASAEA in the high - dimensional hyperparameter configuration of deep convolutional neural networks for classification, the optimization time of PASAEA is compared with NoPA - Dropout without a surrogate model and four latest multi - objective optimization algorithms LERD, MOEADDQN, LMPFE, and HEA. The results are as Figure 12 shown. In the TIDIP and CIFAR - 10 recognition models, PASAEA speeds up by 50.6% and 48.2% on average compared with other multi - objective optimization algorithms respectively. Among them, the speed - up compared with the LERD algorithm is 74.94% and 64.14% respectively, which is more significant. The reason is that during the hyperparameter configuration process, the DVA (decision variable analysis) process in LERD reconstructs the complex variable relationship analysis into an optimization problem with binary decision variables, groups the decision variables. After successful grouping, each hyperparameter configuration is separately applied to the optimization tasks related to convergence or diversity. However, during the grouping process, due to the sparsity of the high - dimensional space, a large number of decision variables and their complex interaction relationships need to be processed, which makes the DVA process consume a large amount of time, and thus leads to a significant increase in the time cost of LERD. The experiment proves the efficiency of PASAEA, and also proves that the PA - Dropout surrogate model can improve the efficiency of high - dimensional hyperparameter configuration.
[0095] At the same time, seven surrogate - assisted evolutionary algorithms are applied to these two groups of recognition models, and the performance comparison is shown in Table 3. Under the setting of 216 maximum evaluation times, PASAEA reaches classification accuracies of 93.52% and 84.8% on the TIDIP recognition model and the CIFAR - 10 recognition model respectively, which are higher than those of other comparison algorithms. At the same time, the training time is shortened by 18.76% and 14.99% on average compared with the comparison algorithms on the two recognition models respectively. In addition, compared with the method using traditional Dropout as a surrogate model, PASAEA shortens the time by 6.59% on average and improves the classification accuracy by 1.66% on average on these two recognition models, verifying the superiority of the PA - Dropout in the hyperparameter configuration scenario of actual classification models. The above experimental results fully prove that the PASAEA method provides a scientific and highly stable solution for the high - dimensional hyperparameter configuration problem of deep convolutional neural networks for classification under the condition of limited computing resources, which not only meets the actual application requirements, but also significantly reduces the human and time costs, and has extremely high engineering application value.
[0096] Performance Comparison of Table 3 and Seven Agent-Assisted Evolutionary Algorithms Applied to Classification Networks
[0097]
[0098] Under the condition of limited computing resources, the present invention addresses the configuration problem of multi-performance architecture schemes in the expensive training problem of neural networks, realizes the paradigm shift from experience-driven to data-driven, breaks through the bottleneck of high subjective design cost and poor interpretability, and provides a scientific and highly stable solution for the expensive training problem of neural networks.
[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A low-cost optimization method for realizing a multi-performance architecture of a neural network, characterized in that: It includes a data set input module, a parameter setting module, a multi-performance architecture hyperparameter search module, and a multi-performance neural network architecture recommendation module; Taking error rate, computational complexity and test time as goals, and considering the hyperparameter combinations in deep convolutional neural networks as individuals in the population, PASAEA is designed. By inputting the corresponding data set, classification network and hyperparameter search range, it can efficiently and automatically search for the corresponding multi-performance architecture solutions.
2. The method for optimizing a neural network multi-performance architecture at low cost according to claim 1, characterized in that: Dataset input module: responsible for reading the input data set to perform classification tasks, so as to evaluate the actual performance of the entire architecture solution; Parameter setting module: covers the classification network hyperparameter range and PASAEA algorithm parameter setting content; this module performs corresponding data conversion operations according to the given hyperparameter range, and then transmits the converted hyperparameters to the multi-performance architecture hyperparameter search module; Multi-performance architecture hyperparameter search module: obtains the characteristic parameter information of the parameter setting module and the data set in the data set input module, and combines it with the input classification network; this module searches for the hyperparameters in the network, that is, for the established network, it performs real calculations on the error rate, computational complexity and test time, builds a PA-Dropout proxy model with the real calculated samples as the training set, and then calculates the non-dominated solution set in the population, relying on PA-Dropout to obtain the performance estimation of the classification network; then through DDIMMS, selects individuals that balance convergence and diversity to conduct real evaluation of the classification network, and updates the training set; then the PASAEA optimization process is continuously iterated, and finally a batch of optimal hyperparameter combinations and corresponding multiple performance indicators are obtained, and these contents are saved in the optimization result path; then, multiple performance indicator data in the .mat file are read from the optimization result path, and a multi-performance target distribution map is drawn based on this; on the distribution map, three points A, B, and C are selected, among which the error rate is the smallest at point A, the network calculation amount is the smallest at point B, and the network has the shortest test time at point C. Through the three solution buttons A, B, and C, the hyperparameters of the three solutions obtained by the search are displayed on the hyperparameter combination configuration solution interface; Multi-performance neural network architecture recommendation module: It can read the corresponding hyperparameter solution from the hyperparameter combination configuration solution based on the user's performance requirements; and substitute this set of solutions into the classification network to recommend multi-performance architecture solutions that meet the performance requirements of the user.
3. The method for optimizing a neural network multi-performance architecture at low cost according to claim 2, characterized in that: In the multi-performance architecture hyperparameter search module, the application objectives: error rate, computational complexity, and test time are set as the optimization goals of PASAEA, and the hyperparameter configuration of the deep convolutional neural network is carried out. The specific configuration process is as follows: First, input the data set, then preset the network architecture of the deep convolutional neural network, and determine the optimal parameter combination, where the hyperparameter types are the number of convolution channels, activation function, gradient descent function, learning rate, and batch size; then initialize the population, and each individual in the population is regarded as a set of hyperparameter configurations. The dimension of the individual is the number of hyperparameters, and the decision variable is the corresponding hyperparameter, which is brought into the neural network; In various application tasks, computational complexity and time are common performance goals. Regarding the measurement of computational complexity, the model's floating-point operation times FLOPs are used to reflect the model's computational efficiency when processing data. FLOPs is calculated using the following formula: Among them, A is the number of convolutional layers, B is the number of fully connected layers, and C is the number of i in is the number of input channels of the i-th convolutional layer, C i out is the number of output channels of the i-th convolutional layer, K i is the side length of the convolution kernel of the i-th convolutional layer, H i out is the height of the output feature map of the i-th convolutional layer, W i out N is the width of the output feature map of the i-th convolutional layer; j in N is the number of input neurons in the jth fully connected layer; j out is the number of output neurons of the jth fully connected layer; Regarding the test time, it is defined as the time it takes for the network to perform an application task on a single data set, and its calculation formula is: T=T end -T start Where T start is the start time of executing the task, T end The end time of the task; There are also unique performance targets for different application tasks: (1) In the classification task, the classification error rate is used as the fitness function and is denoted by F 分类任务 , calculated according to the following formula, t(i) represents the true label of each sample, p(i) represents the corresponding predicted label, and N represents the total number of labels; (2) In the super-resolution task, the perception function PF is used as the fitness function, denoted by F 超分辨率 , v lpips (i) represents the LPIPS value of the i-th image output by the generator and the real image, v pI (i) represents the PI value of the i-th image output by the generator and the real image, batch is the batch size, when F 超分辨率 The minimum value is reached, and the reconstructed image has good perceptual quality. The specific calculation formula is as follows: (3) In the object detection task, with a weight of 10% mAP 0.5 and 90% weight mAP 0.5-0.95 The sum is taken as the fitness function, denoted as F 目标检测 , using the following calculation formula: F 目标检测 =0.1×mAP 0.5 +0.9×mAP 0.5-0.95 Where mAP 0.5 、mAP 0.5-0.95 They represent the average precision from 0.5 to 0.95 with a threshold of IOU = 0.5 and an IOU step of 0.05, respectively.
4. The method for optimizing a neural network multi-performance architecture at low cost according to claim 3, characterized in that: The PA-Dropout proxy model is constructed as follows: First, initialize the weight matrix W i and the bias matrix B i , i = 1, 2, 3, forming the initial network Net(0), dividing the training data into a training set (X d ,Y d ), ensuring that the input is in square matrix form to avoid some irreversible operations in high-dimensional space; The training data passes through a two-layer feedforward neural network. The specific process is the Net update mechanism part. In the initial stage, Net(0) and {X {1} , Y {1} } as the starting input, where X {1} is defined as the input data x o , for the input data x o , and get the predicted output y through forward propagation o , the specific calculation formula is as follows: y h1 =f R ((W1 oM)x o +B1) <h2 style=";text-align:left;direction:ltr">y<h2 style=";text-align:left;direction:ltr"> h2 <h2 style=";text-align:left;direction:ltr"> =f<h2 style=";text-align:left;direction:ltr"> R <h2 style=";text-align:left;direction:ltr"> ((W2 oM)y<h2 style=";text-align:left;direction:ltr"> h1 <h2 style=";text-align:left;direction:ltr"> +B2) <h2 style=";text-align:left;direction:ltr">y<h2 style=";text-align:left;direction:ltr"> o <h2 style=";text-align:left;direction:ltr"> =W3 oy<h2 style=";text-align:left;direction:ltr"> h2 <h2 style=";text-align:left;direction:ltr"> +B3 where y h1 ,y h2 Represent the output of the first and second layers of neurons respectively, W1, W2, W3 and B1, B2, B3 represent the weight matrix and bias vector required for solving each layer, f R Represents the activation mode of the Relu function; the W○M execution process is the weight screening strategy part: first, the weight matrix is expanded into a one-dimensional vector, u is the dimension of this one-dimensional vector, and then the one-dimensional vector is sorted in ascending order so that W i >W i+1 , and retain the original index, and then create a one-dimensional vector M of the same type as the expanded weight vector, the first 1 / 2 of M is M1, and the rest is M2, where the relationship between M and [M1, M2] is as follows: M=[M1,M2] Then ○ represents the Hadamard product, that is, the discarded weights are set to zero, and the retained weights are scaled at a ratio of 1 / (1-h) to ensure that the expected output of the network including the Dropout operation is consistent with the expected output of the network without Dropout; W○M=W○[M1,M2], that is, all the weights of the M1 part are retained, and the weights of the M2 part, where M2 obeys the probability vector of the Bernoulli distribution: M2: Bernourli (h) Wherein the probability value h∈[0.5,1), the weight of the M2 part is Dropout processed according to the probability vector of the Bernoulli distribution; since the output layer neurons correspond one-to-one to the specific prediction or classification results, if W3 is Dropout, the Net prediction result will be unstable, resulting in inconsistent output during training and testing, affecting performance evaluation, so W3 does not perform weight screening; After weight screening, the weight vector is restored to the weight matrix according to the original index, that is, the updated W is obtained. According to the previously calculated y o , where y o Forward propagation data prediction value, Y {1} As data input value Y o , and then calculate the error between the network input and output, and use the error as the loss function. The loss update formula is: Loss=y o -Y o Combine the weights and biases updated by the gradient and chain method for back propagation to obtain Net(1), and then continue to execute the above training process according to the set number of iterations run to obtain Net(run), which is the final proxy model Net.
5. The method for optimizing a neural network multi-performance architecture at low cost according to claim 4, characterized in that: After the training of the proxy model is completed, the multi-performance target solution of the deep convolutional neural network can achieve the approximate equivalence of the target value by using the trained PA-Dropout for one quick calculation. During the test, the closed Dropout neural network is directly fitted without additional parameter adjustment and complex calculation, and the target value y is quickly obtained. obj , which not only improves computing efficiency but also reduces computing resource consumption, and strongly supports the rapid deployment and efficient operation of deep convolutional neural networks in practice.
6. The method for optimizing a neural network multi-performance architecture at low cost according to claim 2, characterized in that: In the multi-performance architecture hyperparameter search module, the dual-drive interactive dynamic model management strategy DDIMMS provides reliable candidate solutions for real evaluation by taking into account the comprehensive evaluation of the diversity of individual convergence of high-dimensional hyperparameter combinations and adaptive weighted calculation, thereby improving the PA-Dropout update and fitting efficiency, and ultimately achieving fast and efficient recommendation of multi-objective performance architectures; the details are as follows: First, the initial population P and the population P after PA-Dropout optimization c , perform normalization processing; then calculate P and P respectively c The minimum distance sum from the individual to the nearest neighbor is recorded as Crowd, where the calculation formula of Crowd is as follows: Crowd=∑ x∈G min{d(x,y)=||x-y||2,|y∈G,y≠x} In the formula, G represents the population, x and y are individuals in the set G, and d(x,y) represents the distance between individuals x and y. Through this formula, the result is the value of Crowd. CrowdOld is calculated based on the population P, and CrowdOld is calculated based on the population P. c CrowdNew is calculated to reflect the density and diversity of individual distribution in the population before and after optimization. Then, P is calculated using the following formula using the Euclidean distance as the metric: c the spatial distance between individuals; Assume that x and y are P c The i-th and j-th individuals in , and k represents the k-th decision variable, Dist_D(i,j) is the diversity evaluation between i and j individuals; in order to comprehensively evaluate P c The performance of individuals in the experiment is introduced, and the calculation of the convergence index is introduced. This index measures the performance of P c The closeness between the non-dominated solution set in and the ideal point z* is designed to capture the individual's ability to approach the optimal solution in the target space; The calculation formula is as follows: Where x represents the i-th individual, k represents the k-th decision variable, and Dist_C(i) is the convergence evaluation of the i-th individual. In order to avoid the calculation error caused by the measurement exceeding the data range, the Dist_D(i,j) and Dist_C(i) are normalized to obtain the normalized Dist_D1(i,j) and Dist_C1(i). Then, the CrowdOld and CrowdNew values are compared. If CrowdOld≤CrowdNew, it indicates that the individual diversity of the optimized population is better. At this time, we should focus on improving convergence, with convergence as the main structure and diversity as the auxiliary structure. First, calculate the attention weight of population diversity. The D matrix is an N×N matrix, D ij represents the diversity attention weight assigned to the jth individual by the i-th individual, D ij Calculated as follows: Calculate the convergence of the i-th individual as the main comprehensive index and update it through the following formula: in Take the inverse of Dist_D(i,j); then select C i The individual with the smallest value is evaluated truthfully; If CrowdOld>CrowdNew, it means that the population diversity is insufficient; At this time, diversity should be set as the main structure and convergence as the auxiliary structure; first calculate the population convergence attention weight and create an N×1 C matrix, C i1 represents the convergent attention weight of the i-th individual, calculated as follows: The diversity of the i-th individual is calculated as the main comprehensive index through the following formula: in Take the reciprocal of Dist_C1(i) and finally select D i The largest individual makes a true evaluation.
Citation Information
Cited By
Intelligent film and television content recommendation method and system based on artificial intelligence
CN120763360A
Intelligent slicing and conveying system for shower room glass products and control method
CN121225298A