Multi-branch network architecture search system and method

By parameterizing the overall network architecture and using hyperparameter optimization technology to search for multi-branch network architecture, the problem of time and computing resources required for AI to design network architecture in emerging application fields is solved, and more efficient model performance and parameter optimization is achieved.

CN119990227APending Publication Date: 2025-05-13IND TECH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410020374.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-10
Filing Date
2024-01-05
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When applying AI in emerging application fields, designing suitable network architectures, data amplification strategies and hyperparameter adjustments requires a lot of time and computing resources, which increases the threshold for small and medium-sized enterprises to introduce AI technology.

Method used

A multi-branch network architecture search system and method are proposed, which parameterizes the overall network architecture, uses blocks as basic elements, and uses hyperparameter optimization technology to optimize multi-branch network architecture search.

Benefits of technology

The optimal multi-branch architecture model obtained through this method has a relatively improved inference accuracy of the test data set by 6.7%, the model parameters are reduced by 79.5%, the inference speed is relatively improved by 37%, and the effect on the test data set is relatively improved by 13%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990227A_ABST
    Figure CN119990227A_ABST
Patent Text Reader

Abstract

A multi-branch network architecture search system and method suitable for an electronic device including a processor, the method comprising: obtaining a training data set including a plurality of input data, obtaining block design elements of a plurality of blocks constituting an architecture of a neural network, the blocks being used for feature extraction of the input data to generate output data, for each hyper-parameter of the neural network, at least one hyper-parameter setting value is obtained, the training data set, the block design element and the at least one hyper-parameter setting value are input to a hyper-parameter optimization algorithm to generate a hyper-parameter combination, and the hyper-parameter combination comprises the hyper-parameter setting value corresponding to each hyper-parameter. And operating the neural network based on the hyper-parameter combination, inputting the test data set to evaluate the model efficiency of the neural network, and outputting the hyper-parameter combination when the model efficiency reaches a threshold value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to neural networks and network architecture search, and in particular to a multi-branch network architecture search system and method. Background Art

[0002] Artificial Intelligence (AI) has been widely used in various fields.

[0003] However, when applying AI in emerging application areas (such as industrial defect detection), even professional AI experts need to spend a lot of time to design network architectures, data augmentation strategies, and related hyperparameter adjustments suitable for this field. The whole process is very time-consuming and laborious. On the other hand, existing automatic machine learning network architecture search technologies require a lot of computing resources and still take a long time to complete, which raises the threshold for small and medium-sized enterprises to introduce AI technology. Summary of the invention

[0004] In view of this, the present invention proposes a multi-branch network architecture search system and method, which uses blocks as basic elements, parameterizes the overall network architecture, and uses hyperparameter optimization (HPO) technology to perform optimized multi-branch network architecture search. The overall network architecture search process adopts a multi-branch architecture search solution.

[0005] A multi-branch network architecture search method according to an embodiment of the present invention is applicable to an electronic device including a processor, and includes the following steps: obtaining a training data set, wherein the training data set includes multiple input data; obtaining block design elements of multiple blocks, wherein the multiple blocks constitute the architecture of a neural network, and the multiple blocks are used to perform feature extraction on the multiple input data to generate output data; for each hyperparameter of the neural network, obtaining at least one hyperparameter setting value; inputting the training data set, the block design elements and at least one hyperparameter setting value into a hyperparameter optimization algorithm to generate a hyperparameter combination, wherein the hyperparameter combination includes the hyperparameter setting value corresponding to each hyperparameter; running the neural network based on the hyperparameter combination, and inputting a test data set to evaluate the model performance of the neural network; and when the model performance reaches a threshold, outputting the hyperparameter combination.

[0006] A multi-branch network architecture search system according to an embodiment of the present invention is suitable for an electronic device including a processor, comprising: an input module, a computing module and a neural network module. The input module is used to obtain a training data set, block design elements of a plurality of blocks, obtain at least one hyperparameter setting value for each hyperparameter of the neural network, and obtain a test data set, wherein the training data set includes a plurality of input data, the plurality of blocks constitute the architecture of the neural network, and the plurality of blocks are used to perform feature extraction on one of the plurality of input data to generate output data. The computing module is communicatively connected to the input module, and the computing module is used to execute a hyperparameter optimization algorithm based on the training data set, the block design elements and at least one hyperparameter setting value to generate a hyperparameter combination, and the hyperparameter combination includes one of at least one hyperparameter setting value corresponding to each hyperparameter. The neural network module is communicatively connected to the input module and the computing module, and the neural network module is used to run a neural network based on a hyperparameter combination, and input a test data set to evaluate the model performance of the neural network, wherein the computing module is further used to output the hyperparameter combination when the model performance reaches a threshold.

[0007] The above description of the content of the present invention and the following description of the implementation modes are used to demonstrate and explain the principles of the present invention, and to provide further explanation for the scope of the patent application of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 is a block diagram of a multi-branch network architecture search system according to an embodiment of the present invention;

[0009] Figure 2 is an example of a multi-branch network architecture; and

[0010] Figure 3 It is a flowchart of a multi-branch network architecture search method according to an embodiment of the present invention.

[0011] 10-Multi-branch network architecture search system;

[0012] 1- Input module;

[0013] 3- Operation module;

[0014] 5-Neural network module;

[0015] S-trunk;

[0016] B-branch;

[0017] P1-P7-Steps. DETAILED DESCRIPTION

[0018] The detailed features and advantages of the present invention are described in detail in the following embodiments, and the contents are sufficient to enable those skilled in the art to understand the technical content of the present invention and implement it accordingly, and according to the contents invented in this specification, the scope of the patent application and the drawings, those skilled in the art can easily understand the relevant purposes and advantages of the present invention. The following examples are to further illustrate the viewpoints of the present invention in detail, but cannot limit the scope of the present invention in any viewpoint.

[0019] Figure 1 FIG. 1 is a block diagram of a multi-branch network architecture search system according to an embodiment of the present invention, wherein the system is applicable to an electronic device including a processor. Figure 1 As shown, the multi-branch network architecture search system 10 includes an input module 1, a calculation module 3 and a neural network module 5.

[0020] The input module 1 is used to obtain a training data set, a validation data set and a test data set of the neural network. Each of these data sets includes a plurality of input data. When the neural network is applied to semiconductor defect detection, the input data includes images of normal semiconductors and images of semiconductors with various defects. Table 1 below is an example of experimental data used in the present invention, and the numbers in the table represent the number of input data (images).

[0021] Table 1. Example of a semiconductor defect dataset

[0022] Defect Type Training Dataset Validation Dataset Test Dataset Total Probe offset 3473 212 422 4107 normal 213457 12462 25067 250986 Dirty 70249 4145 8259 82753 Process defects 10625 640 1271 12536 Particle foreign matter 42909 2513 5062 50484 foreign body 49830 2974 5812 58616 Discoloration 1595 94 187 1876 Total 392238 23040 46080 461358

[0023] The input module 1 is further used to obtain block design elements of multiple blocks. The elements that may be used in the standard block and the dimension reduction block include Conv1×1, Conv3×3, Conv5×5, MaxPool 3×3, AvgPool 3×3, SepConv3×3, SepConv5×5, DilConv3×3, DilConv5×5, etc., but are not limited to this. These blocks constitute the architecture of the neural network. Figure 2 is an example of a multi-branch neural network architecture. Figure 2As shown, this multi-branch neural network includes a trunk S (stem) and three branches B (branch). Except for Conv7×7, stride2, MaxPool 2×2 and Softmax, the rest are composed of blocks. The block is used to extract features from the feature map input to the block to generate output data. In an embodiment, these blocks can be divided into a normal block and a reduction block. The standard block is used to retain the dimension of the input feature map, and the reduction block is used to reduce the dimension of the input feature map. In an embodiment, the generation method of these blocks can be designed by AI engineers themselves, or using residual blocks, or using NASNet, or using the best blocks found by a gradient-based search method, such as Pruning-Based Differentiable Architecture Search (PR-DARTS).

[0024] The neural network includes multiple hyperparameters. In an embodiment, these hyperparameters include at least one of the following: the number of blocks in the trunk S, the number of channels in the trunk S, the number of branches B, the number of blocks in branch B, the number of channels in branch B, and the position of the dimension reduction block in branch B. The input module 1 is further used to obtain at least one hyperparameter setting value for each hyperparameter. All hyperparameter setting values ​​constitute a search space. Table 2 below is an example of the search space, which can generate approximately 1.58 billion different model architectures.

[0025] Table 2. Example of search space

[0026]

[0027] Regarding the position of the dimensionality reduction block, for example, if there are 11 blocks of the branch, numbered from block 0 to block 10, when the position of dimensionality reduction block 1 is 0.1, it means that the position of dimensionality reduction block 1 is the position of block 1 (11*0.1=1.1, rounded to 1).

[0028] In an embodiment, the number of branches is at least 1. Figure 2 For example, the number of branches is set to 3. The present invention does not limit the number of dimension reduction blocks used. Figure 2 For example, generally speaking, in a neural network architecture, five dimensionality reduction operations are usually performed, the trunk S performs three of them, and the branch B performs the remaining two, but the present invention is not limited to the above values.

[0029] Please refer to Figure 1. The operation module 3 is communicatively connected to the input module 1. The operation module 3 is used to execute a hyperparameter optimization algorithm to generate a hyperparameter combination based on a search space consisting of a training data set, block design elements, and hyperparameter setting values. In an embodiment, the hyperparameter optimization algorithm includes at least one of the following: Tree-Structured Parzen Estimation, Bayesian Optimization, Grid Search, Random Optimization, Sequential Model-Based Algorithm Configuration, and Metis. The hyperparameter combination includes a hyperparameter setting value corresponding to each hyperparameter type. Taking Table 2 as an example, an example of a hyperparameter combination may be {2, 16, 1, 8, 16, 0.1, 0.6}.

[0030] Please refer to Figure 1 . The neural network module 5 is communicatively connected to the input module 1 and the operation module 3. The neural network module 5 is used to run the neural network based on the hyperparameter combination, and input the test data set to evaluate the model performance of the neural network. The operation module 3 is further used to output the hyperparameter combination when the model performance reaches a threshold. In an embodiment, the model performance includes at least one of the following: model accuracy, false negative rate (False Negative Rate, FNR), true negative rate (True Negative Rate, TNR), parameter size, and inference speed.

[0031] In an embodiment, the actual operation of the input module 1, the operation module 3 and the neural network module 5 is, for example, run on an electronic device in the form of program code or software. The electronic device may adopt at least one of the following examples: a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller (MCU), an application processor (AP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a digital signal processor (DSP), a system-on-a-chip (SOC), and a deep learning accelerator. However, the present invention is not limited to these examples.

[0032] Figure 3 FIG. 1 is a flow chart of a multi-branch network architecture search method according to an embodiment of the present invention, and the method is applicable to an electronic device including a processor. Figure 3 As shown, the multi-branch network architecture search method includes steps P1 to P6.

[0033] In step P1, input module 1 obtains a training data set. In step P2, input module 1 obtains block design elements of multiple blocks. In step P3, input module 1 obtains at least one hyperparameter setting value of each hyperparameter. In step P4, operation module 3 executes a hyperparameter optimization algorithm based on the training data set, block design elements and hyperparameter setting values ​​to generate a hyperparameter combination. In step P5, neural network module 5 runs a candidate neural network based on the hyperparameter combination generated in step P4, and inputs a test data set to this candidate neural network to evaluate the model performance of the candidate neural network.

[0034] In step P6, the operation module 3 determines whether the model performance reaches the threshold. If it is determined to be yes, the hyperparameter combination used by the current candidate neural network is output. If it is determined to be no, it returns to step P4, and the model performance that does not reach the threshold is fed back to the hyperparameter optimization algorithm, and the hyperparameter optimization algorithm generates another hyperparameter combination, and the process from step P4 to step P6 is repeated until the model performance of the candidate neural network reaches the threshold.

[0035] In summary, the multi-branch network architecture search system and method proposed in the present invention parameterizes the overall network architecture, uses blocks as basic elements, and applies a hyperparameter optimization algorithm to perform overall network architecture search.

[0036] Most of the existing network architecture search technologies do not consider searching the overall network architecture, especially the multi-branch network architecture. The present invention focuses on the neural network with a multi-branch architecture. Through experiments, it is known that the best multi-branch architecture model obtained by the method of the present invention has a 6.7% relative improvement in inference accuracy in the test data set compared with the manually designed neural network model, and the number of model parameters is reduced by 79.5%, and the inference speed is relatively improved by 37%. In addition, compared with the non-multi-branch network architecture search technology, the model searched by the present invention has a 13% relative improvement in the test data set.

Claims

1. A multi-branch neural network architecture search method, applicable to an electronic device including a processor, comprising: Obtain a training data set, wherein the training data set includes a plurality of input data; Obtaining block design elements of a plurality of blocks, wherein the blocks constitute a neural network architecture, and the blocks are used to perform feature extraction on the input data to generate output data; For each of a plurality of hyperparameters of the neural network, obtaining at least one hyperparameter setting value; Inputting the training data set, the block design element and the at least one hyperparameter setting value into a hyperparameter optimization algorithm to generate a hyperparameter combination, wherein the hyperparameter combination includes one of the at least one hyperparameter setting value corresponding to each of the plurality of hyperparameters; Running the neural network based on the hyperparameter combination and inputting a test data set to evaluate the model performance of the neural network; and When the model performance reaches the threshold, the hyperparameter combination is output.

2. The multi-branch neural network architecture search method according to claim 1, wherein the blocks include a plurality of standard blocks and a plurality of dimensionality reduction blocks, each of the standard blocks is used to retain the dimension of one of the input data, and each of the dimensionality reduction blocks is used to reduce the dimension of one of the input data.

3. The multi-branch neural network architecture search method according to claim 2, wherein the architecture of the neural network includes a trunk and branches, and the multiple hyperparameters of the neural network include at least one of the following: the number of blocks in the trunk, the number of channels in the trunk, the number of branches, the number of blocks in the branches, the number of channels in the branches, and the positions of the dimensionality reduction blocks in the branches.

4. The multi-branch neural network architecture search method according to claim 3, wherein the number of branches is at least 1.

5. The multi-branch neural network architecture search method according to claim 1, wherein the hyperparameter optimization algorithm includes at least one of the following: Tree-Structured Parzen Estimation, Bayesian Optimization, Grid Search, Random Optimization, Sequential Model-Based Algorithm Configuration, and Metis.

6. The multi-branch neural network architecture search method according to claim 1, wherein the model performance includes at least one of the following: model accuracy, false negative rate (False Negative Rate, FNR), true negative rate (True Negative Rate, TNR), parameter size and inference speed.

7. A multi-branch neural network architecture search system is applicable to an electronic device including a processor, comprising: An input module, configured to obtain a training data set, block design elements of a plurality of blocks, obtain at least one hyperparameter setting value for each of a plurality of hyperparameters of a neural network, and obtain a test data set, wherein the training data set includes a plurality of input data, the blocks constitute the architecture of the neural network, and the blocks are used to perform feature extraction on one of the input data to generate output data; a computing module, communicatively connected to the input module, for executing a hyperparameter optimization algorithm according to the training data set, the block design element, and the at least one hyperparameter setting value to generate a hyperparameter combination, the hyperparameter combination including one of the at least one hyperparameter setting value corresponding to each of the hyperparameters; and A neural network module is communicatively connected to the input module and the operation module, and is used to run the neural network based on the hyperparameter combination and input the test data set to evaluate the model performance of the neural network, wherein the operation module is further used to output the hyperparameter combination when the model performance reaches a threshold.

8. A multi-branch neural network architecture search system according to claim 7, wherein the blocks include multiple standard blocks and multiple dimensionality reduction blocks, each of the standard blocks is used to retain the dimension of one of the input data, and each of the dimensionality reduction blocks is used to reduce the dimension of one of the input data.

9. The multi-branch neural network architecture search system according to claim 8, wherein the architecture of the neural network includes a trunk and branches, and the multiple hyperparameters of the neural network include at least one of the following: the number of blocks in the trunk, the number of channels in the trunk, the number of branches, the number of blocks in the branches, the number of channels in the branches, and the positions of the dimensionality reduction blocks in the branches.

10. The multi-branch neural network architecture search system according to claim 9, wherein the number of branches is at least 1.

11. The multi-branch neural network architecture search system according to claim 7, wherein the hyperparameter optimization algorithm comprises at least one of the following: tree Parzen estimator, Bayesian optimization, grid search, random optimization, sequence model algorithm configuration and Metis.

12. The multi-branch neural network architecture search system according to claim 7, wherein the model performance includes at least one of the following: model accuracy, false negative rate, true negative rate, parameter size, and inference speed.