A centrally symmetric cross convolutional neural network architecture search method and chip

Through the central symmetric cross-convolutional neural network architecture search method, the problem of accuracy loss of small and medium-sized networks in the existing technology is solved, and high-precision neural network architecture search is realized, which shortens the R&D cycle and reduces enterprise risks.

CN116306845BActive Publication Date: 2025-05-02XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310103316.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-10
Publication Date
2025-05-02
Estimated Expiration
2043-02-10

AI Technical Summary

Technical Problem

When searching for neural network architectures in the prior art, the supernet training method based on progressive shrinkage has serious loss of accuracy in small-size networks, resulting in failure of network search methods, extending product factory time and increasing enterprise costs.

Method used

The central symmetric cross-convolution neural network architecture search method is adopted, and the maximum value of neural network parameter quantity and feature map memory is determined according to the hardware conditions of the target chip, and the hypernet training matrix is ​​randomly generated, and the reference network is extended through the central symmetric selection rules, and iterative training is carried out to obtain the optimal network architecture.

Benefits of technology

The small-size network accuracy during the fine-grained training process of supernet is improved, preventing accuracy losses from being passed to subsequent network searches, shortening the R&D cycle and reducing corporate risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116306845B_ABST
    Figure CN116306845B_ABST
Patent Text Reader

Abstract

The present invention provides a method and chip for searching the architecture of a centrally symmetric cross-type convolutional neural network. For network architecture parameters that require fine-grained search, a centrally symmetric selection rule is first designed based on the parity of the number of parameters, and objects are selected from the supernet training matrix to expand the reference network. The number of rounds of cross-training in each stage is determined based on the parameter size of the expanded network, and a supernet containing all network architecture parameter combinations is obtained through training. Finally, the maximum number of neural network parameters and the peak memory of the feature map allowed by the target chip are used as constraints, and the search is iterated in the search space defined by the supernet to obtain the subnet structure with the highest accuracy for deployment. The present invention can perform fine-grained training on the supernet through centrally symmetric cross-training, improve the accuracy of small-size networks in the process of fine-grained training of the supernet, and prevent the loss of accuracy from being transmitted to the subsequent network search process, causing the network search method to fail, thereby reducing enterprise risks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of neural network design for chip deployment, and specifically relates to a centrally symmetric cross convolutional neural network architecture search method and chip. Background Art

[0002] With the rapid development of science and technology, the performance and cost of chips have become an indispensable central component of various smart devices. The higher the chip cost, the better the chip performance, and vice versa. In the field of R&D, technicians hope that the chip performance is as high as possible, and in enterprise production, enterprises hope that the product cost will not be too high, so the right chip is particularly important.

[0003] In some cases, the chip product design and hardware production have been completed, but it is found that the actual performance of the product cannot meet the requirements. At this time, the cost of changing the chip or hardware design is too high, and the algorithm memory needs to be compressed through software or algorithms. Most algorithm applications involve neural network models. In the actual process, the manual design of convolutional neural networks to meet project requirements is very dependent on the designer's experience. It often takes a lot of time to repeatedly experiment and constantly adjust the parameters of the network architecture. The final result may not be the optimal solution. In recent years, the neural network design method based on neural network architecture search has developed rapidly, and the requirements for computing power have continued to decline. Nowadays, this method can significantly reduce the manpower and time costs while obtaining a high-precision network structure. In the field of micro machine learning, in addition to caring about the highest accuracy that the network structure can achieve, people also pay special attention to indicators such as the total number of model parameters, the peak memory occupied by the input and output feature maps, the amount of multiplication and addition calculations, and the delay time.

[0004] For micro-machine learning, the existing excellent network architecture design example is Once for All (Han Cai et al. "Once for All: Train One Network and Specialize it for Efficient Deployment" (2020).), which mainly includes two parts: supernet training and subnet search. First, a supernet containing all possible network structures is trained by progressive shrinkage, and then the subnet structure is obtained by randomly sampling the supernet, and the optimal network structure under the specified constraints is obtained through evolutionary search. However, this work only sets 3 optional values ​​for each network architecture parameter. If more fine-grained training and search are performed, the effect of the supernet training method based on progressive shrinkage will deteriorate. The supernet training results show that the accuracy loss of small-size networks is serious, which will affect the subsequent search process and even cause the network search method to fail. This requires extending the time it takes for products to leave the factory, which will cause a sudden increase in costs for enterprises. What's worse, the market share will decline and the enterprise will be dragged down. Summary of the invention

[0005] In order to solve the above problems existing in the prior art, the present invention provides a central symmetric cross convolutional neural network architecture search method and chip. The technical problem to be solved by the present invention is achieved through the following technical solutions:

[0006] The present invention provides a method for searching a centrally symmetric cross convolutional neural network architecture, comprising:

[0007] Step 1: Select the target chip according to the design requirements and determine the hardware conditions of the target chip;

[0008] Step 2, determining the maximum value of the neural network parameter and the maximum value of the input and output feature map peak memory according to the hardware conditions of the target chip;

[0009] Step 3, randomly generate a supernet training matrix according to the preset supernet training granularity;

[0010] The dimension of the supernet training matrix is ​​related to the fine granularity, the supernet training matrix is ​​a row or column matrix, and the elements in the matrix are in an arithmetic progression;

[0011] Step 4, selecting a reference network for supernet training and confirming the architecture parameters of the reference network;

[0012] Step 5, selecting a scale parameter item for each training network in the supernet training matrix according to whether the number of elements in the supernet training matrix is ​​even or odd;

[0013] Step 6, multiplying the scale parameter item by the architecture parameter of the reference network during each training, determining the multiplication result as the extension number of the architecture parameter of the reference network, and extending the reference network according to the extension number to obtain an extended network, and iteratively training the extended network obtained each time;

[0014] Step 7, with the maximum value of the neural network parameter amount and the maximum value of the input and output feature map peak memory as constraints, and the highest accuracy of the final network as the goal, the architecture parameters of each extended network after training are used as search objects, and the search is performed, and the search object with the highest accuracy is used as the architecture parameters of each layer of the final network to obtain the final network;

[0015] Step 8: deploy the final network on the target chip.

[0016] Beneficial effects of the present invention:

[0017] The present invention provides a method for searching the architecture of a centrally symmetric cross-type convolutional neural network, which determines the maximum value of the neural network parameter amount and the peak memory occupied by the network input and output feature map according to the hardware characteristics of the target chip. For the architecture parameters of the network that needs to be searched in a fine-grained manner, the centrally symmetric selection rule is first set according to the parity of the number of parameters, the object is selected from the supernet training matrix to expand the reference network, and then the number of rounds of cross-training in each stage is determined according to the parameter size of the extended network, and a supernet containing all network architecture parameter combinations is obtained by training. Finally, with the maximum neural network parameter amount and the peak memory of the feature map allowed by the target chip as constraints, an iterative search is performed in the search space defined by the supernet to obtain the subnet structure with the highest accuracy, and the final network is obtained for deployment. The present invention can perform fine-grained training on the supernet by a centrally symmetric cross-training method. Compared with the existing training method based on a progressive contraction method, the accuracy of the small-size network in the fine-grained training process of the supernet is improved, and the accuracy loss is prevented from being transmitted to the subsequent network search process, resulting in the failure of the network search method, thereby achieving the purpose of shortening the R&D cycle and reducing the risk of the enterprise.

[0018] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flow chart of a method for searching a centrally symmetric cross convolutional neural network architecture provided by the present invention;

[0020] Figure 2 It is a schematic diagram of the centrosymmetric cross training method provided by the present invention. DETAILED DESCRIPTION

[0021] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0022] like Figure 1 As shown, the present invention provides a method for searching a centrally symmetric cross convolutional neural network architecture, including:

[0023] Step 1: Select the target chip according to the design requirements and determine the hardware conditions of the target chip;

[0024] The design requirement can be the product selected for the company project, and the chip that runs the model is selected as the target chip from all the chips contained in the product. This process can be adaptively adjusted according to the needs of the enterprise.

[0025] Step 2, determining the maximum value of the neural network parameter and the maximum value of the input and output feature map peak memory according to the hardware conditions of the target chip;

[0026] The hardware conditions may be the maximum operating capability of the chip, storage capacity, and other hardware parameters that may limit the performance of the chip.

[0027] Step 3, randomly generate a supernet training matrix according to the preset supernet training granularity;

[0028] The dimension of the supernet training matrix is ​​related to the fine granularity, the supernet training matrix is ​​a row or column matrix, and the elements in the matrix are in an arithmetic progression;

[0029] In a specific embodiment of the present invention, step 3 comprises:

[0030] Step 31, determining the dimension of generating a supernet training network according to the preset fine granularity of supernet training;

[0031] Step 32, generating each element in the supernet training matrix by a one-time generation method or a multiple-time generation method according to the equal increase or decrease of the elements, to obtain the supernet training matrix;

[0032] Among them, the multiple generation method is to add one element to the number of elements generated in the previous time until the supernet training matrix reaches the dimension.

[0033] Step 4, selecting a reference network for supernet training and confirming the architecture parameters of the reference network;

[0034] The architecture parameters include at least one of the parameters such as convolution kernel size, number of network layers, number of channels in each layer, etc.;

[0035] The present invention can perform a fine-grained search on the number of channels in each layer, and of course can also perform a fine-grained search on parameters such as the convolution kernel size and the number of network layers, or simultaneously combine two or more network architecture parameters to perform a fine-grained search.

[0036] Step 5, selecting a scale parameter item for each training network in the supernet training matrix according to whether the number of elements in the supernet training matrix is ​​even or odd;

[0037] According to the existing supernet training specifications, the optional architecture parameters of each layer of the network are the baseline network architecture parameters multiplied by the network scale parameter s. The scale parameter will be used as a constraint in the subsequent search process to select the optimal network. Among them, the mathematical relationship between the neural network parameter quantity P and the input and output feature map peak memory M and scale parameter s is expressed as:

[0038] P = f[Net(s)]

[0039] M = g[Net(s)]

[0040] Where f[·] represents the calculation of the weight parameters of the convolution kernel of each layer of the network and outputs the summed result, g[·] represents the calculation of the sum of the input and output feature maps of each layer of the network and outputs the maximum value, and Net represents the benchmark network.

[0041] During the random sampling of subnetworks, the architectural parameters of each layer of the subnetwork can be set to any size in the supernet training matrix. In the multiple generation method, each training adds a parameter to the training matrix, so that the supernet training matrix generated this time has one more scale parameter than the previous supernet training matrix.

[0042] In a specific embodiment of the present invention, step 5 comprises:

[0043] Step 51, determining whether the number of elements of the supernet training matrix is ​​an even number or an odd number;

[0044] Step 52, determining a selection rule for whether the number of elements N of the supernet training matrix is ​​an even number or an odd number, and selecting a scale parameter item for each training network in the generated supernet training matrix according to the selection rule,

[0045] If N is an even number, k=0,1,2,3..., 1≤i≤N, then:

[0046]

[0047] If N is an odd number, k=0,1,2,3…, 1≤i≤N, then:

[0048]

[0049] N represents the N scale parameter items in the supernet training matrix, which are arranged in descending order of size as [x1,x2,…x N ], y(i) represents the scale parameter item selected during the i-th training.

[0050] Step 6, multiplying the scale parameter item by the architecture parameter of the reference network during each training, determining the multiplication result as the extension number of the architecture parameter of the reference network, and extending the reference network according to the extension number to obtain an extended network, and iteratively training the extended network obtained each time;

[0051] In a specific embodiment of the present invention, step 6 includes:

[0052] Step 61, multiplying the scale parameter item by the architecture parameter of the reference network during each training, and determining the multiplication result as the expansion number of the architecture parameter of the reference network;

[0053] Step 62, expanding the reference network according to the expansion number to obtain an extended network;

[0054] Step 63, determining the number of training rounds for iterative training according to the size of the extended network;

[0055] When training the largest-sized network, more training rounds are required. Compared with the sub-network with fewer architectural parameters, the sub-network with larger architectural parameters has a higher similarity in weight parameters with the largest-sized network. Therefore, after inheriting the weights from the largest-sized network, only a small number of training rounds are required to converge.

[0056] When the large-size network architecture parameters are added to the supernet training matrix, a smaller number of training rounds is set, expressed as training round Epoch = E small ; When the small-size network architecture parameters are added to the supernet training matrix, a larger number of training rounds is set, expressed as training round Epoch = E large . E large The value of is usually the same as the number of training rounds for the largest network size. small The value of is related to the similarity between the current training network and the maximum size network.

[0057] Define R(j) as the supernet training matrix [x1,x2,…x N The similarity between the extended network corresponding to the jth scale parameter in ] and the maximum size network is:

[0058]

[0059]

[0060] Arrange the scale parameters of the supernet training matrix in descending order to obtain [x1, x2, …x N ], the number of training rounds Epochs during the extended network training corresponding to the jth scale parameter in the matrix is:

[0061]

[0062] Among them, E large The corresponding extended network size is smaller than E small The corresponding expanded network scale.

[0063] Step 64, iteratively train each extended network using any data set for the target chip until the number of training rounds corresponding to the extended network is reached, thereby obtaining the extended network after training.

[0064] Step 7, with the maximum value of the neural network parameter amount and the maximum value of the input and output feature map peak memory as constraints, and the highest accuracy of the final network as the goal, the architecture parameters of each extended network after training are used as search objects, and the search is performed, and the search object with the highest accuracy is used as the architecture parameters of each layer of the final network to obtain the final network;

[0065] Set the constraints of the neural network parameters and the peak memory of the feature map, take the highest accuracy as the optimization goal, and search for the best network that meets the requirements as the final network. The constraints are:

[0066] maxAccuracy(s)

[0067]

[0068] Among them, P max Represents the maximum value of the neural network parameters, M max Indicates the maximum value of the peak memory of the input and output feature maps.

[0069] The search process executes a search algorithm based on population evolution, and the basic parameters of the search algorithm are set as follows: the number of individuals in the population P = 60, the total number of generations of genetic evolution N = 30, the proportion of individuals in each generation that serve as the mother of the next generation r = 0.25, the proportion of individuals that may mutate in each generation is set to 0.5, and the probability of individual mutation is set to 0.1. After searching, the network structure with the highest accuracy under the constraints of parameter quantity and peak memory is obtained.

[0070] Step 8: deploy the final network on the target chip.

[0071] The present invention provides a chip, which is deployed with the final network obtained by the centrosymmetric cross convolutional neural network architecture search method of the present invention.

[0072] The effect of the present invention can be further illustrated by the following experimental data.

[0073] During the supernet training process, other parameters are kept unchanged, and the results of training the supernet using the central symmetric cross training method proposed by the present invention and the existing progressive shrinkage method are shown in Table 1.

[0074] Table 1 Comparison of the supernet training effects of the present invention and the existing methods

[0075] Network scale parameters x1.5 x1.25 x1.0 x0.75 x0.5 Gradual contraction 71% 69.2% 65.6% 60.3% 48.1% The present invention 71.1% 69.4% 65.9% 61.3% 51.9%

[0076] It can be seen from Table 1 that:

[0077] The method of the present invention is used to perform super-network fine-grained training. On the basis of the large-size network accuracy being equal to the training result of the progressive shrinkage method, the accuracy of the small-size network is improved by 1%-4%. Overall, the network accuracy is higher.

[0078] The results in Table 1 fully demonstrate that, under the condition of fine-grained training, the supernet training effect using the method of the present invention is better, which not only maintains the accuracy of large-size networks, but also improves the accuracy of small-size networks, and avoids the accuracy loss in the supernet training process being transmitted to the subsequent network search process, causing the network search method to fail.

[0079] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality of components or steps.

[0080] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A method for searching a centrally symmetric cross convolutional neural network architecture, characterized in that: include: Step 1: Select the target chip according to the design requirements and determine the hardware conditions of the target chip; Step 2, determining the maximum value of the neural network parameter and the maximum value of the input and output feature map peak memory according to the hardware conditions of the target chip; Step 3, randomly generate a supernet training matrix according to the preset supernet training granularity; The dimension of the supernet training matrix is ​​related to the fine granularity, the supernet training matrix is ​​a row or column matrix, and the elements in the matrix are in an arithmetic progression; Step 4, selecting a reference network for supernet training and confirming the architecture parameters of the reference network; Step 5, selecting a scale parameter item for each training network in the supernet training matrix according to whether the number of elements in the supernet training matrix is ​​even or odd; Step 6, multiplying the scale parameter item by the architecture parameter of the reference network during each training, determining the multiplication result as the extension number of the architecture parameter of the reference network, and extending the reference network according to the extension number to obtain an extended network, and iteratively training the extended network obtained each time; Step 7, with the maximum value of the neural network parameter amount and the maximum value of the input and output feature map peak memory as constraints, and the highest accuracy of the final network as the goal, the architecture parameters of each extended network after training are used as search objects, and the search is performed, and the search object with the highest accuracy is used as the architecture parameters of each layer of the final network to obtain the final network; Step 8: deploy the final network on the target chip.

2. The method for searching a centrally symmetric cross convolutional neural network architecture according to claim 1, characterized in that: Step 3 includes: Step 31, determining the dimension of generating a supernet training network according to the preset fine granularity of supernet training; Step 32, generating each element in the supernet training matrix by a one-time generation method or a multiple-time generation method according to the equal increase or decrease of the elements, to obtain the supernet training matrix; Among them, the multiple generation method is to add one element to the number of elements generated in the previous time until the supernet training matrix reaches the dimension.

3. The method for searching a centrally symmetric cross convolutional neural network architecture according to claim 2, characterized in that: Step 5 includes: Step 51, determining whether the number of elements of the supernet training matrix is ​​an even number or an odd number; Step 52, when determining that the number of elements N of the supernet training matrix is ​​an even number or an odd number, the centrosymmetric selection rule is used to select the scale parameter item for each training network in the generated supernet training matrix according to the centrosymmetric selection rule. If N is an even number, k=0,1,2,3…, 1≤i≤N, then: If N is an odd number, k=0,1,2,3…, 1≤i≤N, then: N represents the N scale parameter items in the supernet training matrix, which are arranged in descending order of size as [x1,x2,…x N ], y(i) represents the scale parameter item selected during the i-th training.

4. The method for searching a centrally symmetric cross convolutional neural network architecture according to claim 2, characterized in that: Step 6 includes: Step 61, multiplying the scale parameter item by the architecture parameter of the reference network during each training, and determining the multiplication result as the expansion number of the architecture parameter of the reference network; Step 62, expanding the reference network according to the expansion number to obtain an extended network; Step 63, determining the number of training rounds for iterative training according to the size of the extended network; Step 64, iteratively train each extended network using any data set for the target chip until the number of training rounds corresponding to the extended network is reached, thereby obtaining the extended network after training.

5. The method for searching a centrally symmetric cross convolutional neural network architecture according to claim 4, characterized in that: Step 63 includes: R(j) is the supernet training matrix [x1,x2,…x N The similarity between the extended network corresponding to the jth scale parameter in ] and the maximum size network is: Arrange the scale parameters of the supernet training matrix in descending order to obtain [x1, x2, …x N ], the number of training rounds Epochs during the extended network training corresponding to the jth scale parameter in the matrix is: Among them, E large The corresponding extended network size is smaller than E small The corresponding expanded network scale.

6. The method for searching a centrally symmetric cross convolutional neural network architecture according to claim 5, characterized in that: The mathematical relationship between the neural network parameter P and the input and output feature map peak memory M and scale parameter s is expressed as: P = f[Net(s)] M = g[Net(s)] Where f[·] represents the calculation of the weight parameters of the convolution kernel of each layer of the network and outputs the summed result, g[·] represents the calculation of the sum of the input and output feature maps of each layer of the network and outputs the maximum value, and Net represents the benchmark network.

7. The method for searching a centrally symmetric cross convolutional neural network architecture according to claim 6, characterized in that: The constraints in step 7 are: Among them, P max Represents the maximum value of the neural network parameters, M max Indicates the maximum value of the peak memory of the input and output feature maps.

8. The method for searching a centrosymmetric cross convolutional neural network architecture according to claim 1, characterized in that: The accuracy of the final network in step 7 is verified by the validation dataset.

9. The method for searching a centrosymmetric cross convolutional neural network architecture according to claim 1, characterized in that: The searching process in step 7 is implemented by a searching algorithm based on population evolution.

10. A chip, characterized in that: The final network as claimed in any one of claims 1 to 9 is deployed.

Citation Information

Patent Citations

  • Infrared target instance segmentation method based on feature fusion and a dense connection network

    CN109584248A

  • Chip surface defect detection method based on sparse space perception and meta learning

    CN115527072A