Neural network model training and application method and device and storage medium
By calculating and selecting the importance of candidate operations in the neural network model, removing unimportant operations and adjusting weight parameters, the search efficiency and accuracy problems of the deep neural network model on resource-constrained devices are solved, and a smaller network structure and higher performance are achieved.
Patent Information
- Application Number
- CN202311863264.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
When the existing deep neural network model runs on embedded devices with limited resources, there is huge model search time and resource overhead, and the performance ranking of the network model after search is largely different from the actual training performance ranking, which reduces the efficiency and accuracy of the automatic search neural network model structure.
A training method of neural network model is adopted to calculate the importance of candidate operations to network output accuracy, select and remove unimportant operations, and adjust weight parameters to obtain an effective neural network model that is continuous with forward propagation and backpropagation processes.
It reduces the time and resource overhead of neural network model structure search, improves the accuracy of search results, and obtains a neural network model with better performance.
Smart Images

Figure CN120235220A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of modeling of deep neural network models (Deep Neural Networks, DNN), and in particular to a training method suitable for a multi-layer low-bit quantized neural network model. Background Art
[0002] The deep neural network model is a model with a complex network architecture in the field of artificial intelligence, and is also one of the most widely used architectures. Common neural network models include the Convolutional Neural Network (CNN) model. The deep neural network model is widely used in the fields of computer vision, computer hearing, and natural language processing, such as image classification, object recognition and tracking, image segmentation, speech recognition, etc. There are a large number of learnable parameters in the deep neural network model. The linear processing units and nonlinear processing units inside the deep neural network model can be interlaced, and their topological relationships can be complex, and they have the ability to represent arbitrarily complex functions. After a specific learning process, the deep neural network model can have strong recognition and generalization capabilities.
[0003] On the other hand, running deep neural network models requires a lot of memory overhead and abundant processor resources. Although deep neural network models can achieve good performance goals on GPU-based workstations or servers, these deep neural network models are usually not suitable for running on resource-constrained embedded devices, such as smartphones, tablets, various handheld devices, etc.
[0004] In order to solve the above problems, the following solutions can usually be used to optimize the model:
[0005] Pruning / sparseness: During the training of the network, unimportant connections are pruned, and most of the weights in the network become 0, storing the model in a sparse mode. Pruning can be implemented at different levels, such as weight level, channel level, layer level, etc., depending on the task.
[0006] Low-rank factorization: Use structured matrices for low-rank factorization, so that the original dense full-rank matrix can be represented as a combination of several low-rank matrices, and the low-rank matrix can be decomposed into the product of small-scale matrices.
[0007] Quantization: Use a lower bitwidth (1 bit, 2 bits, or 8 bits) to represent 32-bit or higher-precision floating-point numbers, thereby mapping the continuous real values in network parameters and feature maps to discrete integer values, significantly reducing the storage space of parameters and memory occupancy, accelerating the operation speed, and reducing device power consumption.
[0008] Knowledge distillation: Different from pruning and quantization in model compression, knowledge distillation is to train a lightweight small model by constructing it and using the supervision information of a large model with better performance, in order to achieve better performance and accuracy. Specifically, transfer the knowledge of a large network with good performance to a small network through transfer learning, so that the small network model can achieve performance comparable to that of the large model, which can reduce the computational cost.
[0009] Design a lightweight model architecture (compact model architecture): Construct network layers with special structures and train from scratch to obtain network performance suitable for deployment on resource-limited devices. There is no need to specifically store pre-trained models, nor to improve performance through fine-tuning, reducing the time cost, and having the characteristics of small storage capacity, low computational complexity, and good network performance.
[0010] Among the above several technical solutions, designing a compact model architecture has received extensive attention. Especially for the technology of automatically searching for neural network model structures, it can not only greatly reduce the time cost of network structure design, but also obtain a network model structure that meets specific constraint conditions. However, due to the huge search space of neural network model structures, the time and resource overhead of sequentially verifying the performance of all neural network model structures are usually unbearable. Although the technology of automatically searching for neural network model structures has the above advantages, it is still difficult to implement. Summary of the Invention
[0011] To solve the problem of huge time and resource overhead in the above model search, the PC-DARTS algorithm proposes a differentiable structure search method, which transforms the discontinuous structure search problem into a continuous structure parameter evolution problem. The specific method is to design corresponding structure parameters for the network candidate structures to be searched. The structure parameters are importance estimation parameters of the corresponding network candidate structures, which participate in the network forward propagation process in the form of weights and are updated and optimized using gradient optimization algorithms. Finally, the searched network structure is deduced using the structure parameters and manually designed rules.
[0012] The SPOS algorithm proposes a method for searching the structure of a two-stage neural network model, decoupling the training process of the neural network and the search process of the network model structure. The first stage is the optimization process of the neural network model. Each sub-structure in the neural network model adopts a single-path construction method. When optimizing the parameters of the neural network model, only one path is activated and updated each time. After experiencing the iterative optimization process, all sub-network models in the search space are optimized simultaneously, and the weight parameters of the neural network can approximately simulate the weight parameters obtained when each sub-network is independently trained, improving the accuracy of the validation set accuracy of the sub-network model. Although the SPOS algorithm decouples the training process of the neural network and the search process of the network model structure, the parameters between different sub-networks are still coupled during multiple iterations, reducing the accuracy of the validation set accuracy of the sub-network; for different application scenarios of neural network models, different search processes are required, which reduces the search efficiency.
[0013] None of the above technical solutions can gradually reduce the search space during the search process and reduce the difficulty of searching the structure of the neural network model. There is usually a large difference between the estimated performance ranking of the network model obtained after the search and the performance ranking of the network model after actual training, reducing the efficiency and accuracy of automatically searching the structure of the neural network model.
[0014] According to one aspect of the present invention, there is provided a method for training a neural network model, characterized in that the method includes: a calculation step of calculating the importance of candidate operations in the neural network model to the accuracy of the network output based on a measurable metric, wherein the candidate operations in the neural network model at least include one of a first type of operation including learnable parameters or a second type of operation not including learnable parameters; a selection step of selecting candidate operations from the neural network model based on the importance of the candidate operations; an update step of removing the selected candidate operations from the neural network model and adjusting the weight parameters in the neural network model to obtain an effective neural network model, wherein the effective neural network model is a neural network model with continuous forward propagation and backward propagation processes.
[0015] According to another aspect of the present invention, there is provided a training apparatus for a neural network model, characterized in that the training apparatus includes: a calculation unit configured to calculate the importance of candidate operations in the neural network model for the accuracy of the network output based on a measurable metric, wherein the candidate operations in the neural network model at least include a first type of operation including learnable parameters and a second type of operation not including learnable parameters; a selection unit configured to select candidate operations from the neural network model based on the importance metrics of the candidate operations; and an update unit configured to remove the selected candidate operations from the neural network model and adjust the weight parameters in the neural network model to obtain an effective neural network model, wherein the effective neural network model is a neural network model with continuous forward propagation and backward propagation processes.
[0016] According to another aspect of the present invention, there is provided a method for applying a neural network model, including: storing the neural network model trained based on the above training method; receiving a data set corresponding to the task requirements that the stored neural network model can execute; and performing operations on the data set layer by layer from top to bottom in the stored neural network model and outputting the result.
[0017] According to another aspect of the present invention, there is provided an application apparatus for a neural network model, including: a storage module configured to store the neural network model trained based on the above training method; a receiving module configured to receive a data set corresponding to the task requirements that the stored neural network model can execute; and a processing module configured to perform operations on the data set layer by layer from top to bottom in the stored neural network model and output the result.
[0018] According to another aspect of the present invention, there is provided a non-transitory computer-readable storage medium storing instructions that, when executed by a computer, cause the computer to perform the training method of the neural network model based on the above.
[0019] Other features of the present invention will become clear from the following description of exemplary embodiments with reference to the accompanying drawings. Description of the Drawings
[0020] The drawings incorporated in the specification and constituting a part of the specification illustrate exemplary embodiments of the present invention and, together with the description of the exemplary embodiments, are used to explain the principles of the present invention.
[0021] Figure 1 A block diagram illustrating the hardware configuration according to an exemplary embodiment of the present invention.
[0022] Figure 2 A flowchart illustrating the training method of a neural network model according to a first exemplary embodiment of the present invention.
[0023] Figure 3 Illustrates a neural network model architecture.
[0024] Figure 4 Illustrates a flowchart of a training method of a neural network model according to the first exemplary embodiment of the present invention.
[0025] Figure 5 Illustrates a flowchart of a training method of a neural network model according to the first exemplary embodiment of the present invention.
[0026] Figure 6 Illustrates a schematic diagram of a training system according to the second exemplary embodiment of the present invention.
[0027] Figure 7 Illustrates a schematic diagram of a training device according to the third exemplary embodiment of the present invention. Detailed implementation manners
[0028] Hereinafter, exemplary embodiments of the present invention will be described in conjunction with the accompanying drawings. For clarity and conciseness, not all features of the embodiments are described in the specification. However, it should be understood that many implementation-specific settings must be made during the implementation of the embodiments in order to achieve the specific goals of the developer, for example, to comply with those limitations related to the device and the business, and these limitations may vary with different implementations. In addition, it should also be understood that although the development work may be very complex and time-consuming, for those skilled in the art who benefit from the content of the present invention, such development work is only a routine task.
[0029] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the processing steps and / or system structures that are closely related to at least the solution of the present invention are shown in the drawings, while other details that are not closely related to the present invention are omitted.
[0030] (Hardware configuration)
[0031] First, reference will be made to Figure 1 Describe the hardware configuration that can implement the technologies described below.
[0032] The hardware configuration 100 includes, for example, a central processing unit (CPU) 110, a random access memory (RAM) 120, a read-only memory (ROM) 130, a hard disk 140, an input device 150, an output device 160, a network interface 170, and a system bus 180. In one implementation, the hardware configuration 100 can be implemented by a computer, such as a tablet computer, a laptop computer, a desktop computer, or other suitable electronic devices.
[0033] In one implementation, the apparatus for training a neural network model according to the present invention is constructed of hardware or firmware and serves as a module or component of the hardware construction 100. In another implementation, the method for training a neural network model according to the present invention is a software construction stored in the ROM 130 or the hard disk 140 and executed by the CPU 110.
[0034] The CPU 110 is any suitable programmable control device (such as, a processor), and can execute various functions to be described below by executing various application programs stored in the ROM 130 or the hard disk 140 (such as, a memory). The RAM 120 is used to temporarily store programs or data loaded from the ROM 130 or the hard disk 140, and is also used as a space where the CPU 110 executes various processes and other available functions. The hard disk 140 stores various information such as an operating system (OS), various applications, control programs, sample images, trained neural network models, predefined data (such as, thresholds (THs)), etc.
[0035] In one implementation, the input device 150 is used to allow a user to interact with the hardware construction 100. In one example, the user can input a sample image and a label of the sample image (such as, region information of an object, category information of an object, etc.) through the input device 150. In another example, the user can trigger the corresponding processing of the present invention through the input device 150. In addition, the input device 150 can take various forms, such as buttons, keyboards, or touchscreens.
[0036] In one implementation, the output device 160 is used to store the finally trained neural network model into, for example, the hard disk 140 or to output the finally generated neural network model to subsequent image processing such as object detection, object classification, image segmentation, etc.
[0037] The network interface 170 provides an interface for connecting the hardware construction 100 to a network. For example, the hardware construction 100 can perform data communication with other electronic devices connected via the network through the network interface 170. Optionally, a wireless interface can be provided for the hardware construction 100 for wireless data communication. The system bus 180 can provide a data transmission path for mutually transmitting data between the CPU 110, the RAM 120, the ROM 130, the hard disk 140, the input device 150, the output device 160, and the network interface 170, etc. Although called a bus, the system bus 180 is not limited to any specific data transmission technology.
[0038] The above-mentioned hardware construction 100 is merely illustrative and is in no way intended to limit the present invention, its applications, or uses. Moreover, for the sake of brevity, Figure 1Only one hardware configuration is shown. However, multiple hardware configurations can also be used as needed, and the multiple hardware configurations can be connected via a network. In this case, the multiple hardware configurations can be implemented, for example, by a computer (e.g., a cloud server), or by an embedded device such as a camera, a video camera, a personal digital assistant (PDA), or other suitable electronic devices.
[0039] Next, various aspects of the present invention will be described.
[0040] <First Exemplary Embodiment>
[0041] The following will refer to Figures 2 to 5 Describe a training method for a neural network model according to the first exemplary embodiment of the present invention. The specific description of this training method is as follows. The first embodiment shows the main workflow of the neural network model structure search based on gradients and feature maps of the present disclosure.
[0042] See Figure 2 and the specific description of this training method is as follows.
[0043] Step S2100: Construct a neural network model and initialize it.
[0044] Specifically, according to specific task requirements, a neural network model is created in this step. The neural network model has a dense topological connection. The backbone network of the neural network model contains two basic structural blocks, namely, a regular block and a downsampling block. In the regular block, the sizes of the input feature map and the output feature map are the same; in the downsampling block, the stride of the first layer of convolution is set to 2, and the width and height dimensions of the input feature map are twice those of the output feature map, and the remaining structure is the same as that of the regular block.
[0045] In the basic structural block of the neural network model, nodes represent feature maps, and the connecting branches between nodes represent operations (operators). The current node is obtained by fusing the information of all shallow nodes. The output feature map of the basic structural block is the sum or concatenated combination of all internal nodes.
[0046] The connecting branches between nodes are composed of all operations in the search space, and the output of each branch is the sum of the outputs of all operations on the branch. An operation (operator) is a set of functions that perform feature map transformation and represents a computational unit in a deep neural network model. Operations that contain learnable parameters include, but are not limited to: convolution, depthwise separable convolution, dilated convolution, etc. Operations that do not contain learnable parameters include, but are not limited to: max pooling, average pooling, channel pooling, channel shuffle, etc.
[0047] Step S2200: Train the neural network model constructed in S2100.
[0048] The training of the neural network model is a cyclic and repetitive process. Each iteration includes three processes: forward calculation, backward calculation, and parameter update. Among them, forward calculation is to input a batch of data to be trained into the network, perform operations layer by layer from top to bottom in the network model, and obtain the output result of the network. Backward calculation is a process of calculating the loss function based on the true value of the batch of training data and the output result of the network, and propagating the gradient of the loss function forward from the last layer of the network. Parameter update mainly calculates the updated value of the current parameter according to the backpropagated gradient value and the corresponding optimization algorithm. This step trains the neural network model until the network converges or meets the exit condition and terminates.
[0049] Figure 3 A simple neural network model architecture is illustrated (the specific network architecture is not shown). After inputting the data x (feature map) to be trained into the neural network model F, x performs operations layer by layer from top to bottom in the neural network model F, and finally outputs the output result y that meets certain distribution requirements from the neural network model F.
[0050] The training process of the neural network model is a cyclic and repetitive process. Each training includes three processes: forward propagation, backward propagation, and parameter update. Among them, forward propagation is a process of inputting the data x to be trained into the neural network model and performing operations layer by layer from top to bottom in the neural network model. The forward propagation process described in this disclosure can be a known forward propagation process, and the forward propagation process may include the quantization process of any bit of weights and feature maps, and this disclosure does not limit this. If the difference between the actual output result and the expected output result of the neural network model does not exceed the predetermined threshold, it means that the weights in the neural network model are the optimal solution, and the performance of the trained neural network model has reached the expected performance, and the training of the neural network model is completed. On the contrary, if the difference between the actual output result and the expected output result of the neural network model exceeds the predetermined threshold, the backward propagation process needs to be continued, that is, based on the difference between the actual output result and the expected output result, perform operations layer by layer from bottom to top in the neural network model to update the weights in the model so that the performance of the network model after weight update is closer to the expected performance.
[0051] The neural network model applicable to the present invention can be any known model, such as a convolutional neural network model, a recurrent neural network model, and a graph neural network model, etc. The present invention does not limit the type of the network model.
[0052] The computational precision applicable to the neural network model of the present disclosure can be any precision, including high precision and low precision. The terms "high precision" and "low precision" represent the relative levels of precision and do not limit specific numerical values. For example, high precision can be 32-bit floating-point type, and low precision can be 1-bit fixed-point type. Of course, other precisions such as 16-bit, 8-bit, 4-bit, and 2-bit are also included in the range of computational precision applicable to the solution of the present disclosure. The term "computational precision" can refer to the precision of the weights in the neural network model or the precision of the feature maps in the neural network model. The present disclosure does not limit this. The neural network model described in the present disclosure can be a binary neural network model (BNNs). Of course, it is not limited to neural network models with other computational precisions.
[0053] Step S2300: Calculate the importance of operations in the neural network model based on measurable metrics, and select operations in the neural network model based on the calculated importance.
[0054] The measurable metrics are calculated based on the information in the neural network model, and the information includes gradient information, feature map information, combined information of feature maps and gradients, and parameter information in the neural network model. The measurable metrics calculated based on gradient information include single-shot network pruning (snip), gradient signal preservation (grasp), joint flow pruning (synflow), and Jacobian determinant. The measurable metrics calculated based on feature map information include batch normalization scale factors and L2 norms. The measurable metrics calculated based on the combined information of feature maps and gradients include Fisher information. The parameter information in the neural network model includes scale factors and their biases in the normalization layer, as well as weights and their biases of filters, etc.
[0055] The measurable metrics may not be normalized or may be normalized by the following information, which includes the number of floating-point operations, multiply-accumulate operation counts, total memory consumption, total computational consumption, etc.
[0056] The importance metric is a measurable metric that characterizes the importance of an operation. In this exemplary embodiment, the importance metrics based on gradients and feature maps include Fisher information, etc. The formula for Fisher information is:
[0057]
[0058] where L is the objective function and f is the output feature map corresponding to a certain operation. It can be obtained during the gradient backpropagation process without additional calculations. Sort the operations in the difference set between the neural network model and the previously derived neural network model according to the importance scores calculated based on gradients and feature maps, and select the operations with the lowest importance.
[0059] The algorithms that the selection operation may be based on include the greedy algorithm, dynamic programming algorithm, backtracking algorithm, branch and bound algorithm, etc. In this step, a densely topologically connected neural network model is used, and the number and position of operations are not restricted, expanding the search space.
[0060] Step S2400: Remove the operations selected in step S2300 from the neural network model and fine-tune the parameters of the remaining operations in the neural network model.
[0061] Remove the operations selected in step S2300 from the neural network model, and use the neural network model after the removal operation as the updated neural network model. The implementation methods for removing the selected operations from the neural network include masking the output of the operation, deleting the candidate operation from the neural network model, etc. According to the task requirements and training set data, fine-tune the weight parameters of the neural network model. In this embodiment, one or a combination of the backbone network, feature pyramid network, or network head in the neural network model can be updated.
[0062] Step S2500: Repeat steps S2300 - S2400 until the updated neural network model in step S2400 meets the constraint conditions and terminates. The constraint condition can be that the number of repeated searches exceeds a predetermined number.
[0063] Step S2600: Use the operations in the neural network model obtained in step S2500 to construct a neural network model.
[0064] The neural network model can be constructed by combining, stacking, and copying the operations in the neural network model obtained in step S2500. The newly constructed neural network model is the result obtained by searching the neural network model.
[0065] Step 2700: Train the neural network model constructed in step S2600.
[0066] According to specific task requirements and training set data, train the neural network model constructed in step S2600 until the network converges or meets the exit conditions and terminates.
[0067] Through the solution of this exemplary embodiment, by using a densely topologically connected neural network model and not restricting the number and position of candidate operations selected, the search space is expanded, a smaller network structure can be gradually searched, and a neural network model with better performance can be obtained.
[0068] Variant 1
[0069] The following will refer to Figure 4The training method of the neural network model of this exemplary embodiment is described. This exemplary embodiment shows the main workflow of gradient-based neural network model structure update. The specific description of the training method is as follows.
[0070] Step S4100: Similar to step S2100, in this step a neural network model is constructed and initialized.
[0071] Step S4200: Similar to step S2200, in this step the neural network model constructed in S4100 is trained.
[0072] Step S4300: Select operations in the neural network model using a gradient-based importance metric.
[0073] In this exemplary embodiment, the importance index is a measurable index that characterizes the importance of an operation. Gradient-based importance indexes include single network pruning (snip), gradient signal preservation (grasp), joint flow pruning (synflow), Jacobian determinant, etc. The calculation formula for single network pruning (snip) is:
[0074]
[0075] The calculation formula for gradient signal preservation (grasp) is:
[0076]
[0077] The calculation formula for joint flow pruning (synflow) is:
[0078]
[0079] Where L is the objective function, θ is the learnable parameter in the operation, and H is the Hessen matrix. Based on the importance index calculated for each operation mentioned above and the set optimization goal, a true subset of the neural network model structure with the highest importance score can be deduced using the dynamic programming algorithm, so that all operations not in the subset are operations that should be removed.
[0080] Step S4400: Similar to step S2400, in this step, the operation selected in step S4300 is removed from the neural network model and the parameters of the remaining operations in the neural network model are fine-tuned.
[0081] The operation selected in step S4300 is removed from the neural network model, and the neural network model after the removal operation is used as the updated neural network model. The weight parameters of the neural network model are fine-tuned according to the task requirements and the training set data. In this embodiment, one or a combination of the backbone network, the feature pyramid network or the network head in the neural network model can be updated.
[0082] Step S4500: Similar to step S2500, in this step, steps S4300 - S4400 are repeated until the updated neural network model in step S4400 meets the constraint condition and terminates. The constraint condition can be that the degree of accuracy degradation of the updated neural network model exceeds the expectation.
[0083] Step S4600: Similar to step S2600, in this step, use the operations in the neural network model obtained in step S4500 to construct a neural network model.
[0084] The neural network model can be constructed by combining, stacking, and replicating the operations in the neural network model obtained in S4500. The newly constructed neural network model is the result obtained by searching the neural network model.
[0085] Step 4700: Similar to step S2700, in this step, train the neural network model constructed in S4600.
[0086] Train the neural network model constructed in step S4600 according to specific task requirements and training set data until the network converges or meets the exit condition and terminates.
[0087] Through the solution of this exemplary embodiment, a smaller network structure can be quickly searched, and at the same time, the output accuracy meets the predetermined requirements.
[0088] Variant 2
[0089] The following will refer to Figure 5 Describe the training method of the neural network model of this exemplary embodiment. This exemplary embodiment shows the main workflow of the neural network model structure search that simultaneously optimizes accuracy and model size. The specific description of the training method is as follows.
[0090] Step S5100: Similar to step S2100, in this step, construct a neural network model and initialize it.
[0091] Step S5200: Similar to step S2200, in this step, train the neural network model constructed in S5100.
[0092] Step S5300: Select operations in the neural network model using an importance metric that contains a model size constraint.
[0093] The importance metric is a measurable metric that characterizes the importance of an operation. In this exemplary embodiment, the importance metric includes Fisher information, single-shot network pruning (snip), gradient signal preservation (grasp), synflow pruning (synflow), Jacobian determinant, and weighted combinations thereof.
[0094] According to the constraints of the network model size requirements, we design and introduce a normalization factor based on the model size to normalize the aforementioned importance metrics. The normalization factors include: the number of floating-point operations, the number of multiply-accumulate operations, etc.
[0095] S = S / (computational cost)......(Formula 5)
[0096] Sort the operations in the difference set between the neural network model and the neural network model obtained from the aforementioned deduction according to the normalized importance scores, and select the operation with the lowest importance after normalization.
[0097] Step S5400: Similar to step S2400, in this step, remove the operation selected in step S5300 from the neural network model and fine-tune the parameters of the remaining operations in the neural network model.
[0098] Remove the operation selected in step S5300 from the neural network model, and use the neural network model after the removal operation as the updated neural network model. According to the task requirements and training set data, fine-tune the weight parameters of the neural network model. In this embodiment, one or a combination of the backbone network, the feature pyramid network, or the network head in the neural network model can be updated.
[0099] Step S5500: Similar to step S2500, in this step, repeat steps S5300 - S5400 until the updated neural network model in step S5400 meets the constraint conditions and terminates. The constraint condition can be whether the size of the updated neural network model meets the predetermined requirements.
[0100] Step S5600: Similar to step S2600, in this step, use the operations in the neural network model obtained in step S4500 to construct a neural network model.
[0101] The neural network model can be constructed by combining, stacking, and copying the operations in the neural network model obtained in S5500. The newly constructed neural network model is the result obtained by searching the neural network model.
[0102] Step 5700: Similar to step S2700, in this step, train the neural network model constructed in S5600.
[0103] According to specific task requirements and training set data, train the neural network model constructed in step S5600 until the network converges or meets the exit conditions and terminates.
[0104] According to this exemplary embodiment, a network structure that meets the target model size can be gradually searched while keeping the accuracy as unchanged as possible.
[0105] According to the solution of the exemplary embodiment of the present invention, first, according to the design requirements and search space specified by the scenario, a neural network model is defined. This model will include the model structures of all candidate networks in the search space and is an integrated model of all candidate network structures. According to the objective of the task, forward calculation and reverse gradient calculation are performed based on the training set and annotation data. The gradient optimization algorithm is used to iteratively update the parameters of the neural network model until the network converges or the exit condition is met and the process terminates.
[0106] Then, inheriting the weight parameters of the neural network model, based on the optimization result, the importance of each candidate operation in the neural network model is calculated based on the specified importance metric and compared and ranked. The specified importance metric can measure the impact of each candidate operation on the overall performance of the neural network model. According to the importance ranking, the candidate operation with the least impact on the overall performance of the system is selected.
[0107] Finally, the system removes the selected candidate operation from the neural network, and the updated neural network will be used as the new neural network to fine-tune the learnable parameters based on the training set and annotation data.
[0108] The above update process is iteratively performed until all network structure models that can meet the task-related constraints are obtained.
[0109] Table 1 and Table 2 show the technical effects of the technical solution of the present invention for face detection compared with the prior art.
[0110] Table 1
[0111] Large face detection (FPPI = 0.01) Accuracy Manually designed model 85.90% PC-DARTS 81.90% The present invention 87.10%
[0112] Table 2
[0113]
[0114]
[0115] Table 3 and Table 4 show the technical effects of the technical solution of the present invention for face landmark detection compared with the prior art.
[0116] Table 3
[0117]
[0118] Table 4
[0119]
[0120] Wherein, P represents the allowable coordinate error of feature points of the system, in pixels.
[0121] It can be seen that, according to the solution of the exemplary embodiment of the present invention, the time and resource overhead of the neural network model structure search are reduced, and the accuracy of the search result is improved.
[0122] <Second Exemplary Embodiment>
[0123] Based on the foregoing first exemplary embodiment, the second exemplary embodiment of the present invention describes a network model training system. The training system includes a terminal, a communication network, and a server. The terminal and the server communicate with each other through the communication network. The server uses the network model stored locally to train the network model stored in the terminal online, so that the terminal can use the trained network model for real-time services. The following describes each part in the training system of the second exemplary embodiment of the present invention.
[0124] The terminal in the training system can be an embedded image acquisition device such as a security camera, or a device such as a smart phone or a PAD. Of course, the terminal can also not be a terminal with relatively weak computing power such as an embedded device, but other terminals with relatively strong computing power. The number of terminals in the training system can be determined according to actual needs. For example, if the training system is to train the security cameras in a shopping mall, all the security cameras in the shopping mall can be regarded as terminals. At this time, the number of terminals in the training system is fixed. For another example, if the training system is to train the smart phones of users in a shopping mall, the smart phones accessing the wireless local area network of the shopping mall can be regarded as terminals. At this time, the number of terminals in the training system is not fixed. In the second exemplary embodiment of the present invention, the type and number of terminals in the training system are not limited, as long as the terminal can store and train the network model.
[0125] The server in the training system can be a high-performance server with relatively strong computing power, such as a cloud server. The number of servers in the training system can be determined according to the number of terminals it serves. For example, if the number of terminals to be trained in the training system is small or the geographical range of the terminal distribution is small, the number of servers in the training system is small, such as only one server. If the number of terminals to be trained in the training system is large or the geographical range of the terminal distribution is large, the number of servers in the training system is large, such as establishing a server cluster. In the second exemplary embodiment of the present invention, the type and number of servers in the training system are not limited, as long as the server can store at least one network model and provide information for training the network model stored in the terminal.
[0126] The communication network in the second exemplary embodiment of the present invention is a wireless network or wired network used to realize information transmission between the terminal and the server. Currently, any network available for uplink / downlink transmission between the network server and the terminal can be used as the communication network in this embodiment. The second exemplary embodiment of the present invention does not limit the type of communication network and the communication method. Of course, the second exemplary embodiment of the present invention is not limited to other communication methods. For example, a third-party storage area is allocated for this training system. When the terminal and the server want to transmit information to each other, the information to be transmitted is stored in the third-party storage area. The terminal and the server periodically read the information in the third-party storage area to realize information transmission between the two.
[0127] Combine the following Figure 6 , the online training process of the training system of the second exemplary embodiment of the present invention is described in detail. Figure 6 An example of a training system is shown, assuming that the training system includes a terminal and a server. The terminal can take real-time photos. Assuming that a network model that can be trained and can process pictures is stored in the terminal, and the server stores the same network model, the training process of the training system is described as follows.
[0128] Step S201: The terminal initiates a training request to the server via the communication network.
[0129] The terminal initiates a training request to the server through the communication network, and the request includes information such as the terminal identification. The terminal identification is information that uniquely indicates the identity of the terminal (for example, the terminal ID or IP address, etc.).
[0130] This step S201 is described by taking one terminal initiating a training request as an example, and of course multiple terminals may initiate training requests in parallel. The processing process for multiple terminals is similar to that for one terminal, and will not be repeated here.
[0131] Step S202: The server receives a training request.
[0132] exist Figure 6 The training system shown includes only one server, so the communication network can transmit the training request initiated by the terminal to the server. If the training system includes multiple servers, the training request can be transmitted to a relatively idle server according to the idle status of the server.
[0133] Step S203: The server responds to the received training request.
[0134] The server determines the terminal that initiated the request based on the terminal identifier included in the received training request, and then determines the network model to be trained stored in the terminal. One optional method is that the server determines the network model to be trained stored in the terminal that initiated the request based on a comparison table of terminals and network models to be trained; another optional method is that the training request contains information about the network model to be trained, and the server can determine the network model to be trained based on the information. Here, determining the network model to be trained includes but is not limited to determining the network architecture, hyperparameters, and other information that characterizes the network model.
[0135] After the server determines the network model to be trained, the method of the first exemplary embodiment of the present invention can be used to train the network model stored in the terminal that initiates the request using the same network model stored locally in the server. Specifically, the server updates the weights in the network model locally according to the method of steps S2100 to S2700 in the first exemplary embodiment, and transmits the updated weights to the terminal, so that the terminal synchronizes the network model to be trained stored in the terminal according to the received updated weights. Here, the network model in the server and the network model to be trained in the terminal can be the same network model, or the network model in the server is more complex than the network model in the terminal, but the outputs of the two are close. The present disclosure does not limit the types of network models used for training in the server and the network models to be trained in the terminal, as long as the updated weights output from the server can synchronize the network model in the terminal, so that the output of the synchronized network model in the terminal is closer to the expected output.
[0136] exist Figure 6 In the training system shown, the terminal actively initiates the training request. Optionally, the second exemplary embodiment of the present invention is not limited to the server broadcasting an inquiry message and then the terminal responding to the inquiry message to perform the above training process.
[0137] Through the training system described in the second exemplary embodiment of the present invention, the server can perform online training on the network model in the terminal, which improves the flexibility of training; at the same time, it also greatly enhances the business processing capabilities of the terminal and expands the business processing scenarios of the terminal. The above second exemplary embodiment describes the training system by taking online training as an example, but the present invention is not limited to the offline training process, which will not be repeated here.
[0138] <Third Exemplary Embodiment>
[0139] The third exemplary embodiment of the present invention describes a training device for a neural network model, which can execute the training method described in the first exemplary embodiment, and when the device is applied in an online training system, it can be the device in the server described in the second exemplary embodiment.Figure 7 Describe the software structure of the device in detail.
[0140] The training device in this third exemplary embodiment includes a calculation unit 11, a selection unit 12, and an update unit 13. Among them, the calculation unit 11 is used to calculate the importance of candidate operations in the neural network model based on measurable metrics. Among them, the candidate operations in the neural network model at least include a first type of operation including learnable parameters and a second type of operation not including learnable parameters. The selection unit 12 is used to select candidate operations from the neural network model based on the importance metrics of the candidate operations. The update unit 13 is used to remove the selected candidate operations from the neural network model and adjust the weight parameters of the remaining candidate operations in the neural network model.
[0141] The training device of this embodiment also has a module that realizes the functions of the server in the training system, such as the function of identifying the received data, the function of data encapsulation, the function of network communication, etc., which will not be elaborated here.
[0142] Other embodiments
[0143] The embodiments of the present invention can also be implemented by a computer of a system or device that reads and executes computer-executable instructions (for example, one or more programs) recorded on a storage medium (which can also be more completely referred to as a "non-transitory computer-readable storage medium") to execute the functions of one or more of the above embodiments and / or include one or more circuits (for example, an application-specific integrated circuit (ASIC)) for executing the functions of one or more of the above embodiments, and can be implemented by a method executed by the computer of the system or device, by, for example, reading and executing computer-readable instructions from the storage medium to execute the functions of one or more of the above embodiments and / or controlling one or more circuits to execute the functions of one or more of the above embodiments. The computer may include one or more processors (for example, a central processing unit (CPU), a microprocessing unit (MPU)), and may include a network of independent computers or independent processors to read and execute computer-executable instructions. The computer-executable instructions may be provided to the computer from, for example, a network or a storage medium. The storage medium may include, for example, one or more of a hard disk, a random access memory (RAM), a read-only memory (ROM), the storage of a distributed computing system, an optical disc (such as a compact disc (CD), a digital versatile disc (DVD), or a Blu-ray disc (BD) (registered trademark)), a flash device, a memory card, etc.
[0144] Embodiments of the present invention can also be implemented by the following method, that is, software (program) that executes the functions of the above embodiments is provided to the system or device through a network or various storage media, and the computer or central processing unit (CPU) or microprocessing unit (MPU) of the system or device reads and executes the program.
[0145] Although the present invention has been described with reference to exemplary embodiments, it should be understood that the present invention is not limited to the disclosed exemplary embodiments. The scope of the appended claims should be given the broadest interpretation so as to cover all modifications, equivalent structures, and functions.
Claims
1. A training method for a neural network model, the method comprising: A calculation step of calculating the importance of candidate operations in the neural network model to the accuracy of the network output based on a measurable metric, wherein the candidate operations in the neural network model at least include one of a first type of operation including learnable parameters or a second type of operation not including learnable parameters; A selection step of selecting candidate operations from the neural network model based on the importance of the candidate operations; An update step of removing the selected candidate operations from the neural network model and adjusting the weight parameters in the neural network model to obtain an effective neural network model, wherein the effective neural network model is a neural network model with continuous forward propagation and backward propagation processes.
2. The method according to claim 1, wherein the composition of the updated neural network model includes the following ways: combining, stacking, and replicating the remaining candidate operations in the updated neural network model.
3. The method according to claim 1, wherein in the update step, at least one of the backbone network, the feature pyramid network, or the network head in the neural network model is updated.
4. The method according to claim 1, wherein the candidate operation is a set of functions for performing feature map transformation and is a computational unit in a deep neural network model, wherein, The first type of operation includes convolution, depthwise separable convolution, dilated convolution, and the second type of operation includes max pooling, average pooling, channel pooling, channel shuffle.
5. The method according to claim 1, wherein the measurable metric is calculated based on the information in the neural network model.
6. The method according to claim 1, wherein the method on which the selection step is based includes sorting algorithms, greedy algorithms, dynamic programming algorithms, backtracking algorithms, branch and bound algorithms.
7. The method according to claim 1, wherein The measurable metric is updated as the neural network model is updated.
8. The method according to claim 1, wherein the method of removing the selected candidate operations from the neural network model includes masking the output of the operation and deleting the candidate operation from the neural network model.
9. The method according to claim 5, wherein the measurable metric may or may not be normalized or normalized with the following information, and the information includes the number of floating-point operations, the number of multiply-accumulate operations, the total memory consumption, and the total computational consumption.
10. The method according to claim 5, wherein the information includes gradient information, feature map information, combined information of feature maps and gradients, and parameter information in the neural network model.
11. The method according to claim 10, wherein, The measurable metrics calculated based on gradient information include single-shot network pruning (snip), gradient signal preservation (grasp), joint flow pruning (synflow), and Jacobian determinant. The measurable metrics calculated based on feature map information include batch normalization scale factors and L2 norms. The measurable metrics calculated based on the combined information of feature maps and gradients include Fisher information.
12. The method according to claim 10, wherein, The parameter information in the neural network model includes the scale factor and its bias in the normalization layer, and the weight and its bias of the filter.
13. A training device for a neural network model, characterized in that, The training device includes: A computing unit configured to calculate the importance of a candidate operation in a neural network model to the accuracy of the network output based on a measurable metric, wherein the candidate operation in the neural network model at least includes a first type of operation including learnable parameters and a second type of operation not including learnable parameters; A selection unit configured to select a candidate operation from the neural network model based on the importance metric of the candidate operation; An update unit configured to remove the selected candidate operation from the neural network model and adjust the weight parameters in the neural network model to obtain an effective neural network model, wherein the effective neural network model is a neural network model with continuous forward propagation and backward propagation processes.
14. A method for applying a neural network model, characterized in that, The application method includes: Storing the neural network model trained by the training method according to any one of claims 1 to 12; Receiving a data set corresponding to the task requirements executable by the stored neural network model; Performing operations on the data set layer by layer from top to bottom in the stored neural network model and outputting the result.
15. An application device for a neural network model, characterized in that, The application device includes: A storage module configured to store the neural network model trained by the training method according to any one of claims 1 to 12; A receiving module configured to receive a data set corresponding to the task requirements executable by the stored neural network model; A processing module configured to perform operations on the data set layer by layer from top to bottom in the stored neural network model and output the result.
16. A non-transitory computer-readable storage medium storing instructions that, when executed by a computer, cause the computer to perform the training method of the neural network model according to any one of claims 1 to 12.