Neural network generation, data processing method, device, electronic device and medium
By amplifying the target network unit based on the importance level of each alternative network unit in the neural network in multiple iteration cycles, the problem of training high-performance neural networks under low hardware configuration conditions is solved, and the effect of reducing hardware overhead and avoiding redundancy problems is achieved.
Patent Information
- Application Number
- CN202210101262.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-01-27
AI Technical Summary
The prior art is difficult to train high-performance neural networks under low hardware configuration conditions, and the structural re-parameterization method has a problem of large hardware overhead.
By determining the target network unit based on the degree of importance of each alternative network unit in the neural network in multiple iteration cycles, the target network unit is determined and amplified using a predetermined alternative operator to build an amplified neural network, and finally a high-performance neural network is trained under the conditions of lower hardware configuration.
Reducing the hardware overhead of training, such as hardware running time and memory usage, can train high-performance neural networks under lower hardware configuration conditions, avoiding redundancy problems caused by excessive network size.
Smart Images

Figure CN114492754B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of deep learning technologies, and in particular, to a method, apparatus, electronic device, and medium for neural network generation and data processing. Background Art
[0002] With the wide application of neural network models in various fields, the requirements for the inference ability of neural networks are also getting higher and higher; the stronger the inference ability of a neural network, the larger its corresponding network scale, and at the same time, the more computing resources are consumed when performing inference tasks; this results in most products based on neural networks being highly dependent on a good operating environment, restricting the application scope of neural network models. For example, current embedded devices are increasingly difficult to support current neural networks in terms of both bearing capacity and computing power due to limited hardware conditions. Therefore, how to enhance the inference ability of neural networks and control the computational resource overhead and inference complexity of neural networks within an acceptable range has become an urgent problem to be solved. Currently, the above problems are solved by the structural re-parameterization method of neural networks, but this method has the problem of large hardware overhead required in the training stage of neural networks. Summary of the Invention
[0003] The embodiments of the present disclosure at least provide a method, apparatus, electronic device, and medium for neural network generation and data processing.
[0004] In a first aspect, an embodiment of the present disclosure provides a neural network generation method. In each iteration cycle of a plurality of iteration cycles, the following operations are performed: obtaining a first neural network of the current iteration cycle; determining a first target network unit based on the first importance degree information respectively corresponding to each first alternative network unit in the first neural network; the first target network unit includes: at least one convolutional layer; the first importance degree information is used to characterize the contribution degree of the first alternative network unit to the performance of the neural network, and the performance of the neural network characterizes the ability of the neural network to process preset type data in hardware that meets preset resource configuration conditions; using a variety of pre-determined alternative operators to perform amplification processing on the first target network unit, constructing an amplified neural network based on the amplified first target network unit; constructing a second neural network of the current iteration cycle based on the amplified neural network; performing branch merging processing on the second neural network of the last iteration cycle to obtain a target neural network.
[0005] In this way, it is possible to specifically perform an amplification process on the first alternative network units in the neural network that contribute more to the inference ability of the neural network, and reduce the number of network branches in the first alternative network units that contribute less to the inference ability in the amplified network, thereby reducing the hardware overhead of training, such as hardware running time, memory occupation, etc., and a high-performance neural network can be trained even under low hardware configuration conditions.
[0006] In a possible implementation manner, in response to the current iteration cycle being the first iteration cycle, the obtaining of the first neural network of the current iteration cycle includes: obtaining the original neural network; training the original neural network with sample data of a preset type to obtain the first neural network of the current iteration cycle.
[0007] In a possible implementation manner, in response to the current iteration cycle not being the first iteration cycle, the obtaining of the first neural network of the current iteration cycle includes: obtaining the second neural network obtained in the previous iteration cycle of the current iteration cycle; training the second neural network obtained in the previous iteration cycle with sample data of a preset type to obtain the first neural network of the current iteration cycle.
[0008] In a possible implementation manner, before determining the first target network unit based on the first importance degree information respectively corresponding to each first alternative network unit in the first neural network, it includes: obtaining the training loss of the first neural network; determining the first importance degree information respectively corresponding to each first alternative network unit based on the training loss and the parameter information respectively corresponding to each first alternative network unit in the first neural network.
[0009] In this way, the contribution of the operator to reducing the training loss can be measured by the gradient information, and the gradient information can characterize the contribution degree of the network branch to the inference ability of the neural network. Furthermore, it is possible to specifically perform an amplification process on the first alternative network units that contribute more to the inference ability of the neural network, and further increase the scale of the first alternative network units in the entire neural network, thereby avoiding the redundancy problem caused by blindly amplifying all first alternative network units resulting in an overly large network scale, but some of the network units contribute very little to the inference ability. Furthermore, it is beneficial to save hardware resources, enabling the neural network to normally process data of a preset type in a hardware environment with a lower configuration.
[0010] In a possible implementation manner, the parameter information includes: a convolution kernel of a convolution operator corresponding to the first alternative network unit; determining, based on the training loss and the parameter information respectively corresponding to each first alternative network unit in the first neural network, the first importance degree information respectively corresponding to each first alternative network unit, includes: for each first alternative network unit in the first neural network, based on the training loss and each convolution element in the convolution kernel of the convolution operator corresponding to each first alternative network unit, determining the parameter significance corresponding to each convolution element; the parameter significance characterizes the degree of change in the performance of the neural network after deleting the corresponding convolution element in the neural network; based on the parameter significance corresponding to multiple convolution elements in the convolution kernel corresponding to the convolution operator, determining the first importance degree information corresponding to each first alternative network unit.
[0011] In a possible implementation manner, determining the first target network unit based on the first importance degree information respectively corresponding to each first alternative network unit in the first neural network includes: determining the first alternative network unit with the highest importance degree as the first target network unit.
[0012] In a possible implementation manner, using a variety of pre-determined alternative operators to perform an amplification process on the first target network unit includes: randomly initializing the parameters of each of the variety of alternative operators to obtain alternative parameter information respectively corresponding to each of the variety of alternative operators; for each of the variety of alternative operators, performing an amplification process on the first target network unit based on the alternative parameter information corresponding to each of the variety of alternative operators to generate an amplified network unit; constructing an amplified neural network based on the amplified first target network unit includes: replacing the first target network unit in the first neural network with the corresponding amplified network unit to obtain the amplified neural network; wherein, the amplified network unit includes the first target network unit.
[0013] In a possible implementation manner, for each of the variety of alternative operators, performing an amplification process on the first target network unit based on the alternative parameter information corresponding to each of the variety of alternative operators to generate an amplified network unit includes: for each of the variety of alternative operators, based on the alternative parameter information corresponding to each of the variety of alternative operators and the conversion relationship information between the alternative operator and the convolution operator corresponding to the first target network unit, converting the alternative operator into a target convolution operator; generating a network layer corresponding to the target convolution operator; constructing an amplified network branch corresponding to each of the variety of alternative operators based on the network layer corresponding to the target convolution operator; constructing the amplified network unit based on the first target network unit and the amplified network branch.
[0014] In this way, when performing network amplification, the candidate operators are first converted into target convolution operators, and then the target convolution operators are used to generate amplified network branches. In addition to being able to perform further amplification processing based on the amplified network branches, it also facilitates the operation of merging the branches of the second neural network obtained in the last iteration cycle.
[0015] In a possible implementation manner, the amplifying the first target network unit by using a plurality of pre-determined candidate operators further includes: performing transformation processing on the original convolution parameters of the convolution layer in the first target network unit based on the candidate parameter information respectively corresponding to the plurality of candidate operators and the original convolution parameters corresponding to the convolution layer in the first target network unit to obtain target convolution parameters; the transformation processing is used to control the output of the first target network unit to be consistent with the output of the corresponding amplified network unit.
[0016] In this way, by performing transformation processing on the original convolution parameters of the convolution layer in the first target network unit based on the candidate parameter information respectively corresponding to the plurality of candidate operators and the original convolution parameters corresponding to the convolution layer in the first target network unit, the output of the amplified network unit is made consistent with the output of the first target network unit with the original convolution parameters, so as to ensure that the model accuracy of the first neural network will not decrease after the amplification processing.
[0017] In a possible implementation manner, the amplified network unit further includes: a batch normalization layer; the performing transformation processing on the original convolution parameters of the convolution layer in the first target network unit based on the candidate parameter information respectively corresponding to the plurality of candidate operators and the original convolution parameters corresponding to the convolution layer in the first target network unit to obtain target convolution parameters includes: performing parameter fusion processing on the candidate parameter information respectively corresponding to the plurality of candidate operators and the parameter information of the batch normalization layer to obtain target parameters; and performing transformation processing on the original convolution parameters based on the target parameters and the original convolution parameters corresponding to the convolution layer in the first target network unit to obtain the target convolution parameters.
[0018] In a possible implementation, constructing a second neural network for the current iteration cycle based on the augmented neural network includes: determining second importance degree information corresponding to network branches in each second alternative network unit of the augmented neural network; the second alternative network unit includes: an augmented network determined in the current iteration cycle and / or a historical iteration cycle, where the second importance degree information represents the contribution degree of the network branches in the second alternative network unit to the performance of the neural network; determining a second target network unit based on the second importance degree information corresponding to the network branches in each second alternative network unit; pruning the second target network unit to generate the second neural network.
[0019] In a possible implementation, determining a second target network unit based on the second importance degree information corresponding to the network branches in each second alternative network unit includes: determining whether each second alternative network unit meets a preset pruning condition based on the second importance degree information corresponding to the network branches in each second alternative network unit; in response to the second alternative network unit meeting the pruning condition, determining the second alternative network unit as the second target network unit.
[0020] In a possible implementation, pruning the second target network unit to generate the second neural network includes: determining network branches to be pruned based on the second importance degree information corresponding to the network branches in the second target network unit; determining a first target network unit corresponding to the network branches to be pruned; where the network branches to be pruned are augmented network branches in an augmented network unit obtained by augmenting the corresponding first target network unit; merging the parameters of the network branches to be pruned into the corresponding first target network unit, and replacing the second target network unit with the first target network unit after merging the parameters to obtain the second neural network.
[0021] In a possible implementation, performing branch merging processing on the second neural network of the last iteration cycle to obtain a target neural network includes: determining a target augmented network unit to be merged from the second neural network of the last iteration cycle; merging the first target network unit and the augmented network branches included in the target augmented network unit to obtain the target neural network.
[0022] In a second aspect, an embodiment of the present disclosure further provides a generating device for a neural network, including: an iteration module configured to: in each iteration cycle of multiple iteration cycles, execute:
[0023] Obtain a first neural network of the current iteration cycle;
[0024] Determine a first target network unit based on the first importance degree information respectively corresponding to each first alternative network unit in the first neural network; the first target network unit includes: at least one convolutional layer; the first importance degree information is used to characterize the contribution degree of the first alternative network unit to the performance of the neural network, and the performance of the neural network characterizes the ability of the neural network to process preset type data in hardware that meets preset resource configuration conditions;
[0025] Use a variety of pre-determined alternative operators to perform amplification processing on the first target network unit, and construct an amplified neural network based on the amplified first target network unit;
[0026] Based on the amplified neural network, construct a second neural network for the current iteration cycle;
[0027] A merging module is used to perform branch merging processing on the second neural network of the last iteration cycle to obtain a target neural network.
[0028] In a third aspect, an embodiment of the present disclosure further provides a data processing method, including:
[0029] Obtain a target neural network; wherein, the target neural network is generated by using the neural network generation method described in any item of the first aspect;
[0030] Use the target neural network to perform preset processing on the data to be processed to obtain a data processing result.
[0031] In a fourth aspect, an embodiment of the present disclosure further provides a data processing device, including:
[0032] An obtaining module is used to obtain a target neural network; wherein, the target neural network is generated by using the neural network generation method described in any item of the first aspect;
[0033] A processing module is used to perform preset processing on the data to be processed by using the target neural network to obtain a data processing result.
[0034] In a fifth aspect, an alternative implementation manner of the present disclosure further provides an electronic device, a processor, and a memory. The memory stores machine-readable instructions executable by the processor. The processor is used to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the machine-readable instructions are executed by the processor to execute the steps in the first aspect, or any possible implementation manner in the first aspect, or execute the steps in the implementation manner of the third aspect.
[0035] In a sixth aspect, an optional implementation of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run, it executes the steps in the above first aspect, or any possible implementation manner in the first aspect, or executes the steps in the implementation manner of the above third aspect.
[0036] For the effect descriptions of the above neural network generation device, data processing method, data processing device, electronic device, and computer-readable storage medium, refer to the description of the above neural network generation method, which will not be elaborated here.
[0037] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0038] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required to be used in the embodiments will be briefly introduced below. The accompanying drawings are incorporated into the specification and constitute a part of this specification. These drawings show the embodiments that conform to the present disclosure and are used together with the specification to explain the technical solutions of the present disclosure. It should be understood that the following drawings only show some embodiments of the present disclosure, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 Shows a flowchart of a neural network generation method provided by an embodiment of the present disclosure;
[0040] Figure 2 Shows an example of directly constructing an amplification network branch using an alternative operator provided by an embodiment of the present disclosure;
[0041] Figure 3 Shows an example of constructing an amplification network branch after converting an alternative operator into a target convolution operator provided by an embodiment of the present disclosure;
[0042] Figure 4 Shows an example of adding a batch normalization layer to the amplification network branch provided by an embodiment of the present disclosure;
[0043] Figure 5 Shows a specific example of pruning the determined second target network unit provided by an embodiment of the present disclosure;
[0044] Figure 6 Shows an example of generating a target neural network provided by an embodiment of the present disclosure;
[0045] Figure 7Shows a flowchart of a data processing method provided by an embodiment of the present disclosure;
[0046] Figure 8 Shows a schematic diagram of a neural network generation device provided by an embodiment of the present disclosure;
[0047] Figure 9 Shows a schematic diagram of a data processing device provided by an embodiment of the present disclosure;
[0048] Figure 10 Shows a schematic diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only some of the embodiments of the present disclosure, rather than all the embodiments. Components of the embodiments of the present disclosure described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the present disclosure claimed, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.
[0050] Through research, it is found that structural re-parameterization of a neural network, as a method to solve the above problems, during the process of training a neural network, expands the original operations in the neural network into multiple branches for training, enabling the neural network to be trained on a larger scale to enhance the inference ability; after training is completed, the expanded branches are fused, and the neural network obtained after branch fusion neither loses the inference ability nor increases the scale of the neural network, thus meeting the requirements of both enhancing the inference ability of the neural network and not incurring excessive computational resource overhead and inference complexity.
[0051] However, the current structural re-parameterization method will amplify all the amplifiable network branches in the neural network to obtain an enhanced network with amplified branches; the number of network branches in the enhanced network is much larger than that of the original neural network; and the more network branches there are, the greater the hardware overhead required during the training phase, such as hardware running time, memory occupancy, etc., and it is impossible to train a neural network with high performance under low hardware configuration conditions.
[0052] Based on the above research, the present disclosure provides a neural network generation method. By using the first importance degree information corresponding to each first alternative network unit in the neural network, the first target network unit that needs to be processed for network branch amplification is determined. This first importance degree information can represent the contribution degree of the corresponding alternative network unit to the performance of the neural network, that is, it can bring more benefits to the ability of the neural network to process preset type data in hardware that meets the preset resource configuration conditions. Thus, it is possible to specifically perform amplification processing on the first alternative network unit in the neural network that contributes more to the inference ability of the neural network, and reduce the number of network branches in the first alternative network unit that contributes less to the inference ability in the amplified network. Thereby, the hardware cost of training is reduced, and a high-performance neural network can be trained under lower hardware configuration conditions.
[0053] It should be noted that: Similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0054] For ease of understanding of this embodiment, first, a neural network generation method disclosed in the embodiments of the present disclosure will be introduced in detail. The execution subject of the neural network generation method provided in the embodiments of the present disclosure is generally an electronic device with certain computing capabilities. Such an electronic device includes, for example: a terminal device, a server, or other processing devices. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the neural network generation method may be implemented by a processor calling computer-readable instructions stored in a memory.
[0055] The generation of the neural network provided in the embodiments of the present disclosure will be described below.
[0056] See Figure 1 As shown, it is a flowchart of the generation of the neural network provided in the embodiments of the present disclosure. The method includes steps S101 to S102, where:
[0057] S101: In each iteration cycle of multiple iteration cycles, execute:
[0058] S1011: Obtain the first neural network of the current iteration cycle;
[0059] S1012: Determine a first target network unit based on the first importance degree information respectively corresponding to each first alternative network unit in the first neural network; the first target network unit includes: at least one convolutional layer; the first importance degree information is used to characterize the contribution degree of the first alternative network unit to the performance of the neural network, and the performance of the neural network characterizes the ability of the neural network to process preset type data in hardware that meets preset resource configuration conditions;
[0060] S1013: Use a variety of pre-determined alternative operators to perform amplification processing on the first target network unit, and construct an amplified neural network based on the amplified first target network unit;
[0061] S1014: Based on the amplified neural network, construct a second neural network for the current iteration cycle;
[0062] S102: Perform branch merging processing on the second neural network of the last iteration cycle to obtain a target neural network.
[0063] In the embodiments of the present disclosure, in each iteration cycle of multiple iteration cycles, the first neural network of the current iteration cycle is obtained, a first target network unit is determined based on the first importance degree information respectively corresponding to each first alternative network unit in the first neural network, and a variety of pre-determined alternative operators are used to perform amplification processing on the first target network unit to obtain an amplified neural network, and then the obtained amplified neural network is used to obtain the second neural network of the current iteration cycle; after multiple iteration cycles, based on the second neural network of the last iteration cycle, branch merging processing is performed to obtain a target neural network, so that it is possible to specifically perform amplification processing on the first alternative network unit in the neural network that contributes more to the inference ability of the neural network, reduce the number of network branches in the first alternative network unit that contributes less to the inference ability in the amplified network, thereby reducing the hardware overhead of training, and a high-performance neural network can also be trained under low hardware configuration conditions.
[0064] The above S101 - S102 will be described in detail below.
[0065] Regarding the above S101:
[0066] In S1011, if the current iteration cycle is the first iteration cycle among multiple iteration cycles, the first neural network of the current iteration cycle is obtained, for example, in the following manner:
[0067] Obtain an original neural network;
[0068] Use sample data of a preset type to train the original neural network to obtain the first neural network of the current iteration cycle.
[0069] In a specific implementation, the original neural network can be determined in the following manner, for example: Determine the number of network layers of the original neural network, and from a variety of optional operators, determine the operators corresponding to each network layer of the original neural network, and perform parameter initialization on the operators corresponding to each network layer to form the original neural network.
[0070] When determining the number of network layers of the original neural network, for example, the optional range of the number of network layers can be determined in advance according to the deployment environment of the target neural network corresponding to the original neural network, and then within this optional range, determine the number of network layers of the original neural network. The deployment environment includes, for example, the hardware environment where the target neural network is to be deployed, and has computing power characteristics, such as memory, chip capabilities, etc. The larger the memory or the stronger the chip capabilities, the more network layers of the target neural network that can be deployed.
[0071] A variety of optional operators include, for example: convolution operators, pooling operators, fully connected operators, batch normalization operators, etc. Specifically according to actual needs, different operators can be selected.
[0072] After obtaining the original neural network, use sample data of a preset type to train the original neural network to obtain the first neural network of the current iteration cycle.
[0073] The type of sample data is related to the specific function of the target neural network to be generated; for example, it can include any one of the types such as images, voices, texts, etc. Or, the sample data type can be a data type related to the scene type or related to the content contained, such as an image reflecting the in-vehicle environment, a face image, and so on.
[0074] Here, for example, the sample data can be input into the original neural network to obtain the inference result of the original neural network for the sample data. This inference result can be different based on the function of the original neural network. For example, if the function of the original neural network is object recognition, then this inference result is the object recognition result of the sample data; if the function of the original neural network is image classification, then this inference result is the image classification result of the sample data; if the function of the original neural network is audio conversion, then this inference result is the audio conversion result of the sample data.
[0075] Then, use the inference result of the sample data and the annotation information of the sample data to determine the training loss of the original neural network, and then use this training loss to adjust the network internal parameters of each network layer in the original neural network to obtain the first neural network of the first iteration cycle.
[0076] The network internal parameters include, for example, at least one of the following: the convolution kernels of the convolutional layer, the fully connected weights of the fully connected layer, etc.
[0077] In another embodiment, if the current iteration is not the first iteration cycle, the first neural network of the current iteration cycle can be obtained in the following manner:
[0078] Obtain the second neural network obtained in the previous iteration cycle of the current iteration cycle;
[0079] Use sample data of a preset type to train the second neural network obtained in the previous iteration cycle to obtain the first neural network of the current iteration cycle.
[0080] Here, the sample data of the preset type is similar to that in the above embodiment and will not be elaborated here.
[0081] For the specific generation method of the second neural network obtained in the previous iteration cycle of the current iteration cycle, reference can be made to the following specific descriptions of S1013 and S1014 and will not be elaborated here.
[0082] After obtaining the second neural network obtained in the previous iteration cycle, use the sample data to train the second neural network obtained in the previous iteration cycle to obtain the first neural network of the current iteration cycle.
[0083] When training the second neural network obtained in the previous iteration cycle, for example, the sample data can be input into the second neural network obtained in the previous iteration cycle to obtain the inference result of the second neural network obtained in the previous iteration cycle for the sample data.
[0084] Then, use the inference result and the annotation information of the sample data to obtain the training loss of the second neural network obtained in the previous iteration cycle, and then use the training loss to adjust the network internal parameters of each network layer in the second neural network obtained in the previous iteration cycle to obtain the first neural network of the current iteration cycle.
[0085] In S1012, when expanding the network branches of the neural network, since branch expansion can only be achieved for the convolutional layer, the first alternative network unit in the embodiments of the present disclosure is, for example, a network unit including a convolutional layer. In each first alternative network unit, there is at least one convolutional layer. Among them, in the first alternative network unit, each network branch is composed of at least one convolutional layer. This network branch can include the network branch including the convolutional layer in the original neural network, or can be a certain amplified branch in the amplified network unit obtained after performing at least one amplification process on the original neural network.
[0086] The first importance degree information corresponding to each first candidate network unit in the first neural network can characterize the performance contribution of the corresponding first candidate network unit to the neural network. The performance of the neural network characterizes the ability of the neural network to process preset types of data in hardware that meets preset resource configuration conditions. Here, the preset resource configuration conditions refer to the conditions in a specific hardware platform where computing power indicators such as memory and number of processors are within a certain range. The preset resource configuration conditions can be set according to the configuration of the hardware platform on which the neural network training is to be performed, so that a hardware platform with a lighter hardware configuration can complete the training of a high-precision neural network.
[0087] Specifically, for example, the first importance degree information corresponding to each first candidate network unit in the first neural network may be determined in the following manner:
[0088] Obtaining a training loss of the first neural network;
[0089] Based on the training loss and the parameter information respectively corresponding to each first candidate network unit in the first neural network, the first importance degree information respectively corresponding to each first candidate network unit is determined.
[0090] Here, the training loss of the first neural network is, for example, the training loss determined when the original neural network or the second neural network obtained in the previous iteration cycle of the current iteration cycle is trained using the sample data in the above S1011.
[0091] The parameter information corresponding to the first candidate network units respectively includes, for example: the convolution kernel of the convolution operator corresponding to the first candidate network unit; in the convolution kernel, at least one convolution element is included; exemplarily, if the size of the convolution kernel is K×K, the corresponding convolution elements also have K×K.
[0092] When determining the first importance degree information corresponding to each first candidate network unit based on the training loss and the parameter information corresponding to each first candidate network unit in the first neural network, for example, for each first candidate network unit in the first neural network, based on the training loss and each convolution element in the convolution kernel of the convolution operator corresponding to each first candidate network unit, the parameter significance corresponding to each convolution element can be determined; the parameter significance represents the degree of change in the performance of the neural network after the corresponding convolution element is deleted in the neural network;
[0093] Based on the parameter significances respectively corresponding to the multiple convolution elements in the convolution kernel corresponding to the convolution operator, the first importance degree information corresponding to each first candidate network unit is determined.
[0094] Among them, the contribution of an operator to reducing the training loss can be measured by gradient information. That is to say, an operator with a small gradient contributes less to the training loss and is therefore more likely to be redundant; on the contrary, an operator with a large gradient contributes more to the training loss and is therefore more likely to be a relatively important operator. Then, the corresponding first alternative network unit is also more likely to be a network unit that contributes more to the performance of the neural network. Therefore, it is possible to specifically perform an amplification process on the first alternative network unit that contributes more to the performance of the neural network, further increasing the scale of the first alternative network unit in the entire neural network, thereby avoiding the redundancy problem caused by blindly amplifying all first alternative network units, resulting in an overly large network scale, but some of the network units contribute very little to the inference ability.
[0095] In a specific implementation, the parameter significance corresponding to a convolutional element, for example, can represent the contribution degree of the convolutional element to the performance during the training of the first neural network. Among them, the greater the contribution degree of the performance, the greater the contribution degree of the convolutional element to the inference ability of the first neural network.
[0096] An embodiment of the present disclosure provides a specific example for determining the parameter significance of a convolutional element based on a gradient. In this example, the parameter significance S p (θ i ) of the i-th convolutional element, for example, satisfies the following formula (1):
[0097]
[0098] Among them, L is the loss function of the first neural network with parameter θ. Through this loss function, the training loss of the first neural network can be determined. S p is the parameter significance of the j-th convolutional element θ j , where θ j ∈θ, and ⊙ represents the Hadamard product.
[0099] Here, θ only includes the parameters of the convolutional operator.
[0100] Extend S p to determine the first importance degree information S o (θ (i) ) corresponding to the operator by summing all the internal parameters in the operator, which satisfies:
[0101]
[0102] Among them, θ (i) is the parameter in the i-th convolutional operator o (i) in the first neural network; represents the i-th convolutional operator o (i)The j-th convolutional element in; m represents the i-th convolutional operator o (i) The total number of the j-th convolutional elements in.
[0103] The first importance degree information of the operator, that is, the first importance degree information of the operator corresponding to the first alternative network unit. Through the above formulas (1) and (2), the first importance degree information corresponding to each first alternative network unit can be obtained.
[0104] After obtaining the first importance degree information corresponding to each first alternative network unit in the first neural network, based on the first importance degree information corresponding to each first alternative network unit, the first alternative network unit with the highest importance degree can be determined from all the first alternative network units, and the first alternative network unit with the highest importance degree is determined as the first target network unit.
[0105] In the above S1013 and S1014, in the embodiments of the present disclosure, since the target neural network is generated by using the principle of structural reparameterization of the neural network, it is necessary to fuse the branches of the neural network obtained after the amplification process; during the fusion, usually the operators corresponding to each branch are converted into operators, and using the compatibility principle between convolutional operators, multiple convolutional operators are merged into the same convolutional operator.
[0106] Exemplarily, for example, a convolutional operator converts the input feature map I into the output feature map O, where the output feature map O satisfies the following formula (3):
[0107]
[0108] Where represents the convolution operation, F represents the convolution kernel of the convolution operation, and b represents the bias term; since the convolution operation performs a multiply-accumulate operation on the feature elements using multiple convolutional elements, making different convolution operations additive, therefore, different convolution operations with the same convolution kernel size and the same number of output channels satisfy the relationship shown in the following formula (4):
[0109]
[0110] Where F (1) represents the convolution operation o (1) of the convolution kernel, and F (2) represents the convolution operation o (2) of the convolution kernel.
[0111] According to the above formula, different convolutional layers can be fused to obtain a merged convolutional layer without affecting the final output result. The above convolution operation o (1) and o (2)After performing the fusion process, a new convolution operator is obtained, and the convolution kernel of the new convolution operator satisfies: F (3) =(F (1) +F (2) ).
[0112] Therefore, in the embodiments of the present disclosure, in order to be able to perform the merging process of the branches on the second neural network in the last iteration cycle in S102, when performing the amplification process on the first target network unit of the first neural network, various alternative operators determined for the amplified network branches can be converted into convolution operators, for example.
[0113] Among them, after converting the alternative operator into a convolution operator, the size information of the converted convolution operator is consistent with the size information of the convolution operator in the corresponding first target network unit. Exemplarily, if the number of output channels of the convolution operator corresponding to the convolution layer in the first target network unit is D, the convolution kernel size is K, and the number of input channels is C, when performing the amplification process on the first target network unit, each alternative operator can be converted into a target convolution operator with the number of output channels being D, the convolution kernel size being K, and the number of input channels being C.
[0114] The operators that can be converted into convolution operators include, for example, at least one of the following: convolution operator, average pooling operator, multi-scale convolution, sequential convolution, residual connection, etc.
[0115] Among them, sequential convolution, for example, is to perform convolution processing on the data to be processed in sequence using convolution kernels with sizes from 1×1 to K×K, and it can be converted into a convolution operation with a convolution kernel of K×K.
[0116] For the average pooling operator, a K×K average pooling operator can be converted into a K×K convolution operator.
[0117] For multi-scale convolution, the convolution kernel of multi-scale convolution is, for example, K H ×K W , where K H <K, K W <K; such as 1×1 convolution and 1×K convolution, they can be converted into K×K convolution by zero-padding the convolution kernel.
[0118] For residual connection, residual connection can be regarded as a special 1×1 convolution with all weights being 1, so it can be converted into a K×K convolution.
[0119] The above optional operators can all be converted into ordinary convolution through certain transformations. Therefore, in the embodiments of the present disclosure, the above alternative operators can be used to perform the amplification process on the first target network unit.
[0120] When performing amplification processing on the first target network unit using a variety of pre-determined alternative operators, for example, the following method can be adopted:
[0121] Randomly initialize the parameters of each of the variety of alternative operators to obtain alternative parameter information corresponding to each of the variety of alternative operators;
[0122] For each alternative operator among the variety of alternative operators, perform amplification processing on the first target network unit based on the alternative parameter information corresponding to each alternative operator to generate an amplified network unit;
[0123] Then, constructing an amplified neural network based on the first target network unit after amplification processing includes:
[0124] Replace the first target network unit in the first neural network with the corresponding amplified network unit to obtain the amplified neural network; wherein, the amplified network unit includes the first target network unit.
[0125] In a specific implementation, in a possible implementation manner, when using an alternative operator to perform amplification processing on the first target network unit to generate an amplified network unit, for example, a network layer can be directly constructed using the alternative operator, and the network layer corresponding to the alternative operator is used to form the amplified network unit, and then the first target network unit in the first neural network is replaced with the corresponding amplified network unit to obtain the amplified neural network.
[0126] When using the network layer corresponding to the alternative operator to form the amplified network unit, for example, for each alternative operator among the variety of alternative operators, a network layer corresponding to the alternative operator is generated, and then based on the network layer corresponding to the alternative operator, an amplified network branch corresponding to each alternative operator is constructed, and then, an amplified network unit is constructed based on the first target network unit and the amplified network branch.
[0127] In this case, when merging the network branches of the second neural network determined in the last iteration cycle in S102, the alternative operators in each amplified network unit can be first converted into convolution operators, and then the merging process of different network branches in the same amplified network unit can be implemented using the above formula (4).
[0128] Such as Figure 2In the shown example, the first target network unit s includes a convolutional layer with a convolutional kernel size of K×K; for the amplification process of the first target network unit s, the alternative operators are respectively: a sequential convolution operator, an average pooling operator, and a convolutional operator with a convolutional kernel size of 1×1. The amplification network branches in the 3 generated amplification network units include: the amplification network branch s1 corresponding to the sequential convolution operator, the amplification network branch s2 corresponding to the average pooling operator, and the amplification network branch s3 corresponding to the convolutional operator with a convolutional kernel size of 1×1. Then the finally formed amplification network unit includes: the first target network unit s, the amplification network branches s1, s2, and s3.
[0129] In addition, for the amplification network branch determined in the current iteration cycle, it can be used as the first alternative network unit in the future iteration cycle for further amplification. Since only the convolutional operator can be amplified, the embodiments of the present disclosure also provide a specific method for amplifying the first target network unit using alternative operators to generate an amplification network unit, including:
[0130] For each of the multiple alternative operators, based on the alternative parameter information corresponding to each alternative operator and the conversion relationship information between the alternative operator and the convolutional operator corresponding to the first target network unit, convert the alternative operator into a target convolutional operator;
[0131] Generate a network layer corresponding to the target convolutional operator;
[0132] Based on the network layer corresponding to the target convolutional operator, construct an amplification network branch corresponding to each alternative operator;
[0133] Based on the first target network unit and the amplification network branch, construct the amplification network unit.
[0134] In this case, when performing network amplification, first convert the alternative operator into a target convolutional operator, and then use the target convolutional operator to construct an amplification network branch. In addition to being able to perform further amplification on the basis of the amplification network branch, it also facilitates the operation of merging branches of the second neural network obtained in the last iteration cycle in S102.
[0135] Such as Figure 3In the illustrated example, the first target network unit s includes a convolutional layer with a convolutional kernel size of K×K; performing an amplification process on the first target network unit s, the alternative operators are respectively: a sequential convolution operator, an average pooling operator, and a convolutional operator with a convolutional kernel size of 1×1. The above three alternative operators are respectively converted into a target convolutional operator 1, a target convolutional operator 2, and a target convolutional operator 3, and the three amplified network branches generated by using the target convolutional operator 1, the target convolutional operator 2, and the target convolutional operator 3 are respectively: the amplified network branch s4 corresponding to the target convolutional operator 1, the amplified network branch s5 corresponding to the target convolutional operator 2, and the amplified network branch s6 corresponding to the target convolutional operator 3. Finally, the generated amplified network unit includes: the first target network unit s, and the amplified network branches s1 to s6.
[0136] In another embodiment of the present disclosure, since the amplification process is performed on the first target network unit in the first neural network, before the amplification process is performed on the first target network unit, the output of the first target network unit satisfies the above formula (3); and after the amplification process is performed on the first target network unit, the output of the amplified network unit composed of the first target network unit and the amplified network branch corresponding to the first target network unit satisfies the following formula (5):
[0137]
[0138] This changes the original output of the first target network unit, and this change will cause the model accuracy of the first neural network to decrease. In order to stabilize the output of the amplified network unit to ensure that the model accuracy of the first neural network does not decrease after the amplification process, when the present disclosure performs the amplification process on the first target network unit, it will also perform a transformation process on the original convolution parameters of the convolutional layer in the first target network unit based on the alternative parameter information respectively corresponding to the multiple alternative operators and the original convolution parameters corresponding to the convolutional layer in the first target network unit, to obtain target original convolution parameters; the transformation process is used to control the output of the first target network unit to be consistent with the output of the corresponding amplified network unit.
[0139] Among them, the amplified network unit includes: the first target network unit, and the corresponding amplified network branch.
[0140] In a specific implementation, the target original convolution parameter F (ori′) For example, it satisfies the following formula (6):
[0141] F (ori′) ←F (ori) -(F (1) +…+F (n) ) (6)
[0142] Among them, F (ori) represents the original convolution parameters corresponding to the convolutional layer in the first target network unit; F (1) ~F (n) represent the alternative parameter information corresponding to n alternative operators respectively.
[0143] In addition, if the values of the alternative parameters in the obtained alternative operators are relatively large when initializing the alternative operators, it may cause relatively large changes in the convolutional elements in the convolutional layer of the first target network unit, interfering with the training of the convolutional layer in the first target network unit. Therefore, in order to stabilize the training process of the convolutional layer in the first target network unit, a batch normalization layer BN can also be added to the network layer corresponding to the target convolutional operator corresponding to each alternative operator, or a batch normalization layer can be added to the network layer corresponding to each alternative operator.
[0144] For example Figure 4 in the example shown, the first target network unit is s, which includes a convolutional operator OP, and this convolutional operator is connected to a batch normalization layer BN. The alternative operators include: OP1, OP2, OP3. Using the alternative operators OP1, OP2, OP3, the first target network unit s is amplified. The amplified network branches generated for OP1, OP2, OP3 are respectively: s1, s2, and s3. Batch normalization layers: BN1, BN2, and BN3 are respectively added to the amplified network branches corresponding to OP1, OP2, OP3. The amplified amplified network unit formed is as shown in Figure 4 the example.
[0145] Among them, the processing of the input x by the batch normalization layer satisfies the following formula (7):
[0146]
[0147] Among them, x represents the input of the batch normalization layer BN; in the embodiments of the present disclosure, the batch normalization layer BN is usually located after the network layer corresponding to the alternative operator. Therefore, the input of the batch normalization layer BN is the output of the network layer corresponding to the corresponding alternative operator. γ and β are learnable weights used to scale and shift the normalized values; if γ takes a small value and β = 0, then the weight F after the transformation of the corresponding amplified network branch is smaller, and the increase of the amplified network branch has a smaller impact on the original convolution parameters of the convolutional layer in the first target network unit. Therefore, the function (representation ability) of the convolutional layer in the first target network unit will be retained.
[0148] In this case, the following method can be used to perform transformation processing on the original convolution parameters of the convolutional layer in the first target network unit to obtain the target convolution parameters:
[0149] Perform parameter fusion processing on the alternative parameter information corresponding to multiple said alternative operators and the parameter information of the batch normalization layer to obtain target parameters;
[0150] Based on the target parameters and the original convolution parameters corresponding to the convolution layer in the first target network unit, perform transformation processing on the original convolution parameters to obtain the target convolution parameters.
[0151] The obtained target parameters satisfy the following formula (8):
[0152]
[0153] where μ and σ respectively represent the cumulative running mean and variance of the batch normalization layer BN. r represents the number of output channels; F r represents the convolution element corresponding to the r-th output channel in the corresponding amplified network branch; F' r represents the convolution element corresponding to the r-th output channel in the amplified network branch.
[0154] After obtaining the target parameters through the above formula (8), then use the above formula (7) to obtain the target convolution parameters.
[0155] In addition, for the randomly initialized amplified network branch, μ and σ in its BN are both initialized with 0 and 1. Therefore, directly using the initialized BN to determine the target convolution parameters will result in inaccurate values of the target convolution parameters. To re-parameterize the original convolution parameters in the first target network unit accurately, before generating the target convolution parameters, the neural network after amplification processing can also be trained to a certain extent using sample data to calibrate the BN statistics. After training the amplified neural network to a certain extent, then perform parameter fusion on the convolution layer and BN of the amplified network branch after a certain amount of training.
[0156] After adding parallel amplified network branches to the first target network unit and using the alternative parameter information of the amplified network branches and the original convolution parameters of the first target network unit to perform transformation processing on the original convolution parameters, the amplified neural network can be obtained.
[0157] After obtaining the amplified neural network, the amplified neural network can be directly used as the second neural network in the current iteration cycle.
[0158] In another embodiment, the following method can also be used to obtain the second neural network in the current iteration cycle based on the amplified neural network:
[0159] Determine the second importance degree information corresponding to the network branches in each second alternative network unit in the amplified neural network;
[0160] Determine a second target network unit based on the second importance degree information corresponding to each network branch in each second alternative network unit;
[0161] Perform pruning processing on the second target network unit to generate the second neural network.
[0162] In a specific implementation, the second alternative network unit includes: an amplified network unit determined by a current iteration period and / or a historical iteration period, where the second importance degree information represents the contribution degree of the network branch in the second alternative network unit to the performance of the neural network. That is, it includes: a first target network unit in a certain amplified network unit, and amplified network branches corresponding to multiple alternative operators.
[0163] The second importance degree information s of the f-th network branch in the second alternative network unit can be represented by the L1 norm of γ of the last batch normalization layer BN in each network branch of the second alternative network unit j , that is:
[0164]
[0165] In the formula, N represents the number of channels of the last batch normalization layer BN in each network branch of the second alternative network unit. γ k represents γ on the k-th channel of the last batch normalization layer BN in the f-th network branch; γ has the same meaning as in the above formula (7). As Figure 2 shown in the example, for a certain amplified network branch S1, it includes two convolutional layers, namely a convolutional layer with a kernel size of 1×1 and a convolutional layer with a kernel size of K×K. A batch normalization layer is set after each convolutional layer. Then the last batch normalization layer in this network branch is the batch normalization layer after the convolutional layer with a kernel size of K×K.
[0166] After obtaining the second importance degree information corresponding to each second alternative network unit, it is possible to determine whether each second alternative network unit meets a preset pruning condition based on the second importance degree information corresponding to each second alternative network unit;
[0167] In response to any second alternative network unit meeting the pruning condition, determine this any second alternative network unit as the second target network unit.
[0168] Exemplarily, the preset pruning condition includes, for example: when the second importance degree information of a certain second alternative network unit satisfies: It means that it is sufficient to distinguish different amplified network branches in the same amplified network unit. Here, H represents the number of network branches in the second alternative network unit. Among them, Var(·) represents variance.
[0169] When pruning the second target network unit to generate a second neural network, for example, the following method can be adopted:
[0170] Based on the second importance degree information corresponding to each network branch in the second target network unit, determine the network branches to be pruned;
[0171] Determine the first target network unit corresponding to the network branch to be pruned; wherein, the network branch to be pruned is the amplified network branch in the amplified network unit obtained by amplifying the corresponding first target network unit;
[0172] Merge the parameters of the network branch to be pruned into the corresponding first target network unit, and replace the second target network unit with the first target network unit after merging the parameters to obtain the second neural network.
[0173] In a specific implementation, the network branch that can be selected as the second target network unit to be pruned, where λ is a threshold, for example, it can be set to λ = 0.02, 0.03, and can be specifically set according to actual needs, which is not limited in the embodiments of the present disclosure. Among them, Mean(·) represents the mean value.
[0174] Exemplarily, as Figure 2 shown in the example, if s1 is determined as the network branch to be pruned, then s is the corresponding first target network branch.
[0175] When merging the parameters of the network branch to be pruned into the corresponding first target network unit,
[0176] For the amplified network unit {o (ori) , o (1) , …, o (j) , …, o (n)}, where o (ori) represents the first target network unit of any network branch in o (1) ~o (n) . If o (j) is determined as the network branch to be pruned and pruned, then o (j) is merged into o (ori) , and the convolutional parameter F (ori′) of the new first target network unit obtained satisfies the following formula (10):
[0177] F (ori′) ←F (ori) +F (j) (10)
[0178] Among them, F (ori) represents the convolution parameter of o (ori) ; F (j) represents the convolution parameter of o (j) .
[0179] Then, the network branch corresponding to o (j) can be safely deleted from the second target network unit.
[0180] After deleting the network branch to be pruned from the second target network unit, the second neural network is obtained.
[0181] As Figure 5 shown in the example, a specific example of pruning the second target network unit is shown. In this example, the second target network unit includes: the first target network unit s, and multiple amplified network branches. The determined network branch to be pruned is s1, and a specific example of pruning this network branch to be pruned (cut) is shown.
[0182] Regarding the above S102: When performing branch merging processing on the second neural network in the last iteration cycle, for example, the target amplified network unit to be merged can be determined from the second neural network in the last iteration cycle; the target amplified network unit includes multiple network branches; the multiple network branches include: the first target network unit determined in any historical iteration cycle, and the amplified network branches obtained by amplifying this first target network unit;
[0183] The first target network unit and the amplified network branches included in the target amplified network unit are merged to obtain the target neural network.
[0184] Among them, in each target amplified network unit, there is included the first target network unit during amplification processing in a historical iteration cycle, and the amplified network branches obtained by amplifying this first target network unit.
[0185] Here, the amplified network branch can include: the amplified network branch that has not been pruned, or the amplified network branch that is confirmed to be retained after pruning.
[0186] When performing merging processing on the multiple network branches included in the target amplified network unit, for example, the weight parameters in the multiple network branches can be fused.
[0187] Exemplarily, for the amplified network unit {o (ori′) , o (1′) , …, o (j′) , …, o (n′)}, the first target network unit in this amplified network unit is o(ori′) , o (1) ~o (n) denotes the amplified network branch obtained by performing amplification processing on the first target network unit (obtained after training).
[0188] After performing fusion processing on multiple network branches, the convolution parameters corresponding to the obtained convolution operator satisfy the following formula (11):
[0189] F ← F (ori′) +(F (1′) +…+F (n′) ) (11)
[0190] Through the above merging process, the target neural network can be obtained.
[0191] As Figure 6 shown, a specific example of obtaining the target neural network is provided. In this example, it shows that in the first iteration cycle, amplification processing is performed on the original neural network to obtain an amplified neural network, and pruning processing is performed on the amplified neural network to obtain the second neural network in the first iteration cycle. On this basis, iterate (N - 1) more times, where N is the number of iteration cycles, to obtain the second neural network in the last iteration cycle. Then, branch merging processing is performed on the second neural network in the last iteration cycle to obtain the target neural network. Among them, Input refers to the data to be processed input to the neural network; Output refers to the inference result of the data to be processed.
[0192] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation to the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0193] Based on the same inventive concept, as shown in Figure 7 , this disclosure embodiment also provides a data processing method, including:
[0194] S701: Obtain a target neural network; wherein, the target neural network is generated by using the neural network generation method described in any embodiment of this disclosure;
[0195] S702: Use the target neural network to perform preset processing on the data to be processed to obtain a data processing result.
[0196] The target neural network of this disclosure embodiment is generated by using the above neural network generation method, and the hardware overhead required for the generation process is smaller, and it has a higher inference accuracy.
[0197] The data to be processed includes, for example, at least one of the following: image data, voice data, and text data. The preset processing includes, for example, at least one of the following: classifying image data, performing target recognition processing; annotating text data, classifying text data; performing speech recognition processing on voice data, etc.
[0198] Based on the same inventive concept, an embodiment of the present disclosure also provides a neural network generation device corresponding to the neural network generation method. Since the principle of solving problems by the device in the embodiment of the present disclosure is similar to that of the above neural network generation method in the embodiment of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0199] Refer to Figure 8 As shown, it is a schematic diagram of a neural network generation device provided by an embodiment of the present disclosure. The device includes: an iteration module 81 and a merging module 82;
[0200] The iteration module 81 is configured to: in each iteration cycle of multiple iteration cycles, execute:
[0201] Obtain the first neural network of the current iteration cycle;
[0202] Based on the first importance degree information respectively corresponding to each first alternative network unit in the first neural network, determine the first target network unit; the first target network unit includes: at least one convolutional layer; the first importance degree information is used to characterize the contribution degree of the first alternative network unit to the performance of the neural network, and the performance of the neural network characterizes the ability of the neural network to process preset type data in hardware that meets the preset resource configuration conditions;
[0203] Use a variety of pre-determined alternative operators to perform amplification processing on the first target network unit, and construct an amplified neural network based on the amplified first target network unit;
[0204] Based on the amplified neural network, construct the second neural network of the current iteration cycle;
[0205] The merging module 82 is configured to perform branch merging processing on the second neural network of the last iteration cycle to obtain a target neural network.
[0206] In a possible implementation manner, in response to the current iteration cycle being the first iteration cycle, when obtaining the first neural network of the current iteration cycle, the iteration module 81 is configured to:
[0207] Obtain an original neural network;
[0208] Use sample data of a preset type to train the original neural network to obtain the first neural network of the current iteration cycle.
[0209] In a possible implementation manner, in response to the current iteration cycle not being the first iteration cycle, when the iteration module 81 obtains the first neural network of the current iteration cycle, it is used for:
[0210] Obtain the second neural network obtained in the previous iteration cycle of the current iteration cycle;
[0211] Use the sample data to train the second neural network obtained in the previous iteration cycle to obtain the first neural network of the current iteration cycle.
[0212] In a possible implementation manner, before the iteration module 81 determines the target network based on the first importance degree information corresponding to each first alternative network unit in the first neural network, it is further used for:
[0213] Obtain the training loss of the first neural network;
[0214] Based on the training loss and the parameter information corresponding to each first alternative network unit in the first neural network, determine the first importance degree information corresponding to each first alternative network unit.
[0215] In a possible implementation manner, the parameter information includes: the convolution kernel of the convolution operator corresponding to the first alternative network unit;
[0216] When the iteration module 81 determines the first importance degree information corresponding to each first alternative network unit based on the training loss and the parameter information corresponding to each first alternative network unit in the first neural network, it is used for:
[0217] For each first alternative network unit in the first neural network, based on the training loss and each convolution element in the convolution kernel of the convolution operator corresponding to each first alternative network unit, determine the parameter significance corresponding to each convolution element; the parameter significance characterizes the degree of change in the performance of the neural network after deleting the corresponding convolution element in the neural network;
[0218] Based on the parameter significance corresponding to multiple convolution elements in the convolution kernel corresponding to the convolution operator, determine the first importance degree information corresponding to each first alternative network unit.
[0219] In a possible implementation manner, when the iteration module 81 determines the first target network unit based on the first importance degree information corresponding to each first alternative network unit in the first neural network, it is used for:
[0220] Determine the first alternative network unit with the highest importance degree as the first target network unit.
[0221] In a possible implementation, the iterative module 81, when amplifying the first target network unit by using a plurality of predetermined alternative operators, is configured to:
[0222] Randomly initialize the parameters of the plurality of alternative operators respectively to obtain alternative parameter information corresponding to the plurality of alternative operators respectively;
[0223] For each alternative operator among the plurality of alternative operators, amplify the first target network unit based on the alternative parameter information corresponding to each alternative operator to generate an amplified network unit;
[0224] Constructing the amplified neural network based on the first target network unit after amplification processing includes:
[0225] Replacing the first target network unit in the first neural network with the corresponding amplified network unit to obtain the amplified neural network; wherein, the amplified network unit includes the first target network unit.
[0226] In a possible implementation, the iterative module 81, when, for each alternative operator among the plurality of alternative operators, amplifying the first target network unit based on the alternative parameter information corresponding to each alternative operator to generate an amplified network unit, is configured to:
[0227] For each alternative operator among the plurality of alternative operators, convert the alternative operator into a target convolution operator based on the alternative parameter information corresponding to each alternative operator and the conversion relationship information between the alternative operator and the convolution operator corresponding to the first target network unit;
[0228] Generate a network layer corresponding to the target convolution operator;
[0229] Based on the network layer corresponding to the target convolution operator, construct an amplified network branch corresponding to each alternative operator;
[0230] Construct the amplified network unit based on the first target network unit and the amplified network branch.
[0231] In a possible implementation, the iterative module 81, when amplifying the first target network unit by using a plurality of predetermined alternative operators, is further configured to:
[0232] The use of a plurality of predetermined alternative operators to amplify the first target network unit further includes:
[0233] Based on the alternative parameter information corresponding to multiple said alternative operators and the original convolution parameters corresponding to the convolution layer in the first target network unit, perform transformation processing on the original convolution parameters of the convolution layer in the first target network unit to obtain target convolution parameters; the transformation processing is used to control the output of the first target network unit to be consistent with the output of the corresponding augmented network unit.
[0234] In a possible implementation manner, the augmented network unit further includes: a batch normalization layer;
[0235] The iterative module 81, when performing transformation processing on the original convolution parameters of the convolution layer in the first target network unit based on the alternative parameter information corresponding to multiple said alternative operators and the original convolution parameters corresponding to the convolution layer in the first target network unit to obtain target convolution parameters, is used for:
[0236] Perform parameter fusion processing on the alternative parameter information corresponding to multiple said alternative operators and the parameter information of the batch normalization layer to obtain target parameters;
[0237] Based on the target parameters and the original convolution parameters corresponding to the convolution layer in the first target network unit, perform transformation processing on the original convolution parameters to obtain the target convolution parameters.
[0238] In a possible implementation manner, the iterative module 81, when constructing the second neural network of the current iteration cycle based on the augmented neural network, is used for:
[0239] Determine the second importance degree information corresponding to the network branches in each second alternative network unit in the augmented neural network; the second alternative network unit includes: the augmented network unit determined in the current iteration cycle and / or the historical iteration cycle, where the second importance degree information characterizes the contribution degree of the network branches in the second alternative network unit to the performance of the neural network;
[0240] Based on the second importance degree information corresponding to the network branches in each second alternative network unit, determine the second target network unit;
[0241] Perform pruning processing on the second target network unit to generate the second neural network.
[0242] In a possible implementation manner, the iterative module 81, when determining the second target network unit based on the second importance degree information corresponding to the network branches in each second alternative network unit, is used for:
[0243] Based on the second importance degree information corresponding to the network branches in each second alternative network unit, determine whether each second alternative network unit meets a preset pruning condition;
[0244] In response to the second alternative network unit satisfying the pruning condition, the second alternative network unit is determined as the second target network unit.
[0245] In a possible implementation, the iteration module 81, when pruning the second target network unit to generate the second neural network, is configured to:
[0246] Determine the network branches to be pruned based on the second importance degree information respectively corresponding to each network branch in the second target network unit;
[0247] Determine a first target network unit corresponding to the network branch to be pruned; wherein, the network branch to be pruned is an amplified network branch in the amplified network unit obtained by amplifying the corresponding first target network unit;
[0248] Merge the parameters of the network branch to be pruned into the corresponding first target network unit, and replace the second target network unit with the first target network unit after merging the parameters to obtain the second neural network.
[0249] In a possible implementation, the merging module 82, when performing branch merging processing on the second neural network in the last iteration cycle to obtain the target neural network, is configured to:
[0250] Determine the target amplified network unit to be merged from the second neural network in the last iteration cycle;
[0251] Merge the first target network unit and the amplified network branch included in the target amplified network unit to obtain the target neural network.
[0252] The description of the processing flow of each module in the device and the interaction flow between each module can refer to the relevant description in the above method embodiments, which will not be elaborated here.
[0253] See Figure 9 As shown, the embodiments of the present disclosure further provide a data processing device, including:
[0254] An acquisition module 91, configured to acquire a target neural network; wherein, the target neural network is generated by using the neural network generation method as described in any embodiment of the present disclosure;
[0255] A processing module 92, configured to perform a preset process on the data to be processed by using the target neural network to obtain a data processing result.
[0256] The embodiments of the present disclosure further provide an electronic device, as Figure 10As shown in the figure, it is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure, including:
[0257] A processor 11 and a memory 12; the memory 12 stores machine-readable instructions executable by the processor 11, and the processor 11 is configured to execute the machine-readable instructions stored in the memory 12. When the machine-readable instructions are executed by the processor 11, the processor 11 performs the following steps:
[0258] Obtain a first neural network for the current iteration cycle;
[0259] Obtain a first neural network for the current iteration cycle;
[0260] Based on the first importance degree information respectively corresponding to each first alternative network unit in the first neural network, determine a first target network unit; the first target network unit includes: at least one convolutional layer; the first importance degree information is used to characterize the contribution degree of the first alternative network unit to the performance of the neural network, and the performance of the neural network characterizes the ability of the neural network to process preset type data in hardware that meets preset resource configuration conditions;
[0261] Use a variety of pre-determined alternative operators to perform amplification processing on the first target network unit, and construct an amplified neural network based on the amplified first target network unit;
[0262] Based on the amplified neural network, construct a second neural network for the current iteration cycle;
[0263] Perform branch merging processing on the second neural network of the last iteration cycle to obtain a target neural network.
[0264] Or perform the following steps:
[0265] Obtain a target neural network; wherein, the target neural network is generated by using the neural network generation method described in any embodiment of the present disclosure;
[0266] Use the target neural network to perform preset processing on the data to be processed to obtain a data processing result.
[0267] The above-mentioned memory 12 includes an internal memory 121 and an external memory 122; here, the internal memory 121 is also called the main memory, which is used to temporarily store the operation data in the processor 11 and the data exchanged with the external memory 122 such as the hard disk. The processor 11 exchanges data with the external memory 122 through the internal memory 121.
[0268] The specific execution process of the above instructions can refer to the steps of the neural network generation method described in the embodiments of the present disclosure, and will not be elaborated here.
[0269] Embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the neural network generation method or the data processing method described in the foregoing method embodiments. Among them, the storage medium may be a volatile or non-volatile computer-readable storage medium.
[0270] Embodiments of the present disclosure also provide a computer program product, which carries program code. The instructions included in the program code can be used to execute the steps of the neural network generation method or the data processing method described in the foregoing method embodiments. For details, please refer to the foregoing method embodiments and will not be elaborated here.
[0271] Among them, the above computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0272] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here. In the several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0273] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0274] In addition, in each embodiment of the present disclosure, the functional units can be integrated into one processing unit, or each unit exists physically alone, or two or more units are integrated into one unit.
[0275] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0276] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A neural network generation method, characterized in that, in each iteration cycle of multiple iteration cycles, the following is performed: Obtain the first neural network of the current iteration cycle; Based on the first importance degree information respectively corresponding to each first alternative network unit in the first neural network, determine the first target network unit; The first target network unit includes: at least one convolutional layer; the first importance degree information is used to characterize the contribution degree of the first alternative network unit to the performance of the neural network, and the performance of the neural network characterizes the ability of the neural network to process preset type data in hardware that meets preset resource configuration conditions; the preset resource configuration conditions are set according to the configuration of the hardware platform for performing neural network training; Use a variety of pre-determined alternative operators to perform amplification processing on the first target network unit, and construct an amplified neural network based on the amplified first target network unit; Based on the amplified neural network, construct the second neural network of the current iteration cycle; Perform branch merging processing on the second neural network of the last iteration cycle to obtain a target neural network; the application of the target neural network belongs to at least one of the following: target recognition, image classification, audio conversion.
2. The neural network generation method according to claim 1, characterized in that, in response to the current iteration cycle being the first iteration cycle, the obtaining of the first neural network of the current iteration cycle includes: Obtain the original neural network; Use the sample data of the preset type to train the original neural network to obtain the first neural network of the current iteration cycle.
3. The neural network generation method according to claim 1 or 2, characterized in that, in response to the current iteration cycle not being the first iteration cycle, the obtaining of the first neural network of the current iteration cycle includes: Obtain the second neural network obtained in the previous iteration cycle of the current iteration cycle; Use the sample data of the preset type to train the second neural network obtained in the previous iteration cycle to obtain the first neural network of the current iteration cycle.
4. The neural network generation method according to claim 1 or 2, characterized in that, before determining the first target network unit based on the first importance degree information respectively corresponding to each first alternative network unit in the first neural network, it includes: Obtain the training loss of the first neural network; Based on the training loss and the parameter information respectively corresponding to each first alternative network unit in the first neural network, determine the first importance degree information respectively corresponding to each first alternative network unit.
5. The neural network generation method according to claim 4, characterized in that, the parameter information includes: the convolution kernel of the convolution operator corresponding to the first alternative network unit; the determining of the first importance degree information respectively corresponding to each first alternative network unit based on the training loss and the parameter information respectively corresponding to each first alternative network unit in the first neural network includes: For each first alternative network unit in the first neural network, based on the training loss and each convolutional element in the convolutional kernel of the convolutional operator corresponding to each first alternative network unit, determine the parameter significance corresponding to each convolutional element, where the parameter significance characterizes the degree of change in the performance of the neural network after deleting the corresponding convolutional element in the neural network; Based on the parameter significances corresponding to multiple convolutional elements in the convolutional kernel corresponding to the convolutional operator, determine the first importance degree information corresponding to each first alternative network unit.
6. The neural network generation method according to claim 1 or 2, wherein, The determining the first target network unit based on the first importance degree information corresponding to each first alternative network unit in the first neural network includes: Determine the first alternative network unit with the highest importance degree as the first target network unit.
7. The neural network generation method according to claim 1 or 2, wherein, The amplifying the first target network unit by using a variety of pre-determined alternative operators includes: Randomly initialize the parameters of each of the variety of alternative operators to obtain the alternative parameter information corresponding to each of the variety of alternative operators; For each of the variety of alternative operators, based on the alternative parameter information corresponding to each alternative operator, amplify the first target network unit to generate an amplified network unit; The constructing an amplified neural network based on the amplified first target network unit includes: Replace the first target network unit in the first neural network with the corresponding amplified network unit to obtain the amplified neural network; wherein, the amplified network unit includes the first target network unit.
8. The neural network generation method according to claim 7, wherein, The amplifying the first target network unit by using a variety of pre-determined alternative operators for each of the variety of alternative operators, based on the alternative parameter information corresponding to each alternative operator, to generate an amplified network unit includes: For each of the variety of alternative operators, based on the alternative parameter information corresponding to each alternative operator and the conversion relationship information between the alternative operator and the convolutional operator corresponding to the first target network unit, convert the alternative operator into a target convolutional operator; Generate a network layer corresponding to the target convolutional operator; Based on the network layer corresponding to the target convolutional operator, construct an amplified network branch corresponding to each of the variety of alternative operators; Construct the amplified network unit based on the first target network unit and the amplified network branch.
9. The neural network generation method according to claim 8, wherein, The amplifying the first target network unit by using a variety of pre-determined alternative operators further includes: Based on the alternative parameter information corresponding to multiple said alternative operators and the original convolution parameters corresponding to the convolutional layer in the first target network unit, perform transformation processing on the original convolution parameters of the convolutional layer in the first target network unit to obtain target convolution parameters; the transformation processing is used to control the output of the first target network unit to be consistent with the output of the corresponding amplified network unit.
10. The neural network generation method according to claim 9, wherein, the amplified network unit further includes: a batch normalization layer; The performing transformation processing on the original convolution parameters of the convolutional layer in the first target network unit based on the alternative parameter information corresponding to multiple said alternative operators and the original convolution parameters corresponding to the convolutional layer in the first target network unit to obtain target convolution parameters includes: Performing parameter fusion processing on the alternative parameter information corresponding to multiple said alternative operators and the parameter information of the batch normalization layer to obtain target parameters; Based on the target parameters and the original convolution parameters corresponding to the convolutional layer in the first target network unit, perform transformation processing on the original convolution parameters to obtain the target convolution parameters.
11. The neural network generation method according to claim 1, wherein, The constructing a second neural network for the current iteration cycle based on the amplified neural network includes: Determining the second importance degree information corresponding to the network branches in each second alternative network unit in the amplified neural network; the second alternative network unit includes: the amplified network units determined in the current iteration cycle and / or historical iteration cycles, wherein the second importance degree information characterizes the contribution degree of the network branches in the second alternative network unit to the performance of the neural network; Based on the second importance degree information corresponding to the network branches in each second alternative network unit, determining a second target network unit; Performing pruning processing on the second target network unit to generate the second neural network.
12. The neural network generation method according to claim 11, wherein, The determining a second target network unit based on the second importance degree information corresponding to the network branches in each second alternative network unit includes: Based on the second importance degree information corresponding to the network branches in each second alternative network unit, determining whether each second alternative network unit meets a preset pruning condition; In response to the second alternative network unit meeting the pruning condition, determining the second alternative network unit as the second target network unit.
13. The neural network generation method according to claim 11 or 12, wherein, The performing pruning processing on the second target network unit to generate the second neural network includes: Based on the second importance degree information corresponding to each network branch in the second target network unit, determining the network branches to be pruned; Determining a first target network unit corresponding to the network branches to be pruned; wherein the network branches to be pruned are the amplified network branches in the amplified network unit obtained by performing amplification processing on the corresponding first target network unit; Merge the parameters of the network branch to be pruned into the corresponding first target network unit, and replace the second target network unit with the first target network unit after merging the parameters to obtain the second neural network.
14. The neural network generation method according to claim 1 or 2, wherein, the branch merging process for the second neural network in the last iteration cycle to obtain the target neural network includes: Determine the target amplified network unit to be merged from the second neural network in the last iteration cycle; Merge the first target network unit included in the target amplified network unit and the amplified network branch to obtain the target neural network.
15. A neural network generation device, wherein, comprising: An iteration module for: in each iteration cycle of multiple iteration cycles, execute: Obtain the first neural network in the current iteration cycle; Determine the first target network unit based on the first importance degree information corresponding to each first alternative network unit in the first neural network; The first target network unit includes: at least one convolutional layer; the first importance degree information is used to characterize the contribution degree of the first alternative network unit to the performance of the neural network, and the performance of the neural network characterizes the ability of the neural network to process preset type data in hardware that meets the preset resource configuration conditions; the preset resource configuration conditions are set according to the configuration of the hardware platform for performing neural network training; Use a variety of pre-determined alternative operators to perform amplification processing on the first target network unit, and construct an amplified neural network based on the amplified first target network unit; Based on the amplified neural network, construct the second neural network in the current iteration cycle; A merging module for performing branch merging processing on the second neural network in the last iteration cycle to obtain the target neural network; the application of the target neural network belongs to at least one of the following: target recognition, image classification, audio conversion.
16. A data processing method, wherein, comprising: Obtain a target neural network; wherein, the target neural network is generated by using the neural network generation method according to any one of claims 1-14; Use the target neural network to perform preset processing on the data to be processed to obtain a data processing result.
17. A data processing device, wherein, comprising: An acquisition module for acquiring a target neural network; wherein, the target neural network is generated by using the neural network generation method according to any one of claims 1-14; A processing module for using the target neural network to perform preset processing on the data to be processed to obtain a data processing result.
18. An electronic device, wherein, comprising: A processor and a memory, the memory stores machine-readable instructions executable by the processor, and the processor is used to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the processor executes the neural network generation method according to any one of claims 1 to 14, or executes the data processing method according to claim 16.
19. A computer-readable storage medium, characterized in that, a computer program is stored on the computer-readable storage medium, and when the computer program is run by an electronic device, the electronic device executes the neural network generation method according to any one of claims 1 to 14, or executes the data processing method according to claim 16.