Construction method and system for computing reusable multi-capacity neural network

RMCNets built through incremental interlayer connections solve the high overhead problem of neural networks when switching between different precisions and resource requirements, realize the computing reusability capability, reduce computing redundancy, and is suitable for multiple application scenarios.

CN120579581APending Publication Date: 2025-09-02湖南工商大学
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510722835.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

When existing neural networks face application switching with different precision and resource requirements, they need to reload the model, resulting in high overhead and system burden, and cannot effectively realize the computing reusable capability.

Method used

Through the incremental inter-layer connection method, computationally reusable multi-capacity neural networks (RMCNets) are built, including progressive inter-layer connections and channel packet connections. The combination of pre-training and multi-capacity joint training ensures that each capacity neural network independently completes tasks and reduces computing redundancy.

Benefits of technology

It realizes the reduction of computing redundancy, reduces system burden in multiple application scenarios, supports the demand switching of resource trade-offs of different precisions, and improves the computing reusability capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579581A_ABST
    Figure CN120579581A_ABST
Patent Text Reader

Abstract

The invention discloses a computing reusable multi-capacity neural network construction method and a computing reusable multi-capacity neural network construction system, which solve the problem by analyzing the reason that a conventional neural network cannot support multiple capacities and reusability and designing an incremental inter-layer connection (IIC). In order to effectively train RMCNets, a training process comprising three stages of pre-training, multi-capacity joint training and fine tuning is provided. Experiments on ImageNet and CIFAR-10 confirm that 4 capacity RMCNets can correctly perform computational reuse, while reducing latency and parameters by about 35%. Experimental analysis on different model training methods shows the efficiency and performance of the training method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer neural networks, and in particular relates to a method and system for constructing a computationally reusable multi-capacity neural network. Background Art

[0002] Neural networks with different architectures and model capacities have their own accuracy-resource tradeoffs. This tradeoff should be determined based on the device's available resources and application requirements (such as accuracy and latency). Furthermore, applications installed on the same mobile device but with different requirements may input the same image (from the same camera). To cope with these varying requirements, the mobile device must recalculate the neural network using different resource-accuracy tradeoffs when switching applications.

[0003] Neural networks can achieve higher accuracy by consuming more compute and memory (GPU memory) to run large models. Alternatively, they can reduce resource consumption and accelerate model inference through model compression methods such as pruning, quantization, and distillation. However, these methods are designed for a static accuracy-resource tradeoff and are difficult to adapt to diverse operating conditions and requirements, such as diverse application requirements, heterogeneous hardware, and dynamically changing available resources at runtime. To support these diverse tradeoffs, numerous neural networks must be designed and trained for different applications and hardware platforms. At runtime, these models are deployed on the same platform. If one tradeoff needs to be switched, i.e., if the model needs to be switched, the neural network must be reloaded. This process results in high overhead during model training, deployment, and runtime. Frequent model switching also places a heavy burden on the system, placing a heavy load not only on mobile devices but also on servers.

[0004] Recently, some researchers have tried to make a single neural network have multiple accuracy-resource trade-off options. These works propose to integrate multiple networks of different model sizes into a multi-capacity neural network (MCNet). In MCNets, each capacity represents a subset of the entire network, completing the target task of the entire network. Parameters on different capacities are fully or partially shared. Through parameter sharing, MCNets can provide a variety of trade-off options using different capacities instead of many separate networks. However, due to the presence of nonlinear components (such as activation functions) in almost all neural networks, current MCNets can only achieve multi-capacity integration in memory, not computationally. In other words, given an input x and a network with two capacities , calculate capacity Make the intermediate calculation results unable to be used in calculation of larger capacity Time reuse.

[0005] The computational reusability between different capacities does not seem to be important. However, on mobile devices, one data source (a camera) usually serves multiple applications, such as Figure 1 The delay-sensitive applications or accuracy-sensitive applications shown are shown. These applications with the same input require backbone networks with different accuracy-resource trade-offs. For example, on a smart car platform, multiple applications such as automatic emergency avoidance and mobile intelligent monitoring can be installed simultaneously. Both applications require target recognition components. However, the difference is that the former requires fast response speed and has strict requirements on latency, while the latter has relatively loose latency requirements and can use larger neural networks to improve accuracy. If the neural network supports computational reusability, the accuracy-sensitive application can infer the larger blue overall network based on the computational results (smaller sub-network) of the delay-sensitive application, thereby avoiding computational waste. This is a problem that neural networks urgently need to solve. Summary of the Invention

[0006] This technical solution analyzes multi-capacity neural networks and the computational reusability between their capacities. It proposes RMCNets, which construct neural networks through incremental inter-layer connections to avoid parameter and output bias. A three-stage training process is designed to ensure that each capacity of RMCNets delivers significant accuracy comparable to individually trained models. Experiments are conducted on two image classification datasets, CIFAR-10 and ImageNet 2012. The results demonstrate that RMCNets can effectively integrate multiple accuracy-resource trade-offs into a single network while providing appropriate computational reusability.

[0007] To this end, the present invention provides the following technical solutions:

[0008] In a first aspect, a computationally reusable multi-capacity neural network construction method is provided, comprising:

[0009] Step 1: Modify the inter-layer neurons of the neural network into incremental inter-layer connections to obtain a reusable multi-capacity neural network architecture;

[0010] Step 2: Pre-training the reusable multi-capacity neural network architecture using a preset training configuration of the neural network to obtain a reusable multi-capacity neural network;

[0011] Step 3: Use the pre-trained reusable multi-capacity neural network to perform multi-capacity joint training, so that when the neural network of each capacity in the reusable multi-capacity neural network independently completes the set target task, and the sum of the loss functions of the neural networks of each capacity reaches the minimum, the trained reusable multi-capacity neural network with a capacity level of K is obtained. .

[0012] The capacity of a neural network refers to the computing power of the neural network. Different capacities correspond to the neural network sub-models formed by different connection relationships of neurons in the neural network. represents a neural network sub-model with a computing power or capacity level of k, , the larger the k is, the more parameters the neural network sub-model has.

[0013] Compared with traditional neural network models, reusable multi-capacity neural networks (RMCNets) can not only integrate multiple capacity sub-models that can perform independent model reasoning into a single model along the width of the network, but also the reasoning process of the high-capacity sub-model can reuse the reasoning results of the low-capacity sub-model, thereby reducing computational redundancy in scenarios where multiple applications coexist.

[0014] The width is the number of neurons in a single hidden layer, and the channel group is a piece cut out along the width of the neural network.

[0015] Furthermore, the incremental inter-layer connection includes two connection methods: progressive inter-layer connection and channel grouping connection;

[0016] The progressive connection is to delete the connection between the high-capacity network layer neurons and the low-capacity network layer neurons in the original neural network, and retain the connection between the low-capacity network layer neurons and the high-capacity network layer neurons;

[0017] The channel grouping connection is to divide all channels of the neural network into independent groups and delete all connections between different groups, wherein the channel group is a piece cut out along the width of the neural network, and the width is the number of neurons in a single hidden layer in the neural network.

[0018] Progressive Inter-layers Connections (PIC) and Channel Grouping (CG) are used to implement reusable multi-capacity neural networks RMCNets. The common point of these two methods is that they need to modify the connection mode of inter-layer neurons, and the neurons are divided into different capacities according to their channel index. In order to ensure that the neurons of the high-capacity sub-model are not connected to the neurons of the low-capacity sub-model in the next neural network layer, these two methods will discard some inter-layer neuron connections. The difference between the two is that the former only deletes the neuron connection from the high-capacity channel group to the low-capacity channel group of the next layer, while the latter needs to delete all cross-capacity connections. Figure 7 As shown in the figure, two model architectures, RMCNetV1 and RMCNetV2, are constructed based on the above two inter-layer neuron connection methods to realize computationally reusable multi-capacity neural networks, where each neural network model has four capacities. , and each capacity has a different number of channels. In the computational reuse scenario, model inference only needs to be performed on channel groups h1, h2, h3, and h4, without redundant inference of the model. .

[0019] The progressive inter-layer connection method PIC deletes the connection between the high-capacity k+1 neurons and the low-capacity (≤k) neurons ( Figure 4 ). Therefore, high-capacity neurons do not affect low-capacity neurons. PIC blocks the flow of information from high-capacity neurons to the next network layer during the forward propagation of the neural network, thereby enabling computational reuse between capacities.

[0020] Furthermore, when the incremental inter-layer connection is a channel group connection, when pre-training the reusable multi-capacity neural network architecture, each channel group is treated as a separate network and the loss function shown in the following formula is used to train each channel group simultaneously ;

[0021] ,

[0022] in, represents the overall loss function of the reusable neural network, represents the cross entropy loss function, represents the output of the k-th channel group, and y represents the label of sample x.

[0023] Furthermore, the pre-classification data and classification labels in the known classification label dataset are used as input and output data respectively to reusable multi-capacity neural networks of various capacities. , using the joint loss function to perform multi-capacity joint training:

[0024] ,

[0025] ,

[0026] in, represents the joint loss function, Representing a reusable multi-capacity neural network with capacity k During training, the loss function when the input data is x and the corresponding classification label data is y; Indicates that a conventional loss function without knowledge distillation is used, and the classification label data used in training is the true label. Represents the distillation loss function using the knowledge distillation method, and the classification label used during training is a reusable multi-capacity neural network with a capacity level of K The predicted labels are obtained, K is the total number of capacity levels of the reusable multi-capacity neural network, and .

[0027] Furthermore, the reusable multi-capacity neural network after multi-capacity joint training is fine-tuned;

[0028] Use parameter freezing to train each channel group The neuron weight parameters are set. In each iteration, the reusable multi-capacity neural network of each capacity is trained in sequence from capacity 1 to K. The training process includes two steps:

[0029] In channel group Perform model inference on , loss calculation, gradient backpropagation and weight update, and then freeze the channel group ;

[0030] In channel group After the weight parameters are updated, the reusable multi-capacity neural network Perform model inference, loss calculation, gradient backpropagation, and weight update, making reusable multi-capacity neural networks Adaptive channel group The weight parameter changes;

[0031] exist After training cycles, the fine-tuning phase is completed and returns the fine-tuned reusable multi-capacity neural network. .

[0032] Freeze means freezing the parameters of some neurons during back propagation (the gradient of the frozen neurons is set to 0)

[0033] The capacity of the reusable multi-capacity neural network is the same as the number of channel groups; the loss is calculated according to the joint loss function; and the weights of the neurons on the channel group are updated using the stochastic gradient descent method.

[0034] Furthermore, the calculation formula of the reusable multi-capacity neural network is as follows:

[0035] ,

[0036] ,

[0037] in and represents a network layer of a reusable multi-capacity neural network with capacity k+1 and k respectively, Represents the k+1th channel group of a network layer in a reusable multi-capacity neural network; represents a computation reuse operation, is a non-negative integer, 、 、 They represent the output results of a network layer sub-model of a reusable multi-capacity neural network with capacities of k+t, k+t-1, and k+t-2 when the input is x; and Represents the output data of neurons in the k+tth and k+t-1th channel groups when the input of the reusable multi-capacity neural network is x; It represents the output of a network layer sub-model of a reusable multi-capacity neural network with capacity k when the first k channel groups of x are used as input, and K represents the total number of capacity levels.

[0038] In the above formula, the first equation indicates that in a multi-capacity model, a low-capacity sub-model is a parameter subset of a higher-capacity sub-model, that is, the low-capacity is a subset of the high-capacity sub-model; the second equation shows that it can be reused The calculation results are used to calculate For each network layer l in RMCNets using incremental inter-layer connections,

[0039] ,

[0040] in, represents the output of the k-th channel group in the l-th hidden layer, represents the output of the k-1th capacity of the lth hidden layer, Indicates a computational reuse operation. concat(a, b) represents the concatenation of a and b.

[0041] Note that due to the incremental inter-layer connections, in PIC, the sub-model The input and The input is different.

[0042] The weight parameter is capacity A subset of weight parameters, for each layer in the neural network, divides the neurons into Channel Groups Here, the channel group Represents a channel group (i.e., a group of neurons) in a hidden layer of a neural network, rather than a channel group for the entire network. represents the kth capacity of a network layer in a reusable multi-capacity neural network, that is, , ; x represents the given input, Represents a computation reuse operation.

[0043] A row represents a hidden layer, a hidden layer has 4 single-layer channel groups h, and a column represents a multi-layer channel group H, f k represents the first k columns of a hidden layer, Fk Represents the first k columns of all rows.

[0044] The capacity of the reusable multi-capacity neural network is the same as the number of channel groups.

[0045] The incremental inter-layer connection refers to:

[0046] Along the dimension of neural network model width, multiple sub-models with different parameter scales are constructed in a single neural network model, which is called capacity. Specifically, the neural network is divided into multiple channel groups. , the kth capacity is expressed as , Indicates the union of sets;

[0047] Next, the low-capacity channel group is connected to the high-capacity channel group of the next hidden layer of the neural network, while the connection between the high-capacity channel group and the low-capacity channel group of the next hidden layer of the neural network is discarded.

[0048] Specifically, each channel group contains L neural network hidden layers, namely , for the hidden layer channel group , The neurons in The neurons in the s are connected two by two. is a non-negative integer, and There are no connections between neurons in .

[0049] Incremental inter-layer connections can also be implemented using a mask matrix similar to that used in unstructured pruning, and the appropriate implementation method can be selected based on the specific hardware.

[0050] In the second aspect, a computationally reusable multi-capacity neural network construction system is provided, comprising:

[0051] Reusable multi-capacity neural network architecture building unit: Modify the inter-layer neurons of the neural network into incremental inter-layer connections to obtain a reusable multi-capacity neural network architecture;

[0052] Pre-training unit: pre-training a reusable multi-capacity neural network architecture using a preset training configuration of the neural network to obtain a reusable multi-capacity neural network;

[0053] Joint training unit: Use the pre-trained reusable multi-capacity neural network to perform multi-capacity joint training, so that when the neural network of each capacity in the reusable multi-capacity neural network independently completes the set target task, and the sum of the loss functions of the neural network of each capacity reaches the minimum, a reusable multi-capacity neural network with a trained capacity level of K is obtained. .

[0054] Furthermore, the incremental inter-layer connection includes two connection methods: progressive inter-layer connection and channel grouping connection;

[0055] The progressive connection is to delete the connection between the high-capacity network layer neurons and the low-capacity network layer neurons in the original neural network, and retain the connection between the low-capacity network layer neurons and the high-capacity network layer neurons;

[0056] The channel grouping connection is to divide all channels of the neural network into independent groups and delete all connections between different groups, wherein the channel group is a piece cut out along the width of the neural network, and the width is the number of neurons in a single hidden layer in the neural network.

[0057] According to a third aspect, an electronic terminal includes:

[0058] one or more processors;

[0059] a memory storing one or more computer programs;

[0060] The processor calls the computer program to implement:

[0061] The above steps are for a computationally reusable method for constructing a multi-capacity neural network.

[0062] In a fourth aspect, a computer-readable storage medium stores a computer program, wherein the computer program is invoked by a processor to implement:

[0063] The above steps are for a computationally reusable method for constructing a multi-capacity neural network.

[0064] The technical effect of the present invention lies in analyzing and studying multi-capacity neural networks and the computational reusability between their capacities. A method for constructing computationally reusable multi-capacity neural networks (RMCNets) is proposed. This method constructs neural networks through incremental inter-layer connections to avoid parameter and output bias. A three-stage training process is designed to ensure the effectiveness of each capacity of the RMCNets neural network. Experiments on two image classification datasets, CIFAR-10 and ImageNet 2012, demonstrate that RMCNets can effectively integrate multiple accuracy-resource trade-offs into a single network and provide accurate computational reusability. Multiple capacities can be integrated into the same neural network, and given the same input, a large-capacity model can reuse the computational results of a small-capacity model, thus avoiding computational waste. To meet different needs, mobile devices no longer need to recalculate using neural networks with different resource-accuracy trade-offs when switching applications. Instead, they can directly complete the computation using the reusable multi-capacity neural network proposed in the technical solution of the present invention, eliminating the need to continuously load different neural networks and significantly reducing the system burden. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 A schematic diagram illustrating the application of computing reusability;

[0066] Figure 2 Schematic diagram of a 4-capacity [0.25x, 0.50x, 0.75x, 1.0x] RMCNet;

[0067] Figure 3 Schematic diagram of the analysis of batch normalization and nonlinear functions in traditional neural networks;

[0068] Figure 4 Schematic diagram comparing incremental intra-layer connection and traditional intra-layer connection;

[0069] Figure 5 is the reuse error R(l) of RMCNet and S-Net under different l;

[0070] Figure 6 Schematic diagram of ablation experiments to verify the effectiveness and efficiency of the training process; (a) shows the training time of four neural networks with different numbers of capacity levels; (b) shows the training time of RMCNet using different training processes; (c) shows the Top-1 accuracy of RMCNet on Image.

[0071] Figure 7 Schematic diagrams of two forms of reusable multi-capacity neural networks provided in the embodiments of the technical solution of the present invention, (a) is RMCNet V1, and (b) is RMCNet V2. DETAILED DESCRIPTION

[0072] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0073] The computational reusability between different capacities does not seem to be important, however, on mobile devices, one data source (a camera) usually serves multiple applications, e.g. Figure 1 The delay-sensitive applications or accuracy-sensitive applications shown. These applications with the same input require backbone networks with different accuracy-resource trade-offs. For example, on a smart car platform, multiple applications such as automatic emergency avoidance and mobile intelligent monitoring can be installed at the same time. Both applications require target recognition components. However, the difference is that the former requires fast response speed and has strict requirements on latency, while the latter has relatively loose latency requirements and can use a larger neural network (blue) to improve accuracy. If the neural network supports computational reusability, the accuracy-sensitive application can infer the larger blue overall network based on the calculation results of the delay-sensitive application (smaller green subnetwork), thereby avoiding computational waste.

[0074] The present invention provides a method for constructing a computationally reusable multi-capacity neural network, comprising:

[0075] Step 1: Modify the inter-layer neurons of the neural network into incremental inter-layer connections to obtain a reusable multi-capacity neural network architecture;

[0076] Step 2: Pre-training the reusable multi-capacity neural network architecture using a preset training configuration of the neural network to obtain a reusable multi-capacity neural network;

[0077] For example, based on the ResNet-50 neural network, a reusable multi-capacity neural network architecture RMCNet is constructed. The default training configuration of ResNet-50 is used to train the reusable multi-capacity neural network architecture. Specifically, with an initial learning rate of 0.1, the CosineAnnealingWarmRestarts() learning rate scheduler in pytorch is used to train for 90 cycles. The relevant parameters are set to and ;

[0078] Step 3: Use the pre-trained reusable multi-capacity neural network to perform multi-capacity joint training, so that when the neural network of each capacity in the reusable multi-capacity neural network independently completes the set target task, and the sum of the loss functions of the neural networks of each capacity reaches the minimum, the trained reusable multi-capacity neural network with a capacity level of K is obtained. .

[0079] The capacity of a neural network refers to the computing power of the neural network. Different capacities correspond to the neural network sub-models formed by different connection relationships of neurons in the neural network. represents a neural network sub-model with a computing power or capacity level of k, , the larger the k is, the more parameters the neural network sub-model has.

[0080] See also Figure 2 and Figure 3 Compared with traditional neural network models, reusable multi-capacity neural networks (RMCNets) can not only integrate multiple capacity sub-models that can perform independent model reasoning into a single model along the width of the network, but also reuse the reasoning results of low-capacity sub-models during the reasoning process of high-capacity sub-models, thereby reducing computational redundancy in scenarios where multiple applications coexist.

[0081] Width is the number of neurons in a single hidden layer, and channel group is a piece cut out along the width of the neural network;

[0082] The incremental inter-layer connection includes two connection methods: progressive inter-layer connection and channel group connection, such as Figure 7 As shown;

[0083] The progressive connection is to delete the connection between the high-capacity network layer neurons and the low-capacity network layer neurons in the original neural network, and retain the connection between the low-capacity network layer neurons and the high-capacity network layer neurons;

[0084] The channel grouping connection is to divide all channels of the neural network into independent groups and delete all connections between different groups, wherein the channel group is a piece cut out along the width of the neural network, and the width is the number of neurons in a single hidden layer in the neural network.

[0085] Progressive Inter-layers Connections (PIC) and Channel Grouping (CG) are used to implement reusable multi-capacity neural networks RMCNets. The common point of these two methods is that they need to modify the connection mode of inter-layer neurons, and the neurons are divided into different capacities according to their channel index. In order to ensure that the neurons of the high-capacity sub-model are not connected to the neurons of the low-capacity sub-model in the next neural network layer, these two methods will discard some inter-layer neuron connections. The difference between the two is that the former only deletes the neuron connection from the high-capacity channel group to the low-capacity channel group of the next layer, while the latter needs to delete all cross-capacity connections. Figure 7As shown in (a) and (b), two model architectures, RMCNetV1 and RMCNetV2, are constructed based on the above two inter-layer neuron connection methods to realize computationally reusable multi-capacity neural networks, where each neural network model has four capacities. , and each capacity has a different number of channels. In the computational reuse scenario, model inference only needs to be performed on channel groups h1, h2, h3, and h4, without redundant inference of the model. .

[0086] See also Figure 4 , the progressive inter-layer connection method PIC deletes the connection between the high-capacity k+1 neurons and the low-capacity (≤k) neurons ( Figure 4 ). Therefore, high-capacity neurons do not affect low-capacity neurons. PIC blocks the flow of information from high-capacity neurons to the next network layer during the forward propagation of the neural network, thereby enabling computational reuse between capacities.

[0087] When the incremental inter-layer connections are channel group connections, when pre-training the reusable multi-capacity neural network architecture, each channel group is treated as a separate network and the loss function shown in the following formula is used to train each channel group simultaneously. ;

[0088] ,

[0089] in, represents the overall loss function of the reusable neural network, represents the cross entropy loss function, represents the output of the k-th channel group, and y represents the label of sample x.

[0090] Using the pre-classification data and classification labels in the known classification label dataset as input and output data respectively, a reusable multi-capacity neural network of each capacity is constructed. , using the joint loss function to perform multi-capacity joint training:

[0091] ,

[0092] ,

[0093] in, represents the joint loss function, Representing a reusable multi-capacity neural network with capacity k During training, the loss function when the input data is x and the corresponding classification label data is y; Indicates that a conventional loss function without knowledge distillation is used, and the classification label data used in training is the true label. Represents the distillation loss function using the knowledge distillation method, and the classification label used during training is a reusable multi-capacity neural network with a capacity level of K The predicted labels obtained are, K is the total number of capacity levels of the reusable multi-capacity neural network, that is, .

[0094] Fine-tune the reusable multi-capacity neural network after multi-capacity joint training;

[0095] Use parameter freezing to train each channel group The neuron weight parameters are set. In each iteration, the reusable multi-capacity neural network of each capacity is trained in sequence from capacity 1 to K. The training process includes two steps:

[0096] In channel group Perform model inference on , loss calculation, gradient backpropagation and weight update, and then freeze the channel group ;

[0097] In channel group After the weight parameters are updated, the reusable multi-capacity neural network Perform model inference, loss calculation, gradient backpropagation and weight update to make reusable multi-capacity neural network networks Adaptive channel group The weight parameter changes;

[0098] exist After training cycles, the fine-tuning phase is completed and returns the fine-tuned reusable multi-capacity neural network. .

[0099] Freezing means freezing the parameters of some neurons during backpropagation (the gradient of frozen neurons is set to 0). The capacity of the reusable multi-capacity neural network is the same as the number of channel groups. The loss is calculated according to the joint loss function. The stochastic gradient descent method is used to update the weights of neurons on the channel group.

[0100] The calculation formula of the reusable multi-capacity neural network is as follows:

[0101] ,

[0102] ,

[0103] in and represents a network layer of a reusable multi-capacity neural network with capacity k+1 and k respectively, Represents the k+1th channel group of a network layer in a reusable multi-capacity neural network; represents a computation reuse operation, is a non-negative integer, 、 、 They represent the output results of the reusable multi-capacity neural network sub-model with capacities of k+t, k+t-1, and k+t-2 when the input is x; and Represents the output data of neurons in the k+tth and k+t-1th channel groups when the input of the reusable multi-capacity neural network is x; It represents the output of capacity k when the first k channel groups of x are used as input, and K represents the total number of capacity levels.

[0104] In the above formula, the first equation indicates that in a multi-capacity model, a low-capacity sub-model is a parameter subset of a higher-capacity sub-model, that is, the low-capacity is a subset of the high-capacity sub-model; the second equation shows that it can be reused The calculation results are used to calculate For each network layer l in RMCNets using incremental inter-layer connections, we have:

[0105] ,

[0106] in, represents the output of the k-th channel group in the l-th hidden layer, represents the output of the k-1th capacity of the lth hidden layer, Indicates a computational reuse operation. concat(a, b) represents the concatenation of a and b.

[0107] Note that due to the incremental inter-layer connections, in PIC, the sub-model The input and The input is different.

[0108] The weight parameter is capacity A subset of weight parameters, for each layer in the neural network, divides the neurons into Channel Groups Here, the channel group Represents a channel group (i.e., a group of neurons) in a hidden layer of a neural network, rather than a channel group for the entire network. represents the kth capacity of a network in a reusable multi-capacity neural network, i.e. , ; x represents the given input, Represents a computation reuse operation.

[0109] A row represents a hidden layer, a hidden layer has 4 single-layer channel groups h, and a column represents a multi-layer channel group H, f k represents the first k columns of a hidden layer, F k represents the first k columns of all rows; the capacity of the reusable multi-capacity neural network is the same as the number of channel groups.

[0110] The incremental inter-layer connection refers to:

[0111] Along the dimension of neural network model width, multiple sub-models with different parameter scales are constructed in a single neural network model, which is called capacity. Specifically, the neural network is divided into multiple channel groups. , the kth capacity is expressed as , Indicates the union of sets;

[0112] Next, the low-capacity channel group is connected to the high-capacity channel group of the next hidden layer of the neural network, while the connection between the high-capacity channel group and the low-capacity channel group of the next hidden layer of the neural network is discarded.

[0113] Specifically, each channel group contains L neural network hidden layers, namely , for the hidden layer channel group , The neurons in The neurons in the s are connected two by two. is a non-negative integer, and There are no connections between neurons in .

[0114] Incremental inter-layer connections can also be implemented using a mask matrix similar to that used in unstructured pruning, and the appropriate implementation method can be selected based on the specific hardware.

[0115] Example 1.

[0116] In the empirical study, the ImageNet 2012 and CIFAR-10 datasets were used to evaluate the performance of RMCNets on image classification tasks. In addition, some ablation experiments were conducted to analyze the performance of the incremental inter-layer connections and multi-stage training process proposed in this solution.

[0117] ImageNet 2012 dataset reference: Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, San-jeevSatheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, MichaelBernstein, et al. Imagenet large scale visual recognition challenge. IJCV, 115(3):211–252, 2015. 5, 6.

[0118] CIFAR-10 dataset reference: Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers offeatures from tiny images. 2009. 6, 7.

[0119] We used independent models of varying widths and a 4-capacity slimmable network as the test baseline, represented by I-Net and S-Net, respectively. We implemented RMCNets on three typical neural network architectures: ResNet-50, MobileNet v1, and MobileNet v2. These models were trained on 1.28 million training images and evaluated on 50k validation images.

[0120] Slimming network see: Jiahui Yu, Linjie Yang, Ning Xu, Jianchao Yang, and Thomas Huang. Slimmable neural networks. In ICLR, 2019. 1, 2, 6, 8.

[0121] ResNet-50 model is available at: Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, June 2016. 1, 2, 6,8.

[0122] MobileNet v1 model is available at: Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, WeijunWang, Tobias Weyand, Marco An-dreetto, and Hartwig Adam. Mobilenets: Efficient convolu-tional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017. 1, 2, 6, 8.

[0123] MobileNet v2 model is available at: Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zh-moginov, andLiang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. InCVPR, pages 4510–4520, 2018. 1, 6, 8.

[0124] RMCNet was empirically investigated on the 1000-class ImageNet image classification dataset using four Nvidia Tesla V100 GPUs. RMCNets were implemented on three representative neural network architectures: ResNet-50, MobileNet v1, and MobileNet v2. These models were trained on 1.28 million training images and validated on 50,000 validation images. The top-5 error rate on 100,000 test images was used as the final result. For I-Nets, the following was followed: Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, June 2016. 1, 2, 6, 8. Default training configuration, including learning rate schedule, weight initialization, weight decay, data augmentation, input image resolution, mini-batch size, number of training iterations, and optimizer. For S-Nets, use the reference: Settings in Jiahui Yu, Linjie Yang, Ning Xu, Jianchao Yang, and Thomas Huang. Slimmable neural networks. In ICLR, 2019. 1, 2, 6, 8.

[0125] The 1000-class ImageNet image classification dataset is available at: Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, San-jeevSatheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, MichaelBernstein, et al. Imagenet large scale visual recognition challenge. IJCV, 115(3):211–252, 2015. 5, 6;

[0126] ResNet-50 model is available at: Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, June 2016. 1, 2, 6,8;

[0127] MobileNet v1 model is available at: Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, WeijunWang, Tobias Weyand, Marco An-dreetto, and Hartwig Adam. Mobilenets: Efficient convolu-tional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017. 1, 2, 6, 8;

[0128] MobileNet v2 model is available at: Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zh-moginov, andLiang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. InCVPR, pages 4510–4520, 2018. 1, 6, 8;

[0129] For RMCNets, a similar training configuration was used, except that: i) for each training phase of RMC-ResNets, a batch size of 512 was used instead of 256. After phase 0 with default training settings, phase 1 was trained for 40 epochs, during which the learning rate was linearly decreased from 0.05 to 0. In phase 2, fine-tuning was performed for 20 epochs, during which the learning rate was linearly decreased from 0.02 to 0. ii) For RMC-MobileNet V1 and RMC-MobileNet V2, default settings were used, except that the batch size was 1024 in phase 0. Phase 1 was trained for 100 epochs, during which the learning rate was linearly decreased from 0.1 to 0. Phase 2 was fine-tuned for 40 epochs, during which the learning rate was linearly decreased from 0.05 to 0.

[0130] Table 1: Experimental results of I-Nets, S-Nets and RMCNets on the ImageNet classification task under the same width configuration

[0131] .

[0132] I-Nets model see: Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, June 2016. 1, 2, 6,8.

[0133] S-Nets model see: Jiahui Yu, Linjie Yang, Ning Xu, Jianchao Yang, and Thomas Huang. Slimmable neural networks. In ICLR, 2019. 1, 2, 6, 8.

[0134] Table 1 shows the Top-5 error rates of I-Nets, S-Nets and RMC-Nets under the same width configuration, as well as their parameters and computational complexity. The test results on ImageNet are slightly lower than those of the independently trained network. However, due to the characteristics of the incremental inter-layer connection method, the number of model parameters and computational complexity of the neural network proposed in the technical solution of the present invention are less, and more importantly, RMCNets has multi-capacity reusability. In addition, some analyses are given to explain the reasons for the slightly lower performance: (1) The incremental inter-layer connection method of RMCNets leads to a reduction in model parameters, thereby affecting model performance; (2) During the backpropagation process, the neurons in layer 1 are Can only receive The gradient of the lower capacity neurons can receive more gradients and learn more useful information, so The neurons cannot be trained well. (3) Currently, the training process of RMCNets is mainly based on joint training, which will cause the neurons with lower capacity to have a greater impact on the loss function, while the neurons with higher capacity have a smaller impact on the loss function, making training more difficult. Although S-Nets also has similar problems, the neurons Can receive All gradients of , while RMCNets cannot. Therefore, how to train RMCNets with high accuracy needs further exploration. (4) The hyperparameters of the training setting (such as the learning rate scheduler) are not well configured.

[0135] Table 2: Comparison of Top-1 error rates of I-Nets, S-Nets, and RMCNets based on ResNet18 on CIFAR-10 at the same width

[0136] .

[0137] RMCNets have broad applicability. Note that RMCNets are implemented on MobileNets and ResNets. However, the reusability and multi-capacity capabilities of RMCNets are universally applicable to a wide variety of networks, including standard convolutions, depthwise separable or group convolutions, pooling layers, fully connected layers, residual connections, feature connections, and many other deep neural network components. When RMCNets are applied to other network architectures, only the inter-layer connections need to be adjusted.

[0138] In addition, an empirical study was conducted on the CIFAR-10 dataset, which contains 10 categories, 50,000 training images and 10,000 test images. I-Net, S-Net, and RMC-Net were implemented on the ResNet18 architecture respectively. N-ResNet18 means training a ResNet18 with the normal training process, but testing it at different widths. The same training configuration was used for all network structures without any other training techniques: 120 cycles were trained with a batch size of 128, stochastic gradient descent (SGD) was used as the optimizer, and weight decay was set to Nesterov momentum was not used. The learning rate was multiplied by 0.1 at the 40th, 80th, and 110th training epochs. Table 2 shows the top-1 error rates on CIFAR-10. It was found that N-ResNet only achieved classification performance as a whole. Compared to the individually trained I-ResNet18, RMC-ResNet18 achieved similar performance and surpassed the performance of S-ResNet18, which lacks inter-capacity reuse.

[0139] CIFAR-10 dataset is available at: Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers offeatures from tiny images. 2009. 6, 7.

[0140] Reusability: We studied the reusability of RMCNets through ablation experiments, confirming that RMCNets can correctly implement computational reuse. The following formula is used as the metric:

[0141] ,

[0142] It means that given the same input Under the condition of The difference between the outputs obtained, where It's the network No. The difference between the layers. Figure 5 As shown, for S-Nets, along with This is due to the accumulated error of the nonlinear components of the neural network. In contrast, for RMCNets, For any All remain zero.

[0143] Multi-capacity training overhead: Multi-capacity networks need to be trained through a suitable training process, which increases the training cost. Therefore, the training overhead of multi-capacity networks under the same training configuration is studied. The time cost of training ResNet50 with the same width configuration for RMC-Nets, US-Nets, and S-Nets is compared. In this experiment, the number of width options is increased from 2 to 16 (multiplied by 2 each time). If the number of width options is greater than 4, the proposed sandwich rule is used to train RMCNets, such as Figure 6 As shown in (a), I-Nets and S-Nets increase the number of width options while increasing their training time. US-Nets and RMCNets, which employ the sandwich rule, can support more width options without requiring additional subnetwork training per epoch. Furthermore, RMCNets require less training time due to reduced latency and parameters. Therefore, while RMCNet requires more epochs to train, the training process does not take significantly longer than that of US-Nets.

[0144] The sandwich rule is described in: Jiahui Yu and Thomas S Huang. Universally slimmable net-works and improved training techniques. In ICCV, pages 1803–1811, 2019. 1, 2, 5, 7, 8.

[0145] This example experiment trains RMC-ResNet50 using different training processes, including normal training, the slimmable network training process (Slimmable), RMCNets training phase 1 (RMC-1), RMCNets pre-trained phase 1 (PreRMC-1), and the full RMCNets training process (wholeRMC). Note that the model that passed the normal training process is used as the pretrained model for PreRMC-1 and wholeRMC. The RMC-1 training configuration is the same as the slimmable network training process.

[0146] The total training time and Top-1 accuracy on Imagenet are used to measure the effectiveness and efficiency of the training process, such as Figure 6 (b) and (c). From this we can find that: i) Figure 6(b), The total training time of RMC-1 and Slimmable takes longer than the standard training process because such training processes need to train multiple capacities in each iteration. In PreRMC-1 and wholeRMC, the number of training epochs can be reduced because a good initialized network is obtained through the pre-trained model. Therefore, even if the number of training epochs increases, the training cost can be similar (wholeRMC) or even reduced (PreRMC-1). ii) In Figure 6 In (c), we can see that using the standard training process to train the entire RMCNets model also results in some performance loss, which may be due to the training configuration and connection characteristics. However, using Slimmable or RMC-1 leads to a significant decrease in accuracy (approximately 7.6% on the 1.0x model). Compared with RMC-1, PreRMC-1 and wholeRMC improve the accuracy of the 1.0x model by 3.2% and 4%, respectively, verifying the effectiveness and efficiency of the pre-training and fine-tuning stages.

[0147] Example 2.

[0148] A computationally reusable multi-capacity neural network construction system, comprising:

[0149] Reusable multi-capacity neural network architecture building unit: Modify the inter-layer neurons of the neural network into incremental inter-layer connections to obtain a reusable multi-capacity neural network architecture;

[0150] Pre-training unit: pre-training a reusable multi-capacity neural network architecture using a preset training configuration of the neural network to obtain a reusable multi-capacity neural network;

[0151] Joint training unit: Use the pre-trained reusable multi-capacity neural network to perform multi-capacity joint training, so that when the neural network of each capacity in the reusable multi-capacity neural network independently completes the set target task, and the sum of the loss functions of the neural network of each capacity reaches the minimum, a reusable multi-capacity neural network with a trained capacity level of K is obtained. .

[0152] The incremental inter-layer connection includes two connection methods: progressive inter-layer connection and channel grouping connection;

[0153] The progressive connection is to delete the connection between the high-capacity network layer neurons and the low-capacity network layer neurons in the original neural network, and retain the connection between the low-capacity network layer neurons and the high-capacity network layer neurons;

[0154] The channel grouping connection is to divide all channels of the neural network into independent groups and delete all connections between different groups, wherein the channel group is a piece cut out along the width of the neural network, and the width is the number of neurons in a single hidden layer in the neural network.

[0155] It should be understood that the specific implementation process of each module please refer to the above method content, the present invention will not go into details here, and the division of the above functional modules is only for example illustration. In some embodiments, some functional modules can be merged, and some functional modules can be split. Each functional module can be implemented in software or hardware or a combination of software and hardware. Among them, the software and hardware equipment includes but is not limited to general-purpose computer equipment, programmable gate arrays, digital signal processors, microprocessors and their corresponding programming or burning software.

[0156] Example 3.

[0157] This embodiment provides an electronic terminal, including one or more processors and a memory storing one or more computer programs;

[0158] Among them, the processing calls a computer program to implement: the steps of a method for constructing a computationally reusable multi-capacity neural network.

[0159] For the specific implementation process of each step, please refer to the description of the above method.

[0160] It should be understood that in the embodiments of the present invention, the processor referred to may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provides instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0161] Example 4.

[0162] The present invention provides a computer-readable storage medium storing a computer program, wherein the computer program is called by a processor to implement: steps of a method for constructing a computationally reusable multi-capacity neural network.

[0163] For the specific implementation process of each step, please refer to the description of the above method.

[0164] The readable storage medium is a computer-readable storage medium, which can be an internal storage unit of the software and hardware device described in any of the aforementioned embodiments, such as a hard disk or memory of a controller. The readable storage medium can also be an external storage device of the controller, such as a plug-in hard disk equipped on the controller, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Furthermore, the readable storage medium can also include both an internal storage unit of the controller and an external storage device. The readable storage medium is used to store the computer program and other programs and data required by the controller. The readable storage medium can also be used to temporarily store data that has been output or is to be output.

[0165] Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes instructions for causing a computer device (such as a personal computer, server, or network device) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned readable storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0166] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present application is a flow chart according to the method, device (system), and computer program product of the embodiment of the present application and / or the instructions executed by the processor to generate a device for realizing the function specified in one flow chart or multiple flows and / or one box or multiple boxes of the block diagram. These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a product comprising an instruction device, which realizes the function specified in one flow chart or multiple flows and / or one box or multiple boxes of the block diagram. These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0167] It should be emphasized that the examples described in the present invention are illustrative rather than restrictive. Therefore, the present invention is not limited to the examples described in the specific embodiments. Any other embodiments derived by those skilled in the art based on the technical solution of the present invention that do not depart from the purpose and scope of the present invention, whether modified or replaced, also fall within the scope of protection of the present invention.

Claims

1. A method for constructing a computationally reusable multi-capacity neural network, characterized in that: include: Step 1: Modify the inter-layer neurons of the neural network into incremental inter-layer connections to obtain a reusable multi-capacity neural network architecture; Step 2: Pre-training the reusable multi-capacity neural network architecture using a preset training configuration of the neural network to obtain a reusable multi-capacity neural network; Step 3: Use the pre-trained reusable multi-capacity neural network to perform multi-capacity joint training, so that when the neural network of each capacity in the reusable multi-capacity neural network independently completes the set target task, and the sum of the loss functions of the neural networks of each capacity reaches the minimum, the trained reusable multi-capacity neural network with a capacity level of K is obtained. .

2. The method according to claim 1, characterized in that The incremental inter-layer connection includes two connection methods: progressive inter-layer connection and channel grouping connection; The progressive connection is to delete the connection between the high-capacity network layer neurons and the low-capacity network layer neurons in the original neural network, and retain the connection between the low-capacity network layer neurons and the high-capacity network layer neurons; The channel grouping connection is to divide all channels of the neural network into independent groups and delete all connections between different groups, wherein the channel group is a piece cut out along the width of the neural network, and the width is the number of neurons in a single hidden layer in the neural network.

3. The method according to claim 2, characterized in that When the incremental inter-layer connections are channel group connections, when pre-training the reusable multi-capacity neural network architecture, each channel group is treated as a separate network and trained simultaneously using the loss function shown in the following formula: , in, represents the overall loss function of the reusable neural network, represents the cross entropy loss function, represents the output of the kth channel group, y represents the label of sample x, and K is the number of channel groups.

4. The method according to claim 1 or 3, characterized in that Using the pre-classification data and classification labels in the known classification label dataset as input and output data respectively, a reusable multi-capacity neural network of each capacity is constructed. , using the joint loss function to perform multi-capacity joint training: , , in, represents the joint loss function, Representing a reusable multi-capacity neural network with capacity k During training, the loss function when the input data is x and the corresponding classification label data is y; Indicates that the loss function without knowledge distillation is used, and the classification label data used in training is the real label. Represents the distillation loss function using the knowledge distillation method, and the classification label used during training is a reusable multi-capacity neural network with a capacity level of K The predicted labels obtained, K is the total number of capacity levels of the reusable multi-capacity neural network.

5. The method according to claim 4, characterized in that Fine-tune the reusable multi-capacity neural network after multi-capacity joint training; Use parameter freezing to train each channel group The neuron weight parameters are set. In each iteration, the reusable multi-capacity neural network of each capacity is trained in sequence from capacity 1 to K. The training process includes two steps: In channel group Perform model inference on , loss calculation, gradient backpropagation and weight update, and then freeze the channel group ; In channel group After the weight parameters are updated, the reusable multi-capacity neural network Perform model inference, loss calculation, gradient backpropagation and weight update to make reusable multi-capacity neural network networks Adaptive channel group The weight parameter changes; exist After training cycles, the fine-tuning phase is completed and returns the fine-tuned reusable multi-capacity neural network. .

6. The method according to claim 1, characterized in that The calculation formula of the reusable multi-capacity neural network is as follows: , , in, and represents a network layer of a reusable multi-capacity neural network with capacity k+1 and k respectively, Represents the k+1th channel group of a network layer in a reusable multi-capacity neural network; represents a computation reuse operation, is a non-negative integer, 、 、 They represent the output results of the reusable multi-capacity neural network sub-model with capacities of k+t, k+t-1, and k+t-2 when the input is x; and Represents the output data of neurons in the k+tth and k+t-1th channel groups when the input of the reusable multi-capacity neural network is x; It represents the output of capacity k when the first k channel groups of x are used as input, and K represents the total number of capacity levels.

7. A computationally reusable multi-capacity neural network construction system, characterized in that: include: Reusable multi-capacity neural network architecture building unit: Modify the inter-layer neurons of the neural network into incremental inter-layer connections to obtain a reusable multi-capacity neural network architecture; Pre-training unit: pre-training a reusable multi-capacity neural network architecture using a preset training configuration of the neural network to obtain a reusable multi-capacity neural network; Joint training unit: Use the pre-trained reusable multi-capacity neural network to perform multi-capacity joint training, so that when the neural network of each capacity in the reusable multi-capacity neural network independently completes the set target task, and the sum of the loss functions of the neural network of each capacity reaches the minimum, a reusable multi-capacity neural network with a trained capacity level of K is obtained. .

8. The system according to claim 1, wherein: The incremental inter-layer connection includes two connection methods: progressive inter-layer connection and channel grouping connection; The progressive connection is to delete the connection between the high-capacity network layer neurons and the low-capacity network layer neurons in the original neural network, and retain the connection between the low-capacity network layer neurons and the high-capacity network layer neurons; The channel grouping connection is to divide all channels of the neural network into independent groups and delete all connections between different groups, wherein the channel group is a piece cut out along the width of the neural network, and the width is the number of neurons in a single hidden layer in the neural network.

9. An electronic terminal, characterized in that: include: one or more processors; a memory storing one or more computer programs; The processor calls the computer program to implement: The steps of the method according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that: A computer program is stored, which is called by a processor to implement: The steps of the method according to any one of claims 1 to 6.