Pruning Method, Data Processing Method and Device

By performing amplitude-based pruning processing on the output channels of the neural network model, the computational volume and memory usage problems caused by excessive model scale are solved, and more efficient data processing is achieved.

CN114254745BActive Publication Date: 2025-06-13NAVINFO
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011021829.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-25
Publication Date
2025-06-13
Estimated Expiration
2040-09-25

AI Technical Summary

Technical Problem

In the prior art, due to the large scale of neural network models, the calculation amount and memory usage are large, which affects the data processing efficiency.

Method used

By obtaining all network layers and their output channels in the initial network model, pruning is performed according to the amplitude of each output channel, the pruning network layer is obtained, thereby obtaining the target network model.

Benefits of technology

It reduces the scale of the neural network model, reduces the computational volume and memory usage, and improves the efficiency of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114254745B_ABST
    Figure CN114254745B_ABST
Patent Text Reader

Abstract

The present invention provides a pruning method, a data processing method and a device. The method includes: obtaining all network layers in an initial network model and output channels corresponding to each network layer; for each network layer, respectively determining amplitudes corresponding to the output channels corresponding to the network layer, and performing pruning processing on the output channels corresponding to the network layer according to the amplitudes corresponding to the output channels to obtain a pruned network layer; obtaining a target network model according to the pruned network layer, realizing compression of the network model, so as to reduce the scale of the network model, and further reduce the amount of computation required when using the target neural network model for data processing, while also occupying less memory, thereby improving data processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of computer technology, and in particular, to a pruning method, a data processing method, and a device. Background Art

[0002] With the continuous development of the information society, neural network models have become increasingly mature and are widely used in data processing (such as image processing) scenarios. Currently, when using a neural network model for data processing, data is input into a pre-trained neural network model so that the neural network model performs corresponding data processing on the data.

[0003] However, when using the above neural network model to process data, due to the overly large scale of the neural network model, problems such as excessive computational complexity and large memory occupancy will occur, thereby affecting the efficiency of data processing. Summary of the Invention

[0004] Embodiments of the present invention provide a pruning method, a data processing method, and a device to solve the problems of large computational complexity and large memory occupancy caused by the large scale of the neural network model in the prior art.

[0005] In a first aspect, an embodiment of the present invention provides an image processing method, including:

[0006] Obtain all network layers in the initial network model and the output channels corresponding to each network layer;

[0007] For each network layer, respectively determine the amplitudes corresponding to the output channels of the network layer, and perform pruning processing on the output channels corresponding to the network layer according to the amplitudes corresponding to the output channels to obtain a network layer after pruning processing;

[0008] Obtain a target network model according to the network layer after pruning processing.

[0009] In a possible design, performing pruning processing on the output channels corresponding to the network layer according to the amplitudes corresponding to the output channels includes:

[0010] Take all output channels corresponding to the network layer as objects to be processed, and determine the maximum variance corresponding to the network layer according to the amplitudes corresponding to the output channels in the objects to be processed;

[0011] Determine whether to perform pruning processing on the objects to be processed according to the maximum variance;

[0012] If pruning processing is to be performed on the objects to be processed, determine the output channels to be pruned from the objects to be processed, and perform pruning processing on the output channels to be pruned.

[0013] In a possible design, determining the maximum variance corresponding to the network layer according to the amplitudes corresponding to the output channels in the object to be processed includes:

[0014] Obtaining the number of all output channels corresponding to the network layer;

[0015] Through Determining the maximum variance corresponding to the network layer, where the mσ 2 Is the maximum variance corresponding to the network layer, the x i Is the amplitude corresponding to the i-th output channel, and the N is the number of all output channels.

[0016] In a possible design, after determining whether to perform pruning processing on the object to be processed according to the maximum variance, it further includes:

[0017] If not performing pruning processing on the object to be processed, then stop performing pruning processing on the object to be processed;

[0018] After performing pruning processing on the output channels to be pruned, it further includes:

[0019] Taking the object to be processed after removing the output channels to be pruned as a new object to be processed, and returning to the step of determining the maximum variance corresponding to the network layer according to the amplitudes corresponding to the output channels in the object to be processed.

[0020] In a possible design, the method further includes:

[0021] Obtaining the number of all pruned output channels and obtaining the number of all output channels corresponding to the network layer;

[0022] Calculating the number of all pruned output channels and the number of all output channels corresponding to the network layer to obtain a pruning rate.

[0023] In a possible design, determining whether to perform pruning processing on the object to be processed according to the maximum variance includes:

[0024] Determining whether the maximum variance is greater than a preset variance value;

[0025] If it is greater than the preset variance value, then determine to perform pruning processing on the object to be processed;

[0026] If it is less than or equal to the preset variance value, then determine not to perform pruning processing on the object to be processed;

[0027] Determining the output channels to be pruned from the object to be processed includes:

[0028] Obtain the output channel with the smallest amplitude in the object to be processed, and use it as the output channel to be pruned.

[0029] In a possible design, obtaining the target network model according to the pruned network layer includes:

[0030] Train the pruned network layer to obtain the target network model.

[0031] In a second aspect, an embodiment of the present invention provides an image processing method, including:

[0032] Obtain the data to be processed;

[0033] Use the target network model to perform corresponding data processing on the data to be processed, where the target network model is obtained by acquiring all network layers in the initial network model and the output channels corresponding to each network layer, for each network layer, respectively determining the amplitudes corresponding to the output channels corresponding to the network layer, pruning the output channels corresponding to the network layer according to the amplitudes corresponding to the output channels to obtain the pruned network layer, and obtaining according to the pruned network layer.

[0034] In a possible design, if the data to be processed includes an image to be processed, then using the target network model to perform corresponding data processing on the data to be processed includes:

[0035] Use the target network model to perform image processing on the image to be processed, where the image processing includes object detection processing and / or image segmentation processing.

[0036] In a possible design, if the image to be processed includes a road image, then using the target network model to perform image segmentation processing on the image to be processed includes:

[0037] Use the target network model to perform image segmentation processing on the road image to obtain a lane image, where the lane image includes lanes in the road image.

[0038] In a third aspect, an embodiment of the present invention provides a pruning device, including:

[0039] An information acquisition module, configured to acquire all network layers in the initial network model and the output channels corresponding to each network layer;

[0040] A first processing module, configured to, for each network layer, respectively determine the amplitudes corresponding to the output channels corresponding to the network layer, and prune the output channels corresponding to the network layer according to the amplitudes corresponding to the output channels to obtain the pruned network layer;

[0041] The first processing module is configured to obtain a target network model according to the pruned network layer.

[0042] In a possible design, the first processing module is further configured to:

[0043] Take all output channels corresponding to the network layer as objects to be processed, and determine the maximum variance corresponding to the network layer according to the amplitudes corresponding to the output channels in the objects to be processed;

[0044] Determine whether to perform pruning processing on the objects to be processed according to the maximum variance;

[0045] If pruning processing is to be performed on the objects to be processed, determine the output channels to be pruned from the objects to be processed, and perform pruning processing on the output channels to be pruned.

[0046] In a possible design, the first processing module is further configured to:

[0047] Obtain the number of all output channels corresponding to the network layer;

[0048] wherein determine the maximum variance corresponding to the network layer, where mσ 2 is the maximum variance corresponding to the network layer, x i is the amplitude corresponding to the i-th output channel, and N is the number of all output channels.

[0049] In a possible design, the first processing module is further configured to: after determining whether to perform pruning processing on the objects to be processed according to the maximum variance, if pruning processing is not to be performed on the objects to be processed, stop performing pruning processing on the objects to be processed;

[0050] In a possible design, the first processing module is further configured to: after performing pruning processing on the output channels to be pruned, take the objects to be processed after removing the output channels to be pruned as new objects to be processed, and return to the step of determining the maximum variance corresponding to the network layer according to the amplitudes corresponding to the output channels in the objects to be processed.

[0051] In a possible design, the first processing module is further configured to:

[0052] Obtain the number of all pruned output channels and obtain the number of all output channels corresponding to the network layer;

[0053] Calculate the ratio of the number of all pruned output channels to the number of all output channels corresponding to the network layer to obtain a pruning rate.

[0054] In a possible design, the first processing module is further configured to:

[0055] Determine whether the maximum value variance is greater than a preset variance value;

[0056] If it is greater than the preset variance value, determine to perform pruning processing on the object to be processed;

[0057] If it is less than or equal to the preset variance value, determine not to perform pruning processing on the object to be processed;

[0058] The first processing module is further configured to:

[0059] Obtain the output channel with the smallest amplitude and use it as the output channel to be pruned.

[0060] Train the network layer after the pruning processing to obtain the target network model.

[0061] Fourthly, an embodiment of the present invention provides a data processing device, including:

[0062] A data acquisition module, configured to acquire data to be processed;

[0063] A second processing module, configured to perform corresponding data processing on the data to be processed by using a target network model, where the target network model is obtained by acquiring all network layers in an initial network model and output channels corresponding to each network layer, for each network layer, respectively determining amplitudes corresponding to the output channels corresponding to the network layer, pruning the output channels corresponding to the network layer according to the amplitudes corresponding to the output channels, obtaining the network layer after the pruning processing, and obtaining according to the network layer after the pruning processing.

[0064] In a possible design, if the data to be processed includes an image to be processed, the second processing module is further configured to:

[0065] Perform image processing on the image to be processed by using a target network model, where the image processing includes target detection processing and / or image segmentation processing.

[0066] In a possible design, if the image to be processed includes a road image, the second processing module is further configured to:

[0067] Perform image segmentation processing on the road image by using the target network model to obtain a lane image, where the lane image includes lanes in the road image.

[0068] Fifthly, an embodiment of the present invention provides an electronic device, including: at least one processor and a memory;

[0069] The memory stores computer-executable instructions;

[0070] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the pruning method described in the first aspect above and various possible designs of the first aspect.

[0071] In a sixth aspect, an embodiment of the present invention provides an electronic device, including: at least one processor and a memory;

[0072] The memory stores computer-executable instructions;

[0073] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the data processing method described in the second aspect above and various possible designs of the second aspect.

[0074] In a seventh aspect, an embodiment of the present invention provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the pruning method described in the first aspect above and various possible designs of the first aspect is implemented.

[0075] In an eighth aspect, an embodiment of the present invention provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the data processing method described in the second aspect above and various possible designs of the second aspect is implemented.

[0076] The pruning method, data processing method and device provided by the present invention obtain the network layers in the initial network model to be pruned and the output channels corresponding to each network layer. For each network layer, the output channels of the network layer are pruned according to the amplitudes corresponding to each output channel of the network layer, that is, the network layer is pruned, and the target network model is determined according to the pruned network layer, so as to compress the network model, reduce the scale of the network model, and further reduce the amount of computation required when using the target neural network model for data processing, while also occupying less memory, thereby improving data processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0078] Figure 1Flow schematic of the pruning method provided by an embodiment of the present invention Figure 1 ;

[0079] Figure 2 Schematic diagram of the weight matrix provided by an embodiment of the present invention;

[0080] Figure 3 Flow schematic of the pruning method provided by an embodiment of the present invention Figure 2 ;

[0081] Figure 4 Flow schematic diagram of the data processing method provided by an embodiment of the present invention;

[0082] Figure 5 Schematic diagram of the road image provided by an embodiment of the present invention;

[0083] Figure 6 Schematic diagram of the structure of the pruning device provided by an embodiment of the present invention;

[0084] Figure 7 Schematic diagram of the structure of the data processing device provided by an embodiment of the present invention;

[0085] Figure 8 Schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0086] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0087] In the prior art, when using a neural network model to process data, the data is input into a pre-trained neural network model so that the neural network model performs corresponding data processing on the data. However, when using a neural network model to process data, due to the overly large scale of the neural network model, problems such as excessive computational amount and large memory occupation will occur, thereby affecting the efficiency of data processing.

[0088] Therefore, in view of the above problems, the technical concept of the present invention is to obtain a trained initial network model, and obtain all network layers in the initial network model and the output channels corresponding to each network layer. For each network layer, determine the amplitude corresponding to each output channel of the network layer, and determine the maximum variance corresponding to the network layer according to the amplitudes corresponding to each output channel. When the maximum variance is less than or equal to a preset variance value, it indicates that the amplitude distributions corresponding to the output channels are uniform, and there is no redundancy in the output channels of the network layer. If pruning is performed, it will affect the performance of the model, that is, affect the accuracy of the model for data processing, so pruning processing is not performed on this network layer. When the maximum variance is greater than the preset variance value, it indicates that the amplitude distributions corresponding to the output channels are non-uniform, and there is redundancy in the output channels of the network layer, so pruning processing is performed on this network layer to achieve compression of the model. When performing pruning processing on a network layer, first prune the output channel with the smallest amplitude, that is, remove the output channel with the smallest amplitude, and then continue to calculate the maximum variance corresponding to the network layer according to the amplitudes corresponding to the remaining output channels. If the maximum variance is still greater than or equal to the preset variance value, prune the remaining output channels according to the pruning rate to achieve accurate pruning of the network layer and avoid the situation of incomplete pruning or excessive pruning. After performing pruning processing on all network layers of the initial network model, train the initial network model to obtain a target network model. Using this target network model to perform corresponding data processing on the data to be processed, since the target network model is a network model after pruning processing, with fewer parameters and a smaller model size, when using this target network model for data processing, the amount of calculation required is smaller, and at the same time, the memory occupied is also smaller, improving the efficiency of data processing.

[0089] It can be understood that the pruning method provided in the embodiments of the present application can be applied to an image processing scenario to perform corresponding image processing. In addition, the pruning method provided in the embodiments of the present application can also be applied to other scenarios that require deep learning, such as semantic understanding of audio data. The embodiments of the present application do not make special restrictions on this.

[0090] The following uses specific examples to elaborate in detail on the technical solutions of the present disclosure and how the technical solutions of the present disclosure solve the above technical problems. These several specific examples can be combined with each other, and the same or similar concepts or processes may not be repeated in some examples. The following will describe the examples of the present disclosure with reference to the accompanying drawings.

[0091] Figure 1 Flow schematic of the pruning method provided in the embodiments of the present invention Figure 1 , the execution subject of this embodiment can be a first device, for example, an electronic device such as a terminal or a server. This embodiment does not make special restrictions here. As Figure 1 shown, the method includes:

[0092] S101. Obtain all network layers in the initial network model and the output channels corresponding to each network layer.

[0093] In this embodiment, the initial network model is composed of multiple network layers, and the network layer includes a convolutional layer and other types of network layers, such as a pooling layer. Each network layer includes corresponding input channels, output channels, and convolutional kernels.

[0094] Among them, the input channel is the channel for input feature data (e.g., feature map), and the output channel is the channel for output feature data (e.g., feature map).

[0095] Among them, the initial network model is a trained basic network model, and the initial network model includes trained parameters. However, since the initial network model is relatively large, that is, the time required for image processing is relatively large, and the occupied space is relatively large, and since only some parameters are involved in the main operation during image processing, redundant parameters of the initial network model can be pruned, that is, the output channels can be pruned.

[0096] S102. For each network layer, respectively determine the amplitudes corresponding to the output channels of the network layer, and perform pruning processing on the output channels corresponding to the network layer according to the amplitudes corresponding to the output channels, to obtain the network layer after pruning processing.

[0097] In this embodiment, for each network layer, respectively determine the amplitudes corresponding to the output channels in the network layer, that is, determine the amplitudes corresponding to each output channel respectively, and then perform pruning processing on the output channels corresponding to the network layer according to the amplitudes corresponding to each output channel respectively, to obtain the network layer after pruning processing.

[0098] In this embodiment, optionally, when determining the amplitudes corresponding to each output channel in the network layer, first determine the weight matrix of the network layer. The determined weight matrix of the network layer has a degree of Cin*Cout*kh*kw, where Cin is the number of input channels corresponding to the network layer, Cout is the number of output channels corresponding to the network layer, and kh and kw are the height and width of the convolutional kernel (kernel) respectively. Then sum the Cin, kh, and kw dimensions of the weight matrix to obtain an initial vector of 1*Cout, and the initial vector is {C 1 , …, C N}, where C 1 is the initial amplitude corresponding to the first output channel, C Nis the initial amplitude corresponding to the last output channel. Normalize the initial amplitudes in this initial vector, that is, obtain the largest initial amplitude in this initial vector. For each initial amplitude in this initial vector, divide this initial amplitude by this largest initial amplitude to obtain the target amplitude, that is, obtain the amplitude corresponding to the output channel corresponding to this initial amplitude, so as to obtain the target vector corresponding to this network layer, and this target vector is {x 1 , …, x N}.

[0099] In addition, optionally, the dimension of a convolution kernel is Cin*Kh*Kw*1, and there are Cout convolution kernels in the network layer. Combining them, the weight matrix of this network layer is Cin*Kh*Kw*Cout. This weight matrix has four axes, namely the Cin axis, the Cout axis, the kh axis, and the kw axis. When summing the Cin, kh, and kw dimensions of the weight matrix, that is, when calculating the initial amplitudes corresponding to each output channel, for each output channel, obtain all the input channels and convolution kernels corresponding to this output channel, and each input channel also corresponds to a convolution kernel. For each input channel, sum all the column elements corresponding to the kh axis in the convolution kernel corresponding to this input channel to eliminate the kh axis and obtain the column matrix corresponding to the kw axis, and then sum the column elements in the column matrix corresponding to the kw axis to eliminate the kw axis and obtain the column element corresponding to this input channel. Sum all the column elements corresponding to all the input channels corresponding to this output channel to eliminate the Cin axis and obtain the initial amplitude corresponding to this output channel.

[0100] Among them, the amplitude represents the importance degree of the convolution kernel. The larger the amplitude, the more important the convolution kernel is.

[0101] In addition, when calculating the column element corresponding to the input channel, first sum all the column elements corresponding to the kh axis in the convolution kernel corresponding to this input channel to eliminate the kh axis and obtain the column matrix corresponding to the kw axis, and then sum the column elements in the column matrix corresponding to the kw axis to eliminate the kw axis. Actually, sum all the weight parameter values in the convolution kernel corresponding to this input channel.

[0102] For example, see Figure 2 . The number of output channels in the network layer is 64, the number of input channels is 32, and the convolution kernel is 3*3. Then the weight matrix corresponding to this network layer is 32*64*3*3. One input channel corresponds to 32 input channels and 32 convolution kernels, and one input channel corresponds to one convolution kernel.

[0103] Optionally, pruning the output channels corresponding to the network layer according to the amplitudes corresponding to each output channel can be seen in Figure 3 as shown Figure 3Flow schematic of the pruning method provided by the embodiment of the present invention Figure 2 Based on the above embodiment, this embodiment elaborates on S202 in detail. The pruning process for the output channels may include:

[0104] S301. Take all the output channels corresponding to the network layer as the objects to be processed, and determine the maximum variance corresponding to the network layer according to the amplitudes corresponding to each output channel in the objects to be processed.

[0105] In this embodiment, when pruning a certain network layer, all the output channels corresponding to this network layer are taken as an object to be processed corresponding to this network layer, that is, the object to be processed includes at least one output channel. Calculate the maximum variance corresponding to this network layer according to the amplitudes corresponding to each output channel in the object to be processed, and take it as the maximum variance corresponding to this network layer.

[0106] Optionally, determining the maximum variance corresponding to the network layer according to the amplitudes corresponding to each output channel in the objects to be processed includes:

[0107] Obtain the number of all output channels corresponding to the network layer, and determine the maximum variance corresponding to the network layer through where mσ 2 is the maximum variance corresponding to the network layer, x i is the amplitude corresponding to the i-th output channel, and N is the number of all output channels, that is, N is the number of output channels included in the object to be processed.

[0108] S302. Determine whether to prune the object to be processed according to the maximum variance.

[0109] In this embodiment, after obtaining the maximum variance corresponding to a certain network layer, determine whether to prune the output channels of this network layer according to this maximum variance, that is, whether to prune the object to be processed.

[0110] In this embodiment, optionally, the implementation manner of S302 is: determine whether the maximum variance is greater than a preset variance value. If it is greater than the preset variance value, determine to prune the object to be processed. If it is less than or equal to the preset variance value, determine not to prune the object to be processed.

[0111] In this embodiment, compare the maximum variance corresponding to the network layer with the preset variance value. When it is determined that the maximum variance is less than or equal to the preset variance value, it indicates that the amplitude distribution of each channel of this network layer is relatively uniform, and there is no redundancy in this network layer. Pruning will damage the model performance, so pruning is not performed to avoid the impact, that is, the object to be processed corresponding to this network layer is not pruned.

[0112] When it is determined that the maximum variance is greater than the preset variance value, it indicates that the amplitude distribution of each channel in this network layer is relatively sparse and there is redundancy. It is necessary to perform pruning on the object to be processed corresponding to this network layer, that is, it is necessary to remove some of the output channels in the output channels corresponding to this network layer. Then it is determined to perform pruning on the output channels corresponding to this network layer, that is, it is determined that it is necessary to perform pruning on the object to be processed corresponding to this network layer.

[0113] Among them, the preset variance value is related to the compression degree of the model. When the preset variance value is smaller, the compression degree of the model is greater, and more output channels are cut off, but it may damage the performance of the model. When the preset variance value is larger, the compression degree of the model is relatively smaller, and the speed-up effect is relatively smaller. Therefore, relevant personnel can set the size of the preset variance value according to the size of the initial network model and the compression expectation.

[0114] S303. If pruning is not performed on the object to be processed, then stop performing pruning on the object to be processed.

[0115] In this embodiment, when it is determined that pruning is not performed on the object to be processed, it indicates that there is no redundancy in the network layer corresponding to this object to be processed. Then stop performing pruning on the object to be processed corresponding to this network layer, that is, stop pruning the output channels corresponding to this network layer.

[0116] S304. If pruning is performed on the object to be processed, then determine the output channels to be pruned from the object to be processed, and perform pruning on the output channels to be pruned.

[0117] In this embodiment, when it is determined that pruning is performed on the object to be processed, it indicates that there is still redundancy in the network layer corresponding to this object to be processed. Then determine the output channels to be removed from the object to be processed corresponding to this network layer, that is, determine the output channels to be pruned from the output channels included in this object to be processed. Then perform pruning on the output channels to be pruned, that is, remove the output channels to be pruned.

[0118] In addition, optionally, when determining the output channels to be pruned from the object to be processed, the output channel with the smallest amplitude in the object to be processed can be obtained and used as the output channel to be pruned. For example, if the amplitude corresponding to output channel 1 is 0.5, the amplitude corresponding to output channel 2 is 0.7, and the amplitude corresponding to output channel 3 is 1, then since the amplitude corresponding to output channel 1 is the smallest, output channel 1 is used as the output channel to be pruned.

[0119] In addition, optionally, after performing pruning on the network channels corresponding to this network layer, determine whether the remaining output channels corresponding to this network layer need to continue to be pruned. Then use the object to be processed after removing the output channels to be pruned as the new object to be processed, and return to the step of determining the maximum variance corresponding to the network layer according to the amplitudes corresponding to the output channels in the object to be processed.

[0120] Specifically, after performing a pruning process on the object to be processed corresponding to the network layer, the object to be processed excluding the output channels to be pruned is used as the new object to be processed. Then, based on the amplitudes corresponding to the output channels in the new object to be processed, it is determined whether pruning needs to be performed on the new object to be processed. For example, after pruning the output channels to be pruned, that is, after removing the output channels to be pruned from all the output channels corresponding to the network layer, the remaining output channels are used as the new object to be processed corresponding to the network layer. The maximum variance corresponding to the network layer is calculated based on the new object to be processed to determine whether pruning still needs to be performed on the network layer, that is, by comparing the maximum variance with a preset variance value. If the maximum variance is greater than the preset variance value, it indicates that pruning can still be performed on the network layer, that is, continue to remove some channels from the new object to be processed. Then, the output channels to be pruned are determined from the new object to be processed, and pruning is performed on the output channels to be pruned, that is, the output channels to be pruned are removed. If the maximum variance is less than or equal to the preset variance value, it indicates that there is no redundancy in the network layer, and pruning is not performed on the network layer, that is, no further channels are removed from the new object to be processed corresponding to the network layer.

[0121] In addition, optionally, after performing a pruning process on the output channels corresponding to the network layer, the current number of pruned output channels is updated, that is, the current number of pruned output channels is incremented by the number of output channels to be pruned this time. The initial value of the number of pruned output channels is 0.

[0122] Taking a specific application scenario as an example, the number of all output channels corresponding to a certain network layer is 10. That is, 10 output channels are used as the objects to be processed corresponding to this network layer. After determining to perform pruning on the objects to be processed corresponding to this network layer, the number of output channels to be pruned is determined to be 1. Then, this output channel to be pruned is removed, and the remaining 9 output channels are used as the new objects to be processed corresponding to this network layer. And the current number of pruned output channels is updated, that is, the current number of pruned output channels (which is 0) is incremented by 1, and the updated current number of pruned output channels is 1. After determining to perform pruning on this new object to be processed, the number of output channels to be pruned is determined to be 1. Then, this output channel to be pruned is removed, and the remaining 8 output channels are used as the new objects to be processed corresponding to this network layer. And the number of pruned output channels is updated, that is, the current number of pruned output channels (which is 1) is incremented by 1, and the updated current number of pruned output channels is 2. After determining that no pruning needs to be performed on this new object to be processed according to the amplitudes of the 8 output channels, the pruning of this new object to be processed is stopped. Then, the number of output channels included in the final object to be processed is 8, and the current number of pruned output channels is 2. That is, all the output channels to be pruned corresponding to this network layer are 2, and a total of 2 output channels are pruned.

[0123] In addition, optionally, obtain the number of all pruned output channels and the number of all output channels corresponding to the network layer, and calculate the ratio of the number of all pruned output channels to the number of all output channels corresponding to the network layer to obtain the pruning rate.

[0124] In this embodiment, obtain the number of all currently pruned output channels of the network layer. The number of all pruned output channels is the current number of pruned output channels, that is, the number of all output channels corresponding to the network layer minus the number of all output channels to be pruned corresponding to the network layer. Divide the number of all pruned output channels by the number of all output channels corresponding to the network layer to obtain the pruning rate corresponding to this network layer, so that the user can know the pruning situation of this network layer.

[0125] Continuing with the above application scenario, the number of all pruned output channels corresponding to the network layer is 2, and the number of all output channels corresponding to the network layer is 10. Then, the pruning rate corresponding to this network layer is 2 / 10 = 0.2.

[0126] S103. Obtain the target network model according to the network layer after pruning.

[0127] In this embodiment, after performing pruning on this network layer, that is, after obtaining the network layer after pruning, determine the corresponding target network model according to the network layer after pruning, so as to be able to process the data to be processed using this target network model.

[0128] In this embodiment, optionally, after pruning each network layer in the initial network model, that is, after obtaining the pruned network layers, each pruned network layer is trained, that is, the initial network model after pruning is trained. The initial network model after pruning is composed of the pruned network layers, and a target network model is obtained to restore the accuracy. The accuracy of the target network model can meet the requirements, so that the target network model can be used for data processing.

[0129] In this embodiment, when pruning a network layer, first prune the output channel with the smallest amplitude, that is, remove the output channel with the smallest amplitude, and then continue to calculate the maximum variance corresponding to this network layer according to the amplitudes corresponding to the remaining output channels. If the maximum variance is still greater than or equal to the preset variance value, prune the remaining output channels according to the pruning rate to achieve accurate pruning of the network layer and avoid the situation of incomplete pruning or excessive pruning.

[0130] In the prior art, when pruning a network model, the amplitudes corresponding to the output channels of each network layer are sorted in descending order, and then the N channels with smaller amplitudes are removed according to the preset pruning ratio to achieve the pruning of the network model. However, when pruning according to the sorted amplitudes, if the amplitudes corresponding to the network layer are not very different, removing the last N channels in this way will greatly damage the model's expression ability, and there will be a large loss in the performance of the pruned model, resulting in too large a loss in the accuracy of the pruned network model.

[0131] In this application, the amplitude distribution of each layer is determined based on the maximum variance. The maximum variance of a network layer reflects the amplitude distribution of that layer, and the optimal structure of the network layer is adaptively determined according to the relationship between the maximum variance corresponding to the network layer and the preset variance value. This avoids the situation where the pruned network model cannot achieve the optimal speed due to insufficient pruning or the accuracy of the pruned network model is greatly lost due to excessive pruning because only sorting the amplitudes of each layer's parameters in descending order and removing the channels with smaller amplitudes according to the preset pruning ratio. Thus, the finally obtained target network model achieves the optimal performance and speed, and further improves the efficiency and accuracy of image processing.

[0132] As can be seen from the above description, the network layers in the initial network model that need to be pruned and the output channels corresponding to each network layer are obtained. For each network layer, the output channels of the network layer are pruned according to the amplitudes corresponding to each output channel of the network layer, that is, the network layer is pruned, and the target network model is determined according to the pruned network layer to achieve the compression of the network model, so as to reduce the scale of the network model. Furthermore, when using the target neural network model for data processing, the required amount of calculation is smaller, and the occupied memory is also smaller, improving the data processing efficiency.

[0133] Figure 4 This is a schematic flowchart of the data processing method provided by the embodiments of the present invention. The execution subject of this embodiment can be a second device, such as an electronic device like a terminal, a server, etc. The second device can be the same device as the above-mentioned first device or a different device, and no special limitation is made here in this embodiment. As Figure 4 shown, the method includes:

[0134] S401. Obtain the data to be processed.

[0135] In this embodiment, the data to be processed is the data that needs to be processed. The data to be processed can be data sent by other terminals or servers, or data imported by the user into the electronic device through an importing device (such as a USB flash drive), or data downloaded by the electronic device from a relevant location (such as a website). Here, the source of the data to be processed is not limited.

[0136] S402. Use the target network model to perform corresponding data processing on the data to be processed, where the target network model is obtained by acquiring all network layers in the initial network model and the output channels corresponding to each network layer. For each network layer, the amplitudes corresponding to the output channels of the network layer are respectively determined, and the output channels of the network layer are pruned according to the amplitudes corresponding to the output channels to obtain the pruned network layer, and the target network model is obtained based on the pruned network layer.

[0137] In this embodiment, after obtaining the data to be processed, it indicates that data processing needs to be performed on the data to be processed. Then, the target network model after pruning is used to process the data, that is, the data to be processed is input into the target network model so that the target network model performs corresponding data processing on the data and generates the required data processing result.

[0138] Among them, when the data to be processed includes an image to be processed, the implementation manner of S402 is:

[0139] Use the target network model to perform image processing on the image to be processed, where the image processing includes object detection processing and / or image segmentation processing.

[0140] In this embodiment, after obtaining the image to be processed, it indicates that image processing needs to be performed on the image to be processed. Then, the target network model after pruning is used to process the image, that is, the image to be processed is input into the target network model so that the target network model performs corresponding image processing on the image and generates the required image processing result, and the image processing includes object detection processing and / or image segmentation processing.

[0141] Since the target network model is a network model after pruning, that is, a compressed network model, using the compressed network model to perform image processing on an image can improve the speed of image processing on the basis of ensuring the accuracy of image processing, thereby improving the efficiency of image processing.

[0142] Optionally, the image to be processed includes a road image, an electronic map image, or other types of images, such as an image in front of a vehicle during driving. Here, the type of the image to be processed is not limited.

[0143] Optionally, when the image to be processed includes a road image and the image processing includes image segmentation processing, the target network model is used to perform image segmentation processing on the road image to obtain a lane image, where the lane image includes lanes in the road image.

[0144] Specifically, when a lane needs to be obtained, the road image is input into the target network model, and the target network model performs image segmentation on the road image to obtain a lane image including lanes in the road image. The lane image does not include background information in the road image, such as pedestrians on the road (see Figure 5 ), trees, etc., to achieve fast and accurate lane segmentation.

[0145] Optionally, when the image processing includes target detection processing, the target network model is used to perform target detection processing on the road image to obtain a target detection result, where the target detection result is that there is a target object in the image or there is no target object in the image.

[0146] Specifically, during the automatic driving of a vehicle, the road image in front of the vehicle during driving can be input into the target network model, so that the target network performs target detection processing on the road image. For example, it detects whether there are pedestrians in the road image and outputs a corresponding target detection result. When it is detected that there are pedestrians in the road image, the target detection result is that there is a target object in the image, that is, there are pedestrians. When it is detected that there are no pedestrians in the road image, the target detection result is that there is no target object in the image, that is, there are no pedestrians. At the same time, the vehicle can also be guided to avoid pedestrians according to the target detection result.

[0147] It can be understood that the above use of the target network model to perform target detection processing and image segmentation processing on the road image is only an example, and the target network model can process any type of image, and this application is not limited thereto.

[0148] In this embodiment, when performing image processing on the image to be processed, a target network model that has been pruned, that is, a compressed target network model, is used to perform corresponding image processing on the image. Since the target network model is a compressed target network model and the scale of the target network model is small, when using this neural network model for image processing, the amount of computation required is small, and the memory occupied is also small, improving the image processing efficiency, so that the problem of low existing image processing efficiency will not occur.

[0149] Figure 6 FIG. is a schematic structural diagram of a pruning device provided by an embodiment of the present invention. As Figure 6 shown, the image pruning device 60 includes: an information acquisition module 601 and a first processing module 602.

[0150] Among them, the information acquisition module 601 is used to acquire all network layers in the initial network model and the output channels corresponding to each network layer.

[0151] The first processing module 602 is used to, for each network layer, respectively determine the amplitudes corresponding to the output channels corresponding to the network layer, and perform pruning processing on the output channels corresponding to the network layer according to the amplitudes corresponding to the output channels, so as to obtain the network layer after pruning processing.

[0152] The first processing module 602 is used to obtain a target network model according to the network layer after pruning processing.

[0153] In a possible design, the first processing module 602 is further used to:

[0154] Take all the output channels corresponding to the network layer as the objects to be processed, and determine the maximum variance corresponding to the network layer according to the amplitudes corresponding to the output channels in the objects to be processed.

[0155] Determine whether to perform pruning processing on the objects to be processed according to the maximum variance.

[0156] If pruning processing is to be performed on the objects to be processed, determine the output channels to be pruned from the objects to be processed, and perform pruning processing on the output channels to be pruned.

[0157] In a possible design, the first processing module 602 is further used to:

[0158] Obtain the number of all output channels corresponding to the network layer.

[0159] By determine the maximum variance corresponding to the network layer, where mσ 2 is the maximum variance corresponding to the network layer, x i is the amplitude corresponding to the i-th output channel, and N is the number of all output channels.

[0160] In a possible design, the first processing module 602 is further configured to: after determining whether to perform pruning on the object to be processed according to the maximum value variance, if pruning is not performed on the object to be processed, stop performing pruning on the object to be processed.

[0161] In a possible design, the first processing module 602 is further configured to: after performing pruning on the output channels to be pruned, use the object to be processed after removing the output channels to be pruned as a new object to be processed, and return to the step of determining the maximum value variance corresponding to the network layer according to the amplitudes corresponding to the output channels in the object to be processed.

[0162] In a possible design, the first processing module 602 is further configured to:

[0163] Obtain the number of all pruned output channels and obtain the number of all output channels corresponding to the network layer.

[0164] Calculate the number of all pruned output channels and the number of all output channels corresponding to the network layer to obtain a pruning rate.

[0165] In a possible design, the first processing module 602 is further configured to:

[0166] Determine whether the maximum value variance is greater than a preset variance value.

[0167] If it is greater than the preset variance value, determine to perform pruning on the object to be processed.

[0168] If it is less than or equal to the preset variance value, determine not to perform pruning on the object to be processed.

[0169] The first processing module 602 is further configured to:

[0170] Obtain the output channel with the minimum amplitude and use it as the output channel to be pruned.

[0171] Train the network layer after pruning to obtain a target network model.

[0172] The device provided in this embodiment can be used to execute the technical solutions of the above pruning method embodiments, and its implementation principles and technical effects are similar, which will not be elaborated here in this embodiment.

[0173] Figure 7 It is a schematic structural diagram of a data processing device provided in an embodiment of the present invention. As Figure 7 shown, the data processing device 70 includes: a data acquisition module 701 and a second processing module 702.

[0174] The data acquisition module 701 is configured to acquire data to be processed.

[0175] A second processing module 702, configured to perform corresponding data processing on the data to be processed by using a target network model, where the target network model is obtained by acquiring all network layers in an initial network model and output channels corresponding to each network layer, determining, for each network layer, amplitudes corresponding to the output channels corresponding to the network layer, pruning the output channels corresponding to the network layer according to the amplitudes corresponding to the output channels to obtain a pruned network layer, and obtaining the target network model according to the pruned network layer.

[0176] In a possible design, if the data to be processed includes an image to be processed, the second processing module 702 is further configured to:

[0177] Perform image processing on the image to be processed by using the target network model, where the image processing includes object detection processing and / or image segmentation processing.

[0178] In a possible design, if the image to be processed includes a road image, the second processing module 702 is further configured to:

[0179] Perform image segmentation processing on the road image by using the target network model to obtain a lane image, where the lane image includes lanes in the road image.

[0180] The device provided in this embodiment can be used to execute the technical solutions of the above data processing method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.

[0181] Figure 8 It is a schematic hardware structure diagram of an electronic device provided in an embodiment of the present invention. As Figure 8 shown, the electronic device 80 in this embodiment includes: a processor 801 and a memory 802;

[0182] Among them, the memory 802 is used to store computer execution instructions;

[0183] The processor 801 is configured to execute the computer execution instructions stored in the memory to implement each step executed by the receiving device in the above embodiment. For details, please refer to the relevant descriptions in the foregoing method embodiment.

[0184] Optionally, the memory 802 can be either independent or integrated with the processor 801.

[0185] When the memory 802 is independently provided, the electronic device further includes a bus 803 for connecting the memory 802 and the processor 801.

[0186] An embodiment of the present invention also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the pruning method and / or data processing method described above are implemented.

[0187] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in an electrical, mechanical or other form.

[0188] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0189] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a unit. The units formed by the above modules can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0190] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods described in various embodiments of the present application.

[0191] It should be understood that the above-mentioned processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed and completed by a hardware processor, or can be executed and completed by a combination of hardware and software modules in the processor.

[0192] The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.

[0193] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0194] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0195] An exemplary storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a master control device.

[0196] Those of ordinary skill in the art will understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments; and the foregoing storage medium includes various media that can store program codes such as ROM, RAM, magnetic disks, or optical discs.

[0197] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A pruning method, characterized in that, it includes: Obtain all network layers in the initial network model and the output channels corresponding to each network layer; wherein, the output channels are the channels for outputting feature data; For each network layer, determine the weight matrix of the network layer, sum the number of input channels, the height of the convolution kernel, and the width dimension of the convolution kernel in the weight matrix to obtain an initial vector, normalize the initial amplitude in the initial vector, obtain the largest initial amplitude in the initial vector, divide the initial amplitude by the target amplitude to obtain the target amplitude, use the target amplitude as the amplitude corresponding to each output channel corresponding to the network layer, and perform pruning processing on the output channels corresponding to the network layer according to the amplitudes corresponding to each output channel to obtain the network layer after pruning processing; Obtain the target network model according to the network layer after pruning processing, wherein the target network model is used for target detection processing and / or image segmentation processing of the image to be processed; The performing pruning processing on the output channels corresponding to the network layer according to the amplitudes corresponding to each output channel includes: Take all output channels corresponding to the network layer as the objects to be processed, and determine the maximum value variance corresponding to the network layer according to the amplitudes corresponding to each output channel in the objects to be processed; Determine whether to perform pruning processing on the objects to be processed according to the maximum value variance; If pruning processing is to be performed on the objects to be processed, determine the output channels to be pruned from the objects to be processed, and perform pruning processing on the output channels to be pruned.

2. The method according to claim 1, characterized in that, Determining the maximum value variance corresponding to the network layer according to the amplitudes corresponding to each output channel in the objects to be processed includes: Obtain the number of all output channels corresponding to the network layer; By determining the maximum variance corresponding to the network layer, where the is the maximum variance corresponding to the network layer, and the is the amplitude corresponding to the i-th output channel, and N is the number of all output channels.

3. The method according to claim 1, characterized in that, After determining whether to perform pruning processing on the objects to be processed according to the maximum value variance, it further includes: If pruning processing is not to be performed on the objects to be processed, stop performing pruning processing on the objects to be processed; After performing pruning processing on the output channels to be pruned, it further includes: Take the objects to be processed after removing the output channels to be pruned as the new objects to be processed, and return to the step of determining the maximum value variance corresponding to the network layer according to the amplitudes corresponding to each output channel in the objects to be processed.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Obtain the number of all pruned output channels and obtain the number of all output channels corresponding to the network layer; Calculate the number of all pruned output channels and the number of all output channels corresponding to the network layer to obtain the pruning rate.

5. The method according to claim 1, characterized in that, Determining whether to perform pruning processing on the objects to be processed according to the maximum value variance includes: Determine whether the maximum value variance is greater than the preset variance value; If it is greater than the preset variance value, determine that pruning processing is to be performed on the objects to be processed; If it is less than or equal to the preset variance value, it is determined that no pruning process is performed on the object to be processed; The determining the output channels to be pruned from the object to be processed includes: Obtaining the output channel with the minimum amplitude in the object to be processed and using it as the output channel to be pruned.

6. A data processing method, Characterized in that, It includes: Obtaining the data to be processed; Using a target network model to perform corresponding data processing on the data to be processed, where the target network model is obtained by acquiring all network layers in the initial network model and the output channels corresponding to each network layer. For each network layer, determining the weight matrix of the network layer, summing the number of input channels, the height of the convolutional kernel, and the width dimension of the convolutional kernel in the weight matrix to obtain an initial vector, normalizing the initial amplitude in the initial vector, obtaining the maximum initial amplitude in the initial vector, dividing the initial amplitude by the target amplitude to obtain the target amplitude, using the target amplitude as the amplitude corresponding to each output channel corresponding to the network layer, and performing pruning processing on the output channels corresponding to the network layer according to the amplitudes corresponding to each output channel to obtain the network layer after pruning processing, and obtaining according to the network layer after pruning processing; where the output channel is the channel for outputting feature data; If the data to be processed includes an image to be processed, then using the target network model to perform corresponding data processing on the data to be processed includes: Using the target network model to perform image processing on the image to be processed, where the image processing includes target detection processing and / or image segmentation processing; The performing pruning processing on the output channels corresponding to the network layer according to the amplitudes corresponding to each output channel includes: Regarding all output channels corresponding to the network layer as the object to be processed, and determining the maximum variance of the network layer according to the amplitudes corresponding to each output channel in the object to be processed; Determining whether to perform pruning processing on the object to be processed according to the maximum variance; If pruning processing is performed on the object to be processed, determining the output channels to be pruned from the object to be processed and performing pruning processing on the output channels to be pruned.

7. A pruning device, Characterized in that, It includes: An information acquisition module for acquiring all network layers in the initial network model and the output channels corresponding to each network layer; where the output channel is the channel for outputting feature data; A first processing module for, for each network layer, determining the weight matrix of the network layer, summing the number of input channels, the height of the convolutional kernel, and the width dimension of the convolutional kernel in the weight matrix to obtain an initial vector, normalizing the initial amplitude in the initial vector, obtaining the maximum initial amplitude in the initial vector, dividing the initial amplitude by the target amplitude to obtain the target amplitude, using the target amplitude as the amplitude corresponding to each output channel corresponding to the network layer, and performing pruning processing on the output channels corresponding to the network layer according to the amplitudes corresponding to each output channel to obtain the network layer after pruning processing; The first processing module is further configured to obtain a target network model according to the pruned network layer, where the target network model is used to perform object detection processing and / or image segmentation processing on an image to be processed; The first processing module is specifically configured to use all output channels corresponding to the network layer as an object to be processed, and determine the maximum variance corresponding to the network layer according to the amplitudes corresponding to the output channels in the object to be processed; Determine whether to perform pruning processing on the object to be processed according to the maximum variance; If pruning processing is performed on the object to be processed, determine the output channels to be pruned from the object to be processed, and perform pruning processing on the output channels to be pruned.

8. A data processing device, characterized in that, it includes: a data acquisition module, configured to acquire data to be processed; a second processing module, configured to use a target network model to perform corresponding data processing on the data to be processed, where the target network model is obtained by acquiring all network layers in an initial network model and output channels corresponding to each network layer. For each network layer, determine the weight matrix of the network layer, sum the number of input channels, the height of the convolution kernel, and the width dimension of the convolution kernel in the weight matrix to obtain an initial vector, normalize the initial amplitude in the initial vector, obtain the maximum initial amplitude in the initial vector, divide the initial amplitude by a target amplitude to obtain a target amplitude, use the target amplitude as the amplitude corresponding to each output channel corresponding to the network layer, and perform pruning processing on the output channels corresponding to the network layer according to the amplitudes corresponding to the output channels, obtain the pruned network layer, and is obtained according to the pruned network layer, where the output channel is a channel for outputting feature data; The data to be processed includes an image to be processed, and the second processing module is configured to use a target network model to perform image processing on the image to be processed, where the image processing includes object detection processing and / or image segmentation processing; The second processing module is specifically configured to: use all output channels corresponding to the network layer as an object to be processed, and determine the maximum variance corresponding to the network layer according to the amplitudes corresponding to the output channels in the object to be processed; Determine whether to perform pruning processing on the object to be processed according to the maximum variance; If pruning processing is performed on the object to be processed, determine the output channels to be pruned from the object to be processed, and perform pruning processing on the output channels to be pruned.

Citation Information

Patent Citations

  • Convolutional neural network trimming method and device, and storage medium

    CN110232436A