Digital media processing method and system based on improved channel pruning algorithm
The channel pruning algorithm is improved through local sensitive hashing algorithms, and the pruning input and output channels are synchronized, which solves the problem of insufficient correlation in the existing technology and realizes efficient digital media processing.
Patent Information
- Application Number
- CN202510726218.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The channel pruning method of existing convolutional neural networks fails to effectively maintain the correlation between input and output channels, resulting in negative impact on model performance and high computational complexity.
The locally sensitive hashing algorithm is used to improve the channel pruning algorithm, and synchronize the pruning input and output channels through the block division, block selection and fine-tuning stages to maintain correlation and expand the search space, and reduce the computational burden.
It significantly improves the accuracy and efficiency of digital media processing, reduces memory usage, reduces pruning time and negative impact on model performance.
Smart Images

Figure CN120258076A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital media processing, and in particular to a digital media processing method and system based on an improved channel pruning algorithm. Background Art
[0002] In the prior art, the algorithms of digital media processors are too large and the computational complexity is relatively high. Therefore, to reduce unnecessary computations and improve computational efficiency, redundant parameters can be removed through pruning to avoid processing a large amount of useless information and achieve the optimization of digital media processors.
[0003] The general channel pruning of the prior art's Convolutional Neural Networks (CNN) will first consider a pre-trained model based on a convolutional neural network, which consists of L convolutional layers and some fully connected layers. Represent all the full subsets of the layers as , is the four-dimensional weight tensor of the th convolutional layer, R is the real tensor space, is the number of convolutional layers, is the number of output channels, which is also the number of filters in the current layer, , , are respectively the number of input channels, filter height, and filter width of the th layer. For the fully connected layer, the weight tensor can be generalized to a 1 × 1 convolution, where and . Here, we mainly focus on the convolutional layer.
[0004] Generally speaking, the purpose of channel pruning is to create a lightweight model by deleting redundant channels. In this process, the representation of all the full subsets of the layers is , where , is the four-dimensional weight tensor after pruning, is the number of output channels after pruning, is the number of output channels after pruning, where and . Since this application only focuses on channel pruning, the weight tensor can be generalized to a 1 × 1 convolution, where and So is redefined as , is the original weight matrix. Therefore, each element in is a filter channel with the shape of a , denotes the index of the convolutional layer (the -th layer), j denotes the index of the current convolutional kernel (the j-th convolutional kernel), k denotes the index of the input channel (the k -th input channel), :,: denotes the height and width of the convolutional kernel. The -th layer's j -th convolutional filter can be expressed as . For the -th layer, the target channel pruning task can be formulated as: ; where r is the expected pruning rate, is the loss function, D is the dataset, is a binary mask vector containing only 0 or 1, indicating whether the corresponding filter is pruned, is taking the minimum value in the set, is the j-th output channel of the l-th layer, is the convolutional operation, is the zero norm, representing the number of non-zero elements, is the multiplication between elements, is under what conditions. The above goal is to find a suitable mask vector where the proportion of zero elements is equal to the pruning rate while minimizing the impact on model performance. Obviously, pruning mainly targets the output channels of the current layer, and its input channels are adjusted according to the pruning results of the output channels of the previous layer.
[0005] However, this method of only pruning the output channels cannot maintain the correlation between the input and output channels. Secondly, this pruning method also limits the selection of channels to the output channels, which has a great negative impact on model performance. SUMMARY OF THE INVENTION
[0006] To solve the above problems, the present invention proposes a digital media processing method and system based on an improved channel pruning algorithm. The local sensitive hashing pruning algorithm helps to maintain the correlation between the input and output channels, expand the candidate space for channel selection, optimize the digital media processing method, and improve the accuracy and efficiency of digital media processing.
[0007] A digital media processing method based on an improved channel pruning algorithm provided by the present invention includes the following steps: S1. Obtain the digital media data to be processed, including image data, video data, and audio data; S2. Process the digital media data to be processed through an improved CNN convolutional neural network to obtain target information; the improved CNN convolutional neural network is improved based on a channel pruning algorithm of locality-sensitive hashing, and the channel pruning algorithm of locality-sensitive hashing includes: a block division stage, a block selection stage, and a fine-tuning stage.
[0008] Preferably, the block division stage includes the following steps: S211. Obtain the original weight matrix of the filter in the pre-trained convolutional neural network and a predetermined block size; the expression of the original weight matrix is: ; Where is the original weight matrix, R is the real number tensor space, is the current layer number, is the input channel number of the th layer, S212. Divide the original weight matrix of the model of the pre-trained convolutional neural network according to the predetermined block size to obtain target blocks; the expression of the target blocks is: ; Where is the index in the row direction, is the predetermined block size, is the index in the column direction.
[0009] Preferably, the block selection stage includes the selection of output channels and the selection of input channels; after the block selection stage performs the selection of output channels and the selection of input channels on the target blocks, a pruned convolutional neural network model is obtained.
[0010] Preferably, the selection of output channels includes the following steps: S2211. Define each target block as: ; Where is a set formed by splicing multiple vector blocks for output channels, is the operation of converting a matrix into a vector, is the connection of vectors, is the dimensional parameter of the th layer, is the filter height, S2212. Define b hash values and introduce a fixed random matrix , define a locality - sensitive hashing function ; The locality - sensitive hashing function has the following expression: ; wherein, is the input vector, is to take the maximum value from the set; S2213. Through the locality - sensitive hashing function perform hash mapping on each target block so that each target block obtains a hash value; S2214. Randomly select one and retain its corresponding output channel, while deleting the channels corresponding to other channels.
[0011] Preferably, in S2214, randomly select one and retain its corresponding output channel, while deleting the channels corresponding to other channels, specifically including: S22141. Determine similar blocks according to the hash value; the hash values of the similar blocks are the same; S22142. Retain one similar block and the output channel corresponding to the similar block, delete the redundant blocks and the output channels corresponding to the redundant blocks, to obtain a convolutional neural network model with pruned output channels; the redundant blocks have the same hash value as the similar blocks.
[0012] Preferably, the selection of the input channels includes the following steps: S2221. Define each target block as: ; wherein, is a set formed by splicing multiple vector blocks of the input channels; S2222. Define b hash values and introduce a fixed random matrix , define a locality - sensitive hashing function ; S2223. Through the locality - sensitive hashing function perform hash mapping on each target block so that each target block obtains a hash value; S2224. Randomly select one and retain its corresponding output channel, while deleting the channels corresponding to other channels.
[0013] Preferably, the fine - tuning stage includes: fine - tuning the pruned model based on the convolutional neural network to obtain a compact convolutional neural network model.
[0014] A digital media processing system based on an improved channel pruning algorithm, comprising: An acquisition module for acquiring digital media data to be processed, including image data, video data, and audio data; A digital media processing module for processing the digital media data to be processed through a channel pruning algorithm based on locality-sensitive hashing to obtain processed target information; the channel pruning algorithm based on locality-sensitive hashing includes: a block partitioning stage, a block selection stage, and a fine-tuning stage.
[0015] An electronic device includes a memory and a processor. A computer program is stored in the memory. When the processor calls the computer program in the memory, the content of the digital media processing method based on the improved channel pruning algorithm as described above is implemented.
[0016] A storage medium stores computer-executable instructions. When the computer-executable instructions are loaded and executed by a processor, the content of the digital media processing method based on the improved channel pruning algorithm as described above is implemented.
[0017] In summary, for a digital media processing method and system based on an improved channel pruning algorithm of the present invention, compared with traditional technologies, the solution of the present invention has the following advantages: 1. The present invention improves the pruning algorithm based on the locality-sensitive hashing algorithm, introduces a block partitioning strategy into structured pruning, significantly expands the search space of pruning, and effectively reduces the impact of pruning on the model performance, which helps to maintain the correlation between input and output channels and expand the candidate space for channel selection; 2. The present invention does not rely on an iterative algorithm to solve the sparse optimization problem, which not only reduces the computational burden but also reduces the pruning time, and further minimizes the negative impact of pruning on the model performance; 3. The present invention improves the channel pruning algorithm through a locality-sensitive hashing function, realizes the optimization of the digital media processing method, reduces the memory occupied by the processing program, and improves the speed of digital media processing.
[0018] Next, through the drawings and embodiments, the technical method of the present invention will be further described in detail. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic flowchart of the channel pruning algorithm based on locality-sensitive hashing of the present invention; Figure 2 It is a schematic flowchart of the channel pruning algorithm based on locality-sensitive hashing acting on the original weight matrix of a pre-trained convolutional neural network model of. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The technical method of the present invention will be further described below with reference to the accompanying drawings and embodiments. It should be noted that: unless otherwise specifically stated, the relative arrangements, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present application.
[0021] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way a limitation on the present application, its application or use.
[0022] Technologies, systems and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, systems and devices should be regarded as part of the specification.
[0023] In all the examples shown and discussed here, any specific values should be construed as merely exemplary and not as a limitation. Accordingly, other examples of the exemplary embodiments may have different values.
[0024] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention pertains.
[0025] Convolutional neural networks have shown excellent performance in computer vision tasks, but their high computational and memory requirements limit their deployment in the Internet of Things and edge devices. Model compression techniques (such as quantization, pruning, and knowledge distillation) that generate regularized and compact models from structural units (such as filters, channels) have become a research hotspot in the field of model compression. Their core advantages are reflected in hardware compatibility (the pruned model can be directly adapted to standard hardware and computing libraries), computational efficiency (significantly reducing storage and computational amounts), and technical synergy (combining with quantization and distillation methods can improve the compression effect). There are three mainstream methods for structured pruning: one is attribute-based pruning, which is divided into weight-dependent and activation-dependent. Weight-dependent pruning evaluates importance through filter norms (the sum of the absolute values of each element) or correlation analysis, while activation-dependent pruning uses the reconstruction error of the channel activation map to screen redundant filters. This method that requires manual setting of the pruning rate and iterative fine-tuning is time-consuming and highly dependent on experts. The second is regularization-based pruning, which automatically identifies redundant parameters by introducing sparse constraints, but requires retraining from scratch and fine-tuning, resulting in a high computational cost. The third is neural architecture search-based pruning, which takes the pruning rate search scale as an optimization problem to reduce manual intervention, but the huge search space is likely to lead to a decline in efficiency.
[0026] Therefore, the present invention provides a digital media processing method and system based on an improved channel pruning algorithm, which improves the channel pruning algorithm through a locality-sensitive hashing function to optimize the digital media processing method and improve the accuracy and efficiency of digital media processing.
[0027] A digital media processing method based on an improved channel pruning algorithm provided by the present invention includes the following steps: S1. Obtain the digital media data to be processed, which includes image data, video data, and audio data.
[0028] In CNN, most existing channel pruning methods focus on either the input or output channels, and the unpruned channels are adjusted according to the pruning results of adjacent layers. The channel pruning algorithm based on locality sensitive hashing proposed in this application is a novel collaborative channel pruning method that synchronously prunes the input and output channels in the same neural network layer.
[0029] S2. Process the digital media data to be processed through an improved CNN convolutional neural network to obtain target information.
[0030] The improved CNN convolutional neural network is improved based on the channel pruning algorithm of locality sensitive hashing. Different from the traditional method, the method of this application introduces a block partitioning strategy into structured pruning, significantly expanding the search space of pruning and effectively reducing the impact of pruning on the model performance.
[0031] The channel pruning algorithm based on locality sensitive hashing includes: a block partitioning stage, a block selection stage, and a fine-tuning stage.
[0032] The channel pruning algorithm based on locality sensitive hashing (Locality Sensitive Hashing-Pruner, LSH-Pruner) is a more efficient and faster pruning algorithm, including a block partitioning stage, a block selection stage, and a fine-tuning stage. The overall process is as Figure 1 shown. First, feature vectors are extracted from the original convolutional neural network-based model, then hash mapping is performed based on the locality sensitive hash function, and then through redundant channel clustering, randomly retaining key channels, and fine-tuning the pruned model in sequence, an optimized lightweight convolutional neural network model, that is, a lightweight convolutional neural network-based model, is obtained.
[0033] This method first divides the input and output channels into blocks, then extracts high-dimensional feature vectors from the weight tensors of the channels within each block. Based on this, the locality sensitive hash algorithm maps the high-dimensional feature vectors to low-dimensional hash codes through a random projection hash function, so that blocks with high similarity are mapped to the same hash bucket, synchronously removing redundant blocks of the input and output channels. This can not only achieve fast screening of the channels to be pruned, but also avoid the high computational overhead of multiple iterative solutions to the sparsification problem in the traditional method, thus significantly improving the pruning efficiency and retaining the advantages of structured pruning while efficiently selecting blocks.
[0034] The specific process is as follows: Among them, S21. The block partitioning stage includes the following steps: S211. Obtain the original weight matrix of the filters in the pre-trained convolutional neural network-based model and a predetermined block size.
[0035] The expression of the original weight matrix is: (1).
[0036] Where is the original weight matrix, R is the real number tensor space, is the current layer number, is the input channel number of the th layer,
[0037] S212. Divide the original weight matrix of the pre-trained convolutional neural network-based model according to the predetermined block size to obtain target blocks.
[0038] The expression of the target block is: (2).
[0039] Where is the index in the row direction, is the predetermined block size, is the index in the column direction.
[0040] S22. The block selection stage includes the selection of output channels and the selection of input channels. After the block selection stage performs the selection of output channels and input channels on the target blocks, a pruned lightweight convolutional neural network-based model is obtained.
[0041] Among them, S221. The selection of output channels includes the following steps: S2211. Define each target block as: (3).
[0042] Where is the operation of converting a matrix to a vector, is the operation of converting a matrix to a vector, is the concatenation of vectors, is the dimension parameter of the th layer, is the filter height, is the filter width.
[0043] The vector , corresponding to in s consecutive columns, is the basic unit of the hash map.
[0044] S2212. Definition b Introduce a fixed random matrix for each hash value and define a locality-sensitive hashing function . Among them, the elements of the fixed random matrix are generated by a Gaussian distribution.
[0045] Locality-sensitive hashing function has the following expression: (4).
[0046] Among them, is the input vector, and taking the maximum value from the set.
[0047] S2213. Perform hash mapping on each target block through the locality-sensitive hashing function so that each target block obtains a hash value.
[0048] S2214. Randomly select one and retain its corresponding output channel, while deleting the channels corresponding to other channels, specifically including: S22141. Determine similar blocks according to the hash value. The hash values of the similar blocks are the same.
[0049] S22142. Retain one similar block and the output channel corresponding to the similar block, and delete the redundant blocks and the output channels corresponding to the redundant blocks, obtaining a convolutional neural network model with trimmed output channels. The hash values of the redundant blocks are the same as those of the similar blocks.
[0050] The present invention uses random projection for efficient hash coding. First, given b a number of hash values, by introducing a fixed random matrix , define a locality-sensitive hashing function , and then after performing hash mapping through the locality-sensitive hashing function, each is assigned a hash value .
[0051] The segmentation vectors with the same hash value are considered highly similar and contain a large amount of redundant information. Therefore, it is necessary to randomly select one and retain its corresponding output channel, while deleting the channels corresponding to other channels, which effectively trims the output channels.
[0052] S222. Selection of input channels, including the following steps: S2221. Define each target block as: (5).
[0053] Among them, is a set composed of multiple vector blocks of the input channel.
[0054] S2222. Definition b Introduce a fixed random matrix for each hash value and define a locality - sensitive hashing function .
[0055] S2223. Through the locality - sensitive hashing function perform hash mapping on each target block so that each target block gets a hash value.
[0056] S2224. Randomly select one and retain its corresponding output channel, while deleting the channels corresponding to other channels.
[0057] Similarly, the selection of input channels is also based on the locality - sensitive hashing method. Apply the hash mapping to using the locality - sensitive hashing function and delete the redundant blocks sharing the same hash value.
[0058] S23. The fine - tuning stage includes: fine - tuning the pruned convolutional neural network - based model to obtain a lightweight convolutional neural network - based model.
[0059] The flow of the channel pruning algorithm based on locality - sensitive hashing is shown in Table 1:[[]]END]] Table 1 ;
[0060] The channel pruning algorithm based on locality - sensitive hashing of the present invention can force the construction of a pruning weight matrix, and at the same time achieve the sparsity of input and output channels. As Figure 2 shown, the channel pruning algorithm based on locality - sensitive hashing specifically acts on the matrix. First, divide the weight matrix of the pre - trained convolutional neural network model into smaller blocks, then apply the locality - sensitive hashing pruning algorithm for hash mapping and clustering, and prune the redundant blocks with the same hash value. Finally, merge the pruned model to obtain a convolutional neural network model.
[0061] The computational complexity of the channel pruning algorithm based on locality - sensitive hashing is , where is the computational complexity, L is the number of layers of the model, is the number of input channels of the th layer, is the th layer, and
[0062] The channel pruning algorithm based on locality-sensitive hashing in the present invention helps to maintain the correlation between input and output channels, expand the candidate space of channel selection, and does not rely on iterative algorithms to solve the sparse optimization problem, thus reducing both the computational burden and the pruning time, and minimizing the negative impact of pruning on model performance.
[0063] A digital media processing system based on an improved channel pruning algorithm, comprising: An acquisition module, configured to acquire digital media data to be processed, including image data, video data, and audio data.
[0064] A digital media processing module, configured to process the digital media data to be processed through a channel pruning algorithm based on locality-sensitive hashing to obtain processed target information. The channel pruning algorithm based on locality-sensitive hashing includes: a block partitioning stage, a block selection stage, and a fine-tuning stage.
[0065] An electronic device, comprising a memory and a processor. When the processor calls the computer program stored in the memory, the content of the digital media processing method based on the improved channel pruning algorithm as described above is implemented.
[0066] A storage medium stores computer-executable instructions. When the computer-executable instructions are loaded and executed by a processor, the content of the digital media processing method based on the improved channel pruning algorithm as described above is implemented.
[0067] Finally, it should be noted that the above embodiments are only used to illustrate the technical method of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical method of the present invention, and these modifications or equivalent replacements do not enable the modified technical method to deviate from the spirit and scope of the technical method of the present invention.
Claims
1. A digital media processing method based on an improved channel pruning algorithm, characterized in that, It includes the following steps: S1. Obtain the digital media data to be processed, which includes image data, video data, and audio data; S2. Process the digital media data to be processed through an improved CNN convolutional neural network to obtain target information; The improved CNN convolutional neural network is improved based on the channel pruning algorithm of locality-sensitive hashing. The channel pruning algorithm based on locality-sensitive hashing includes: a block division stage, a block selection stage, and a fine-tuning stage.
2. The digital media processing method based on the improved channel pruning algorithm according to claim 1, characterized in that, The block division stage includes the following steps: S211. Obtain the original weight matrix of the filters in the pre-trained convolutional neural network and a predetermined block size; the expression of the original weight matrix is: ; Among them, is the original weight matrix, R is the real number tensor space, is the current layer number, is the input channel number of the layer, and is the number of output channels; It should be noted that the original text seems to have some incomplete or unclear parts in the description, especially the relationship between some variables and expressions is not very clear. The translation is based on the existing text as accurately as possible. S212. Divide the original weight matrix of the model of the pre-trained convolutional neural network according to the predetermined block size to obtain target blocks; the expression of the target blocks is: ; Among them, is the index in the row direction, is the predetermined block size, is the index in the column direction.
3. A digital media processing method based on an improved channel pruning algorithm according to claim 1, characterized in that, The block selection stage includes the selection of output channels and the selection of input channels; after the block selection stage performs the selection of output channels and input channels on the target blocks, a pruned convolutional neural network structure is obtained.
4. A digital media processing method based on an improved channel pruning algorithm according to claim 3, characterized in that, The selection of output channels includes the following steps: S2211. Define each target block as: ; Among them, is a set formed by splicing multiple vector blocks for the output channel, is the operation of converting a matrix into a vector, is the concatenation of vectors, is the dimensional parameter of the is the filter height, is the filter width; S2212. Definition b Introduce a fixed random matrix for each hash value, and define a locality-sensitive hashing function ; The expression of the locality-sensitive hashing function is as follows: ; Among them, is the input vector, is to take the maximum value from the set; S2213. Through a locality-sensitive hashing function Perform hash mapping on each target block so that each target block obtains a hash value; S2214. Randomly select one and retain its corresponding output channel, and at the same time delete the channels corresponding to other channels.
5. A digital media processing method based on an improved channel pruning algorithm according to claim 4, characterized in that In S2214, randomly select one and retain its corresponding output channel, and at the same time delete the channels corresponding to other channels, specifically including: S22141. Determine similar blocks according to the hash values; the hash values of the similar blocks are the same; S22142. Retain one similar block and the output channel corresponding to the similar block, and delete the redundant blocks and the output channels corresponding to the redundant blocks to obtain a convolutional neural network model with pruned output channels; the hash values of the redundant blocks are the same as those of the similar blocks.
6. A digital media processing method based on an improved channel pruning algorithm according to claim 4, characterized in that, The selection of input channels includes the following steps: S2221. Define each target block as: ; Among them, is a set formed by splicing multiple vector blocks of the input channel; S2222. Definition b Introduce a fixed random matrix into a hash value, and define a locality-sensitive hashing function ; S2223. Through a locality-sensitive hashing function Perform hash mapping on each target block so that each target block obtains a hash value; S2224. Randomly select one and retain its corresponding output channel, and at the same time delete the channels corresponding to other channels.
7. A digital media processing method based on an improved channel pruning algorithm according to claim 3, characterized in that The fine-tuning stage includes: fine-tuning the model based on the convolutional neural network after pruning to obtain a lightweight model.
8. A digital media processing system based on an improved channel pruning algorithm, characterized in that, It includes: An acquisition module that acquires digital media data to be processed, which includes image data, video data, and audio data; A digital media processing module for processing the digital media data to be processed through a channel pruning algorithm based on locality-sensitive hashing to obtain target information; the channel pruning algorithm based on locality-sensitive hashing includes: a block division stage, a block selection stage, and a fine-tuning stage.
9. An electronic device, characterized in that, It includes a memory and a processor. A computer program is stored in the memory. When the processor calls the computer program in the memory, it implements the content of the digital media processing method based on the improved channel pruning algorithm according to any one of claims 1 to 7.
10. A storage medium, characterized in that, Computer-executable instructions are stored in the storage medium. When the computer-executable instructions are loaded and executed by the processor, they implement the content of the digital media processing method based on the improved channel pruning algorithm according to any one of claims 1 to 7.
Citation Information
Patent Citations
Image processing method and device based on greedy strategy reverse channel pruning
CN116992945A