Residual network construction method based on discrete wavelet transform
By introducing discrete wavelet transforms to decompose and reconstruct features into the bottleneck layer of the residual network, the problem of excessive residual network parameters and high computational complexity in resource-constrained environments is solved, and the parameter quantity and computational complexity are reduced, while improving the performance of the model's processing details characteristics.
Patent Information
- Application Number
- CN202510317145.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In resource-constrained environments, the residual network model has too many parameters and high computational complexity, making it difficult to reduce the parameter quantity and computational complexity without sacrificing model performance.
Using the residual network construction method based on discrete wavelet transform, by introducing one-dimensional and two-dimensional discrete wavelet transforms into the bottleneck layer of the residual network, the input features are decomposed into low-frequency and high-frequency information, and the features are reconstructed through convolution operations and inverse wavelet transforms to reduce the amount of parameters and calculation complexity.
It effectively reduces the amount of network parameters and reduces the computational complexity. It is especially suitable for embedded and mobile devices with limited resources. At the same time, it retains high-frequency details of the image and improves the performance of the model's processing details characteristics.
Smart Images

Figure CN120068951A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and more particularly to a method for constructing a residual network based on discrete wavelet transform. Background Art
[0002] With the continuous development of computer vision technology, deep convolutional neural networks have been widely used in tasks such as image classification, object detection, and semantic segmentation. Among them, the residual network is one of the current mainstream deep networks. With its bottleneck structure and skip connection mechanism, it effectively alleviates the problem of gradient disappearance and significantly improves the training effect of deep networks.
[0003] In many practical application scenarios, such as embedded systems, Internet of Things devices, mobile devices, and drones, computing resources, memory, and battery power are extremely limited. These resource-constrained devices need to process complex image recognition tasks, but their hardware limitations cannot support large-scale residual network models. For example, smart cameras in Internet of Things devices need to process video data in real time, but their memory and computing capabilities are limited; artificial intelligence applications in mobile devices need to process high-complexity tasks while maintaining low power consumption and efficient operation. Therefore, in these scenarios, how to reduce the number of parameters and computational complexity of the residual network without sacrificing model performance has become an urgent problem to be solved. Therefore, we propose a method for constructing a residual network based on discrete wavelet transform. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for constructing a residual network based on discrete wavelet transform to solve the problems proposed in the above background art: to solve the problems of excessive parameters and high computational complexity of the residual network model in resource-constrained environments in the prior art.
[0005] To achieve the above object, the present invention provides the following technical solutions: A method for constructing a residual network based on discrete wavelet transform, comprising the following steps: Step 1, establish a wavelet bottleneck layer based on discrete wavelet transform; Step 1.1 In the channel domain, perform one-dimensional discrete wavelet transform on the input feature X, decompose the input feature into low-frequency information and high-frequency information, respectively denoted as X low and X high , perform 1×1 convolution operations on the low-frequency information and high-frequency information respectively, and use the ReLU activation function and batch normalization processing; then reconstruct the feature through one-dimensional inverse discrete wavelet transform to obtain a feature tensor X 1 consistent with the dimension of the input feature; Step 1.2: In the spatial domain, perform two-dimensional discrete wavelet transform on the feature tensor X 1 to decompose it into a low-frequency subband XLL and three high-frequency subbands X LH 、X HL 、X HH ; Apply a 3×3 grouped convolution operation to each subband to extract multi-scale spatial features, and use the ReLU activation function and batch normalization; finally, reconstruct the features through two-dimensional inverse discrete wavelet transform to obtain the feature tensor X 2 ; Step 1.3: In the channel domain, perform one-dimensional discrete wavelet transform on the feature tensor X 2 again, extract the low-frequency information and high-frequency information in the feature tensor X 2 , and perform 1×1 convolution processing to restore the channel dimension; finally, reconstruct the features through one-dimensional inverse discrete wavelet transform to obtain the feature tensor X 3 .
[0006] Step 1.4: Using the residual connection, directly add the input feature X to the feature tensor X 3 through a skip connection to obtain the output feature tensor X out ; Step 2: Use the wavelet bottleneck layer established in Step 1 to construct a wavelet transform-based residual network. The construction of the residual network includes: Step 2.1: Construct an input layer to receive the input image data and perform normalization processing; Step 2.2: Construct at least one convolutional layer, use a convolutional kernel of size 7×7 and a stride of 2 to perform local feature extraction on the input image data, output a feature map, and perform processing through the ReLU activation function and batch normalization; Step 2.3: Construct at least one max pooling layer, use a pooling kernel of size 3×3 and a stride of 2 to perform downsampling on the output feature map of the convolutional layer; Step 2.4: Construct multiple wavelet bottleneck layers, and perform feature extraction on the output of the pooling layer through the combination of wavelet decomposition and convolution operations; The multiple wavelet bottleneck layers are connected in sequence, and the size of the feature map is gradually reduced through downsampling operations; Construct at least one average pooling layer to perform global pooling on the final output of the wavelet bottleneck layer to obtain a feature vector of a fixed size; Step 2.5: Construct a fully connected layer, flatten the output of the average pooling layer, and output a 1×1000 feature vector; Step 2.6: Construct a Softmax layer to convert the output of the fully connected layer into a probability distribution of 1000 classes for predicting the final class; Step 3: Train and test the constructed residual network model on a classical image classification dataset.
[0007] Preferably, in step 1.2, the grouped convolution operation is performed separately on the low-frequency subband and high-frequency subband features to make full use of the wavelet high-frequency information and reduce the number of model parameters.
[0008] Preferably, in step 2.2, after each convolution operation, a non-linear transformation is performed through the ReLU activation function to enhance the expression ability of the residual network. The output of the convolutional layer is normalized through batch normalization to prevent gradient vanishing or explosion.
[0009] Preferably, the number of wavelet bottleneck layers in the residual network is adjusted according to specific tasks and datasets to form residual network models with different depths.
[0010] Preferably, the Haar wavelet function is used for wavelet transform.
[0011] Preferably, in step 2.6, according to the number of categories generated by the Softmax layer for classification, the 1×1000 output by flattening the features in the fully connected layer is replaced with 1×the number of categories for classification.
[0012] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) By introducing discrete wavelet transform into the bottleneck block structure of the residual network, the present invention can effectively reduce the number of parameters in the network. Compared with the traditional ResNet architecture, the number of parameters is reduced by one-fifth to one-fourth, which is particularly suitable for resource-constrained embedded and mobile devices.
[0013] (2) In the feature extraction process, the method proposed by the present invention decomposes the input features into low-frequency and high-frequency subbands through discrete wavelet transform, effectively retaining the high-frequency detail information of the image and improving the performance of the residual network model in processing detail features.
[0014] (3) By using discrete wavelet transform, the present invention can achieve a reversible downsampling operation without loss of information, reduce the feature resolution, thereby reducing the computational cost and improving the running efficiency of the residual network model.
[0015] (4) By performing multi-scale feature extraction in the channel domain and spatial domain, the present invention enables the residual network model to better capture important features in the image, enhancing the performance of the model in tasks such as image classification and object detection. Description of the Drawings
[0016] Figure 1 It is a schematic diagram of the wavelet bottleneck layer of the present invention; Figure 2 It is a schematic diagram of the residual network based on discrete wavelet transform of the present invention. Detailed Embodiment
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0018] Embodiment: Please refer to Figure 1-2 , a method for constructing a residual network based on discrete wavelet transform, comprising the following steps: Step 1, as Figure 1 shown, establish a wavelet bottleneck layer based on discrete wavelet transform; Step 1.1 In the channel domain, perform one-dimensional discrete wavelet transform on the input feature X. By performing channel dimensionality reduction in the one-dimensional wavelet domain of the input feature, the correlation and important features between channels are effectively captured. The specific operation is to use one-dimensional discrete wavelet transform to decompose the input feature into low-frequency information and high-frequency information, respectively denoted as X low and X high . Perform 1×1 convolution operations on the low-frequency information and high-frequency information respectively, and use the ReLU activation function and batch normalization processing. The convolution layer uses the ReLU activation function for nonlinear processing, and combines batch normalization (BN) to improve the stability of training; then reconstruct the feature through one-dimensional inverse discrete wavelet transform to obtain a feature tensor X 1 with the same dimension as the input feature; Step 1.2: In the spatial domain, perform two-dimensional discrete wavelet transform on the feature tensor X 1 . By performing feature extraction in the two-dimensional wavelet domain, important feature extraction in the spatial domain is realized. The specific operation is to use two-dimensional discrete wavelet transform to perform spatial decomposition on X 1 , decompose it into a low-frequency subband X LL and three high-frequency subbands X LH , X HL , X HH ; apply 3×3 grouped convolution operations to each subband to extract multi-scale spatial features, and use the ReLU activation function and batch normalization processing; finally, reconstruct the feature through two-dimensional inverse discrete wavelet transform to obtain the feature tensor X 2 ; among them, the grouped convolution operation is performed separately on the low-frequency subband and high-frequency subband information features to make full use of wavelet high-frequency information while reducing the number of parameters.
[0019] Step 1.3: In the channel domain, restore the original channel dimension by performing channel dimensionality increase in the one-dimensional wavelet domain of the input feature. The specific operation is: perform one-dimensional discrete wavelet transform on the feature tensor X 2 again, and extract the feature tensor X 2The low-frequency information and high-frequency information in it are subjected to 1×1 convolution processing for restoring the channel dimension; finally, the feature is reconstructed through one-dimensional inverse discrete wavelet transform to obtain the feature tensor X 3 。
[0020] Step 1.4: Using the residual connection, directly add the input feature X to the feature tensor X through the skip connection 3 to obtain the output feature tensor X out 。
[0021] Among them, the wavelet bottleneck layer specifically includes: One-dimensional discrete wavelet transform module, which is used to decompose the input feature in the channel domain to obtain low-frequency information and high-frequency information; 1×1 convolution module, which performs convolution operations on the low-frequency information and high-frequency information respectively to extract channel domain features; One-dimensional inverse discrete wavelet transform module, which is used to reconstruct the feature to obtain a feature tensor with the same dimension as the input feature; Two-dimensional discrete wavelet transform module, which is used to decompose the reconstructed feature tensor in the spatial domain to obtain a low-frequency sub-band and three high-frequency sub-bands; 3×3 grouped convolution module, which performs convolution operations on each sub-band to extract multi-scale features in the spatial domain; Two-dimensional inverse discrete wavelet transform module, which is used to reconstruct the feature to obtain a feature tensor containing rich spatial features; Channel dimension recovery module, which restores the channel dimension of the feature tensor by performing one-dimensional discrete wavelet transform and 1×1 convolution operation again.
[0022] Step 2: As Figure 2 shown, using the wavelet bottleneck layer established in Step 1, construct a wavelet transform-based residual network, and the construction of the residual network includes: Step 2.1: Construct an input layer, which is used to receive the input image data and perform normalization processing; Step 2.2: Construct at least one convolutional layer, use a convolutional kernel with a size of 7×7 and a stride of 2 to perform local feature extraction on the input image data, output a feature map, and process it through the ReLU activation function and batch normalization; Step 2.3: Construct at least one max pooling layer, use a pooling kernel with a size of 3×3 and a stride of 2 to perform downsampling on the output feature map of the convolutional layer to reduce the number of parameters; Step 2.4: Construct multiple wavelet bottleneck layers, and through the combination of wavelet decomposition and convolution operations, perform feature extraction on the output of the pooling layer, effectively reducing the computational amount and retaining detailed information (i.e., texture and edge information); The multiple wavelet bottleneck layers are connected in sequence, and the size of the feature map is gradually reduced through downsampling operations; Construct at least one average pooling layer to perform global pooling on the final output of the wavelet bottleneck layer to obtain a feature vector of a fixed size; Step 2.5: Construct a fully connected layer, flatten the output of the average pooling layer, and output a 1×1000 feature vector; Step 2.6: Construct a Softmax layer to convert the output of the fully connected layer into a probability distribution of 1000 categories for predicting the final category; Step 3: Train and test the constructed residual network model on a classical image classification dataset.
[0023] In this application, in step 1.2, In this application, in step 2.2, after each convolution operation, a ReLU activation function is used for non-linear transformation to enhance the expression ability of the residual network. The output of the convolution layer is normalized through batch normalization to prevent gradient disappearance or explosion.
[0024] In this application, the number of wavelet bottleneck layers in the residual network is adjusted according to specific tasks and datasets to form residual network models with different depths.
[0025] In this application, the Haar wavelet function is used for wavelet transform, and the Haar wavelet function can also be replaced with other wavelet functions such as db and bior.
[0026] In this application, according to the number of categories generated by the Softmax layer, the 1×1000 output by flattening the features of the fully connected layer is replaced with 1×the number of categories for classification.
[0027] In the present invention, a wavelet bottleneck structure is designed to reduce the number of parameters and the amount of computation. The wavelet bottleneck structure introduces discrete wavelet transform into the bottleneck structure of the residual network structure and can seamlessly replace the bottleneck structure of the original network. As Figure 1 shown, in the channel domain, feature extraction is performed through one-dimensional wavelet transform of the input data, effectively capturing the correlation and important features between channels. In the spatial domain, important feature extraction in the spatial domain is achieved through feature extraction in the two-dimensional wavelet domain. The convolution of the present invention uses grouped convolution for wavelet high and low frequency information features, so the wavelet high and low frequency information is fully utilized while reducing the number of parameters.
[0028] As Figure 1 shown, the input feature is X, where [H, W, C], H and W are the height and width respectively, and C is the number of channels; Conv is the abbreviation of "Convolution"; Batch Normalization (BN), also known as batch normalization; ReLU is the rectified linear unit.
[0029] In the present invention, the discrete wavelet transform can completely recover the original signal or image through the inverse transform, ensuring that no information is lost during the image processing process. Performing convolution in the wavelet domain not only reduces the number of parameters but also does not lose information. The bottleneck structure based on the discrete wavelet transform theoretically reduces the number of parameters by one-fifth to one-fourth.
[0030] Example 1: Taking an image size of 224×224×3 as an example; Step 2.1: First, take the image of 224×224×3 as the input data. The input data is first subjected to normalization processing to ensure that during the subsequent convolution and feature extraction processes, the distribution of the data can be effectively processed and optimized.
[0031] Step 2.2: Perform a convolution operation on the normalized input image data in the convolutional layer. The convolutional layer uses a convolutional kernel of size 7×7, a stride of 2, an input depth of 3, and an output depth of 64 feature maps. After the convolution operation, the size of the feature map is 112×112×64. After each convolution operation, a non-linear transformation is performed through the ReLU activation function to increase the expression ability of the residual network model. Batch normalization processing is applied after the convolutional layer. The batch normalization layer normalizes the feature maps output by the convolutional layer to accelerate the convergence speed of the network and stabilize the training process. This step helps prevent gradient vanishing or explosion and improves the stability of training.
[0032] Step 2.3: Perform a pooling operation on the batch-normalized feature maps. The pooling layer uses the maximum pooling (MaxPooling) operation, with a pooling kernel size of 3×3 and a stride of 2, and outputs a feature map with a size of 56×56×64. After the pooling operation, the size of the feature map is reduced through downsampling while retaining important feature information, thereby reducing the computational amount.
[0033] Step 2.4: Then perform wavelet bottleneck processing. The wavelet bottleneck layer extracts features by combining the discrete wavelet transform and convolution operations. This layer includes multiple bottleneck modules, and the input and output sizes of each module are 56×56. The wavelet bottleneck layer performs downsampling operations after multiple wavelet bottleneck modules, reducing the feature map size from 56×56 to 28×28, 14×14, and 7×7 respectively to extract more abstract features.
[0034] Among them, different wavelet bottlenecks constitute different residual network models. For the residual networks based on the discrete wavelet transform - 50, 101, 152, n1, n2, n3, n4 are [3, 4, 6, 3], [3, 4, 23, 3], [3, 8, 36, 3] respectively. Among them, Figure 2 in and the above n1, n2, n3, n4 all represent the number of repetitions of each layer of the network.
[0035] Step 2.5: The downsampled feature map is passed through a pooling layer, then flattened and input into a fully connected layer. The pooling layer performs average pooling on the feature map to further reduce the number of parameters and is used for regularization to prevent overfitting. After the feature map is flattened, it is input into the fully connected layer, which outputs a 1×1000 feature vector for the classification task. The fully connected layer transforms the extracted features into a probability distribution of the classification results.
[0036] Step 2.6: Apply Softmax activation to the output of the fully connected layer. The Softmax layer generates a probability distribution over 1000 classes, representing the prediction probability of the residual network for each class. The final classification result is the class with the highest probability, determining the classification label of the image.
[0037] Example 2: Use classical classification datasets (mini-imagenet, cifar-100, cifar-10) for training and validation. Use an adamw optimizer with a momentum of 0.9. Fix the learning rate and weight decay to 0.001 and 0.05, and the batch size to 128. Train the residual network model for 500 epochs, and use data augmentation techniques such as random cropping, horizontal flipping, and normalization to enhance the robustness of the residual network model.
[0038] And compare different models of residual networks based on discrete wavelet transform with the baseline network. As can be seen from Table 1 and Table 2, with the use of the wavelet bottleneck structure, the Top-1 accuracy generally decreases slightly, especially on the Mini-ImageNet dataset, where the accuracy decrease is more obvious. However, the overall decrease is not significant. The top-5 accuracy changes very little, and in some cases, there are improvements, such as ResNet50 on CIFAR-10 and ResNet152 on CIFAR-10. Although the accuracy decreases slightly considering the significant reduction in the number of parameters, in many practical applications, especially on resource-constrained devices, this trade-off is acceptable. For applications that need to run with limited resources, this method can significantly improve the efficiency of the residual network model, providing an effective means of compressing the residual network model.
[0039] Table 1: Comparison of different residual network models on different datasets; Table 2: Comparison of the number of parameters of different residual network model structures; The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and the above embodiments and the descriptions in the specification are only preferred examples of the present invention, which are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A residual network construction method based on discrete wavelet transform, characterized in that: The steps include: Step 1: Establish a wavelet bottleneck layer based on discrete wavelet transform, including: Step 1.1 In the channel domain, perform a one-dimensional discrete wavelet transform on the input feature X and decompose the input feature into low-frequency information and high-frequency information, represented as X low and X high , perform 1×1 convolution operations on low-frequency information and high-frequency information respectively, and use ReLU activation function and batch normalization processing; then reconstruct the features through one-dimensional inverse discrete wavelet transform to obtain the feature tensor X1 that is consistent with the input feature dimension; Step 1.2: In the spatial domain, perform a two-dimensional discrete wavelet transform on the feature tensor X1 and decompose it into low-frequency subbands X LL and three high frequency sub-bands X LH , X HL , X HH ; Apply 3×3 grouped convolution operation to each subband to extract multi-scale spatial features, and use ReLU activation function and batch normalization processing; finally, reconstruct the features through two-dimensional inverse discrete wavelet transform to obtain the feature tensor X2; Step 1.3: In the channel domain, perform one-dimensional discrete wavelet transform on the feature tensor X2 again to extract the low-frequency and high-frequency information in the feature tensor X2, and perform 1×1 convolution to restore the channel dimension; finally, reconstruct the features through one-dimensional inverse discrete wavelet transform to obtain the feature tensor X3; Step 1.4: Using residual connections, add the input feature X directly to the feature tensor X3 through skip connections to obtain the output feature tensor X out ; Step 2: Using the wavelet bottleneck layer established in step 1, construct a residual network based on wavelet transform. The construction of the residual network includes: Step 2.1: Construct an input layer to receive input image data and perform standardization processing; Step 2.2: Construct at least one convolutional layer, use a convolution kernel of size 7×7 and stride 2, extract local features from the input image data, output feature maps, and process them through ReLU activation function and batch normalization; Step 2.3: Construct at least one maximum pooling layer, using a pooling kernel of size 3×3 and stride 2 to downsample the output feature map of the convolutional layer. Step 2.4: Construct multiple wavelet bottleneck layers, and extract features from the output of the pooling layer by combining wavelet decomposition and convolution operations; Multiple wavelet bottleneck layers are connected in sequence, and the feature map size is gradually reduced through downsampling operations; Construct at least one average pooling layer to globally pool the final output of the wavelet bottleneck layer to obtain a fixed-size feature vector; Step 2.5: Build a fully connected layer, flatten the output of the average pooling layer, and output a 1×1000 feature vector; Step 2.6: Construct a Softmax layer to convert the output of the fully connected layer into a probability distribution of 1000 categories for predicting the final category. Step 3: Train and test the constructed residual network on the classic image classification dataset.
2. The method for constructing a residual network based on discrete wavelet transform according to claim 1, characterized in that: In step 1.2, the group convolution operation is performed on the low-frequency sub-band and the high-frequency sub-band respectively to make full use of the high-frequency information and reduce the number of model parameters.
3. The method for constructing a residual network based on discrete wavelet transform according to claim 1, characterized in that: In step 2.2, each convolution operation is followed by a nonlinear transformation using the ReLU activation function to improve the expressive power of the residual network. The output of the convolution layer is standardized using batch normalization to prevent the gradient from vanishing or exploding.
4. The method for constructing a residual network based on discrete wavelet transform according to claim 1, characterized in that: The number of wavelet bottleneck layers in the residual network is adjusted according to specific tasks and data sets to form residual network models of different depths.
5. The method for constructing a residual network based on discrete wavelet transform according to claim 1, characterized in that: The wavelet transform adopts the Haar wavelet function.
6. The method for constructing a residual network based on discrete wavelet transform according to claim 1, characterized in that: In step 2.6, according to the number of categories generated by the Softmax layer, the fully connected layer replaces the 1×1000 output of the feature flattening with 1× the number of categories of the classification.
Citation Information
Patent Citations
Discrete wavelet-based visual converter network classification method
CN119313998A
Cross-modal pedestrian re-identification method and system based on wavelet transform
CN119445621A