An Image Classification Method Based on Shearlet Network and Direction Attention Mechanism

Through the combination of shear wave network and direction attention module, the multi-directional and multi-scale features of the image are extracted, which solves the problem of ignoring texture, edge and direction information in the existing methods, and realizes efficient image classification.

CN116109855BActive Publication Date: 2025-07-29XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211044682.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2025-07-29
Estimated Expiration
2042-08-30

AI Technical Summary

Technical Problem

Existing image classification methods ignore the texture, edge and direction information of the image, make it difficult to process the geometric transformation of the image, and have a large demand for computing resources, resulting in insufficient classification accuracy and speed.

Method used

The shear wave network is used for image decomposition, combining the direction attention module and the lightweight convolutional neural network, multi-directional and multi-scale features of the image are extracted, important features are enhanced through adaptive weights, and lightweight network model is constructed.

Benefits of technology

It improves the accuracy and speed of image classification, reduces computing costs, and is suitable for platforms with limited computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116109855B_ABST
    Figure CN116109855B_ABST
Patent Text Reader

Abstract

The present invention discloses an image classification method based on a shear wave network and a directional attention mechanism, mainly solving the problem that existing methods ignore discriminative decomposition features, resulting in poor running speed and accuracy of image classification. The solution includes: 1) constructing a training sample set, preprocessing and padding it to obtain a set of images to be decomposed; 2) constructing a lightweight network model composed of a shear wave network, a directional attention module, and a lightweight convolutional neural network; wherein the shear wave network is used to extract features in different directions of the image, the directional attention module assigns adaptive weights according to the importance of features in different directions of the image, and the lightweight convolutional neural network performs abstract feature extraction; 3) training the constructed lightweight network model; 4) using the trained network model to predict the category of the image to be classified to complete the classification. The present invention can effectively reduce the computational cost while improving the running speed and accuracy of image classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and image processing, and further relates to an image classification method, specifically an image classification method based on a shearlet network and a directional attention mechanism, which can be applied to face recognition, remote sensing image classification, and natural scene classification. Background Art

[0002] Image classification is a fundamental task in computer vision, and its task is to predict the category to which a given image belongs through an algorithm. Image classification has a wide range of applications in scenarios such as face recognition, remote sensing image classification, and medical image classification. The general method of image classification includes two steps: one is to first extract features from the image to obtain a discriminative feature representation of the image, and the other is that the classifier uses these discriminative features to calculate a classification prediction vector, and the category corresponding to the element with the largest value in the vector is the predicted category.

[0003] With the development of deep learning, these two steps of the image classification algorithm are often combined in a network structure to complete classification. The current popular methods use deep neural networks based on convolutional neural network CNN or based on transformer to extract features and classify images, and have achieved excellent performance. However, some problems in deep learning have led to some limitations in its image classification tasks. Specifically: 1) The deep learning-based methods do not particularly pay attention to the texture, edge, and direction information of the image, which are important image features, in the image classification task; 2) The intra-class geometric transformation of the image makes it difficult for the algorithm to extract important features in the image, but it is extremely difficult to study these geometric transformations in the deep neural network; 3) The deep neural network generally directly extracts image features in the spatial domain, but some features that are difficult to extract in the spatial domain can be more easily obtained if the image is decomposed in the frequency domain; 4) The deep neural network is often used as a black box, with poor mathematical interpretability, and the training of the deep neural network often requires a large amount of computing resources, which is very unfriendly to platforms with limited computing resources.

[0004] In order to avoid the limitations of the above deep learning methods in the image classification task, the existing methods add an image decomposition module between the original image and the deep neural network. This module decomposes the image in the frequency domain, can extract the texture, edge, and direction information in the image, provides invariance to geometric transformations, and helps the neural network reduce parameter overhead. The commonly used methods to implement this module include wavelet transform, variants of wavelet transform, and multi-scale geometric transform, etc. These transforms are often used in tasks such as texture extraction, contour edge extraction, image fusion, and image denoising. In recent years, due to the development of deep learning, some studies have combined these transforms with deep learning to achieve better performance.

[0005] Existing mainstream methods use wavelet transform or variants of wavelet transform to decompose images. However, the support interval of wavelet transform is square, which cannot sparsely represent singular curves in images. Moreover, wavelet transform can only capture directional features in horizontal, vertical, and diagonal directions in images. These limitations lead to low efficiency of wavelet transform in processing image edges, textures, and directional features, and it is difficult to obtain a high classification accuracy rate with less computational consumption. Additionally, wavelet transform is difficult to handle geometric transformations in images, which is not conducive to the extraction of the essential features of images by the algorithm. And as a multi-scale geometric transform, contourlet transform will have sub-band spectral aliasing phenomenon, which weakens the directional selectivity of contourlet transform and is not conducive to image feature extraction.

[0006] In addition, existing methods only use image transformation methods to decompose images, without paying attention to more discriminative decomposition features, and often use large and complex deep neural networks for subsequent feature extraction, resulting in extremely high demand for computing resources by the algorithm and being difficult to be deployed to platforms with low computing power. Therefore, it is urgent to study a lightweight feature extraction network to reduce the computational cost, improve the running speed and accuracy of image classification. Summary of the Invention

[0007] The purpose of the present invention is to propose an image classification method based on shearlet network and directional attention mechanism in view of the above deficiencies of the existing technologies, so as to enhance the ability of the neural network to extract different directional features in images, and reduce the training parameters and computational cost of the network.

[0008] The technical solution of the present invention is: to construct a lightweight network that can extract texture contours and geometric features in different directions of an image, and can adaptively increase the weight of the directional features beneficial to classification in the image; first, construct a 3-layer shearlet network as the image decomposition module of the present invention to extract features in different directions of the image; then use the directional attention module to assign adaptive weights according to the importance of different directional features of the image, and obtain a lightweight CNN for extracting more abstract feature representations of the image and outputting a classification prediction vector. The present invention can achieve a high-precision image classification effect with less computing resources.

[0009] The specific steps for the present invention to achieve the above object are as follows:

[0010] (1) Construct a training sample set, and perform preprocessing and padding on it to obtain a set of images to be decomposed:

[0011] (1.1) Construct a training sample set with m image samples and c categories by using a public dataset or taking images of specific scenes according to the needs of the classification scenario Its corresponding true label is where, H z and Wz They are the height and width of the z-th image sample respectively;

[0012] (1.2) Normalize all pixels of each image in the training sample set X to obtain a normalized image with pixel values in the range [0, 1];

[0013] (1.3) Horizontally flip, edge pad, and randomly crop the normalized image to obtain a preprocessed training sample set where H and W represent the height and width of the cropped image sample;

[0014] (1.4) Perform edge padding of length p (p ≥ 0) on the preprocessed training sample set to obtain a set of images to be decomposed

[0015] (2) Construct a lightweight network model composed of a shearlet network, a direction attention module, and a lightweight convolutional neural network:

[0016] (2.1) Construct a shearlet network with a 3-layer filter structure, and the implementation steps are as follows;

[0017] (2.1.1) Create a Gaussian low-pass filter As the filter of the first layer of the shearlet network, it is used to perform low-pass filtering on the input image to obtain the low-frequency information of the image;

[0018] (2.1.2) Create a shearlet filter bank ψ according to the shearlet transform as the filter of the second layer of the shearlet network:

[0019]

[0020] where a represents the filter parameter that controls the scale of the shearlet transform, and s represents the filter parameter that controls the direction of the shearlet transform;

[0021] (2.1.3) Create 9 shearlet filter banks identical to ψ according to the shearlet transform, and cascade each of them with each filter in the shearlet filter bank ψ created in step (2.1.2). These 9 shearlet filter banks together serve as the filter of the third layer of the shearlet network;

[0022] (2.1.4) The filters of the first, second, and third layers together form the shearlet network, which is used to obtain the decomposition feature B of the image to be decomposed;

[0023] (2.2) Construct a direction attention module, and the specific steps are as follows:

[0024] (2.2.1) Select the channel attention module of SENet to form the image adaptive weight vector extraction part, whose structure is GAP→FC1→activation function→FC2→activation function, where GAP represents global average pooling, FC1 represents the fully connected layer for compressing the vector length, and FC2 represents the fully connected layer for expanding the vector length;

[0025] (2.2.2) Use D two-dimensional convolutions to form the image multi-scale information extraction part, D≥1, and its structure is a series structure, that is: the first two-dimensional convolution→the second two-dimensional convolution→···→the Dth two-dimensional convolution;

[0026] (2.2.3) The image adaptive weight vector extraction part and the multi-scale information extraction part together form the direction attention module, which is used to further process the decomposed feature B to obtain the weighted multi-directional and multi-scale feature set F;

[0027] (2.3) Build a lightweight convolutional neural network, and the specific steps are as follows:

[0028] (2.3.1) Build a 1×1 convolution block, whose structure is the first two-dimensional convolution→activation function→Dropout→the second two-dimensional convolution→Dropout; Dropout represents the random inactivation layer; the convolution kernel sizes of the first and second two-dimensional convolutions are both 1×1;

[0029] (2.3.2) Connect n 1×1 convolution blocks constructed in step (2.3.1) in series, where n≥D, and then connect a global average pooling layer, a fully connected layer, and a softmax layer in sequence after the last 1×1 convolution block to obtain a lightweight convolutional neural network, which is used to extract abstract features from the multi-directional and multi-scale feature set F to obtain the final classification prediction vector where c is the number of categories of images in the training sample set;

[0030] (2.4) Take the output of the shearlet network as the input of the direction attention module, take the output of the first layer of two-dimensional convolution of the direction attention module as the input of the first 1×1 convolution of the lightweight convolutional neural network, and add the outputs of the remaining D - 1 layers of two-dimensional convolution to the outputs of the first D - 1 layers of the lightweight convolutional neural network to obtain a lightweight network model;

[0031] (3) Train the lightweight network model:

[0032] (3.1) Input the z-th sample image x z ” in the image set X” to be decomposed into the lightweight network model, and the model outputs its classification prediction vector According to and x z ” corresponding image sample x zThe true label y z Calculate the prediction loss L of the lightweight network model:

[0033]

[0034] in, Represents the lightweight network model predicting the sample image x z The probability of belonging to category i; y z_i For the sample image x z True label y z Components, if the sample image x z The true label of is the i-th category, then y z_i =1, otherwise y z_i =0;

[0035] (3.2) Use the backpropagation algorithm to update the trainable parameters of the lightweight network model, and use the SGD algorithm as the optimizer to train the network during the training process;

[0036] (3.3) Take z = 1, 2, ..., m and repeat steps (3.1) to (3.2) until all images in the image set to be decomposed X' are trained to obtain a trained lightweight network model;

[0037] (4) Input the image to be classified into the trained lightweight network model to obtain the classification prediction vector corresponding to the classified image, select the category with the largest element value in the prediction vector as the predicted category of the image, and complete the image classification.

[0038] Compared with the prior art, the present invention has the following advantages:

[0039] First, because the present invention uses a shearlet network for image decomposition, the geometric transformations in the image can be well represented, solving the problem of insufficient representation of image geometric transformations by traditional convolutional neural networks, increasing the network's ability to represent images, and thus improving image classification accuracy;

[0040] Second, the present invention uses a directional attention module to assign greater weight to important directional features in the image, while also giving directional features different scale information, which facilitates the subsequent neural network to extract the essential features of the image, significantly reducing the complexity of the network and thus lowering the computational cost of image classification.

[0041] Third, the present invention uses a 1×1 convolution module to construct a lightweight CNN to fuse information between channels, so that the entire network has fewer trainable parameters, thereby improving the speed of image classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is an implementation flow chart of the present invention;

[0043] Figure 2 It is a schematic diagram of the shear wave network structure diagram in the present invention;

[0044] Figure 3 It is a schematic diagram of the direction attention module and the lightweight convolutional neural network structure diagram of the present invention. Specific implementation manners

[0045] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0046] Refer to Figure 1 A method for image classification based on a shear wave network and a direction attention mechanism proposed by the present invention constructs a lightweight network model that can extract texture contours and geometric features in different directions of an image, and can adaptively increase the weight of the direction features beneficial to classification in the image, uses it to extract the features for classification in the image, and outputs a classification prediction vector. The specific implementation steps are as follows:

[0047] Step 1. Construct a training sample set, and perform preprocessing and padding on it to obtain a set of images to be decomposed:

[0048] (1.1) Construct a training sample set containing m image samples and c categories by using a public data set or taking images of a specific scene according to the needs of the classification scenario Its corresponding true label is wherein, H z and W z are respectively the height and width of the z-th image sample;

[0049] (1.2) Normalize all pixels of each image in the training sample set X to obtain a normalized image with pixel values in the interval [0, 1];

[0050] (1.3) Horizontally flip, edge pad, and randomly crop the normalized image to obtain a preprocessed training sample set wherein, H and W represent the height and width of the cropped image sample; in this embodiment, the edge padding uses mirror padding with a length of 4, and the size of the random cropping is H×W;

[0051] (1.4) Perform edge padding with a length of p on the preprocessed training sample set, p≥0, to obtain a set of images to be decomposed

[0052] Step 2. Construct a lightweight network model composed of a shear wave network, a direction attention module, and a lightweight convolutional neural network:

[0053] (2.1) Construct a shear wave network with a 3-layer filter structure, such as Figure 2As shown below, the implementation steps are as follows;

[0054] (2.1.1) Create a Gaussian low-pass filter As the filter for the first layer of the shearlet network, it is used to perform low-pass filtering on the input image to obtain the low-frequency information of the image;

[0055] (2.1.2) Create a shearlet filter bank ψ according to the shearlet transform as the filter for the second layer of the shearlet network:

[0056]

[0057] Among them, a represents the filter parameter that controls the scale of the shearlet transform, and s represents the filter parameter that controls the direction of the shearlet transform. Here, the two filter parameters a and s are respectively used to control the scale and direction of the shearlet transform. In this embodiment, taking a = 1 means the scale of the shearlet filter is 1, and when s = 0, it means is a low-pass filter, and when s takes values from 1 to 8, it means has direction sensitivity and can extract the direction information of the image at 0°, 30°, 60°,..., 150° with the center coordinate as the center of symmetry respectively.

[0058] (2.1.3) Create 9 shearlet filter banks identical to ψ according to the shearlet transform, and cascade each of them with each filter in the shearlet filter bank ψ created in step (2.1.2). These 9 shearlet filter banks together serve as the filter for the third layer of the shearlet network;

[0059] (2.1.4) The filters of the first, second, and third layers together form the shearlet network, which is used to obtain the decomposition feature B of the image to be decomposed; obtaining the decomposition feature B of the image to be decomposed is achieved as follows:

[0060] (2.1.4a) Input the image x” to be decomposed into the low-pass filter φ of the first layer of the shearlet network for filtering to obtain the response b1, and then perform central cropping on it to obtain the output of the first layer of the shearlet network

[0061] (2.1.4b) Input the image x” to be decomposed into the shearlet filter bank ψ of the second layer of the shearlet network for decomposition to obtain a set of direction feature maps where b 2_0 is the low-frequency feature map of the image, and b 2_1 to b 2_8 are the feature maps of 8 different directions of the image. α represents the index of the direction feature map obtained by the shearlet filter bank of the second layer; then perform central cropping on b2 to obtain the output of the second layer of the shearlet network

[0062] (2.1.4c) Input the 9 feature maps in the set b2 of directional feature maps into the 9 filter banks in the third layer of the shearlet network for decomposition to obtain the decomposition feature set of the third layer of the shearlet network. Among them, β represents the index of the filter bank in the third layer, and αβ represents the index of the directional feature map obtained by the shearlet filter bank in the third layer; it contains 9 2 = 81 features in different directions, and then perform central cropping on b3 to obtain the output of the third layer of the shearlet network.

[0063] (2.1.4d) Concatenate the outputs of the first, second, and third layers obtained in steps (2.1.4a)-(2.1.4b) along the channels to obtain the output of the shearlet network. Among them, R = 3×(1 + 9 + 9 2 ) is the number of filters used for decomposition in the shearlet network.

[0064] The shearlet network constructed in the present invention is composed of cascaded shearlet filters. The main purpose is to extract the texture and directional information of the image, and to obtain a robust representation of the geometric transformation in the image. Among them, the shearlet filter can also be replaced by filters of other image transformations, and the number of cascaded layers can also be replaced by 2 layers, 4 layers or others according to actual conditions. What the present invention mainly highlights in this step is the construction method of the cascaded filter bank.

[0065] (2.2) Construct a directional attention module, and the specific steps are as follows:

[0066] (2.2.1) Select the channel attention module of SENet to form the part for extracting the image adaptive weight vector. Its structure is GAP→FC1→activation function→FC2→activation function, where GAP represents global average pooling, FC1 represents the fully connected layer for compressing the vector length, and FC2 represents the fully connected layer for expanding the vector length; in this embodiment, FC1 is a fully connected layer with the number of channels compressed by 10 times, FC2 is a fully connected layer with the number of channels expanded by 10 times, and activation functions such as ReLU and Sigmoid are used.

[0067] (2.2.2) Use two-dimensional convolutions in layer D to form the image multi-scale information extraction part, where D ≥ 1, and its structure is a cascaded structure, that is: the first two-dimensional convolution → the second two-dimensional convolution → ··· → the D-th two-dimensional convolution; in this embodiment, taking the case of using four two-dimensional convolutions to form the image multi-scale information extraction part as an example, its structure is the first two-dimensional convolution → the second two-dimensional convolution → the third two-dimensional convolution → the fourth two-dimensional convolution, where the convolution kernel size of each two-dimensional convolution is 3×3, the stride is 1, and the edge padding length is 1. The number of input channels of the first two-dimensional convolution is 3R, and the number of output channels is 64. The number of input channels of the remaining two-dimensional convolutions is 64, and the number of output channels is 64.

[0068] (2.2.3) The image adaptive weight vector extraction part and the multi-scale information extraction part together form the direction attention module, which is used to further process the decomposed feature B to obtain the weighted multi-direction and multi-scale feature set F; the implementation steps are as follows:

[0069] (2.2.3a) Input the decomposed feature B into the adaptive weight vector extraction part in the direction attention module, and output the adaptive weight vector of B Each element in this vector corresponds to the weight of a feature in B. The larger the weight, the greater the influence of the direction feature of the corresponding channel on the classification result, and the network should pay attention to it;

[0070] (2.2.3b) Expand the adaptive weight vector v0 into a tensor with the same size as the decomposed feature B The value of the two-dimensional tensor of each channel is the value of the corresponding element of v0;

[0071] (2.2.3c) Obtain the direction feature B' with adaptive weights according to the tensor V and the decomposed feature B:

[0072] B' = V * B,

[0073] where * represents element-wise multiplication;

[0074] (2.2.3d) Input the direction feature B' into the multi-scale information extraction part in the direction attention module to obtain the output set of D two-dimensional convolutions, that is, the weighted multi-direction and multi-scale feature set Because the convolution kernel of each layer of two-dimensional convolution is 3×3, different convolution layers have different receptive fields, and their corresponding outputs are given different scale information;

[0075] In the adaptive weight vector extraction part of the direction attention module constructed by the present invention, the main purpose is to calculate the adaptive weights of the direction features extracted by the shear wave network, and this module can also be replaced by other attention modules; for the number of layers of the two-dimensional convolution in the image multi-scale information extraction part, this embodiment gives 4 layers, and of course, any number of layers more than 1 layer can also be used to implement it.

[0076] (2.3) Construct a lightweight convolutional neural network, and the specific steps are as follows:

[0077] (2.3.1) Construct a 1×1 convolutional block, and its structure is the first two-dimensional convolution → activation function → Dropout → the second two-dimensional convolution → Dropout; Dropout represents a random inactivation layer; the convolutional kernel sizes of the first and second two-dimensional convolutions are both 1×1; in this embodiment, the convolutional kernel sizes of the two-dimensional convolutions are set to 3×3, the stride is 1, there is no padding, the number of input channels of the first two-dimensional convolution is 64, the number of output channels is 128, the number of input channels of the second two-dimensional convolution is 128, the number of output channels is 64, the GELU activation function is used as the activation function, and the inactivation probability of the random inactivation layer Dropout is 0.5;

[0078] (2.3.2) Connect n 1×1 convolutional blocks constructed in step (2.3.1) in series, where n≥D, and this embodiment preferably takes n as 18; and then connect a global average pooling layer, a fully connected layer, and a softmax layer in series after the last 1×1 convolutional block to obtain a lightweight convolutional neural network for abstract feature extraction of the multi-direction and multi-scale feature set F to obtain the final classification prediction vector where c is the number of categories of images in the training sample set; the implementation steps are as follows:

[0079] (2.3.2a) Use the output f1 of the first layer of two-dimensional convolution in the weighted multi-direction and multi-scale feature set F as the input of the first 1×1 convolutional block of the lightweight convolutional neural network, and output the first response feature

[0080] (2.3.2b) Add the output f d of the d-th layer of two-dimensional convolution in the weighted multi-direction and multi-scale feature set F pixel by pixel with the (d - 1)-th response feature f d-1 ', and use the addition result as the input of the d-th 1×1 convolutional block of the lightweight convolutional neural network to output the d-th response feature map Let d take 2, 3,..., D in turn to obtain the D-th response feature map

[0081] (2.3.2c) Use f D ' as the subsequent input of the lightweight convolutional neural network to obtain the final response feature map

[0082] (2.3.2e) Map f c to a vector through global average pooling

[0083] (2.3.2f) Map l through a fully connected layer to a vector with the same length as the number of classes c where each element of the vector l' corresponds to the predicted probability value of the network output for each class;

[0084] (2.3.2g) Use the softmax function to normalize the vector l' to obtain the classification prediction vector p.

[0085] (2.4) Take the output of the shearlet network as the input of the direction attention module, take the output of the first layer of two-dimensional convolution of the direction attention module as the input of the first 1×1 convolution of the lightweight convolutional neural network, and add the outputs of the remaining D - 1 layers of two-dimensional convolution to the outputs of the first D - 1 layers of the lightweight convolutional neural network to obtain the lightweight network model;

[0086] Step 3. Train the lightweight network model:

[0087] (3.1) Input the z-th sample image x z ” in the set X” of images to be decomposed into the lightweight network model, and the model outputs its classification prediction vector According to and x z ” and the true label y z of the corresponding image sample x z in the training sample set, calculate the prediction loss L of the lightweight network model:

[0088]

[0089] where, represents the probability that the lightweight network model predicts the sample image x z belongs to class i; y z_i is the component of the true label y z of the sample image x z If the true label of the sample image x z is the i-th class, then y z_i = 1, otherwise y z_i = 0;

[0090] (3.2) Use the backpropagation algorithm to update the trainable parameters of the lightweight network model, and use the SGD algorithm as the optimizer to train the network during the training process;

[0091] (3.3) Let \(z = 1, 2, \cdots, m\), and repeat steps (3.1) to (3.2) until all the images in the image set \(X''\) to be decomposed are trained, obtaining a trained lightweight network model;

[0092] Step 4. Input the image to be classified into the trained lightweight network model, obtain the classification prediction vector corresponding to the classified image, and select the category with the largest element value in the prediction vector as the predicted category of the image, completing image classification.

[0093] Example 2: Refer to Figure 3 , the overall steps of the image classification method proposed in this example are the same as those in Example 1. Now, for the process of steps (2.3.2a) to (2.3.2c), taking the case of using a 4-layer two-dimensional convolution to form the image multi-scale information extraction part as an example, the process of obtaining the final response feature map is described in detail as follows:

[0094] 1) Take \(f_1\) in \(F\) as the input of the first \(1\times1\) convolution block of the lightweight CNN, and output the response feature

[0095] 2) Add \(f_1'\) and \(f_2\) pixel by pixel, and use the result as the input of the second \(1\times1\) convolution block of the lightweight CNN, outputting the response feature map

[0096] 3) Add \(f_2'\) and \(f_3\) pixel by pixel, and use the result as the input of the third \(1\times1\) convolution block of the lightweight CNN, outputting the response feature map

[0097] 4) Add \(f_3'\) and \(f_4\) pixel by pixel, and use the result as the input of the fourth \(1\times1\) convolution block of the lightweight CNN, outputting the response feature map

[0098] 5) Take \(f_4'\) as the input of the subsequent lightweight CNN to obtain the final response feature map

[0099] Example 2: The overall steps of the image classification method proposed in this example are the same as those in Example 1. Now, for the process of training the lightweight network model in step 3, specific parameter settings are described in detail as follows:

[0100] a) Construct a stochastic gradient descent (SGD) optimizer to perform supervised training on the network. The parameter settings of SGD are: the initial learning rate is 0.04, the momentum is 0.9, and the learning rate decays by 0.2 times at the 200th and 300th rounds of training respectively;

[0101] b) Input an image \(x\) into the network, and the network outputs the classification prediction vector \(p\) of \(x\);

[0102] c) Calculate the prediction loss L of the deep neural network according to the classification prediction vector p of x and the true label y of x. The loss function is the cross-entropy loss:

[0103]

[0104] where c is the number of image classes in the dataset, and y i is 0 or 1. If the true label of x is the i-th class, then y i = 1; otherwise, y i = 0, and p i is the probability that the sample x belongs to class i;

[0105] d) Use the backpropagation algorithm to update the trainable parameters of the network;

[0106] e) Repeat steps b) to d) until all images in the training dataset are trained;

[0107] f) Input the image to be classified into the trained deep neural network model to obtain the classification prediction vector of the image. Select the class with the largest element value in the prediction vector as the predicted class of the image to complete image classification.

[0108] The effects of the present invention can be further illustrated by the following simulation results:

[0109] 1. Simulation experiment conditions

[0110] The hardware platform used in the present invention is as follows: The CPU uses an eight-core and eight-thread Intel Core i7-9700k with a main frequency of 3.6 GHz and a memory of 64 GB; the GPU uses two Nvidia RTX 3090s with a video memory of 24 GB each. The software platform used is as follows: The operating system uses Ubuntu16.04LTS, the deep learning computing framework uses PyTorch 1.8.1, and the programming language uses Python 3.7.11.

[0111] In this simulation experiment, the three datasets CIFAR10, CIFAR100, and STL10 are used to test the present invention. CIFAR10 and CIFAR100 are two datasets with different difficulties, both containing 50,000 training images and 10,000 test images. Each image is 32×32 in size and is a color image. The difference is that CIFAR10 contains a total of 10 categories, while CIFAR100 contains 100 categories. Therefore, relatively speaking, the classification difficulty of CIFAR100 is greater. The STL10 dataset has a total of 10 categories, containing 5,000 training images and 8,000 test images. Each image is 96×96 in size, which is larger than the CIFAR10 and CIFAR100 datasets. The number of images provided for each category is small, so it is a more challenging dataset.

[0112] 2. Simulation Content and Result Analysis

[0113] The present invention and two existing methods, WaveMix and ViN, are used to test CIFAR10 under the above simulation conditions, and the results are shown in Table 1.

[0114] Table 1 Classification Accuracy of the Present Invention and Comparative Methods on the CIFAR10 Dataset

[0115] Number of parameters Accuracy of the test set WaveMix 2.30M 83.63 ViN 0.53M 65.06 The method of the present invention 0.55M 84.71

[0116] As can be seen from Table 1, when the number of parameters of the present invention is less than or approximately equal to that of the comparative methods, the classification accuracy is relatively high, indicating that the classification effect of the present invention is better than that of the two existing methods. This is because the present invention uses a shearlet network to extract the texture contour features of the image in different directions, has a geometrically invariant representation ability for the image, and the direction attention module further enhances the representation ability of the method for the discriminative features of the image. Therefore, compared with the comparative methods, it has a stronger image feature representation ability and classification ability.

[0117] The present invention and two existing methods, WaveMix and ViN, are used to test CIFAR100 under the above simulation conditions, and the results are shown in Table 2.

[0118] Table 2 Classification Accuracy of the Present Invention and Comparative Methods on the CIFAR100 Dataset

[0119] Number of parameters Accuracy of the test set WaveMix 2.87M 52.93 ViN 0.53M 35.44 The method of the present invention 0.55M 59.25

[0120] As can be seen from Table 1, when classifying the CIFAR100 dataset with the present invention, under the condition that the number of parameters is less than or approximately equal to that of the comparative methods, its classification accuracy is higher than that of the two existing methods.

[0121] The STL10 test was conducted using the present invention and two existing methods, WaveMix and ViN, under the above simulation conditions, and the results are shown in Table 3.

[0122] Table 3 Classification accuracy of the present invention and comparative methods on the STL10 dataset

[0123] Number of parameters Accuracy of the test set WaveMix 2.87M 61.62 ViN 0.53M 43.55 The method of the present invention 0.55M 65.32

[0124] As can be seen from Table 1, when classifying the STL10 dataset using the present invention, with the number of parameters less than or approximately equal to that of the comparative methods, its classification accuracy is higher than that of the two existing methods.

[0125] In summary, the method proposed by the present invention, through the design of the shearlet network, direction attention module, and lightweight convolutional neural network, can achieve high-accuracy image classification with fewer network parameters, while reducing the computational amount of the classification model and accelerating the model classification speed, and has advantages in classification performance compared with existing methods.

[0126] The above simulation analysis proves the correctness and effectiveness of the method proposed by the present invention.

[0127] The parts not detailed in the present invention belong to the common general knowledge of those skilled in the art.

[0128] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Obviously, for professionals in the field, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention. However, these modifications and changes based on the idea of the present invention are still within the scope of protection of the claims of the present invention.

Claims

1. An image classification method based on a shear wave network and a direction attention mechanism, characterized in that Construct a lightweight network model that can extract texture contours and geometric features in different directions of an image, and can adaptively increase the weights of the directional features beneficial for classification in the image, use it to extract the features for classification in the image, and output a classification prediction vector; specifically including the following steps: (1) Construct a training sample set, and perform preprocessing and padding on it to obtain a set of images to be decomposed: (1.1) Construct a training sample set with \(m\) image samples and \(c\) categories by using public data sets or capturing images of specific scenarios according to the needs of classification scenarios. Its corresponding true label is Among them, \(H\) z and \(W\) z are the height and width of the \(z\)-th image sample respectively. (1.2) Normalize all pixels of each image in the training sample set X to obtain a normalized image with pixel values in the interval [0, 1]; (1.3) Horizontally flip, edge pad, and randomly crop the normalized image to obtain the preprocessed training sample set where H and W represent the height and width of the cropped image samples; (1.4) Perform edge padding of length p on the preprocessed training sample set, where p ≥ 0, to obtain the set of images to be decomposed (2) Construct a lightweight network model composed of a shearlet network, a directional attention module, and a lightweight convolutional neural network: (2.1) Construct a shearlet network with a 3-layer filter structure, and the implementation steps are as follows; (2.1.1) Create a Gaussian low-pass filter As the filter for the first layer of the shearlet network, it is used to perform low-pass filtering on the input image to obtain the low-frequency information of the image; (2.1.2) Create a shearlet filter bank ψ according to the shearlet transform as the filter of the second layer of the shearlet network: Among them, a represents the filter parameter that controls the scale of the shearlet transform, and s represents the filter parameter that controls the direction of the shearlet transform; (2.1.3) Create 9 shearlet filter banks identical to ψ according to the shearlet transform, and cascade each of them with each filter in the shearlet filter bank ψ created in step (2.1.2). These 9 shearlet filter banks together serve as the filter of the third layer of the shearlet network; (2.1.4) The filters of the first, second, and third layers together form the shearlet network, which is used to obtain the decomposition feature B of the image to be decomposed; (2.2) Construct a directional attention module, and the specific steps are as follows: (2.2.1) Select the channel attention module of SENet to form the part for extracting the image adaptive weight vector, and its structure is GAP→FC1→activation function→FC2→activation function, where GAP represents global average pooling, FC1 represents the fully connected layer for compressing the vector length, and FC2 represents the fully connected layer for expanding the vector length; (2.2.2) Use D two-dimensional convolutions to form the part for extracting image multi-scale information, D≥1, and its structure is a series structure, that is: the first two-dimensional convolution→the second two-dimensional convolution→···→the Dth two-dimensional convolution; (2.2.3) The part for extracting the image adaptive weight vector and the part for extracting multi-scale information together form the directional attention module, which is used to further process the decomposition feature B to obtain a weighted multi-directional and multi-scale feature set F; (2.3) Construct a lightweight convolutional neural network, and the specific steps are as follows: (2.3.1) Construct a 1×1 convolution block, and its structure is the first two-dimensional convolution→activation function→Dropout→the second two-dimensional convolution→Dropout; Dropout represents the random inactivation layer; the convolution kernel sizes of the first and second two-dimensional convolutions are both 1×1; (2.3.2) Connect the 1×1 convolutional blocks constructed in n steps (2.3.1) in series, where n ≥ D, and then connect a global average pooling layer, a fully connected layer, and a softmax layer in sequence after the last 1×1 convolutional block to obtain a lightweight convolutional neural network for abstracting feature extraction from the multi-direction and multi-scale feature set F to obtain the final classification prediction vector. where c is the number of categories of images in the training sample set; (2.4) Take the output of the shearlet network as the input of the directional attention module, the output of the first layer of two-dimensional convolution of the directional attention module as the input of the first 1×1 convolution of the lightweight convolutional neural network, and add the outputs of the remaining D - 1 layers of two-dimensional convolutions to the outputs of the first D - 1 layers of the lightweight convolutional neural network to obtain the lightweight network model; (3) Train the lightweight network model: (3.1) Input the z-th sample image x in the set X” of images to be decomposed z” into the lightweight network model, and the model outputs its classification prediction vector p xz” ; According to p xz” and x z” corresponding to the true label y of the image sample x in the training sample set z Calculate the prediction loss L of the lightweight network model: z ​ Among them, represents the probability that the lightweight network model predicts the sample image x z belongs to class i; y z_i is the sample image x z true label y z component. If the true label of the sample image x z is the i-th class, then y z_i = 1, otherwise y z_i = 0; (3.2) Update the trainable parameters of the lightweight network model using the backpropagation algorithm, and use the SGD algorithm as the optimizer to train the network during the training process; (3.3) Let z = 1, 2,..., m, and repeat steps (3.1) to (3.2) until all images in the set X" of images to be decomposed are trained, obtaining a trained lightweight network model; (4) Input the image to be classified into the trained lightweight network model, obtain the classification prediction vector corresponding to the classified image, and select the class with the largest element value in the prediction vector as the predicted class of the image to complete image classification.

2. The method according to claim 1, characterized in that: In step (2.1.4), the decomposition feature B of the image to be decomposed is obtained as follows: (2.1.4a) Input the image \(x''\) to be decomposed into the low-pass filter \(\varphi\) of the first layer of the shearlet network for filtering to obtain the response \(b_1\), and then perform central cropping on it to obtain the output of the first layer of the shearlet network (2.1.4b) Input the image \(x''\) to be decomposed into the shearlet filter bank \(\psi\) of the second layer of the shearlet network for decomposition to obtain a set of directional feature maps where \(b\) 2_0 is the low-frequency feature map of the image, and \(b\) 2_1 to \(b\) 2_8 are the feature maps of 8 different directions of the image. \(\alpha\) represents the index of the directional feature map obtained by the shearlet filter bank of the second layer; then perform central cropping on \(b_2\) to obtain the output of the second layer of the shearlet network (2.1.4c) Input the 9 feature maps in the set b2 of directional feature maps into the 9 filter banks in the third layer of the shearlet network for decomposition, obtaining the decomposition feature set of the third layer of the shearlet network where β represents the index of the filter bank in the third layer, and αβ represents the index of the directional feature map obtained from the shearlet filter bank in the third layer; it contains 9 2 = 81 features in different directions, and then b3 is centrally cropped to obtain the output of the third layer of the shearlet network (2.1.4d) Concatenate the outputs of the first layer, the second layer, and the third layer obtained in steps (2.1.4a)-(2.1.4b) along the channels to obtain the output of the shearlet network where R = 3×(1 + 9 + 9 2 ) is the number of filters used for decomposition in the shearlet network.

3. The method according to claim 1, characterized in that: In step (2.2.3), the decomposed feature B is further processed to obtain the weighted multi-directional and multi-scale feature set F. The implementation steps are as follows: (2.2.3a) Input the decomposed feature B into the adaptive weight vector extraction part in the direction attention module, and output the adaptive weight vector of B (2.2.3b) Expand the adaptive weight vector v0 into a tensor with the same size as the decomposed feature B where the values of the two-dimensional tensors in each channel are the corresponding element values of v0; (2.2.3c) Obtain the directional feature B' with adaptive weights according to the tensor V and the decomposed feature B: B' = V * B, where * represents element-wise multiplication; (2.2.3d) Input the direction feature B' into the multi-scale information extraction part of the direction attention module to obtain the output set of the 2D convolution of layer D, that is, the weighted multi-direction and multi-scale feature set 4. The method according to claim 3, wherein: In step (2.3.2), abstract feature extraction is performed on the multi-directional and multi-scale feature set F to obtain the final classification prediction vector p. The implementation steps are as follows: (2.3.2a) Use the output f1 of the first-layer two-dimensional convolution in the weighted multi-directional and multi-scale feature set F as the input to the first 1×1 convolution block of the lightweight convolutional neural network, and output the first response feature (2.3.2b) Add the output \(f\) of the 2D convolution in the \(d\)-th layer in the weighted multi-directional and multi-scale feature set \(F\) d to the \(d - 1\) response feature \(f\) d-1 pixel by pixel, and use the addition result as the input of the \(d\)-th \(1\times1\) convolution block of the lightweight convolutional neural network, and output the \(d\)-th response feature map Let \(d\) take values of 2, 3, …, \(D\) in sequence to obtain the \(D\)-th response feature map (2.3.2c) Use f D ' as the input for the subsequent lightweight convolutional neural network to obtain the final response feature map (2.3.2e) Map f c to a vector through global average pooling (2.3.2f) Map l to a vector with the same length as the number of classes c through a fully connected layer Each element of the vector l' corresponds to the predicted probability value of the network output for each category; (2.3.2g) Use the softmax function to normalize the vector l' to obtain the classification prediction vector p.

Citation Information

Patent Citations

  • Remote sensing image classification method based on attention mechanism deep Contourlet network

    CN110728224A

  • Method for objective, noninvasive staging of diffuse liver disease from ultrasound shear-wave elastography

    WO2019191059A1