Image Classification Method and System Based on High-Confidence Local Feature and Global Feature Learning

The network is constructed by deeply separating the convolutional layer and nested downsampling layer, combining multi-scale feature extraction and high confidence local area of ​​interest discrimination, and solving the problem of high inter-class similarity and large in-class differences in remote sensing image scene classification, improving classification accuracy and speed.

CN115346071BActive Publication Date: 2025-07-04NANJING UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211002091.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-20
Publication Date
2025-07-04
Estimated Expiration
2042-08-20

AI Technical Summary

Technical Problem

There are problems in the classification of remote sensing image scenes with high similarity between classes, large differences within classes, and large scale differences. The existing methods are difficult to effectively deal with, especially the lack of effective mechanisms in local feature selection, which affects the classification accuracy.

Method used

The network is constructed by deep separation convolutional layer and nested downsampling layer. Through multi-scale feature extraction and non-maximum suppression, a high-confidence local areas of interest are generated. The network is trained in combination with the sort consistency loss function, and local and global features are aggregated for classification.

Benefits of technology

The accuracy and speed of remote sensing image classification are improved, especially the classification performance on small sample images, and the problems of high similarity between classes and large differences within classes are effectively dealt with.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115346071B_ABST
    Figure CN115346071B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for image classification by learning high-confidence local features and global features. The method includes: adopting a stacked form of depthwise separable convolutional layers, nesting downsampling layers for feature extraction; extracting multi-scale features to form a feature tensor; using sliding windows with different strides and non-maximum suppression methods to remove redundant regions and generate candidate local regions of interest; calculating the probability that a candidate local region of interest belongs to a class label as a confidence value, constructing a ranking consistency loss function based on the confidence value to train the network, and performing ranking on the candidate local regions of interest to discriminate high-confidence regions of interest; aggregating the features of high-confidence local regions of interest and global features to obtain a feature concatenation tensor; and performing final image classification according to the aggregated features. The present invention can effectively handle classification problems with high inter-class similarity and large intra-class differences, and improve the classification accuracy and speed of small-sample images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to remote sensing scene classification technology, in particular to a picture classification method and system for learning high-confidence local features and global features. Background Art

[0002] Remote sensing scene classification belongs to a direction of picture classification, which aims to attach semantic labels to scene pictures through algorithms and is one of the iconic tasks in the field of computer vision. In recent years, it has been a hot research field in both computer vision and pattern recognition. Scene classification for remote sensing images is of great significance for various civilian applications such as geographical information system mapping, agriculture, traffic planning, and navigation. Due to the wide spatial coverage of remote sensing images and the usually numerous objects in the images, there are problems of high inter-class similarity, large intra-class difference, and large scale difference. The above factors make remote sensing image scene classification a challenging task.

[0003] With the rapid development of deep learning, many remote sensing image classification methods based on deep learning have been proposed. Remote sensing scene images are different from natural images usually taken from a horizontal perspective. Remote sensing scene images are usually bird's-eye views, which means that the images always contain different types of ground objects, and the complex semantic information increases the difficulty of scene classification. Penatti et al. [Penatti O AB, Nogueira K, Dos Santos J A. Do deep features generalize from everyday objects to remote sensing and aerial scenes domains? [C] / / Proceedings of the IEEE conference on computer vision and pattern recognition workshops. 2015:44-51.] studied the generalization ability of pre-trained convolutional neural networks in remote sensing scene classification. Nogueira K et al. [Nogueira K, Penatti O A B, dos Santos J A. Towards better exploiting convolutional neural networks for remote sensing scene classification[J]. Pattern Recognition, 2017, 61:539-556.] used four different strategies to classify remote sensing scenes based on different convolutional neural network models in order to improve the classification accuracy. Finally, it was concluded that the best effect was obtained by fine-tuning. Chen Yadang et al. [An image classification method based on Resnet50 combined with attention mechanism and feature pyramid [P]. Chinese Patent: CN114494805A, 2022-05-13.] adopted a channel and spatial attention fusion mechanism to focus on the detailed information of the target, and used an optimization algorithm of a pyramid layer three-classifier. After the full connection of each feature layer, classification processing was carried out, and a probability value was output for each category; the classification result with the largest probability value was taken as the final output prediction value. This method improves the feature extraction effect of ResNet-50, pays attention to the combination of local features and global features, and achieves better classification results. However, for the selection of local features, there is no good selection mechanism. Summary of the Invention

[0004] The object of the present invention is to propose a method and system for image classification that learns high-confidence local features and global features, fully considering the impact of the aggregation of local features and global features on the classification task, capable of effectively handling image classification problems with high inter-class similarity, large intra-class differences, and large scale differences, and having excellent classification performance.

[0005] The technical solution for achieving the object of the present invention is as follows: In the first aspect, the present invention provides a method for image classification that learns high-confidence local features and global features, including the following steps:

[0006] In the first step, a deep network composed of data preprocessing and a depthwise separable convolution module is used to extract image features;

[0007] In the second step, taking the features extracted by the depthwise separable convolution feature learning module in the first step as input, a three-layer pyramid structure is constructed using depth convolution operations to extract multi-scale deep features;

[0008] In the third step, on the multi-scale deep features, according to different scales, sliding windows with different strides are adopted, and non-maximum suppression is performed on the multi-scale feature regions to reduce regional redundancy, generating a list representing the multi-scale feature regions, and extracting a specified number of candidate local regions of interest as the input of the high-confidence local region of interest discrimination and feature learning module;

[0009] In the fourth step, the sizes of the local region of interest features are normalized to the same standard, and then through the depthwise separable convolution feature learning module, the probability of each region being a class label is calculated as the confidence value, and a list of confidence values is output; according to this confidence value, a ranking consistency loss function is adopted to adjust the network training, re-rank the list of candidate local region of interest feature regions to make it consistent with the ranking of the list of confidence values, and extract the top M high-confidence local regions of interest;

[0010] In the fifth step, the M high-confidence local regions of interest extracted are adjusted into tensors with the same size as the context features through pooling operations, and are concatenated with the context features for feature aggregation;

[0011] In the sixth step, the features after aggregating the high-confidence local region of interest features and the global context features are connected to a fully connected layer and a Softmax classifier for final classification.

[0012] In the second aspect, the present invention provides a system for image classification that learns high-confidence local features and global features, including:

[0013] A depthwise separable convolution feature learning module, which adopts a stacked form of depthwise separable convolution layers, nests downsampling layers, and performs feature extraction through pre-training;

[0014] The multi-scale feature learning module uses convolutions of three different scales to extract multi-scale features and form a feature tensor.

[0015] The candidate local region of interest generation module uses sliding windows with different strides and non-maximum suppression methods to remove redundant regions and generate candidate local regions of interest.

[0016] The high-confidence local region of interest discrimination and feature learning module calculates the probability that a local region of interest candidate belongs to a class label as a confidence value, constructs a ranking consistency loss function based on the confidence value to train the network, and ranks the candidate local regions of interest to discriminate high-confidence regions of interest.

[0017] The feature aggregation module aggregates the features of the high-confidence local regions of interest and the global features to obtain a feature concatenation tensor.

[0018] The classification module performs the final image classification based on the aggregated features.

[0019] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0020] In a fourth aspect, the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0021] Compared with the prior art, the present invention has the following remarkable features: (1) By stacking depthwise separable convolutional layers and embedding downsampling layers to construct a network, and extracting features through pre-training, the extracted features are more abundant; (2) The multi-scale information feature extraction module constructs a pyramid structure to obtain features of different spatial scales, and the extracted multi-scale feature tensor is more accurate; (3) By adopting non-maximum suppression loss, redundancy can be removed; (4) Through the high-confidence local region of interest discrimination and feature learning module, a ranking consistency loss function is constructed to discriminate high-confidence local region of interest features, strengthening the acquisition of local effective information and aggregating with context features, which can effectively handle classification problems and improve the classification accuracy of small-sample images.

[0022] The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a structural diagram of the method of the present invention.

[0024] Figure 2 is a structural diagram of a depthwise separable convolution block.

[0025] Figure 3It is the downsampling structure diagram of the depthwise separable convolution module.

[0026] Figure 4 It is the structure diagram of the candidate local region of interest generation module.

[0027] Figure 5 It is the classification confusion matrix diagram of the method of the present invention for AID30.

[0028] Figure 6 It is the classification confusion matrix diagram of the method of the present invention for NWPU-RESISC45. Detailed implementation manners

[0029] Compared with the existing methods in the background art, the present invention proposes a method and system for image classification of high-confidence local features and global features learning. Using a pre-trained convolutional network as a feature extractor, combined with a pyramid object recognition structure, a sliding window with multiple scales and aspect ratios is constructed to obtain a local region of interest feature map. Through a high-confidence local region of interest discrimination and feature learning module, high-confidence local regions are located, and then the extracted high-confidence local region of interest feature map is aggregated with the context semantics and sent into the convolutional network feature extractor for final classification, thereby improving the classification accuracy.

[0030] Next, in combination with Figure 1 , the implementation process of the present invention will be described in detail.

[0031] A method for image classification of high-confidence local features and global features learning includes the following steps:

[0032] In the first step, a deep network composed of data preprocessing and a depthwise separable convolution module is used to extract image features, including feature extraction operations such as depthwise separable convolution, downsampling, and activation functions. The specific process is as follows:

[0033] (1) Perform data augmentation on the original image, crop it to a size of 224×224, randomly flip the image horizontally with a probability of 0.5, and then convert the image data into a standard normal distribution mean = [0.485, 0.456, 0.406], std = [0.229, 0.224, 0.225]. The normalization formula is as follows:

[0034] input[channel] = (input[channel] - mean[channel]) / std(channel)

[0035] For an image of 224×224×3, pass through a convolutional layer with a convolutional kernel size of 4*4, 96 convolutional filters, and a stride of 4, and perform layer normalization to achieve 4-fold downsampling operation. The image is adjusted to 56×56×96, where 96 is the feature dimension.

[0036] (2) The composition of a depthwise separable convolution block is as follows:

[0037] As Figure 2 , first, there is a depthwise separable convolution layer with a kernel size of 7*7, a stride of 1, and a padding of 3, followed by layer normalization; then, there is a convolution layer with a kernel of 1*1, followed by a Gaussian Error Linear Unit (GELU) activation function to achieve dimensionality increase; then, it passes through another convolution layer with a kernel of 1*1, followed by channel scaling and random channel dropout to restore the initial dimension, and is added to the initial feature tensor as the final output.

[0038] (3) The composition of the downsampling layer is as follows:

[0039] As Figure 3 , first, through layer normalization, and then connected to a convolution layer with a kernel of 2*2 and a stride of 2 to achieve 2x downsampling.

[0040] (4) A total of four stages of depthwise separable convolution blocks are stacked, and a downsampling layer is embedded between every two stages of stacking. In the first stage, the depthwise separable convolution blocks are stacked 3 times, with an output size of 56×56×96; after passing through a downsampling layer, it enters the second stage, where the depthwise separable convolution blocks are stacked 3 times, with an output size of 28×28×192; after passing through the downsampling layer, it enters the third stage, where the depthwise separable convolution blocks are stacked 9 times, with an output size of 14×14×384; after passing through the downsampling layer, it enters the fourth stage, where the depthwise separable convolution blocks are stacked 3 times, with an output size of 7×7×768.

[0041] The output of this step is used as the input to both the multi-scale feature learning module and the global feature extraction module.

[0042] In the second step, using the features extracted by the depthwise separable convolution feature learning module in the first step as the input, a three-layer pyramid structure is constructed using depth convolution operations to extract multi-scale deep features, as Figure 4 , and the specific process is as follows:

[0043] (1) For the input features, first, pass through a convolution layer with a kernel size of 3*3, a stride of 1, and 128 convolution kernels to obtain a first-layer feature map with a size of 7×7×128 and a receptive field of 3×3.

[0044] (2) For the feature map with a size of 7×7×128, pass through a convolution layer with a kernel size of 3*3 and a stride of 2 to obtain a second-layer feature map with a size of 14×14×128 and a receptive field of 5×5.

[0045] (3) For the feature map with a size of 7×7×128, pass through a convolution with a kernel size of 3*3 and a stride of 2 to obtain a third-layer feature map with a size of 28×28×128 and a receptive field of 9×9.

[0046] (4) The three-layer feature maps respectively pass through a 1×1 convolutional layer to achieve horizontal connection and obtain the foreground information score.

[0047] Thirdly, on the multi-scale deep features, according to different scales, sliding windows with different strides are adopted, and non-maximum suppression is performed on the multi-scale feature regions to reduce regional redundancy, generating a list representing the multi-scale feature regions, and extracting a specified number of candidate local regions of interest as the input of the high-confidence local region of interest discrimination and feature learning module. The specific process is as follows:

[0048] (1) Based on the multi-scale feature layers generated by the multi-scale feature learning module, sliding windows with different strides are assigned, and a pixel area of 48 2 , 96 2 , 192 2 is set for each pixel point, and the strides correspond to 32, 64, and 128 respectively, and 9 types of region boxes with aspect ratios of 1 / 1, 3 / 2, and 2 / 3 are generated. Among them, 1 type of pixel area corresponds to generating 3 region boxes with different aspect ratios, and a total of 9 types of region boxes are generated for 3 types of pixel areas, and they are mapped to the corresponding positions of the feature map, generating a list of the amount of information of a specific number of local regions of interest features.

[0049] (2) Non-maximum suppression is performed according to the amount of information to remove duplicate region boxes. First, the list L of the amount of information of the specified number of local regions of interest features output in step (1) is sorted according to the amount of information, and the region box with the largest amount of information is taken out and stored in the final reserved list D; calculate the IOU between the remaining region boxes in L and the current region box, and delete them when the IOU between the two is greater than the fixed threshold u. After such screening, until the first 6 local regions of interest {R1, R2... R6} and the corresponding information {I1, I2... I6} are stored in the final reserved list D. IOU represents the intersection over union of the current region box and other region boxes of the same target, and is defined as:

[0050]

[0051] where area(·) represents the area calculation operator of the set, b i and b j represent two different region boxes.

[0052] Fourthly, the sizes of the local regions of interest features are normalized to the same standard, and then through the depthwise separable convolutional feature learning module, the probability of each region as a class label is calculated as the confidence value, and a list of confidence values is output. According to this confidence value, a sorting consistency loss function is adopted to adjust the network training, and the list of candidate local regions of interest features is re-sorted to make it consistent with the sorting of the list of confidence values, and the first M high-confidence local regions of interest are extracted. The specific process is as follows:

[0053] (1) Upsample the M regions {R1, R2... R6} using bilinear interpolation method to convert them into the size of 224×224. Through data preprocessing and depthwise separable convolution module, calculate the probability that each local region of interest is the true value category as the confidence value, and the confidence levels {C1, C2... C6}. At the same time, optimize this step by minimizing the cross-entropy loss between each class label and the confidence value, and calculating the cross-loss function of the global image X:

[0054]

[0055] Where C(·) is the confidence value calculation function. The first part of the formula is the sum of the cross-losses of all regions, and the second part is the cross-entropy loss of the entire image.

[0056] (2) According to the confidence levels {C1, C2... C6} of each region of interest in (1), and the information content of the M regions of interest in the candidate local regions of interest, construct a ranking consistency loss function. The specific rules of this loss function are as follows. Let the information content ranking be {I1, I2... I6}. When I s > I i And C s > C i , the corresponding label is 0. When the input is the opposite, that is, I s < I i And C s > C i , the corresponding label is 1. The ranking consistency loss function is defined as follows:

[0057]

[0058] Where f(·) uses the hinge loss function: f(x) = max{1 - x, 0}.

[0059] (3) Through the ranking consistency loss function in (2), train the network to guide the reordering of the candidate local regions of interest feature list to make it consistent with the confidence value column ranking, and extract the top 6 high-confidence local regions of interest.

[0060] The fifth step is to adjust the 6 high-confidence local regions of interest extracted into tensors with the same size as the context features through pooling operations, and splice them with the context features for feature aggregation. The specific process is as follows:

[0061] (1) Adjust the outputs {R1, R2... R6} of the high-confidence local region of interest discrimination and feature learning module into 224×224×786.

[0062] (2) Use the output of the first step as the input for global feature extraction. After global average pooling and layer normalization operations, it is output as context features.

[0063] (3) Aggregate the high-confidence local features in (1) with the context features in (2) as the input for the classification module.

[0064] In the sixth step, the features obtained by aggregating the high-confidence local region of interest features and the global features are connected to a fully connected layer and a Softmax classifier for final classification. The specific process is as follows:

[0065] (1) Use the output of the fifth step as the input and connect it to a fully connected layer.

[0066] (2) After the output in (1), connect a Softmax classifier to predict the final classification result.

[0067] Based on the same concept, the present invention also provides an image classification system for learning high-confidence local features and global features, including:

[0068] A depthwise separable convolution feature learning module, which adopts a stacked form of depthwise separable convolution layers, nests downsampling layers, and performs feature extraction through pre-training;

[0069] A multi-scale feature learning module, which adopts convolutions of three different scales to extract multi-scale features and form a feature tensor;

[0070] A candidate local region of interest generation module, which uses sliding windows with different strides and non-maximum suppression methods to remove redundant regions and generate candidate local regions of interest;

[0071] A high-confidence local region of interest discrimination and feature learning module, which calculates the probability that a local region of interest candidate belongs to a class label as a confidence value, constructs a ranking consistency loss function to train the network based on the confidence value, and ranks the candidate local regions of interest to discriminate high-confidence regions of interest;

[0072] A feature aggregation module, which aggregates the extracted high-confidence local region of interest features and global features to obtain a feature concatenation tensor;

[0073] A classification module, which performs final image classification based on the aggregated features.

[0074] The specific implementation manners of the above-mentioned modules correspond to the contents of the first to sixth steps of the foregoing image classification method, and will not be elaborated here.

[0075] Furthermore, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the foregoing image classification method.

[0076] Furthermore, the present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the foregoing picture classification method.

[0077] The present invention fully establishes a multi-scale local region of interest high-confidence ranking and selection mechanism, effectively integrates global context feature information, uses a pre-trained network for feature extraction, enhances the feature learning ability, can effectively handle classification problems with high inter-class similarity and large intra-class differences, and improves the classification accuracy and classification speed of small-sample pictures.

[0078] The effects of the present invention can be further illustrated by the following simulation experiments:

[0079] Simulation conditions

[0080] Two sets of optical remote sensing image data are used in the simulation experiments: the AID dataset and the NWPU-RESISC45 dataset. All images in the AID dataset are released by Wuhan University and Huazhong University of Science and Technology, and it contains a total of 30 scene images, about 220 - 420 images for each class, with a total of 10,000 images, and the training ratio is 50%. All images in the NWPU-RESISC45 dataset are created by Northwestern Polytechnical University. This dataset contains 31,500 images, covering 45 scene classes, with 700 images for each class. The data covers more than 100 countries and regions around the world and is of a relatively large scale. Except for islands, lakes, mountains, etc. with relatively low spatial resolution, the spatial resolution of most scene classes ranges from 0.2 to 30m. At the same time, this dataset takes into account the influence of different weather, seasons, lighting and other natural conditions, and has rich image variations in terms of background, occlusion, etc. According to the training ratio of 20%, the training set and the test set contain 6300 and 25200 images respectively, and the images are 256×256 pixels in size. In the experiment, all images in the AID and NWPU-RESISC45 datasets are adjusted to a size of 224*224. The overall classification accuracy is used as the evaluation index for the two sets of experiments. In addition, the comparison methods include: basic convolutional neural networks, such as: AlexNet, GoogLeNet, VGG_16, the method of "bag of color features" (Bag of Color Features) on the basic convolutional neural network, the method of multi-scale feature aggregation (Self-attention-based deep feature fusion, SAFF), Capsule Network (CapsNet), and a deep separable convolutional feature learning network without combining high-confidence local features.

[0081] In the experiment, the optimizer of the deep separable convolutional feature learning network is the Adam optimizer, with an initial learning rate of 0.001, which is divided by 10 after 60 epochs, and the momentum and weight decay are 0.9 and 1e-4 respectively. In addition, the network trains the model within the first 50 epochs of AID and the first 50 epochs of NWPU-RESISC45. Other network hyperparameter configurations are summarized in Table 1. The simulation experiments are all completed using Python3.8 + pytorch1.8 + cuda11.2 under the Linux operating system.

[0082] Table 1 Network Hyperparameter Configuration

[0083]

[0084]

[0085] Analysis of Simulation Experiment Results

[0086] Tables 2 - 3 show the classification accuracies (%) of the method of the present invention for the NWPU-RESISC45 and AID data sets in the simulation experiment.

[0087] Table 2 Classification Results of Different Methods for the AID Data Set

[0088]

[0089] Table 3 Classification Results of Different Methods for the NWPU-RESISC45 Data Set

[0090]

[0091]

[0092] From the experimental results, it can be seen that the classification accuracy of the two data sets can be significantly improved by using the method of the present invention. On the AID data set, the classification accuracy of the method of the present invention is 95.70 ± 0.17%, and the classification confusion matrix obtained by the method of the present invention is as Figure 5 shown. Compared with other methods, the method of the present invention has a better classification effect on the two classes with large scale changes, namely basketball court and school, which benefits from the regional information extraction module combined in the present invention, and this module can extract multi-scale feature information more accurately. On the NWPU-RESISC45 data set, the average precision of the method of the present invention is 91.93 ± 0.18%, and the classification confusion matrix obtained by the method of the present invention is as Figure 6 shown. Compared with other methods, the method of the present invention can obtain better classification results, mainly due to the feature extraction of the depthwise separable convolutional neural network and the combination of high-confidence local region features and global features. The above results fully show that the method of the present invention can effectively learn the feature information of remote sensing images and has high classification performance.

Claims

1. A method for image classification by learning high-confidence local features and global features, characterized in that It includes the following steps: In the first step, a deep network composed of a data preprocessing and a depthwise separable convolution module is used to extract image features; In the second step, taking the features extracted by the depthwise separable convolution feature learning module in the first step as the input, a three-layer pyramid structure is constructed by using depth convolution operations to extract multi-scale deep features; In the third step, on the multi-scale deep features, according to different scales, sliding windows with different strides are adopted, and non-maximum suppression is performed on the multi-scale feature regions to reduce regional redundancy, generating a list representing the multi-scale feature regions, and extracting a specified number of candidate local regions of interest as the input of the high-confidence local region of interest discrimination and feature learning module; In the fourth step, the sizes of the local regions of interest features are normalized to the same standard, and then through the depthwise separable convolution feature learning module, the probability of each region as a class label is calculated as the confidence value, and a list of confidence values is output; according to this confidence value, a sorting consistency loss function is adopted to adjust the network training, re-sorting the list of candidate local regions of interest features to make it consistent with the sorting of the list of confidence values, and extracting the top M high-confidence local regions of interest; In the fifth step, the M high-confidence local regions of interest extracted are adjusted into tensors with the same size as the context features through pooling operations, and are concatenated with the context features for feature aggregation; In the sixth step, the features after aggregating the high-confidence local region of interest features and the global context features are connected to a fully connected layer and a Softmax classifier for final classification.

2. The image classification method based on learning of high-confidence local features and global features according to claim 1, wherein In the first step, a deep network composed of a data preprocessing and a depthwise separable convolution module is used to extract image features, including feature extraction operations such as depthwise separable convolution, downsampling, and activation functions. The specific process is as follows: (1) First, an initial preprocessing operation is performed on the image H×W×N, where H represents the height of the picture, W represents the width of the picture, and N represents the number of channels of the picture; through a convolutional layer with a kernel size of 4*4, C convolutional kernels, and a stride of 4, and layer normalization, a 4-fold downsampling operation is achieved, and the image is adjusted to H / 4×W / 4×C, where C is the feature dimension; (2) The composition of a depthwise separable convolution block is as follows: First, there is a depthwise separable convolutional layer with a kernel size of 7*7, a stride of 1, and a padding of 3, and layer normalization is performed; then there is a convolutional layer with a kernel of 1*1, and a Gaussian error linear unit activation function is connected to achieve dimensionality increase; then it passes through a convolutional layer with a kernel of 1*1, and channel scaling and random channel dropout are connected to restore the initial dimension, and it is added to the initial feature tensor as the final output; (3) The composition of the downsampling layer is as follows: First, through layer normalization, and then connected to a convolutional layer with a kernel of 2*2 and a stride of 2 to achieve 2-fold downsampling; (4) A total of four stages of stacking of depthwise separable convolution blocks are performed, and a downsampling layer is embedded between every two stages of stacking; in the first stage, the depthwise separable convolution blocks are stacked 3 times, and the output size is H / 4×W / 4×C; After passing through a downsampling layer, it enters the second stage, and the depthwise separable convolution blocks are stacked 3 times, and the output size is H / 8×W / 8×2C; After passing through the downsampling layer, it enters the third stage, where the depthwise separable convolution blocks are stacked 9 times, and the output size is H / 16×W / 16×4C; After passing through the downsampling layer, it enters the fourth stage, where the depthwise separable convolution blocks are stacked 3 times, and the output size is H / 32×W / 32×8C; The output of step (4) serves as the input to both the multi-scale feature learning module and the global feature extraction module.

3. The method for image classification by learning high-confidence local features and global features according to claim 2, wherein Taking the features extracted by the first-step depthwise separable convolution feature learning module as the input, a three-layer pyramid structure is constructed using depth convolution operations to extract multi-scale deep features. The specific process is as follows: (1) For the input features, first pass through a convolution layer with a kernel size of 3*3, a stride of 1, and a channel number of F to obtain a first-layer feature map with a size of H / 32×W / 32×F and a receptive field of 3×3; (2) For the feature map with a size of H / 32×W / 32×F, pass through a convolution layer with a kernel size of 3*3 and a stride of 2 to obtain a second-layer feature map with a size of H / 16×W / 16×F and a receptive field of 5×5; (3) For the feature map with a size of H / 16×W / 16×F, pass through a convolution layer with a kernel size of 3*3 and a stride of 2 to obtain a third-layer feature map with a size of H / 8×W / 8×F and a receptive field of 9×9; (4) The three-layer feature maps then pass through 1*1 convolution layers respectively to achieve horizontal connection and obtain the foreground information score.

4. The image classification method based on learning of high-confidence local features and global features according to claim 3, wherein In the third step, on the multi-scale deep features, according to different scales, sliding windows with different strides are adopted, and non-maximum suppression is applied to the multi-scale feature regions to reduce regional redundancy, generating a list representing the multi-scale feature regions. Extract a specified number of candidate local regions of interest as the input to the high-confidence local region of interest discrimination and feature learning module. The specific process is as follows: (1) Based on the multi-scale feature layers generated by the multi-scale feature learning module, sliding windows with different strides are assigned, and the pixel area of 48 is set for each pixel point 2 , 96 2 , 192 2 . The strides correspond to 32, 64, and 128 respectively, and the aspect ratios are 1 / 1, 3 / 2, and 2 / 3; one of the pixel areas corresponds to generating 3 region boxes with different aspect ratios, and a total of 9 types of region boxes are generated for 3 pixel areas, and they are mapped to the corresponding positions of the feature map to generate a list of local region of interest feature information with a specific quantity; (2) Apply non-maximum suppression according to the information content to remove duplicate region boxes; first, sort the information content list L of a specific number of local regions of interest feature regions output in step (1), take out the region box with the largest information content, and store it in the final retained list D; Calculate the IOU between the remaining region boxes in L and the current region box, and delete them when the IOU between the two is greater than the fixed threshold u; perform such screening until the first M candidate local regions of interest {R1, R2... R M} and the corresponding information {I1, I2... I M} are stored in the final retained list D; IOU represents the intersection over union of the current region box of the same object and other region boxes, and is defined as: Among them, area(·) represents the area calculation operator of a set, and b i and b j represent two different bounding boxes.

5. The method for classifying pictures by learning high-confidence local features and global features according to claim 4, characterized in that, In the fourth step, standardize the sizes of the local regions of interest features to the same standard, and then through depthwise separable convolution feature learning, calculate the probability of each region as the class label as the confidence value, and output a list of confidence values; based on this confidence value, adopt a ranking consistency loss function to adjust the network training, re-rank the list of candidate local regions of interest features to make it consistent with the ranking of the confidence value list, and extract the top M high-confidence local regions of interest. The specific process is as follows: (1) Upsample the M regions {R1, R2, …, R M}} using bilinear interpolation method to convert them into the size of H'×C'. Through data preprocessing and depthwise separable convolution feature learning, calculate the probability of each local region of interest as the class label as the confidence value, and the confidences {C1, C2, …, C M}}. At the same time, by minimizing the cross-entropy loss between each class label and the confidence value, and calculating the cross-loss function of the global image X, optimize this step: Where C(·) is the confidence value calculation function, the first part of the formula is the sum of the cross losses of all regions, and the second part is the cross-entropy loss of the entire image; (2) According to the confidence levels {C1, C2... C M} of each region of interest in (1), and the information content of M regions of interest in the candidate local regions of interest, construct a sorting consistency loss function; the specific rules of this loss function are as follows. Let the information content sorting be {I1, I2... I M}, when I s > I i and C s > C i , the corresponding label is 0. When the input is the opposite, that is, I s < I i and C s > C i , the corresponding label is 1; the sorting consistency loss function is defined as L I (·) as follows: Where f(·) uses the hinge loss function: f(x) = max{1 - x, 0}; (3) Through the (2) ranking consistency loss function, train the network to guide the re-ranking of the list of candidate local regions of interest features to make it consistent with the ranking of the confidence value list, and extract the top M high-confidence local regions of interest.

6. The image classification method for learning high-confidence local features and global features according to claim 5, wherein In the fifth step, the M highly confident local regions of interest extracted are adjusted into tensors with the same size as the context features through pooling operations, and then concatenated with the context features for feature aggregation; The specific process is as follows: (1) Resize the outputs {R1, R2... R M}} of the high-confidence local region discriminant and feature learning module to the size of H'×W'×8C, where H' is the height of the image, W' is the width of the image, and 8C is the feature dimension; (2) The output of the first step is used as the input for global feature extraction. After global average pooling and layer normalization operations, it is output as the context features; (3) The highly confident local features in step (1) are aggregated with the context features in step (2) and used as the input for the classification module.

7. The method for image classification by learning high-confidence local features and global features according to claim 6, characterized in that In the sixth step, the features obtained by aggregating the highly confident local features of interest and the global features are connected to a fully connected layer and a Softmax classifier for final classification. The specific process is as follows: (1) The output of the fifth step is used as the input and connected to a fully connected layer; (2) After the output of step (1), a Softmax classifier is connected to predict the final classification result.

8. An image classification system for learning high-confidence local features and global features, characterized in that It includes: A depthwise separable convolution feature learning module, which adopts a stacked form of depthwise separable convolution layers, nests downsampling layers, and performs feature extraction through pre-training; A multi-scale feature learning module, which uses convolutions of three different scales to extract multi-scale features and form a feature tensor; A candidate local region of interest generation module, which uses sliding windows with different strides and non-maximum suppression methods to remove redundant regions and generate candidate local regions of interest; A highly confident local region of interest discrimination and feature learning module, which calculates the probability that a candidate local region of interest belongs to a class label as the confidence value, constructs a ranking consistency loss function to train the network based on the confidence value, and ranks the candidate local regions of interest to discriminate highly confident regions of interest; A feature aggregation module, which aggregates the features of the extracted highly confident local regions of interest and the global features, to obtain a feature concatenation tensor; A classification module, which performs final image classification based on the aggregated features.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method described in any one of claims 1-7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1-7.

Citation Information

Patent Citations

  • Image classification method based on Resnet50 in combination with attention mechanism and feature pyramid

    CN114494805A

  • A multi-scale target detection method fusing context information

    CN109816012A

  • Oracle carved text detection method combining local prior features and deep convolution features

    CN111310760A