Remote Sensing Image Thumbnail Abnormal Data Recognition Method Based on Convolutional Neural Network
By constructing a remote sensing image thumbnail abnormal data recognition data set and using convolutional neural network technology, the problem of remote sensing image data relies on manual identification is solved, and the rapid and accurate abnormal data recognition of remote sensing image thumbnails is realized, reducing labor costs and resource waste.
Patent Information
- Application Number
- CN202310375451.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-10
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-04-10
AI Technical Summary
In the prior art, the abnormal identification of remote sensing image data mainly relies on manual methods, which is slow and costly, and lacks an abnormal data identification method for remote sensing image thumbnails, resulting in waste of satellite resources and data processing burden.
A remote sensing image thumbnail exception data recognition data set is constructed, and a convolutional neural network technology is used, combining channel and spatial attention mechanisms, an identification algorithm is designed, and binary cross entropy, Focal Loss and L2 Loss loss functions are trained to achieve automated recognition.
It realizes fast and accurate abnormal data recognition of remote sensing image thumbnails, reduces labor costs, and improves recognition efficiency and recognition accuracy.
Smart Images

Figure CN116385880B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital image processing, and in particular, to a method for identifying abnormal data in remote sensing image thumbnails based on a convolutional neural network. Background Art
[0002] Remote sensing data of an on-orbit satellite needs to go through different data transmission and processing stages from collection to ground reception. Inevitably, internal and external environmental impacts of instruments exist in these stages, which makes the remote sensing image data obtained on the ground have data anomalies. Due to the huge amount of remote sensing image data, abnormal remote sensing image data generally has no use value. These abnormal data not only increase the disk load on the satellite, but also increase unnecessary data transmission between the satellite and the ground, as well as increase unnecessary data processing work in the industry such as image quality evaluation. Currently, the identification of the vast majority of remote sensing image abnormal data is completed manually, which has the disadvantages of slow speed and high labor cost. Therefore, it is of great significance for satellite production management departments to effectively and quickly identify abnormal data from a large amount of remote sensing image data accurately.
[0003] In recent years, with the rapid development of machine learning and deep learning related technologies, the performance of related algorithms has become higher and higher. However, many methods based on deep learning technologies are developed for the scene classification of satellite images, and lack research on methods for identifying abnormal data in remote sensing images. On the other hand, some researchers in the remote sensing field have developed a set of remote sensing data quality detection systems by comparing the data received on the ground with the data sent by remote sensing satellites. However, it mainly targets the abnormal data generated during the process from the satellite sending to the ground receiving, and ignores the anomalies of the sensors that have been running on the remote sensing satellite for a long time.
[0004] Compared with the original image, the remote sensing image thumbnail has the advantages of small image scale and fast transmission speed. At the same time, when the remote sensing image data is abnormal, the data anomaly will also be clearly shown in the remote sensing image thumbnail. Therefore, deep learning technology can be used to identify abnormal data in remote sensing image thumbnails, so as to quickly and accurately identify abnormal data in a large amount of remote sensing image data. Based on this, the present invention proposes a method for identifying abnormal data in remote sensing image thumbnails based on a convolutional neural network. Summary of the Invention
[0005] The purpose of the present invention is to establish a dataset for identifying abnormal data in remote sensing image thumbnails, realize the automatic identification of abnormal data from a large amount of remote sensing image data, and on this basis, propose a method for identifying abnormal data in remote sensing image thumbnails based on a convolutional neural network.
[0006] In order to achieve the above purpose, the present invention adopts the following technical solutions:
[0007] A method for identifying abnormal data in remote sensing image thumbnails based on convolutional neural network, specifically including the following steps:
[0008] S1. Establishment of the dataset for identifying abnormal data in remote sensing image thumbnails:
[0009] S1.1. Data collection: Use satellites to establish the dataset, and fuse the collected multi-spectral data of remote sensing images to obtain remote sensing image thumbnails with three channels (R, G, B);
[0010] S1.2. Data annotation: Manually annotate all the remote sensing image thumbnails obtained in S1.1 according to the abnormal situations they present;
[0011] S1.3. Data cropping: Crop all the remote sensing image thumbnail data obtained in S1.1 to a size of 224×224 pixels for the training of subsequent recognition algorithms;
[0012] S1.4. Data division: Randomly divide all the remote sensing image thumbnail data obtained in S1.1 according to the ratio of training set: validation set: test set = 8:1:1;
[0013] S1.5. Data augmentation: For the images in the training set, perform data augmentation by flipping, translating, rotating, and adding noise respectively. Without changing the labels, expand the scale of the training set to improve the generalization ability of subsequent recognition algorithms;
[0014] S2. Construction of the algorithm for identifying abnormal data in remote sensing image thumbnails: Based on convolutional neural network technology, combined with the dataset for identifying abnormal data in remote sensing image thumbnails established in S1, construct an algorithm for identifying abnormal data in remote sensing image thumbnails. The specific content of the algorithm is as follows:
[0015] ① Backbone network: Select ResNet-50 as the backbone network to extract features from the input remote sensing image thumbnails;
[0016] ② Channel attention module: Add channel attention modules after the first convolutional layer and the last convolutional layer in the backbone network respectively;
[0017] ③ Spatial attention module: Add a spatial attention module after each channel attention module;
[0018] ④ Classification layer: Add two fully connected layers and a sigmoid layer in sequence at the end of the backbone network to map the feature map output by the backbone network to a vector y with a length of 1×4, where y = (y1, y2, y3, y4) respectively represent the scores of each type of abnormality;
[0019] ⑤ Loss function and hyperparameters: Use three loss functions, binary cross-entropy Loss, Focal Loss, and L2Loss, to jointly supervise the output of the convolutional neural network;
[0020] S3. Algorithm training: Based on the deep learning framework Pytorch, train the model using the Adam optimizer. Iteratively train for N epochs on the entire training set, and set the learning rate to f; when training reaches N / 2 epochs, adjust the learning rate to f / 10; after training is completed, obtain the final stable remote sensing image thumbnail anomaly data recognition model;
[0021] S4. Model testing and result output: Input the remote sensing image thumbnail data in the test set into the stable remote sensing image thumbnail anomaly data recognition model obtained in S3, and output the corresponding anomaly data recognition results.
[0022] Preferably, the abnormal situations described in S1.2 are specifically divided into four types: data missing, data garbled, data color cast, and CCD stitching; combined with the abnormal situations, specifically manually labeled as where k represents the kth image in the dataset, y i k ∈{0,1}, i = 1,2,3,4; y i k = 1 indicates that the image has the ith type of anomaly, y i k = 0 indicates that there is no such anomaly; indicates that the image has no anomalies and is a normal remote sensing image.
[0023] Preferably, each residual block in the ResNet-50 consists of three convolutional layers, three BatchNormalization layers, two ReLU operations, and residual connections; the feature representation output by the residual block is:
[0024]
[0025]
[0026]
[0027]
[0028] where F in represents the input feature of the residual block; Conv and BN respectively represent the convolutional layer and the BatchNormalization layer; F residual is the output feature of the residual block.
[0029] Preferably, the channel attention module specifically includes the following content:
[0030] First, perform a global average pooling operation on the input features to obtain a 1×1×C vector, where C represents the number of channels of the input features; then, successively pass through a multi-layer perceptron and a sigmoid operation to obtain the weights for each channel; finally, multiply the weights by the input features.
[0031] The feature representation output by the channel attention module is:
[0032] F channel attention = σ(MLP(GAP(F in )))·F in
[0033] where GAP represents global average pooling; MLP represents a multi-layer perceptron, consisting of a fully connected layer and a ReLU operation; σ(·) represents the sigmoid function.
[0034] Preferably, the spatial attention module specifically includes the following:
[0035] First, calculate the average value of the input features in the channel dimension to obtain an H×W×1 vector, where H and W represent the scale of the input features; then, successively pass through a convolutional layer and a sigmoid operation to obtain the weights in the spatial range; finally, multiply the weights by the input features.
[0036] The feature output by the spatial attention module can be expressed as:
[0037] F spatial attention = σ(Conv(Mean c (F in )))·F in
[0038] where Mean c represents calculating the average value in the channel dimension.
[0039] Preferably, the three loss functions of binary cross-entropy Loss, Focal Loss, and L2 Loss are specifically expressed as:
[0040]
[0041]
[0042]
[0043] where L BCE represents binary cross-entropy Loss, L FocalFocal Loss is denoted as, L2 Loss is denoted as L2; y represents the label of the current image, represents the output of the convolutional neural network, and γ represents the hyperparameter in FocalLoss;
[0044] Based on the above function representation, the final loss function can be expressed as:
[0045] L = a1L BCE + a2L Focal + a3L2
[0046] where a1, a2, and a3 represent hyperparameters.
[0047] Compared with the prior art, the present invention provides a method for identifying abnormal data in remote sensing image thumbnails based on a convolutional neural network, having the following beneficial effects:
[0048] (1) The present invention constructs a dataset for identifying abnormal data in remote sensing image thumbnails.
[0049] (2) Based on the constructed dataset for identifying abnormal data in remote sensing image thumbnails, the present invention further proposes an algorithm and model for identifying abnormal data in remote sensing image thumbnails, and uses a convolutional neural network and channel and spatial attention mechanisms to give the identified abnormal results for the input remote sensing image thumbnails.
[0050] (3) Experiments show that the method proposed by the present invention can accurately and efficiently identify abnormal data in remote sensing image thumbnails. Description of the Drawings
[0051] Figure 1 is the flowchart of the method for identifying abnormal data in remote sensing image thumbnails based on a convolutional neural network proposed by the present invention. Detailed Embodiments
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0053] Embodiment 1:
[0054] Please refer to Figure 1 , the present invention proposes a method for identifying abnormal data in remote sensing image thumbnails based on a convolutional neural network, including the following content:
[0055] ① Establishment of the dataset for identifying abnormal data in remote sensing image thumbnails:
[0056] The process of establishing the dataset mainly includes the following 5 steps:
[0057] Step 1, data collection:
[0058] The present invention uses the GF-1 BCD satellite and the GF-2 satellite to establish a data set, and fuses the collected multi-spectral data of remote sensing images to obtain a thumbnail of a remote sensing image with three channels (R, G, B).
[0059] Step 2, data annotation:
[0060] The abnormal phenomena in the data after collection and fusion can be directly distinguished by the naked eye. According to the different abnormal situations in the thumbnail of the remote sensing image, it can be divided into four types of abnormalities: data missing, data garbled, data color cast, and CCD splicing. The same image may be affected by multiple abnormalities. Therefore, for each thumbnail of the remote sensing image, according to the abnormal situation it shows, its label can be manually annotated as where k represents the k-th image in the data set, and y i k ∈{0, 1}, i = 1, 2, 3, 4; y i k = 1 indicates that the image has the i-th type of abnormality, and y i k = 0 indicates that there is no such abnormality; indicates that the image has no any abnormality and is a normal remote sensing image.
[0061] Step 3, data cropping:
[0062] All the thumbnail data of the remote sensing images are cropped to a size of 224×224 pixels for the training of subsequent recognition algorithms.
[0063] Step 4, data division:
[0064] All the thumbnail data of the remote sensing images are randomly divided according to the ratio of training set: validation set: test set = 8:1:1.
[0065] Step 5, data augmentation:
[0066] For the images in the training set, data augmentation is performed. The data augmentation methods of flipping, translation, rotation, and adding noise are respectively adopted to expand the scale of the training set without changing the labels, so as to improve the generalization ability of the subsequent recognition algorithm.
[0067] ② Flow of the abnormal data recognition algorithm for remote sensing image thumbnails:
[0068] The present invention uses the input image x k and its corresponding label y k to design an algorithm for identifying abnormal data, which mainly includes the following contents:
[0069] 1) Backbone network:
[0070] ResNet-50 is selected as the backbone network to extract features from the input remote sensing image thumbnail. Each residual block consists of three convolutional layers, three BatchNormalization layers, two ReLU operations, and a residual connection. The features output by the residual block can be expressed as:
[0071]
[0072]
[0073]
[0074]
[0075] Among them, F in represents the input features of the residual block; Conv and BN represent the convolutional layer and the BatchNormalization layer respectively; F residual is the output feature of the residual block.
[0076] 2) Channel attention module:
[0077] Channel attention modules are added after the first convolutional layer and the last convolutional layer in the backbone network respectively. Specifically, first perform a global average pooling operation on the input features to obtain a 1×1×C vector, where C represents the number of channels of the input features; then pass through a multi-layer perceptron and a sigmoid operation in sequence to obtain the weight of each channel; finally, multiply the weight by the input features. The features output by the channel attention can be expressed as:
[0078] F channel attention = σ(MLP(GAP(F in )))·F in
[0079] Among them, GAP represents global average pooling; MLP represents a multi-layer perceptron, which consists of a fully connected layer and a ReLU operation; σ(·) represents the sigmoid function,
[0080] 3) Spatial attention module:
[0081] Add a spatial attention module after each channel attention module. Specifically, first calculate the average value of the input features in the channel dimension to obtain a vector of size H×W×1, where H and W represent the scale of the input features; then, successively pass through a convolutional layer and a sigmoid operation to obtain the weights in the spatial range; finally, multiply the weights by the input features. The features output by the channel attention can be expressed as:
[0082] F spatial attention = σ(Conv(Mean c (F in )))·F in
[0083] where Mean c represents calculating the average value in the channel dimension.
[0084] 4) Classification layer:
[0085] Since a remote sensing image may be affected by multiple anomalies simultaneously, two fully connected layers and a sigmoid layer are successively added at the end of the backbone network to map the feature map output by the backbone network to a vector y of length 1×4, where y = (y1, y2, y3, y4), representing the scores of each type of anomaly respectively.
[0086] 5) Loss function and hyperparameters:
[0087] This invention uses three loss functions to jointly supervise the output of the convolutional neural network: Binary Cross Entropy (BCE) Loss, Focal Loss, and L2 Loss, which are respectively expressed as:
[0088]
[0089]
[0090]
[0091] where L BCE represents Binary Cross Entropy Loss, L Focal represents Focal Loss, L2 represents L2 Loss; y represents the label of the current image, represents the output of the convolutional neural network, γ represents the hyperparameter in Focal Loss, and in this invention, γ = 2 is set. The final loss function is expressed as follows:
[0092] L = a1L BCE + a2L Focal + a3L2
[0093] Among them, a1, a2, and a3 represent hyperparameters. In the present invention, a1 = 1.5, a2 = 1.5, and a3 = 1 are set.
[0094] ③ Algorithm training.
[0095] In the present invention, the Adam optimizer is used, and its parameter settings are as follows: the exponential decay rate β1 of the first moment estimate is 0.9, the exponential decay rate β2 of the second moment estimate is 0.999, and the numerical stability parameter ∈ = 10 -8 . The initial learning rate is set to 0.0001. The present invention uses the deep learning framework Pytorch to train the model. A total of 300 epochs are iteratively trained on the entire training set. When training reaches 150 epochs, the learning rate is adjusted to 1 / 10 of the previous one.
[0096] ④ Model testing and result output.
[0097] Input the thumbnail data of remote sensing images in the test set into the convolutional neural network for inference to obtain the corresponding abnormal data recognition results.
[0098] Statistically analyze the results of the method proposed in the present invention on the test set and compare them with other common classification algorithms. The specific result indicators are compared as shown in Table 1.
[0099] Table 1 Comparison of result indicators
[0100]
[0101] It can be seen from Table 1 that the method proposed in the present invention can accurately and efficiently identify abnormal data in the thumbnail of remote sensing images.
[0102] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A method for identifying abnormal data in remote sensing image thumbnails based on a convolutional neural network, characterized in that, Specifically, it includes the following steps: S1. Establishment of the remote sensing image thumbnail abnormal data recognition dataset: S1.
1. Data collection: Use satellites to establish the dataset, and fuse the collected multi-spectral data of remote sensing images to obtain remote sensing image thumbnails with three channels; S1.
2. Data annotation: Manually annotate all the thumbnail remote sensing images obtained in S1.1 according to the anomalies they present; the anomalies are specifically divided into four types: data loss, data garbling, data color cast, and CCD stitching; combined with the anomalies, the specific manual annotation is , where k represents the k th image in the dataset, , i = 1, 2, 3, 4; indicates that the image has the i th anomaly, indicates that this type of anomaly does not exist; indicates that the image has no anomalies and is a normal remote sensing image; S1.
3. Data cropping: Crop all the remote sensing image thumbnail data obtained in S1.1 to a size of 224×224 pixels; S1.
4. Data division: Randomly divide all the remote sensing image thumbnail data obtained in S1.1 according to the ratio of training set: validation set: test set = 8:1:1; S1.
5. Data augmentation: For the images in the training set, perform data augmentation by flipping, translation, rotation, and adding noise respectively to expand the scale of the training set without changing the labels; S2. Construction of the remote sensing image thumbnail abnormal data recognition algorithm: Based on convolutional neural network technology, combine the remote sensing image thumbnail abnormal data recognition dataset established in S1 to construct a remote sensing image thumbnail abnormal data recognition algorithm. The algorithm specifically includes the following content: ① Backbone network: Select ResNet-50 as the backbone network to extract features from the input remote sensing image thumbnails; Each residual block in the ResNet-50 consists of three convolutional layers, three Batch Normalization layers, two ReLU operations, and a residual connection; The feature output by the residual block is expressed as: Among them, F in represents the input feature of the residual block; Conv and BN respectively represent the convolutional layer and the Batch Normalization layer; F residual is the output feature of the residual block; ② Channel attention module: Add a channel attention module after the first convolutional layer and the last convolutional layer in the backbone network respectively; ③ Spatial attention module: Add a spatial attention module after each channel attention module; ④ Classification layer: Two fully-connected layers and a sigmoid layer are sequentially added at the end of the backbone network to map the feature map output by the backbone network to a vector of length 1×4, where y on it, where y =( y 1, y 2, y 3, y 4) represent the scores for each type of anomaly respectively; ⑤ Loss function and hyperparameters: Use three loss functions, binary cross-entropy Loss, Focal Loss, and L2 Loss, to jointly supervise the output of the convolutional neural network; S3. Algorithm training: Train the model based on the deep learning framework Pytorch, use the Adam optimizer, and iterate and train for N cycles on the entire training set. Set the learning rate to f; When training reaches N / 2 cycles, adjust the learning rate to f / 10; After the training ends, obtain the final stable remote sensing image thumbnail anomaly data recognition model; S4. Model testing and result output: Input the remote sensing image thumbnail data in the test set into the stable remote sensing image thumbnail abnormal data recognition model obtained in S3, and output the corresponding abnormal data recognition results.
2. The method for identifying abnormal data of remote sensing image thumbnails based on convolutional neural network according to claim 1, wherein The channel attention module specifically includes the following content: First, perform a global average pooling operation on the input features to obtain a 1×1× C vector, where C represents the number of channels of the input features; Then, it passes through a multi-layer perceptron and a sigmoid operation in sequence to obtain the weight of each channel; Finally, multiply the weight by the input feature; The feature output by the channel attention module is expressed as: F channel attention = σ ( MLP ( GAP ( F in )))· F in Among them, GAP represents global average pooling; MLP represents a multi-layer perceptron, which consists of a fully connected layer and ReLU operations; σ (·) represents the sigmoid function, .
3. The method for identifying abnormal data in remote sensing image thumbnails based on a convolutional neural network according to claim 1, characterized in that, The spatial attention module specifically includes the following content: First, calculate the average value of the input features in the channel dimension to obtain a H × W ×1 vector, where H and W represent the scale of the input features; Then, it passes through a convolutional layer and a sigmoid operation in sequence to obtain the weight in the spatial range; Finally, multiply the weight by the input feature; The feature output by the spatial attention module is expressed as: F spatial attention = σ ( Conv ( Mean c ( F in )))· F in Among them, Mean c represents taking the average in the channel dimension.
4. The method for identifying abnormal data of remote sensing image thumbnails based on a convolutional neural network according to claim 1, characterized in that The three loss functions, binary cross-entropy Loss, Focal Loss, and L2 Loss, are specifically expressed as: Among them, L BCE represents binary cross-entropy Loss, L Focal represents Focal Loss, L 2 represents L2 Loss; y represents the label of the current image, represents the output of the convolutional neural network, γ represents the hyperparameter in Focal Loss; Based on the above function representation, the final loss function can be expressed as: L = a 1 L BCE + a 2 L Focal + a 3 L 2 Among them, a 1、 a 2、 a 3 represents a hyperparameter.
Citation Information
Patent Citations
PCB surface defect classification method based on improved ResNet34 network
CN114820569A
Computerized systems and methods for generating models for identifying thumbnail images to promote videos
US20150131967A1