High-resolution satellite remote sensing image general change detection method based on deep learning
By combining Segformer and FastSAM feature extractors and using technical means such as multi-level feature aggregation modules, the problem that a single feature extractor is difficult to take into account both global and local features is solved, and the accurate change detection of high-resolution satellite remote sensing images is achieved.
Patent Information
- Application Number
- CN202510002955.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
AI Technical Summary
The existing single feature extractor is difficult to take into account both global and local features, resulting in limited accuracy and efficiency in the change detection of high-resolution satellite remote sensing images.
The synergy between Segformer and FastSAM dual feature extractors is adopted to perform pixel-by-pixel fusion of features through multi-level feature aggregation module, and combine differential feature modules and decoder modules to achieve efficient processing of image object detection.
It realizes the effective fusion of global semantic information and detailed features, improves the accuracy and robustness of change detection, and can be applied to satellite remote sensing images in different scenarios.
Smart Images

Figure CN119942147A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of satellite remote sensing technology, and in particular to a general change detection method for high-resolution satellite remote sensing images based on deep learning. Background Art
[0002] General change detection based on satellite remote sensing images can monitor changes in land objects and is widely used in urban planning, non-agriculturalization of cultivated land, protection of forest and grassland resources, and other fields. Traditional change detection methods mostly rely on manual visual interpretation or algorithms based on simple image difference, which have high manual participation and low efficiency. It is difficult to meet the needs of large-scale data sets and high-frequency updates for accurate identification of land object change information. Especially when facing high-resolution satellite remote sensing images, the richness of image details brings more types of changes and complex land object features, which limits the applicability and accuracy of traditional methods. In addition, in general change detection tasks, there are complex factors such as multiple land object change categories, large differences in change areas, and imbalanced change categories. The usual semantic segmentation network is prone to low accuracy in the extraction results of change areas.
[0003] Segformer is an image segmentation model based on the Transformer architecture, which can effectively capture long-range contextual information and help the model better understand the spatial relationship in the scene, thereby improving sensitivity to changes in objects. FastSAM is a fast and efficient large-scale image segmentation visual model that can better capture complex image patterns and edge details and improve image segmentation accuracy. However, although Segformer is good at capturing global semantic information, it may be insufficient for detailed features (such as small targets or changing boundaries), especially in tasks such as change detection that require high-resolution features. FastSAM emphasizes local edges and target segmentation, but lacks modeling capabilities for global semantic context, and may misjudge changes between regions in complex scenes (such as remote sensing images with complex objects). Moreover, existing single feature extractors find it difficult to efficiently process global semantic information and detailed features at the same time.
[0004] In summary, it is difficult for existing single feature extractors to take into account both global and local features at the same time. Summary of the invention
[0005] The present invention solves the problem that the existing single feature extractor is difficult to take into account both global and local features at the same time.
[0006] The present invention provides a method for detecting an image target based on a deep learning network, specifically:
[0007] The images are divided into two paths. One path is extracted by FastSAM feature extractor and the feature map of the first scale is output. The other path is extracted by Segformer feature extractor and the feature map of the second scale is output.
[0008] The feature map of the first scale and the feature map of the second scale are input into the multi-level feature aggregation module for pixel-by-pixel fusion, and the fused feature map is output;
[0009] The fused feature map is passed through the difference feature module for absolute difference calculation, and the difference feature map is output;
[0010] The difference feature map is decoded by the decoder module and the image target detection result is output.
[0011] Further, in one embodiment of the present invention, the feature map of the first scale and the feature map of the second scale are input into a multi-level feature aggregation module for pixel-by-pixel fusion, and the fused feature map is output, specifically:
[0012] The feature map of the first scale and the feature map of the second scale are respectively input into the channel conversion unit for channel dimension conversion, and the feature map of the first scale and the feature map of the second scale with the same channel dimension are output;
[0013] The feature map of the first scale and the feature map of the second scale with the same channel dimension are input into the channel connection unit for connection to obtain a multi-level feature map;
[0014] Multi-level feature maps are aggregated through the hierarchical aggregation unit to output the fused feature map.
[0015] Furthermore, in one embodiment of the present invention, the multi-level feature graphs are aggregated by a hierarchical aggregation unit, specifically:
[0016] The multi-level feature maps are aggregated through the upsampling layer, channel connection layer and 3×3 convolutional block in sequence.
[0017] Further, in one embodiment of the present invention, the difference feature map is decoded by a decoder module, specifically:
[0018] The difference feature map is decoded by multiple residual blocks, 3×3 convolution blocks and classifier modules in sequence.
[0019] Furthermore, in one embodiment of the present invention, an attention module is connected in series after the second BN layer of each of the plurality of residual blocks.
[0020] Furthermore, in one embodiment of the present invention, the attention module is specifically:
[0021] The feature map is divided into two paths. One feature map passes through the global average pooling layer for spatial feature compression and outputs the first feature map. The other feature map passes through the global maximum pooling layer for spatial feature compression and outputs the second feature map.
[0022] After both the first feature map and the second feature map are convolved, the convolved first feature map and the convolved second feature map are output respectively;
[0023] The first feature map after convolution and the second feature map after convolution are added in sequence and operated with Sigmoid function, and then the third feature map is output;
[0024] The feature map is multiplied point by point with the third feature map, and a feature map with channel attention is output.
[0025] Furthermore, in one embodiment of the present invention, the classifier module is specifically:
[0026] The feature map passes through the upsampling layer, 3×3 convolution block and upsampling layer in sequence, and outputs a binary change map.
[0027] The present invention provides a general change detection method for high-resolution satellite remote sensing images based on deep learning, wherein the method is implemented by using any of the above-mentioned methods for detecting image targets based on a deep learning network, and comprises the following steps:
[0028] Step S1, selecting a plurality of satellite remote sensing image pairs taken at different times in the same geographical area, wherein the plurality of satellite remote sensing image pairs all include a pre-phase image and a post-phase image;
[0029] Step S2, preprocessing the multiple groups of satellite remote sensing images respectively;
[0030] Step S3, respectively outlining the change spots in the multiple sets of satellite remote sensing images and generating corresponding label maps;
[0031] Step S4, segmenting and slicing the plurality of satellite remote sensing image pairs and their corresponding label images to produce a slice sample set;
[0032] Step S5, after amplifying the slice sample set, input it into the deep learning network for training;
[0033] In step S6, the multiple groups of satellite remote sensing images to be detected are input into the trained deep learning network for change detection, and multi-scale prediction is adopted to obtain the intersection of the multi-scale prediction results to obtain a general change detection result.
[0034] Further, in one embodiment of the present invention, in the step S3, the step of outlining the change spots in the pre-phase image and the post-phase image respectively is as follows:
[0035] The change spots in the previous phase image and the next phase image are outlined respectively, and the outlined lines need to fit the boundaries of the change spots.
[0036] Further, in one embodiment of the present invention, in the step S3, the label map includes changed spots and unchanged spots;
[0037] The pixel value of the changed spot is marked as 1, and the pixel value of the unchanged spot is marked as 0;
[0038] The area of the changed pattern is greater than 20 square meters.
[0039] The present invention solves the problem that the existing single feature extractor is difficult to take into account both global and local features at the same time. Specific beneficial effects include:
[0040] 1. The method for detecting image targets based on a deep learning network described in the present invention mainly uses a single feature extractor in the prior art, which has the problem of being unable to take into account both global and local features at the same time. In order to solve the above technical problems, the present invention proposes a deep learning network, which includes a feature extraction module, a multi-level feature aggregation module, a difference feature module and a decoder module. The feature extraction module fully utilizes the advantages of the SegFormer and FastSAM dual feature extractors in global semantic extraction and local detail capture through the synergistic effect of the two, and can achieve efficient feature fusion and accurate prediction of change areas through channel connection and optimized decoder design, solving the problem that a single model is difficult to take into account both global and local features at the same time.
[0041] 2. The image target detection method based on a deep learning network described in the present invention can achieve the complementarity of global and local features through the combination of the Segformer feature extractor and the FastSAM feature extractor, thereby effectively solving the problem that the existing single feature extractor is difficult to take into account both global and local features at the same time. However, the channel dimensions of the feature maps output by the Segformer feature extractor and the FastSAM feature extractor are different, and the feature maps cannot be better fused. In order to solve the above technical problems, the present invention proposes a multi-level feature aggregation module, which is composed of a channel conversion unit, a channel connection unit and a hierarchical aggregation unit. The multi-level feature aggregation module fuses the features extracted by the Segformer feature extractor and the FastSAM feature extractor pixel by pixel, retaining both the global semantic information and the detail features;
[0042] 3. The present invention discloses a general change detection method for high-resolution satellite remote sensing images based on deep learning. In the field of satellite remote sensing image change detection, for high-resolution satellite remote sensing images with diverse types of objects and complex changes, the use of a single model will result in the problem of difficulty in taking into account both global and local features at the same time. To solve the above technical problems, the present invention provides a more comprehensive solution for the feature expression of remote sensing images through the synergy of the SegFormer and FastSAM dual feature extractors.
[0043] The present invention provides a general change detection method for high-resolution satellite remote sensing images based on deep learning, which can be applied to satellite remote sensing images of different scenes and meet the needs of high-resolution change detection in different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0045] Figure 1 It is a general change detection flow chart of high-resolution satellite remote sensing images based on deep learning as described in Implementation Mode 7;
[0046] Figure 2 It is the front phase image diagram described in the seventh embodiment;
[0047] Figure 3 The post-phase image diagram described in Embodiment 7;
[0048] Figure 4 It is the label diagram described in the seventh embodiment;
[0049] Figure 5 is the attention module diagram described in Implementation Mode 5;
[0050] Figure 6 The pre-phase image, post-phase image and label image described in Embodiment 7;
[0051] Figure 7 This is a comparison chart of the test results described in Implementation Method Seven. DETAILED DESCRIPTION
[0052] The following will clearly and completely describe various embodiments of the present invention in conjunction with the accompanying drawings. The embodiments described with reference to the accompanying drawings are exemplary and intended to be used to explain the present invention, but should not be understood as limiting the present invention.
[0053] Implementation method 1: A method for detecting an image target based on a deep learning network described in this implementation method is specifically as follows:
[0054] The images are divided into two paths. One path is extracted by FastSAM feature extractor and the feature map of the first scale is output. The other path is extracted by Segformer feature extractor and the feature map of the second scale is output.
[0055] The feature map of the first scale and the feature map of the second scale are input into the multi-level feature aggregation module for pixel-by-pixel fusion, and the fused feature map is output;
[0056] The fused feature map is passed through the difference feature module for absolute difference calculation, and the difference feature map is output;
[0057] The difference feature map is decoded by the decoder module and the image target detection result is output.
[0058] In the prior art, a single feature extractor has the problem of being difficult to take into account both global and local features at the same time.
[0059] In order to solve the above technical problems, this embodiment proposes a deep learning network, based on which image targets are detected, specifically:
[0060] A feature extraction module is used to extract features from images, and a multi-level feature aggregation module is used to enhance the model's sensitivity to small and large area changes. At the same time, the accuracy and generalization ability of change detection are improved through difference feature extraction and channel attention strategy.
[0061] The feature extraction module includes a Segformer feature extractor and a FastSAM feature extractor, specifically:
[0062] Taking high-resolution satellite remote sensing images as an example, we obtain the pre-phase image A and the post-phase image B. The image size input to the feature extraction module is 512×512×3 (i.e. H×W×C 0 , where H is the image height, W is the image width, and C 0 is the number of channels).
[0063] The FastSAM feature extractor is used to input the front phase image A and the back phase image B, and output the feature maps of A at four different levels (or scales): 128×128×160, 64×64×320, 32×32×640, 16×16×640; and 4 feature maps of different scales of B: 128×128×160, 64×64×320, 32×32×640, 16×16×640.
[0064] The Segformer feature extractor is used to input the front phase image A and the back phase image B. After passing through different Transformer blocks, it outputs four feature maps of different scales of A: 128×128×C 1 , 64×64×C 2 , 32×32×C 3 , 16×16×C 4 ; And 4 feature maps of different scales of B: 128×128×C 1 , 64×64×C 2 , 32×32×C 3 , 16×16×C 4 . Number of channels C 1 to C 4 You can set it as needed. Here, C 1 to C 4 Set to 64, 128, 320 and 512 respectively.
[0065] Through the above two multi-level feature extractors, feature information at different angles and levels can be mined. The combination of Segformer feature extractor and FastSAM feature extractor can achieve the complementarity of global and local features, thus effectively solving the problem that the existing single feature extractor is difficult to take into account both global and local features at the same time.
[0066] Implementation method 2: This implementation method is a further limitation of the image target detection method based on a deep learning network described in implementation method 1. The feature map of the first scale and the feature map of the second scale are jointly input into a multi-level feature aggregation module for pixel-by-pixel fusion, and the fused feature map is output, specifically:
[0067] The feature map of the first scale and the feature map of the second scale are respectively input into the channel conversion unit for channel dimension conversion, and the feature map of the first scale and the feature map of the second scale with the same channel dimension are output;
[0068] The feature map of the first scale and the feature map of the second scale with the same channel dimension are input into the channel connection unit for connection to obtain a multi-level feature map;
[0069] Multi-level feature maps are aggregated through the hierarchical aggregation unit to output the fused feature map.
[0070] In this implementation, the multi-level feature graphs are aggregated by a hierarchical aggregation unit, specifically:
[0071] The multi-level feature maps are aggregated through the upsampling layer, channel connection layer and 3×3 convolutional block in sequence.
[0072] In this embodiment, the feature extraction module described in the first embodiment can achieve the complementarity of global and local features through the combination of the Segformer feature extractor and the FastSAM feature extractor, thereby effectively solving the problem that the existing single feature extractor is difficult to take into account both global and local features at the same time. However, the channel dimensions of the feature maps output by the Segformer feature extractor and the FastSAM feature extractor are different, and the feature maps cannot be better fused.
[0073] In order to solve the above technical problems, this embodiment proposes a multi-level feature aggregation module, which is composed of a channel conversion unit, a channel connection unit and a hierarchical aggregation unit, specifically:
[0074] The channel conversion unit transforms the FastSAM feature map and Segformer feature map of the same level to the same channel dimension by using a one-dimensional convolution layer, and can reduce the number of model parameters. Specifically, by setting the number of input and output channels of the one-dimensional convolution layer, the feature maps of different levels of the front-phase image A are transformed as follows (the same applies to the back-phase image B):
[0075] 128×128:
[0076] 64×64:
[0077] 32×32:
[0078] 16×16:
[0079] The channel connection unit is used to connect the FastSAM feature map and the Segformer feature map of the same level in the channel direction. After the channel connection operation (Concat), the feature maps of each level of the front phase image A and the back phase image B are obtained. and Then the dimensions of the feature maps of each level of the previous phase image A are: The dimension of the feature map at the same level of the later-phase image B is consistent with that of A.
[0080] The hierarchical aggregation unit consists of an upsampling layer, a channel connection layer, and a 3×3 convolutional block, which is used to aggregate feature maps of different levels. The upsampling layer is used to double the height and width of the feature map while keeping the number of channels unchanged.
[0081] The specific implementation process of the hierarchical aggregation unit: for the feature map First, after the upsampling layer, it is transformed into Then through the channel connection layer and Perform channel connection to obtain a dimension of 32×32×640 Then, after a 3×3 convolution block (the number of input and output channels is set to 640 and 160 respectively), we get Will After further upsampling, it is transformed into Then through the channel connection layer and Perform channel connection to obtain a dimension of 64×64×320 Then, after a 3×3 convolution block (with the number of input and output channels set to 320 and 80 respectively), we get Will After further upsampling, it is transformed into Then through the channel connection layer and Perform channel connection to obtain a dimension of 128×128×160 Then, after a 3×3 convolution block (the number of input and output channels is set to 160 and 80 respectively), the aggregated feature map is obtained. Similarly, we get the aggregate feature map of the later phase image B:
[0082] Therefore, this embodiment uses channel connection to fuse the features extracted by the Segformer feature extractor and the FastSAM feature extractor pixel by pixel, which not only retains the global semantic information but also the detail features. Compared with other fusion methods (such as simple weighted summation), channel connection can effectively avoid information loss and provide richer feature input for subsequent decoders.
[0083] Implementation method three: This implementation method is a further limitation of the image target detection method based on a deep learning network described in implementation method one. The difference feature map is decoded by a decoder module, specifically:
[0084] The difference feature map is decoded by multiple residual blocks, 3×3 convolution blocks and classifier modules in sequence.
[0085] In this implementation manner, the difference feature map is obtained by performing absolute difference calculation by the difference feature module, specifically:
[0086] The difference feature calculation method at any position (m,n) in the image is:
[0087]
[0088] Therefore, the difference feature module is used to perform an absolute difference operation on the aggregated feature maps of the before and after phase images, thereby obtaining a difference feature map D.
[0089] After the difference feature map D passes through the decoder module, the change detection prediction map F is obtained out The decoder module consists of a residual block, a 3×3 convolutional block, and a classifier module, specifically:
[0090] Multiple residual blocks do not change the feature map dimension (128 × 128 × 80). A 3 × 3 convolution block (with 80 input channels and 8 output channels) is connected in series after the residual block, which transforms the feature map dimension from 128 × 128 × 80 to 128 × 128 × 8. The classifier module is used to output the binary change map of the network model.
[0091] Implementation method 4: This implementation method is a further limitation of the image target detection method based on a deep learning network described in implementation method 3, and an attention module is connected in series after the second BN layer of the multiple residual blocks.
[0092] In this implementation, an attention module is connected in series after the second BN layer of each residual block, thereby avoiding gradient vanishing, increasing network depth, and improving model expression ability, and multiple residual blocks (for example, 5) are stacked.
[0093] Embodiment 5: This embodiment further limits the method for detecting an image target based on a deep learning network described in Embodiment 4. The attention module is specifically:
[0094] The feature map is divided into two paths. One feature map passes through the global average pooling layer for spatial feature compression and outputs the first feature map. The other feature map passes through the global maximum pooling layer for spatial feature compression and outputs the second feature map.
[0095] After both the first feature map and the second feature map are convolved, the convolved first feature map and the convolved second feature map are output respectively;
[0096] The first feature map after convolution and the second feature map after convolution are added in sequence and operated with Sigmoid function, and then the third feature map is output;
[0097] The feature map is multiplied point by point with the third feature map, and a feature map with channel attention is output.
[0098] In this embodiment, if Figure 5 As shown in FIG. 1 , the attention module includes a global average pooling layer, a global maximum pooling layer, a one-dimensional convolution layer, a sigmoid function layer, an addition calculation layer, and a multiplication calculation layer. The global average pooling (GAP) layer and the global maximum pooling (GMP) layer are used to input feature maps F 0 Perform spatial feature compression. The global average pooling operation obtains a feature map F with a dimension of 1×1×C 1 , the global maximum pooling obtains the feature map F with a dimension of 1×1×C 2 The one-dimensional convolutional layer is used to transform the feature map F 1 and feature map F 2 Perform a one-dimensional convolution operation to obtain F 12 and F 22 Among them, the size k of the one-dimensional convolution kernel is set to 3. The Sigmoid function layer is used to transform the feature map F 12 and feature map F 22 The sum of the F 3 Perform Sigmoid function operation and obtain the feature map F 4 . 4 Expand to F 0 size, and with the input feature map F 0 Multiply point by point to get the feature map F with channel attention 5 The attention module of this patent can learn the input feature map F 0 The dependency between channels is used to obtain the weight information of different channels and improve the performance of the change detection model.
[0099] Implementation method 6: This implementation method further limits the method for detecting an image target based on a deep learning network described in implementation method 3. The classifier module is specifically:
[0100] The feature map passes through the upsampling layer, 3×3 convolution block and upsampling layer in sequence, and outputs a binary change map.
[0101] In this embodiment, the classifier module is composed of an upsampling layer, a 3×3 convolution block, and an upsampling layer in sequence. The upsampling layer is used to upsample the feature map from 128×128×8 to 256×256×8, the 3×3 convolution block (the number of input channels is 8 and the number of output channels is 1) is used to convert the 256×256×8 feature map into a binary change map of 256×256×1, and the upsampling layer is used to sample the binary change map back to the input image size of 512×512.
[0102] Therefore, the present invention combines the two feature extraction models, SegFormer and FastSAM, to enhance the ability of global semantic modeling and local feature capture. The optimized design for change detection tasks provides new ideas for multi-scale feature expression and robustness detection of dual-phase images. The innovation of feature fusion method and decoder structure uses channel connection and decoder optimization to achieve efficient feature fusion and decoding.
[0103] Implementation method 7. A general change detection method for high-resolution satellite remote sensing images based on deep learning described in this implementation method, wherein the method is implemented by a method for detecting image targets based on a deep learning network described in any one of implementation methods 1 to 6, and includes the following steps:
[0104] Step S1, selecting a plurality of satellite remote sensing image pairs taken at different times in the same geographical area, wherein the plurality of satellite remote sensing image pairs all include a pre-phase image and a post-phase image;
[0105] Step S2, preprocessing the multiple groups of satellite remote sensing images respectively;
[0106] Step S3, respectively outlining the change spots in the multiple sets of satellite remote sensing images and generating corresponding label maps;
[0107] Step S4, segmenting and slicing the plurality of satellite remote sensing image pairs and their corresponding label images to produce a slice sample set;
[0108] Step S5, after amplifying the slice sample set, input it into the deep learning network for training;
[0109] In step S6, the multiple groups of satellite remote sensing images to be detected are input into the trained deep learning network for change detection, and multi-scale prediction is adopted to obtain the intersection of the multi-scale prediction results to obtain a general change detection result.
[0110] In this embodiment, in the step S3, the changing spots in the previous phase image and the next phase image are delineated respectively, specifically:
[0111] The change spots in the previous phase image and the next phase image are outlined respectively, and the outlined lines need to fit the boundaries of the change spots.
[0112] In this embodiment, in step S3, the label image includes changed spots and unchanged spots;
[0113] The pixel value of the changed spot is marked as 1, and the pixel value of the unchanged spot is marked as 0;
[0114] The area of the changed pattern is greater than 20 square meters.
[0115] In this embodiment, a method for detecting an image target based on a deep learning network described in any one of embodiments 1 to 6 is applied to satellite remote sensing images, such as Figure 1 As shown in the figure, a general change detection method for high-resolution satellite remote sensing images based on deep learning is proposed, which includes the following steps:
[0116] Step S1, select multiple groups of satellite remote sensing image pairs taken at different times in the same geographical area, each group of image pairs includes a pre-phase image and a post-phase image. The selected images must meet the following requirements: there is basically no cloud, snow or other obstruction in the image, the image contains changes in various basic landform types such as buildings, cultivated land, woodland, water bodies, roads, greenhouses, photovoltaics, bare soil, etc., and the image resolution is sub-meter level (such as 0.75m or 0.5m).
[0117] Step S2, preprocessing the remote sensing image, including radiation correction, geometric registration, image fusion, image mosaicking, image cropping, etc.
[0118] Step S3, outline the change spots in the previous and next phase images and generate corresponding label maps. Outlining is performed according to the following rules: outline the area where the basic landform type has changed, and the outline line needs to fit the boundary of the change area; in the label map, the pixel value of the change area is marked as 1, and the pixel value of the unchanged area is marked as 0; only the change area with an area greater than 20 square meters is outlined; shadows, roof color changes, seasonal changes in forests and grasses, etc. are regarded as non-real changes (also called pseudo changes), and they also need to be marked, and the corresponding pixels in the label map are marked as 0.
[0119] Step S4: Based on the above outline results, a slice sample set is prepared. The sample set only retains the RGB band of the original satellite image, and synchronously segments and slices the pre-phase image, post-phase image and label image according to a 512×512 window. The sample set slice example is as follows: Figure 2-Figure 4 Data augmentation is performed on the before and after phase images and label image slices, including horizontal flipping, vertical flipping, random rotation, brightness change, saturation change, noise addition, etc.
[0120] Step S5, input the sample set after data amplification into the deep learning model for training. In the training stage of the model, the overall network loss function selects Cross-Entropy (CE), such as:
[0121]
[0122] Among them, N represents the number of categories, y i is the true label value, is the predicted value.
[0123] The Adam algorithm is used for optimization training, with weight decay set to 0.01 and beta value set to (0.9, 0.999). The learning rate is initially set to 0.0001 and linear decay is 0. The network model is trained according to the above parameter settings until convergence.
[0124] Step S6, in the prediction stage of the model, multi-scale prediction (scaling factor is [0.5, 1.0, 1.5, 2]) is used for the before and after phase images input into the model, and the prediction results at different scales are intersected to obtain the final general change detection binary raster image result.
[0125] Therefore, after inputting the front-phase and the back-phase images, this embodiment will perform feature extraction through two different feature extractors, FastSAM and Segformer; the multi-level feature aggregation module will fuse the feature maps of different categories and levels; the difference feature module will perform absolute difference calculation on the aggregated features of the front-phase image and the back-phase image to obtain the difference features containing multi-scale information; the decoder will process the difference features to obtain the change detection extraction results.
[0126] In order to better illustrate the general change detection method for high-resolution satellite remote sensing images based on deep learning described in this embodiment, a detailed description is given through the following examples:
[0127] like Figure 6 As shown in FIG. 1 , the before-phase image, the after-phase image and the label image are shown. It can be seen that the after-phase image is in a weaker lighting condition and has smaller changing targets such as houses.
[0128] like Figure 7 As shown, the results of the detection by the general change detection method of high-resolution satellite remote sensing images based on deep learning described in this embodiment are compared with the results of the detection by the prior art. The results of the detection by the method described in this embodiment are closer to those of Figure 6 As shown in the label diagram, the result of using the existing technology for detection will result in the omission of some small target change spots.
[0129] Therefore, the general change detection method for high-resolution satellite remote sensing images based on deep learning described in this embodiment is significantly superior to existing solutions in terms of change detection accuracy, robustness and applicability, especially in scenes with large illumination differences and significant changes in small targets, which can reduce false detections and missed detections. This technology is suitable for processing large-scale remote sensing images and has wide application value and good social and economic benefits.
[0130] The above is a detailed introduction to a general change detection method for high-resolution satellite remote sensing images based on deep learning proposed in the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A method for detecting image targets based on a deep learning network, characterized in that: Specifically: The images are divided into two paths. One path is extracted by FastSAM feature extractor and the feature map of the first scale is output. The other path is extracted by Segformer feature extractor and the feature map of the second scale is output. The feature map of the first scale and the feature map of the second scale are input into the multi-level feature aggregation module for pixel-by-pixel fusion, and the fused feature map is output; The fused feature map is passed through the difference feature module for absolute difference calculation, and the difference feature map is output; The difference feature map is decoded by the decoder module and the image target detection result is output.
2. According to the method for detecting image targets based on deep learning network in claim 1, it is characterized in that: The feature map of the first scale and the feature map of the second scale are input into the multi-level feature aggregation module for pixel-by-pixel fusion, and the fused feature map is output, specifically: The feature map of the first scale and the feature map of the second scale are respectively input into the channel conversion unit for channel dimension conversion, and the feature map of the first scale and the feature map of the second scale with the same channel dimension are output; The feature map of the first scale and the feature map of the second scale with the same channel dimension are input into the channel connection unit for connection to obtain a multi-level feature map; Multi-level feature maps are aggregated through the hierarchical aggregation unit to output the fused feature map.
3. The method for detecting an image target based on a deep learning network according to claim 2, characterized in that: The multi-level feature graphs are aggregated through the hierarchical aggregation unit, specifically: The multi-level feature maps are aggregated through the upsampling layer, channel connection layer and 3×3 convolutional block in sequence.
4. The method for detecting an image target based on a deep learning network according to claim 1, characterized in that: The difference feature map is decoded by the decoder module, specifically: The difference feature map is decoded by multiple residual blocks, 3×3 convolution blocks and classifier modules in sequence.
5. The method for detecting an image target based on a deep learning network according to claim 4, characterized in that: An attention module is connected in series after the second BN layer of each of the multiple residual blocks.
6. The method for detecting an image target based on a deep learning network according to claim 5, characterized in that: The attention module is specifically: The feature map is divided into two paths. One feature map passes through the global average pooling layer for spatial feature compression and outputs the first feature map. The other feature map passes through the global maximum pooling layer for spatial feature compression and outputs the second feature map. After both the first feature map and the second feature map are convolved, the convolved first feature map and the convolved second feature map are output respectively; The first feature map after convolution and the second feature map after convolution are added in sequence and operated with Sigmoid function, and then the third feature map is output; The feature map is multiplied point by point with the third feature map, and a feature map with channel attention is output.
7. The method for detecting an image target based on a deep learning network according to claim 4, characterized in that: The classifier module is specifically: The feature map passes through the upsampling layer, 3×3 convolution block and upsampling layer in sequence, and outputs a binary change map.
8. A general change detection method for high-resolution satellite remote sensing images based on deep learning, wherein the method is implemented by using the image target detection method based on a deep learning network as described in any one of claims 1 to 7, characterized in that: The following steps are involved: Step S1, selecting a plurality of satellite remote sensing image pairs taken at different times in the same geographical area, wherein the plurality of satellite remote sensing image pairs all include a pre-phase image and a post-phase image; Step S2, preprocessing the multiple groups of satellite remote sensing images respectively; Step S3, respectively outlining the change spots in the multiple sets of satellite remote sensing images and generating corresponding label maps; Step S4, segmenting and slicing the plurality of satellite remote sensing image pairs and their corresponding label images to produce a slice sample set; Step S5, after amplifying the slice sample set, input it into the deep learning network for training; In step S6, the multiple groups of satellite remote sensing images to be detected are input into the trained deep learning network for change detection, and multi-scale prediction is adopted to obtain the intersection of the multi-scale prediction results to obtain a general change detection result.
9. The method of general change detection of high-resolution satellite remote sensing images based on deep learning according to claim 8, characterized in that: In the step S3, the changing spots in the previous phase image and the next phase image are delineated respectively, specifically: The change spots in the previous phase image and the next phase image are outlined respectively, and the outlined lines need to fit the boundaries of the change spots.
10. The method of general change detection of high-resolution satellite remote sensing images based on deep learning according to claim 8, characterized in that: In the step S3, the label map includes changed spots and unchanged spots; The pixel value of the changed spot is marked as 1, and the pixel value of the unchanged spot is marked as 0; The area of the changed pattern is greater than 20 square meters.
Citation Information
Cited By
Industrial anomaly detection method and device based on multi-feature fusion
CN120833310A