Pyramid-imitated remote sensing image intelligent interpretation method
Through a multi-scale model that imitates pyramids, multi-scale scaling of the human eye is simulated, the shortcomings of shape and spatial relationship feature extraction in remote sensing images are solved, and the fine and accurate recognition of remote sensing elements is achieved, and the training efficiency and prediction accuracy of the model are improved.
Patent Information
- Application Number
- CN202510455303.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-25
AI Technical Summary
The existing intelligent interpretation technology lacks multi-scale vision and is difficult to extract features such as shape and spatial relationships in remote sensing images, resulting in incomplete land objects during the cropping process, too small receptive field, low edge accuracy, and huge video memory occupies when model training, and an upper limit for input size.
Using the method of imitating pyramids, multi-scale information is reflected through multi-layer pyramids, multi-scale scaling is simulated when human eyes interpret targets, and multi-scale models are established, including large-scale coarse segmentation layers and small-scale sub-segmentation layers. Pyramid index and layer-by-layer freezing training are used to achieve the acquisition and integration of multi-scale information.
The fine and accurate identification of remote sensing element types is achieved, and the problems of incomplete land and small receptive fields in remote sensing image segmentation are solved, and the training efficiency and prediction accuracy of the model are improved.
Smart Images

Figure CN120375060A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image interpretation, and specifically to an intelligent remote sensing image interpretation method imitating a pyramid. Background Art
[0002] The development of remote sensing image interpretation technology has gone through stages such as manual visual interpretation, image processing algorithm-assisted interpretation, and intelligent interpretation. In recent years, with the rapid development of artificial intelligence technology, it has evolved from "not applicable" and "not easy to use" to the stage of "usable". Therefore, intelligent interpretation technology has also begun to be widely applied to remote sensing information interpretation.
[0003] Compared with traditional manual visual interpretation methods and image processing algorithm-assisted interpretation methods, intelligent interpretation methods combine the high accuracy of visual interpretation and the high efficiency of image processing methods. In the prior art, remote sensing intelligent interpretation technology has played an important role in the interpretation of information on typical ground objects such as cultivated land and buildings.
[0004] In remote sensing image interpretation technology, the core is always the remote sensing interpretation marks. Among them, the visual interpretation method is that experts interpret according to the interpretation marks in their experience, and the image processing method is that experts solidify different indicators of remote sensing interpretation marks into algorithms for information interpretation, such as methods like NDVI and decision trees.
[0005] Remote sensing intelligent interpretation is for a neural network to learn the features required for categories from samples, that is, remote sensing interpretation marks, similar to the process of humans learning features. Among remote sensing image information interpretation marks, interpretation indicators such as shape and spatial relationship are as important as spectral indicators. During the visual interpretation process, the human eye can view target information at different scales by zooming, so as to judge the category of ground objects at multiple scales, and interpretation indicators such as spectrum, shape, and spatial relationship at different scales play important roles.
[0006] When performing intelligent interpretation of remote sensing information, neural networks can learn effective feature information in samples, naturally including interpretation indicators such as spectra, shapes, and spatial relationships. However, existing intelligent interpretation technologies lack a multi-scale perspective, making it difficult to extract features such as shapes and spatial relationships and relying more on spectral information. When existing intelligent interpretation technologies are applied to remote sensing images, especially in the field of remote sensing image segmentation, due to the large size of remote sensing images and the number of pixels in a single image exceeding hundreds of millions, they cannot be directly input into the model for training and prediction. The input sizes of existing models are mostly in dimensions such as 256×256 or 512×512. When performing intelligent interpretation operations such as training and inference on remote sensing images, the entire image usually needs to be cropped into standard sizes to meet the input requirements. However, during the cropping process, large features will be cropped into multiple parts, making it difficult to input the features completely into the model. The small segmented images face problems such as too small receptive fields and incomplete features on the image. At the same time, due to the lack of neighborhood information at the image edges, the accuracy of the edge parts is often low. For the problems of incomplete features and too small receptive fields, techniques such as increasing the input size and dilated convolution are often used to solve them. However, increasing the size of the input to the model will also lead to an exponential increase in the model, and the video memory occupancy during training is huge. Constrained by the video memory, there is an upper limit to the input size.
[0007] Therefore, it is necessary to design a pyramid-like intelligent interpretation method for remote sensing images, which can simulate the multi-scale zoom retrieval of the human eye when interpreting targets and reflect multi-scale information through multiple layers of pyramids, enabling the model to utilize information at multiple scales during training and prediction, thereby fully obtaining information such as shapes and spatial relationships and achieving fine and accurate identification of feature types. Summary of the Invention
[0008] The object of the present invention is to overcome the deficiencies of the prior art and provide a pyramid-like intelligent interpretation method for remote sensing images, which can simulate the multi-scale zoom retrieval of the human eye when interpreting targets and reflect multi-scale information through multiple layers of pyramids, enabling the model to utilize information at multiple scales during training and prediction, thereby fully obtaining information such as shapes and spatial relationships and achieving fine and accurate identification of feature types.
[0009] To achieve the above object, the present invention provides a pyramid-like intelligent interpretation method for remote sensing images: It includes the following steps: S1, Interpretation element selection and image preprocessing: Select typical target areas and corresponding high-quality remote sensing images for remote sensing intelligent interpretation of target information, specifically including: S1-1, Data selection: Select multi-source high-spatial-resolution remote sensing images with similar resolutions and the same bands as the data source; S1-2, Data preprocessing: Perform atmospheric correction and geometric correction on multi-source original high-resolution images to generate high-quality ortho-images; S2, Feature information extraction: Interpret the information of each image, extract all target regions on the image, and prepare for subsequent sample production, specifically including: S2-1, Image multi-scale segmentation: The process of generating meaningful polygons with the least heterogeneity and the greatest homogeneity at any scale. Starting from any pixel, use a bottom-up region merging method to form objects; S2-2, Assignment and manual editing of segmented regions: Assign attributes to the segmented results, assign the target region as 1, and other regions as 0; if there are multiple features, the features are assigned as 1, 2, 3... in sequence, and other regions are assigned as 0; during the assignment process, check the boundaries of the segmentation, correct the automatic segmentation results, and fine-tune the boundaries; S2-3, Rasterization of feature information: Rasterize the extracted feature information. The rasterized feature information has the same size and corresponding position as the original image, and the numerical values in the feature attributes are converted into the pixel values of the raster; S3, Establish pyramid samples and indexes: Refer to the pyramid loading technology used when displaying large images. When the scale is small, display small-scale thumbnails. During the zooming process, when the scale reaches a certain layer of the pyramid, display that layer of the pyramid for fast display, avoiding lags during the display process, and establish indexes, specifically including: S3-1, Determination of the number of pyramid layers: The number of pyramid layers is determined by the inter-layer ratio and the sample size. The inter-layer ratio is the scale ratio between two layers of the pyramid. The inter-layer ratio t is a non-zero power of 2 for the needs of model input size and index establishment; for features greatly affected by scale, t is taken as 2 when it is small; for features less affected by scale, t is taken as 4 when it is large. The relationship between the number of layers n, the inter-layer ratio t, and the sample size L is as follows: L>t n-1 ; Among them, if the model input size is 512×512, the inter-layer ratio t is taken as 4, and the number of layers n is taken as 3, then the pixel size of each layer is 4 times that of the previous layer, and the pixel size of the third layer of the pyramid is 16 times that of the original pixel size; S3-2, Establish a pyramid: According to the determined number of pyramid layers, perform downsampling from the bottom layer and build the pyramid layer by layer; that is, each layer is 1 / t of the previous layer; the downsampling method selects average pooling, and the same pyramid construction is performed on the remote sensing image data and the corresponding rasterized feature information to ensure the one-to-one correspondence between the pyramid of the remote sensing image and the feature information; S3-3, Pyramid index: According to the established pyramid, establish index numbers from bottom to top; S3-4. Establish a sample: The sample is an n-layer pyramid sample, with the size of each layer being L, and the middle pixel of each layer being the target area of this sample; S4. Establish a multi-scale model: The multi-scale model consists of two parts. The first part is a large-scale rough segmentation layer, and the second part is a small-scale fine segmentation; The large-scale rough segmentation layer adopts a network structure with relatively simple structure, fast convergence speed, and relatively low sample dependence; The small-scale fine segmentation layer adopts a network structure with high segmentation accuracy; The large-scale rough segmentation layer is responsible for the segmentation of the upper-layer pyramid, and inputs the segmentation result together with the lower-layer pyramid into the small-scale fine segmentation layer; The fine segmentation layer combines the result of the upper-layer pyramid with the image of this layer pyramid as input data for segmentation, so as to realize adding the information of the upper-layer pyramid during segmentation. Specifically, it includes: S4-1. Large-scale rough segmentation layer: The large-scale rough segmentation layer adopts a simplified U-shaped network structure, which has the characteristics of fast convergence speed and small samples, and is suitable for quickly reaching a relatively high segmentation level; The overall structure is to first encode through downsampling, then decode, and regress to pixel point classification of the same size as the original image; The structure of the large-scale rough segmentation layer is divided into three parts: downsampling, upsampling, and skip connection; First, divide the large-scale rough segmentation layer network into left and right parts for analysis. The left side is the compression process, that is, encoding; The image size is reduced through convolution and downsampling to extract some superficial features; The right part is the decoding process, that is, decoding; Some deep features are obtained through convolution and upsampling; The output result is a normalized heat map; S4-2. Small-scale fine segmentation layer: The small-scale fine segmentation layer adopts the DeepLabV3+ network with the CBAM dual attention mechanism added to fuse multi-scale information on the sample; The CBAM dual attention mechanism is connected in parallel with the ASPP structure; S4-3. Overall model structure: The overall model consists of a group of large-scale rough segmentation layers and small-scale fine segmentation layers. The large-scale segmentation layer removes the classification module and outputs a probability layer; The small-scale segmentation layer removes the shallow convolution results of the convolutional neural network and is replaced by the output result of the large-scale segmentation layer; Since the range of the upper-layer pyramid is 4 times that of the lower-layer pyramid, an index area number is added to realize the integration of large-scale information and small-scale information; S5. Model multi-scale training: The training of the multi-scale model adopts a training method of layer-by-layer freezing; Since the model spans multiple pyramids and the hierarchy is relatively deep after multiple nestings of the model, during training, start training from the bottom-layer pyramid, freeze other layers, and then train layer by layer of the pyramid. Every time the pyramid rises one layer, the field of view expands 4 times, and finally conduct overall training; S6. Multi-scale inference: The trained model needs to be split and split according to the corresponding pyramid layers for easy multi-scale inference. The inference process is as follows: S6-1. Read the model parameters and split them into corresponding network layers according to the number of pyramid levels. S6-2. Obtain the nth-level pyramid, crop it to a size of 512×512, and use the network layer corresponding to the nth-level pyramid for inference. The inference results are merged into the nth-level predicted feature map. S6-3. Obtain the (n - 1)th-level pyramid, crop it to a size of 512×512, and at the same time, crop the nth-level predicted feature map according to the index. The cropped nth-level predicted feature map and the (n - 1)th-level pyramid are used as new inputs, and the network layer corresponding to the (n - 1)th-level pyramid is used for inference. The inference results are merged into the (n - 1)th-level predicted feature map. S6-4. Repeat steps S6-2 and S6-3 until the bottom-level predicted feature map is inferred. S6-5. Use the classifier of the model, input the bottom-level predicted feature map, and generate the segmentation result.
[0010] Compared with the prior art, the present invention has the following beneficial effects: By introducing the pyramid technology for accelerating image display, the present invention uses the pyramid to simulate the zooming when the human eye recognizes target elements, digitizes the continuous zooming into multiple scales, and thus simulates the extraction of features such as position and spatial relationship by the human eye during the image zooming process, thereby realizing the multi-scale interpretation of remote sensing element information. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is a schematic diagram of the mangrove sample structure according to an embodiment of the present invention.
[0012] Figure 2 It is a schematic diagram of the mangrove rough segmentation layer structure according to an embodiment of the present invention.
[0013] Figure 3 It is a schematic diagram of the mangrove fine segmentation layer structure according to an embodiment of the present invention.
[0014] Figure 4 It is a schematic diagram of the overall model structure of the mangrove according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] Refer to Figures 1 to 4 , and now the present invention will be further described in conjunction with the accompanying drawings. In this embodiment, mangroves are used as the remote sensing target area for further illustration: It includes the following steps: 1. Selection of interpretation elements and image preprocessing. Select typical mangrove areas and corresponding high-quality remote sensing images for the remote sensing intelligent interpretation of mangrove information.
[0016] 1.1 Data selection: Select multi-source high-spatial-resolution remote sensing images with similar resolutions and the same bands as the data source; 1.2 Data preprocessing: Perform atmospheric correction and geometric correction on multi-source original high-resolution images to generate high-quality ortho-images. 2. Feature information extraction. Interpret the information of each image, extract all mangrove areas on the image, and prepare for subsequent sample production.
[0017] 2.1 Image multi-scale segmentation. The multi-scale segmentation method is a process of generating meaningful polygons with the least heterogeneity and the greatest homogeneity at any scale. Starting from any pixel, a bottom-up region merging method is used to form objects.
[0018] 2.2 Assignment and manual editing of segmented regions. Assign attributes to the segmented results, assign the mangrove area as 1 and other areas as 0. If there are multiple features, the features are assigned as 1, 2, 3... in sequence, and other areas are assigned as 0. During the assignment process, check the boundaries of the segmentation, correct the automatic segmentation results, and fine-tune the boundaries.
[0019] 2.3 Rasterization of feature information. Rasterize the extracted feature information. The rasterized feature information has the same size and corresponding position as the original image, and the numerical values in the feature attributes are converted into the pixel values of the raster.
[0020] 3. Establish pyramid samples and indexes. Referring to the pyramid loading technology used when displaying large images, when the scale is small, a small-scale thumbnail is displayed. During the zooming process, when the scale reaches a certain layer of the pyramid, that layer of the pyramid is displayed, thus achieving fast display and avoiding lags during the display process. This method borrows this technology to establish pyramid-style samples and establish indexes.
[0021] 3.1 Determination of the number of pyramid layers. The number of pyramid layers in this method is determined by the inter-layer ratio and the sample size. The inter-layer ratio is the scale ratio between two layers of the pyramid. Due to the model input size and the need to establish indexes, the inter-layer ratio t is a non-zero power of 2. For features that are more affected by scale, t is smaller and can be taken as 2; for features that are less affected by scale, t is larger and can be taken as 4. The relationship between the number of layers n, the inter-layer ratio t, and the sample size L is as follows: L>tn-1 In this method, the model input size is 512x512, the inter-layer ratio t is taken as 4, and the number of layers n is taken as 3. Then the pixel size of each layer is 4 times that of the previous layer, and the pixel size of the third layer of the pyramid is 16 times that of the original pixel size. 3.2 Establish the pyramid. According to the determined number of pyramid layers, downsample from the bottom layer and construct the pyramid layer by layer. That is, each layer is 1 / t of the previous layer. The downsampling method selects average pooling. The same pyramid construction is performed on the remote sensing image data and the corresponding rasterized feature information to ensure that the pyramids of the remote sensing image and the feature information correspond one by one.
[0022] 3.3 Pyramid Index. According to the established pyramid, index numbers are established from bottom to top.
[0023] 3.4 Establishing Samples. The samples are n-layer pyramid samples, with the size of each layer being L, and the middle pixel of each layer being the target area of this sample, as Figure 1 shown.
[0024] 4. Establishing a Multi-scale Model. The multi-scale model consists of two parts. The first part is the large-scale rough segmentation layer, and the second part is the small-scale fine segmentation. The large-scale segmentation layer adopts a network structure with relatively simple structure, fast convergence speed, and relatively low sample dependence. The small-scale segmentation layer adopts a network structure with high segmentation accuracy. The large-scale rough segmentation layer is mainly responsible for the segmentation of the upper-layer pyramid. The segmentation result and the lower-layer pyramid are input into the small-scale fine segmentation layer together. The fine segmentation layer combines the result of the upper-layer pyramid with the image of this layer pyramid as input data for segmentation, so as to add the information of the upper-layer pyramid during segmentation.
[0025] 4.1 Large-scale Rough Segmentation Layer. The large-scale rough segmentation layer adopts a simplified U-shaped network structure, which has the characteristics of fast convergence speed and small samples, and is suitable for quickly reaching a relatively high segmentation level. The overall structure is to encode (downsample) first, and then decode, and return to pixel classification of the same size as the original image. Its structure is as Figure 2 shown.
[0026] It can be seen that this network structure is mainly divided into three parts: downsampling, upsampling, and skip connection. First, the network is divided into left and right parts for analysis. The left side is the compression process, that is, encoding. The image size is reduced through convolution and downsampling to extract some superficial features. The right part is the decoding process, that is, decoding. Some deep features are obtained through convolution and upsampling. The last layer performs classification through 1x1 convolution. This method does not need to perform classification at this step. The last layer will be removed, and the output result is a normalized heat map.
[0027] 4.2 Small-scale Fine Segmentation Layer. The small-scale fine segmentation layer adopts the DeepLabV3+ network with an attention mechanism added. Deeplab v3+ has an Atrous Spatial Pyramid Pooling (ASPP) structure and an Encoder-Decoder structure, which can better fuse multi-scale information on samples. The CBAM double attention mechanism has good use effects on middle-level features or high-level features. It is added to the ASPP structure of the DeepLabv3+ network and is connected in parallel with the ASPP structure. The model structure of the fine segmentation layer is as Figure 3 shown.
[0028] 4.4 Overall model structure. The overall model consists of a set of large-scale coarse segmentation layers and small-scale fine segmentation layers. The large-scale segmentation layers remove the classification module and output a probability layer. The small-scale segmentation layers remove the shallow convolutional results of the convolutional neural network and are replaced by the output results of the large-scale segmentation layers. Since the range of the upper pyramid is 4 times that of the lower pyramid, the index area number is increased to integrate large-scale information and small-scale information. The modified structure is as shown in Figure 4 the following figure.
[0029] 5. Multi-scale training of the model. The training of the multi-scale model adopts a layer-by-layer freezing training method. Since the model spans multiple pyramids and has a relatively deep hierarchy after multiple nestings, during training, start training from the bottommost pyramid, freeze other layers, and then train layer by layer for each pyramid. For each layer the pyramid rises, the field of view expands by 4 times, and finally, overall training is carried out.
[0030] 6. Multi-scale inference. The trained model needs to be split and split according to the corresponding pyramid levels for convenient multi-scale inference. The inference process is as follows: (1) Read the model parameters and split them into corresponding network layers according to the pyramid levels; (2) Obtain the nth layer pyramid, crop it to a size of 512x512, and use the network layer corresponding to the nth layer pyramid for inference. The inference results are merged into the nth layer predicted feature map.
[0031] (3) Obtain the (n - 1)th layer pyramid, crop it to a size of 512x512, and at the same time, according to the index, crop the nth layer predicted feature map. The cropped nth layer predicted feature map and the (n - 1)th layer pyramid are used as new inputs, and the network layer corresponding to the (n - 1)th layer pyramid is used for inference. The inference results are merged into the (n - 1)th layer predicted feature map.
[0032] (4) Repeat steps (2) and (3) until the bottommost predicted feature map is inferred.
[0033] (5) Use the classifier of the model, input the bottommost predicted feature map, and generate a segmentation result.
[0034] The above is only the preferred implementation manner of the present invention, which is only used to help understand the method and its core idea of the present application. The protection scope of the present invention is not limited to the above embodiments. Any technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements should also be regarded as within the protection scope of the present invention.
[0035] The above embodiments can be implemented in whole or in part by software or hardware. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium.
[0036] The present invention solves the problems in the prior art that the intelligent interpretation technology lacks a multi-scale vision, is difficult to extract features such as shape and spatial relationship, relies more on spectral information, is difficult to input the ground objects completely into the model when cropping remote sensing images, the small images after segmentation face the problems of too small receptive fields, incomplete ground objects on the map, and relatively low accuracy in the edge part. When using techniques such as increasing the input size and dilated convolution to solve the problems of incomplete ground objects and too small receptive fields, it brings problems such as an exponential increase in the model, huge video memory occupation during training, being restricted by the video memory, and an upper limit on the input size. By extracting the information to be extracted from the entire image, then building a pyramid and index for the image and the extracted elements, making multi-scale element samples, and finally building a multi-scale model, after training, a multi-scale interpretation model is formed to realize the multi-scale interpretation of remote sensing element information. The present invention can simulate the multi-scale zoom retrieval of the human eye when interpreting targets, and reflect the multi-scale information through multiple layers of pyramids, enabling the model to utilize information at multiple scales during training and prediction, thereby fully obtaining information such as shape and spatial relationship, and realizing the fine and accurate recognition of element types.
Claims
1. An intelligent interpretation method for remote sensing images imitating a pyramid, characterized in that, It includes the following steps: S1, Interpretation Element Selection and Image Preprocessing: Select typical target areas and corresponding high-quality remote sensing images for the intelligent remote sensing interpretation of target information, specifically including: S1-1, Data Selection: Select multi-source high-spatial-resolution remote sensing images with similar resolutions and the same bands as the data source; S1-2, Data Preprocessing: Perform atmospheric correction and geometric correction on the multi-source original high-resolution images to generate high-quality ortho-images; S2, Element Information Extraction: Interpret the information of each image, extract all target areas on the image, and prepare for subsequent sample production, specifically including: S2-1, Image Multi-scale Segmentation: The process of generating meaningful polygons with the least heterogeneity and the greatest homogeneity at any scale. Starting from any pixel, use a bottom-up region merging method to form objects; S2-2, Assignment and Manual Editing of Segmented Regions: Assign attributes to the segmented results, assign the target area as 1 and other areas as 0; if there are multiple elements, the elements are assigned as 1, 2, 3... in sequence, and other areas are assigned as 0; during the assignment process, check the boundaries of the segmentation, correct the automatic segmentation results, and fine-tune the boundaries; S2-3, Rasterization of Element Information: Rasterize the extracted element information. The rasterized element information has the same size and corresponding position as the original image, and the numerical values in the element attributes are converted into the pixel values of the raster; S3, Establishing Pyramid Samples and Indexes: Referring to the pyramid loading technology used when displaying large images, when the scale is small, display small-scale thumbnails. During the zooming process, when the scale reaches a certain layer of the pyramid, display that layer of the pyramid for fast display, avoiding lags during the display process, and establish an index, specifically including: S3-1, Determination of the Number of Pyramid Layers: The number of pyramid layers is determined by the inter-layer ratio and the sample size. The inter-layer ratio is the scale ratio between two layers of the pyramid. The inter-layer ratio t is a non-zero power of 2 for the needs of the model input size and index establishment; for elements greatly affected by the scale, t is taken as 2 when it is small; for elements less affected by the scale, t is taken as 4 when it is large. The relationship between the number of layers n, the inter-layer ratio t, and the sample size L is as follows: L>t n-1 ; Among them, the size of the model input is 512×512, the inter-layer ratio t is taken as 4, and the number of layers n is taken as 3. Then the number of pixels in each layer is 4 times that of the previous layer, and the pixel size of the third layer of the pyramid is 16 times that of the original pixel size; S3-2, Establishing the Pyramid: According to the determined number of pyramid layers, perform downsampling from the bottom layer and construct the pyramid layer by layer; that is, each layer is 1 / t of the previous layer; the downsampling method selects average pooling, and the same pyramid construction is performed on the remote sensing image data and the corresponding rasterized element information to ensure the one-to-one correspondence between the pyramid of the remote sensing image and the element information; S3-3, Pyramid Index: Establish index numbers from bottom to top according to the established pyramid; S3-4, Establishing Samples: The samples are n-layer pyramid samples, the size of each layer is L, and the middle pixel of each layer is the target area of this sample; S4. Establish a multi-scale model: The multi-scale model consists of two parts. The first part is the large-scale rough segmentation layer, and the second part is the small-scale fine segmentation. The large-scale rough segmentation layer adopts a network structure with relatively simple structure, fast convergence speed, and relatively low sample dependence. The small-scale fine segmentation layer adopts a network structure with high segmentation accuracy. The large-scale rough segmentation layer is responsible for the segmentation of the upper-layer pyramid. The segmentation result and the lower-layer pyramid are input into the small-scale fine segmentation layer together. The fine segmentation layer combines the result of the upper-layer pyramid with the image of the current layer pyramid as input data for segmentation, so as to add the information of the upper-layer pyramid during segmentation. Specifically, it includes: S4-1. Large-scale rough segmentation layer: The large-scale rough segmentation layer adopts a simplified U-shaped network structure, which has the characteristics of fast convergence speed and small samples, and is suitable for quickly reaching a relatively high segmentation level. The overall structure is to first encode through downsampling, then decode, and return to pixel point classification of the same size as the original image. The structure of the large-scale rough segmentation layer is divided into three parts: downsampling, upsampling, and skip connection. First, divide the large-scale rough segmentation layer network into left and right parts for analysis. The left part is the compression process, that is, encoding; the image size is reduced through convolution and downsampling to extract some superficial features. The right part is the decoding process, that is, decoding; some deep features are obtained through convolution and upsampling. The output result is a normalized heat map. S4-2. Small-scale fine segmentation layer: The small-scale fine segmentation layer adopts the DeepLabV3+ network with the CBAM double attention mechanism to fuse multi-scale information on the samples. The CBAM double attention mechanism is connected in parallel with the ASPP structure. S4-3. Overall model structure: The overall model consists of a group of large-scale rough segmentation layers and small-scale fine segmentation layers. The large-scale segmentation layer removes the classification module and outputs the probability layer. The small-scale segmentation layer removes the shallow convolution results of the convolutional neural network and is replaced by the output result of the large-scale segmentation layer. Since the range of the upper-layer pyramid is 4 times that of the lower-layer pyramid, the index area number is increased to realize the integration of large-scale information and small-scale information. S5. Model multi-scale training: The training of the multi-scale model adopts a training method of layer-by-layer freezing. Since the model spans multiple pyramids and the hierarchy is relatively deep after multiple nestings, during training, start training from the bottom-layer pyramid, freeze other layers, and then train layer by layer for each pyramid. When the pyramid rises one layer, the field of view expands 4 times, and finally conduct overall training. S6. Multi-scale inference: The trained model needs to be split and split according to the corresponding pyramid layers for convenient multi-scale inference. The inference process is as follows: S6-1. Read the model parameters and split them into corresponding network layers according to the pyramid layers. S6-2. Obtain the nth-layer pyramid, crop it to a size of 512×512, and use the network layer corresponding to the nth-layer pyramid for inference. The inference results are merged into the nth-layer predicted feature map. S6-3. Obtain the (n - 1)-th layer pyramid, crop it to a size of 512×512, and at the same time, according to the index, crop the predicted feature map of the n-th layer; The cropped predicted feature map of the n-th layer and the (n - 1)-th layer pyramid are used as new inputs, and the network layer corresponding to the (n - 1)-th layer pyramid is used for inference, and the inference results are merged into the predicted feature map of the (n - 1)-th layer; S6-4. Repeat steps S6-2 and S6-3 until the predicted feature map of the bottom layer is inferred; S6-5. Use the classifier of the model, input the predicted feature map of the bottom layer, and generate the segmentation result.
Citation Information
Patent Citations
Method and system for establishing multisource geospatial information correlation model
CN103488736A
Remote sensing object interpretation method based on focusing weight matrix and variable-scale semantic segmentation neural network
CN110490081A
Open-pit mine area land utilization identification method based on improved DeepLabV3+
CN113435411A
Remote sensing image semantic segmentation method and system based on multi-scale information fusion
CN113780296A
Remote sensing image semantic segmentation method based on pyramid segmentation attention module
CN113807210A