Remote sensing image cross-domain semantic segmentation method based on bionic eye movement mode
By adopting the fovea-eye-eye bionic attention calculation model based on the bionic eye movement mode in the semantic segmentation of high-resolution remote sensing images, the accuracy and reliability problems in cross-domain applications are solved, and efficient visual perception is achieved, which is suitable for the processing of complex remote sensing image features.
Patent Information
- Application Number
- CN202311812257.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-06-27
AI Technical Summary
The semantic segmentation of high-resolution remote sensing images has accuracy and reliability problems in cross-domain applications, and traditional deep learning algorithms lack visual perception capabilities, making it difficult to effectively process complex remote sensing image features.
The fovea-yeol-light bionic attention calculation model based on the bionic eye movement mode is adopted. By constructing the gaze model, the saccade model exceeding the edge and the saccade model limited by the edge, the attention of the eye movement model is calculated, and the cross-domain semantic segmentation of high-resolution remote sensing images is achieved.
It improves the accuracy and reliability of cross-domain semantic segmentation of high-resolution remote sensing images, has visual perception capabilities, and can more effectively process complex remote sensing image features, which are suitable for different sensors and urban market scenarios.
Smart Images

Figure HSA0000297649470000011 
Figure HSA0000297649470000021 
Figure HSA0000297649470000031
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of remote sensing and also relates to the image analysis of high-resolution images. Background Art
[0002] With the increasing attention of more and more researchers to remote sensing technology and machine learning, high-resolution remote sensing images are becoming increasingly easy to obtain. The purpose of semantic segmentation is to divide an image into regions with different category semantic features to obtain a pixel-level semantic annotation image. The semantic segmentation of high-resolution remote sensing images is widely applied in fields such as urban planning, crop assessment, and intelligent transportation. In this context, based on geospatial big data, using artificial intelligence technology to mine its deep information and endow it with more application modes will become the long-term theme for the development of the future geospatial data analysis application field. However, the complexity and diversity of feature information pose great challenges to the semantic segmentation of high-resolution remote sensing images.
[0003] Different from other image datasets, due to the often different data distributions of remote sensing images with different imaging sensors and geographical locations, current general deep learning algorithms are not applicable to the semantic segmentation of high-resolution remote sensing images. The semantic segmentation of high-resolution remote sensing images needs to work reliably and accurately in different sensors and other urban scenarios. After being well-trained on its original source domain, when a traditional deep learning algorithm is applied to a new target domain with data distribution differences, its generalization ability will ultimately degrade.
[0004] With the development of artificial intelligence, more and more deep learning network models are applied in the field of image processing. As a new type of deep learning model, the vision transformer allows the model to capture dependencies in the input sequence at different positions, can efficiently process long-sequence samples, and capture details and complex relationships in the text layer by layer. The disadvantages of the vision transformer are: the structure of the vision transformer lacks visual perception ability, lacks the guidance of the attention center, and will have the problem of deep degradation during the long training process. Summary of the Invention
[0005] In order to solve the problems such as inaccurate and unreliable results and lack of visual perception ability existing in the current cross-domain semantic segmentation of high-resolution remote sensing images, the present invention provides an accurate and reliable cross-domain semantic segmentation method for high-resolution remote sensing images based on a bionic eye movement pattern and with visual perception ability.
[0006] In order to achieve the above invention purpose, the technical solution adopted by the present invention is: a cross-domain semantic segmentation method for remote sensing images based on a bionic eye movement pattern, and the steps include:
[0007] Step 1: Use two remote sensing image datasets as the source domain and the target domain respectively, and obtain training set samples and test set samples after cropping.
[0008] Step 2: Construct a fovea-peripheral vision bionic attention calculation model;
[0009] Step 3: Calculate the complexity centroid of the attention map;
[0010] Step 4: Construct a fixation model;
[0011] Step 5: Construct a saccade model beyond the edge;
[0012] Step 6: A saccade model restricted by the edge;
[0013] Step 7: Calculate the attention of the eye movement model.
[0014] Step 8: Use the trained model to test and analyze the test set data.
[0015] In the above Step 1, the training set samples are composed of labeled source domain samples and a part of unlabeled target domain samples, and the test set samples are composed of unlabeled target domain samples. In Step 1 of the present invention, the training set samples and the test set samples are divided. Steps 2 to 7 are to construct a corresponding fovea-peripheral vision bionic attention calculation model with eye movement attention for the training set, and Step 8 is to perform semantic segmentation of high-resolution remote sensing images on the test set samples.
[0016] In the above Step 2, when constructing the fovea-peripheral vision bionic attention calculation model, first input the original image, perform layer normalization operation on the feature map, and then perform self-attention calculation and eye movement attention calculation on the obtained feature map. The attention maps obtained from the self-attention calculation and the eye movement attention calculation are respectively fused, and after layer normalization, they are input into a multi-layer perceptron, and then the initial original feature map and the current attention map are fused in terms of features. Repeat the above operation 12 times.
[0017] In the above Step 3, in the process of calculating the complexity centroid of the attention map, first take the weights of the network one layer before the classification layer for the attention image, and after global average pooling, obtain the category mapping feature map. Take the category with the highest positive excitation as the label with the most pixels, that is, take the pixel category with the largest sum of positive weights as the label with the most pixels. Take the center position of the pixels corresponding to this label as the complexity centroid of the attention map.
[0018] In the above Step 4, when constructing the fixation model, fix the visual center position of the attention map in the exact middle of the image, and use this center as the center point of the subsequent feature maps to be made. Assume that the side length of the feature map is 8*L. Taking the visual center as the attention center point, divide the attention map with side lengths of 2*L, 4*L, and 6*L to obtain 4 attention maps P1, P2, P3, and P4 with side lengths of 2*L, 4*L, 6*L, and 8*L respectively.
[0019] In step five, for the saccade model beyond the edge, the center of gravity of the complexity of the attention map is used as the visual center. Assuming the side length of the feature map is 8*L, with the visual center as the attention center point, the attention map is divided with side lengths of 2*L, 4*L, and 6*L to obtain four attention maps P1, P2, P3, and P4 with side lengths of 2*L, 4*L, 6*L, and 8*L respectively. For the pixel points of the attention map that exceed the initial attention map, zero-padding operations are performed.
[0020] In step six, for the saccade model restricted by the edge, that is, the center of gravity of the complexity of the attention map is used as the visual center. Assuming the side length of the feature map is 8*L, with the visual center as the attention center point, the attention map is divided with side lengths of 2*L, 4*L, and 6*L to obtain four attention maps P1, P2, P3, and P4 with side lengths of 2*L, 4*L, 6*L, and 8*L respectively. If, when making the attention map with the visual center as the center point, it exceeds the edge, then the attention map is restricted by the edge and cannot exceed the edge.
[0021] In step seven, calculate the attention of the eye movement model, that is, for the four attention maps P1, P2, P3, and P4 obtained from the above fixation model, saccade model beyond the edge, and saccade model beyond the edge, the final attention is obtained and updated through the following formula.
[0022] P2 a1 = P2 a1 + e * P1
[0023] P3 a2 = P3 a2 + e * P2
[0024] P4 a3 = P4 a3 + e * P3
[0025] Where a1 is the area where the pixel points of P2 coincide with those of P1, a2 is the area where the pixel points of P3 coincide with those of P2, a3 is the area where the pixel points of P4 coincide with those of P3, e is the eye movement coefficient, and its value range is [0, 1], representing the influence of the pixel points in the central area on the edge pixels.
[0026] Step eight, use the trained model to test and analyze the test set data, that is, use the trained fovea-peripheral vision bionic attention calculation model with an eye movement model to perform semantic segmentation on the high-resolution remote sensing images in the test set.
[0027] The beneficial effects of the present invention are as follows:
[0028] (1) The present invention addresses the problems such as inaccurate and unreliable results in cross - domain semantic segmentation of high - resolution remote sensing images, and provides an accurate and reliable method for cross - domain semantic segmentation of high - resolution remote sensing images that simulates human vision;
[0029] (2) The present invention is applicable to cross - domain semantic segmentation of high - resolution images in the field of remote sensing, and has good operability and practicability;
[0030] (3) The present invention addresses the problems such as the lack of visual perception ability in current vision transformers, and provides three different eye movement models: a fixation model, a saccade model beyond the edge, and a saccade model restricted by the edge;
[0031] (4) The existing research on semantic segmentation of remote sensing images focuses on pixel - level semantic segmentation of high - resolution images. The present invention tries to distinguish different classifications of different - category pixels as much as possible. In reality, samples from different domains cannot be accurately classified at the pixel level using traditional methods, which is highly related to differences in imaging machines, locations, and times. Our method focuses on pixel differences between different domains, extracts domain - similar features as much as possible, and achieves accurate segmentation of samples from different domains. This method is applicable not only to image segmentation in the field of remote sensing, but also to the inspection and segmentation of lesions in medical images of different patients.
[0032] (5) The method of the present invention is based on the remote sensing big data environment, obtains more accurate pixel - level category segmentation, and can be widely applied to semantic segmentation of images with various different resolutions. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is the flowchart for constructing the fovea - peripheral vision bionic attention calculation model of the present invention;
[0034] Figure 2 is the flowchart for calculating the attention center of gravity of the present invention;
[0035] Figure 3 is the schematic diagram of the fixation model of the present invention;
[0036] Figure 4 is the schematic diagram of the saccade model beyond the edge of the present invention;
[0037] Figure 5 is the schematic diagram of the saccade model restricted by the edge of the present invention;
[0038] Figure 6 is the schematic diagram of the attention update of the eye movement model of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0039] As an attention calculation model, the vision transformer lacks visual perception ability. High-resolution remote sensing images have redundant and complex feature information, and it is difficult for traditional vision transformers to accurately extract the important features of high-resolution remote sensing images. To address the above problems, the present invention provides an accurate and reliable cross-domain semantic segmentation method for high-resolution remote sensing images based on a bionic eye movement pattern and with visual perception ability.
[0040] The cross-domain semantic segmentation method for remote sensing images based on a bionic eye movement pattern provided by the present invention comprises the following steps:
[0041] Step 1: Use two remote sensing image datasets as the source domain and the target domain respectively. After cropping, training set samples and test set samples are obtained.
[0042] Step 2: Construct a fovea-peripheral vision bionic attention calculation model;
[0043] Step 3: Calculate the complexity centroid of the attention map;
[0044] Step 4: Construct a fixation model;
[0045] Step 5: Construct a saccade model beyond the edge;
[0046] Step 6: A saccade model restricted by the edge;
[0047] Step 7: Calculate the attention of the eye movement model.
[0048] Step 8: Use the trained model to test and analyze the test set data.
[0049] In the above Step 1, the training set samples are composed of labeled source domain samples and a part of unlabeled target domain samples, and the test set samples are composed of unlabeled target domain samples, where the training set and the test set samples do not overlap.
[0050] In Step 2, to construct the fovea-peripheral vision bionic attention calculation model, as Figure 1 shown, Figure 1 is the flowchart for constructing the fovea-peripheral vision bionic attention calculation model. First, the original image is input, and layer normalization is performed on the feature map. Then, self-attention calculation and eye movement attention calculation are performed on the obtained feature map. The attention maps obtained from self-attention calculation and eye movement attention calculation are respectively fused, and after layer normalization, they are input into a multi-layer perceptron. Then, the initial original feature map and the current attention map are fused in terms of features. The above operations are repeated 12 times.
[0051] In Step 3, the process of calculating the complexity centroid of the attention map is as Figure 2 shown, Figure 2It is a flowchart for calculating the center of gravity of attention. First, take the weights of the network layer before the classification layer for the attention image. After global average pooling, obtain the class mapping feature map. Take the class with the highest positive excitation as the label with the most pixels, that is, take the pixel class with the largest sum of positive weights as the label with the most pixels. Take the center position of the pixels corresponding to this label as the complexity center of gravity of this attention map.
[0052] In step four, construct a fixation model, as Figure 3 shown. Figure 3 It is a schematic diagram of the fixation model. Fix the visual center position of the attention map in the exact middle of the image, and use this center as the center point of the subsequent feature map to be made. Assume the side length of the feature map is 8*L. With the visual center as the attention center point, divide the attention map with side lengths of 2*L, 4*L, and 6*L to obtain four attention maps P1, P2, P3, and P4 with side lengths of 2*L, 4*L, 6*L, and 8*L respectively.
[0053] In step five, construct a saccade model beyond the edge, as Figure 4 shown. Figure 4 It is a schematic diagram of the saccade model beyond the edge. Use the complexity center of gravity of the attention map as the visual center. Assume the side length of the feature map is 8*L. With the visual center as the attention center point, divide the attention map with side lengths of 2*L, 4*L, and 6*L to obtain four attention maps P1, P2, P3, and P4 with side lengths of 2*L, 4*L, 6*L, and 8*L respectively. Perform a zero-padding operation on the pixel points of the attention map that exceed the initial attention map.
[0054] In the said step six, a saccade model restricted by the edge, as Figure 5 shown. Figure 5 It is a schematic diagram of the saccade model restricted by the edge. Use the complexity center of gravity of the attention map as the visual center. Assume the side length of the feature map is 8*L. With the visual center as the attention center point, divide the attention map with side lengths of 2*L, 4*L, and 6*L to obtain four attention maps P1, P2, P3, and P4 with side lengths of 2*L, 4*L, 6*L, and 8*L respectively. If when making an attention map with the visual center as the center point, it will exceed the edge, then this attention map is restricted by the edge and cannot exceed the edge.
[0055] In the said step seven, calculate the attention of the eye movement model, as Figure 6 shown. Figure 6 It is a flowchart for updating the attention of the eye movement model. For the four attention maps P1, P2, P3, and P4 obtained from the above fixation model, saccade model beyond the edge, and saccade model restricted by the edge, obtain and update the final attention through the following formula.
[0056] P2 a1 = P2 a1 + e * P1
[0057] P3 a2 = P3 a2 + e * P2
[0058] P4 a3 = P4 a3 + e * P3
[0059] Where a1 is the area in P2 that coincides with the pixel position of P1, a2 is the area in P3 that coincides with the pixel position of P2, a3 is the area in P4 that coincides with the pixel position of P3, e is the eye movement coefficient, and its value range is [0, 1], representing the influence of the pixels in the central region on the edge pixels.
[0060] Step 8: Use the trained model to test and analyze the test set data, that is, use the trained fovea-peripheral vision bionic attention calculation model with an eye movement model to perform semantic segmentation on the high-resolution remote sensing images in the test set.
Claims
1. A cross-domain semantic segmentation method for remote sensing images based on bionic eye movement patterns, characterized in that It includes the following steps: First: Take two remote sensing image datasets as the source domain and the target domain respectively. After cropping, training set samples and test set samples are obtained. The training set samples consist of labeled source domain samples (Source Data, abbreviated as SD) and a part of unlabeled target domain samples (Target Data1, abbreviated as TD1), and the test set samples consist of unlabeled target domain samples (Target Data2, abbreviated as TD2). The samples in TD1 and TD2 do not overlap. Then: Construct an FPAM model with the training set samples and perform semantic segmentation on the TD2 samples in the test set.
2. The high-resolution remote sensing image cross-domain semantic segmentation method based on the bionic eye movement pattern according to claim 1, wherein: The construction of the FPAM model includes the following steps: Construct an eye movement attention model with a fovea-peripheral structure; Among them: Constructing an eye movement attention model with a fovea-peripheral structure includes: Calculate the complexity center of gravity of the attention map and use it as the attention center; Then introduce 3 different eye movement models to calculate the attention of the feature map.
3. The high-resolution remote sensing image cross-domain semantic segmentation method based on the bionic eye movement pattern according to claim 2, characterized in that: The calculation of the complexity center of gravity of the attention map includes: The center of gravity of the region with the most pixels of the same label is used as the center of gravity of the feature map.
4. The high-resolution remote sensing image cross-domain semantic segmentation method based on the bionic eye movement pattern according to claim 2, characterized in that: The 3 different eye movement models include: a fixation model, a saccade model beyond the edge, and a saccade model restricted by the edge.
5. The high-resolution remote sensing image cross-domain semantic segmentation method based on the bionic eye movement pattern according to claim 3, characterized in that: The center of gravity of the region with the most pixels of the same label is: calculated through an affinity network to obtain a category mapping feature map, and the category with the highest positive excitation is taken as the label with the most pixels, that is, the pixel category with the largest sum of positive weights is taken as the label with the most pixels.
6. The method for cross-domain semantic segmentation of high-resolution remote sensing images based on bionic eye movement patterns according to claim 3, wherein: Calculating the attention of the feature map is: Using the self-attention module of the vision transformer to calculate the feature map attention respectively.
7. The high-resolution remote sensing image cross-domain semantic segmentation method based on the bionic eye movement pattern according to claim 4, characterized in that: The fixation model is: Fix the center of gravity position, take the spatial center of gravity of the feature map, that is, the center point, as the visual center. Assume the side length of the feature map is 8*L. Take the visual center as the attention center point and divide the attention map with side lengths of 2*L, 4*L, 6*L to obtain 4 attention maps with side lengths of 2*L, 4*L, 6*L, 8*L respectively.
8. The method for cross-domain semantic segmentation of high-resolution remote sensing images based on bionic eye movement patterns according to claim 4, characterized in that: The saccade model beyond the edge is: Take the complexity center of gravity of the attention map as the visual center. Assume the side length of the feature map is 8*L. Take the visual center as the attention center point and divide the attention map with side lengths of 2*L, 4*L, 6*L to obtain 4 attention maps with side lengths of 2*L, 4*L, 6*L, 8*L respectively.
9. The high-resolution remote sensing image cross-domain semantic segmentation method based on the bionic eye movement pattern according to claim 4, characterized in that: The saccade model restricted by the edge is: Take the complexity center of gravity of the attention map as the visual center. Assume the side length of the feature map is 8*L. Take the visual center as the attention center point and divide the attention map with side lengths of 2*L, 4*L, 6*L to obtain 4 attention maps with side lengths of 2*L, 4*L, 6*L, 8*L respectively. If making an attention map with the visual center as the center point will exceed the edge, then this attention map is restricted by the edge and cannot exceed the edge.
10. The high-resolution remote sensing image cross-domain semantic segmentation method based on the bionic eye movement pattern according to claim 6, wherein: The calculation of the feature map attention is as follows: for each attention map, calculate its query matrix, key matrix, and value matrix, which have the same dimension. After multiplying them, attention maps of 2*L, 4*L, 6*L, and 8*L will be obtained. Then, multiply each pixel of the attention map by an eye movement coefficient and finally sum them up to obtain the final attention map of the feature map.
11. The method for cross-domain semantic segmentation of high-resolution remote sensing images based on bionic eye movement patterns according to any one of claims 1-10, characterized in that, The described remote sensing image samples are replaced with other remote sensing image samples.