Remote sensing image segmentation method and system based on modal balance knowledge distillation framework
Through the method based on the modal balance knowledge distillation framework, the image-level and pixel-level virtual modal generation technology is used to realize modal unbiased knowledge transfer between optical and SAR images, solving the problem of poor transfer learning effect in the existing technology, and improving the general semantic segmentation ability of the model.
Patent Information
- Application Number
- CN202510132266.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-06
AI Technical Summary
The prior art is difficult to realize effective transfer learning between optical and SAR images, resulting in a significant decline in the performance of the model on unlabeled SAR images and it is difficult to obtain a general semantic segmentation model with equalization performance for different modal images.
Using a modal equilibrium knowledge distillation framework, the virtual samples with modal symmetrical optical images and SAR images are obtained through image-level virtual mode generation and pixel-level virtual mode generation to achieve modal unbiased knowledge distillation. The framework includes dual-modal supervised learning, hybrid modal supervised learning and dual-modal knowledge inference, and transforms cross-modal transfer learning tasks into homomodal knowledge transfer tasks.
The unbiased knowledge distillation of optical and SAR images is realized, which improves the model's segmentation ability of images on different modes and ensures the model's equalization performance on optical and SAR images.
Smart Images

Figure CN120014277A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a remote sensing image segmentation method and system based on a modal balance knowledge distillation framework. Background Art
[0002] Land cover / land use (LULC) is basic geographic information, making important contributions to climate change research, ecological environment monitoring and land resource management. Remote sensing images are the main carrier of earth surface information. In recent years, with the continuous advancement of remote sensing imaging technology, in addition to a large number of optical satellites, synthetic aperture radar (SAR) satellites have also flourished, such as Germany's TerraSAR-X, Japan's ALOS-2, ESA's Sentinel-1 and China's GaoFen-3 satellite. SAR is an active remote sensing and has the advantage of not being affected by clouds and fog. When the optical sensor fails, SAR can still continue to record surface information. The comprehensive use of optical and SAR satellite images for semantic segmentation has become an important way to achieve all-day and all-weather LULC mapping.
[0003] Currently, there are a large number of optical image LULC classification datasets in the field, such as ISPRS, AID, GID, etc. Due to the limitations of different imaging principles, the representation of objects in SAR images is quite different from that in optics, and the annotation of objects in SAR images requires a high degree of professionalism. Therefore, the current LULC classification datasets for SAR images are far less than those for optical images, such as WHU-OPT-SAR and SEN12MS, which makes it difficult to support the growing LULC classification research and application of SAR images. The completely different imaging mechanisms of optical and SAR images lead to huge differences in their object characteristics. When the CNN-based classification model is trained using optical images with existing labels (source domain), its performance on unlabeled SAR images (target domain) will drop significantly, such as Figure 1 As shown in the figure. To address this problem, it is expected that a certain transfer learning method can be designed to break the modality barrier and obtain a universal semantic segmentation model with similar segmentation capabilities for optical and SAR images by training the model using only optical images. This universality means that the model achieves similar and balanced performance for images of different modalities.
[0004] Optical and SAR image features are quite different. Currently, the cross-modal segmentation capability of semantic segmentation models is mainly achieved through transfer learning. The transfer learning semantic segmentation method mainly achieves the model transfer effect by reducing the distribution difference between the source domain and the target domain. Among them, the more classic methods include domain adaptation and knowledge distillation. However, the models obtained by the above methods may be biased towards the source domain or the target domain, and the models often do not have the common segmentation capabilities for the source domain and the target domain. Overall, there are still the following difficulties and challenges in using transfer learning technology to obtain common semantic segmentation capabilities for optical and SAR images: (1) Controlling the migration direction: The same type of ground objects have very different representations in optical and SAR images. How to drive the semantic segmentation model trained on optical images to migrate in a direction that is effective for SAR images is a problem that needs to be overcome. (2) Controlling migration balance: In order to obtain universal segmentation capabilities for optical and SAR images, it is necessary to accurately control the migration direction of the model to achieve a balance. How to ensure that the model does not bias towards either the source domain or the target domain is a problem that needs to be overcome. Summary of the invention
[0005] The present invention provides a remote sensing image segmentation method and system based on a modal balance knowledge distillation framework, so as to solve the defects in the prior art in the transfer learning between optical images and SAR images.
[0006] In a first aspect, the present invention provides a remote sensing image segmentation method based on a modality balance knowledge distillation framework, comprising: Acquire source domain optical images and target domain SAR images; Performing image-level virtual modality generation on the source domain optical image and the target domain SAR image to obtain a virtual SAR image and a virtual optical image; Inputting the source domain optical image and the virtual SAR image into dual-modal supervised learning to obtain a source domain optical image pseudo label and a virtual SAR image pseudo label; Inputting the target domain SAR image and the virtual optical image into bimodal knowledge reasoning to obtain a target domain SAR image pseudo label and a virtual optical image pseudo label; Based on the source domain real label, the target domain SAR image pseudo label, the source domain optical image and the target domain SAR image, pixel-level virtual modality generation is performed taking into account the proportion of ground object categories to obtain a mixed image and a mixed pseudo label; Based on the virtual SAR image pseudo-label, the virtual optical image pseudo-label, the virtual SAR image and the virtual optical image, pixel-level virtual modality generation is performed taking into account the proportion of ground object categories to obtain a mixed virtual image and a mixed virtual pseudo-label; The mixed image and the mixed virtual image are input into mixed modality supervised learning to obtain a remote sensing image segmentation result.
[0007] According to a remote sensing image segmentation method based on a modality balance knowledge distillation framework provided by the present invention, image-level virtual modality generation is performed on the source domain optical image and the target domain SAR image to obtain a virtual SAR image and a virtual optical image, including: The source domain optical image and the target domain SAR image are processed respectively using a pix2pixHD-based style transfer method to generate the virtual SAR image and the virtual optical image.
[0008] According to a remote sensing image segmentation method based on a modality balance knowledge distillation framework provided by the present invention, the source domain optical image and the virtual SAR image are input into dual-modality supervised learning to obtain a source domain optical image pseudo label and a virtual SAR image pseudo label, including: Determining a supervised encoder and a supervised decoder for the bimodal supervised learning; The source domain optical image is sequentially subjected to multi-scale fusion by the supervised encoder and the supervised decoder to obtain a pseudo label of the source domain optical image; The virtual SAR image is sequentially passed through the encoder and the decoder for multi-scale fusion to obtain a pseudo label of the virtual SAR image; Determine that the loss of the source domain optical image pseudo label is a first loss function, the loss of the virtual SAR image pseudo label is a second loss function, and the first loss function and the second loss function constitute a dual-modal supervised learning loss function; Among them, the first loss function includes the cross entropy loss between the source domain true label and the virtual SAR image pseudo label, and the softened Dice loss between the source domain true label and the virtual SAR image pseudo label, and the second loss function includes the cross entropy loss between the source domain true label and the virtual SAR image pseudo label, and the softened Dice loss between the source domain true label and the virtual SAR image pseudo label.
[0009] According to a remote sensing image segmentation method based on a modal balance knowledge distillation framework provided by the present invention, the target domain SAR image and the virtual optical image are input into dual-modal knowledge reasoning to obtain a target domain SAR image pseudo label and a virtual optical image pseudo label, including: Determining an inference encoder and an inference decoder for the bimodal knowledge reasoning; The knowledge acquired in the bimodal supervised learning phase is transferred using exponential moving average weights; The target domain SAR image is sequentially subjected to multi-scale fusion by the inference encoder and the inference decoder to obtain a pseudo label of the target domain SAR image; The virtual optical image is sequentially passed through the inference encoder and the inference decoder for multi-scale fusion to obtain the virtual optical image pseudo label.
[0010] According to a remote sensing image segmentation method based on a modality balance knowledge distillation framework provided by the present invention, pixel-level virtual modality generation taking into account the proportion of ground object categories is performed based on the real label of the source domain, the pseudo label of the target domain SAR image, the source domain optical image and the target domain SAR image to obtain a mixed image and a mixed pseudo label, including: Determine the background and foreground in the image based on the true label of the source domain and the proportion of preset ground object categories; Mapping the background pixel value in the source domain true label to 0 and the foreground pixel value to 1 to generate a first mask; Mapping background pixel values in the target domain SAR image pseudo-label to 1 and foreground pixel values to 0 to generate a second mask; Performing a dot product of the first mask and the second mask to obtain a final mask; Based on the final mask, the source domain optical image and the target domain SAR image, the mixed image is calculated; The mixed pseudo label is calculated based on the final mask, the source domain true label and the target domain SAR image pseudo label.
[0011] According to a remote sensing image segmentation method based on a modality balance knowledge distillation framework provided by the present invention, pixel-level virtual modality generation taking into account the proportion of ground object categories is performed based on the virtual SAR image pseudo-label, the virtual optical image pseudo-label, the virtual SAR image and the virtual optical image to obtain a mixed virtual image and a mixed virtual pseudo-label, including: Based on the final mask, the virtual SAR image and the virtual optical image, the mixed virtual image is calculated; The mixed virtual pseudo label is calculated based on the final mask, the virtual SAR image pseudo label and the virtual optical image pseudo label.
[0012] According to a remote sensing image segmentation method based on a modality balance knowledge distillation framework provided by the present invention, the mixed image and the mixed virtual image are input into mixed modality supervised learning to obtain a remote sensing image segmentation result, including: Determining that the hybrid modality supervised learning adopts the supervised encoder and the supervised decoder in the bimodal supervised learning; The mixed image is sequentially passed through the supervised encoder and the supervised decoder for multi-scale fusion to obtain a mixed pseudo-label segmentation result; The mixed virtual image is sequentially passed through the supervised encoder and the supervised decoder for multi-scale fusion to obtain a mixed virtual pseudo-label segmentation result; The mixed pseudo-label segmentation result is supervised by the mixed pseudo-label to construct a third loss function, wherein the third loss function includes a cross entropy loss between the mixed pseudo-label and the mixed pseudo-label segmentation result, and a softened Dice loss between the mixed pseudo-label and the mixed pseudo-label segmentation result; The mixed virtual pseudo-label segmentation result is supervised by the mixed virtual pseudo-label to construct a fourth loss function, wherein the fourth loss function includes a cross entropy loss between the mixed virtual pseudo-label and the mixed virtual pseudo-label segmentation result, and a softened Dice loss between the mixed virtual pseudo-label and the mixed virtual pseudo-label segmentation result.
[0013] In a second aspect, the present invention further provides a remote sensing image segmentation system based on a modality balance knowledge distillation framework, comprising: An acquisition module is used to acquire source domain optical images and target domain SAR images; An image-level generation module, used for performing image-level virtual modality generation on the source domain optical image and the target domain SAR image to obtain a virtual SAR image and a virtual optical image; A dual-modal supervised learning module, used for inputting the source domain optical image and the virtual SAR image into dual-modal supervised learning to obtain a source domain optical image pseudo label and a virtual SAR image pseudo label; A bimodal knowledge reasoning module, used for inputting the target domain SAR image and the virtual optical image into bimodal knowledge reasoning to obtain a target domain SAR image pseudo label and a virtual optical image pseudo label; A first pixel-level generation module is used to generate a pixel-level virtual modality taking into account the proportion of ground object categories based on the source domain real label, the target domain SAR image pseudo label, the source domain optical image and the target domain SAR image to obtain a mixed image and a mixed pseudo label; A second pixel-level generation module is used to generate a pixel-level virtual modality taking into account the proportion of ground object categories based on the virtual SAR image pseudo-label, the virtual optical image pseudo-label, the virtual SAR image and the virtual optical image to obtain a mixed virtual image and a mixed virtual pseudo-label; The mixed modality supervised learning module is used to input the mixed image and the mixed virtual image into the mixed modality supervised learning to obtain the remote sensing image segmentation result.
[0014] In a third aspect, the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the remote sensing image segmentation method based on the modal balance knowledge distillation framework as described above is implemented.
[0015] In a fourth aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a remote sensing image segmentation method based on a modal balance knowledge distillation framework as described in any one of the above.
[0016] The remote sensing image segmentation method and system based on the modal balanced knowledge distillation framework provided by the present invention obtain virtual samples with modal symmetry of optical images and SAR images by adopting image-level virtual modality generation strategy and pixel-level virtual modality generation strategy, so as to provide support for modality-unbiased knowledge distillation. Through the modality-balanced optical and SAR image knowledge distillation framework, a cross-modal transfer learning task is converted into two same-modal transfer learning tasks through bimodal supervised learning, mixed-modal supervised learning and bimodal knowledge reasoning, so as to realize the unbiased knowledge distillation of optical and SAR image modalities. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is a comparison chart of the prediction results of the optical image training model provided by the prior art on the optical image and the prediction results on the SAR image; Figure 2 It is a comparison between the knowledge extraction in the traditional mode provided by the present invention and the knowledge extraction in the present invention; Figure 3 It is a flow chart of a remote sensing image segmentation method based on a modal balance knowledge distillation framework provided by the present invention; Figure 4 It is a framework diagram of modal balance knowledge distillation provided by the present invention; Figure 5 is a schematic diagram of a pixel-level virtual modality generation process provided by the present invention; Figure 6 is a pixel-level virtual modality generation result diagram provided by the present invention; Figure 7 It is a visualization result diagram of different methods provided by the present invention on the SAR image test set; Figure 8 It is a qualitative evaluation result diagram of different methods provided by the present invention; Fig. 9 It is a large-scale visualization map of the entire area of a county in region B provided by the present invention. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] In view of the problems existing in the prior art, this paper proposes a modality balance knowledge distillation framework consisting of three parts: bimodal supervised learning, mixed modality virtual learning, and bimodal knowledge reasoning, for the cross-modal segmentation task of optical and SAR images. This framework innovatively transforms the cross-modal problem into the same-modal knowledge transfer problem by establishing a virtual modality. Figure 2 The traditional model of domain adaptation and knowledge distillation shown in (a) is as follows: Figure 2 The modal balance knowledge distillation framework shown in (b) above, in which the source domain and target domain are transformed into virtual modal representations close to the target domain and source domain, respectively. This approach aims to eliminate the problem of misaligned migration direction caused by modal imbalance in cross-modal migration tasks, and achieve the ability to learn different modal balances in a more friendly way.
[0021] The main ones involved include knowledge distillation and domain adaptation. The knowledge distillation method uses the source domain with supervised information to generate pseudo labels for the target domain to iteratively optimize the segmentation model, thereby obtaining the ability to segment the target domain. This type of method generally designs the structure of the student-teacher network, mainly including typical structures such as offline distillation, online distillation and self-distillation: (1) Offline distillation: This method uses a pre-trained teacher model with fixed parameters to guide the student model. During the distillation process, the student network obtains fixed knowledge after each round of training. For example, models such as FSP, SSKD, and SemCKD. This model is simple to implement, but it usually transfers knowledge from the teacher to the student in a one-way direction. When the difference between the student and teacher networks is too large, it is difficult for the student network to learn useful knowledge.
[0022] (2) Online distillation: This network trains the student network and the teacher network together, updates the parameters at the same time, and the knowledge is continuously updated during the transmission process. The entire model is trained end-to-end. Classic networks include Rocket-KD, DCM, and ACNs. This is a method that can be highly efficient and parallel. However, the differences in structure and size will cause the model capacity between the teacher network and the student network to be different, which makes it impossible to effectively transfer knowledge to the student network.
[0023] (3) Self-distillation: The student and teacher networks of self-distillation use the same structure for distillation, which mainly improves the capacity mismatch problem between the student and teacher networks caused by different structural capacities. It can be regarded as a special form of online distillation. The pseudo labels generated by general distillation models have more noise interference. Therefore, researchers proposed the Mean Teacher (MT) model to improve the accuracy of the inference stage. Methods such as SePiCo, Dual-Teacher++, DACS and DAFormer all introduced the MT model. In the field of remote sensing, Luo et al. proposed a two-stage domain adaptive cross-temporal classification method to achieve transfer learning between optical images of different time and space. Wang et al. proposed a cross-sensor land cover framework to migrate between aerial images and satellite images to solve the problems of inconsistent spatial resolution and spectral differences. However, these models are all based on optical images for migration, and the difference between the source domain and the target domain is much smaller than that of optical and SAR images. Therefore, in this scenario, the model is easier to migrate in a direction that is beneficial to the target domain. However, for multimodal images such as optical and SAR, the difficulty of the model controlling the migration direction is greatly increased.
[0024] Most of the above methods are aimed at application scenarios where the source domain and the target domain belong to the same modality. When these methods are applied to images of different modalities, the versatility of the model will be greatly reduced. It is not difficult to find that too much or too little guidance from the teacher network to the student network will cause the model to shift toward the target domain or the source domain, which is likely to cause distillation deviation. Therefore, the direction of knowledge distillation is unstable. The method of the present invention improves this distillation instability by designing a series of balancing strategies in view of the characteristics of excessive differences in multimodal images, thereby improving the versatility of the model.
[0025] The core of domain adaptation is to solve the impact of inconsistent data distribution on the performance of learning models. It reduces the domain offset between the source domain and the target domain so that they are in the same feature space as much as possible. It is usually applied to the actual scenario of using the model trained in the source domain to directly solve the target domain task. It is mainly divided into two categories: image level and feature level: (1) Image-level domain adaptation: It generally uses GAN to minimize the appearance difference at the image level to align the style distribution between the source domain and the target domain. This type of method is suitable for alleviating the differences in color, texture and lighting conditions at the image level, and is mainly suitable for the differences in the overall appearance of the image caused by different imaging conditions. Typical works in the field of computer vision include CycleGAN, AgGAN and pix2pixHD. Inspired by the field of vision, remote sensing image domain adaptation methods have also been gradually proposed. This type of method is to translate remote sensing images of different phases and different sensors into images of the same style to reduce domain shift. The advantage of this type of method is that it can learn complex data distributions and thus transform the underlying features of the image. For the scenario of optical and SAR image migration, it can narrow the distance between the underlying feature spaces of the two types of images and directly alleviate the feature gap caused by the imaging mode. However, its disadvantage is that the training process is sometimes unstable.
[0026] (2) Feature-level domain adaptation: It is to apply the knowledge of the source domain to the target domain by using the similarity between the source domain and the target domain at the feature level. This type of method is suitable for scenes with object-level differences, such as object posture and spatial distribution. Classic methods in computer vision include CyCADA, ADVENT, and LITR. In addition, researchers have also proposed a series of methods for remote sensing images. For example, ColorMapGANs, TriADA, joint MLP-GNN, SDA, and the method proposed by Ye et al. These methods are usually implemented using adversarial methods with discriminators. When the discriminator cannot distinguish the features of the source domain and the target domain, the model believes that the features of the two domains are aligned. However, adversarial methods often generate some specific non-representative features in order to deceive the discriminator. For multi-modal images such as optical and SAR images with essential underlying differences, the effect of simply forcibly narrowing the distance at the high-level feature level is very limited.
[0027] The above methods attempt to learn domain-invariant features between the source domain and the target domain from the image level and the feature level, and use this feature to apply the knowledge of the source domain to the target domain. However, the feature distribution in the global space is discrete and diffuse. Even if the discriminator achieves the optimal discrimination effect at the high-level feature level, the model is still prone to bring the features of different categories of the source domain and the target domain closer, resulting in negative transfer. At present, some scholars have proposed to combine the domain adaptation at the image level and the feature level. However, this method is only for single-channel optical images rather than RGB images, and its scope of application is relatively small. Therefore, the present invention introduces image-level domain adaptation to perform modality transfer between RGB images and SAR images at the image level.
[0028] Figure 3 is a flow chart of a remote sensing image segmentation method based on a modal balance knowledge distillation framework provided by an embodiment of the present invention, such as Figure 3 As shown, including: Step 100: Acquire source domain optical images and target domain SAR images; Step 200: performing image-level virtual modality generation on the source domain optical image and the target domain SAR image to obtain a virtual SAR image and a virtual optical image; Step 300: inputting the source domain optical image and the virtual SAR image into dual-modal supervised learning to obtain a source domain optical image pseudo label and a virtual SAR image pseudo label; Step 400: inputting the target domain SAR image and the virtual optical image into bimodal knowledge reasoning to obtain a target domain SAR image pseudo label and a virtual optical image pseudo label; Step 500: generating a pixel-level virtual modality taking into account the proportion of ground object categories based on the source domain real label, the target domain SAR image pseudo label, the source domain optical image and the target domain SAR image to obtain a mixed image and a mixed pseudo label; Step 600: generating a pixel-level virtual modality taking into account the proportion of ground object categories based on the virtual SAR image pseudo-label, the virtual optical image pseudo-label, the virtual SAR image and the virtual optical image to obtain a mixed virtual image and a mixed virtual pseudo-label; Step 700: input the mixed image and the mixed virtual image into mixed modality supervised learning to obtain a remote sensing image segmentation result.
[0029] Specifically, the modal balanced knowledge distillation framework proposed in the embodiment of the present invention introduces virtual modality generation so that the model trained only with optical images has universal segmentation capabilities for optical and SAR images. The student network of the framework is composed of bimodal supervised learning and mixed modality supervised learning, and the teacher network is composed of bimodal knowledge reasoning. It is implemented by three stages: In the first stage, balanced bimodal supervised learning is performed using real optical images and virtual SAR images in the source domain, so that the model obtains similar interpretation capabilities for the two modalities. In the second stage, the weights of the first stage are transferred to the bimodal knowledge reasoning module, and then it is used to obtain pseudo-labels for real SAR images and virtual optical images in the target domain. In the third stage, mixed modality supervised learning is performed using mixed images and mixed labels. This process is supervised and optimized by the pseudo-labels output by the second stage to ultimately achieve the purpose of balanced knowledge distillation. It is worth noting that the input of each stage is composed of two parallel modalities. The overall framework structure is as follows Figure 4 shown.
[0030] The symbol definitions are shown in Table 1: Table 1
[0031] The first stage is bimodal supervised learning.
[0032] Bimodal supervised learning is the first part of the student network in the balanced knowledge distillation framework. It aims to use source domain labels to guide the training of image-level modality balanced samples to establish the model's initial segmentation capability for balanced optical and SAR images.
[0033] (1) Image-level virtual modality generation Image-level modality differences may cause overall distribution shifts, weakening the possibility of comprehensive utilization of multimodal data in transfer learning. Inspired by domain adaptation, this paper designs image-level virtual modality generation (IVMG) to try to solve the modality imbalance problem in the knowledge transfer process from the image level. Source domain optical image and target domain SAR images The appearance of the two images is transferred to each other while retaining their respective content information to obtain a virtual SAR image. and virtual optical images , as shown in formula (1): (1) After the above operations, and Belong to the same mode, and Belong to the same modality. Therefore, IVMG is a prerequisite for converting the cross-modal transfer task into two same-modal transfer subtasks. The style transfer method used in the embodiment of the present invention is pix2pixHD. Compared with the classic CycleGAN method, this model can not only generate high-resolution images more stably, but also more fully capture and reproduce the details and textures of real images.
[0034] (2) Bimodal supervised learning like Figure 4 As shown in (a), and The input is into dual-modal supervised learning (DMSL), and the dual-modal features are extracted through the same encoder to obtain the ability to perceive the balance of different modalities.
[0035] Encoder The structure of is Mix transformer. It learns and The features of the decoder gradually master the ability to extract features from the source domain and the target domain. Perform multi-scale fusion to obtain segmentation results and , as shown in formula (2); (2) and By label Carry out supervision. The loss is . The loss is . By cross-entropy loss and soft dice loss The sum of is calculated as shown in formula (3): (3) in: (4) The second stage is bimodal knowledge reasoning.
[0036] The optimization of the student network requires the guidance of the teacher network. Therefore, the model also needs to obtain the knowledge of real SAR images in the target domain to guide the learning of the student network. To achieve this goal, this section uses real SAR images to generate virtual optical images, and then designs dual-modal knowledge inference (DMKI) to predict these two types of images in parallel to obtain pseudo labels as supervision information for the training of the student network.
[0037] like Figure 4 As shown in (c), DMKI has the same network structure as DMSL, but it does not update the gradient. It uses the exponential moving average (EMA) weight to transfer the knowledge obtained in the DMSL stage. Then it directly predicts the SAR image in the target domain. and virtual optical images Pseudo labels and , as shown in formula (5): (5) in, is the encoder of the DMKI stage, It is the decoder of DMKI stage.
[0038] It should be noted that since DMKI does not update the gradient, there is no need to design an additional loss function. The baseline of the entire network framework is the student-teacher structure of DAFormer.
[0039] The third stage is mixed modality supervised learning.
[0040] The pseudo labels generated by common knowledge distillation methods have certain errors. Using such pseudo labels as supervision information for the student network may cause the student network to be difficult to converge and generalize poorly. In this embodiment, a mixed modality supervised learning method is proposed. This method takes into account the proportion of object categories, mixes the images of the source domain and their corresponding labels with the images of the target domain and their corresponding pseudo labels at the pixel level, and then uses pseudo labels with partial true values for supervised learning to obtain a modality-unbiased knowledge distillation model.
[0041] (1) Generation of pixel-level virtual modes taking into account the proportion of ground object categories This embodiment proposes a pixel-level virtual modality generation method (PVMG), which mixes multimodal images at the pixel level based on the distribution characteristics of remote sensing objects. It has two goals: on the one hand, to achieve modality balance within an image; on the other hand, to increase the proportion of minority categories and reduce the impact of long-tail distribution.
[0042] By counting the true labels in the source domain, we calculate the proportion of pixels in each category to the total number of pixels, and define the category with a proportion of more than 10% as the background, and the category with a proportion of less than 10% as the foreground. Figure 5 As shown in the figure, since forest, cultivated land and water bodies account for more than 10% of the labels, these categories are defined as background, while urban, rural and road are defined as foreground. The background pixel value is mapped to 0, and the foreground pixel value is mapped to 1 to generate a binary mask . Pseudo labels The mask is generated by mapping in the opposite way (i.e., foreground pixel values are mapped to 0 and background pixel values are mapped to 1). . Get the final mask .
[0043] use Respectively , Image pairs and , The image pair is calculated to obtain the mixed image and the corresponding labels , as shown in formula (6). By observing It can be seen that the same scene has the characteristics of both optical and SAR images. (6) Similarly, , , and Do the same thing and get mixed results and .
[0044] like Figure 6 As shown in the figure, (a), (e), (h) and (l) are optical images, (b), (d), (i) and (k) are SAR images, and (c) and (g) are images generated by random mixing. We can see that the foreground objects are severely blocked or destroyed by the background objects, causing the relative relationship between the distribution of objects to deviate from the real scene. Figure 6 As shown in (j) and (m), PVMG takes into account the characteristics of remote sensing and reasonably mixes multi-modal images to avoid occlusion and damage of foreground objects.
[0045] (2) Virtual modality supervised learning After processing, the pseudo-labels are mixed Replace some wrong labels with the true values of the categories with a smaller proportion. On this basis, this embodiment designs virtual modal supervised learning, which is the second part of the student network in the balanced knowledge distillation framework. As the supervisory information of the student network, it can significantly improve the accuracy of the training process, especially beneficial for improving the segmentation accuracy of small categories that are more difficult to train.
[0046] Leveraging Encoders in Bimodal Supervised Learning Extract pixel-level virtual modality samples and The features of the model consolidate the model's ability to extract features that balance the two modalities. Then, the decoder in the bimodal supervised learning is used The segmentation result is obtained by decoding, as shown in formula (7): (7) Segmentation results and Depend on and Supervise separately to improve the accuracy of the model in learning the unlabeled target domain. The loss function is shown in formula (8): (8) Therefore, the loss function of the entire model is shown in formula (9): (9) In order to demonstrate the effectiveness and practicality of the modal balance knowledge distillation framework proposed in the present invention, this embodiment carried out application experiments on the large public dataset WHU-OPT-SAR and 6 regions. The results showed that the method of the present invention is significantly superior to other methods in terms of performance and stability.
[0047] First, the dataset is described. Dataset I is the WHU-OPT-SAR dataset, which is an open source optical and SAR image segmentation dataset published by a university. The optical image is acquired by the GaoFen-1 satellite and has red, green and blue bands. SAR is acquired by the GaoFen-3 satellite. The WHU-OPT-SAR dataset contains 100 optical images of 5556×3704 pixels and the same number and size of SAR images. The sampling resolution is 5m. The categories of annotated data include cultivated land, urban, rural, water, forest, road and others. These categories account for 35%, 5%, 6%, 14%, 38%, 1% and 1% respectively. We use dataset I to verify the effectiveness of the proposed method in the task of semantic segmentation of optical and SAR image transfer learning in the same region.
[0048] In this embodiment, the images contained in the WHU-OPT-SAR dataset are cropped to 512×512 pixels, and a total of 7000 patches are obtained, including 5652 training sets and 1348 test sets, as shown in Table 2. The same operation is performed on SAR images. Dataset I The optical image test set and the SAR image test set are prepared to test the migration performance of the model in the same area.
[0049] Table 2 Division of training set and test set
[0050] The symbol “*” denotes target domain image patches (whose labels are not used).
[0051] Dataset II contains 6 different regions, see Table 2. Optical images were collected by GaoFen-2 satellite with a resolution of 1m. SAR images were collected by GaoFen-3's fine strip II with a resolution of 10m. Annotation data comes from the third land cover survey. To ensure consistency, these images were resampled to 5m resolution. The dataset includes 5 categories, namely cultivated land, buildings, water bodies, forests, and roads. The proportions of these categories in order are shown in the last column of Table 2. It can be observed that cultivated land and water bodies account for a large proportion in region A, while forests account for a large proportion in region B. The proportion of categories in region C in the test city is relatively balanced. However, the proportion of forests in regions D and E is very large, and the categories are very unbalanced. The proportion of cities and roads in region E is small, and the source domain is far away from the target domain and the test city. The source domain is located in a large river basin, and the target domain is located in another river basin. The test city is different from these two regions and is widely distributed. The topography of these regions varies greatly, and the proportion of categories varies, which makes the classification tasks in different regions much more challenging than dataset I, and puts higher requirements on the transferability, generalization, and robustness of the model.
[0052] Table 3. Areas covered by optical and SAR images in different regions
[0053] The training and test sets of optical images in Dataset II are from region A. The training and test sets of SAR images are from region B and the other four regions in Table 3. The optical image test set of region A and the SAR image test set of region B in Dataset II are prepared to test the migration performance of the model in different regions. Regions C, D, E, and F in Dataset II are prepared to test the generalization of the model. According to the same cropping method, the data set division of Dataset II is shown in Table 4.
[0054] Table 4 Division of training set and test set
[0055] The symbol “*” denotes target domain image patches (whose labels are not used).
[0056] In order to evaluate the performance of the BL proposed in this paper, eight common DASS methods are selected, which are mainly divided into two categories: (1) Domain Adaptation Method (DA). Use classic image-based methods such as CycleGAN, AgGAN, and Pix2PixHD to transfer optical images to the style of SAR images. Use it as a training set to train Deeplabv2 and directly predict SAR images. In addition, there are also feature-based methods such as ADVENT, LTIR, and ADVENT. Use them to transfer the source domain and the target domain. Unlike the other two methods, the input of LTIR also includes the source domain of the target domain style.
[0057] (2) Knowledge distillation methods. Including DACS, DAFormer and DACST. In the cross-domain semantic segmentation task, DACS introduced Cutmix for the first time. DAFormer used transformer as the network skeleton for the first time. DACST is a method in the remote sensing field, which uses the source domain with the style of the target domain instead of the original source domain to influence the input to the network.
[0058] To ensure fairness, the backbone of the above reference method uses Deeplabv2. The method of the present invention uses Deeplabv2 as the backbone and is called BL-D. In other cases, the method of the present invention uses MiT-B5 as the pre-training model.
[0059] The experimental environment is a CentOS 7.5 Linux platform in a university supercomputing center. The model was trained using the Adam optimizer on an Nvidia Tesla V100. The hyperparameters are set as follows: batch size is 4, number of iters is 60000, initial learning rate is 6×10-5, and the optimization method is stochastic gradient descent (SGD) solver with momentum 0.9 and weight decay 5× 10-4. When the error rate stops decreasing, we divide the learning rate by 10 and use the new value to update the parameters.
[0060] The migration performance of the model for optical and SAR images in the same area was tested using Dataset I, and the results are shown in Table 5. Due to the huge difference in features between optical and SAR images, the model trained by the No adaptation method is invalid in the target domain. Domain adaptation and knowledge distillation methods have obvious improvement effects compared to the No adaptation method. For example, the F1Score of Pix2PixHD is 49% higher than that of No adaptation. However, the accuracy of ADVENT, DACS and DAFormer methods is lower than that of the image-level domain adaptation method. This is because the imaging principles of optical and SAR images are completely different, and it is impossible to completely rely on implicit learning to close the features between them. LTIR, DACST and BL all combine the image level and the feature level to close the modality. Among them, the method of the present invention has achieved the highest accuracy. The F1 Score is 7.40% higher than that of LTIR. OA is 7.21% higher than that of DACST. The LTIR method uses the idea of confrontation, which is easy to cause mode collapse. DACST only uses the source domain after style conversion as input, resulting in the model being completely unable to learn the modal invariant features contained in the original source domain. The idea of modal balance of BL enables the model to learn dual-modal information in a balanced manner to prevent information bias. In addition, compared with the Deeplabv2 structure, the method of the present invention uses the MiT_B5 structure as the Backbone F1 Score to improve by 7%. This proves that the transformer structure is helpful for learning large-scale remote sensing image context information and meets the needs of remote sensing object interpretation.
[0061] Table 5
[0062] Figure 7 The visualization results of different methods on the SAR image test set are shown. Figure 7As shown in (d), the optically trained model is ineffective on SAR images. The results of the domain adaptation method have serious road omissions and false detections ( Figure 7 (e)-(h)). This is because a lot of detail information will inevitably be lost during the domain adaptation process, making it difficult for the model to extract small objects. Figure 7 (i) and (k) are knowledge distillation methods. It can be observed that the results of these methods produce a large number of false positives. Due to the huge differences in the characteristics of optical and SAR images, these methods cannot achieve good migration effects only through knowledge transfer at the feature level. Figure 7 (l) and (m) are the methods of the present invention. The backbone of the Transformer structure makes the segmentation result in (m) richer in details and higher in road continuity than that in (l). The model of the present invention even predicts small roads that are not marked in the label. Compared with the reference method, the method of the present invention achieves the best visual effect with the help of the modal balance knowledge distillation framework and the virtual modal generation strategy.
[0063] Dataset II is used to test the model's migration performance on optical and SAR images in different regions. The optical images of region A are the source domain, and the SAR images of region B are the target domain. The results are shown in Table 6. Unlike the same region, due to the inconsistency between the source and target domains, the image-level domain adaptation method will produce negative migration, resulting in unsatisfactory segmentation accuracy. For the knowledge distillation method, the completely different style and content distribution between the source and target domains will make knowledge transfer more difficult. Such methods also perform poorly. Our method uses a hierarchical hybrid strategy to generate mixed data with consistent content and different modalities, forcing the model to learn multimodal features in a balanced manner, thereby achieving satisfactory accuracy. The F1 Score of BL is 17%, 19%, and 10% higher than that of LTIR, DACST, and BL-D, respectively. This further illustrates the effectiveness of our proposed method.
[0064] Table 6
[0065] Qualitative evaluation such as Figure 8 As shown in the figure, there is no correspondence between the optical and SAR image pixels in different regions, which leads to serious misclassification in most domain adaptation and knowledge distillation methods. The LTIR method improves the visual effect to a certain extent. The method of the present invention achieves the best visual effect, with less noise at the boundary of the object and smoother edges. In addition, the large-scale visualization of the entire area of a county in region B is shown in the figure. Fig. 9As shown. Compared with the best LTIR method in Table 5, the overall classification effect of the method of the present invention is more accurate and closer to the real label. From the enlarged image, the road edge extracted by the method of the present invention is more refined. This once again proves the effectiveness of pixel-level modal mixing for enhancing the relatively small proportion of ground objects. In experiments in different regions, BL showed consistent advantages in both qualitative and quantitative results.
[0066] The above experiments have proved the superiority of the modality balance knowledge distillation framework. This may be due to the unique properties of remote sensing images. The present invention has made innovative designs in the following three aspects: (1) The idea of modal balance symmetry was proposed. Multimodal images have completely different features. Currently, most transfer learning networks only operate at the feature level, which leads to a significant bias of the model towards a certain modality. We designed a modal balance strategy that progresses from the image level to the pixel level and from the overall to the local level, so that the model has the ability to segment different modalities similarly.
[0067] (2) Improved the problem of the target domain not having accurate labels in knowledge distillation. Currently, most knowledge distillation networks rely solely on pseudo labels of the target domain to judge the accuracy of knowledge transfer during training. However, pseudo labels are usually inaccurate, which affects the generalization performance of the model. We use the mutual conversion of source and target domain styles and pixel-level mixing to make the target domain also have some real labels, which makes the supervision information more accurate during the knowledge transfer process and promotes the transfer of knowledge from the target domain to the source domain, thereby improving the segmentation accuracy.
[0068] (3) Pixel-level mixing rules that are consistent with remote sensing scenes are formulated. In large-scale real scenes, the long-tail distribution of objects is particularly significant. This is a key feature that distinguishes remote sensing images from natural images. The virtual modal mixing operation formulates mixing rules that are consistent with remote sensing images, achieving category enhancement while avoiding destroying the actual distribution of foreground objects.
[0069] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0070] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0071] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A remote sensing image segmentation method based on a modality balance knowledge distillation framework, characterized in that: include: Obtain source domain optical images and target domain SAR images; Performing image-level virtual modality generation on the source domain optical image and the target domain SAR image to obtain a virtual SAR image and a virtual optical image; Inputting the source domain optical image and the virtual SAR image into dual-modal supervised learning to obtain a source domain optical image pseudo label and a virtual SAR image pseudo label; Inputting the target domain SAR image and the virtual optical image into bimodal knowledge reasoning to obtain a target domain SAR image pseudo label and a virtual optical image pseudo label; Based on the source domain real label, the target domain SAR image pseudo label, the source domain optical image and the target domain SAR image, pixel-level virtual modality generation is performed taking into account the proportion of ground object categories to obtain a mixed image and a mixed pseudo label; Based on the virtual SAR image pseudo-label, the virtual optical image pseudo-label, the virtual SAR image and the virtual optical image, pixel-level virtual modality generation is performed taking into account the proportion of ground object categories to obtain a mixed virtual image and a mixed virtual pseudo-label; The mixed image and the mixed virtual image are input into mixed modality supervised learning to obtain a remote sensing image segmentation result.
2. The remote sensing image segmentation method based on the modal balance knowledge distillation framework according to claim 1 is characterized in that: The source domain optical image and the target domain SAR image are subjected to image-level virtual modality generation to obtain a virtual SAR image and a virtual optical image, including: The source domain optical image and the target domain SAR image are processed respectively using a pix2pixHD-based style transfer method to generate the virtual SAR image and the virtual optical image.
3. The remote sensing image segmentation method based on the modal balance knowledge distillation framework according to claim 1 is characterized in that: Inputting the source domain optical image and the virtual SAR image into dual-modal supervised learning to obtain a source domain optical image pseudo label and a virtual SAR image pseudo label, including: Determining a supervised encoder and a supervised decoder for the bimodal supervised learning; The source domain optical image is sequentially subjected to multi-scale fusion by the supervised encoder and the supervised decoder to obtain a pseudo label of the source domain optical image; The virtual SAR image is sequentially passed through the encoder and the decoder for multi-scale fusion to obtain a pseudo label of the virtual SAR image; Determine that the loss of the source domain optical image pseudo label is a first loss function, the loss of the virtual SAR image pseudo label is a second loss function, and the first loss function and the second loss function constitute a dual-modal supervised learning loss function; Among them, the first loss function includes the cross entropy loss between the source domain true label and the virtual SAR image pseudo label, and the softened Dice loss between the source domain true label and the virtual SAR image pseudo label, and the second loss function includes the cross entropy loss between the source domain true label and the virtual SAR image pseudo label, and the softened Dice loss between the source domain true label and the virtual SAR image pseudo label.
4. The remote sensing image segmentation method based on the modal balance knowledge distillation framework according to claim 1 is characterized in that: Inputting the target domain SAR image and the virtual optical image into bimodal knowledge reasoning to obtain a target domain SAR image pseudo label and a virtual optical image pseudo label, including: Determining an inference encoder and an inference decoder for the bimodal knowledge reasoning; The knowledge acquired in the bimodal supervised learning phase is transferred using exponential moving average weights; The target domain SAR image is sequentially subjected to multi-scale fusion by the inference encoder and the inference decoder to obtain a pseudo label of the target domain SAR image; The virtual optical image is sequentially passed through the inference encoder and the inference decoder for multi-scale fusion to obtain the virtual optical image pseudo label.
5. The remote sensing image segmentation method based on the modal balance knowledge distillation framework according to claim 1, characterized in that: Based on the source domain real label, the target domain SAR image pseudo label, the source domain optical image and the target domain SAR image, pixel-level virtual modality generation taking into account the proportion of ground object categories is performed to obtain a mixed image and a mixed pseudo label, including: Determine the background and foreground in the image based on the true label of the source domain and the proportion of preset ground object categories; Mapping the background pixel value in the source domain true label to 0 and the foreground pixel value to 1 to generate a first mask; Mapping background pixel values in the target domain SAR image pseudo-label to 1 and foreground pixel values to 0 to generate a second mask; Performing a dot product of the first mask and the second mask to obtain a final mask; Based on the final mask, the source domain optical image and the target domain SAR image, the mixed image is calculated; The mixed pseudo label is calculated based on the final mask, the source domain true label and the target domain SAR image pseudo label.
6. The remote sensing image segmentation method based on the modal balance knowledge distillation framework according to claim 5 is characterized in that: Based on the virtual SAR image pseudo-label, the virtual optical image pseudo-label, the virtual SAR image and the virtual optical image, pixel-level virtual modality generation taking into account the proportion of ground object categories is performed to obtain a mixed virtual image and a mixed virtual pseudo-label, including: Based on the final mask, the virtual SAR image and the virtual optical image, the mixed virtual image is calculated; The mixed virtual pseudo label is calculated based on the final mask, the virtual SAR image pseudo label and the virtual optical image pseudo label.
7. The remote sensing image segmentation method based on the modal balance knowledge distillation framework according to claim 6 is characterized in that: Inputting the mixed image and the mixed virtual image into mixed modality supervised learning to obtain a remote sensing image segmentation result, including: Determining that the hybrid modality supervised learning adopts the supervised encoder and the supervised decoder in the bimodal supervised learning; The mixed image is sequentially passed through the supervised encoder and the supervised decoder for multi-scale fusion to obtain a mixed pseudo-label segmentation result; The mixed virtual image is sequentially passed through the supervised encoder and the supervised decoder for multi-scale fusion to obtain a mixed virtual pseudo-label segmentation result; The mixed pseudo-label segmentation result is supervised by the mixed pseudo-label to construct a third loss function, wherein the third loss function includes a cross entropy loss between the mixed pseudo-label and the mixed pseudo-label segmentation result, and a softened Dice loss between the mixed pseudo-label and the mixed pseudo-label segmentation result; The mixed virtual pseudo-label segmentation result is supervised by the mixed virtual pseudo-label to construct a fourth loss function, wherein the fourth loss function includes a cross entropy loss between the mixed virtual pseudo-label and the mixed virtual pseudo-label segmentation result, and a softened Dice loss between the mixed virtual pseudo-label and the mixed virtual pseudo-label segmentation result.
8. A remote sensing image segmentation system based on a modal balance knowledge distillation framework, characterized in that: include: An acquisition module is used to acquire source domain optical images and target domain SAR images; An image-level generation module, used for performing image-level virtual modality generation on the source domain optical image and the target domain SAR image to obtain a virtual SAR image and a virtual optical image; A dual-modal supervised learning module, used for inputting the source domain optical image and the virtual SAR image into dual-modal supervised learning to obtain a source domain optical image pseudo label and a virtual SAR image pseudo label; A bimodal knowledge reasoning module, used for inputting the target domain SAR image and the virtual optical image into bimodal knowledge reasoning to obtain a target domain SAR image pseudo label and a virtual optical image pseudo label; A first pixel-level generation module is used to generate a pixel-level virtual modality taking into account the proportion of ground object categories based on the source domain real label, the target domain SAR image pseudo label, the source domain optical image and the target domain SAR image to obtain a mixed image and a mixed pseudo label; A second pixel-level generation module is used to generate a pixel-level virtual modality taking into account the proportion of ground object categories based on the virtual SAR image pseudo-label, the virtual optical image pseudo-label, the virtual SAR image and the virtual optical image to obtain a mixed virtual image and a mixed virtual pseudo-label; The mixed modality supervised learning module is used to input the mixed image and the mixed virtual image into the mixed modality supervised learning to obtain the remote sensing image segmentation result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the remote sensing image segmentation method based on the modal balance knowledge distillation framework as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the remote sensing image segmentation method based on the modal balance knowledge distillation framework as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Virtual modal imaging calculation method based on multi-level consistency
CN118866320A
SAR image classification method based on multi-modal knowledge distillation transmission
CN119131477A
Cited By
Modal missing scene image processing method for transfer learning
CN120563946A
Optical and SAR (Synthetic Aperture Radar) collaborative domain adaptive segmentation method for flood and ponding
CN122115873A
Floodwater-oriented optical and sar collaborative domain adaptive segmentation method
CN122115873B