Method for segmenting radio frequency interference by using Swin Transform to embed U2-Net

Through Swin Transformer embedded in the U2-Net model, combined with auxiliary encoder and main encoder, the missed detection and missed detection problems of narrowband RFI in RF interference segmentation are solved, and higher precision radio frequency interference segmentation is achieved.

CN120354892AActive Publication Date: 2025-07-22KUNMING UNIV OF SCI & TECH

Patent Information

Application Number
CN202510847387.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

Existing deep learning methods have problems of missed detection and missed detection when segmenting radio frequency interference (RFI) in radio telescope observation data, especially narrowband RFI, which affects the integrity and accuracy of the observation data.

Method used

Swin Transformer is used to embed the U2-Net model, combined with auxiliary encoder and main encoder, and through spatial interaction modules, feature compression modules and relational aggregation modules, the feature representation and segmentation capability of RF interference, especially narrowband RFI.

Benefits of technology

It improves the segmentation accuracy of radio frequency interference, reduces false detection and missed detection, and improves the overall segmentation performance of radio observation data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354892A_ABST
    Figure CN120354892A_ABST
Patent Text Reader

Abstract

The invention discloses a method for segmenting radio frequency interference by embedding U2-Net by using Swin Transform, and belongs to the technical field of radio astronomical image segmentation based on deep learning. The method comprises the following steps: simultaneously inputting a time-frequency image of radio observation data into a main encoder and an auxiliary encoder in a trained ST-U2Net model, and respectively extracting local detail features and global context information; and fusing the auxiliary encoder intermediate characteristic pattern output by each stage in the auxiliary encoder and the main encoder intermediate characteristic pattern output by each stage in the main encoder through a relationship aggregation module in the main encoder, inputting the fused image into a decoder, and outputting a prediction segmentation mask image containing a radio frequency interference region and a non-interference region through the decoder. Compared with an existing deep learning radio frequency interference segmentation method, the method has obvious advantages in the aspects of narrow-band RFI segmentation, overall segmentation precision and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for using Swin Transformer to embed U 2 -Net to segment radio frequency interference, belonging to the technical field of radio astronomical image segmentation based on deep learning. Background Art

[0002] Any unwanted signal received by a radio telescope can be called radio frequency interference. The data observed by a radio telescope is crucial in radio astronomy research. However, when it captures radio signals from the universe, it inevitably encounters radio frequency interference (RFI). RFI usually originates from various artificial signals, such as electromagnetic radiation from wireless communication, satellite transmission, and other electronic devices. These interferences not only occupy an increasingly large frequency band but also seriously affect the quality of the observed data of cosmic signals, posing a great challenge to scientific research. With the improvement of the sensitivity and resolution of radio telescopes, the problem of RFI has become more prominent, and there is an urgent need to develop effective segmentation methods to ensure the accuracy of scientific data.

[0003] Traditional RFI segmentation methods include singular value decomposition of the time-frequency matrix, VarThreshold, and wavelet transform, etc., which have been widely applied to various actual observed data and shown good results. However, when we apply these traditional methods to the actual observed data collected by the 40 m radio telescope of the Yunnan Observatory, it is found that these methods perform well in segmenting broadband RFI signals covering multiple frequency channels in the data, but perform poorly in segmenting narrowband RFI with fewer channels and shorter duration. This limitation results in the easy missed detection of narrowband RFI, thus affecting the integrity and accuracy of the observed data.

[0004] To overcome the limitations of traditional methods, in recent years, semantic segmentation of RFI images based on convolutional neural networks (CNNs) has received extensive attention and has gradually become an important means to solve interference problems in radio data processing. For example, early research by Akeret et al. used the U-Net structure to segment RFI and achieved good results. Subsequently, Kerrigan et al. further designed a deep fully convolutional network (DFCN) to segment RFI in the Hydrogen Epoch of Reionization Array (HERA) by fusing amplitude and phase information and achieved excellent results. R-Net was trained on simulated data and then fine-tuned on real data, achieving superior performance compared to U-Net. Although these methods have shown better results than traditional methods on simulated data, when applying these deep learning methods to the actual observational data collected by the 40 m radio telescope at the Yunnan Observatory, it is found that there are still obvious deficiencies in segmenting narrowband RFI signals, manifested as the omission of narrowband RFI and the misdetection of non-RFI as RFI, resulting in inaccurate segmentation of RFI. Summary of the Invention

[0005] The present invention provides a method for segmenting radio frequency interference by embedding Swin Transformer into U 2 -Net, for constructing an ST-U 2 Net model, and segmenting radio frequency interference in the time-frequency image of the radio observational data to be predicted based on the trained ST-U 2 Net model.

[0006] The technical solution of the present invention is as follows:

[0007] A method for segmenting radio frequency interference by embedding Swin Transformer into U 2 -Net, comprising the following steps:

[0008] Obtain a trained ST-U 2 Net model; the ST-U 2 Net model includes an auxiliary encoder based on Swin Transformer and a main encoder and decoder based on U 2 -Net;

[0009] Obtain the radio observational data to be predicted; input the time-frequency image of the radio observational data into the trained ST-U 2The main encoder and the auxiliary encoder in the Net model extract local detailed features and global context information respectively; the intermediate feature maps of the auxiliary encoder output at each stage in the auxiliary encoder and the intermediate feature maps of the main encoder output at each stage in the main encoder are fused through the relationship aggregation module at each stage in the main encoder and then input into the decoder; the decoder gradually restores to the original resolution feature map through multi-level upsampling and feature fusion operations, and finally outputs a predicted segmentation mask image containing the radio interference area and the non-interference area through convolution operations and the Sigmoid activation function.

[0010] Further, in the auxiliary encoder, the time-frequency image of the radio observation data is first divided into overlapping blocks by the image partitioning module; subsequently, these overlapping blocks are flattened and projected into the embedding dimension through the linear embedding layer; then, the output of the linear embedding layer is processed through four consecutive first downsampling stages and the feature compression module to obtain the output of each downsampling stage, that is, the intermediate feature maps of the auxiliary encoder at each stage; the outputs of the four downsampling stages are used as the inputs of the four second downsampling stages in the main encoder; each of the four first downsampling stages includes a plurality of sequentially connected transformers, and the transformers are constructed by integrating the spatial interaction module into the improved Swin Transformer block. The input of the first downsampling stage is used as the input of the spatial interaction module and the improved Swin Transformer block in the first transformer. The outputs of the spatial interaction module and the improved Swin Transformer block in the first transformer are added as the output of the first transformer, and the output of the first transformer is used as the input of the second transformer.

[0011] Further, the improved Swin Transformer block is constructed by taking the traditional Swin Transformer block as the framework and introducing KAN to replace the MLP in the traditional Swin Transformer block.

[0012] Further, in the main encoder, the time-frequency image of the radio observation data is first input into the residual U-shaped structure (RSU) module with a first preset height, and then passes through the efficient multi-scale attention mechanism module to generate the first-scale feature map. Next, the first-scale feature map further passes through four consecutive second downsampling stages. In the second downsampling stage, it is processed by the RSU module and the efficient multi-scale attention mechanism module in the second downsampling stage to obtain the intermediate feature maps of the main encoder at each stage. The intermediate feature maps of the main encoder at each stage and the intermediate feature maps of the auxiliary encoder at each stage are fused through the relationship aggregation module in the corresponding second downsampling stage to obtain the second, third, fourth, and fifth-scale feature maps. The fifth-scale feature map is used as the input of the RSU-4F module to obtain the RSU-4F output feature map. The first, second, third, fourth, and fifth-scale feature maps and the RSU-4F output feature map are used as the input of the decoder.

[0013] Further, in the four consecutive second downsampling stages from the input to the output direction, the height of the RSU module in each stage decreases sequentially.

[0014] Further, the height of the RSU module with the first preset height is the same as the height of the RSU module in the second downsampling stage of the RSU module in the four consecutive second downsampling stages that is close to the RSU module with the first preset height.

[0015] Further, the ST-U 2 Net model uses the MIou as the primary evaluation index.

[0016] The beneficial effects of the present invention are:

[0017] The present invention proposes an ST-U 2 Net model with a dual encoder structure combining Swin Transformer and residual U-shaped structure for segmenting RFI in the actual observation data captured by radio telescopes. It introduces an efficient multi-scale attention mechanism through the main encoder, strengthens the recombination of feature channels and pixel-level relationship modeling, and improves the segmentation accuracy of RFI under complex backgrounds. The auxiliary encoder enhances the feature representation of narrowband RFI and reduces detail loss by introducing a spatial interaction module and a feature compression module respectively, and uses the Kolmogorov-Arnold network (KAN) to replace the multi-layer perceptron (MLP) in Swin Transformer to improve the modeling ability of the narrowband RFI features. The relationship aggregation module in the main encoder fuses the intermediate features of the main encoder and the auxiliary encoder to effectively combine the local details and global information of RFI. Compared with the existing deep learning radio frequency interference segmentation methods, the present invention has obvious advantages in aspects such as segmenting narrowband RFI and overall segmentation accuracy. Description of the Drawings

[0018] Figure 1 This is the structure diagram of the ST-U 2 Net network model proposed by the present invention based on a dual encoder structure;

[0019] Figure 2 This is the structure diagram of the traditional Swin Transformer and the improved Swin Transformer block of the present invention;

[0020] Figure 3 This is the real observation data of the pulsar PSR J0032 +5434 observed by the 40-meter radio telescope of the Yunnan Observatory used in the present invention;

[0021] Figure 4 This is the experimental result diagram of the fourth group of simulations in Example 2. Detailed implementation manners

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other arbitrarily.

[0023] Example 1: As Figures 1 - 4 shown, according to the first aspect of the embodiment of the present invention, a method for using SwinTransformer to embed U 2 -Net to segment radio frequency interference is provided, including: obtaining a trained ST-U 2 Net model; the ST-U 2 Net model includes an auxiliary encoder based on Swin Transformer, a main encoder based on U 2 -Net, and a decoder; obtaining radio observation data to be predicted; simultaneously inputting the time-frequency image of the radio observation data into the trained ST-U 2The main encoder and the auxiliary encoder in the Net model extract local detailed features and global context information respectively; the auxiliary encoder intermediate feature maps output at each stage in the auxiliary encoder and the main encoder intermediate feature maps output at each stage in the main encoder are fused through the relationship aggregation module at each stage in the main encoder and then input into the decoder; the decoder gradually restores to the original resolution feature map through multi-level upsampling and feature fusion operations, and finally outputs a predicted segmentation mask image including the radio frequency interference area and the non-interference area through convolution operations and the Sigmoid activation function. The model of the present invention uses the Swin Transformer module as the auxiliary encoder to inject multi-scale context prior information into the main encoder of the U 2 -Net, effectively making up for the deficiencies of existing methods in extracting narrowband RFI features. By constructing a dual-encoder structure, it can fuse the local details and global features of RFI, and finally achieve an improvement in the RFI segmentation performance in the time-frequency images of complex radio observation data.

[0024] Further, in the auxiliary encoder, the time-frequency image of the radio observation data is first divided into overlapping 8×8 blocks through the image partitioning module Patch Partition, and an overlapping rate of 50% is set to ensure obtaining a feature map with pixel continuity and semantic relevance; subsequently, these overlapping blocks are flattened and projected into the embedding dimension through the linear embedding layer; then, the output of the linear embedding layer is processed through four consecutive first downsampling stages and the feature compression module to obtain the output of each downsampling stage, that is, the auxiliary encoder intermediate feature map of each stage; the outputs of the four downsampling stages are used as the inputs of the four second downsampling stages in the main encoder; each of the four first downsampling stages includes a plurality of sequentially connected transformers, and the transformers are constructed by integrating the spatial interaction module into the improved Swin Transformer block. The input of the first downsampling stage is used as the input of the spatial interaction module and the improved Swin Transformer block in the first transformer. The outputs of the spatial interaction module and the improved Swin Transformer block in the first transformer are added as the output of the first transformer. The output of the first transformer is used as the input of the second transformer (that is, as the input of the spatial interaction module (SIM) and the improved Swin Transformer block in the second transformer. The outputs of the spatial interaction module and the improved Swin Transformer block in the second transformer are added as the output of the second transformer. The output of the second transformer is used as the input of the third transformer), and so on. The output of the last transformer is used as the input of the feature compression module.

[0025] Further, as Figure 2 shown ( Figure 2(a) is a traditional Swin Transformer block, Figure 2 (b) is an improved Swin Transformer block. The improved Swin Transformer block uses the traditional Swin Transformer block as a framework and introduces KAN to replace the MLP in the traditional Swin Transformer block, so as to improve the model's ability to characterize radio frequency interference signals through a more flexible activation function.

[0026] In the traditional Swin Transformer block structure, it includes LN, MLP, and window multi-head self-attention (W-MSA) and shifted window multi-head self-attention (SW-MSA) used alternately. Based on this structure, for the prediction of radio frequency interference, the present invention considers the following aspects:

[0027] First, to solve the problem of fine feature extraction and precise segmentation of narrowband RFI in complex backgrounds, the present invention introduces the Kolmogorov-Arnold Network (KAN) to replace the traditional MLP. Different from MLP using a fixed activation function at nodes, KAN uses a learnable activation function in weights. By configuring the adjustable network size and piecewise polynomial order, KAN can generate a fine activation grid in the range of [-1,1]. In addition, KAN parameterizes each weight using univariate spline functions, thus getting rid of the dependence on linear weights. This non-linear weight representation method significantly improves the model's ability to express complex functional relationships. Due to the lightweight structure of KAN, it can achieve or exceed the accuracy of the traditional MLP with a smaller network scale. At the same time, during the network expansion process, the accuracy improvement speed of KAN is also better than that of MLP, and it can achieve higher segmentation accuracy with lower computational costs. Based on these design concepts, the present invention replaces MLP with KAN to further enhance the model's feature extraction and segmentation ability for narrowband RFI in complex backgrounds.

[0028] In a second aspect, although the traditional Swin Transformer block structure effectively reduces the memory overhead by establishing relationships between image patches and adopts the strategy of alternating execution of window attention (W-MSA) and shifted window attention (SW-MSA), this modeling method weakens its global modeling ability to some extent. In addition, due to the small space occupancy and weak features of narrowband RFI, some spatial information is required to help the model learn the features of narrowband RFI. Therefore, the present invention decides to use a spatial interaction module (SIM module) to strengthen spatial information exchange and encode more accurate spatial information at the same time. The SIM module uses a 3×3 convolution with a dilation ratio of 2 to reconstruct the spatial structure of the input features and compress the channel dimension to optimize the technical efficiency; the SIM module introduces two spatial dimensions, extracts the vertical and horizontal statistical features of the reconstructed features through global average pooling, not only pays attention to the association between tokens, but also models the relationship between pixels, thereby enhancing the applicability of the Swin Transformer module in image segmentation tasks. Immediately afterwards, the statistical features generate a position-sensitive attention map according to the tensor product method. Finally, the attention map is added to the output of the traditional Swin Transformer block as the output of the transformer.

[0029] In a third aspect, the traditional Swin Transformer constructs a transformer by flattening image patches and projecting the image patches or merging adjacent 2×2 image patches and performing linear processing. However, these methods tend to cause a large amount of loss of detail and structural information, limiting the segmentation of dense and small-scale objects. Given that narrowband RFI also has the characteristics of being dense and small-scale. Therefore, the present invention introduces a feature compression module (FCM) in the downsampling process of image patches to avoid the above problems, thereby improving the segmentation effect of narrowband RFI objects. The FCM adopts a dual-branch collaborative design: the first branch expands the receptive field through hierarchical dilated convolution to extract the features of narrowband RFI; the second branch adopts a soft pooling operation based on exponential weighting to retain more details of the feature map after downsampling. By performing channel alignment and element-wise fusion on the outputs of the two branches, the FCM can effectively enhance the segmentation effect of narrowband RFI.

[0030] Further, in the main encoder, the time-frequency image of the radio observation data is first input into the residual U-shaped structure module (RSU module with a height of 7), and then passes through the efficient multi-scale attention mechanism module to generate the first-scale feature map. Then, the first-scale feature map further passes through four consecutive second downsampling stages. In the second downsampling stage, it is processed by the residual U-shaped structure module and the efficient multi-scale attention mechanism module in the second downsampling stage to obtain the intermediate feature maps of the main encoder at each stage. The intermediate feature maps of the main encoder at each stage and the intermediate feature maps of the auxiliary encoder at each stage are fused through the relationship aggregation module in the corresponding second downsampling stage to obtain the second, third, fourth, and fifth-scale feature maps. The fifth-scale feature map is used as the input of the RSU-4F module to obtain the RSU-4F output feature map. The first, second, third, fourth, and fifth-scale feature maps and the RSU-4F output feature map are used as the input of the decoder.

[0031] Further, in the four consecutive second downsampling stages from the input to the output direction, the height of the residual U-shaped structure RSU module in each stage decreases sequentially (that is, from the input to the output direction, they are RSU-7, RSU-6, RSU-5, and RSU-4 in turn).

[0032] Further, the height of the residual U-shaped structure RSU module with the first preset height is the same as the height of the residual U-shaped structure RSU module in the second downsampling stage of the residual U-shaped structure RSU module close to the first preset height in the four consecutive second downsampling stages (that is, the height of the residual U-shaped structure RSU module with the first preset height is 7).

[0033] Applying the above technical solution, it can be seen that the innovation of the main encoder of the present invention lies in: on the basis of the traditional U 2 -Net, the well-known efficient multi-scale attention mechanism module (i.e., the EMA module) and the relationship aggregation module (RAM module) are introduced. For each of the introduced modules, the following is further described:

[0034] First, models based on channel or spatial attention mechanisms have been widely applied to segmentation tasks. However, traditional methods usually rely on channel dimension reduction to model cross-channel relationships, which may affect the extraction of deep visual features. In addition, there is a trade-off between computational complexity and feature retention in existing attention mechanisms, making it difficult to efficiently capture local and global features simultaneously. To this end, the present invention introduces the EMA module to reduce computational complexity and retain the expressive ability of global features, thereby better supporting the RFI segmentation task of this study. First, EMA remaps some channels of the input features to the batch dimension, enabling the network to efficiently model cross-channel information interaction within different sub-feature groups. In the feature extraction stage, EMA adopts parallel 1×1 and 3×3 convolution branches to perform cross-channel information interaction and local spatial feature extraction respectively. Among them, the 1×1 branch focuses on cross-channel information interaction and encodes long-range dependencies along the horizontal and vertical directions using global average pooling; the 3×3 branch focuses on the extraction of local spatial features and enhances the multi-scale feature modeling ability through a larger receptive field. Subsequently, through the cross-space learning mechanism, EMA fuses the feature maps of the two parallel branches using matrix dot product and captures pixel-level pairing relationships with the help of global average pooling, thereby strengthening the correlation of features globally. Finally, the fused features are calculated with the Sigmoid activation function to obtain the attention weights, which are then applied to the input feature map to generate the output of EMA.

[0035] Second, although the main encoder based on the CNN architecture can extract rich local features in the spatial dimension through convolutional kernels, it lacks explicit modeling of the relationships between channel dimensions. Existing CNN-based encoders mainly rely on the local receptive ability of convolutional kernels in the spatial dimension, but the explicit modeling of the relationships between channel dimensions is still insufficient. When similar distribution patterns appear in different channels, this limitation may lead to confusion in feature representation. This will reduce the discriminative ability of the model and cause the problem that the model misdetects non-RFI as RFI. To address this issue, the present invention introduces the RAM module. The RAM module first extracts the channel dependencies in the global features from the intermediate features of the auxiliary encoder and fuses them into the intermediate features extracted by the main encoder. In addition, RAM introduces deformable convolution and soft pooling operations to adapt to the diversity of the shapes of segmentation targets, thereby further refining the feature extraction of the main encoder. Specifically, RAM first applies deformable convolution and channel dimension transformation to the outputs of the main encoder and the auxiliary encoder respectively to generate geometrically robust features. Subsequently, average pooling, max pooling, and soft pooling are used to quantify the importance of the features of the auxiliary encoder in the channel dimension. Then, channel attention weights are generated through a cascaded fully connected layer. Finally, the result of element-wise multiplication of the attention weights and the features of the main encoder processed by deformable convolution is element-wise added to the result of the features of the auxiliary encoder processed by 1×1 convolution to obtain the output features of RAM.

[0036] In addition, it should be noted that in the present invention, the traditional U 2 -Net is used as the framework, and its main RSU module integrates the hierarchical feature aggregation ability of the U-shaped encoding and decoding structure and the efficient feature transfer mechanism of residual learning. In the encoding stage, the RSU module extracts local details through multiple layers of convolution and gradually integrates global context information to enhance the feature expression ability. In the decoding stage, a cross-level feature splicing and convolution refinement strategy is adopted to reconstruct the spatial boundary of the target with higher precision. Compared with the classical CNN architectures (such as VGG, ResNet, and DenseNet, etc.) that mainly rely on small convolution kernels for feature extraction. The symmetric encoding and decoding characteristics of the RSU module effectively alleviate the problem of the receptive field limitation of the shallow network, making it show stronger robustness to background noise in the RFI segmentation task of radio observation data.

[0037] Furthermore, the decoder is respectively composed of RSU-4, RSU-5, RSU-6, and RSU-7 decoding modules; the fifth-scale feature map and the output feature map of RSU-4F are spliced and upsampled as the input of the RSU-4 decoding module, the fourth-scale feature map and the output of the RSU-4 decoding module are spliced and upsampled as the input of the RSU-5 decoding module, the third-scale feature map and the output of the RSU-5 decoding module are spliced and upsampled as the input of the RSU-6 decoding module, the second-scale feature map and the output of the RSU-6 decoding module are spliced and upsampled as the input of the RSU-7 decoding module, and the first-scale feature map and the output of the RSU-7 decoding module are spliced and upsampled to obtain the radio time-frequency feature map of the original scale; each of the decoding modules, the output feature map of RSU-4F, and the radio time-frequency feature map of the original scale are first processed through a 3×3 convolutional layer, and then upsampled through bilinear interpolation to restore their resolution to the size of the time-frequency image of the original input radio observation data as a prediction mask. The prediction masks are spliced and then processed through a 1×1 convolution and a Sigmoid activation function, and then each pixel's category is predicted through a segmentation head, and finally a predicted segmentation mask image containing interference regions and non-interference regions is output.

[0038] It should be noted that in the RSU-4, RSU-5, RSU-6, and RSU-7 decoding modules, "-4, -5, -6, -7" represent heights of 4, 5, 6, and 7. In the RSU-4F decoding module, "-4F" means that dilated convolution is used to prevent the feature map from being further downsampled, resulting in too low resolution and losing features, and the height is 4.

[0039] Applying the above technical solution, it can be seen that the present invention adopts a dual-encoder structure to balance the segmentation tasks of significant broadband RFI and narrowband RFI in the image. The auxiliary encoder introduces a spatial interaction module and a feature compression module to enhance the feature representation of narrowband RFI and reduce detail loss respectively, thereby improving the segmentation accuracy of small narrowband RFI. In addition, the multi-layer perceptron (MLP) in the traditional SwinTransformer is replaced by (KAN) to enhance the modeling ability of narrowband RFI features. The main encoder is dedicated to accurately extracting the significant broadband RFI features in the radio observation image during the downsampling process, which is crucial for the overall segmentation performance. In the main encoder, the feature representation is enhanced through an efficient multi-scale attention (EMA) mechanism, the channel information is reorganized, pixel-level relationships are captured, and the segmentation accuracy in complex backgrounds is improved. Further, the introduced relationship aggregation module fuses the features of the two encoders to achieve more accurate segmentation of narrowband RFI in radio observation data and reduce the misdetection of non-RFI as RFI.

[0040] According to the second aspect of the embodiments of the present invention, there is provided a device for segmenting radio frequency interference by embedding Swin Transformer into U 2 -Net, including: an acquisition module for acquiring a trained ST-U 2 Net model; the ST-U 2 Net model includes an auxiliary encoder based on Swin Transformer, a main encoder based on U 2 -Net, and a decoder; a prediction module for acquiring radio observation data to be predicted; for simultaneously inputting the time-frequency image of the radio observation data into the main encoder and the auxiliary encoder to extract local detail features and global context information respectively; inputting the auxiliary encoder intermediate feature maps output at each stage in the auxiliary encoder and the main encoder intermediate feature maps output at each stage in the main encoder into the decoder after fusion through the relationship aggregation module at each stage in the main encoder; the decoder gradually restores to the original resolution feature map through multi-level upsampling and feature fusion operations, and finally outputs a predicted segmentation mask image including radio frequency interference regions and non-interference regions through convolution operations and Sigmoid activation functions.

[0041] According to the third aspect of the embodiments of the present invention, there is provided a processor for running a program, wherein when the program runs, it executes the method for segmenting radio frequency interference by embedding Swin Transformer into U 2 -Net as described in any one of the above.

[0042] Embodiment 2: A method for segmenting radio frequency interference by embedding Swin Transformer into U 2 -Net, including the following steps:

[0043] S1. Obtain the observation data of the 40-meter radio telescope at the Yunnan Observatory in Southwest China from September 2016 to October 2022, covering three different pulsars: PSR J0332+5434, J00358+5413, and J0437-4715. Screen out the observation data with sub-integration numbers greater than 90 and channel numbers greater than 128 for research, and a total of 1383 time-frequency image samples and corresponding labels are produced; to ensure the fairness and scientificity of data division, randomly divide the training set, validation set, and test set according to a ratio of 7:2:1. Among them, 970 samples are used for model training, 276 samples are used for model validation, and 138 samples are used for model testing. It should be noted that in the training set, validation set, and test set, any time-frequency image contains both useful pulsar signals and useless radio frequency interference signals, while the corresponding masked label image only marks the radio frequency interference signal part.

[0044] Exemplarily, an example of the time-frequency image of PSR J0332+5434 is as Figure 3 shown, from Figure 3 it can be seen that the RFI is brighter than the background noise or astronomical signals.

[0045] S2. Construct the ST-U Figure 1 Net model as shown in 2 ; The ST-U 2 Net network is built using the Pytorch framework and trained using the AdamW optimizer with a weight decay parameter of 1e-4. The initial learning rate is set to 0.001, and a warm-up strategy is introduced at the beginning of training. In the first 2 epochs, it is gradually increased from a smaller learning rate to 0.001 linearly. After the warm-up stage, the learning rate will be dynamically adjusted by a custom scheduler according to the training progress and gradually decreased to improve the convergence efficiency and final performance of the model. In addition, the training process supports mixed precision training (Mixed Precision Training), which effectively improves the calculation efficiency and reduces the video memory occupancy. In the experiment, the training batch size is set to 4, and the data augmentation strategies include scaling the image to a size of 256×256 pixels, randomly flipping horizontally (with a probability of 0.5), and normalization processing. During the training process, the model is trained for up to 500 epochs and the model performance is evaluated on the validation set every 10 epochs. Finally, the model weights with the best MIoU index performance in the validation set are retained for subsequent testing. All experiments are run on an NVIDIA GeforceRTX 4050 6-GB GPU.

[0046] To enhance the model's segmentation ability for RFI, the present invention selects to use the BCE loss function. The loss function is expressed as:

[0047]

[0048] where L is the total loss, n is the number of classes output by the network, represents the output and the target of the binary cross-entropy loss.

[0049] The present invention uses the F1 score, recall, precision, and MIou to evaluate the model performance. These four evaluation metrics are based on the confusion matrix, which contains four elements: true positive (TP), false positive (FP), true negative (TN), and false negative (FN). For each class, Iou is defined as the intersection over union of the predicted value and the true value, and is calculated as follows:

[0050]

[0051] The F1 score is calculated as follows:

[0052]

[0053] Recall is calculated as follows:

[0054]

[0055] Precision is calculated as follows:

[0056]

[0057] where MIou represents the average value of Iou for all classes.

[0058] S3. Simulation description: For the following groups of simulations, the training set in the present example S1 is used to train the model, and the time-frequency images used in training are uniformly scaled to 256×256 pixels; the model is trained for at most 500 epochs, and the model performance is evaluated on the validation set every 10 epochs, and the results of each round are recorded. Finally, the model with the best MIoU performance on the validation set is used as the best model and is used for subsequent simulation tests.

[0059] The first group of simulations: To evaluate the performance of the proposed ST-U 2 Net model and its various modules, the present invention uses U 2The ablation experiment of -Net as the baseline model was carried out on the dataset. In addition, it was also compared with the ST-UNet model that proposed SIM, FCM, and RAM (that is, an architecture with an auxiliary encoder based on Swin Transformer and a main encoder and decoder based on residual network and RAM. The difference from the model of the present invention is that in the auxiliary encoder, the transformer integrates the spatial interaction module into the traditional Swin Transformer block for construction, and the main encoder and decoder adopt the architecture based on residual network and RAM). Then, a structure was built that obtains the output features of each stage by element-wise addition encoding of the features of the main encoder and the auxiliary encoder in each stage. This structure is simply referred to as STRSU (that is, an auxiliary encoder based on Swin Transformer, based on U 2 -Net's main encoder and decoder architecture. The difference from the model of the present invention is that the entire auxiliary encoder uses the traditional Swin Transformer, and the main encoder and decoder use the traditional U 2 -Net). STRSU was compared with the baseline model and the ST-UNet model on the above-mentioned evaluation metrics. The experimental results are shown in Table 1. The experimental results show that STRSU, which introduces Swin Transformer based on U 2 -Net, can effectively improve the segmentation performance of the baseline model U 2 -Net for RFI compared with ST-UNet that introduces Swin Transformer based on the residual network. In addition, the dual-encoder architecture aggregates more information beneficial to RFI segmentation hierarchically.

[0060] Table 1 Ablation experiment of the dual-encoder structure on the dataset

[0061]

[0062] The second group of simulations: The ablation experiment carried out in this stage aims to explore the influence of introducing SIM, FCM, and RAM modules into the STRSU structure on the segmentation performance. The experimental results are shown in Table 2 (the highest score of each index is marked in bold).

[0063] Table 2 Ablation experiment of SIM, FCM, and RAM modules on the dataset

[0064]

[0065] The experimental results shown in Table 2 indicate that (since the present invention is directed to a segmentation task, and the most important indicator reflecting the segmentation performance is MIou, thus for the convenience of discussion, the present invention mainly focuses on MIou), with the introduction of the SIM, FCM, and RAM modules, the segmentation performance of the model is gradually improved. Specifically, when the SIM module is introduced into STRSU, the MIou is increased by 1.6%. After the introduction of the FCM module, the MIou is increased by 2.0%. After the introduction of the RAM module, the MIou is increased by 1.1%. When the SIM and FCM modules are introduced simultaneously, the MIou is increased by 2.2%. When the SIM and RAM modules are introduced simultaneously, the MIou is increased by 2.3%. When the FCM and RAM modules are introduced simultaneously, the MIou is increased by 2.2%. Finally, when the SIM, FCM, and RAM modules are introduced simultaneously, the MIou is increased by 2.5%.

[0066] The above experimental results show that in the STRSU structure, the combined use of the SIM, FCM, and RAM modules can significantly improve the segmentation performance, especially the improvement in key indicators such as MIou, Precision, and F1 is the most significant. Based on this discovery, in the subsequent ablation experiments, the present invention regards the SIM, FCM, and RAM modules as an overall module for improving the model segmentation performance, and simply refers to it as the Spatial Feature Aggregation Module (SFAM).

[0067] The third group of simulations: The experimental results of the effects of EMA and KAN are shown in Table 3. Among them, the highest score of each indicator is marked in bold (for the same reason as above, this group of simulations mainly focuses on MIou). After the EMA module is introduced after the RSU output of each stage of the main encoder in the STRSU structure, the MIou is increased by 2.3%. This result shows that EMA can enhance the ability of STRSU to extract local features of RFI, thereby improving its segmentation performance. Further, after the SFAM and EMA modules are introduced into STRSU simultaneously, the MIou is increased by 2.9%. This shows that the introduction of EMA in the existing network structure can enhance the segmentation ability of the model. Through the above ablation experiments, the effectiveness of the EMA module in improving the segmentation performance is verified.

[0068] Replacing the MLP with KAN in STRSU improved the MIou by 1.9%. This result indicates that KAN helps enhance the segmentation ability of STRSU. Further, after simultaneously introducing the SFRM and KAN modules in STRSU, the MIou increased by 2.5%. Through the above ablation experiments, the effectiveness of the KAN module in improving the segmentation performance was verified. In addition, after simultaneously introducing KAN and EMA modules in STRSU, the MIou increased by 3.7%. Finally, applying SFRM, EMA, and KAN simultaneously to STRSU, this complete configuration enabled the model to achieve the best performance in both the MIou and Recall evaluation metrics, and the Precision and F1 reached sub-optimal performance. Among them, MIou, Recall, Precision, and F1 increased by 4.0%, 4.7%, 2.4%, and 3.7% respectively.

[0069] Table 3 Ablation experiments of the introduced modules on the dataset

[0070]

[0071] The fourth group of simulations: Comparing the proposed ST-U 2 Net with a variety of existing methods, focusing on evaluating the performance of these models in the RFI segmentation task. The compared models include U 2 -Net based on the CNN architecture, EMSCA-UNet, and DCA-MSFF-UNet, as well as MS-TransUNet based on the Transformer architecture. It is worth noting that MS-TransUNet adopts a serial structure of CNN and Transformer in the encoding stage, while ST-U 2 Net adopts a parallel structure.

[0072] Table 4 shows the various evaluation metrics of the above models in the RFI segmentation task. The experimental results show that ST-U 2 Net is superior to other models in terms of MIou, Recall, Precision, and F1 score. U 2 -Net adopts an innovative U-shaped structure to learn multi-scale features to achieve high-precision image segmentation. The experimental data shows that ST-U 2 Net is superior to U 2- Net in terms of RFI multi-scale feature aggregation. Compared with U 2- Net, ST-U 2The MIou, Recall, Precision, and F1 of Net increased by 4.9%, 7.7%, 2.4%, and 5.4% respectively. MS-TransUNet is based on TransUNet and utilizes a multi-scale convolutional attention mechanism. MS-TransUNet outperforms U in all evaluation metrics. 2- Net. Compared with MS-TransUNet, ST-U 2 Net increased the MIou, Recall, Precision, and F1 by 3.0%, 5.8%, 1.5%, and 3.9% respectively. This result proves that the proposed dual-encoder structure has better performance than the serial models of Transformer and CNN in the RFI segmentation task. DCA-MSFF-UNet is an RFI segmentation model that combines a dual cross-attention mechanism with multi-scale features. The experimental results show that the RFI segmentation performance of DCA-MSFF-UNet is better than that of U 2- Net and MS-TransUNet. However, compared with DCA-MSFF-UNet, ST-U 2 Net increased the MIou, Recall, Precision, and F1 by 2.5%, 3.7%, 0.5%, and 2.3% respectively. EMSCA-UNet adopts a dual attention mechanism and continuously integrates the spatial information of low-level features through skip connections. The experimental data show that the RFI segmentation performance of EMSCA-UNet is better than that of U 2 -Net, DCA-MSFF-UNet, and MS-TransUNet. However, compared with EMSCA-UNet, ST-U 2 Net increased the MIou, Recall, Precision, and F1 by 2.1%, 3.0%, 0.1%, and 1.7% respectively.

[0073] Table 4 Comparison of segmentation results with other methods

[0074]

[0075] Figure 4 Show the prediction results of all segmentation methods in Table 4 ( Figure 4 In (a) represents the original visibility time-frequency image; (b) represents the prediction result of ST-U 2 Net; (c) represents U 2-Net prediction results; (d) represents the prediction results of EMSCA-UNet; (e) represents the prediction results of MS-TransUNet; (f) represents the prediction results of DCA-MSFF-UNet; (g) represents the ground truth label of (a); (h) represents the difference map between (b) and (g); (i) represents the difference map between (c) and (g); (j) represents the difference map between (d) and (g); (k) represents the difference map between (e) and (g); (l) represents the difference map between (f) and (g)). The first column of this figure shows the visibility data and its corresponding ground truth label respectively. From the second column to the last column, it describes the output masks of the above models and the difference maps between the output masks of the above models and the ground truth label. We use the difference maps to more intuitively show the differences between the prediction masks of different models and the ground truth label. The specific method is to calculate the difference between the GT and the output mask of the corresponding model and take the absolute value. The yellow area in the difference map represents the absolute value of the difference between the prediction mask and the ground truth label. From Figure 4 Among these difference maps, it can be seen that the difference between the output mask of the model of the present invention and the ground truth label is the smallest. Compared with the segmentation results of the above methods, ST-U 2 Net not only outperforms other methods in evaluation metrics, but also shows stronger narrowband RFI segmentation ability and lower misjudgment rate in the difference map, meeting the expectations.

[0076] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A method for using Swin Transformer to embed U 2 -Net to segment radio frequency interference, characterized in that Including: Obtain the trained ST-U 2 Net model; the ST-U 2 Net model includes an auxiliary encoder based on Swin Transformer and a main encoder and decoder based on U 2 -Net; Obtain the radio observation data to be predicted; simultaneously input the time-frequency image of the radio observation data into the main encoder and the auxiliary encoder in the trained ST-U 2 Net model to extract local detailed features and global context information respectively; input the auxiliary encoder intermediate feature maps output at each stage in the auxiliary encoder and the main encoder intermediate feature maps output at each stage in the main encoder into the decoder after fusing them through the relationship aggregation module at each stage in the main encoder; The decoder gradually restores the feature map to the original resolution through multi-level upsampling and feature fusion operations, and finally outputs a predicted segmentation mask image containing the radio interference area and the non-interference area through convolution operations and the Sigmoid activation function.

2. Method for separating radio frequency interference by using Swin Transformer embedded in U 2 -Net, characterized in that In the auxiliary encoder, the time-frequency image of the radio observation data is first divided into overlapping blocks by an image partitioning module; subsequently, these overlapping blocks are flattened and projected into the embedding dimension through a linear embedding layer; then, the output of the linear embedding layer is processed by four consecutive first downsampling stages and a feature compression module to obtain the outputs of each downsampling stage, that is, the intermediate feature maps of the auxiliary encoder at each stage; the outputs of the four downsampling stages are used as the inputs of four second downsampling stages in the main encoder; each of the four first downsampling stages includes a plurality of sequentially connected transformers, and the transformers are constructed by integrating a spatial interaction module into an improved Swin Transformer block. The input of the first downsampling stage is used as the input of the spatial interaction module and the improved Swin Transformer block in the first transformer. The outputs of the spatial interaction module and the improved Swin Transformer block in the first transformer are added as the output of the first transformer, and the output of the first transformer is used as the input of the second transformer.

3. The method for separating radio frequency interference by using Swin Transformer embedded in U 2 -Net according to claim 2, characterized in that The improved Swin Transformer block is constructed by taking the traditional Swin Transformer block as a framework and introducing KAN to replace the MLP in the traditional Swin Transformer block.

4. Method for using Swin Transformer to embed U-Net to segment radio frequency interference according to claim 1, characterized in that 2 - In the main encoder, the time-frequency image of the radio observation data is first input into a residual U-shaped structure RSU module with a first preset height, and then passes through an efficient multi-scale attention mechanism module to generate a first-scale feature map; then, the first-scale feature map further passes through four consecutive second downsampling stages. In the second downsampling stage, it is processed by the residual U-shaped structure RSU module and the efficient multi-scale attention mechanism module in the second downsampling stage to obtain the intermediate feature maps of the main encoder at each stage; the intermediate feature maps of the main encoder at each stage and the intermediate feature maps of the auxiliary encoder at each stage are fused through a relationship aggregation module in the corresponding second downsampling stage to obtain second, third, fourth, and fifth-scale feature maps; the fifth-scale feature map is used as the input of the RSU-4F module to obtain the RSU-4F output feature map; the first, second, third, fourth, and fifth-scale feature maps and the RSU-4F output feature map are used as the inputs of the decoder.

5. Method for using Swin Transformer to embed U-Net to segment radio frequency interference according to claim 4, characterized in that 2 In the four consecutive second downsampling stages, from the input to the output direction, the height of the residual U-shaped structure RSU module in each stage decreases sequentially. ​ 6. The method for using Swin Transformer to embed U 2 -Net to divide radio frequency interference, characterized in that The height of the residual U-shaped structure RSU module with the first preset height is the same as the height of the residual U-shaped structure RSU module in the second downsampling stage of the residual U-shaped structure RSU module in the four consecutive second downsampling stages that is close to the first preset height.

7. Method for using Swin Transformer to embed U-Net to divide radio frequency interference according to claim 1, characterized in that, 2 - The ST-U 2 Net model uses MIou as the primary evaluation metric.

Citation Information

Patent Citations

  • Method of image segmentation using CN

    CN111373439A

  • Surgical instrument image intelligent segmentation method and system based on multi-scale feature fusion

    CN113763386A

  • Brain tumor MR image segmentation method

    CN114519719A

  • Transform and U-Net combined medical image liver segmentation method and system

    CN115965633A

  • Monocular depth prediction method based on multi-path feature extraction and multi-scale feature fusion

    CN116758130A

Cited By

  • Mine equipment fault intelligent diagnosis and prediction system based on deep learning

    CN121456614A