Metalearning-based small sample space target key component segmentation method and equipment
By applying a small sample space target key component segmentation method based on meta-learning in the spatial target ISAR image segmentation task, combining MAML and an improved dual-branch Swin-UNet network, the problem of insufficient training efficiency and generalization ability in small sample scenarios is solved, and efficient segmentation effect and stability are achieved.
Patent Information
- Application Number
- CN202510224943.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
AI Technical Summary
The existing deep learning key component segmentation method is difficult to effectively adapt to new tasks in small sample scenarios, and it is difficult to obtain measured data in spatial target ISAR images, high label production costs, and limited sample size, which makes the model difficult to train and lack of generalization capabilities.
Using a small sample space target key component segmentation method based on meta-learning, combining model-independent meta-learning (MAML) and an improved dual-branch Swin-UNet network, we adapt to new tasks under a small number of labeled samples through rapid gradient updates, improving training efficiency and generalization capabilities.
It significantly improves the training efficiency in a small sample environment, reduces the model's dependence on large-scale data sets, improves segmentation accuracy and stability, and enhances the model's generalization ability in a small sample environment.
Smart Images

Figure CN120163833A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and particularly relates to a method and device for segmenting key components of space targets based on meta-learning. Background Art
[0002] Inverse Synthetic Aperture Radar (ISAR) can perform high-resolution imaging on space targets, and obtain key information such as the size and projection structure of the target through image feature extraction and interpretation, which has important applications in the fields of situation awareness, target recognition, etc. The segmentation of key components in ISAR images of space targets is to perform pixel-level precise segmentation on specific components (such as solar panels, satellite bodies, etc.) in the imaging results to assist in feature extraction, attitude estimation, and structural analysis of the target.
[0003] Existing ISAR image segmentation techniques can be divided into traditional segmentation methods and deep learning-based segmentation methods. Traditional segmentation methods mainly rely on simple edge extraction techniques. These methods are only applicable to grayscale images, cannot meet the requirements of component-level segmentation, and have poor application effects under complex scenarios or low signal-to-noise ratio conditions. Due to the lack of effective capture of global semantic information in images, these methods cannot cope with the structural changes of complex components in space target images.
[0004] Classic segmentation models based on deep learning, such as the Fully Convolutional Network (FCN), Deeplab series, U-Net and its variant networks, although have achieved certain results in some image segmentation tasks, still have significant limitations in ISAR image segmentation. First of all, the network structures of these models are relatively shallow, resulting in insufficient ability to extract global semantic information in images. Most existing networks can only establish a mapping relationship between a single pixel value and the label image value in the image, and cannot deeply understand the complex structure of space targets and the subtle differences between components, resulting in low segmentation accuracy. In addition, existing fully supervised learning methods rely on a large amount of labeled data for training, while the cost of making labels for ISAR images is high and the sample size is limited, resulting in a strong dependence of the model on label data. Traditional networks highly rely on the accurate one-to-one correspondence between ISAR images and labels during the training process, which limits the generalization ability of the model under small sample conditions. Especially in the case of scarce labeled data, traditional segmentation methods are difficult to use limited samples for effective training and learning, resulting in the model being unable to achieve the expected segmentation effect in practical applications.
[0005] In summary, existing key component segmentation methods for deep learning usually rely on large-scale labeled data for training, and it is difficult to effectively adapt to new tasks in small-sample scenarios. For spatial target ISAR images, it is difficult to obtain measured data in real application scenarios, the cost of making imaging result labels is high, and the sample size is limited. Therefore, it leads to problems such as difficulty in training deep learning models for key component segmentation of spatial targets and insufficient generalization ability. Summary of the Invention
[0006] To solve the above problems existing in the prior art, the present invention provides a small-sample spatial target key component segmentation method and device based on meta-learning.
[0007] The technical problems to be solved by the present invention are achieved through the following technical solutions:
[0008] The present invention provides a small-sample spatial target key component segmentation method based on meta-learning, including:
[0009] Obtain the image to be segmented, and preprocess the image to be segmented to obtain a preprocessed image;
[0010] Use the trained improved dual-branch Swin-UNet network to perform segmentation processing on the preprocessed image to obtain the key component segmentation result of the target in the image to be segmented;
[0011] Among them, the trained improved dual-branch Swin-UNet network is trained based on the model-agnostic meta-learning method; among them, the improved dual-branch Swin-UNet network has multi-scale feature extraction ability and global feature modeling ability.
[0012] The present invention also provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. The processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0013] The memory is used to store a computer program;
[0014] When the processor is used to execute the program stored on the memory, it implements the steps of the above-mentioned small-sample spatial target key component segmentation method based on meta-learning.
[0015] Compared with the prior art, the beneficial effects of the present invention:
[0016] 1) The present invention combines Model-Agnostic Meta-Learning (MAML) for meta-learning optimization, enabling the network model to adapt to new tasks through rapid gradient updates relying on only a small number of labeled samples, significantly improving the training efficiency in the few-shot scenario and reducing the model's dependence on large-scale datasets. It reduces the overfitting risk of traditional deep learning methods in few-shot ISAR image tasks and overcomes the limitation of high data volume requirements of existing methods.
[0017] 2) By introducing MAML and an improved dual-branch Swin-UNet method, the present invention enables the model to efficiently learn the key component information of spatial targets with extremely few samples, improving the segmentation accuracy and stability, achieving fast training and adaptive parameter adjustment with a small number of samples, and significantly enhancing the generalization ability of the model in the few-shot scenario;
[0018] 3) The improved dual-branch Swin-UNet network of the present invention extracts features using image patches of different scales, capturing both coarse-grained and fine-grained features simultaneously, enabling the segmentation network to maintain a balance between global information and local details, effectively improving the segmentation effect in the few-shot scenario. In addition, by introducing a Spatial Feature Fusion (SFF) module, global dependencies can be established between multi-scale features, realizing efficient information interaction between different-scale features and further enhancing the model's expressive ability;
[0019] 4) The present invention designs an adaptive hybrid loss function, aligning semantic information at different scales through a dual-scale feature consistency loss, while combining a spatial contrast loss to enhance the distinguishability between components, and using a Gaussian adjustment factor that changes over time to adaptively adjust the loss weight, making the training process more stable and efficient.
[0020] The following will further elaborate on the present invention in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings
[0021] Figure 1 is a schematic flowchart of a few-shot spatial target key component segmentation method based on meta-learning provided by an embodiment of the present invention;
[0022] Figure 2 is a schematic structural diagram of a Swin-UNet network model;
[0023] Figure 3 is a schematic diagram of the working principle and structure of each Swin transformer module;
[0024] Figure 4It is a schematic diagram of the working principle and structure of the improved dual-branch Swin-UNet network model provided by the embodiments of the present invention;
[0025] Figure 5 It is a schematic diagram of the working principle of the exemplary SFF module provided by the embodiments of the present invention;
[0026] Figure 6 It is a schematic diagram of the model of the main body and solar panels of a satellite provided by the embodiments of the present invention;
[0027] Figure 7 It is a schematic diagram of the training effect of the network model after training the network model proposed by the present invention using the method proposed by the present invention. Detailed implementation manners
[0028] The following further describes the present invention in detail with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0029] The present invention proposes a method for segmenting key components of small-sample space targets based on meta-learning, aiming to solve the problems that the deep learning model is difficult to train and has insufficient generalization ability due to difficulties in obtaining measured data, high cost of imaging result labeling, and limited sample size in the real application scenario of space target ISAR images, so as to improve the performance of segmenting key components of space targets. In addition, aiming at the problems of high computational complexity and large consumption of computing resources of the Transformer network structure in the segmentation task, the present invention combines MAML and an improved dual-branch Swin-UNet network to achieve fast training and parameter adaptive adjustment using a small number of samples, while effectively reducing the computational burden. The method of the present invention realizes the segmentation of key components of space targets through image preprocessing, an improved dual-branch Swin-UNet segmentation network, and small-sample training based on meta-learning, which are respectively used to improve the effectiveness of small-sample data, enhance the generalization ability of the model in the small-sample environment, and combine the global feature modeling advantages of the Transformer network. The following further describes the present invention.
[0030] Figure 1 It is a schematic flowchart of a method for segmenting key components of small-sample space targets based on meta-learning provided by the embodiments of the present invention. As Figure 1 shown, the method includes:
[0031] S101. Obtain the image to be segmented, and preprocess the image to be segmented to obtain a preprocessed image.
[0032] Here, the preprocessing may be to perform size cropping and / or data augmentation processing on the image. For example, the data augmentation processing may include random image scaling, random flipping, etc.
[0033] S102. Use the trained improved dual-branch Swin-UNet network to segment the preprocessed image to obtain the segmentation result of the key components of the target in the image to be segmented. Among them, the trained improved dual-branch Swin-UNet network is obtained by training based on the model-agnostic meta-learning method. Among them, the improved dual-branch Swin-UNet network has the ability of multi-scale feature extraction and global feature modeling.
[0034] In the present invention, the network architecture of the improved dual-branch Swin-UNet network is an improvement on Figure 2 the Swin-UNet segmentation network architecture shown. The Swin-UNet segmentation network includes an encoder and a decoder structure composed of Swin transformer modules. Different from the multi-head self-attention (MSA) proposed in traditional transformers, Swin transformer proposes window multi-head self-attention (W-MSA) for training parameter and image information fusion. The core idea of this method is to divide the image into multiple non-overlapping small windows, and then process each window separately. Specifically, each calculation of multi-head self-attention (MSA) is only performed within the current window, which is different from the way Vision Transformer (ViT) directly performs MSA calculation and feed-forward neural network processing on the entire image. After processing, the results of each window are integrated to obtain the final output of the entire image. In this way, the number of model parameters and the computational complexity are significantly reduced, thereby greatly reducing the training time and improving the training efficiency. Subsequently, in the next layer, the strategy of shifted window multi-head self-attention (SW-MSA) is adopted to move the window partition to generate new windows, and cross-window connections are added between the new windows and the old windows to ensure information interaction between different windows. The working principle and structure of each Swin transformer module are as Figure 3 shown, and its corresponding calculation formula is as follows:
[0035]
[0036]
[0037] Among them, represents the output feature of the l-th (S)W-MSA module, z lrepresents the output features of the l-th MLP module, where LN(·) is the function corresponding to Layer Normalization, and MLP(·) is the function corresponding to Multi-Layer Perceptron. z l+1 z l-1 Similarly.
[0038] The formula for the self-attention mechanism is as follows:
[0039]
[0040] Among them, Q (Query), K (Key), and V (Value) are the matrices of query, key, and value respectively. Self-attention calculates the similarity between the query and all keys through dot product to weight each value. M and d represent the sizes of patches and Q or K in the window respectively. B is the value in the self-bias matrix, which is used to adjust the output.
[0041] Although W-MSA and SW-MSA can effectively establish long-range dependencies between patches and the shifted window allows for more interaction of local information, there is still a problem that the pixel information within each window (such as edge and line features) may be diluted and the original pixel-level features cannot be fully retained. The improved dual-branch Swin-UNet network structure provided by the present invention is used to effectively extract global and local features of images. This network structure performs feature extraction in the encoder and decoder through Swin Transformer and adopts a dual-branch architecture to process feature maps of different scales respectively. The high-resolution feature maps provide fine-grained spatial information, while the low-resolution feature maps enhance global semantic information. By processing features of different scales in parallel, the improved dual-branch Swin-UNet can better capture the complex features of spatial targets, especially accurately identify the edges and details of targets in component segmentation tasks. Specifically, in order to improve the segmentation performance and enhance the robustness of the model, the present invention uses multi-scale Swin Transformer for feature extraction. Patches of different scales can complement each other in feature extraction. Large-scale images can capture more extensive context information and are especially suitable for processing the global structure of images. By dividing the image into larger blocks, the model can establish long-range dependencies and improve the understanding of the entire image. Small-scale images can more precisely capture local details in the image, including pixel-level information such as edges and lines. At a small scale, more detailed information within each patch is retained, enabling better processing of the details. Therefore, in the present invention, the improved dual-branch Swin-UNet network includes: a first image patch segmentation layer, a second patch segmentation layer, a first linear embedding layer, a second linear embedding layer, a dual-branch encoder, multiple feature fusion modules, a decoder, and a segmentation module.The segmentation scale of the first patch segmentation layer is larger than that of the second patch segmentation layer. The inputs of both the first patch segmentation layer and the second patch segmentation layer are images. The output of the first patch segmentation layer is connected to the input of the first linear embedding layer, and the output of the first linear embedding layer is connected to the input of the first branch in the dual-branch encoder; the output of the second patch segmentation layer is connected to the input of the second linear embedding layer, and the output of the second linear embedding layer is connected to the input of the second branch in the dual-branch encoder; the outputs of the first branch and the second branch are respectively connected to the two inputs of a feature fusion module, and moreover, the output of this feature fusion module is connected to the input of the decoder; each of the remaining feature fusion modules performs a skip connection between one Swin Transformer block in the first branch and the corresponding one Swin Transformer block in the decoder, where one Swin Transformer block consists of two cascaded Swin Transformer modules; the output of the decoder is connected to the segmentation module, and the segmentation module outputs the segmentation result of the key components of the target in the image. Exemplarily, the segmentation scale S of the first patch segmentation layer is 8, and the segmentation scale S of the second patch segmentation layer is 4, for extracting features at different spatial levels. Thus, in the improved dual-branch Swin-UNet network proposed in the present invention, when the length and width of the input image are H and W respectively, the resolutions of the small-scale image patches are respectively. and The resolutions of the large-scale image patches are respectively and
[0042] The feature fusion module in the present invention is an SFF module. In the SFF module, the low-resolution feature map is extended to the same size as the high-resolution feature map through an upsampling operation, and then the two are fused pixel by pixel with weights. This process enables the network to utilize both local details and global semantic information simultaneously, improving the recognition accuracy of the key target components in the image. In addition, the SFF module models the relationship between features of different scales through a spatial feature interaction mechanism, enhancing the effective fusion between features, and further improving the accuracy and robustness of image segmentation.
[0043] Exemplarily, Figure 4 is a network architecture diagram of the improved dual-branch Swin-UNet network provided by the present invention. As Figure 4 shown, this network includes 4 SFFs. Correspondingly, the first branch of the dual-branch encoder includes 3 first network blocks and 1 second network block cascaded in sequence, the second branch of the dual-branch encoder includes 3 first network blocks and 1 second network block cascaded in sequence, the decoder includes 3 third network blocks cascaded in sequence, and the segmentation module is not shown in Figure 4 AsFigure 4 As shown, each first network block includes a Swin Transformer block and a cascaded downsampling layer, and a Swin Transformer block is composed of two cascaded Swin Transformer modules; each second network block includes a Swin Transformer block; each third network block includes an upsampling layer and a cascaded Swin Transformer block. The two inputs of the first SFF are respectively connected to the outputs of the Swin Transformer blocks in the first first network block in the first branch and the second branch, and the output is connected to an input of the Swin Transformer block in the third third network block in the decoder. The two inputs of the second SFF are respectively connected to the outputs of the Swin Transformer blocks in the second first network blocks in the first branch and the second branch, and the output is connected to an input of the Swin Transformer block in the second third network block in the decoder. The two inputs of the third SFF are respectively connected to the outputs of the second network blocks in the first branch and the second branch, and the output is connected to the input of the first third network block in the decoder.
[0044] For example, Figure 4 As shown, the segmentation scale of the first patch segmentation layer is S=8, the segmentation scale of the second patch segmentation layer is S=4, and the input of the first patch segmentation layer and the second patch segmentation layer is a satellite image. Figure 4 The working principle of the network architecture diagram of the improved dual-branch Swin-UNet network provided by the present invention is described. Figure 4 As shown in the figure, after the image to be segmented is input, the first and second patch segmentation layers are used to divide the image into 16 and 64 non-overlapping patches of fixed sizes, i.e., if S=4, the patch is 4×4, and if S=8, the patch is 8×8. Then, the first and second linear embedding layers are used to map the feature dimensions of each corresponding patch to an embedding space suitable for Transformer processing. Specifically, each patch is regarded as a token and mapped to the specified dimension C through a fully connected layer to form a feature vector sequence. If the input image size is H×W, when the segmentation scale is S, then tokens, each of which is projected into C dimensions through a linear embedding layer. Subsequently, the first linear embedding layer inputs the obtained embedded features into the first branch of the dual-branch encoder, and the second linear embedding layer inputs the obtained embedded features into the second branch of the dual-branch encoder; the first branch and the second branch both contain four processing stages. Specifically, in the first stage, when the input resolution is After the embedded features, the adjacent 2×2 patch features are concatenated through the Patch Merging layer, and applied to the linear layer to compress the number of channels. Finally, the output resolution is The number of channels is expanded to 2C. After each downsampling, the feature dimension becomes twice the original, and the spatial dimension is halved. Therefore, the feature maps output by the four processing stages in the first branch are large-scale and low-resolution feature maps, and the resolutions of the output feature maps are and respectively, with the number of channels being C, 2C, 4C, and 8C; the feature maps output by the four processing stages in the second branch are small-scale and high-resolution feature maps, and the resolutions of the output feature maps are and respectively, with the number of channels being C, 2C, 4C, and 8C. Continuing to refer to Figure 4 , one feature map output by each processing stage in the first branch and the second branch is processed by a corresponding SFF module. Specifically, Figure 5 is the working principle diagram of the SFF module. Combining Figure 4 and Figure 5 , it can be seen that each SFF module is used to perform feature fusion on a high-resolution feature map and a low-resolution feature map of the input. Specifically, assuming that the size of a low-resolution feature map of the input is and the size of a high-resolution feature map of the input then the low-resolution feature map is first upsampled to through bilinear interpolation operation to match its size with that of the high-resolution feature map; then, the upsampled low-resolution feature map and the high-resolution feature map are concatenated in the channel dimension to form a fused feature map with a size of The fused feature map is independently passed through 1×1 convolutional layers (W q and W k ) to map to a unified semantic space, and the number of channels is mapped from 2C to C to generate Query and Key. The high-resolution feature map directly passes through a 1×1 convolutional layer (W v ) to generate Value. After that, through the reshape step, the spatial dimension is flattened into a sequence form, and the shape of the feature map becomes Next, the dependence of each position on the global feature is calculated through the attention weight e i'j' , and the specific formula is: where, (K j' W K ) T represents the transpose operation of Key, which adjusts the dimension from to Then Softmax normalization is performed again, and the specific formula is: Then, the weighted sum of Value and the normalized attention weights is calculated, and the fused feature map obtained from the sum is connected to the original high-resolution feature map by residual connection to form a more comprehensive feature representation, which helps the model to have better robustness in complex scenarios. The specific formula is as follows: Output = Z + F high , where F high represents the original high-resolution feature map. Continuing to refer to the above Figure 4 , each of the first three SFF modules outputs the more comprehensive feature representation to the Swin transformer block in a corresponding network block in the decoder for processing, and the fourth SFF module outputs the more comprehensive feature representation to the sampling layer in a network block of the decoder for processing. Similarly, the decoder also includes three processing stages. In the first processing stage, the upsampling layer uses nearest neighbor interpolation (NearestUpsampling) to double the resolution of the feature map output by the fourth SFF module. Subsequently, the output features of the corresponding SFF module are concatenated with the upsampled feature map along the channel dimension through skip connection to obtain multi-scale semantic information. These concatenated features are further processed by two cascaded Swin Transformer modules, which can capture the global and local information of the image and model the global dependencies through the self-attention mechanism to improve the detail restoration ability. Specifically, as Figure 4 shown, the sizes of the feature maps output by the first, second, and third processing stages of the decoder are The decoder not only includes three processing stages but also includes a resolution restoration stage. As shown in the above Figure 4 , after the feature map with a size of is output in the third processing stage, it is also sampled and restored to the original resolution H×W×C at once through a 4× bilinear interpolation layer for the feature and output. Finally, the feature map with a resolution of H×W×C output by the decoder is input into the segmentation module. The segmentation module applies a 1×1 convolutional layer to the Figure 1 feature map to adjust the number of channels to the number of classes, and generates the probability maps of three classes for each pixel through softmax activation. Finally, the index is extracted through the class to map to an RGB image, and then multiplied by the original input pixels to obtain the final segmentation result. As shown, different colors are used in the obtained segmentation result to segment the regions where the key components of the satellite (satellite body, solar panel) are located. Figure 4
[0045] The SFF module adopted in the present invention enhances the expression ability of multi-scale features in the image segmentation task by fusing feature information of different scales. By introducing feature maps of different scales and combining global and local information, the SFF module can significantly improve the recognition accuracy of the network for key components of spatial targets. By fusing and modeling features at different scales, the module improves the multi-scale perception ability of the model, thereby enhancing the network's understanding ability of complex image structures. In the ISAR image segmentation network, the low-resolution image may contain high-level feature information such as the overall contour of the spatial target after feature extraction, while the high-resolution image may contain low-level feature information such as edge textures after feature extraction. The attention mechanism dynamically combines the information of both to ensure that the restored details are consistent with the global semantics.
[0046] In the process of training the improved dual-branch Swin-UNet network of the present invention, MAML is cited. During the learning process, the training data is an image with a true label of the segmentation target. Given a model P i , a set of training samples is collected which is the support set in meta-learning. This model is represented as f(x; θ0), where x is the input image and θ0 represents the initialization parameters of the model. After updating the parameters of the model k times using gradient descent, the parameters are θ k , which can be expressed as: where L is the loss function, y is the label corresponding to the input image x, is the number of samples in the current support set, and α is the learning rate of the internal optimization. The iterative process shown above is called the internal optimization process. After internal optimization, it is also necessary to verify the generalization ability of the optimized model. On the same model P i , another sample set is generated which is the query set of meta-learning and contains the number of samples. Calculate the average loss L of the current model on the query set mean : where y is the label corresponding to the input image x. Calculate the loss after one iteration and update the initialization parameter θ0 with the following formula: where N” represents the number of models, β is the learning rate of the outer loop, and θ * represents the updated parameter. The process shown above is the external optimization process. The gradient of the external optimization is backpropagated through the gradient of the internal optimization. By the above steps, the required model can be obtained using the gradient descent method.
[0047] Specifically, in the present invention, the training process of the above-mentioned trained improved dual-branch Swin-UNet network is the following steps S1 to S4:
[0048] S1. Obtain a dataset; the dataset contains multiple labeled sample images; the labels are used to represent the regions where the key components of the targets in the sample images are located.
[0049] Exemplarily, when making a dataset of space targets, first use each space target model as input, and use electromagnetic calculation methods to obtain a highly realistic ISAR image of the space target as a sample image. Then, use the models of the respective key components of the space model as input, and use two-dimensional spatial projection to obtain accurate labels for the highly realistic ISAR image of the space target. Use this method to generate multiple sample images and their labels; afterwards, perform data preprocessing on all sample images and their labels. For example, first reconstruct the resolution of all sample images and their labels to 128*128, and then apply the same data augmentation strategy to all sample images and their labels, including random image scaling, random flipping, etc. Finally, obtain the dataset.
[0050] S2. Divide the dataset into a support set D s , a query set D v and a test set D t , and sample the support set and the query set respectively according to the number of categories N and the number of tasks N s and N v in the few-shot training task to obtain different task sets T s and T v .
[0051] In the present invention, the number of categories N = 1, and the number of tasks N s and N v can be set according to actual needs. For example, N s = 3, N v = 5, then 1×3 labeled sample images can be sampled from the support set, and these 3 labeled sample images constitute the task set T s , and 1×5 labeled sample images can be sampled from the query set, and these 5 labeled sample images constitute the task set T v .
[0052] S3. Construct two improved dual-branch Swin-UNet networks, denoted as the first network f A (θ) and the second network f B (φ).
[0053] S4. Use the task sets T s , T v and the loss function to perform multiple rounds of iterative training on the first network and the second network, and use the first network obtained in the last round of training as the trained improved dual-branch Swin-UNet network, where each round of training includes one inner-loop training and one outer-loop training.
[0054] Specifically, the steps of the j-th round of training include:
[0055] 1a) When j is 1, initialize the parameters of the first network f A (θ) to obtain a set of initial parameters θ0. At the same time, set the parameters φ of the second network f B (φ) to 0. Then, assign the set of initial parameters θ0 to the second network f B (φ) to obtain the second network f B (θ0) with parameter assignment;
[0056] 1b) When j is greater than 1, obtain the first network f A (θ j-1 ) and the second network f B (φ j-1 ) obtained from the (j - 1)-th round of training. Assign the parameters θ A (θ j-1 ) of the first network f j-1 obtained from the (j - 1)-th round of training to the second network f B (φ) to obtain the second network f B (θ j-1 ) with parameter assignment;
[0057] 2) Use the task set T s , the loss function, and the second network f B (θ0) / f B (θ j-1 ) with parameter assignment to perform one update on the parameters φ / φ B (φ) of the second network f j-1 to complete one inner-loop training and obtain a new set of parameters φ1 / φ B ; j ;
[0058] Specifically, input each sample image in the task set T s into f B (θ0) / f B (θ j-1 ). According to the labels of these sample images and the feature maps generated by f B (θ0) / f B (θ j-1 ), use the loss function to calculate the loss corresponding to each sample After that, update the parameters φ / φ j-1 by performing gradient descent on the loss to correspondingly obtain the new parameters φ1 / φ j .
[0059] In the present invention, the loss function is a loss function based on the dual-scale feature consistency loss; wherein the dual-scale feature consistency loss is the consistency loss between the low-resolution feature map output by the first branch and the high-resolution feature map output by the second branch. Exemplarily, the expression of the loss function is as follows:
[0060]
[0061] Where N' is the number of pixels in the feature map, F H Represents a high-resolution feature map containing detailed information, F L represents a low-resolution feature map containing global structural information, represents the low-resolution feature map output by the i-th network block in the first branch of the decoder, represents the high-resolution feature map output by the ith network block in the second branch of the decoder, φ(·) and ψ(·) are the transformation functions used to align features of different scales, Represents the L2 norm, which is used to measure the consistency of two scale features. is the dual-scale feature consistency loss of the i-th layer, which is used to align feature representations of different scales to ensure the consistency of segmentation results in the global and local range. Specifically, is the dual-scale feature consistency loss between the low-resolution feature map output by the i-th network block in the first branch and the high-resolution feature map output by the i-th network block in the second branch. λ(i) is the level weight, which can be set according to actual needs.
[0062] Multi-layer fusion loss function designed by the present invention By calculating the dual-scale feature consistency loss at different levels Constraining features at different levels Figure 1 The consistency enables the model to capture both local details and global structures, improving the segmentation robustness in complex scenes.
[0063] For example, φ1 and φ j The expressions are:
[0064]
[0065] in, and Both represent gradient descent of the loss, and α is the learning rate of the inner loop.
[0066] 3) A new set of parameters φ1 / φ j Assign to the second network f B (φ), and obtain the second network f obtained by the jth round of training B (φ1) / f B (φj ) Assign a new set of parameters φ1 and φ j to the first network f A (θ), obtaining the first network f A (φ1) / f A (φ j );
[0067] 4) Use the task set T v , the loss function, and the first network f A (φ1) / f A (φ j ) after parameter assignment, and update the parameters θ0 / θ A of the first network f j-1 (θ) once to complete one outer loop training, obtaining a new set of parameters θ1 / θ A of the first network f j ;
[0068] Specifically, input each sample image in the task set T s into f A (φ1) / f A (φ j ). According to the labels of these sample images and the feature maps generated by f A (φ1) / f A (φ j ), calculate the loss corresponding to each sample using the loss function After that, update the parameter θ0 / θ j-1 by performing gradient descent on the loss, correspondingly obtaining the new parameter θ1 / θ j .
[0069] Exemplarily, the expressions of θ1 and θ j are respectively:
[0070]
[0071] where and both represent performing gradient descent on the loss, and β is the learning rate of the outer loop.
[0072] 5) Assign the new parameter θ1 / θ j to the first network f A (θ), obtaining the first network f A (θ1) / f A (θ j ) obtained from the j-th round of training.
[0073] The present invention also provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. The processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0074] The memory is used for storing a computer program;
[0075] The processor is used for implementing the steps of the above-mentioned small-sample space target key component segmentation method based on meta-learning when executing the program stored on the memory.
[0076] The present invention has the following advantages:
[0077] 1) The present invention proposes an improved dual-branch Swin-UNet architecture, which uses patches of different scales for feature extraction, realizes dual-scale feature learning, and can capture coarse-grained and fine-grained feature information simultaneously, thereby improving the accuracy of semantic segmentation. In addition, the present invention combines a spatial feature fusion module, effectively models the global dependence relationship between different-scale features, and realizes efficient information interaction among multi-scale features, thereby enhancing the segmentation ability and robustness of the model.
[0078] 2) The present invention introduces model-agnostic meta-learning. Through rapid adaptive training with a small number of samples, efficient spatial target component segmentation is realized in an environment with low data volume and limited parameters. The present invention can enable the model to quickly adjust parameters, improve the generalization ability, and reduce the computational complexity under the condition of extremely few labeled data, thereby effectively coping with the challenges of data scarcity and limited computing resources in the ISAR image segmentation task.
[0079] 3) The present invention designs a multi-layer fusion loss function. By calculating the consistency loss of dual-scale features at different levels, the consistency of features at different levels is constrained, enabling the model to capture local details and global structures simultaneously, and enhancing the segmentation robustness in complex scenarios. Figure 1 The superiority of the present invention is further illustrated by the following simulation experiments.
[0080] The following further illustrates the superiority of the present invention through simulation experiments.
[0081] (1) Experimental conditions
[0082] The spatial target selects a satellite, and the model of the target component is as Figure 6 shown. To fully verify the effectiveness of the method proposed by the present invention, high-fidelity electromagnetic computational imaging is used to obtain ISAR images. The horizontal and vertical axes of the images are the Doppler axis and the range axis respectively. The improved dual-branch Swin-UNet network is used as the basic segmentation network. The experiment selects different components of the spatial target as segmentation tasks for comparison to verify the effectiveness of the network under different scales and structures under small-sample conditions.
[0083] (2) Experimental content and result analysis
[0084] (2.1) Simulation steps
[0085] Introduce the MAML training segmenter. The input image size is 128×128 pixels, and a data augmentation strategy is used to obtain the segmentation dataset. When training the model, the tasks are divided into two subtask sets. In order to find the relatively optimal initial parameters for each task, a relatively large learning rate must be used, while the learning rate for optimizing the true model parameters should be relatively small. In the experiment, the internal loop learning rate is 0.01, and the external loop uses a learning rate of 0.001.
[0086] Use the MAML method to train the model for 300 rounds. At the beginning of each round, the random generator randomly selects 3 and 5 samples from the dataset as the support set and the query set respectively; during the testing process, since one experiment is random and accidental and cannot accurately reflect the model performance, the random generator is used to repeatedly iterate and generate 100 groups of experimental data from the test set, and one sample is randomly selected independently for each test.
[0087] (2.2) Evaluation metrics
[0088] To quantify the segmentation results of the present invention, the Intersection over Union (IoU), Dice Similarity Coefficient (Dsc), Sensitivity (Sen), Specificity (Spe), Accuracy (Acc), Mean Absolute Error (MAE), and Mean Squared Error (MSE) are used as evaluation metrics. The definitions of different evaluation metrics are as follows:
[0089]
[0090]
[0091] Among them, TP (True Positive) represents the number of samples correctly predicted as positive examples, FP (False Positive) represents the number of samples wrongly predicted as positive examples, TN (True Negative) represents the number of samples correctly predicted as negative examples, and FN (False Negative) represents the number of samples wrongly predicted as negative examples. A is the number of pixels in the predicted region; B is the number of pixels in the true label; N is the total number of pixels.
[0092] (2.3) Analysis of experimental results
[0093] The training results of the small-sample space target key component segmentation method based on meta-learning proposed by the present invention are as follows Figure 7 shown. Among them, (a) is the test image, (b) is the accurately annotated image, and (c) is the segmentation results of the target main body and solar panels obtained from a test image. It can be seen from the visualization results that the segmentation network proposed by the present invention clearly segments components, can effectively identify and extract the main body and the panels, can also achieve good segmentation of electromagnetic images with small inter-class differences, and has good segmentation performance under small-sample conditions. The experimental indicators are shown in Table 1.
[0094] Table 1
[0095] Task Category IoU Dsc Sen Spe Acc MAE MSE Satellite Main Body 72.12% 81.38% 88.20% 90.43% 92.65% 0.065 0.072 Solar Panel 78.72% 85.51% 89.05% 92.12% 95.18% 0.048 0.058
[0096] The improved Swin-Unet segmentation network based on meta-learning proposed by the present invention adopts MAML training and a dual-branch structure, combines the SFF method, and effectively improves the small-sample target key component segmentation performance of ISAR images. By improving indicators such as the accuracy, sensitivity, specificity, and accuracy of the model, and at the same time reducing the MAE, the effectiveness of the proposed method is proved. Especially under challenging conditions such as small samples and complex backgrounds, the model can maintain high performance and low errors. This method has strong feasibility in practical applications, especially suitable for situations where data is scarce or annotation is difficult.
[0097] It should be noted that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the present invention, the meaning of "a plurality" is two or more unless otherwise specifically defined.
[0098] In the description of this specification, the description referring to terms such as "an embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0099] In the specification, the word "including" does not exclude other components or steps, and "one" or "a" does not exclude the case of a plurality. Certain measures are described in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0100] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A small sample space target key component segmentation method based on meta-learning, characterized in that: include: Acquire an image to be segmented, and preprocess the image to be segmented to obtain a preprocessed image; Using the trained improved dual-branch Swin-UNet network, the pre-processed image is segmented to obtain a key component segmentation result of the target in the image to be segmented; The trained improved two-branch Swin-UNet network is obtained by training based on a model-independent meta-learning method; the improved two-branch Swin-UNet network has multi-scale feature extraction capabilities and global feature modeling capabilities.
2. The method for segmenting key components of small sample space objects based on meta-learning according to claim 1, characterized in that: The improved dual-branch Swin-UNet network includes: a first image block segmentation layer, a second image block segmentation layer, a first linear embedding layer, a second linear embedding layer, a dual-branch encoder, multiple feature fusion modules, a decoder and a segmentation module; Among them, the segmentation scale of the first image block segmentation layer is larger than that of the second image block segmentation layer, the inputs of the first image block segmentation layer and the second image block segmentation layer are both images, the output of the first image block segmentation layer is connected to the input of the first linear embedding layer, and the output of the first linear embedding layer is connected to the input of the first branch in the dual-branch encoder; the output of the second image block segmentation layer is connected to the input of the second linear embedding layer, and the output of the second linear embedding layer is connected to the input of the second branch in the dual-branch encoder; the outputs of the first branch and the second branch are respectively connected to the two inputs of a feature fusion module, and the output of the feature fusion module is connected to the input of the decoder; each of the remaining feature fusion modules jumps a Swin Transformer block in the first branch and the second branch with a corresponding Swin Transformer block in the decoder; the output of the decoder is connected to the segmentation module, and the segmentation module outputs the key component segmentation result of the target in the image.
3. The small sample space target key component segmentation method based on meta-learning according to claim 2 is characterized in that: The improved dual-branch Swin-UNet network includes: 4 feature fusion modules; the first branch includes 3 first network blocks and 1 second network block cascaded in sequence; the second branch includes 3 first network blocks and 1 second network block cascaded in sequence; the decoder includes 3 third network blocks cascaded in sequence; Each first network block includes a Swin Transformer block and a cascaded downsampling layer; each second network block includes a Swin Transformer block; each third network block includes an upsampling layer and a cascaded Swin Transformer block, and a Swin Transformer block is two cascaded Swin Transformer modules; The two inputs of the first feature fusion module are respectively connected to the outputs of the Swin Transformer blocks in the first first network block in the first branch and the second branch, and the output is connected to one input of the Swin Transformer block in the third third network block in the decoder; The two inputs of the second feature fusion module are respectively connected to the outputs of the Swin Transformer blocks in the second first network blocks in the first branch and the second branch, and the output is connected to one input of the Swin Transformer block in the second third network block in the decoder; The two inputs of the third feature fusion module are respectively connected to the outputs of the second network blocks in the first branch and the second branch, and the output is connected to the input of the first third network block in the decoder.
4. The small sample space target key component segmentation method based on meta-learning according to claim 2 or 3, characterized in that: Each feature fusion module is a SFF module.
5. The method for segmenting key components of small sample space objects based on meta-learning according to claim 2 or 3, characterized in that: The first image block segmentation layer is used to divide an input image into 16 non-overlapping image blocks with the same width and height according to the width and height of the image; The second image block segmentation layer is used to divide an input image into 64 non-overlapping image blocks with the same width and height according to the width and height of the image.
6. The method for segmenting key components of small sample space objects based on meta-learning according to claim 3, characterized in that: The loss function used when training the improved two-branch Swin-UNet network is a loss function based on a dual-scale feature consistency loss; wherein the dual-scale feature consistency loss is the consistency loss between the low-resolution feature map output by the first branch and the high-resolution feature map output by the second branch.
7. The method for segmenting key components of small sample space objects based on meta-learning according to claim 6, characterized in that: The expression of the loss function used when training the improved two-branch Swin-UNet network is as follows: in, is the dual-scale feature consistency loss between the low-resolution feature map output by the i-th network block in the first branch and the high-resolution feature map output by the i-th network block in the second branch, and λ(i) is the level weight.
8. The method for segmenting key components of small sample space objects based on meta-learning according to claim 1, characterized in that: The trained improved dual-branch Swin-UNet network is trained using the following method: Acquire a data set; the data set includes a plurality of sample images with labels; the labels are used to indicate the areas where key components of the targets in the sample images are located; The dataset is divided into a support set, a query set, and a test set, and the number of categories N and the number of tasks N in the small sample training task are calculated. s and N v The support set and the query set are sampled respectively to obtain different task sets T s and T v ; Construct two improved dual-branch Swin-UNet networks, respectively denoted as the first network f and A (θ) and the second network f B (φ); Using task set T s , T v And the loss function, for the first network f A (θ) and the second network f B (φ) performing multiple rounds of iterative training, and using the first network obtained from the last round of training as the trained improved two-branch Swin-UNet network, wherein each round of training includes one inner loop training and one outer loop training.
9. The method for segmenting key components of small sample space objects based on meta-learning according to claim 8, characterized in that: The steps of the j-th round of training include: When j is 1, the first network f is initialized A (θ) parameters, get a set of initialization parameters θ0, and at the same time, make the second network f B The parameter φ of (φ) is 0, and then the set of initialization parameters θ0 is assigned to the second network f B (φ), and obtain the second network f after parameter assignment B (θ0); When j is greater than 1, get the first network f obtained from the j-1th round of training A (θ j-1 ) and the second network f B (φ j-1 ), the first network f obtained by the j-1th round of training A (θ j-1 ) parameter θ j-1 Assign to the second network f B (φ), and obtain the second network f after parameter assignment B (θ j-1 ); Using task set T s , the second network f after the loss function and the parameter assignment B (θ0) / f B (θ j-1 ), for the parameter φ / φ j-1 Perform an update to complete an inner loop training and obtain the second network f B A new set of parameters φ1 / φ of (φ) j ; The new set of parameters φ1 / φ j Assign to the second network f B (φ), and obtain the second network f obtained by the jth round of training B (φ1) / f B (φ j ); The new set of parameters φ1 / φ j Assign to the first network f A (θ), get the first network f after parameter assignment A (φ1) / f A (φ j ); Using task set T v , the first network f after the loss function and the parameter assignment A (φ1) / f A (φ j ), for the parameter θ0 / θ j-1 Perform an update to complete an outer loop training and obtain the first network f A A new set of parameters θ1 / θ for (θ) j ; The new set of parameters θ1 / θ j Assign to the first network f A (θ), and get the first network f obtained by the jth round of training A (θ1) / f A (θ j ).
10. An electronic device comprising a processor, a communication interface, a memory and a communication bus, characterized in that: The processor, the communication interface and the memory communicate with each other via the communication bus; The memory is used to store computer programs; The processor is used to implement the method steps described in any one of claims 1-9 when executing the program stored in the memory.
Citation Information
Cited By
Gravity inversion imaging method based on DT-UNet
CN121033323A