Self-supervised SAR target identification method and system based on MSP-MAE model
Through the self-supervised SAR target recognition method based on the MSP-MAE model, the problem of the impact of spot noise in SAR images is solved, high-quality feature learning and target recognition without labeled data are achieved, and the accuracy of recognition and generalization ability of the model are improved.
Patent Information
- Application Number
- CN202510181330.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The spot noise present in SAR images seriously interferes with the semantic interpretation of the image, resulting in overfitting of deep learning models in target recognition, reducing the accuracy of recognition.
The self-supervised SAR target recognition method based on the MSP-MAE model is adopted, and the MSP-MAE and MSP-CT models are constructed and optimized through steps such as data augmentation, self-supervised training, weight migration and supervised fine-tuning to improve the accuracy of target recognition.
Through self-supervised training and weight transfer, the model can learn high-quality feature representation without large amounts of labeled data, reduce dependence on labeled data, and improve the generalization ability and recognition accuracy of the model.
Smart Images

Figure CN120071001A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of radar automatic target recognition, and in particular, to a self-supervised SAR target recognition method and system based on the MSP-MAE model. Background Art
[0002] Synthetic Aperture Radar (SAR), as an important microwave remote sensing technology, plays a key role in many fields such as topographic mapping and target monitoring with its all-weather and all-time working ability. For example, in natural disaster monitoring, SAR can penetrate clouds and dark environments to obtain images of the affected areas, providing important basis for rescue decisions. However, the inherent speckle noise problem in SAR images seriously interferes with the semantic interpretation of the images, making it extremely challenging to accurately identify targets. At present, with the rapid development of computer vision technology, the application of Automatic Target Recognition (ATR) technology in SAR image interpretation is becoming increasingly widespread. Machine learning, especially deep learning methods, has gradually become the mainstream technology for SAR ATR due to its powerful feature learning ability. Deep learning models can automatically learn features from raw data without the need for cumbersome manual feature extraction and classifier design, greatly improving the efficiency and accuracy of SAR target recognition.
[0003] However, the training of deep learning models highly depends on a large amount of labeled data. In the field of SAR images, due to the high difficulty and cost of data annotation, the labeled data is relatively scarce, which makes deep learning models prone to overfitting, severely restricting their performance in SAR target recognition. SAR images have severe speckle noise due to electromagnetic wave backscattering, which greatly interferes with the semantic interpretation of the images, making it difficult to accurately identify targets and hindering the effective application of SAR in practical scenarios such as terrain monitoring and target surveillance. At present, a self-supervised SAR target recognition method and system based on the MSP-MAE model are needed. Summary of the Invention
[0004] In order to solve the problems of the influence of speckle noise in SAR images and the low accuracy of target recognition, the present invention provides a self-supervised SAR target recognition method and system based on the MSP-MAE model. The self-supervised SAR target recognition based on MSP-MAE is carried out, covering multiple links such as data processing, model construction and training, model fine-tuning, and recognition performance evaluation, effectively solving the problem of SAR image target recognition and improving the recognition accuracy.
[0005] In the first aspect, a self-supervised SAR target recognition method based on the MSP-MAE model provided by the present invention adopts the following technical solutions: A self-supervised SAR target recognition method based on the MSP-MAE model includes: Obtain SAR image data for training and testing, including obtaining image data through the OpenSAR - Ship, MSTAR, and FUSAR - Ship datasets; Perform data augmentation based on the obtained SAR image data, including using random cropping, random horizontal flipping, and Mixup methods for data augmentation; Construct the MSP - MAE model, and use the augmented image data to perform self - supervised training on the MSP - MAE model to obtain the model weight parameters of MSP - MAE; Obtain the encoder weight parameters in the trained MSP - MAE model, and use the encoder weight parameters to initialize some weight parameters of the MSP - CT model; Perform parameter fine - tuning on the initialized MSP - CT model, and perform supervised training on the MSP - CT model; Use the fine - tuned MSP - CT model for SAR target recognition, and the output is the recognition result of each test image.
[0006] Further, the constructing the MSP - MAE model and using the augmented image data to perform self - supervised training on the MSP - MAE model includes constructing the MSP - MAE model composed of an encoder and a decoder. The MSP - MAE model combines a multi - scale convolution module and a multi - head self - attention module, and uses a window masking strategy to mask the input image. Among them, the encoder uses the MSP - CT encoder, and a window masking layer is added between the linear embedding layer and the MSP - CT Block in the first stage.
[0007] Further, the constructing the MSP - MAE model and using the augmented image data to perform self - supervised training on the MSP - MAE model also includes that the decoder uses a lightweight decoder divided into three stages. In each stage, first, the input features are expanded twice through a Patch Expanding layer composed of a linear layer and a normalization layer, and the prediction head is used to map the token dimension for reconstructing the original image according to the latent vector output by the encoder. For the prediction error of the masked part in the image, calculate the mean square error between the original image and the reconstructed image as the loss function, and use the loss function to pre - train the constructed MSP - MAE model.
[0008] Further, the part of the weight parameters of the MSP-CT model is initialized using the encoder weight parameters, including discarding the decoder part after the MSP-MAE model training is completed, extracting the weight parameters of the encoder from the trained MSP-MAE model, where the weight parameters are stored in each layer of the encoder in the form of a numerical matrix, determining the corresponding network layer based on the network structures of the MSP-MAE encoder and the MSP-CT model, and assigning the weight parameters of the corresponding layer extracted from the MSP-MAE encoder to the corresponding network layer in the MSP-CT model.
[0009] Further, the parameter fine-tuning of the initialized MSP-CT model includes inputting the complete unmasked image data into the encoder in the assigned MSP-CT to obtain the output features of the MSP-CT encoder, performing global pooling on the output features, inputting the pooled feature vectors into the classifier for classification, and fine-tuning the MSP-CT using the MSTAR and FUSAR-Ship datasets respectively.
[0010] Further, the supervised training of the MSP-CT model includes adopting a structure in which the MSDC layer and the MHSA layer are alternately stacked in four stages of the MSP-CT, performing linear embedding operations through convolutional kernels, and downsampling each stage except the first stage. Among them, the MSDC layer is composed of an MSCF module, a channel attention module, and a spatial attention module. The MSCF module uses a convolutional branch to improve the feature representation, and the spatial attention module enhances the features through pooling, convolution, and activation operations. Finally, the SE module is used by the attention module to obtain channel weights.
[0011] Further, the supervised training of the MSP-CT model also includes using a multi-scale inverted feed-forward network composed of a convolutional layer and a feature fusion layer as the feed-forward network, then inserting a convolutional branch to form a residual connection. The output features pass through a global average pooling layer, a layer normalization layer, a Dropout layer, and a fully connected layer in sequence, and the fully connected layer outputs classification prediction values. The cross-entropy loss function is used to compare the predicted probabilities of each category with the corresponding true labels, and finally the MSP-CT model is trained by the backpropagation algorithm.
[0012] In a second aspect, a self-supervised SAR target recognition system based on the MSP-MAE model includes: A data acquisition module, configured to: acquire SAR image data for training and testing, including acquiring image data through the OpenSAR-Ship, MSTAR, and FUSAR-Ship datasets; A preprocessing module, configured to: perform data augmentation based on the acquired SAR image data, including using random cropping, random horizontal flipping, and the Mixup method for data augmentation; A model module, configured to: build an MSP-MAE model, and perform self-supervised training on the MSP-MAE model using the enhanced image data to obtain the model weight parameters of MSP-MAE; A conversion module, configured to: obtain the encoder weight parameters in the trained MSP-MAE model, and use the encoder weight parameters to initialize some weight parameters of the MSP-CT model; A training module, configured to: perform parameter fine-tuning on the initialized MSP-CT model, and perform supervised training on the MSP-CT model; An output module, configured to: perform SAR target recognition using the fine-tuned MSP-CT model, and the output is the recognition result of each test image.
[0013] In a third aspect, the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded and executed by a processor of a terminal device for the self-supervised SAR target recognition method based on the MSP-MAE model.
[0014] In a fourth aspect, the present invention provides a terminal device, including a processor and a computer-readable storage medium, the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for the self-supervised SAR target recognition method based on the MSP-MAE model.
[0015] In summary, the present invention has the following beneficial technical effects: 1. The present invention uses the MSP-MAE model for self-supervised training, which can make full use of a large amount of unlabeled SAR image data. Through the window masking strategy and the mean square error loss function, the model can focus on learning the features of the masked parts in the image, mine the internal structure and semantic information of the data, so as to learn high-quality feature representations without a large amount of manual annotation, laying a solid foundation for subsequent target recognition tasks.
[0016] 2. The present invention combines the MSP-MAE model with a multi-scale convolution module and a multi-head self-attention module. The multi-scale convolution module can capture the detailed information of the target and improve the quality of the reconstructed image, while the multi-head self-attention module can obtain the global information of the target and help the model better reconstruct the information of the masked block. This way of multi-module cooperation enables the model to understand the features of SAR images more comprehensively and accurately.
[0017] 3. The present invention fine-tunes the MSP-CT model using the MSTAR and FUSAR-Ship datasets, enabling the model to be optimized according to specific target recognition tasks. By processing and classifying complete unmasked image data, the model can better adapt to actual application scenarios and improve the accuracy and reliability of recognition.
[0018] 4. The MSP-CT model of the present invention adopts a structure with alternating stacking of MSDC layers and MHSA layers, as well as a multi-scale reverse feed-forward network, enabling the model to effectively fuse local and global features. The MSDC layer realizes the fusion of multi-scale information and the attention to key information through the MSCF module, channel attention module, and spatial attention module; the MSI-FFN module improves the feature extraction ability and training stability of the model through the fusion of multi-scale features and residual connections, thereby increasing the clarity of image recognition.
[0019] 5. The present invention migrates the weight parameters of the MSP-MAE encoder to the MSP-CT model, achieving effective knowledge transfer. This way of transfer learning allows the MSP-CT model to have good feature extraction ability at the initial stage, reduces the dependence on labeled data, speeds up the convergence rate of the model, and also improves the performance of the model under small-sample data, effectively solving the problem of SAR image target recognition and improving the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is the overall framework diagram of a self-supervised SAR target recognition method based on the MSP-MAE model in Embodiment 1 of the present invention; Figure 2 is the specific structure diagram of MSP-CT in Embodiment 1 of the present invention. Among them, the leftmost part of the picture is the overall structure diagram of MSP-CT, and the right side is the detailed structure diagram of the modules in this network; Figure 3 is the effect diagram of the dilated convolution at different dilation rates in Embodiment 1 of the present invention; Figure 4 is the structure diagram of the SE module in Embodiment 1 of the present invention; Figure 5 is the structure diagram of MSI-FFN in Embodiment 1 of the present invention; Figure 6 is the attention heat map of each stage in MSF-CT in Embodiment 1 of the present invention; Figure 7 is the flowchart of the Patch Merging module in Embodiment 1 of the present invention; Figure 8 is the network structure diagram of the MSP-MAE in Embodiment 1 of the present invention; Figure 9It is a flowchart of window masking in Embodiment 1 of the present invention; Figure 10 It is the effect diagram of different masking strategies in Embodiment 1 of the present invention. Among them, (a) is the original image, (b) is the image after using random masking, and (c) is the image after using window masking; Figure 11 It is five types of SAR ship datasets in OpenSARShip of Embodiment 1 of the present invention, (a) Cargo data map, (b) Fishing data map, (c) Passenger data map, (d) Tanker data map, (e) Tug data map; Figure 12 It is ten types of SAR vehicle datasets in MSTAR of Embodiment 1 of the present invention, (a) 2S1 data map, (b) BMP2 data map, (c) BRDM2 data map, (d) BTR60 data map, (e) BTR70 data map, (f) D7 data map, (g) T62 data map, (h) T72 data map, (i) ZIL131 data map, (j) ZSU234 data map; Figure 13 It is four types of SAR ship datasets in FUSAR-Ship of Embodiment 1 of the present invention, (a) Cargo data map, (b) Fishing data map, (c) Tanker data map, (d) Other data map; Figure 14 It is the effect diagram of the random cropping method in Embodiment 1 of the present invention. Among them, (a) is the original image, and (b), (c), and (d) are the images after random cropping; Figure 15 It is the flow and effect diagram of the Mixup method in Embodiment 1 of the present invention; Figure 16 It is the reconstruction effect diagram of MSF-MAE in Embodiment 1 of the present invention under different masking rates, (a) original image, (b) 75%, (c) 50%, (d) 25%; Figure 17 It is the schematic diagram of the recognition accuracy of each method in Embodiment 1 of the present invention on the MSTAR dataset; Figure 18 It is the schematic diagram of the recognition accuracy of each method in Embodiment 1 of the present invention on the FUSAR-Ship dataset. Detailed implementation manners
[0021] The present invention will be further described in detail below with reference to the accompanying drawings.
[0022] Embodiment 1 Referring to Figure 1 , a self-supervised SAR target recognition method based on the MSP-MAE model in this embodiment includes: Obtain SAR image data for training and testing, including obtaining image data through the OpenSAR - Ship, MSTAR, and FUSAR - Ship datasets; Perform data augmentation based on the obtained SAR image data, including using random cropping, random horizontal flipping, and Mixup methods for data augmentation; Construct the MSP - MAE model, and perform self - supervised training on the MSP - MAE model using the augmented image data to obtain the model weight parameters of MSP - MAE; Obtain the encoder weight parameters in the trained MSP - MAE model, and use the encoder weight parameters to initialize some of the weight parameters of the MSP - CT model; Fine - tune the parameters of the initialized MSP - CT model, and perform supervised training on the MSP - CT model; Use the fine - tuned MSP - CT model for SAR target recognition, and the output is the recognition result of each test image.
[0023] Specifically, a self - supervised SAR target recognition method based on the MSP - MAE model includes the following steps: S1. Obtain SAR image data for training and testing, including obtaining image data through the OpenSAR - Ship, MSTAR, and FUSAR - Ship datasets; As Figure 11 , Figure 12 , Figure 13 shown, collect three types of datasets: OpenSAR - Ship, MSTAR, and FUSAR - Ship. The OpenSAR - Ship dataset is derived from Sentinel - 1 images with a resolution of 2.5m×1.0m, and select the ship target data of 5 categories to provide rich samples for the pre - training of MSP - MAE; the MSTAR dataset is collected by X - band spotlight SAR with a resolution of 0.3m×0.3m and HH polarization, and contains 10 categories of vehicle targets, which are mainly used for the fine - tuning and testing of MSP - CT; the FUSAR - Ship dataset is prepared by the Key Laboratory of Electromagnetic Wave Information Science of Fudan University based on the Gaofen - 3 satellite (GF - 3), and select the ship target data of 4 categories. Among them, the OpenSAR - Ship dataset is used for the self - supervised training of MSP - MAE, and the MSTAR and FUSAR - Ship datasets are used for the fine - tuning and testing of MSP - CT.
[0024] S2. Perform data augmentation based on the obtained SAR image data, including using random cropping, random horizontal flipping, and Mixup methods for data augmentation; AsFigure 14 , Figure 15 As shown, based on the acquired SAR image data, randomly select some regions from the original SAR image for cropping, and then adjust the cropped image to a specific size (such as 224×224). When cropping, the cropping area ratio is randomly selected between [0.6, 1.0], which can simulate the presentation of targets of different sizes in the image. At the same time, the aspect ratio of the width and height of the cropped image is limited to vary within the range of [3 / 4, 4 / 3] to avoid overly distorted images. After cropping, the image is adjusted to the specified size by bicubic interpolation, which can retain the detail information of the image to a certain extent and reduce image distortion. This operation simulates different perspectives and zoom levels, enabling the model to encounter SAR images at different scales and perspectives during training, which helps the model learn richer features. Then, random horizontal flipping and the Mixup method are used for data augmentation. Random horizontal flipping performs a horizontal flipping operation on the SAR image with a probability of 50%. Since the orientation of the targets in the SAR image in the actual scenario is random, random horizontal flipping can increase the diversity of the data and enable the model to learn the features of the targets in different orientations. Finally, the Mixup method generates new samples by mixing two images and their labels in a certain proportion. In SAR image data augmentation, the probability of performing the Mixup operation on each image is set to 50%. Specifically, when implementing, two different SAR images are linearly combined according to specific weights, and their corresponding labels are also mixed in the same weight ratio. This method not only increases the diversity of the data but also enables the model to learn the transitional features between samples during training, improving the robustness of the model. For example, when performing the Mixup operation on the SAR image of a cargo ship and the SAR image of a fishing boat, the generated new sample contains partial features of both types of ships. The model can better understand the differences and connections between different categories during the learning process, so as to more accurately perform target recognition when facing SAR images in complex scenarios.
[0025] S3. Construct the MSP-MAE model, and use the enhanced image data to perform self-supervised training on the MSP-MAE model to obtain the model weight parameters of MSP-MAE; As Figure 1 shown, the MSP-MAE of the present invention is obtained by improving MAE, and the network structure is as Figure 8As shown in the figure. As a self-supervised learning method, MSP-MAE, similar to other autoencoder methods, can be divided into two parts: an encoder and a decoder. The encoder in the network is responsible for encoding the image into corresponding tokens and mapping them to a high-dimensional semantic space, while the decoder will refer to these latent representations and restore the original image. In the method of the present invention, the OpenSAR-Ship dataset is used as an unlabeled data to pre-train the MSP-MAE network. The input image will be randomly masked by 75% of the area and input into the encoder, and the decoder will reconstruct the original image according to the latent vector output by the encoder.
[0026] Pre-train MSP-MAE using an unlabeled dataset. In this process, each input image will be randomly masked by 75% of the area and input into the encoder. In the masking process, the input image will first be divided into windows one by one; then, the first sub-window of each window will be obtained as the index of each window; next, all the obtained sub-windows will be shuffled to disrupt their order; then, a certain proportion of sub-windows will be randomly selected from all the shuffled sub-windows as unmasked windows, and the others that are not selected will be used as masked windows; finally, all the unselected windows will be masked to obtain the masked input data of the model.
[0027] Input the masked input data into the encoder of MSP-MAE. The encoder of MSP-MAE can infer the features of the masked part based on the unmasked data therein; in this process, by analyzing the information in the unmasked blocks, the encoder can establish context relationships. This enables it to better understand the latent features of the masked blocks, so as to perform accurate reconstruction during decoding; moreover, by randomly masking part of the input during the training process, the encoder learns how to still effectively generate high-quality representations under incomplete information, thereby improving the robustness and generalization ability of the model.
[0028] In MSP-MAE, the decoder is responsible for converting the latent representation generated by the encoder back to the original data form, and its main task is to reconstruct the masked part. It uses context information to restore the missing details, ensuring that the generated output is as consistent as possible with the original input in terms of features and structure. By continuously optimizing the mapping relationship, the decoder can effectively handle the situation of missing information and improve the accuracy of reconstruction.
[0029] MSP-MAE combines a multi-scale convolution module and a multi-head self-attention module. The multi-scale convolution module can help the model obtain the detailed information of the target, which helps to improve the quality of the reconstructed image; the multi-head self-attention is used to obtain the global information of the target, so that the model can better reconstruct the information of the masked block through the visible blocks. The loss function used for training calculates the mean square error between the original image and the reconstructed image: , Among them, is the true value of the image data, is the predicted value of the model, N is the number of samples. Only masked patchs are involved in the loss calculation, so that the network training focuses on the prediction of masked patchs. Calculate the mean square error between the original image and the reconstructed image as the loss function, and use the loss function to pre-train the constructed MSP-MAE model.
[0030] As Figure 7 shown, the encoder of the present invention uses the MSP-CT proposed above. Different from MAE, the window masking layer in MSP-MAE does not remove the masked tokens, but directly and uniformly replaces them with learnable vectors, thus keeping the number of tokens unchanged. Deleting these masked tokens will result in not having enough tokens to perform the next downsampling operation (Patch Merging) in the end. Under this strategy of not removing the masked tokens, the relative position encoding is independent in the MSP-CT block, and using both relative and absolute position encodings will not theoretically bring performance improvement. If the masked tokens are not removed, there is no need to add absolute position encoding before the window masking layer. Therefore, there is no need to add a position embedding layer. The only difference between the MSP-MAE encoder and MSP-CT is that in the first stage, this window masking layer is added between the linear embedding layer and the MSP-CT Block, masking 75% of the image area.
[0031] Since only the encoder weights are used in the transfer learning of downstream tasks, the decoder can be flexibly designed. The lightweight decoder design can, on the one hand, reduce the use of computation and memory, which can reduce the training time and allow the use of larger batches. On the other hand, it can also make the training of the network more focused on the encoder, thus producing better transfer results in downstream tasks. As mentioned above, since the reconstruction target of the decoder is the original pixels, as long as the final reconstruction size can be reshaped to the original image size. The decoder design of the present invention is as Figure 8 shown. It only needs three stages to finally reconstruct the encoded features of the encoder to the original image size. In each stage, first perform downsampling (PatchExpanding) on the input to double the features. Patch Expanding consists of a linear layer and a layer normalization. The prediction head of the decoder directly uses the predictor prediction layer to map the dimension of the tokens to 16.
[0032] The principles of the encoder and decoder of MSP-MAE are described; The encoder of MSP-MAE uses the same encoder as MSP-CT. The window masking layer in MSP-MAE does not remove the masked tokens, but directly and uniformly replaces them with learnable vectors, thus keeping the number of tokens unchanged. MSP-MAE does not require adding a position embedding layer. The only difference between the encoder of MSP-MAE and that of MSP-CT is that in the first stage, a window masking layer is added between the Linear Embedding and the encoder, masking 75% of the image area.
[0033] The decoder of MSP-MAE is designed to be more lightweight. And since the reconstruction target of the decoder is the original pixels, as long as the final reconstruction size can be reshaped to the original image size, the decoder of MSP-MAE only has three stages, and each stage contains an MSP-CT block. After passing through these three MSP-CT blocks, the encoded features of the encoder can be reconstructed to the original image size. In each stage, first, Patch Expanding is performed on the input to double the features. Patch Expanding consists of a linear layer and a layer normalization. The predictor head of the decoder directly uses the Predictor Projection layer to map the dimension of the tokens to 16.
[0034] When using the random masking method in MAE, even when the masking rate is 0.75, since the patch size in MSP-MAE becomes 4×4, the area of each patch is very small, resulting in the model still being able to easily use interpolation to reconstruct the masked patches, which easily leads to MSP-MAE learning shortcut solutions. The process of this method is as Figure 9 shown. In the selection of the masking rate, a smaller masking rate will make the task too simple, thus affecting the learning of higher-dimensional semantic features. On the contrary, a larger masking rate will make the task too complex, resulting in underfitting during training, obtaining the effect diagram as Figure 10 shown.
[0035] S4. Obtain the encoder weight parameters in the trained MSP-MAE model, and use the encoder weight parameters to initialize some weight parameters of the MSP-CT model; As Figure 2As shown in the figure, after the MSP-MAE training is completed, since the downstream task mainly uses the feature extraction ability of the encoder, the decoder part can be discarded. Extract the weight parameters of the encoder from the trained MSP-MAE model. These parameters are stored in the various layers of the model in the form of numerical matrices. Initialize some weight parameters of the MSP-CT model using the encoder weight parameters. Analyze the network structures of the MSP-MAE encoder and the MSP-CT model, match the corresponding network layers, and find the corresponding network layers in both. Since the MSP-MAE encoder is improved based on the MSP-CT, there are some similar network layer structures, and the weight parameters of these similar network layers can be migrated.
[0036] Weight parameter assignment: Directly assign the weight parameters of the corresponding layer extracted from the MSP-MAE encoder to the corresponding network layer in the MSP-CT model. For example, the weight parameters of a certain convolutional layer in the MSP-MAE encoder can be directly copied to the convolutional layer with a similar structure in the MSP-CT model. In this way, before the MSP-CT model starts supervised training, some network layers already have the feature extraction ability learned by the MSP-MAE encoder.
[0037] Subsequent training fine-tuning: After completing the weight initialization, use the labeled dataset (such as the MSTAR and FUSAR-Ship datasets) to perform supervised training on the MSP-CT model. During the training process, the model will further adjust the weight parameters according to the labeled data, so that the model can better adapt to the specific SAR target recognition task.
[0038] S5. Fine-tune the parameters of the initialized MSP-CT model and perform supervised training on the MSP-CT model; Transfer some weight parameters of the encoder of the pre-trained MSP-MAE to the encoder of the MSP-CT, and then use the labeled training samples to fine-tune the entire MSP-CT network. During this process, the input image data is not masked, but directly input into the encoder of the MSP-CT; then, after the output features of the MSP-CT encoder pass through the global average pooling layer, the global average features are input into the classifier for the final SAR target classification.
[0039] As Figure 3 shown, in each iteration, use the loss function to update all the parameters of the network through the backpropagation algorithm. Minimize the loss value and optimize the network in this process. Among them, the model uses the cross-entropy loss to compare the predicted probability of each class with the corresponding true label, and the formula is as follows: ,
[0040] Among them, represents the true label of the th sample, represents the predicted probabilities of the model for the th sample belonging to each category, represents the predicted probability that the sample belongs to the i-th category under the condition of the output of the last layer of the given model , is the output result obtained at the last layer after the model undergoes a series of forward propagation calculations. c represents the total number of categories. The multi-scale dilated convolution module (MSDC) proposed in this embodiment consists of the following three parts: a multi-scale convolution fusion module (MSCF), a channel attention module, and a spatial attention module. The combination of these three structures enables the entire MSDC to effectively fuse and process multi-scale information, and none of them can be missing. Next, these three modules will be introduced in detail.
[0041] MSCF: This module uses a convolution with a kernel size of 3×3, and by controlling the dilation rate, convolutions equivalent to those with kernel sizes of 3×3, 5×5, and 7×7 are respectively constructed. These three convolution modules are complementary during the feature fusion process, so as to better depict the local texture information and global contour information of the target, and thus improve the completeness of the feature representation; A branch with a kernel size of 1×1 is also added in this module. The purpose of this branch is to further improve the feature representation with the least number of parameters and computational cost. This to a certain extent improves the problem of discontinuity at the convolution center point, creates conditions for subsequent feature fusion, and plays a boosting role; Then, the features of the four branches are concatenated as follows: , , Among them, is the input feature map, and H and W respectively represent the height and width of the feature map, represents performing a convolution operation on the input feature map using a convolution with a kernel size of 1×1, is a dilated convolution with a kernel size of 3×3, and d is the dilation rate, represents feature concatenation, represents a batch normalization operation, represents the Gaussian error linear unit activation function. After that, the multi-scale features are effectively integrated through a convolution with a kernel size of 1×1 and compressed by four times. After each convolution, a batch normalization layer and a Gelu activation function are added to accelerate the convergence speed of the model and enhance the nonlinear ability of the model. The overall structure of this module is as follows: , Spatial attention: This spatial attention uses two methods, average pooling and max pooling, to aggregate spatial features, and concatenates the two parts of features to obtain a 2×W×H feature map. Then, a convolution with a kernel size of 7×7 is used to compress this feature to obtain a 1×W×H feature map. Finally, the sigmoid function is used to activate this feature map, and the activated feature map is multiplied by the initial fused feature. The spatial attention module can be expressed as: , , where, and represent average pooling and max pooling operations respectively, represents the activation function, represents the spatial attention map, represents the concatenated feature map. As Figure 4 shown, channel attention: The channel attention module used here is the squeeze-and-excitation (SE) attention. SE has two key operations, squeezing and excitation. Among them, squeezing is achieved through global average pooling, which directly compresses the C×W×H feature map containing global information into a C×1×1 feature vector. Excitation is achieved through two fully connected layers. The first fully connected layer compresses the C channels into C / r channels to reduce the computational complexity, and then introduces non-linearity through a Relu. The second fully connected layer restores the number of channels back to C channels, and then obtains the weights through sigmoid activation. Finally, the attention weights obtained above are weighted to the features of each channel. The formula of the SE module is as follows:
[0042] , , where, represents directly compressing the feature map into a C×1×1 feature vector, represents inputting the feature vector obtained through the squeezing operation into the first fully connected layer , represents inputting the activated feature vector into the second fully connected layer , represents the channel attention weight vector. Based on the effective combination of the above three modules, the overall architecture of the MSDC module is described as follows: , where, represents that the MSCF module processes the input feature map . Denotes the output of the spatial attention module. Under the weighting of spatial attention and channel attention, this MSDC module can obtain multi-scale target information, effectively focus on the key information areas of the target, and model the dependencies between these features, which further improves the model's understanding ability of complex scenes and thus enhances the overall performance.
[0043] As Figure 2 shown, MSP-CT uses multi-head self-attention to model the global information of the target. MHSA allows the model to jointly attend to information from different representation subspaces at different positions. Therefore, the MHSA module plays a crucial role in the final performance of this model. In the four stages of MSP-CT, the MHSA module is connected behind the MSDC module. The number of heads of MHSA is set to 4, 8, 16, and 32 in the four stages respectively to further obtain global features, thereby realizing the interaction between local and global information and making the features learned by the model more comprehensive. Specifically, for the input feature , first enhance the input feature through a fully connected layer and obtain the query q, key k, and value v; then, reshape k and perform a dot product with q to generate the attention map ; perform a dot product on A and v to obtain the attention-weighted feature, and finally integrate all features through another fully connected layer. Generally speaking, the attention is defined as: , , , where, is the output feature map, is a learnable scale parameter used to control softmax, denotes processing the input feature through a fully connected layer ( ), denotes the reshaping operation, denotes the softmax function, denotes calculating the attention-weighted feature, denotes another fully connected layer. To achieve multi-head attention, q, k, and v are respectively split into h heads along the feature channel dimension, and the dimension of each head is P. Therefore, separate attention maps can be efficiently learned in parallel.
[0044] As Figure 5 、 Figure 6 shown, FFN consists of two linear layers separated by the Gelu activation. The first layer expands the size by 4 times, and the second layer shrinks the size by the same ratio: , Among them, and represent the weights of two linear layers, and are bias terms, represents the input feature map. As shown in Figure 5 , it is the structural diagram of the MSI-FFN designed by the present invention. Different from the standard FFN, this module consists of two 1×1 convolutional layers and a feature fusion module. The overall structure is similar to an inverted residual block. The multi-scale inverted feed-forward network (MSI-FFN) proposed by the present invention replaces the FFN structure. This module obtains more refined features by fusing multi-scale features to achieve better performance. This module can be expressed as: , , , where the activation layer and batch normalization are omitted, represents that this feature Figure X fuses multi-scale feature information, represents multi-scale feature fusion and residual connection operations, which helps the model to comprehensively understand the input features. This module inserts a 1×1 convolutional branch to expand the channels by 3 times to form a residual connection with the fused features, so as to promote the cross-layer propagation ability of gradients. Finally, a 1×1 convolution is used to compress the overall features.
[0045] All four stages of MSP-CT adopt a structure of alternating stacking of MSDC blocks and MHSA blocks. MSP-CT uses a convolution with a kernel size of 4×4 and a stride of 4 for Linear Embedding. Except for the first stage, in each stage, a Patch Merging is first performed for downsampling. Specifically, the overall architecture is described as follows: , , Among them, is the input image, is the output of the previous stage, is the linear embedding (LinearEmbedding), is the downsampling operation (Patch Merging), is the output of each stage, Is the final output of the model. MHSA represents the multi-head attention block. After the model goes through four stages, the output features are fed into the global average pooling layer, which generates a low-dimensional feature vector by calculating the spatial average of each feature map. This step effectively reduces the number of parameters and computational complexity while retaining important global information. Next, layer normalization is applied to the feature vector after global average pooling to standardize the mean and variance of the input features, which helps alleviate the problem of internal covariate shift, accelerate the convergence of the model, and improve training stability. Subsequently, the normalized feature vector passes through a Dropout layer. Dropout is a regularization technique that reduces overfitting by randomly discarding neurons and enhances the generalization ability of the model. In this stage, a dropout ratio of 0.5 is adopted to ensure that the model can effectively learn robust feature representations during training. Finally, the feature vector processed by dropout is fed into a fully connected layer, and the output result is the classification prediction value of the model.
[0046] S6. Use the fine-tuned MSP-CT model for SAR target recognition, and the output is the recognition result of each test image.
[0047] Use the test set images in the MSTAR and FUSAR-Ship datasets as input data. These test images should maintain the original format and size and have not undergone special processing such as masking during the training stage to simulate the SAR images to be recognized in a real scenario.
[0048] Input the prepared test images into the fine-tuned MSP-CT model one by one. The model will process the input images according to its established network structure.
[0049] Encoder extracts features: The input image first enters the encoder part of MSP-CT. The encoder extracts features from the image through a series of operations such as convolution and attention mechanisms, such as the multi-scale dilated convolution (MSDC) module and the multi-head self-attention (MHSA) module. These modules can capture local and global feature information from the image and convert the input image into a feature representation with rich semantics.
[0050] Global pooling: The features output by the encoder pass through the global average pooling layer. This layer calculates the spatial average of each feature map and converts the high-dimensional feature map into a low-dimensional feature vector. This can reduce the number of parameters and computational complexity while retaining important global information.
[0051] Normalization processing: The feature vector after global pooling passes through the layer normalization operation to standardize the mean and variance of the input features. This helps alleviate the problem of internal covariate shift, accelerate the convergence of the model, and improve training stability.
[0052] Dropout Regularization: The feature vectors after layer normalization will pass through a Dropout layer. Dropout is a regularization technique that randomly discards neurons during training with a certain probability (such as 0.5) to reduce overfitting and enhance the generalization ability of the model. During the testing phase, the role of the Dropout layer is to perform a certain degree of random perturbation on the feature vectors to make the model's predictions more robust. After the above feature processing steps, the feature vectors will be fed into the fully connected layer. The fully connected layer performs linear transformation and classification prediction based on the input feature vectors, and outputs the probabilities of each test image belonging to each category. These probability values represent the confidence of the model in the image belonging to different target categories.
[0053] According to the probability values output by the fully connected layer, the category with the highest probability is selected as the recognition result of the test image. Finally, the target category corresponding to each test image is obtained. These results can be used to evaluate the performance of the model in the SAR target recognition task, such as calculating metrics like accuracy and recall, or can be directly applied to actual SAR image analysis scenarios, such as military reconnaissance and ocean monitoring.
[0054] Experimental Verification: In the method proposed in the present invention, MSF-CT serves as the backbone network of MSF-MAE, and this method is implemented in the Pytorch framework. All experiments use the Adam optimizer. When pre-training MSF-MAE, the masking rate of MSF-MAE is set to 0.75, the batch size is 32, the initial learning rate is 0.0001, and the cosine annealing strategy is used to gradually reduce the learning rate to 0.000001, and a total of 400 epochs are trained. During the classification training phase of MSF-CT, the batch size is 32, the initial learning rate is 0.0001, and the cosine annealing strategy is used to gradually reduce the learning rate to 0.00001, and a total of 200 epochs are trained. In the experiment, each recognition result obtained is the average of five experiments. The experiment was tested and evaluated on a server configured with two RTX A4000s. The operating system of this server is Windows 10 and the memory is 64GB.
[0055] Data Augmentation: To make the experiment more persuasive, the same data augmentation was performed on all comparison methods, including random cropping, random horizontal flipping, and Mixup
[44] . In random cropping, the input image is randomly cropped and then resized to the specified target size of 224×224. The cropping area ratio is randomly selected from [0.6, 1.0], and the aspect ratio of the cropped image varies within the range of [3 / 4, 4 / 3]. Finally, the image is resized to the specified size by bicubic interpolation. The cropping effect is as shown in Figure 14 Figure []. This method can simulate different perspectives and zoom levels, enabling the model to perform better when facing images of different scales and perspectives. In this way, the model is less likely to rely solely on certain fixed parts of the image for prediction. Random horizontal flipping randomly flips the input image horizontally with a probability of 50% to increase the diversity of the data. Mixup generates new samples by mixing two images and their labels in a certain proportion, further increasing the data diversity and the robustness of the model. The effect is as shown in Figure 15 Figure []. The probability of performing the Mixup operation on each image is also 50%. These data augmentation methods are applied to each input sample during the training process, introducing more variability, effectively preventing overfitting, and improving the performance of the model on different datasets. This enables the model to generalize better when facing unseen test data.
[0056] In the reconstruction task of MSF-MAE, the ideal masking rate should find a balance between simplicity and complexity so that the model can effectively extract complex semantic features without lacking challenges due to the task being too simple. In this experiment, by setting the masking rates to 75%, 50%, and 25% and testing their performance in downstream tasks respectively, the most suitable masking rate strategy was found. The above experiments were verified on the MSTAR and FUSAR-Ship datasets respectively. In the experiment on the MSTAR dataset, the recognition situations were verified when the number of training samples for each category was 200, 100, 50, 20, and 10 respectively. In the experiment on the FUSAR-Ship dataset, only the recognition situation under full data volume was verified.
[0057] As shown in Figure 16As shown, the reconstructed structures of the ship targets in OpenSAR-Ship by the model at different masking rates are presented. It can be seen from the figure that as the masking rate decreases, the reconstruction effect becomes closer to the real image. In the case of a 25% masking rate, the model basically reconstructs the detailed structure of the original image. However, from the experimental data in Tables 4 and 5, it can be seen that at this masking rate, the performance of the model in the downstream recognition task is worse than that of the other two masking strategies. This result indicates that a smaller masking rate simplifies the reconstruction task, and the model is exposed to more original information during training, resulting in relatively simple features learned by the model, which can only capture low-dimensional or surface features and cannot deeply understand higher-dimensional semantic features, thus limiting the expressive ability of the model.
[0058] Table 1 Recognition Rates (%) of MSF-MAE at Different Masking Rates on the MSTAR Dataset Result Explanation: It can also be seen from Figure 16 that when the masking rate is 75%, the model can still reconstruct the general outline of the model, but it cannot reconstruct the detailed features of the model. However, in the performance of the downstream task, the recognition rate of the model under this masking rate condition has been greatly improved. It can be seen from Table 4 that as the training data becomes less and less, the recognition performance of the model under the 75% masking rate in the downstream task is better and better than that of the other two masking rate strategies. This indicates that a larger masking rate complicates the task, and the model needs to recover the complete structure from limited data, which enables the model to learn higher-dimensional semantic features from it. From the experiment, it can be concluded that this learned high-dimensional feature is very beneficial to the recognition task.
[0059] Experimental Verification 2: To verify the improvement of MSF-MAE pre-training for the low recognition rate caused by the lack of training data for the model, in this experiment, a certain number of samples are selected from each category of the ten-class vehicle data of MSTAR and the four-class ship data of FUSAR as training samples to test the recognition of each method under different numbers of training sample strategies. The data of MSTAR is the same as the experiment in Table 4 above, and the number of training samples for each category is 200, 100, 50, 20, and 10 respectively; for FUSAR, the number of training samples for each category is 150, 75, 20, and 10 respectively.
[0060] The experimental results are shown in Tables 2 and 3. It can be seen that when there are only 10 labeled samples in each category, the proposed method achieves an accuracy of 74.33% on the MSTAR data and 80.28% on the FUSAR data, which is significantly higher than other methods. For example, the accuracies of the MA method under the same conditions are 70.2% and 74.73% respectively, and the accuracies of DenseNet121 are 57.69% and 78.08% respectively. Similarly, when there are 20 labeled samples in each category, the proposed method achieves accuracies of 87.44% and 84.90% respectively under the two datasets, which are 8.61% and 3.78% higher than those of the MAE method (78.83% and 81.12%) respectively, fully demonstrating the significant advantages of the proposed method under small-sample conditions. Additionally, it is worth noting that on the MSTAR dataset, when there are 200 and 100 labeled samples in each category, the proposed method still performs excellently, achieving high recognition rates of 99.84% and 99.33% respectively. In comparison, the MSF-CT without pre-training can also achieve relatively satisfactory recognition results, reaching 99.63% when there are 200 samples in each category, performing excellently but slightly lower than the proposed method. ResNet18 and DenseNet121 also perform well when the number of samples increases, but are slightly inferior to the proposed method under small-sample conditions. The above experimental results show that compared with other methods, the proposed method performs optimally in all indicators, proving its potential and excellent performance, which also fully demonstrates the advantages of the self-supervised pre-training in this paper. It can learn effective representations from multi-scale features and, through the masking mechanism, enable the model to learn potential feature representations, effectively improving the generalization ability of the model.
[0061] Table 3 Recognition rates (%) of each method under different training sample conditions in the SOC scenario of MSTAR Table 2 Recognition rates (%) of each method under different training sample conditions in the FUSAR dataset Result description: As Figure 17 、 Figure 18 shown, the line charts of the classification accuracies of different methods in the MSTAR and FUSAR datasets are presented. The horizontal axis represents the number of labeled samples in each category, and the vertical axis represents the classification accuracy. The performance curves of eight methods are included in the figure, and the upper right corner is locally enlarged to more detailedly show the accuracy differences of each method. It can be seen from the figure that the accuracy of all methods increases with the increase in the number of samples. In the MSTAR dataset, the accuracy of most methods increases significantly when the number of samples ranges from 10 to 50 and tends to be stable when the number of samples is from 100 to 200.
[0062] Embodiment 2 The difference between this embodiment and Embodiment 1 is that this embodiment provides a self-supervised SAR target recognition system based on the MSP-MAE model, including: A data acquisition module, configured to: acquire SAR image data for training and testing, including acquiring image data through the OpenSAR-Ship, MSTAR, and FUSAR-Ship data sets; A preprocessing module, configured to: perform data augmentation based on the acquired SAR image data, including performing data augmentation using random cropping, random horizontal flipping, and the Mixup method; A model module, configured to: construct an MSP-MAE model, and perform self-supervised training on the MSP-MAE model using the augmented image data to obtain the model weight parameters of MSP-MAE; A conversion module, configured to: obtain the encoder weight parameters in the trained MSP-MAE model, and initialize some weight parameters of the MSP-CT model using the encoder weight parameters; A training module, configured to: perform parameter fine-tuning on the initialized MSP-CT model, and perform supervised training on the MSP-CT model; An output module, configured to: perform SAR target recognition using the fine-tuned MSP-CT model, and the output is the recognition result of each test image.
[0063] A computer-readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded and executed by a processor of a terminal device for the self-supervised SAR target recognition method based on the MSP-MAE model described above.
[0064] A terminal device, including a processor and a computer-readable storage medium, where the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for the self-supervised SAR target recognition method based on the MSP-MAE model described above.
[0065] The above are all preferred embodiments of the present invention. The protection scope of the present invention is not limited hereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention shall be covered within the protection scope of the present invention.
Claims
1. A self-supervised SAR target recognition method based on the MSP-MAE model, characterized in that: include: Acquisition of SAR image data for training and testing, including through the OpenSAR-Ship, MSTAR, and FUSAR-Ship datasets; Perform data enhancement based on the acquired SAR image data, including using random cropping, random horizontal flipping and Mixup methods to enhance data; Construct the MSP-MAE model, and use the enhanced image data to perform self-supervised training on the MSP-MAE model to obtain the model weight parameters of the MSP-MAE; Obtain the encoder weight parameters in the trained MSP-MAE model, and use the encoder weight parameters to initialize some weight parameters of the MSP-CT model; Fine-tune the parameters of the initialized MSP-CT model and perform supervised training on the MSP-CT model; The fine-tuned MSP-CT model is used for SAR target recognition, and the output obtained is the recognition result of each test image.
2. The self-supervised SAR target recognition method based on the MSP-MAE model according to claim 1, characterized in that: The method comprises constructing an MSP-MAE model and performing self-supervisory training on the MSP-MAE model using enhanced image data, including constructing an MSP-MAE model composed of an encoder and a decoder, wherein the MSP-MAE model combines a multi-scale convolution module and a multi-head self-attention module, and uses a window masking strategy to mask the input image, wherein the encoder uses an MSP-CT encoder, and adds a window masking layer between the linear embedding layer of the first stage and the MSP-CT Block.
3. The self-supervised SAR target recognition method based on the MSP-MAE model according to claim 2, characterized in that: The method comprises constructing an MSP-MAE model and performing self-supervisory training on the MSP-MAE model using enhanced image data, wherein the decoder adopts a lightweight decoder composed of multiple stages, wherein each stage first expands the input feature by two times through a Patch Expanding layer composed of a linear layer and a normalization layer, maps the label dimension using a prediction head, and is used to reconstruct the original image according to a latent vector output by the encoder, calculates the mean square error between the original image and the reconstructed image as a loss function for the prediction error of the masked part in the image, and pre-trains the constructed MSP-MAE model using the loss function.
4. The self-supervised SAR target recognition method based on the MSP-MAE model according to claim 3 is characterized in that: The method of using the encoder weight parameters to initialize some weight parameters of the MSP-CT model includes discarding the decoder part after the MSP-MAE model training is completed, extracting the encoder weight parameters from the trained MSP-MAE model, storing the weight parameters in each layer of the encoder in the form of a numerical matrix, determining the corresponding network layer based on the network structure of the MSP-MAE encoder and the MSP-CT model, and assigning the weight parameters of the corresponding layer extracted from the MSP-MAE encoder to the corresponding network layer in the MSP-CT model.
5. The self-supervised SAR target recognition method based on the MSP-MAE model according to claim 4 is characterized in that: The parameter fine-tuning of the initialized MSP-CT model includes inputting complete and unmasked image data into the encoder of the assigned MSP-CT to obtain the output features of the MSP-CT encoder, globally pooling the output features, inputting the pooled feature vector into the classifier for classification, and fine-tuning the MSP-CT using the MSTAR and FUSAR-Ship datasets respectively.
6. The self-supervised SAR target recognition method based on the MSP-MAE model according to claim 5, characterized in that: The MSP-CT model is subjected to supervised training, including a structure in which MSDC layers and MHSA layers are alternately stacked in four stages of MSP-CT, a linear embedding operation is performed through a convolution kernel, and each stage except the first stage is downsampled, wherein the MSDC layer is composed of an MSCF module, a channel attention module, and a spatial attention module, the MSCF module uses a convolution branch to improve the feature representation, the spatial attention module performs feature enhancement through pooling, convolution, and activation operations, and finally the SE module is used through the attention module to obtain the channel weight.
7. The self-supervised SAR target recognition method based on the MSP-MAE model according to claim 6, characterized in that: The supervised training of the MSP-CT model also includes using a multi-scale inverted feedforward network composed of a convolutional layer and a feature fusion layer as a feedforward network, then inserting a convolutional branch to form a residual connection, the output features sequentially pass through a global average pooling layer, a layer normalization, a Dropout layer and a fully connected layer, the fully connected layer outputs a classification prediction value, the cross-entropy loss function is used to compare the predicted probability of each category and the corresponding true label, and finally the MSP-CT model is trained by a back-propagation algorithm.
8. A self-supervised SAR target recognition system based on the MSP-MAE model, characterized in that: include: A data acquisition module is configured to: acquire SAR image data for training and testing, including acquiring image data through OpenSAR-Ship, MSTAR and FUSAR-Ship datasets; The preprocessing module is configured to: perform data enhancement based on the acquired SAR image data, including data enhancement using random cropping, random horizontal flipping and Mixup methods; The model module is configured to: construct an MSP-MAE model, and use the enhanced image data to perform self-supervised training on the MSP-MAE model to obtain the model weight parameters of the MSP-MAE; The conversion module is configured to: obtain the encoder weight parameters in the trained MSP-MAE model, and use the encoder weight parameters to initialize some weight parameters of the MSP-CT model; The training module is configured to: fine-tune the parameters of the initialized MSP-CT model and perform supervised training on the MSP-CT model; The output module is configured to use the fine-tuned MSP-CT model to perform SAR target recognition, and the output obtained is the recognition result of each test image.
9. A computer-readable storage medium storing a plurality of instructions, characterized in that: The instructions are suitable for being loaded by a processor of a terminal device and executing the self-supervised SAR target recognition method based on the MSP-MAE model as claimed in claim 1.
10. A terminal device, comprising a processor and a computer-readable storage medium, wherein the processor is used to implement each instruction; and the computer-readable storage medium is used to store multiple instructions, characterized in that: The instructions are suitable for being loaded by a processor and executing the self-supervised SAR target recognition method based on the MSP-MAE model as claimed in claim 1.
Citation Information
Patent Citations
Image classification method and device in combination with CNN and Transform, and computer storage medium
CN116912552A
Mask enhancement and self-attention mechanism-based unsupervised domain adaptive SAR image water body extraction method, system and equipment and medium
CN116935187A
Power transformation equipment infrared image recognition method based on multi-scale perception and related device
CN117237587A
Gearbox fault diagnosis method based on MSDC-Swinin-T
CN117516925A
Synthetic aperture radar target identification method based on self-supervised learning basic model
CN118155013A
Cited By
Airspace self-supervised learning method and system for rail transit engineering video images
CN120783177A
Spatial self-supervised learning method and system for rail transit engineering video images
CN120783177B