A Small-Sample PolSAR Image Classification Method Based on DBCL-3DDC
Through a multi-level contrast learning and fine-tuning method based on DBCL-3DDC, PolSAR images are classified, which solves the problem of limited performance in small samples and multi-instance tasks, realizes high-precision image classification and reduces labeling costs.
Patent Information
- Application Number
- CN202510396197.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The prior art relies on the high number of tag samples in PolSAR image classification and is not highly targeted for image characteristics, resulting in limited performance in small sample and multi-instance tasks.
Using a DBCL-3DDC-based method, PolSAR images are sliced, and multi-level comparison learning and fine-tuning of comparative branches, mixed branches and classified branches are weakened to reduce the dependence on the number of tag samples and enhance the feature learning ability of multi-instance tasks.
It realizes high-precision PolSAR image classification under very small number of label samples, reducing the time and cost of remote sensing data annotation, and is suitable for the PolSAR field where label samples are scarce.
Smart Images

Figure CN119919815B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of radar remote sensing image interpretation, and more specifically, to a small-sample PolSAR image classification method based on DBCL-3DDC. Background Art
[0002] Synthetic Aperture Radar (SAR) has the advantage of obtaining information all-weather and all-day, and its products have important application values in fields such as the national economy. The early SAR imaging systems were single-frequency and single-polarization, and this imaging method could not fully obtain the key information of the target scattering process contained in the scattering signal, and lacked the description of target scattering characteristics such as phase information and polarization at the same time.
[0003] Polarimetric Synthetic Aperture Radar (PolSAR) is a multi-channel and multi-parameter radar imaging system. Different from optical imaging systems, PolSAR can work stably under various weather conditions and has a strong penetration ability. More importantly, compared with traditional single-polarization SAR imaging systems, in addition to retaining traditional advantages, PolSAR can record four different polarization states by using the polarization characteristics of electromagnetic waves, so as to obtain richer target information and provide more accurate data support for subsequent research and interpretation.
[0004] Remote sensing technology, with its characteristics of wide coverage, strong timeliness, economy and convenience, has shown unique advantages in large-scale monitoring based on remote sensing images and has great application potential in fine classification. At present, the remote sensing classification technology based on PolSAR images aims to accurately divide each pixel in the image into a specific terrain category.
[0005] In recent years, Deep Learning (DL) methods have made breakthrough progress in many fields. It realizes non-linear fitting through a large amount of data-driven methods, avoids complex mathematical modeling and parameter selection, and transfers the computational overhead from the prediction end to the training end, effectively improving the model performance. With the continuous improvement of SAR sensors and the further development of remote sensing technology, DL technology can automatically learn the deep features of images in PolSAR image classification, and perform effective feature extraction and classification, showing its superiority in processing complex data. However, in PolSAR image classification, the existing technologies face the following problems and deficiencies:
[0006] (1) High dependence on the number of labeled samples
[0007] PolSAR images usually have high-dimensional features and complex scattering characteristics. However, compared with optical remote sensing images, the annotation of PolSAR images is not only dense but also requires professional domain knowledge, making it difficult to obtain high-quality labels. Most existing deep learning methods belong to supervised learning, such as CV-CNN-SE (two-dimensional complex-valued network), and the quality of its classification accuracy depends on the number of label samples, resulting in limited performance in small-sample PolSAR image classification tasks.
[0008] (2)Lack of strong pertinence to the characteristics of PolSAR images
[0009] PolSAR image classification belongs to a pixel-level classification task. Usually, a regular rectangular sliding window method is used to cut the area around the central pixel point into image slices as the input of the network model. Therefore, it can be regarded as a special semantic segmentation problem. However, at the category edges or in regions with irregular shapes, the image slices often contain two or even more categories, rich in multiple semantic information, and this characteristic limits the application of the model in multi-instance tasks. Summary of the Invention
[0010] The present invention aims to overcome the defects of high dependence on the number of label samples and lack of strong pertinence to the characteristics of PolSAR images in the above-mentioned existing technologies, and provides a small-sample PolSAR image classification method based on DBCL-3DDC.
[0011] To solve the above technical problems, the technical solution of the present invention is as follows:
[0012] In a first aspect, a small-sample PolSAR image classification method based on DBCL-3DDC includes:
[0013] Slice the pixel points in the PolSAR image to obtain image slices as the input of the DBCL-3DDC network; wherein, the DBCL-3DDC network includes a contrast branch, a hybrid branch, and a classification branch;
[0014] In the contrast branch and the hybrid branch, a random augmentation strategy is adopted to preprocess the image slices to obtain slice augmented views, and the contrast branch and the hybrid branch are pre-trained on the slice augmented views, and the network parameters of the contrast branch and the hybrid branch are updated based on a multi-level contrast learning update strategy; wherein, the contrast branch includes an online network and a target network, and the hybrid branch includes a hybrid network with the same structure as the online network and sharing network parameters;
[0015] Let the classification branch perform fine-tuning based on small-sample classification on the image slices, and update the network parameters of the classification branch;
[0016] The updated DBCL-3DDC network is used for the classification task of the entire PolSAR image.
[0017] In a second aspect, a computer program product includes a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the method described in the first aspect is implemented.
[0018] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0019] The present invention discloses a small-sample PolSAR image classification method based on DBCL-3DDC. The proposed DBCL-3DDC consists of a multi-network structure and multi-level contrast learning. By combining with small-sample PolSAR images, it weakens the model's dependence on the number of labeled samples, enhances the feature learning ability for multi-instance tasks, and fully considers data diversity, reducing the time and cost of remote sensing data annotation. It is especially suitable for the PolSAR field with scarce labeled samples. In addition, the present invention can achieve high-precision small-sample PolSAR classification in both single-temporal and multi-temporal phases, meeting the monitoring requirements in dynamic environments (such as agricultural dynamic monitoring, disaster warning, environmental protection, etc.), and improving the feasibility of actual model applications. Compared with the prior art, the performance of the present invention is more stable when the labeled samples are extremely few, and it has higher classification accuracy. Description of the Drawings
[0020] Figure 1 It is a schematic flowchart of the small-sample PolSAR image classification method in Embodiment 1 of the present application;
[0021] Figure 2 It is another schematic flowchart of the small-sample PolSAR image classification method in Embodiment 1 of the present application;
[0022] Figure 3 It is a schematic structural diagram of the FEM component in Embodiment 1 of the present application;
[0023] Figure 4 It is a schematic structural diagram of the FPM component in Embodiment 1 of the present application;
[0024] Figure 5 It is a schematic structural diagram of the Predictor component in Embodiment 1 of the present application;
[0025] Figure 6 It is a schematic structural diagram of the Classification component in Embodiment 1 of the present application;
[0026] Figure 7 It is the visual classification result of different methods on the Flevoland 1989 dataset in Embodiment 2 of the present application. Detailed implementation manners
[0027] The terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinction adopted when describing objects with the same attributes in the embodiments of this application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device comprising a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices. The term "determine" broadly covers a variety of actions, which may include obtaining, calculating, computing, processing, deriving, researching, looking up (e.g., looking up in a table, database or other data structure), ascertaining, and similar actions, and may also include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory) and similar actions, and may also include generating, creating, establishing and similar actions, as well as parsing, selecting, picking and similar actions, etc. The relevant definitions of other terms will be given in the following description.
[0028] It should be noted that when an element is considered to be "connected" to another element, it can be directly connected to the other element or connected to the other element through an intermediate element. In addition, in the following embodiments, "connection", if there is a transfer of electrical signals or data between the connected objects, should be understood as "electrical connection", "communication connection", etc.
[0029] The drawings are only for illustrative purposes and cannot be construed as a limitation of this patent;
[0030] To better illustrate this embodiment, some components in the drawings are omitted, enlarged or reduced, which does not represent the size of the actual product;
[0031] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0032] The technical solutions of the present invention will be further described below in conjunction with the drawings and embodiments.
[0033] Embodiment 1
[0034] This embodiment provides a few-shot PolSAR image classification method based on DBCL-3DDC (Dense Bootstrap Contrastive Learning Method With 3-Dimensional Dynamic Convolution), referred to as the "DBCL-3DDC method". Referring to FIGS. 1 and 2, it includes:
[0035] Slice the pixel points in the PolSAR image to obtain image slices as the input of the DBCL-3DDC network; wherein, the DBCL-3DDC network includes a contrast branch, a mixing branch, and a classification branch;
[0036] In the contrast branch and the mixing branch, a random augmentation strategy is adopted to preprocess the image slices to obtain slice-augmented views, and the contrast branch and the mixing branch are pre-trained on the slice-augmented views, and the network parameters of the contrast branch and the mixing branch are updated based on a multi-level contrast learning update strategy; wherein, the contrast branch includes an Online network and a Target network, and the mixing branch includes a Mix network with the same structure as the Online network and sharing network parameters;
[0037] Fine-tune the classification branch on the image slices for few-shot classification, and update the network parameters of the classification branch;
[0038] Use the updated DBCL-3DDC network for the classification task of the entire PolSAR image.
[0039] It should be noted that DBCL-3DDC in the above embodiment is composed of a multi-network structure and multi-level contrast learning. It learns global and local representations through global features and local dense features respectively, heuristically extracts the feature representations of multi-instance data, and combines few-shot PolSAR images. Through a small amount of labeled data (only 0.2% of the samples per class), high-precision classification can be achieved, reducing the dependence of the model on the number of labeled samples, demonstrating excellent few-shot learning ability and efficient online learning ability, and effectively solving the problem of poor specificity of PolSAR image characteristics.
[0040] In some preferred embodiments, the Online network It includes an FEM (Feature Extraction Module) component for extracting dense features based on 3DDC (3-Dimensional Dynamic Convolution), an FPM (Feature Projection Module) component for feature modeling, and a Predictor component for prediction;
[0041] The hybrid network includes an FEM component, an FPM component, and a Predictor component;
[0042] The target network includes an FEM component and an FPM component.
[0043] In the above embodiments, dense features of an image are extracted by the FEM containing 3D dynamic convolution (3DDC), and the FPM component is used to model global features and local dense features to enhance the feature representation ability of the network in complex data.
[0044] In some specific implementation processes, each pixel point in a PolSAR image is defined by a sliding window of size, which is called an image slice, and its central pixel is located at position. Therefore, each pixel in the PolSAR image can be converted into an image slice with a size of [16×16×number of channels], and the set of these image slices is used as the input of the FEM module.
[0045] In some alternative embodiments, referring to FIG. 3, the FEM component includes three connected convolutional blocks, and an attention mechanism (Squeeze-and-Excitation, SE) is also provided as a bypass between adjacent convolutional blocks; wherein, the convolutional blocks sequentially include a 3DDC layer, a BN layer, a ReLU layer, and a max pooling layer; and / or,
[0046] Referring to FIG. 4, the FPM component includes a global feature projection sub-branch for outputting global features and a dense feature projection sub-branch for outputting local dense features in parallel; wherein, the global feature projection sub-branch sequentially includes a global average pooling layer, an FC layer, a BN (BatchNorm) layer, a ReLU layer, and an FC layer, and the dense feature projection sub-branch sequentially includes a 3DDC layer, a BN layer, a ReLU layer, a 3DDC layer, and an FC layer.
[0047] In some specific implementation processes, the 3DDC layer in the FEM component uses a 3DDC with a convolution kernel of 3×3×3, denoted as 3DDC-3.
[0048] Further, between adjacent convolutional blocks of the FEM component, an attention mechanism is connected as a bypass between the global average pooling layer of the previous convolutional block and the 3DDC-3 layer of the subsequent convolutional block.
[0049] In some other specific implementation processes, the 3DDC layer in the FPM component uses 3DDC with a convolution kernel of 1×1×1, denoted as 3DDC-1.
[0050] It should be noted that the FEM component takes 3DDC as the core, aiming to retain local dense features to enhance the model's ability to aggregate similar local representations, solve the problem of weak pertinence to the characteristics of PolSAR images in the prior art, and finally, the FEM in each network outputs dense features , which are respectively represented as , and in each network during the pre-training stage, and are respectively used as the inputs to the FPM components in the online network, the target network, and the hybrid network. The FPM consists of two parallel sub-branches (i.e., the global feature projection sub-branch and the dense feature projection sub-branch). The two sub-branches receive the same input, enabling the model to simultaneously learn and preserve global features and local dense features, and finally output the global feature and the local dense feature , which are respectively represented as the first global feature and the first local dense feature in the online network, the second global feature and the second local dense feature in the target network, and the hybrid global feature and the hybrid local dense feature in the hybrid network during the pre-training stage.
[0051] It should also be noted that in the FPM component, the dense feature projection sub-branch removes the global average pooling layer in the global feature projection sub-branch and replaces the FC with a three-dimensional dynamic convolutional layer (such as 3DDC-1). This convolutional layer performs local weighting in the spatial dimension, thereby more flexibly processing feature information of different dimensions. This design enables the model to independently process features in the spatial and channel dimensions, thereby enhancing the ability of feature extraction and combination.
[0052] In the above embodiments, the FEM component retains dense features (also known as "dense features") based on 3D dynamic convolution, and then the FPM component respectively learns global and local representations in a ratio of 30% and 70%, heuristically extracting the feature representation of multi-instance PolSAR images, making up for the problems existing in general contrast learning methods in multi-instance data.
[0053] Furthermore, the 3DDC layer is implemented by conditional parametric convolution, and its convolution kernel is parameterized as a linear combination of 4 experts, with the expression as follows:
[0054]
[0055] In the formula, represents the activation function; each represents a scalar weight related to the input, which is calculated by a routing function with learning parameters:
[0056]
[0057] It should be noted that in the above embodiments, 3DDC is designed to construct the network in the contrast structure, which can improve the feature extraction ability of the network in complex data and reduce the interference of redundant information.
[0058] In some alternative embodiments, as shown in FIG. 5, the Predictor component consists of two layers of FC, ReLU activation function, and BN layer.
[0059] It should be noted that Predictor is an important component in the contrast learning of this embodiment, located at the end of each branch network (i.e., the online network and the hybrid network), and is used to predict the feature representation. Using two layers of FC, ReLU activation function, and BN layer to form Predictor, this design enhances the representation ability in the contrast learning process, so that the model can more effectively extract and compare the feature representations between samples.
[0060] In some specific implementation processes, for the global features and local dense features output by the FPM component in each network, including and and and , the corresponding global feature representations and local dense feature representations are respectively generated by the Predictor component in each network. Specifically, it includes the first global representation and the first local representation in the online network, as well as the hybrid global representation and the hybrid local representation in the hybrid network.
[0061] In some alternative embodiments, the sliced augmented views include the first sliced augmented view and the second sliced augmented view ;
[0062] The comparison branch and the hybrid branch are pre-trained on the slice-enhanced view, including:
[0063] In the comparison branch:
[0064] Using the first slice-enhanced view as the input of the online network to generate the first projection feature and predict the corresponding online network representation; wherein, the first projection feature includes the first global feature and the first local dense feature , and the online network representation includes the first global representation and the first local representation ;
[0065] Using the second slice-enhanced view as the input of the target network and outputting the corresponding second projection feature for providing a regression target for the training of the online network; wherein, the second projection feature includes the second global feature and the second local dense feature ;
[0066] In the hybrid branch:
[0067] Mix the first slice-enhanced view and the second slice-enhanced view to obtain a mixed slice-enhanced view ;
[0068] Using the mixed slice-enhanced view as the input of the hybrid network, and outputting the hybrid network representation; wherein, the hybrid network representation includes the hybrid global representation and the hybrid local representation ;
[0069] In each iteration, when the first projection feature, the second projection feature and the hybrid network representation are determined, based on the multi-level contrast learning strategy, determine the multi-level contrast loss function according to the online network representation and the second projection feature , and update the network parameters of the online network and the target network based on the multi-level contrast loss function .
[0070] It should be emphasized that in the above embodiments, the corresponding global representation and local representation are obtained through the comparison branch and the hybrid branch in the pre-training stage, and then the multi-level contrast learning is used to improve the feature capture ability of the model for multi-instance problems in the pre-training stage.
[0071] It should be noted that most of the existing contrast learning methods focus on self-supervised training based on the global representation of images, often ignoring the relationships between local representations in multi-semantic images. To this end, this embodiment proposes a new local dense contrast loss, which performs multi-level contrast learning by utilizing the spatial information of dense features. Refer to Figure 2. Specifically, the multi-level contrast learning includes global-to-global contrast loss and local-to-local contrast loss, where the global contrast loss is derived from global features and the local contrast loss is derived from local dense features.
[0072] Furthermore, the multi-level contrast loss function includes a global loss and a local dense loss , and its expression is:
[0073]
[0074] Among them, the global loss includes a global contrast loss and a global mixing loss , and its expression is:
[0075]
[0076] In the formula, the global contrast loss is determined based on the first global representation and the second global feature , and the global mixing loss is determined based on the mixed global representation , the first global feature and the second global feature ;
[0077] The local dense loss includes a local dense contrast loss and a local dense mixing loss , and its expression is:
[0078]
[0079] In the formula, the local dense contrast loss is determined based on the first local representation and the second local dense feature , and the local dense mixing loss is based on the mixed local representation , the first local dense feature and the second local dense feature .
[0080] It should be noted that in the above embodiments, the contrast loss function is optimized on the global representation and the local representation, and the two parallel projection heads of FEM and FPM are trained end-to-end.
[0081] Furthermore, the global contrast loss is expressed as:
[0082]
[0083] In the formula, denotes the norm, and
[0084] denotes the inner product; The global hybrid loss
[0085]
[0086] is expressed as: In the formula, denotes the feature vector generated by calculating the maximum value of the -dimensional features and ;
[0087] The local dense contrast loss is expressed as:
[0088]
[0089] The local dense hybrid loss is expressed as:
[0090]
[0091] In the formula, denotes the feature vector generated by calculating the maximum value of the -dimensional features and ;
[0092] In the above embodiments, the global-global and local-local loss functions are established by learning the global features and the local dense features. The loss function aims to optimize the model's ability to capture features of images containing various semantic information, and to more effectively learn and understand the complex characteristics of the data, so as to achieve high-precision PolSAR image classification under the condition of extremely few labeled samples.
[0093] It should be noted that in order to enable the network model to directly compare the most discriminative features, in the above embodiments, for the -dimensional features and Generate a new feature vector by calculating the maximum value of its elements , for -dimensional features and a new feature vector generated by calculating the maximum value of its elements .
[0094] It should be emphasized that in the above embodiments, by simultaneously learning the representations of global features and local dense features, the classification accuracy of downstream PolSAR image tasks can be significantly improved.
[0095] It should also be emphasized that in the above embodiments, a new contrast learning paradigm is designed, which can perform dense contrast learning on local representations, so as to comprehensively learn the instance representations at the global and local levels. This multi-level contrast learning strategy can maximize the use of filtered instance information and enhance information learning from multiple perspectives.
[0096] Furthermore, the network parameter update processes of the online network and the target network are as follows:
[0097]
[0098] In the formula, represents the online network parameters regarding the online network; represents the optimizer; represents the batch; represents the learning rate; represents the target network parameters regarding the target network, and is the exponential moving average EMA (Exponential Moving Average) of the online network parameters ; represents the decay rate of the EMA.
[0099] Furthermore, the random augmentation strategy adopted to preprocess the image slices to obtain slice-augmented views includes:
[0100] Adopt two sets of random augmentation strategies and , and respectively perform random cropping on the image slices at the first ratio and the second ratio , and perform random horizontal flipping with probability to obtain the initial augmented views;
[0101] Adjust the initial augmented views to the image slices to its original size, and the number of channels C remains unchanged throughout the process, to obtain the slice-enhanced view, including the first slice-enhanced view (also referred to as the first randomly augmented view) and the second slice-enhanced view (also referred to as the second randomly augmented view).
[0102] In some specific implementation processes, in the random augmentation strategy, the first ratio , the second ratio .
[0103] In some alternative embodiments, the classification branch includes a FEM component and a Classification component; wherein, the Classification component is responsible for capturing global features, and the goal is to predict the class to which the central pixel point of the image slice belongs;
[0104] The fine-tuning of the classification branch for few-shot classification on the image slice includes:
[0105] Initializing the FEM component in the classification branch with the network parameters corresponding to the FEM component in the pre-trained online network;
[0106] Establishing a few-shot classification training set based on the few-shot image slices with labels, and making the classification branch perform fine-tuning training on the few-shot classification training set, and sequentially outputting the predicted classes of the corresponding central pixel points through the FEM component and the Classification component;
[0107] Updating the network parameters of the classification branch based on the cross-entropy loss between the predicted classes and the labels.
[0108] As a non-limiting example, as shown in FIG. 6, the Classification component sequentially includes a global average pooling layer, a FC layer, a BN layer, a ReLU layer, and another FC layer.
[0109] PolSAR image classification belongs to a pixel-level task, so the spatial information of dense features may interfere with the prediction of sample classes. Therefore, in the few-shot classification stage, the network structure includes a FEM component and a Classification component. In the classification branch, first, a small number of labeled samples jointly act on the FEM component under the guidance of the weight information obtained in the pre-training stage to obtain high-quality feature representations; second, the Classification component is responsible for capturing global features and predicting the classes of sample pixel points; finally, the optimization of the model at this stage is achieved by calculating the cross-entropy loss between the labels and the predicted classes.
[0110] It should be noted that the goal of the Classification component is to predict the category to which the central pixel of the image slice belongs, that is, only the global features of the image slice need to be retained. The image spatial features extracted by the FEM component in the classification branch are input into the Classification component, and the high-dimensional feature map is converted into low-dimensional global features through global average pooling. Subsequently, the global features are finally classified through the FC, ReLU activation function, and BN layer.
[0111] It should also be noted that on a small number of unlabeled samples, after the pre-training stage, the model will have the ability to aggregate the features of similar instances and generate a pre-trained weight file of the model, which can be used to initialize the FEM component in the classification branch; in the few-shot classification fine-tuning stage, a very small number of image slices that have not undergone data preprocessing are input into the FEM component, and the model is fine-tuned jointly with the Classification component. Through the fine-tuning of a small number of samples, the model can be quickly updated and adapted to the semantic information of different categories, which makes this embodiment particularly suitable for real-time monitoring tasks, such as agricultural dynamic monitoring, disaster warning, and other scenarios.
[0112] In some specific implementation processes, after the DBCL-3DDC network completes fine-tuning and updating, its classification branch (i.e., the network model composed of the FEM component and the Classification component) is applied to perform the classification task of the entire PolSAR image.
[0113] Exemplarily, by applying the method described in the above embodiment, through the classification and dynamic monitoring of crops, the application of precision agriculture is realized, such as the monitoring of crop growth.
[0114] Exemplarily, by applying the method described in the above embodiment, the change of land cover type is monitored, and the impact of environmental change on the ecosystem is evaluated.
[0115] Exemplarily, by applying the method described in the above embodiment, through ground object classification and change detection, support is provided for observation and planning.
[0116] Embodiment 2
[0117] This embodiment conducts a comparative experiment based on the Flevoland 1989 dataset (L-band). 5% of unlabeled samples and 0.2% of labeled samples are randomly selected from the constructed image dataset at a certain ratio. These samples are used as the pre-training dataset and the few-shot classification training dataset of the DBCL-3DDC method proposed in Embodiment 1, respectively.
[0118] Flevoland 1989 dataset (L-band): This data is the fully polarized airborne SAR data publicly available from NASA / JPL Laboratory. It was collected in 1989, and the study area is located in the Flevoland region of the Netherlands. The image size is . It contains 15 different land cover classes, namely Stem Beans, Peas, Rapeseed, Beet, Forest, Lucerne, Wheat 1, Bare Soil, Grass, Water, Barley, Wheat 2, Wheat 3, Potatoes, and Buildings. This image is widely used for the performance verification of PolSAR image classification methods.
[0119] During the experiment, the data processing and model construction are as follows:
[0120] 1) Data processing: The input data is a PolSAR image. Image patches of size [16×16×number of channels] corresponding to each pixel are constructed using a sliding window.
[0121] 2) Model training (pre-training stage): The Dense-guided Contrastive Learning (DBCL) framework is adopted to train the pre-training dataset.
[0122] 3) Model fine-tuning (few-shot classification stage): The FEM and classification module are adopted to fine-tune the few-shot classification training set.
[0123] 4) Model testing: In the testing stage, the full-image data is classified to evaluate the classification accuracy of the model.
[0124] The existing technologies for comparison are: (a) CV-CNN (Complex-Valued Convolutional Neural Network), (b) CV-CNN-SE (Two-Dimensional Complex-Valued Network), (c) RV-CNN (Real-Valued Convolutional Neural Network), (d) SSPRL (Self-Supervised PolSAR Representation Learning), (e) SDF2Net (Shallow-to-Deep Feature Fusion Network), (f) HybridCVNet (Hybrid Complex-Valued Network), and (g) 3D-CNN (Three-Dimensional Convolutional Network).
[0125] The results of the comparative experiments are shown in Table 1 and Figure 7.
[0126] Table 1 Classification results of different methods on the Flevoland 1989 dataset
[0127]
[0128] In Figure 7, (a) shows the visualization classification results of the CV-CNN method, (b) shows the visualization classification results of the CV-CNN-SE method, (c) shows the visualization classification results of the RV-CNN method, (d) shows the visualization classification results of the SSPRL method, (e) shows the visualization classification results of the SDF2Net method, (f) shows the visualization classification results of the HybridCVNet method, (g) shows the visualization classification results of the 3D-CNN method, and (h) shows the visualization classification results of the DBCL-3DDC (Ours) method. Among them, the blue rectangular box highlights the area with dense categories, the red rectangular box highlights the area with balanced numbers of adjacent categories, while the white rectangular box emphasizes the area with unbalanced numbers of adjacent categories.
[0129] Combined with Table 1 and Figure 7, it can be seen that the proposed DBCL-3DDC method in Example 1 can achieve high-precision classification (97.29%) with an extremely small number of labeled samples (0.2%), significantly superior to existing deep learning methods, and the classification accuracy under small sample conditions has increased by approximately 15%. This is due to the fact that DBCL can perform dense contrast learning on local representations, thereby comprehensively learning instance representations at both the global and local levels. In addition, due to the introduction of local dense features, multi-level contrast learning, and 3DDC, the model can effectively extract key features of PolSAR images, heuristically perform dense classification tasks on multi-instance PolSAR images, has stronger small sample learning ability, and significantly reduces the time and cost of remote sensing data annotation, especially suitable for the PolSAR field where labeled samples are scarce.
[0130] It can be understood that the optional items in the above Example 1 also apply to this example, so they will not be described repeatedly here.
[0131] Example 3
[0132] This example provides a computer-readable storage medium, on which at least one instruction, at least one program, a code set, or an instruction set is stored. The at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by a processor, so that the processor executes some or all of the steps of the method provided in Example 1 or 2 of this application.
[0133] It can be understood that the storage medium can be transient or non-transient. Exemplarily, the storage medium includes, but is not limited to, various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0134] Exemplarily, the processor may be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or the like.
[0135] Exemplarily, the read-only memory includes, but is not limited to, MASK ROM, PROM, EPROM, EEPROM, Flash, etc.
[0136] Exemplarily, the random access memory includes, but is not limited to, DRAM, SRAM, SDRAM, DDR SDRAM, etc.
[0137] In some examples, a computer program product is provided, which may be implemented specifically by means of hardware, software, or a combination thereof. As a non-limiting example, the computer program product may be embodied as the storage medium, and may also be embodied as a software product, such as an SDK (Software Development Kit), etc.
[0138] As a non-limiting example, a computer program product is provided, which includes a computer program or computer-executable instructions, and the computer program or computer-executable instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes some or all of the steps of the method described in the embodiments of the present application.
[0139] In some examples, a computer program is provided, including computer-readable code. When the computer-readable code runs on a computer device, the processor in the computer device executes some or all of the steps for implementing the method.
[0140] This embodiment also proposes an electronic device, including a memory and a processor. The memory stores at least one instruction, at least one program, a code set, or an instruction set. When the processor executes the at least one instruction, at least one program, the code set, or the instruction set, some or all of the steps of the method described in Embodiment 1 or 2 are implemented.
[0141] In some examples, a hardware entity of the electronic device is provided, including: a processor, a memory, and a communication interface; wherein, the processor generally controls the overall operation of the electronic device; the communication interface is used to enable the electronic device to communicate with other terminals or servers through a network; the memory is configured to store instructions and applications executable by the processor, and can also cache data to be processed or already processed by the processor and each module in the electronic device (including but not limited to image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or random access memory (RAM).
[0142] The processor may include one or more processing elements. Therefore, the processor may include one or more integrated circuits (ICs) configured to execute the functions of the processor. In addition, each integrated circuit may include circuits (e.g., a first circuit, a second circuit, and other circuits, etc.) configured to execute the functions of the processor.
[0143] Furthermore, data transmission may be carried out among the processor, the communication interface, and the memory through a bus, and the bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together.
[0144] It can be understood that the optional items in the above-mentioned Embodiment 1 or 2 are equally applicable to this embodiment, so they will not be described repeatedly here.
[0145] The same or similar reference numerals correspond to the same or similar components;
[0146] The terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be construed as a limitation to this application;
[0147] It should be noted that, without conflict, the embodiments and features in the embodiments of this application may be combined with each other.
[0148] In different specific implementations, the method or system described in this application may be implemented in software, hardware, or a combination thereof. In addition, the order of the steps of the method may be changed, and various elements may be added, reordered, combined, omitted, modified, etc.
[0149] Obviously, the above embodiments of the present application are merely examples for clearly illustrating the present application, rather than limitations on the implementation manners of the present application, and are not used to limit the present application. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. Each discrete structure / functional module or unit can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part. The structures and functions of the discrete components can be implemented as a combined structure or component. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the claims of the present application.
Claims
1. A small sample PolSAR image classification method based on DBCL-3DDC, characterized in that: include: Slicing the pixels in the PolSAR image to obtain image slices as input of a DBCL-3DDC network; wherein the DBCL-3DDC network includes a comparison branch, a mixing branch, and a classification branch; In the comparison branch and the hybrid branch, a random enhancement strategy is adopted to preprocess the image slice to obtain a slice enhancement view, the comparison branch and the hybrid branch are pre-trained on the slice enhancement view, and the network parameters of the comparison branch and the hybrid branch are updated based on a multi-level contrast learning update strategy; wherein the comparison branch includes an online network and a target network, and the hybrid branch includes a hybrid network with the same structure as the online network and shared network parameters; Allowing the classification branch to perform fine-tuning based on small sample classification on the image slice, and updating the network parameters of the classification branch; The updated DBCL-3DDC network is used for the classification task of the entire PolSAR image; Wherein, the online network f θ It includes FEM component for extracting dense features based on 3DDC, FPM component for feature modeling and Predictor component for prediction; The hybrid network includes a FEM component, a FPM component and a Predictor component; The target network f ξ Including FEM components and FPM components; The FEM component includes three connected convolution blocks, and an attention mechanism is also provided between adjacent convolution blocks as a bypass; wherein the convolution block includes a 3DDC layer, a BN layer, a ReLU layer and a maximum pooling layer in sequence; The FPM component includes a parallel global feature projection sub-branch for outputting global features and a dense feature projection sub-branch for outputting local dense features; wherein the global feature projection sub-branch includes a global average pooling layer, an FC layer, a BN layer, a ReLU layer and an FC layer in sequence, and the dense feature projection sub-branch includes a 3DDC layer, a BN layer, a ReLU layer, a 3DDC layer and an FC layer in sequence.
2. The small sample PolSAR image classification method based on DBCL-3DDC according to claim 1 is characterized in that: The 3DDC layer is implemented by conditional parameterized convolution, whose convolution kernel is parameterized as a linear combination of four experts, expressed as follows: Output(x)=σ(α1(W1*x)+α2(W2*x)+α3(W3*x)+α4(W4*x)) In the formula, σ represents the activation function; each α i =γ i (x) represents a scalar weight associated with the input, which is calculated by a routing function with learning parameters: γ i (x)=Sigmoid(FC(GlobalAveragePool(x)))。 3. A small sample PolSAR image classification method based on DBCL-3DDC according to any one of claims 1-2, characterized in that: The slice enhancement view comprises a first slice enhancement view x i and the second slice enhanced view x j ; The contrast branch and the mixing branch are pre-trained on the slice enhanced view, including: In the comparison branch: Enhance the view x with the first slice i For the online network θ The first projection feature is generated and the corresponding online network representation is predicted; wherein the first projection feature includes the first global feature z i and the first local dense feature The online network representation includes a first global representation q i and the first local representation Enhance the view x with the second slice j The second projection feature is input to the target network and outputs the corresponding second projection feature, which is used to provide a regression target for the training of the online network; wherein the second projection feature includes the second global feature z j and the second local dense feature In the hybrid branch: Enhance the view x for the first slice i With the second slice enhanced view x j Mix and get the mixed slice enhanced view x m ; Enhanced view with mixed slices x m As the input of the hybrid network, a hybrid network representation is output; wherein the hybrid network representation includes a hybrid global representation q m and hybrid local representation In each iteration, after determining the first projection feature, the second projection feature and the hybrid network representation, a multi-level contrast loss function L is determined based on the online network representation and the second projection feature based on a multi-level contrast learning strategy. total , and based on the multi-level contrast loss function L total Update the online network θ With the target network f ξ network parameters.
4. The small sample PolSAR image classification method based on DBCL-3DDC according to claim 3 is characterized in that: The online network θ With the target network f ξ The network parameter update process is as follows: In the formula, θ represents the online network parameters of the online network; O represents the optimizer; B represents batch; η represents the learning rate; ξ represents the target network parameters of the target network, and ξ is the exponential moving average EMA of the online network parameters θ; τ represents the decay rate of EMA.
5. The small sample PolSAR image classification method based on DBCL-3DDC according to claim 4 is characterized in that: The multi-level contrast loss function L total Including the global loss L global and the local dense loss L dense , whose expression is: THE total =(1-λ)L global +λL dense Among them, the global loss L global Including global contrast loss and global hybrid loss Its expression is: In the formula, the global contrast loss Based on the first global representation q i and the second global feature z j Determine that the global mixing loss Based on the hybrid global representation q m , the first global feature z i and the second global feature z j Sure; The local density loss L dense Including local dense contrast loss and local dense mixing loss Its expression is: In the formula, the local dense contrast loss Based on the first local representation and the second local dense feature Determine that the local dense mixing loss Based on hybrid local representation The first local dense feature and the second local dense feature 6. The small sample PolSAR image classification method based on DBCL-3DDC according to claim 5, characterized in that: The global contrast loss The expression is: In the formula, ‖·‖ represents the l2 norm, and <·> represents the inner product; The global mixing loss The expression is: In the formula, z r It means that by calculating the d-dimensional feature z i and z j The eigenvector generated by the maximum value of the element, The local dense contrast loss The expression is: The local dense mixing loss The expression is: In the formula, Indicates that by calculating the d-dimensional features and The eigenvector generated by the maximum value of the elements.
7. The small sample PolSAR image classification method based on DBCL-3DDC according to claim 3 is characterized in that: The step of adopting a random enhancement strategy to preprocess the image slice to obtain a slice enhancement view includes: Two sets of random enhancement strategies V and V′ are used to slice the image Perform the first ratio P h and the second ratio P w The random cropping with probability P f Perform random horizontal flipping to obtain the initial enhanced view; The initial enhanced view is adjusted to the original size of the image slice x, and the number of channels C remains unchanged during the whole process, to obtain the slice enhanced view, including the first slice enhanced view x i and the second slice enhanced view x j .
8. A small sample PolSAR image classification method based on DBCL-3DDC according to any one of claims 1-2, characterized in that: The classification branch includes a FEM component and a Classification component; wherein the Classification component is responsible for capturing global features, and the goal is to predict the category to which the central pixel of the image slice belongs; The step of fine-tuning the classification branch on the image slice based on small sample classification includes: Initializing the FEM components in the classification branch using the pre-trained network parameters corresponding to the FEM components in the online network; Establish a small sample classification training set based on the small sample image slices with labels, make the classification branch perform fine-tuning training on the small sample classification training set, and output the predicted category of the corresponding central pixel point through the FEM component and the Classification component in turn; The network parameters of the classification branch are updated based on the cross entropy loss between the predicted category and the label.
Citation Information
Patent Citations
Semi-supervised polarimetric SAR image classification method based on multi-branch network
CN113869136A
Small sample remote sensing image classification method and system based on multi-source domain self-attention
CN115019104A