A Semi-Supervised Efficient Video SAR Shadow Tracking Method Based on Small Samples
Through the small sample semi-supervised learning method, combined with deep learning and semi-supervised learning technology, the data dependence problem of shadow tracking algorithm in small sample scenarios is solved, and efficient and accurate tracking of shadowed areas in SAR images is achieved.
Patent Information
- Application Number
- CN202411880886.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-12-19
AI Technical Summary
In the prior art, in small sample scenarios, shadow tracking algorithms rely on a large amount of labeled data, which makes it difficult to identify and track shadow areas, especially in SAR images, signal attenuation makes it difficult to accurately process low-brightness areas.
Using a semi-supervised learning method based on small samples, combined with deep learning and semi-supervised learning technology, by optimizing the model structure and introducing advanced semi-supervised learning strategies, the tracking accuracy and robustness of shadowed areas are improved by utilizing limited annotated samples and a large number of unlabeled samples.
In the case of reducing labeled data, the detection accuracy and robustness of shadowed areas are significantly improved, and the learning efficiency and generalization ability of the model are improved.
Smart Images

Figure CN119904771B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of radar target recognition, and particularly to a semi-supervised efficient video SAR shadow tracking method based on small samples. Background Art
[0002] With the rapid development of synthetic aperture radar (SAR) technology, SAR images are increasingly widely used in fields such as surface monitoring, disaster assessment, and military reconnaissance. Especially in complex environments, accurate tracking of shadow areas is crucial for target recognition and ground object classification. However, in SAR images, shadow areas usually suffer from signal attenuation due to imaging geometry and reflection characteristics, presenting as low-brightness areas, which poses challenges to subsequent target detection and classification tasks. Therefore, accurately identifying and tracking shadow areas has become the key to improving the performance of SAR image processing.
[0003] Currently, most shadow tracking algorithms rely on a large amount of labeled data for training, which becomes particularly difficult in small-sample scenarios. Therefore, how to effectively utilize limited labeled samples and a large amount of unlabeled samples to improve the accuracy of shadow tracking has become an urgent problem to be solved.
[0004] In recent years, semi-supervised learning techniques have gradually attracted the attention of researchers because they can combine the advantages of labeled and unlabeled data. At the same time, image processing techniques based on semi-supervised learning have made significant progress in multiple fields. However, methods for shadow tracking in video SAR images are still relatively scarce, especially in application scenarios under small-sample conditions. Summary of the Invention
[0005] To solve the problem that existing traditional shadow tracking algorithms rely on a large amount of labeled data for training, this application provides a semi-supervised efficient video SAR shadow tracking method based on small samples. By optimizing the model structure and introducing advanced semi-supervised learning techniques, the tracking accuracy and robustness of shadow areas are improved to meet the actual application requirements in various complex environments.
[0006] The technical solution adopted in this application is as follows:
[0007] A semi-supervised efficient video SAR shadow tracking method based on small samples, comprising the following steps:
[0008] Step 1: Obtain SAR images and perform image preprocessing on them (to adapt to the set semantic segmentation model, such as the DeepLabV3+ model), and obtain an original dataset based on several SAR images after image preprocessing; wherein, the image preprocessing includes: image binarization processing, and image size normalization processing on the binarized image.
[0009] In this step, the binarization process is as follows: set the pixel points of the background to 0 and the pixel points of the target to 1. The unified image size can be set to 512×512.
[0010] Step 2: Divide the original dataset into a labeled original dataset and an unlabeled original dataset.
[0011] Step 3: Divide the labeled original dataset into a first training set and a first validation set. Then, based on the first training set, perform supervised training on the set semantic segmentation model to optimize the network parameters of the model (such as weights and bias terms), and tune the hyperparameters of the semantic segmentation model based on the first validation set. The loss function used during supervised training is a consistency loss, such as cross-entropy loss.
[0012] Among them, the semantic segmentation model is used to identify multiple targets and perform segmentation. Its output layer is used to perform N+1 classification and segmentation on the input data, that is, it includes N required target categories and one unknown category.
[0013] Step 4: Perform semi-supervised training on the supervised-trained semantic segmentation model based on the unlabeled original dataset.
[0014] Perform image weak augmentation processing on each unlabeled original data sample in the unlabeled original dataset to obtain at least one weakly augmented sample for each unlabeled original data sample.
[0015] Input the weakly augmented sample into the supervised-trained semantic segmentation model, and obtain the pseudo-label of the current weakly augmented sample based on the predicted category output by it.
[0016] Perform image strong augmentation processing on each unlabeled original data sample in the unlabeled original dataset to obtain at least one strongly augmented sample for each unlabeled original data sample.
[0017] For each unlabeled original data sample j, obtain a pair of weak-strong sample pairs corresponding to j based on one weakly augmented sample and one strongly augmented sample.
[0018] Based on the labeled original dataset and all weak-strong sample pairs, form a mixed set, and then use this mixed set as the second training set to perform secondary training on the supervised-trained semantic segmentation model.
[0019] The total loss function used during secondary training is: the weighted sum of the first consistency loss corresponding to the labeled original dataset and the second consistency loss based on all weak-strong sample pairs. Among them, for the second consistency loss, the strongly augmented sample in the weak-strong sample pair is used as the model input, and the pseudo-label of the corresponding weakly augmented sample is regarded as the true label.
[0020] Step 5: Extract a fine-tuning dataset from the mixed set to fine-tune the sematic segmentation model after secondary training. The loss function used during fine-tuning is the consistency loss. In this fine-tuning dataset, the weak and strong sample pairs are regarded as one model input sample respectively, and their corresponding pseudo-labels are regarded as the true labels of each sample.
[0021] Based on the fine-tuned semantic segmentation model, a segmentation model for SAR shadow tracking is obtained, and based on this segmentation model, the shadow area prediction of the SAR image to be tracked is realized. That is, it is necessary to preprocess the SAR image to be tracked to adapt to the input of the semantic segmentation model.
[0022] Furthermore, step 4 further includes:
[0023] Regard the pseudo-labels with prediction confidence (output by the semantic segmentation model) greater than the preset threshold as reliable pseudo-labels.
[0024] For each original data sample with a reliable pseudo-label, a weak-strong sample pair of the current original data sample j is formed based on a strongly augmented sample and a weakly augmented sample corresponding to the reliable pseudo-label.
[0025] Furthermore, the weak-strong sample pairs in the fine-tuning dataset extracted in step 5 are weak-strong sample pairs with reliable pseudo-labels.
[0026] Furthermore, in step 2, the ratio of labeled data to unlabeled data is 2:8 to form a small sample dataset.
[0027] Furthermore, in step 3, the ratios of the first training set and the first validation set are 80% and 10% respectively.
[0028] Furthermore, the semantic segmentation model uses the DeepLabV3+ model, its backbone network is MobileNet, the optimizer selects SGD (stochastic gradient descent), and the learning rate decay method based on the set initial learning rate, minimum learning rate lr min 、maximum learning rate lr max and the total number of training epochs T is the cosine annealing method.
[0029] Furthermore, in step 4, the weak image enhancement processing includes: randomly flipping the data (e.g., with a probability of 50%), randomly cropping, and randomly rotating (e.g., by 10°).
[0030] Furthermore, in step 4, the strong image enhancement processing includes: randomly occluding the image, generating a new image by linearly combining two images in a specified ratio, and adding noise.
[0031] Further, in step 4, the second consistency loss is the categorical cross-entropy loss with the strongly augmented image as the input and the corresponding pseudo-label / reliable pseudo-label regarded as the true label.
[0032] Further, in step 4, the total loss function is: Total Loss = Ls + λ * Lu, where Ls and Lu are the first and second consistency losses respectively, and λ is a preset weight coefficient representing the contribution of the unlabeled loss to the total loss. Preferably, it can be set to λ = 0.67.
[0033] Further, step 5 further includes calculating the intersection over union between the shadow region predicted by the fine-tuned semantic segmentation model and the true shadow region. When it is greater than 0.5, it is regarded as successful SAR shadow tracking. Furthermore, the tracking rate can be calculated to evaluate whether the model achieves the goal.
[0034] The technical solution provided by this application at least brings the following beneficial effects:
[0035] This application can reduce the labeled data while improving the accuracy of shadow tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The above and / or additional aspects and advantages of this application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0037] Figure 1 is a flowchart of a semi-supervised efficient video SAR shadow tracking method based on small samples provided by an embodiment of this application;
[0038] Figure 2 is a schematic diagram of SAR shadow tracking in an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this application will be described in detail and completely in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described by referring to the drawings are exemplary and are intended to explain this application, and should not be construed as a limitation to this application.
[0040] The embodiment of the present application provides a semi-supervised efficient video SAR shadow tracking method based on small samples. By combining deep learning and semi-supervised learning in a new way under small sample conditions, it solves the limitations of traditional shadow tracking techniques. Through an improved segmentation model (such as the DeepLabV3+ model) and a semi-supervised strategy (such as the FixMatch strategy), it improves the detection accuracy and robustness of the shadow area, and based on the designed new hybrid training strategy, effectively integrates labeled and unlabeled data, significantly improving the learning efficiency and generalization ability of the model. By combining deep learning technology and semi-supervised learning strategies, it makes full use of limited labeled data and rich unlabeled data to solve the deficiencies of traditional methods under small sample conditions, and effectively improves the detection and tracking accuracy of shadow areas in SAR images.
[0041] In one embodiment, referring to Figure 1 , a semi-supervised efficient video SAR shadow tracking method based on small samples provided by the present application includes the following steps:
[0042] S1. Obtain a SAR remote sensing image dataset and perform image preprocessing on it to match the input of the segmentation model; wherein, the segmentation model is used to indirectly identify multiple targets and perform segmentation by distinguishing different categories, and classify the targets, that is, the segmentation model is used to identify and segment multiple targets.
[0043] In this embodiment, the image preprocessing of each SAR image in the obtained dataset includes: setting the pixel points of the background to 0 and the pixel points of the target to 1. In addition, all dataset images (composed of several obtained SAR remote sensing images) are cut into pixel blocks of 512×512 size to ensure adaptability to network training. At the same time, the data is enhanced by random cropping, rotating by 15°, and horizontal flipping to increase the diversity of samples, thereby enhancing the generalization ability of the model. In this embodiment, the segmentation model uses an improved DeepLabV3+ model.
[0044] S2. Divide the original dataset (the dataset composed of each segmented pixel block) into labeled data and unlabeled data;
[0045] In this embodiment, the original dataset is divided into labeled data and unlabeled data. Among them, the labeled data is divided into a training set, a validation set, and a test set according to a ratio of 80%:10%:10% to ensure the representativeness of each subset. The ratio of labeled data to unlabeled data is 2:8 to construct a small sample dataset.
[0046] Among them, the training set is the data samples for model fitting; the validation set is the sample set set aside separately during the model training process, which is used to adjust the hyperparameters of the model and to preliminarily evaluate the capabilities of the model. It is usually used during iterative model training to verify the generalization ability (accuracy, recall, etc.) of the current model to decide whether to stop further training. The test set is used to evaluate the generalization ability of the final model (the trained one). However, it cannot be used as the basis for algorithm-related selections such as tuning parameters and feature selection.
[0047] S3. The training phase is divided into supervised training and semi-supervised training. Use the labeled dataset to train the DeepLabV3+ model, set relevant parameters, and obtain the consistency loss Ls of this dataset.
[0048] In this embodiment, in the fully-supervised phase, the DeepLabV3+ model is trained using labeled data, and the relevant parameters of the model are configured. The backbone network of this model selects MobileNet, the number of classes to be distinguished is 6, the downsampling multiple is set to 8, the optimizer selects SGD (stochastic gradient descent), the initial learning rate is set to 0.02, and the weight decay is set to 0.0005 to prevent overfitting and increase the generalization ability of the model. The total number of training epochs T is set to 500, the minimum learning rate is 0.002, the maximum learning rate is 0.02, and the learning rate decay method adopts the cosine annealing method (Cosine Annealing):
[0049]
[0050] where t is the current training epoch, T is the total number of training epochs, lr min represents the minimum learning rate, and lr max represents the maximum learning rate.
[0051] The cross-entropy loss Ls of the test set and the validation set is obtained through training, and its calculation formula is:
[0052]
[0053] where y i is the true label, is the model prediction output, N is the total number of samples, and i is the sample number.
[0054] Retain the best model weights, and use the test set to verify the model to ensure the generalization ability of the model on unseen data. That is, the segmentation model is fully-supervised trained based on the training set and the validation set, and the loss function during the fully-supervised training process is the cross loss.
[0055] Then, perform weak augmentation on the unlabeled data (there is at least one weakly augmented image for each original image), and use the trained model to train this set of data to generate pseudo-labels. At the same time, set a threshold. When the prediction confidence of the model for the data is greater than the threshold, trust this set of pseudo-labels, that is, the pseudo-labels greater than the threshold are reliable pseudo-labels. Then, perform strong augmentation on the same set of unlabeled data (at least one strong augmentation result), that is, perform image strong augmentation on the original images with reliable pseudo-labels, and use this network to train the augmented unlabeled data to obtain the consistency loss. For each original image with a reliable pseudo-label, a weak-strong sample pair of the current original image can be formed based on one strongly augmented image (if there are multiple, randomly select one) and a weakly augmented image corresponding to a reliable pseudo-label.
[0056] Among them, in this embodiment, the weak augmentation image processing is specifically set as follows: perform random flipping (with a probability of 50%), random cropping, and random rotation (10°) on the data, and then use the trained DeepLabV3+ model to predict these weakly augmented images to obtain pseudo-labels. At the same time, set a confidence threshold of 0.95 to be used for screening the reliability of the pseudo-labels.
[0057] The calculation formula for the confidence is:
[0058]
[0059] Among them, is the class probability predicted by the model.
[0060] Only when the prediction confidence of the current model for these unlabeled samples is greater than this threshold, the pseudo-labels of this data are considered reliable.
[0061] Perform RandAugment strong augmentation on the same set of unlabeled data. Set the number of transformations to 2, the intensity parameter to 7, randomly occlude 25% of the image, linearly combine two images at a ratio of 0.5 to generate a new image. In addition, add random noise to the image, and set the noise standard deviation to 0.13. Also use the trained DeepLabV3+ model to predict the strongly augmented images to obtain the prediction results of this set of data.
[0062] S4. Calculate the consistency loss L between the prediction results of the strongly augmented images and the pseudo-labels of the corresponding weakly augmented images u ;
[0063] At this time, the strongly augmented samples in each weak-strong sample pair can be used as the input of the semantic segmentation model, and the reliable pseudo-labels of the corresponding weakly augmented samples are regarded as the true labels of the strongly augmented samples. Based on the prediction output of the classification output by the model, the corresponding consistency loss can be calculated, denoted as L u , and the classification is such as cross-entropy loss.
[0064] In this embodiment, the calculation formula of the cross-entropy loss Lu between the two is as follows:
[0065]
[0066] where y j is the pseudo-label of the weakly augmented image, is the predicted output of the strongly augmented image, M is the total number of samples (corresponding unlabeled data), j represents the number of the unlabeled data (specifically, it can refer to the number of unlabeled data with reliable pseudo-labels), and weak augmentation and strong augmentation can be set in advance.
[0067] S5. Combine the unlabeled data and the labeled data to form a new mixed dataset, and then combine the loss Ls of the labeled data obtained in S3 and the consistency loss Lu of the unlabeled data to form a total loss function to update the DeepLabV3+ model.
[0068] Based on the labeled original dataset and all weak-strong sample pairs to form a mixed set, and then use this mixed set as the second training set to perform secondary training on the supervised-trained semantic segmentation model;
[0069] The total loss function used in the secondary training is: the weighted sum of the first consistency loss corresponding to the labeled original dataset and the second consistency loss based on all weak-strong sample pairs; among them, for the second consistency loss, the strongly augmented sample in the weak-strong sample pair is used as the model input, and the reliable pseudo-label of the corresponding weakly augmented sample is regarded as the true label;
[0070] In this step, it is also possible to directly take one strongly augmented and one weakly augmented data for each unlabeled data as a pair of data, and use the strongly augmented data in this pair as the model input, and regard the pseudo-label of the weakly augmented data as the true label to calculate the consistency loss Lu of all unlabeled data.
[0071] In this embodiment, the unlabeled data and the labeled data are combined to form a new mixed dataset, and then the loss Ls of the labeled data obtained in S32 and the consistency loss Lu of the unlabeled data obtained in S35 are combined to form a total loss function, and the calculation formula is
[0072] Total Loss = Ls + λ * Lu
[0073] where λ is the weight coefficient, indicating the contribution of the unlabeled loss to the total loss. Preferably, λ = 0.67 can be set.
[0074] S6. Train the updated model on the mixed set and conduct validation. Fine-tune the model according to the validation results to obtain a convergent result. Finally, use the adjusted model to predict on an independent labeled dataset, and compare it with the true labels to obtain the tracking effect.
[0075] At this time, for the unlabeled data, only select those unlabeled data with a confidence level not lower than 0.95 corresponding to the pseudo-labels from the mixed set, and then regard these pseudo-labels as the true labels of the samples. Each of these unlabeled samples has a weakly augmented sample and a strongly augmented sample. Based on these weakly augmented samples, strongly augmented samples, and the labeled dataset, form a fine-tuning dataset for the semantic segmentation model trained in S5. Based on the true labels and the regarded true labels, and the consistency loss, fine-tune the semantic segmentation model trained in S5 again. When the training converges, stop. Thus, obtain a segmentation model for SAR shadow tracking based on the fine-tuned semantic segmentation model, and predict the shadow area of the target object based on this segmentation model.
[0076] That is, in this application, finally use the fine-tuned model to predict on an independent labeled dataset, and compare it with the true labels to obtain the tracking effect.
[0077] In this embodiment, train the updated model on the mixed set. Set the number of training rounds to 500 rounds, validate once every 50 rounds, and set the batch size to 32. Fine-tune the model according to the validation results. Finally, predict on an independent labeled dataset, and then compare it with the true labels to obtain the tracking effect, as Figure 2 shown. The early stopping method can be used to prevent overfitting during the training process. Determine whether the target is tracked successfully through the IoU (Intersection over Union) value. If IoU is greater than 0.5, it is considered that the tracking is successful. The calculation formula is where A is the predicted shadow area and B is the true shadow area.
[0078] Then calculate the tracking rate to evaluate whether the model achieves the goal.
[0079] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0080] In addition, the descriptions such as "first", "second", etc. are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. can explicitly or implicitly include at least one of such features.
[0081] Any process or method description shown in a flowchart or described in other ways in this specification can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a customized logic function or process. And the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a way that is not in the order shown or discussed, including in a substantially simultaneous manner according to the involved functions or in a reverse order, which should be understood by those skilled in the art to which the embodiments of this application belong.
[0082] It should be understood that each part of this application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware as in another embodiment, any one of the following technologies well-known in the art or their combinations can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0083] Those of ordinary skill in the art in this technical field can understand that all or part of the steps carried by the method for implementing the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the application, rather than to limit them; although the application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
[0085] The above are only some embodiments of the present application. For those of ordinary skill in the art, without departing from the creative concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application.
Claims
1. A semi-supervised efficient video SAR shadow tracking method based on small samples, characterized in that, It includes the following steps: Step 1: Obtain SAR images and perform image preprocessing on them, and obtain an original dataset based on several SAR images after image preprocessing; wherein, the image preprocessing includes: image binarization processing, and image size normalization processing on the binarized image; Step 2: Divide the original dataset into a labeled original dataset and an unlabeled original dataset; Step 3: Divide the labeled original dataset into a first training set and a first validation set, then perform supervised training on the set semantic segmentation model based on the first training set to optimize the network parameters of the model, and tune the hyperparameters of the semantic segmentation model based on the first validation set; and the loss function during supervised training uses consistency loss; Among them, the semantic segmentation model is used to identify multiple targets and perform segmentation, and its output layer is used to perform N + 1 classification and segmentation on the input data, that is, it includes N required target categories and an unknown category; Step 4: Perform semi-supervised training on the semantically segmented model that has been supervised-trained based on the unlabeled original dataset; Perform image weak augmentation processing on each unlabeled original data sample in the unlabeled original dataset to obtain at least one weak augmentation sample for each unlabeled original data sample; Input the weak augmentation sample into the supervised-trained semantic segmentation model, and obtain the pseudo-label of the current weak augmentation sample based on the predicted category output by it; Perform image strong augmentation processing on each unlabeled original data sample in the unlabeled original dataset to obtain at least one strong augmentation sample for each unlabeled original data sample; For each unlabeled original data sample j, obtain a pair of weak-strong sample pairs corresponding to j based on a weak augmentation sample and a strong augmentation sample; Based on the labeled original dataset and all weak-strong sample pairs, form a mixed set, and then use this mixed set as the second training set to perform secondary training on the semantically segmented model that has been supervised-trained; The total loss function used during secondary training is: the weighted sum of the first consistency loss corresponding to the labeled original dataset and the second consistency loss based on all weak-strong sample pairs; wherein, for the second consistency loss, the strong augmentation sample in the weak-strong sample pair is used as the model input, and the pseudo-label of the corresponding weak augmentation sample is regarded as the true label; Step 5: Extract a fine-tuning dataset from the mixed set to perform fine-tuning on the semantically segmented model after secondary training, and the loss function during fine-tuning uses consistency loss; and in this fine-tuning dataset, the weak-strong sample pairs are respectively regarded as a model input sample, and their corresponding pseudo-labels are regarded as the true labels of each sample; Based on the fine-tuned semantic segmentation model, obtain a segmentation model for SAR shadow tracking, and based on this segmentation model, realize the prediction of the shadow area of the SAR image to be tracked.
2. The method according to claim 1, characterized in that, In Step 1, the binarization processing is: set the pixel points of the background to 0 and the pixel points of the target to 1.
3. The method according to claim 1, wherein In Step 2, the ratio of labeled data to unlabeled data is 2:
8.
4. The method according to claim 1, wherein In Step 3, the ratios of the first training set and the first validation set are 80% and 10% respectively.
5. The method according to claim 1, wherein The semantic segmentation model uses the DeepLabV3+ model, with its backbone network being MobileNet. The optimizer is selected as SGD, based on the set initial learning rate, minimum learning rate , maximum learning rate and the total number of training epochs The learning rate decay method adopted is the cosine annealing method.
6. The method according to claim 1, characterized in that, In step 4, the weak image enhancement processing includes: randomly flipping the data, randomly cropping the data, and randomly rotating the data; the strong image enhancement processing includes: randomly occluding the image, generating a new image by linearly combining two images at a specified ratio, and adding noise.
7. The method according to claim 1, wherein In the said step 4, the second consistency loss is the categorical cross-entropy loss with the strongly enhanced image as the input and the corresponding pseudo-label / reliable pseudo-label regarded as the true label.
8. The method according to claim 1, wherein In step 4, the total loss function is: , where and are the first and second consistency losses respectively. λ is a preset weight coefficient, representing the contribution of the unlabeled loss to the total loss, and it is set to λ = 0.
67.
9. The method according to claim 1, characterized in that, Step 4 further includes: Regarding the pseudo-labels with prediction confidence greater than the preset threshold as reliable pseudo-labels; For each original data sample with a reliable pseudo-label, forming a weak-strong sample pair of the current original data sample j based on a strongly enhanced sample and the weakly enhanced sample corresponding to the reliable pseudo-label.
10. The method according to claim 1, wherein The weak-strong sample pairs in the fine-tuning dataset extracted in step 5 are the weak-strong sample pairs with reliable pseudo-labels, where the reliable pseudo-label refers to the pseudo-label with the prediction confidence of a certain weakly enhanced sample greater than the preset threshold.
Citation Information
Patent Citations
Remote sensing video weak supervision target detection method based on motion alignment and dynamic updating
CN116229257A
Semi-supervised remote sensing image semantic segmentation method, device, equipment, medium and product
CN118710901A