Remote sensing image segmentation method and device based on semi-supervised learning, and electronic equipment

Through the semi-supervised learning method, the remote sensing image is segmented, which solves the problem of high time cost of manual labeling, and realizes efficient remote sensing image segmentation, improving the efficiency of land cover classification.

CN120070469AActive Publication Date: 2025-05-30NAT INST OF NATURAL HAZARDS MINISTRY OF EMERGENCY MANAGEMENT OF CHINA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510133859.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

In the prior art, the interpretation and analysis of meter-level high-resolution remote sensing images require a large amount of manual annotation, resulting in high time cost and low efficiency, which limits its application in land cover classification.

Method used

The remote sensing image segmentation method based on semi-supervised learning is used to train the target segmentation model through a small number of labeled samples, and the trained model is used to quickly and efficiently segment the remote sensing image to be measured.

Benefits of technology

This method can significantly save time in manual segmentation and labeling, improve efficiency, and enhance the application potential of meter-level high-resolution remote sensing images in land cover classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070469A_ABST
    Figure CN120070469A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a remote sensing image segmentation method and device based on semi-supervised learning, and electronic equipment. The method comprises the following steps: acquiring an annotation set based on a preset remote sensing image, wherein the annotation set comprises a training set and a verification set; training the target segmentation model by adopting the training set; inputting the verification set into a target segmentation model to obtain a prediction segmentation result of the verification set, and when the prediction segmentation result of the verification set does not meet a preset verification requirement, updating a verification image to which a verification sample of which the sample segmentation result meets the preset prediction requirement in the verification set belongs and a prediction category label corresponding to the verification sample to a training set, returning to execute the step of training the at least one target segmentation model by adopting the training set; and inputting the to-be-detected remote sensing image into the at least one target segmentation model to obtain a prediction segmentation result of the to-be-detected remote sensing image. According to the scheme, the to-be-detected remote sensing image can be quickly and efficiently segmented, and the manual segmentation and labeling time is saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a remote sensing image segmentation method, apparatus, electronic device, storage medium, and computer program product based on semi-supervised learning. Background Art

[0002] The land cover classification method based on remote sensing images has important application prospects in aspects such as land resource management, urban planning, agricultural yield estimation, and environmental protection. In recent years, the spatial resolution of remote sensing images has been continuously improved. The meter-level high-resolution remote sensing images have very rich details of ground objects such as texture, shape, and spatial distribution, and are particularly significant for the recognition of fine land classes. However, the interpretation and analysis of meter-level high-resolution remote sensing images rely on a large number of manual annotations. Since the image size of remote sensing images is large and the number of pixels can reach hundreds of millions, using this manual annotation method will result in high time costs and low efficiency, which to a certain extent restricts its application in land cover classification. Summary of the Invention

[0003] In view of the above problems, the present invention is proposed. The present invention provides a remote sensing image segmentation method, apparatus, electronic device, storage medium, and computer program product based on semi-supervised learning. This solution can train one or more target segmentation models based on a small number of labeled samples, and the trained target segmentation models can be used to quickly and efficiently segment the remote sensing images to be measured, thereby saving a large amount of time for manual segmentation and annotation.

[0004] According to one aspect of the present invention, a remote sensing image segmentation method based on semi-supervised learning is provided. The method includes: obtaining an annotation set based on a preset remote sensing image, the annotation set including a training set and a validation set, the training set including at least one training image and class labels of each image region on at least one training image, the validation set including at least one validation image and class labels of each image region on at least one validation image; training at least one target segmentation model using the training set, the at least one target segmentation model being used to predict the class of each image region of an input image; inputting the validation set into the at least one target segmentation model to obtain a predicted segmentation result of the validation set, when the predicted segmentation result of the validation set does not meet the preset validation requirements, updating the validation image to which the validation sample whose sample segmentation result in the validation set meets the preset prediction requirements belongs and the predicted class label corresponding to the validation sample to the training set, and returning to execute the step of training the at least one target segmentation model using the training set until the predicted classification result of the validation set meets the preset validation requirements, wherein the predicted segmentation result of the validation set includes the sample segmentation results corresponding to each validation sample in the validation set, the sample segmentation result being used to indicate the predicted class label corresponding to the validation sample, the validation sample being an image region in the validation image, the preset prediction requirements including that the prediction probability of the predicted class label corresponding to the validation sample meets a first preset confidence level, and the preset validation requirements including that the prediction probability of the predicted class label corresponding to each validation sample in the validation set meets a second preset confidence level; inputting the to-be-detected remote sensing image into the at least one target segmentation model to obtain a predicted segmentation result of the to-be-detected remote sensing image.

[0005] Optionally, the number of target segmentation models is greater than 1, and the network architectures of the respective target segmentation models are different. Inputting the validation set into the at least one target segmentation model to obtain a predicted segmentation result of the validation set includes: inputting the validation set into each of the target segmentation models respectively to obtain the model segmentation results output by each of the target segmentation models, the model segmentation result being used to indicate the prediction probability that each pixel in the validation set belongs to each preset class; determining the respective target weights of each of the target segmentation models according to the model segmentation results output by each of the target segmentation models; for each preset class among a plurality of preset classes, for each pixel in the validation set, weighting and summing the prediction probabilities that the pixel belongs to the preset class according to the respective target weights of each of the target segmentation models, and determining the preset class with the largest weighted result as the predicted class label of the pixel, the predicted class label corresponding to the pixel belonging to one of the plurality of preset classes, and the predicted class label corresponding to the validation sample being the predicted class labels of the respective pixels included in the validation sample.

[0006] Optionally, determining the target weights of the target segmentation models according to the model segmentation results output by the target segmentation models respectively includes: determining the overall accuracy rates of the target segmentation models respectively according to the model segmentation results output by the target segmentation models respectively; determining the target weights of the target segmentation models respectively according to the overall accuracy rates of the target segmentation models respectively.

[0007] Optionally, for any pixel, the prediction probability of the pixel belonging to the preset category is weighted through the following formula, and the preset category with the largest weighted result is determined as the prediction category label of the pixel:

[0008]

[0009] In the formula, A represents the prediction category label of the pixel, n represents the number of target segmentation models, W i represents the target weight of the i-th target segmentation model, and P i (x) represents the probability that the pixel belongs to the x-th preset category in the model segmentation result output by the i-th target segmentation model.

[0010] Optionally, the training set includes a first training sample and a second training sample. The first training sample and the second training sample are image regions in the training image, and the area of the first training sample is larger than the area of the second training sample. Training at least one target segmentation model using the training set includes: performing a first training operation, including: for each target segmentation model, inputting the training set into the target segmentation model; using the focal loss function to adjust the parameters of the target segmentation model until the target segmentation model meets the first optimization requirement, where during the parameter adjustment process, the learning rate of the target segmentation model remains unchanged at a preset standard learning rate, and the first optimization requirement includes that the prediction probability corresponding to the prediction category label obtained by predicting the first training sample meets the third preset confidence level; after performing the first training operation, performing a second training operation, including: for each target segmentation model, inputting the training set into the target segmentation model; using the focal loss function to adjust the parameters of the target segmentation model until the target segmentation model meets the second optimization requirement, where the parameters involved in the adjustment during the parameter adjustment process include the learning rate, and the second optimization requirement includes that the prediction probability corresponding to the prediction category label obtained by predicting the second training sample meets the fourth preset confidence level.

[0011] Optionally, the balancing factor adopted by the focal loss function in the first training operation is smaller than the balancing factor adopted in the second training operation. The formula of the focal loss function is as follows:

[0012] FL(p t )=-α t (1-p t ) γlog(p t )

[0013] where p t represents the sample rate of correct prediction. The sample rate of correct prediction is equal to the ratio of the number of training samples whose predicted probability corresponding to the predicted class label obtained by prediction conforms to the corresponding preset confidence level to the total number of training samples. α t represents the balance factor, and γ represents the preset adjustment factor.

[0014] Optionally, obtaining an annotation set based on a preset remote sensing image includes: determining reliable sample data and sample data to be annotated based on the preset remote sensing image. The reliable sample data includes target reliable samples and the class labels of the target reliable samples. The sample data to be annotated includes samples to be annotated. The target reliable samples and the samples to be annotated are image regions in the preset remote sensing image; training an annotation model using the image to which the target reliable samples belong; inputting the image to which the samples to be annotated belong into the annotation model to obtain re-annotated samples. When the number of target reliable samples does not meet the preset quantity requirement, adding the reliable samples in the re-annotated samples to the target reliable samples, and returning to execute the step of training the annotation model using the image to which the target reliable samples belong until the number of target reliable samples meets the preset quantity requirement, and then determining the image to which the target reliable samples belong as the image included in the annotation set; wherein, the reliable samples in the re-annotated samples include samples whose predicted class labels output by the annotation model conform to the preset label requirements, and samples obtained after manual re-annotation of samples whose predicted class labels output by the annotation model do not conform to the preset label requirements.

[0015] Optionally, training an annotation model using the image to which the target reliable samples belong includes: performing image enhancement on the image to which the target reliable samples belong; inputting the image to which the target reliable samples belong before image enhancement and the image to which the target reliable samples belong after image enhancement into the annotation model to calculate the consistency loss value of the target reliable samples before and after image enhancement; using the consistency loss value to adjust the parameters of the annotation model.

[0016] Optionally, determining reliable sample data and sample data to be labeled based on a preset remote sensing image includes: segmenting the preset remote sensing image into multiple image patches; inputting the multiple image patches into a pre-trained segmentation model to obtain the image patch segmentation results of the multiple image patches output by the pre-trained segmentation model, where the image patch segmentation results are used to indicate the predicted class labels of each image region included in each of the multiple image patches; in response to a relabeling operation by the user on at least some of the multiple image patches based on the predicted class labels predicted by the pre-trained segmentation model, determining new predicted class labels for at least some of the image regions in at least some of the image patches; in response to the user's determination operation, determining at least some of the image regions with existing predicted class labels as target reliable samples, and at least some of the image regions without existing predicted class labels as samples to be labeled.

[0017] Optionally, the training set includes first training samples and second training samples, where the first training samples and the second training samples are image regions in a training image, and the area of the region of the first training samples is larger than the area of the region of the second training samples; after determining the image to which the target reliable samples belong as an image included in the annotation set, the method further includes: performing image enhancement on the second training samples in the training set; by randomly embedding the image-enhanced second training samples into any target background sample in the training set, obtaining a new training image based on the initial training image to which the target background sample belongs, and adding the new training image to the training set or replacing the initial training image with the new training image, where the area of the region occupied by the target background sample in the initial training image is the largest or greater than a preset area threshold.

[0018] According to another aspect of the present invention, there is also provided an image segmentation device, which includes: an acquisition module, configured to obtain an annotation set based on a preset remote sensing image, the annotation set including a training set and a validation set, the training set including at least one training image and class labels of each image region on at least one training image, and the validation set including at least one validation image and class labels of each image region on at least one validation image; a training module, configured to train at least one target segmentation model using the training set, where the at least one target segmentation model is used to predict the class of each image region of an input image; a first input module, configured to input the validation set into the at least one target segmentation model to obtain a predicted segmentation result of the validation set. When the predicted segmentation result of the validation set does not meet the preset validation requirements, update the validation image to which the validation sample whose sample segmentation result in the validation set meets the preset prediction requirements and the predicted class label corresponding to the validation sample to the training set, and return to execute the step of training the at least one target segmentation model using the training set until the predicted classification result of the validation set meets the preset validation requirements. Wherein, the predicted segmentation result of the validation set includes the sample segmentation results corresponding to each validation sample in the validation set, the sample segmentation result is used to indicate the predicted class label corresponding to the validation sample, the validation sample is an image region in the validation image, the preset prediction requirements include that the prediction probability of the predicted class label corresponding to the validation sample meets a first preset confidence level, and the preset validation requirements include that the prediction probability of the predicted class label corresponding to each validation sample in the validation set meets a second preset confidence level; a second input module, configured to input the to-be-detected remote sensing image into the at least one target segmentation model to obtain a predicted segmentation result of the to-be-detected remote sensing image.

[0019] According to still another aspect of the present invention, there is also provided an electronic device, including: a processor and a memory. Wherein, computer program instructions are stored in the memory, and when the computer program instructions are run by the processor, they are used to execute the above-mentioned remote sensing image segmentation method based on semi-supervised learning.

[0020] According to yet another aspect of the present invention, there is also provided a storage medium, on which program instructions are stored, and when the program instructions are run, they are used to execute the above-mentioned remote sensing image segmentation method based on semi-supervised learning.

[0021] According to yet another aspect of the present invention, there is also provided a computer program product, including computer program instructions, and when the computer program instructions are run, they are used to execute the remote sensing image segmentation method based on semi-supervised learning as described above.

[0022] The above technical solution trains the target segmentation model using a training set and verifies the performance of the target segmentation model using a validation set. This helps ensure that the accuracy of the trained target segmentation model in predicting the categories of each image region of the input remote sensing image meets the actual prediction requirements. In particular, when the predicted segmentation result obtained after inputting the validation set into the target segmentation model does not meet the preset validation requirements, that is, when the performance of the target segmentation model does not meet the actual prediction requirements, by updating the validation images to which the validation samples with segmentation results meeting the preset prediction requirements in the validation set belong and the corresponding predicted category labels of these validation samples to the training set, without having to obtain additional training images and the category labels of each image region on the training images, this training method based on semi-supervised learning can effectively save the time for obtaining training images and the category labels of each image region again to expand the training set, and this way of updating the training set helps expand the number of training images in the training set and the number of category labels of each image region on the training images. By updating the training set multiple times to train the target segmentation model multiple times, the performance of the target segmentation model can be effectively improved. On the other hand, by inputting the remote sensing image to be measured into the trained target segmentation model, the categories corresponding to each image region on the remote sensing image to be measured can be quickly and accurately determined, greatly saving the time for manual segmentation and annotation, and the efficiency is relatively high.

[0023] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically given below. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] By describing the embodiments of the present invention in more detail in conjunction with the drawings, the above and other purposes, features and advantages of the present invention will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the description, and are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings, the same reference numerals generally represent the same components or steps.

[0025] Figure 1 FIG. shows a schematic flowchart of a remote sensing image segmentation method based on semi-supervised learning according to an embodiment of the present invention;

[0026] Figure 2 FIG. shows a schematic block diagram of a remote sensing image segmentation device based on semi-supervised learning according to an embodiment of the present invention;

[0027] Figure 3 FIG. shows a schematic block diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] In order to make the purpose, technical scheme and advantages of the present invention more obvious, the exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described herein. Based on the embodiments of the present invention described in the present invention, all other embodiments obtained by those skilled in the art without creative work should fall within the protection scope of the present invention.

[0029] To at least partially solve the above technical problems, the embodiments of the present invention provide a remote sensing image segmentation method, device, electronic device, storage medium and computer program product based on semi-supervised learning. This solution can train one or more target segmentation models based on a small number of labeled samples, and the trained target segmentation model can be used to quickly and efficiently segment the remote sensing image to be tested, thereby greatly saving the time of manual segmentation and labeling.

[0030] See also Figure 1 As shown, it is a schematic flow chart of a remote sensing image segmentation method based on semi-supervised learning according to an embodiment of the present invention. According to one aspect of the present invention, a remote sensing image segmentation method based on semi-supervised learning is provided, and the method includes step S110, step S120, step S130 and step S140.

[0031] In step S110, a labeling set is obtained based on a preset remote sensing image, the labeling set includes a training set and a verification set, the training set includes at least one training image and a category label for each image region on at least one training image, and the verification set includes at least one verification image and a category label for each image region on at least one verification image.

[0032] Exemplarily, the preset remote sensing image can be, for example, a remote sensing image collected by devices such as an optical sensor or a radar sensor on a satellite platform, or can be, for example, a remote sensing image collected by devices such as an optical camera, a multispectral camera, or a thermal imager on a manned aircraft or an unmanned aerial vehicle on an aerial platform. The number of preset remote sensing images can be one or more, and an annotation set can be obtained based on the preset remote sensing images. Specifically, the annotation set can include a training set, and the training set can include one or more training images. At least part of the image regions of each training image can each have a corresponding class label. The training image can be the entire preset remote sensing image, or can also be an image block segmented from the preset remote sensing image. Similarly, the annotation set can also include a validation set, and the validation set can include one or more validation images. At least part of the image regions of each validation image can each have a corresponding class label. The validation image can be the entire preset remote sensing image, or can also be an image block segmented from the preset remote sensing image. The training images included in the training set and the validation images included in the validation set can be the same or can be different. When the training images and the validation images are the same, at least part of the image regions with class labels in the training images and the image regions with class labels in the validation images are different. The class labels can include, for example: cultivated land, forest land, grassland, urban and rural residential and public facility land, industrial and mining land, etc., which can be defined by the user according to actual classification needs. Preferably, the class labels can be defined with reference to the classification method of "Current Classification of Land Use". The class label of a single image region can be one or more. For example, the class label of a single image region can be paddy field, or for another example, the class label of a single image region can include cultivated land and paddy field, where cultivated land can be used as a first-level class label and paddy field can be used as a second-level class label. Each image region can be a single pixel of the image, or can also be a connected domain composed of multiple pixels of the image.

[0033] In step S120, at least one target segmentation model is trained using the training set, and the at least one target segmentation model is used to predict the class of each image region of the input image.

[0034] Exemplarily, the target segmentation model can be, for example, any one or more of a Pyramid Scene Parsing Network (PSPNet), a DeepLabv3+ (a dilated convolutional semantic segmentation network based on an encoder-decoder), a High-Resolution Net (HRNet), a Fully Convolutional Networks (FCN), a Dual Attention Network (DANet), etc., which are semantic segmentation models that can be used to predict the categories of each image region of the input image. In some embodiments, the initial target segmentation model can be an open-source semantic segmentation model pre-trained on a large dataset. The training set in the annotation set can be used to train the target segmentation model to further optimize the target segmentation model.

[0035] In step S130, the validation set is input into at least one target segmentation model to obtain the predicted segmentation results of the validation set. When the predicted segmentation results of the validation set do not meet the preset validation requirements, the validation images to which the validation samples whose sample segmentation results in the validation set meet the preset prediction requirements belong, and the predicted class labels corresponding to the validation samples are updated to the training set, and the step of training at least one target segmentation model using the training set is returned and executed until the predicted classification results of the validation set meet the preset validation requirements. Among them, the predicted segmentation results of the validation set include the sample segmentation results corresponding to each validation sample in the validation set, the sample segmentation results are used to indicate the predicted class labels corresponding to the validation samples, the validation samples are image regions in the validation images, the preset prediction requirements include that the prediction probability of the predicted class label corresponding to the validation sample meets the first preset confidence level, and the preset validation requirements include that the prediction probability of the predicted class label corresponding to each validation sample in the validation set meets the second preset confidence level.

[0036] Exemplarily, for each validation image in the validation set, some or all of the image regions of the validation image may have corresponding class labels, and the image regions with class labels can be used as validation samples. After each training of the target segmentation model using the training set, the validation set can be input into the target segmentation model, and the target segmentation model can output the predicted segmentation results of the validation set. Specifically, the predicted segmentation results of the validation set may include the sample segmentation results corresponding to each validation sample in the validation set, and the sample segmentation results can indicate the predicted class labels corresponding to the validation samples. Specifically, for each validation sample, the target segmentation model can determine and output the predicted class label of the validation sample according to the predicted probabilities of various class labels predicted for the validation sample. For example, the target segmentation model can determine the class label with the highest predicted probability as the predicted class label of the validation sample, and optionally output the predicted probability corresponding to the class label. For another example, the target segmentation model can output a specific number of class labels as the predicted class labels of the validation sample in the order of priority from the highest predicted probability to the lowest, and optionally output the predicted probabilities corresponding to the output class labels. For another example, the target segmentation model can determine the class labels with predicted probabilities greater than or equal to a specific value as the predicted class labels of the validation sample, and optionally output the predicted probabilities corresponding to the output class labels. In a specific embodiment, for a certain validation sample, the sample segmentation result of the validation sample may include: forest land (predicted probability = 96.7%), grassland (predicted probability = 2.2%), farmland (predicted probability = 0.4%). In another specific embodiment, for a certain validation sample, the sample segmentation result of the validation sample may include: forest land (predicted probability = 96.7%). In yet another specific embodiment, for a certain validation sample, the sample segmentation result of the validation sample may include: arable land (predicted probability = 99.3%), paddy field (predicted probability = 97.4%), irrigated land (predicted probability = 1.4%), forest land (predicted probability = 0.4%), marshland (predicted probability = 0.33%). In this embodiment, arable land and forest land are used as first-level class labels, paddy field and irrigated land are second-level class labels belonging to arable land, and marshland is a second-level class label belonging to forest land. Each validation sample may correspond to the predicted class label output by the target segmentation model and may correspond to the class label it itself has. It can be understood that for each validation sample, among the predicted class labels output by the target segmentation model for the validation sample, the greater the predicted probability of the predicted class label that is consistent with the class label that the validation sample itself has, the better the performance of the target segmentation model.

[0037] Exemplarily, for each verification sample, when the prediction probability of the prediction class label corresponding to the verification sample meets the first preset confidence level, it can be considered that the sample segmentation result corresponding to the verification sample meets the preset prediction requirements. In some embodiments, when the prediction class label with the highest prediction probability corresponding to the verification sample is consistent with the class label of the verification sample itself, and the prediction probability of the prediction class label with the highest prediction probability is greater than or equal to the first preset probability threshold, it can be considered that the prediction probability of the prediction class label corresponding to the verification sample meets the first preset confidence level. The first preset probability thresholds corresponding to the class labels belonging to different preset categories may be the same or may not be the same. In other embodiments, when the prediction class label with the highest prediction probability corresponding to the verification sample is consistent with the class label of the verification sample itself, and the prediction probability of any prediction class label other than the prediction class label with the highest prediction probability is less than the second preset probability threshold, it can be considered that the prediction probability of the prediction class label corresponding to the verification sample meets the first preset confidence level.

[0038] Exemplarily, when the predicted segmentation result of the validation set does not meet the preset validation requirements, the validation images to which the validation samples with segmentation results meeting the preset prediction requirements in the validation set belong and the corresponding predicted class labels of the validation samples can be updated to the training set. Specifically, when the validation image to which the validation sample belongs is the same as a certain training image in the training set, the class label of the image region corresponding to the validation sample in the training image can be updated, and after the update, the class label of the image region in the training image can be the predicted class label with the maximum predicted probability of the corresponding validation sample or the predicted class label with a predicted probability greater than or equal to a specific probability threshold. Alternatively, the validation image to which the validation sample belongs can be directly added to the training set as a training image, and the class label of the image region corresponding to the validation sample in the training image can be updated to the predicted class label with the maximum predicted probability of the validation sample or the predicted class label with a predicted probability greater than or equal to a specific probability threshold. When the validation image to which the validation sample belongs is different from any training image in the training set, the validation image to which the validation sample belongs can be directly added to the training set as a training image, and the class label of the image region corresponding to the validation sample in the training image can be updated to the predicted class label with the maximum predicted probability of the validation sample or the predicted class label with a predicted probability greater than or equal to a specific probability threshold. Exemplarily, when the predicted segmentation result of the validation set meets the preset validation requirements, it can be considered that the target segmentation model meets the preset optimization requirements. The preset validation requirements may include: the predicted probabilities of the predicted class labels corresponding to the validation samples in the validation set meet the second preset confidence level. For each validation sample in the validation set, to determine whether the predicted probability of the predicted class label corresponding to the validation sample meets the second preset confidence level, reference can be made to the situation where the predicted probability of the predicted class label corresponding to the validation sample meets the first preset confidence level as described above, which will not be elaborated here. It should be noted that the first preset probability threshold and / or the second preset probability threshold in the second preset confidence level and the first preset probability threshold and / or the second preset probability threshold in the first preset confidence level may have the same or different numerical values. When the predicted probabilities of the predicted class labels corresponding to each validation sample meet the second preset confidence level, it can be considered that the predicted segmentation result of the validation set meets the preset validation requirements.

[0039] In step S140, the remote sensing image to be measured is input into at least one target segmentation model to obtain the predicted segmentation result of the remote sensing image to be measured.

[0040] Exemplarily, after completing the training of the target segmentation model, the remote sensing image to be measured can be input into the target segmentation model to obtain the predicted segmentation result of the remote sensing image to be measured. The predicted segmentation result of the remote sensing image to be measured can refer to the relevant description of the predicted segmentation result of the validation set in the foregoing embodiments, which will not be elaborated here. In some embodiments, the remote sensing image to be measured and the preset remote sensing image may be related images. For example, the remote sensing image to be measured and the preset remote sensing image may be the same image. In this case, the union of the image regions of each image included in the annotation set may be a partial image region of the preset remote sensing image (i.e., the remote sensing image to be measured). For another example, the preset remote sensing image is an image patch in the remote sensing image to be measured. For another example, the remote sensing image to be measured and the preset remote sensing image belong to a set of remote sensing images collected by the same device for the target region. The set of remote sensing images may include multiple images. Some of the multiple images may be used as the remote sensing image to be measured, and some of the other images may be used as the preset remote sensing image. In other embodiments, the remote sensing image to be measured and the preset remote sensing image may be independently collected remote sensing images. Exemplarily, when there are multiple target segmentation models, the predicted segmentation result of the remote sensing image to be measured can be comprehensively determined according to the model segmentation results output by each target segmentation model. For example, for each pixel of the remote sensing image to be measured, each target segmentation model can output the predicted probability that the pixel belongs to each preset category. Each target segmentation model may have its own target weight. The predicted probability that the pixel belongs to the preset category is weighted and summed according to the target weights of each target segmentation model, and the preset category with the largest weighted result is determined as the predicted category label of the pixel. In this case, the predicted category label corresponding to each image region of the remote sensing image to be measured is the predicted category label of each pixel included in the image region. For another example, for each pixel of the remote sensing image to be measured, each target segmentation model can output the predicted probability that the pixel belongs to each preset category. The average value of the predicted probabilities that each target segmentation model outputs for the pixel belonging to the preset category can be calculated, and the preset category with the largest average value is determined as the predicted category label of the pixel. Similarly, the predicted category label corresponding to each image region of the remote sensing image to be measured is the predicted category label of each pixel included in the image region.

[0041] The above technical solution trains the target segmentation model using a training set and verifies the performance of the target segmentation model using a validation set. This helps ensure that the accuracy of the classes of each image region of the remotely sensed image predicted by the trained target segmentation model meets the actual prediction requirements. In particular, when the predicted segmentation result obtained after inputting the validation set into the target segmentation model does not meet the preset validation requirements, that is, when the performance of the target segmentation model does not meet the actual prediction requirements, by updating the validation images to which the validation samples with sample segmentation results meeting the preset prediction requirements in the validation set belong and the predicted class labels corresponding to the validation samples to the training set, without having to obtain additional training images and the class labels of each image region on the training images, this training method based on semi-supervised learning can effectively save the time for obtaining training images and the class labels of each image region again to expand the training set. Moreover, this way of updating the training set helps expand the number of training images in the training set and the number of class labels of each image region on the training images. By updating the training set multiple times to train the target segmentation model multiple times, the performance of the target segmentation model can be effectively improved. On the other hand, by inputting the remotely sensed image to be measured into the trained target segmentation model, the classes corresponding to each image region on the remotely sensed image to be measured can be determined quickly and accurately, greatly saving the time for manual segmentation and annotation, with relatively high efficiency.

[0042] Optionally, the number of target segmentation models is greater than 1, and the network architectures of the target segmentation models are different. Inputting the validation set into at least one target segmentation model to obtain the predicted segmentation result of the validation set includes: inputting the validation set into each target segmentation model respectively to obtain the model segmentation results output by each target segmentation model, where the model segmentation results are used to indicate the predicted probabilities of each pixel in the validation set belonging to each preset class; determining the target weights of each target segmentation model according to the model segmentation results output by each target segmentation model; for each preset class among multiple preset classes, for each pixel in the validation set, performing a weighted sum of the predicted probabilities of the pixel belonging to the preset class according to the target weights of each target segmentation model, and determining the preset class with the largest weighted result as the predicted class label of the pixel. The predicted class label corresponding to the pixel belongs to one of the multiple preset classes, and the predicted class label corresponding to the validation sample is the predicted class labels of each pixel included in the validation sample.

[0043] Exemplarily, the number of target segmentation models can be multiple. In this case, the network architectures of the respective target segmentation models can be different. In step S130, for each target segmentation model, the validation set can be input into the target segmentation model to obtain the model segmentation result output by the target segmentation model. The model segmentation result can indicate the prediction probability that each pixel in the validation set belongs to each preset category. Specifically, for the same pixel, each target segmentation model can output multiple prediction probabilities that the pixel belongs to multiple preset categories, and each preset category corresponds to one prediction probability. The preset categories can be the categories indicated by the category labels in the foregoing embodiments, and can include, for example: arable land, forest land, grassland, urban and rural residential and public facilities land, industrial and mining land, etc. According to the model segmentation results output by the respective target segmentation models, the respective target weights of the respective target segmentation models can be determined. For example, for each target segmentation model, the preset category with the maximum prediction probability output by the target segmentation model for each pixel can be used as the prediction category of the pixel. According to the prediction categories of the respective pixels and the category labels corresponding to the respective pixels, the accuracy rate of the target segmentation model for predicting each preset category can be determined. The respective target weights of the respective target segmentation models can be determined according to the minimum accuracy rates of the respective target segmentation models, and the respective target weights of the respective target segmentation models can also be determined according to the overall accuracy rates of the respective target segmentation models. The overall accuracy rate can be, for example, equal to the ratio of the number of pixels with consistent prediction categories and category labels to the total number of pixels predicted. Exemplarily, after determining the respective target weights of the respective target segmentation models, for each preset category, according to the respective target weights of the respective target segmentation models, the prediction probabilities that each pixel in the validation set belongs to the preset category can be weighted and summed, and the preset category with the largest weighted result can be determined as the prediction category label of the pixel. For example, at least one target segmentation model can include a first target segmentation model, a second target segmentation model, and a third target segmentation model, and their respective target weights are 0.5, 0.3, and 0.2 respectively. For a certain pixel, the model segmentation result output by the first target segmentation model is: preset category a = 70%, preset category b = 20%, preset category c = 10%; the model segmentation result output by the second target segmentation model is: preset category a = 60%, preset category b = 10%, preset category d = 30%; the model segmentation result output by the third target segmentation model is: preset category a = 80%, preset category b = 20%. Then the weighted result of the prediction probability that the pixel belongs to preset category a = 0.5 * 70% + 0.3 * 60% + 0.2 * 80% = 69%, the weighted result of the prediction probability that the pixel belongs to preset category b = 0.5 * 20% + 0.3 * 10% + 0.2 * 20% = 17%, the weighted result of the prediction probability that the pixel belongs to preset category c = 0.5 * 10% = 5%, and the prediction probability that the pixel belongs to preset category d = 0.3 * 30% = 9%.In this embodiment, if the weighted result of the predicted probability of the preset category a is the largest, the preset category a can be determined as the predicted category label of this pixel. It can be understood that the predicted category label corresponding to each pixel belongs to one of the multiple preset categories, and the predicted category label corresponding to the verification sample can be the predicted category labels of the respective pixels included.

[0044] In the above technical solution, the verification set is input into multiple target segmentation models respectively, and the target weights of each target segmentation model are determined according to the model segmentation results output by each target segmentation model. The predicted probabilities of each pixel in the verification set belonging to each preset category are weighted and summed to determine the predicted category labels of each pixel. This can comprehensively determine the predicted category labels of each verification sample by combining the model segmentation results of each target segmentation model. The determined predicted category labels can reflect the comprehensive performance of multiple target segmentation models to a certain extent. When training each target segmentation model, based on the predicted category labels of the verification samples determined comprehensively, it is verified whether the performance of the target segmentation model meets the actual prediction requirements. Compared with verifying whether the performance of the target segmentation model meets the actual prediction requirements based on the predicted category labels of the verification samples determined by this target segmentation model alone, the requirement for the self-performance of the target segmentation model is lower, the training difficulty is smaller, and the training time can be saved. On the other hand, this way of determining the predicted category labels of the verification samples is beneficial to improving the fault tolerance rate of each target segmentation model that has completed training, and the accuracy of the finally determined predicted segmentation result of the remote sensing image to be measured is relatively high.

[0045] Optionally, determining the target weights of each target segmentation model according to the model segmentation results output by each target segmentation model includes: determining the overall accuracy of each target segmentation model one by one according to the model segmentation results output by each target segmentation model; determining the target weights of each target segmentation model one by one according to the overall accuracy of each target segmentation model.

[0046] Exemplarily, for each target segmentation model, the overall accuracy of the target segmentation model can be determined according to the model segmentation result output by the target segmentation model. For example, the ratio of the number of pixels with consistent predicted categories and category labels in the model segmentation result output by any target segmentation model to the total number of pixels included in the validation sample can be used as the overall accuracy of the target segmentation model. The relevant descriptions of the predicted categories can refer to the foregoing embodiments and will not be elaborated here. According to the respective overall accuracies of the target segmentation models, the respective target weights of the target segmentation models can be determined one by one. For example, at least one target segmentation model may include a first target segmentation model, a second target segmentation model, and a third target segmentation model. The overall accuracy of the first target segmentation model is f1, the overall accuracy of the second target segmentation model is f2, and the overall accuracy of the third target segmentation model is f3. Then, the target weight of the first target segmentation model can be f1 / (f1 + f2 + f3), the target weight of the second target segmentation model can be f2 / (f1 + f2 + f3), and the target weight of the third target segmentation model can be f3 / (f1 + f2 + f3).

[0047] The above technical solution determines the respective target weights of the target segmentation models according to the respective overall accuracies of the target segmentation models, which can assign higher weights to the target segmentation models with better performance. This helps the predicted category labels of each pixel obtained to better reflect the comprehensive performance of each target segmentation model, thereby helping to accurately determine whether the performance of each target segmentation model meets the actual prediction requirements.

[0048] Optionally, for any pixel, the predicted probability of the pixel belonging to the preset category is weighted by the following formula, and the preset category with the largest weighted result is determined as the predicted category label of the pixel:

[0049]

[0050] In the formula, A represents the predicted category label of the pixel, n represents the number of target segmentation models, W i represents the target weight of the i-th target segmentation model, and P i (x) represents the probability that the pixel belongs to the x-th preset category in the model segmentation result output by the i-th target segmentation model.

[0051] Exemplarily, taking n = 3 as an example, the multiple target segmentation models include the first target segmentation model (i = 1), the second target segmentation model (i = 2), and the third target segmentation model (i = 3), where W 1 = 0.4, W 2 = 0.3, W 3 = 0.3. The number of preset categories is, for example, 2, including the first preset category (x = 1) and the second preset category (x = 2). P1 P(1) = 0.8 1 P(2) = 0.2 2 P(1) = 0.7 2 P(2) = 0.3 3 P(1) = 0.75 3 P(2) = 0.25, then A = argmax(x = 1, 0.4 * 0.8 + 0.3 * 0.7 + 0.3 * 0.75 = 0.755, x = 2, 0.4 * 0.2 + 0.3 * 0.3 + 0.3 * 0.25 = 0.245) = the second preset category.

[0052] The above technical solution can quickly and accurately determine the preset category with the largest weighted result and use this preset category as the predicted label category of the corresponding pixel. In particular, when there are many preset categories, the calculation amount using this weighted summation formula is small, and the predicted label category of each pixel can be efficiently determined.

[0053] Optionally, the training set includes a first training sample and a second training sample. The first training sample and the second training sample are image regions in the training image, and the area of the first training sample is larger than the area of the second training sample. Training at least one target segmentation model using the training set includes: performing a first training operation, including: for each target segmentation model, inputting the training set into the target segmentation model; using the focal loss function to adjust the parameters of the target segmentation model until the target segmentation model meets the first optimization requirement. Among them, during the parameter adjustment process, the learning rate of the target segmentation model remains unchanged at the preset standard learning rate, and the first optimization requirement includes that the prediction probability corresponding to the predicted category label obtained by predicting the first training sample meets the third preset confidence level. After performing the first training operation, perform a second training operation, including: for each target segmentation model, inputting the training set into the target segmentation model; using the focal loss function to adjust the parameters of the target segmentation model until the target segmentation model meets the second optimization requirement. Among them, during the parameter adjustment process, the parameters involved in the adjustment include the learning rate, and the second optimization requirement includes that the prediction probability corresponding to the predicted category label obtained by predicting the second training sample meets the fourth preset confidence level.

[0054] Exemplarily, the training set may include a first training sample and a second training sample. The first training sample and the second training sample are image regions in the training images. In other words, each first training sample or each second training sample is an image region of one of the training images in the training set. The area of any first training sample is larger than the area of any second training sample. That is to say, the area of each first training sample among the first training samples in the corresponding training image is relatively large and can be regarded as a large target region in the corresponding training image. Correspondingly, the area of each second training sample among the second training samples in the corresponding training image is relatively small and can be regarded as a small target region in the corresponding training image. The first training sample and the second training sample can be determined by pixel statistics and setting an area threshold. Specifically, the area of each image region can be determined through pixel statistics. By setting a specific area threshold, the image regions greater than or equal to the area threshold can be determined as the first training samples, and the image regions smaller than the area threshold can be determined as the second training samples. In step S120, for each target segmentation model, the training of the target segmentation model may include a first training operation and a second training operation. Specifically, the first training operation can be executed first. When executing the first training operation, the training set can be input into the target segmentation model, and the focal loss function can be used to adjust the parameters of the target segmentation model. During the parameter adjustment process, the learning rate of the target segmentation model can always be maintained at a preset standard learning rate. One or more parameters among other parameters except the learning rate can be adjusted according to the loss value obtained based on the focal loss function until the target segmentation model meets the first optimization requirement. The first optimization requirement may include: the prediction probability corresponding to the predicted class label obtained for the first training sample conforms to the third preset confidence level. For each first training sample in the training set, to determine whether the prediction probability of the predicted class label corresponding to the first training sample conforms to the third preset confidence level, reference can be made to the situation where the prediction probability of the predicted class label corresponding to each of the foregoing verification samples conforms to the first preset confidence level, which will not be elaborated here. It should be noted that the first preset probability threshold and / or the second preset probability threshold in the third preset confidence level may be the same as or different from the first preset probability threshold and / or the second preset probability threshold in the first preset confidence level. When the prediction probability of the predicted class label corresponding to each first training sample conforms to the third preset confidence level, it can be considered that the target segmentation model meets the first optimization requirement. More specifically, since the first training sample belongs to the large target region, it can be understood that when the target segmentation model meets the first optimization requirement, it can be considered that the target segmentation model can accurately segment the large target region in the image. In other words, the first training operation is mainly used to enable the target segmentation model to accurately segment the large target region in the image.Since the segmentation task for a large target area is relatively simple, the target segmentation model can be easily optimized to meet the first optimization requirement even without adjusting the learning rate. Using a fixed learning rate can simplify the first training operation, reduce the time for finding the optimal learning rate, and save the computing resources required to find the optimal learning rate.

[0055] Exemplarily, after executing the first training operation, the second training operation can be executed. When executing the second training operation, for any target segmentation model, similarly, the training set can be input into the target segmentation model, and the focus loss function can be used to adjust the parameters of the target segmentation model. In the parameter adjustment process, the adjustable parameters of the target segmentation model include the learning rate, and one or more parameters of the target segmentation model including the learning rate can be adjusted according to the loss value obtained based on the focus loss function until the target segmentation model meets the second optimization requirement. The second optimization requirement may include: the predicted probability corresponding to the predicted category label obtained for the second training sample prediction meets the fourth preset reliability. For each second training sample in the training set, it is judged whether the predicted probability of the predicted category label corresponding to the second training sample meets the fourth preset reliability, and the predicted probability of the predicted category label corresponding to each verification sample mentioned above meets the first preset reliability, which will not be repeated here. It should be noted that the first preset probability threshold and / or the second preset probability threshold in the fourth preset reliability can be the same as or different from the numerical value of the first preset probability threshold and / or the second preset probability threshold in the first preset reliability. When the predicted probabilities of the predicted category labels corresponding to each second training sample meet the fourth preset confidence, it can be considered that the target segmentation model meets the second optimization requirement. More specifically, since the second training sample belongs to the small target area, it can be understood that when the target segmentation model meets the second optimization requirement, it can be considered that the target segmentation model can accurately segment out the small target area in the image. In other words, the second training operation is mainly used to enable the target segmentation model to accurately segment out the small target area in the image. Since the segmentation task for small target areas is more difficult, adjusting the learning rate can help the target segmentation model converge faster, and can better improve the performance of the target segmentation model, which is conducive to ensuring that the target segmentation model can segment out small target areas more accurately, so as to optimize the target segmentation model to meet the second training requirement. In some embodiments, in order to enhance feature extraction for small target areas, a target segmentation model including an attention mechanism module or a pyramid pooling module can be used.

[0056] The above technical solution optimizes the target segmentation model by executing the first training operation and the second training operation, which can take into account the performance of the target segmentation model in segmenting large target areas and segmenting small target areas. The trained target segmentation model has better performance and is not easy to lose small target areas.

[0057] Optionally, the balancing factor adopted by the focal loss function in the first training operation is smaller than the balancing factor adopted in the second training operation. The formula of the focal loss function is as follows:

[0058] FL(p t )=-α t (1-p t ) γ log(p t )

[0059] In the formula, p t represents the sample rate of correct predictions. The sample rate of correct predictions is equal to the ratio of the number of training samples whose predicted probability corresponding to the predicted class label obtained by prediction conforms to the corresponding preset confidence level to the total number of training samples. α t represents the balancing factor, and γ represents the preset adjustment factor.

[0060] Exemplarily, the first training samples whose predicted probability corresponding to the predicted class label obtained by prediction conforms to the third preset confidence level, and the second training samples whose predicted probability corresponding to the predicted class label obtained by prediction conforms to the fourth preset confidence level can be used as samples with correct predictions. The ratio between the number of samples with correct predictions and the total number of the first training samples and the second training samples can be used as the sample rate of correct predictions. The preset adjustment factor γ can be defined by the user according to the actual situation, and the specific value of the preset adjustment factor γ is not limited in the embodiments of the present invention. In both the first training operation and the second training operation, the above-mentioned focal loss function can be used to calculate the loss values of the first training samples and the second training samples. The balancing factor α t adopted by the focal loss function in the first training operation is smaller than the α t adopted in the second training operation.

[0061] In the above technical solution, the focal loss function is used to optimize the target segmentation model, which can better adapt to the problem of the imbalance in the number of the first training samples and the second training samples, and is beneficial to improving the robustness of the target segmentation model. On the other hand, by setting the preset adjustment factor of the focal loss function higher in the second training operation, this helps to make the target segmentation model more sensitive to the second training samples with smaller regional areas, so as to ensure that the trained target segmentation model can pay attention to the small-area regions in the image.

[0062] Optionally, obtain an annotation set based on a preset remote sensing image, including: determining reliable sample data and sample data to be annotated based on the preset remote sensing image, where the reliable sample data includes target reliable samples and class labels of the target reliable samples, the sample data to be annotated includes samples to be annotated, and the target reliable samples and the samples to be annotated are image regions in the preset remote sensing image; training an annotation model using the images to which the target reliable samples belong; inputting the images to which the samples to be annotated belong into the annotation model to obtain re-annotated samples, and when the number of target reliable samples does not meet the preset quantity requirement, adding the reliable samples in the re-annotated samples to the target reliable samples, and returning to execute the step of training the annotation model using the images to which the target reliable samples belong until the number of target reliable samples meets the preset quantity requirement, and then determining the images to which the target reliable samples belong as the images included in the annotation set; where the reliable samples in the re-annotated samples include samples whose predicted class labels output by the annotation model meet the preset label requirements, and samples obtained after manual re-annotation of samples whose predicted class labels output by the annotation model do not meet the preset label requirements.

[0063] Exemplarily, in step S110, reliable sample data and to-be-annotated sample data can be determined based on a preset remote sensing image. The reliable sample data can include target reliable samples and class labels of the target reliable samples. The target reliable samples can be image regions in the preset remote sensing image, and the class labels of the target reliable samples can indicate the preset classes corresponding to the target reliable samples. The to-be-annotated sample data can include to-be-annotated samples, and the to-be-annotated samples can be image regions in the preset remote sensing image. It can be understood that the target reliable samples and the to-be-annotated samples are different image regions in the preset remote sensing image. The image to which the target reliable samples belong can be used to train the annotation model. The image to which the target reliable samples belong can be an image patch in the preset remote sensing image, and the target reliable samples can be image regions in the image patch. The annotation model can be a deep learning network, such as EfficientNet, Vision Transformers, etc. After each training of the annotation model, the to-be-annotated samples can be input into the annotation model to obtain re-annotated samples. In the re-annotated samples, reliable samples and unreliable samples can be included. Specifically, the reliable samples can include samples whose predicted class labels output by the annotation model meet the preset label requirements. In some embodiments, for each re-annotated sample, the preset label requirements can include, for example, that each pixel in the re-annotated sample has a corresponding predicted class label, and / or the predicted class labels of the pixels in the re-annotated sample are all the same, and / or the ratio between the number of pixels with the predicted class label with the largest number of corresponding pixels in the re-annotated sample and the total number of pixels in the re-annotated sample is greater than or equal to a preset ratio threshold. By detecting the predicted class labels of the re-annotated samples, it can be determined whether the re-annotated samples belong to reliable samples. In other embodiments, whether the predicted class labels output by the annotation model for the re-annotated samples meet the preset label requirements can also be determined by the user himself / herself. Specifically, the user can determine by himself / herself whether there are errors or omissions in the predicted class labels of the re-annotated samples. The reliable samples in the re-annotated samples can also include samples obtained after manual re-annotation of samples whose predicted class labels output by the annotation model do not meet the preset label requirements. The manually re-annotated samples can be part of all the samples whose predicted class labels output by the annotation model do not meet the preset label requirements. Exemplarily, when the number of target reliable samples does not meet the preset quantity requirement, the reliable samples in the re-annotated samples can be added to the target reliable samples, and the step of training the annotation model using the image to which the target reliable samples belong can be returned for execution. It can be understood that during the process of training the annotation model, the number of target reliable samples gradually increases. When the number of target reliable samples meets the preset quantity requirement, the image to which the target reliable samples belong can be used as the image included in the annotation set. The preset quantity requirement can include, for example, that the number of target reliable samples is greater than or equal to a preset sample quantity threshold.

[0064] The above technical solution trains the annotation model by using the images to which the target reliable samples belong. When the number of target reliable samples does not meet the preset quantity requirement, the reliable samples in the relabeled samples are added to the target reliable samples. This can optimize the annotation model multiple times when the initial number of target reliable samples is small, and the number of target reliable samples used to optimize the annotation model gradually increases each time. This is beneficial to improving the performance of the target model. In addition, after each optimization of the annotation model, more target reliable samples can be obtained, which can quickly and efficiently obtain target reliable samples that meet the preset quantity requirement, so that an annotation set can be obtained quickly and efficiently.

[0065] Optionally, training the annotation model by using the images to which the target reliable samples belong includes: performing image enhancement on the images to which the target reliable samples belong; inputting the images to which the target reliable samples belong before image enhancement and the images to which the target reliable samples belong after image enhancement into the annotation model to calculate the consistency loss value of the target reliable samples before and after image enhancement; and using the consistency loss value to adjust the parameters of the annotation model.

[0066] Exemplarily, during the process of training the annotation model, image enhancement can be performed on the images to which the target reliable samples belong. The image enhancement method can specifically include one or more of methods such as rotation, cropping, translation, scaling, color transformation, etc. It can be understood that the predicted class labels of the corresponding image regions in the images to which the target reliable samples belong before and after image enhancement are the same. Inputting the images to which the target reliable samples belong before image enhancement and the images to which the target reliable samples belong after image enhancement into the annotation model can calculate the consistency loss value of the target reliable samples before and after image enhancement. The consistency loss value can be used to adjust the parameters of the annotation model. Specifically, when training the annotation model by using the images to which the target reliable samples belong each time, when the consistency loss value is less than or equal to the preset loss threshold, it can be determined that the current training is completed.

[0067] The above technical solution performs image enhancement on the images to which the target reliable samples belong and optimizes the annotation model according to the consistency loss value of the target reliable samples before and after image enhancement. This is beneficial to ensuring that the predicted class labels output by the trained annotation model for the images after image enhancement are highly consistent with the predicted class labels output for the original images, so that the stability of the trained annotation model can be relatively high.

[0068] Optionally, determining reliable sample data and sample data to be labeled based on a preset remote sensing image includes: dividing the preset remote sensing image into a plurality of image patches; inputting the plurality of image patches into a pre-trained segmentation model to obtain the image patch segmentation results of the plurality of image patches output by the pre-trained segmentation model, where the image patch segmentation results are used to indicate the predicted class labels of each image region included in each of the plurality of image patches; in response to a relabeling operation of the user on at least some of the plurality of image patches based on the predicted class labels predicted by the pre-trained segmentation model, determining new predicted class labels for at least some of the image regions in at least some of the image patches; in response to the user's determination operation, determining at least some of the image regions with existing predicted class labels as target reliable samples, and determining at least some of the image regions without existing predicted class labels as samples to be labeled.

[0069] Exemplarily, when determining reliable sample data and sample data to be labeled based on a preset remote sensing image, the preset remote sensing image can be segmented into multiple image patches. For example, the preset remote sensing image can be segmented into image patches of 1024*1024. After obtaining the multiple image patches segmented from the preset remote sensing image, the multiple image patches can be input into a pre-trained segmentation model to obtain the image patch segmentation results of the multiple image patches output by the pre-trained segmentation model. Specifically, the pre-trained segmentation model can be, for example, a convolutional neural network such as a Residual Network (ResNet) or a U-Net that can be used for image segmentation. The image patch segmentation results can indicate the predicted class labels of each image region included in each image patch. It should be noted that the predicted class labels of each image region output by the pre-trained segmentation model may not be the true classes of the corresponding image regions. In other words, the predicted class labels output by the pre-trained segmentation model can be regarded as "pseudo-labels". The user can perform a relabeling operation on at least some of the multiple image patches based on the predicted class labels predicted by the pre-trained segmentation model. For example, predicted class labels can be supplemented for image regions lacking predicted class labels. For another example, predicted class labels whose indicated classes are inconsistent with the actual classes of the corresponding image regions can be corrected. For another example, the boundaries of the image regions corresponding to a certain predicted class label can be adjusted. In response to the user's relabeling operation, new predicted class labels for at least some image regions in at least some of the image patches can be determined. The user can perform a first determination operation to determine the image regions corresponding to at least some of the predicted class labels output by the pre-trained segmentation model, and / or the image regions corresponding to at least some of the predicted class labels re-determined through the relabeling operation as the image regions currently having predicted class labels, and other image regions can be determined as the image regions currently having no predicted class labels. In response to the user's second determination operation, the image regions indicated by the second determination operation in the image regions currently having predicted class labels can be determined as target reliable samples, and the image regions indicated by the second determination operation in the image regions currently having no predicted class labels can be determined as samples to be labeled. In some embodiments, the first determination operation and the second determination operation can be the same operation, and the determination operation can include: determining the image regions corresponding to at least some of the predicted class labels output by the pre-trained segmentation model, and / or the image regions corresponding to at least some of the predicted class labels re-determined through the relabeling operation as target reliable samples, and determining at least some of the other image regions other than the target reliable samples as samples to be labeled.

[0070] The above technical solution can use the predicted class labels output by the pre-trained segmentation model for reference, which can greatly reduce the workload of the user for segmenting and labeling each image region one by one, thereby saving the time for obtaining reliable sample data and sample data to be labeled.

[0071] Optionally, the training set includes a first training sample and a second training sample, where the first training sample and the second training sample are image regions in the training images, and the area of the region of the first training sample is larger than that of the second training sample. After determining the image to which the target reliable sample belongs as the image included in the annotation set, the method further includes: performing image enhancement on the second training sample in the training set; obtaining a new training image based on the initial training image to which the target background sample belongs by randomly embedding the image-enhanced second training sample into any target background sample in the training set, and adding the new training image to the training set or replacing the initial training image with the new training image, where the area of the region occupied by the target background sample in the initial training image is the largest or larger than a preset area threshold.

[0072] Exemplarily, when the training set includes a first training sample and a second training sample, image enhancement can be performed on the second training sample in the training set, and the image enhancement method can include one or more of methods such as rotation, flipping, scaling, cropping, adjusting brightness, adjusting contrast, and adjusting saturation. After obtaining the image-enhanced second training sample, the image-enhanced second training sample can be randomly embedded into any target background sample in the training set, and a new training image can be obtained based on the initial training image to which the target background sample belongs. It can be understood that the second training sample is embedded in the corresponding target background sample in the new training image. Specifically, the target background sample can be the image region with the largest area occupied in the initial training image, or the target background sample can be the image region with an area larger than a preset area threshold in the initial training image.

[0073] The above technical solution can effectively solve the problem that the number of second training samples with a small area in the training set is small, which is beneficial to enhancing the diversity of training samples in the training set. By randomly selecting a target background sample in the initial training image and embedding the image-enhanced second training sample therein, more real scenarios can be simulated, and it helps the annotation model to pay attention to small target regions.

[0074] Please refer to Figure 2 shown, which is a schematic block diagram of an image segmentation device according to an embodiment of the present invention. On the other hand, according to the present invention, a remote sensing image segmentation device based on semi-supervised learning is also provided. The device 200 includes:

[0075] An acquisition module 210, configured to obtain an annotation set based on a preset remote sensing image, where the annotation set includes a training set and a validation set. The training set includes at least one training image and class labels of each image region on at least one training image, and the validation set includes at least one validation image and class labels of each image region on at least one validation image;

[0076] A training module 220 is configured to train at least one target segmentation model using a training set, and the at least one target segmentation model is used to predict the categories of the respective image regions of the input image;

[0077] A first input module 230 is configured to input a validation set into the at least one target segmentation model to obtain a predicted segmentation result of the validation set. When the predicted segmentation result of the validation set does not meet the preset validation requirements, the validation images to which the validation samples whose sample segmentation results in the validation set meet the preset prediction requirements belong, and the predicted class labels corresponding to the validation samples are updated to the training set, and the step of training the at least one target segmentation model using the training set is returned to be executed until the predicted classification result of the validation set meets the preset validation requirements. Wherein, the predicted segmentation result of the validation set includes the sample segmentation results corresponding to the respective validation samples in the validation set, the sample segmentation result is used to indicate the predicted class label corresponding to the validation sample, the validation sample is an image region in the validation image, the preset prediction requirements include that the prediction probability of the predicted class label corresponding to the validation sample meets a first preset confidence level, and the preset validation requirements include that the prediction probabilities of the predicted class labels corresponding to the respective validation samples in the validation set meet a second preset confidence level;

[0078] A second input module 240 is configured to input the remote sensing image to be measured into the at least one target segmentation model to obtain a predicted segmentation result of the remote sensing image to be measured.

[0079] Please refer to Figure 3 As shown, it is a schematic block diagram of an electronic device according to an embodiment of the present invention. According to another aspect of the present invention, an electronic device is further provided. The electronic device 300 includes: a processor 310 and a memory 320. Wherein, computer program instructions are stored in the memory 320, and when the computer program instructions are run by the processor 310, they are used to execute the above-mentioned remote sensing image segmentation method based on semi-supervised learning.

[0080] According to yet another aspect of the present invention, a storage medium is further provided. Program instructions are stored on the storage medium, and when the program instructions are run by a computer or a processor, the computer or the processor is made to execute the corresponding steps of the above-mentioned remote sensing image segmentation method based on semi-supervised learning according to the embodiments of the present invention, and is used to implement the corresponding modules in the above-mentioned image segmentation device according to the embodiments of the present invention or the corresponding modules in the above-mentioned device for image segmentation. The storage medium may include, for example, a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.

[0081] According to another aspect of the present invention, there is also provided a computer program product, including computer program instructions which are used to execute the remote sensing image segmentation method based on semi-supervised learning as described above when running.

[0082] Those of ordinary skill in the art can understand the specific implementation and beneficial effects of the above-mentioned remote sensing image segmentation device, electronic device, storage medium and computer program product based on semi-supervised learning by reading the above specific description. For the sake of brevity, they will not be elaborated here.

[0083] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely exemplary and are not intended to limit the scope of the present invention thereto. Those of ordinary skill in the art can make various changes and modifications therein without departing from the scope and spirit of the present invention. All such changes and modifications are intended to be included within the scope of the present invention as claimed in the appended claims.

[0084] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0085] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0086] In the specification provided here, a large number of specific details are described. However, it can be understood that the embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures and technologies are not shown in detail so as not to obscure the understanding of this specification.

[0087] Similarly, it should be understood that, for the purpose of streamlining the present invention and assisting in understanding one or more of the various inventive aspects, in the description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, the method of the present invention should not be construed as reflecting an intention that the claimed invention requires more features than those expressly recited in each claim. Rather, as reflected by the corresponding claims, the inventive point lies in that the corresponding technical problems can be solved by features fewer than all the features of a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim itself serves as a separate embodiment of the present invention.

[0088] Those skilled in the art can understand that, except for mutual exclusion between features, any combination can be adopted for all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.

[0089] In addition, those skilled in the art can understand that, although some of the embodiments herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means being within the scope of the present invention and forms different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.

[0090] The various component embodiments of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some of the modules in the image segmentation device according to the embodiments of the present invention. The present invention can also be implemented as a device program (e.g., a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0091] It should be noted that the above embodiments are illustrative of the present invention rather than restrictive of the present invention, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names.

[0092] The above is only a specific embodiment or an illustration of the specific embodiment of the present invention, and the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. The protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A remote sensing image segmentation method based on semi-supervised learning, characterized in that: The method comprises: Acquire an annotation set based on a preset remote sensing image, the annotation set comprising a training set and a verification set, the training set comprising at least one training image and a category label of each image region on the at least one training image, and the verification set comprising at least one verification image and a category label of each image region on the at least one verification image; Using the training set to train at least one target segmentation model, the at least one target segmentation model is used to predict the category of each image region of the input image; The verification set is input into the at least one target segmentation model to obtain the predicted segmentation result of the verification set. When the predicted segmentation result of the verification set does not meet the preset verification requirement, the verification image to which the verification sample whose sample segmentation result in the verification set meets the preset prediction requirement and the predicted category label corresponding to the verification sample are updated to the training set, and the step of training at least one target segmentation model using the training set is returned to be executed until the predicted classification result of the verification set meets the preset verification requirement, wherein the predicted segmentation result of the verification set includes the sample segmentation result corresponding to each verification sample in the verification set, the sample segmentation result is used to indicate the predicted category label corresponding to the verification sample, the verification sample is an image area in the verification image, the preset prediction requirement includes that the predicted probability of the predicted category label corresponding to the verification sample meets the first preset confidence, and the preset verification requirement includes that the predicted probability of the predicted category label corresponding to each verification sample in the verification set meets the second preset confidence; The remote sensing image to be measured is input into the at least one target segmentation model to obtain a predicted segmentation result of the remote sensing image to be measured.

2. The method according to claim 1, characterized in that: The number of the target segmentation models is greater than 1, the network architectures of the target segmentation models are different, and the inputting the validation set into the at least one target segmentation model to obtain the predicted segmentation result of the validation set includes: Inputting the verification set into each target segmentation model respectively to obtain a model segmentation result output by each target segmentation model, wherein the model segmentation result is used to indicate the predicted probability that each pixel in the verification set belongs to each preset category; Determine the target weight of each target segmentation model according to the model segmentation results output by each target segmentation model; For each preset category among multiple preset categories, for each pixel in the verification set, a weighted sum is performed on the predicted probability that the pixel belongs to the preset category according to the target weights of each target segmentation model, and the preset category with the largest weighted result is determined as the predicted category label of the pixel, the predicted category label corresponding to the pixel belongs to one of the multiple preset categories, and the predicted category label corresponding to the verification sample is the predicted category label of each pixel contained in the verification sample.

3. The method according to claim 2, characterized in that The determining of the target weights of the target segmentation models according to the model segmentation results outputted by the target segmentation models respectively includes: Determine the overall accuracy of each target segmentation model according to the model segmentation results output by each target segmentation model; The target weight of each target segmentation model is determined one-to-one according to the overall accuracy of each target segmentation model.

4. The method according to claim 2, characterized in that: For any pixel, the predicted probability that the pixel belongs to the preset category is weighted by the following formula, and the preset category with the largest weighted result is determined as the predicted category label of the pixel: Where A represents the predicted category label of the pixel, n represents the number of target segmentation models, and W i represents the target weight of the i-th target segmentation model, P i (x) represents the probability that the pixel belongs to the xth preset category in the model segmentation result output by the i-th target segmentation model.

5. The method according to any one of claims 1 to 4, characterized in that: The training set includes a first training sample and a second training sample, the first training sample and the second training sample are image regions in the training image, and the region area of ​​the first training sample is larger than the region area of ​​the second training sample; The adopting the training set to train at least one target segmentation model comprises: Performing a first training operation includes: for each target segmentation model, Inputting the training set into the target segmentation model; The target segmentation model is adjusted using a focal loss function until the target segmentation model meets a first optimization requirement, wherein during the parameter adjustment process, the learning rate of the target segmentation model maintains a preset standard learning rate unchanged, and the first optimization requirement includes that the prediction probability corresponding to the predicted category label obtained by predicting the first training sample meets a third preset confidence: After performing the first training operation, performing the second training operation includes: for each target segmentation model, Inputting the training set into the target segmentation model; The target segmentation model is adjusted using the focal loss function until the target segmentation model meets the second optimization requirement, wherein the parameters adjusted in the parameter adjustment process include the learning rate, and the second optimization requirement includes that the predicted probability corresponding to the predicted category label obtained for the second training sample prediction meets the fourth preset confidence.

6. The method according to claim 5, characterized in that The balance factor used by the focus loss function in the first training operation is less than the balance factor used in the second training operation, and the formula of the focus loss function is as follows: FL(p t )=-a t (1-p t ) γ log(p t ) In the formula, p t It represents the correct prediction sample rate, which is equal to the ratio of the number of training samples whose prediction probability corresponding to the predicted category label obtained by prediction meets the corresponding preset confidence to the total number of training samples, α t represents the balance factor, and γ represents the preset adjustment factor.

7. The method according to any one of claims 1 to 4, characterized in that: The step of obtaining a label set based on a preset remote sensing image includes: Determine reliable sample data and sample data to be labeled based on the preset remote sensing image, wherein the reliable sample data includes target reliable samples and category labels of the target reliable samples, and the sample data to be labeled includes samples to be labeled, and the target reliable samples and the samples to be labeled are image areas in the preset remote sensing image; Using the image of the target reliable sample to train the annotation model; Input the image to which the sample to be labeled belongs into the labeling model to obtain a re-labeled sample; when the number of the target reliable samples does not meet the preset number requirement, add the reliable samples in the re-labeled sample to the target reliable sample, and return to the step of training the labeling model with the image to which the target reliable sample belongs, until the number of the target reliable samples meets the preset number requirement, and determine the image to which the target reliable sample belongs as the image included in the labeling set; Among them, the reliable samples in the re-labeled samples include samples whose predicted category labels output by the labeling model meet the preset label requirements, and samples obtained after manual re-labeling of samples whose predicted category labels output by the labeling model do not meet the preset label requirements.

8. The method according to claim 7, characterized in that The step of training the annotation model using the image to which the target reliable sample belongs includes: Performing image enhancement on the image to which the target reliable sample belongs; Inputting the image to which the target reliable sample belongs before image enhancement and the image to which the target reliable sample belongs after image enhancement into the annotation model to calculate the consistency loss value of the target reliable sample before and after image enhancement; The consistency loss value is used to adjust the parameters of the annotation model.

9. The method according to claim 7, characterized in that: The determining of reliable sample data and sample data to be labeled based on the preset remote sensing image includes: Segmenting the preset remote sensing image into a plurality of image blocks; Inputting the plurality of image blocks into a pre-trained segmentation model to obtain image block segmentation results of the plurality of image blocks output by the pre-trained segmentation model, wherein the image block segmentation results are used to indicate predicted category labels of each image region included in each image block of the plurality of image blocks; In response to a re-labeling operation of at least part of the plurality of image blocks by a user based on the predicted category labels predicted by the pre-trained segmentation model, determining new predicted category labels for at least part of the image regions in the at least part of the image blocks; In response to a user's determination operation, at least a portion of the image region currently having a predicted category label is determined as the target reliable sample, and at least a portion of the image region currently not having a predicted category label is determined as the sample to be labeled.

10. The method according to claim 7, characterized in that The training set includes a first training sample and a second training sample, the first training sample and the second training sample are image regions in the training image, and the region area of ​​the first training sample is larger than the region area of ​​the second training sample; After determining that the image to which the target reliable sample belongs is an image included in the annotation set, the method further includes: Performing image enhancement on the second training sample in the training set; A new training image is obtained based on an initial training image to which the target background sample belongs by randomly embedding the image-enhanced second training sample into any target background sample in the training set, and the new training image is added to the training set or the initial training image is replaced by the new training image, wherein the area occupied by the target background sample in the initial training image is the largest or greater than a preset area threshold.

11. A remote sensing image segmentation device based on semi-supervised learning, characterized in that: The device comprises: An acquisition module is used to acquire a label set based on a preset remote sensing image, wherein the label set includes a training set and a verification set, wherein the training set includes at least one training image and a category label of each image region on the at least one training image, and the verification set includes at least one verification image and a category label of each image region on the at least one verification image; A training module, used to train at least one target segmentation model using the training set, wherein the at least one target segmentation model is used to predict the category of each image region of the input image; A first input module is used to input the verification set into the at least one target segmentation model to obtain a predicted segmentation result of the verification set. When the predicted segmentation result of the verification set does not meet the preset verification requirement, the verification image to which the verification sample whose sample segmentation result in the verification set meets the preset prediction requirement and the predicted category label corresponding to the verification sample are updated to the training set, and the step of training at least one target segmentation model using the training set is returned to be executed until the predicted classification result of the verification set meets the preset verification requirement, wherein the predicted segmentation result of the verification set includes the sample segmentation result corresponding to each verification sample in the verification set, the sample segmentation result is used to indicate the predicted category label corresponding to the verification sample, the verification sample is an image area in the verification image, the preset prediction requirement includes that the predicted probability of the predicted category label corresponding to the verification sample meets the first preset reliability, and the preset verification requirement includes that the predicted probability of the predicted category label corresponding to each verification sample in the verification set meets the second preset reliability; The second input module is used to input the remote sensing image to be measured into the at least one target segmentation model to obtain the predicted segmentation result of the remote sensing image to be measured.

12. An electronic device comprising a processor and a memory, characterized in that: The memory stores computer program instructions, which are used by the processor to execute the remote sensing image segmentation method based on semi-supervised learning as described in any one of claims 1 to 10 when the processor is running.

13. A storage medium having program instructions stored thereon, characterized in that: The program instructions are used to execute the remote sensing image segmentation method based on semi-supervised learning as described in any one of claims 1 to 10 when running.

14. A computer program product comprising computer program instructions, characterized in that The computer program instructions are used to execute the remote sensing image segmentation method based on semi-supervised learning as described in any one of claims 1 to 10 when running.

Citation Information

Patent Citations

  • Model iterative training method and system based on automatic labeling

    CN112001407A

  • Image segmentation method, model training method, related device and electronic equipment

    CN115908444A

  • Automatic driving model generation method and device, vehicle and storage medium

    CN116626670A

  • Low-cost cell segmentation method, system and device based on active learning

    CN119295480A

  • Processing a database

    WO2008029154A1