Semi-supervised remote sensing image semantic segmentation model training method and device

By calculating the consistency score of remote sensing image data and filtering labelless data, the problem of performance degradation of semi-supervised remote sensing image semantic segmentation model under large-scale labelless data is solved, and higher robustness and generalization performance are achieved.

CN120047948AActive Publication Date: 2025-05-27WUHAN UNIV

Patent Information

Application Number
CN202510117766.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The existing semi-supervised remote sensing image semantic segmentation models have deteriorated performance in large-scale unlabeled data scenarios, mainly due to uneven unlabeled data quality, the model is susceptible to interference from unknown categories, and improper training strategies, resulting in semantic drift.

Method used

By calculating the consistency scores of labeled remote sensing image data and labelless remote sensing image data, filtering and sorting labelless remote sensing image data, giving priority to using data with high consistency scores for training, and gradually introducing data with low consistency scores to alleviate the semantic drift problem.

Benefits of technology

It effectively alleviates the performance degradation of semantic segmentation models in large-scale label-free data scenarios, improves the robustness and generalization performance of the model, reduces interference from unknown categories, and ensures the accuracy of decision boundaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047948A_ABST
    Figure CN120047948A_ABST
Patent Text Reader

Abstract

The invention discloses a semi-supervised remote sensing image semantic segmentation model training method and device, and belongs to the technical field of computer vision. The method comprises the following steps: calculating consistency scores of unlabeled remote sensing image data based on labeled remote sensing image data, and screening the consistency scores; inputting the screened label-free remote sensing image data into a teacher module to obtain a teacher prediction result; inputting the labeled remote sensing image data and the screened unlabeled remote sensing image data into a student module, and supervising the student module through a teacher prediction result to obtain a student prediction result; and based on the teacher prediction result and the student prediction result, constructing a loss function of the semantic segmentation model so as to update parameters of the semantic segmentation model. According to the method, the training effect of the semantic segmentation model is improved from the three aspects of no-label training data screening, model design and training strategies, and the problem that the performance of the semantic segmentation model is reduced along with increase of no-label data can be effectively relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and relates to a semi-supervised remote sensing image semantic segmentation model training method and device. Background Art

[0002] Semantic segmentation of remote sensing images has important applications in tasks such as land cover classification, but its annotation process is expensive and time-consuming, resulting in a shortage of labeled data. To address this problem, semi-supervised semantic segmentation provides a good solution. Due to limited data resources, current research simulates the optimal allocation under fixed resources by dynamically adjusting the ratio of labeled data to unlabeled data to achieve the best performance. However, this approach limits the exploration of the performance improvement potential of the model in large-scale unlabeled data scenarios to a certain extent. For example, as the amount of unlabeled data increases, the model often cannot effectively process more unlabeled information.

[0003] Studies have shown that the main reasons for the performance degradation are the uneven quality of unlabeled data, the model is easily disturbed by unknown categories, and improper training strategies lead to semantic drift. First, the semi-supervised semantic segmentation task assumes that labeled data and unlabeled data share the same distribution. However, in practical applications, this assumption is often difficult to hold. Unlabeled data may contain samples with large distribution differences from the training data, which interferes with the pseudo-label generation and feature learning of the model; second, the existing semi-supervised semantic segmentation model generates labels for unlabeled data through a pseudo-label mechanism and uses these labels in training. However, since unlabeled data may contain unknown categories, the model is prone to confuse these unknown categories with known categories, thereby affecting the classification and segmentation accuracy; finally, the current training method may lead to the accumulation of erroneous pseudo-labels, causing the decision boundary of the model to gradually deviate from the actual data distribution, thereby exacerbating the semantic drift problem. The above three problems do not exist in isolation throughout the training process, but affect each other. As the model is iteratively trained, these problems will gradually accumulate at different stages, eventually affecting the overall performance of the model. Summary of the invention

[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and to provide a semi-supervised remote sensing image semantic segmentation model training method and device, which can effectively alleviate the problem of semantic segmentation model performance degradation caused by the increase of unlabeled data.

[0005] To achieve the above object, the present invention is implemented by adopting the following technical solutions:

[0006] In a first aspect, the present invention provides a semi-supervised remote sensing image semantic segmentation model training method, comprising:

[0007] Based on the labeled remote sensing image data, the consistency score of the unlabeled remote sensing image data is calculated;

[0008] The unlabeled remote sensing image data are screened according to their consistency scores;

[0009] Input the filtered unlabeled remote sensing image data into the teacher module of the semantic segmentation model to obtain the teacher prediction results;

[0010] Inputting the labeled remote sensing image data and the screened unlabeled remote sensing image data into a student module of a semantic segmentation model, and supervising the student module through the teacher prediction results to obtain the student prediction results;

[0011] Based on the teacher prediction results and the student prediction results, construct a loss function of the semantic segmentation model;

[0012] The parameters of the semantic segmentation model are updated through the loss function of the semantic segmentation model.

[0013] Furthermore, based on the labeled remote sensing image data, the consistency score of the unlabeled remote sensing image data is calculated, including:

[0014] Training a fully supervised semantic segmentation network model using the labeled remote sensing image data;

[0015] Freezing decoder parameters of the fully supervised semantic segmentation network model;

[0016] Inputting the labeled remote sensing image data into the encoder of the fully supervised semantic segmentation network model to obtain features of the labeled remote sensing image data;

[0017] Inputting the unlabeled remote sensing image data into the encoder of the fully supervised semantic segmentation network model to obtain features of the unlabeled remote sensing image data;

[0018] By measuring the consistency of the features of the labeled remote sensing image data and the features of the unlabeled remote sensing image data, a consistency score of the unlabeled remote sensing image data is obtained.

[0019] Furthermore, by measuring the consistency of the features of the labeled remote sensing image data and the features of the unlabeled remote sensing image data, a consistency score of the unlabeled remote sensing image data is obtained, including:

[0020] ,

[0021] in, score [ i ] Indicates The consistency score of unlabeled remote sensing image data; Represents mapping the eigenvector to the reproducing kernel Hilbert space ; Represents the features of labeled remote sensing image data, Indicates Features of unlabeled remote sensing image data; Represents the number of unlabeled remote sensing image data.

[0022] Furthermore, the unlabeled remote sensing image data are screened according to the consistency scores of the unlabeled remote sensing image data, including:

[0023] The unlabeled remote sensing image data set selected by the teacher module and the student module in the current training round for:

[0024] ,

[0025] ,

[0026] in, Represents unlabeled remote sensing image data arranged in descending order according to consistency scores. i ∈ [ m , m + increment ] ; Represents a collection of unlabeled remote sensing image data; represents the initial data volume, represents the amount of new data in each training round, Indicates the current training round.

[0027] Furthermore, the teacher module and the student module each include an encoder and a decoder;

[0028] The decoder of the student module includes a main decoder and an auxiliary decoder;

[0029] The prediction results of the main decoder are supervised by pseudo labels generated by the teacher module;

[0030] The main decoder classifies the pixels of the unlabeled remote sensing image data input to the encoder of the student module into known categories and unknown categories;

[0031] The auxiliary decoder is used to determine the location of the unknown class.

[0032] Furthermore, it also includes: constraining the features of the unlabeled remote sensing image data output by the encoder of the student module, including:

[0033] If the input to the auxiliary decoder is the features of labeled remote sensing image data, then the supervision signal of the auxiliary decoder ;

[0034] If the input to the auxiliary decoder is the features of unlabeled remote sensing image data, then the supervision signal of the auxiliary decoder ,

[0035] in, Indicates confidence, represents the confidence threshold;

[0036] Based on the location of unknown categories determined by the auxiliary decoder, the features of unlabeled remote sensing image data are divided into known category features and unknown category features ;

[0037] Use K-means clustering algorithm to calculate the regional center of each known category feature and regional centers of unknown class characteristics ;

[0038] Will The clustering result of is regarded as an unknown category and regarded as a positive sample. The clustering results of are regarded as negative samples, then the feature loss function between known categories and unknown categories is for:

[0039] ,

[0040] ,

[0041] ,

[0042] in, represents the feature similarity loss between positive samples, Represents the feature similarity loss between negative samples; express and The cosine similarity of represents a hyperparameter.

[0043] Furthermore, the loss function of the semantic segmentation model is:

[0044] ,

[0045] in, represents the supervision loss function; represents the loss function of the auxiliary decoder; represents the consistency loss function; and Represents the weight of the corresponding loss;

[0046] The auxiliary decoder is supervised by both labeled remote sensing image data and unlabeled remote sensing image data, and its loss function The calculation formula is:

[0047] ,

[0048] ,

[0049] L u =− 1 N ∑ i = 1 N [ y i log( p i ) + ( 1 − y i )log( 1 − p i )] ,

[0050] in, Represents the loss function of labeled remote sensing image data in the auxiliary decoder, represents the loss function of unlabeled remote sensing image data in the auxiliary decoder; represents the predicted value, represents the true value;

[0051] Consistency loss function The calculation formula is:

[0052] ,

[0053] in, Represents the unlabeled remote sensing image data set participating in the training in the current training round, represents the parameters of the student model, Represents the segmentation map output grid The pixel address of Indicated in The segmentation predictions from the teacher module, p θ s ( x )( ω ) ∈ [ 0 , 1 ] Y Indicated in The segmentation predictions from the student model are Indicated in The confidence of the segmentation prediction from the teacher module.

[0054] In a second aspect, the present invention further provides a training device for a semi-supervised remote sensing image semantic segmentation model, the device comprising:

[0055] A consistency score calculation module is used to calculate the consistency score of unlabeled remote sensing image data based on labeled remote sensing image data;

[0056] The unlabeled data screening module is used to screen the unlabeled remote sensing image data according to the consistency score of the unlabeled remote sensing image data;

[0057] The teacher prediction result acquisition module is used to input the filtered unlabeled remote sensing image data into the teacher module of the semantic segmentation model to obtain the teacher prediction result;

[0058] A student prediction result acquisition module is used to input the labeled remote sensing image data and the screened unlabeled remote sensing image data into the student module of the semantic segmentation model, and supervise the student module through the teacher prediction result to obtain the student prediction result;

[0059] A loss function construction module, used to construct a loss function of a semantic segmentation model based on the teacher prediction result and the student prediction result;

[0060] The model parameter updating module is used to update the parameters of the semantic segmentation model through the loss function of the semantic segmentation model.

[0061] In a third aspect, the present invention further provides a computer device, comprising:

[0062] Memory for storing computer programs;

[0063] A processor is used to execute the computer program to implement the steps of the above-mentioned semi-supervised remote sensing image semantic segmentation model training method.

[0064] In a fourth aspect, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the steps of the above-mentioned semi-supervised remote sensing image semantic segmentation model training method.

[0065] Compared with the prior art, the present invention has the following beneficial effects:

[0066] The semi-supervised remote sensing image semantic segmentation model training method provided by the present invention screens and sorts the unlabeled remote sensing image data by calculating the consistency degree of the labeled remote sensing image data and the unlabeled remote sensing image data to evaluate the distribution similarity between the labeled data and the unlabeled data. The present invention proposes a training strategy for the problem of semantic drift. In the early stage of training, the unlabeled remote sensing image data closely related to the distribution of the labeled remote sensing image data is given priority, that is, the unlabeled remote sensing image data with a high consistency score, ensuring that the semantic segmentation model learns in a clear and clear semantic context, which helps to establish an accurate decision boundary. As the training proceeds, more complex unlabeled remote sensing image data, that is, unlabeled remote sensing image data with a low consistency score, are gradually introduced. These data are relatively inconsistent with the distribution of the labeled remote sensing image data, which can enhance the robustness of the model and effectively reduce the risk of semantic drift when the semantic segmentation model encounters complex unlabeled data. In addition, the present invention develops a semi-supervised semantic segmentation network architecture based on consistency, and the decoder can identify unknown category areas and distinguish their features from known category features, thereby reducing the interference of unknown categories. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 A flowchart of a semi-supervised remote sensing image semantic segmentation model training method provided by an embodiment of the present invention;

[0068] Figure 2 A schematic diagram of a structure for obtaining a consistency score of unlabeled remote sensing image data in an embodiment of the present invention;

[0069] Figure 3 A schematic diagram of a framework for training a semantic segmentation model in an embodiment of the present invention;

[0070] Figure 4 A schematic diagram of the structure of a semi-supervised remote sensing image semantic segmentation model training device provided by an embodiment of the present invention;

[0071] Figure 5 An internal structure diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0072] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. The same reference numerals in the accompanying drawings indicate the same or similar components or parts. It should be understood by those skilled in the art that these drawings are not necessarily drawn to scale. The embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. In the absence of conflict, the embodiments of the present application and the technical features in the embodiments can be combined with each other.

[0073] The term "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0074] Embodiment 1:

[0075] like Figures 1 to 3 As shown, an embodiment of the present invention provides a semi-supervised remote sensing image semantic segmentation model training method. Figure 1 The flowchart is a flow chart of the semi-supervised remote sensing image semantic segmentation model training method. This flowchart only shows the logical sequence of the method described in this embodiment. Under the premise of no conflict, in other possible embodiments of the present invention, different Figure 1 The steps shown or described are accomplished in the order shown.

[0076] The semi-supervised remote sensing image semantic segmentation model training method provided in this embodiment can be applied to a terminal, and can be executed by a semi-supervised remote sensing image semantic segmentation model training device, which can be implemented by software and / or hardware, and the device can be integrated in a terminal.

[0077] The embodiment of the present invention is implemented using the Pytorch deep learning framework under the Windows operating system. The specific experimental environment configuration is as follows:

[0078]

[0079] See also Figure 1 The method of the embodiment of the present invention specifically includes the following steps:

[0080] Step 1: Based on the labeled remote sensing image data, calculate the consistency score of the unlabeled remote sensing image data.

[0081] Labeled remote sensing image data set of the present invention , unlabeled remote sensing image data set , where M represents the number of labeled remote sensing image data, N represents the number of unlabeled remote sensing image data, M< <N。

[0082] When it comes to semantic segmentation of remote sensing images in large-scale and complex scenes, the quality of unlabeled data varies, and not all unlabeled data need to be involved in training. Therefore, before model training, all unlabeled remote sensing image data are first evaluated for consistency.

[0083] like Figure 2 As shown, step 1 specifically includes the following steps:

[0084] Step 11: training a fully supervised semantic segmentation network model using the labeled remote sensing image data;

[0085] Step 12: Freeze the decoder parameters of the fully supervised semantic segmentation network model;

[0086] Step 13: Inputting the labeled remote sensing image data into the encoder of the fully supervised semantic segmentation network model to obtain the features of the labeled remote sensing image data;

[0087] Step 14: inputting the unlabeled remote sensing image data into the encoder of the fully supervised semantic segmentation network model to obtain the features of the unlabeled remote sensing image data;

[0088] Step 15: Obtaining a consistency score of the unlabeled remote sensing image data by measuring the consistency of the features of the labeled remote sensing image data and the features of the unlabeled remote sensing image data, including:

[0089] ,

[0090] in, score [ i ] Indicates The consistency score of unlabeled remote sensing image data; Represents mapping the eigenvector to the reproducing kernel Hilbert space ; Represents the features of labeled remote sensing image data, Indicates Features of unlabeled remote sensing image data; Represents the number of unlabeled remote sensing image data.

[0091] Step 2: Filter the unlabeled remote sensing image data according to their consistency scores.

[0092] According to the consistency scores of the unlabeled remote sensing image data obtained in step 1, all unlabeled remote sensing image data are sorted from high to low according to the consistency scores, and the unlabeled remote sensing image data are input into the semantic segmentation model for training in this order.

[0093] The set of unlabeled remote sensing image data selected by the semantic segmentation model in the current training round for:

[0094] ,

[0095] ,

[0096] in, Represents unlabeled remote sensing image data arranged in descending order according to consistency scores. i ∈ [ m , m + increment ] ; Represents a collection of unlabeled remote sensing image data; Indicates the initial data volume, represents the amount of new data in each training round, Indicates the current training round.

[0097] The present invention proposes a training strategy for the problem of semantic drift. In the early stage of semantic segmentation model training, priority is given to unlabeled remote sensing image data that is closely related to the distribution of labeled remote sensing image data, that is, unlabeled remote sensing image data with a high consistency score, to ensure that the semantic segmentation model is learned in a clear and unambiguous semantic context, which helps to establish an accurate decision boundary. As the training proceeds, more complex unlabeled remote sensing image data, that is, unlabeled remote sensing image data with a low consistency score, are gradually introduced. These data are relatively inconsistent with the distribution of labeled remote sensing image data, which can enhance the robustness of the semantic segmentation model. Through this progressive training strategy, starting from simple training data and gradually transitioning to more challenging training data, the risk of semantic drift when the model encounters complex unlabeled data is effectively reduced, and the adaptability to potential data changes can be improved while maintaining the stability of the semantic segmentation model, ultimately improving the generalization performance of the semantic segmentation model.

[0098] Step 3: Input the filtered unlabeled remote sensing image data into the teacher module of the semantic segmentation model to obtain the teacher prediction results.

[0099] Step 4: Input the labeled remote sensing image data and the screened unlabeled remote sensing image data into the student module of the semantic segmentation model, and supervise the student module through the teacher prediction results to obtain the student prediction results.

[0100] like Figure 3 As shown, the semantic segmentation model of the present invention includes a teacher module and a student module. The student module provides parameter updates to the teacher module, and the teacher module provides pseudo labels to the student module. The two interact with each other to achieve better training purposes.

[0101] Both the teacher module and the student module include an encoder and a decoder.

[0102] In an embodiment of the present invention, the labeled remote sensing image data and the filtered unlabeled remote sensing image data need to be preprocessed before being input into the model, including: cropping the labeled remote sensing image data and the filtered unlabeled remote sensing image data according to a size of 512×512; performing data enhancement on the cropped unlabeled remote sensing image data, specifically including weak enhancement (rotation, translation, etc.) and strong enhancement (color transformation, etc.), and inputting the weakly enhanced data into the encoder of the student module, and inputting the strongly enhanced data into the encoder of the teacher module.

[0103] In addition, before inputting the training data, the encoder of the teacher module and the encoder of the student module need to be randomly initialized to ensure that they are different.

[0104] The decoder of the student module is divided into a main decoder and an auxiliary decoder. The prediction result of the main decoder is supervised by the pseudo-label generated by the teacher module, and the output of the main decoder is , where B represents the batch size, K represents the number of known categories plus 1, H represents the height of the feature, and W represents the width of the feature. The main decoder of the present invention divides the pixel points of the unlabeled remote sensing image data input into the encoder of the student module into known categories and unknown categories, and designs a channel to store the unknown categories. In order to alleviate the interference of the unknown categories, an auxiliary decoder is also designed to determine the position of the unknown categories.

[0105] Specifically, the filtered unlabeled remote sensing image data is input into the teacher module of the semantic segmentation model to obtain the teacher prediction results. :

[0106] ,

[0107] in, represents the encoder of the teacher module, Represents the decoder of the teacher module.

[0108] For the student module, it is trained by both labeled remote sensing image data and filtered unlabeled remote sensing image data. First, it is necessary to identify the unknown category area:

[0109] ,

[0110] in, The mask representing the unknown category, represents the encoder of the student module, Represents the auxiliary decoder of the student module.

[0111] use Find the location of the unknown category and use it as a benchmark to classify the features of the unlabeled remote sensing image data into known category features and unknown category features :

[0112] [ F 1 , F 2 ,..., F k ] = F S u ⋅ mask ,

[0113] ,

[0114] in, The features of the unlabeled remote sensing image data output by the encoder of the student module are input into the main encoder and combined with Calculate the final prediction :

[0115] .

[0116] The present invention also constrains the features of the unlabeled remote sensing image data output by the encoder of the student module, including:

[0117] If the input to the auxiliary decoder is the features of labeled remote sensing image data, then the supervision signal of the auxiliary decoder ;

[0118] If the input to the auxiliary decoder is the features of unlabeled remote sensing image data, then the supervision signal of the auxiliary decoder ,

[0119] in, Indicates confidence, represents the confidence threshold;

[0120] Based on the location of the unknown category determined by the auxiliary decoder, the features of the unlabeled remote sensing image data are divided into known category features. and unknown category features ;

[0121] Use K-means clustering algorithm to calculate the regional center of each known category feature and regional centers of unknown class characteristics ;

[0122] Will The clustering result of is regarded as an unknown category and regarded as a positive sample. The clustering results of are regarded as negative samples, then the feature loss function between known categories and unknown categories is for:

[0123] ,

[0124] ,

[0125] ,

[0126] in, represents the feature similarity loss between positive samples, Represents the feature similarity loss between negative samples; express and The cosine similarity of Represents a hyperparameter.

[0127] Step 5: Based on the teacher prediction results and the student prediction results, construct a loss function of the semantic segmentation model.

[0128] The loss function of the semantic segmentation model of the present invention is:

[0129] ,

[0130] in, represents the supervision loss function; represents the loss function of the auxiliary decoder; represents the consistency loss function; and Represents the weight of the corresponding loss.

[0131] Supervised Loss Function The student module is mainly supervised by labeled data, which is calculated by multi-class cross entropy loss, which is defined as follows:

[0132] ,

[0133] Where M represents the number of samples, Indicates the number of categories; samples If the true category is equal to ,but Take 1, otherwise take 0; Represents an observation sample Belongs to category The loss guides the feature extraction of the student model by comparing the difference between the predicted probability distribution of the network and the true label distribution.

[0134] The auxiliary decoder is supervised by both labeled remote sensing image data and unlabeled remote sensing image data, and its loss function The calculation formula is:

[0135] ,

[0136] ,

[0137] L u =− 1 N ∑ i = 1 N [ y i log( p i ) + ( 1 − y i )log( 1 − p i )] ,

[0138] in, Represents the loss function of labeled remote sensing image data in the auxiliary decoder, represents the loss function of unlabeled remote sensing image data in the auxiliary decoder; represents the predicted value, represents the true value;

[0139] The teacher's prediction results supervise the student module, which is reflected in the consistency loss function , and its calculation formula is:

[0140] ,

[0141] in, Represents the unlabeled remote sensing image data set participating in the training in the current training round, represents the parameters of the student model, Represents the segmentation map output grid The pixel address of Indicated in The segmentation predictions from the teacher module, p θ s ( x )( ω ) ∈ [ 0 , 1 ] Y Indicated in The segmentation predictions from the student model are Indicated in The confidence of the segmentation prediction from the teacher module.

[0142] Step 6: Update the parameters of the semantic segmentation model through the loss function of the semantic segmentation model.

[0143] The batch size of the embodiment of the present invention is set to 16, which includes 8 labeled remote sensing image data and 8 unlabeled remote sensing image data. The teacher module and the student module are pre-trained on the labeled data for 5 cycles before starting training. The cluster learning method uses a stochastic gradient descent optimizer with the following settings: the base learning rate is 0.01, the momentum is 0.9, the weight decay is 0.0001, the training process lasts for 40 cycles, and the initial learning rate is 0.01; after each iteration, the learning rate is updated by the following formula: ; The momentum coefficient is set to 0.9 and the weight decay coefficient is set to 0.0001.

[0144] Embodiment 2:

[0145] Based on the same inventive concept as Example 1, the embodiment of the present invention also provides a semi-supervised remote sensing image semantic segmentation model training device for implementing the above-mentioned semi-supervised remote sensing image semantic segmentation model training method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above-mentioned method, so the specific limitations in the embodiment of the semi-supervised remote sensing image semantic segmentation model training device provided below can be referred to the limitations of the above-mentioned method for enhancing ionospheric data over the ocean, which will not be repeated here.

[0146] like Figure 4 As shown, an embodiment of the present invention provides a semi-supervised remote sensing image semantic segmentation model training device, comprising:

[0147] A consistency score calculation module is used to calculate the consistency score of unlabeled remote sensing image data based on labeled remote sensing image data;

[0148] The unlabeled data screening module is used to screen the unlabeled remote sensing image data according to the consistency score of the unlabeled remote sensing image data;

[0149] The teacher prediction result acquisition module is used to input the filtered unlabeled remote sensing image data into the teacher module of the semantic segmentation model to obtain the teacher prediction result;

[0150] A student prediction result acquisition module is used to input the labeled remote sensing image data and the screened unlabeled remote sensing image data into the student module of the semantic segmentation model, and supervise the student module through the teacher prediction result to obtain the student prediction result;

[0151] A loss function construction module, used to construct a loss function of a semantic segmentation model based on the teacher prediction result and the student prediction result;

[0152] The model parameter updating module is used to update the parameters of the semantic segmentation model through the loss function of the semantic segmentation model.

[0153] Embodiment 3:

[0154] The embodiment of the present invention further provides a computer device, which may be a server, and its internal structure diagram may be as shown in FIG. Figure 5 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the steps of the semi-supervised remote sensing image semantic segmentation model training method in the aforementioned embodiment are implemented.

[0155] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0156] Embodiment 4:

[0157] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the following method are implemented:

[0158] Based on the labeled remote sensing image data, the consistency score of the unlabeled remote sensing image data is calculated;

[0159] The unlabeled remote sensing image data are screened according to their consistency scores;

[0160] Input the filtered unlabeled remote sensing image data into the teacher module of the semantic segmentation model to obtain the teacher prediction results;

[0161] Inputting the labeled remote sensing image data and the screened unlabeled remote sensing image data into a student module of a semantic segmentation model, and supervising the student module through the teacher prediction results to obtain the student prediction results;

[0162] Based on the teacher prediction results and the student prediction results, construct a loss function of the semantic segmentation model;

[0163] The parameters of the semantic segmentation model are updated through the loss function of the semantic segmentation model.

[0164] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems, or computer program products, and therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0165] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0166] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0167] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0168] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the enlightenment of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, which all fall within the protection of the present invention.

Claims

1. A semi-supervised remote sensing image semantic segmentation model training method, characterized in that: include: Based on the labeled remote sensing image data, the consistency score of the unlabeled remote sensing image data is calculated; The unlabeled remote sensing image data are screened according to their consistency scores; Input the filtered unlabeled remote sensing image data into the teacher module of the semantic segmentation model to obtain the teacher prediction results; Inputting the labeled remote sensing image data and the screened unlabeled remote sensing image data into a student module of a semantic segmentation model, and supervising the student module through the teacher prediction results to obtain the student prediction results; Based on the teacher prediction results and the student prediction results, construct a loss function of the semantic segmentation model; The parameters of the semantic segmentation model are updated through the loss function of the semantic segmentation model.

2. The semi-supervised remote sensing image semantic segmentation model training method according to claim 1, characterized in that: Based on labeled remote sensing image data, the consistency score of unlabeled remote sensing image data is calculated, including: Training a fully supervised semantic segmentation network model using the labeled remote sensing image data; Freezing decoder parameters of the fully supervised semantic segmentation network model; Inputting the labeled remote sensing image data into the encoder of the fully supervised semantic segmentation network model to obtain features of the labeled remote sensing image data; Inputting the unlabeled remote sensing image data into the encoder of the fully supervised semantic segmentation network model to obtain features of the unlabeled remote sensing image data; By measuring the consistency of the features of the labeled remote sensing image data and the features of the unlabeled remote sensing image data, a consistency score of the unlabeled remote sensing image data is obtained.

3. The semi-supervised remote sensing image semantic segmentation model training method according to claim 2, characterized in that: By measuring the consistency of the features of the labeled remote sensing image data and the features of the unlabeled remote sensing image data, a consistency score of the unlabeled remote sensing image data is obtained, including: , in, Indicates The consistency score of unlabeled remote sensing image data; Represents mapping the eigenvector to the reproducing kernel Hilbert space ; Represents the features of labeled remote sensing image data, Indicates Features of unlabeled remote sensing image data; Represents the number of unlabeled remote sensing image data.

4. The semi-supervised remote sensing image semantic segmentation model training method according to claim 1, characterized in that: According to the consistency score of the unlabeled remote sensing image data, the unlabeled remote sensing image data is screened, including: The unlabeled remote sensing image data set selected by the teacher module and the student module in the current training round for: , , in, Represents unlabeled remote sensing image data arranged in descending order according to consistency scores. ; Represents a collection of unlabeled remote sensing image data; Indicates the initial data volume, represents the amount of new data in each training round, Indicates the current training round.

5. The semi-supervised remote sensing image semantic segmentation model training method according to claim 1, characterized in that: The teacher module and the student module both include an encoder and a decoder; The decoder of the student module includes a main decoder and an auxiliary decoder; The prediction results of the main decoder are supervised by pseudo labels generated by the teacher module; The main decoder classifies the pixels of the unlabeled remote sensing image data input to the encoder of the student module into known categories and unknown categories; The auxiliary decoder is used to determine the location of the unknown class.

6. The semi-supervised remote sensing image semantic segmentation model training method according to claim 5, characterized in that: Also includes: Constrain the features of the unlabeled remote sensing image data output by the encoder of the student module, including: If the input to the auxiliary decoder is the features of labeled remote sensing image data, then the supervision signal of the auxiliary decoder ; If the input to the auxiliary decoder is the features of unlabeled remote sensing image data, then the supervision signal of the auxiliary decoder , in, Indicates confidence, represents the confidence threshold; Based on the location of unknown categories determined by the auxiliary decoder, the features of unlabeled remote sensing image data are divided into known category features and unknown category features ; Use K-means clustering algorithm to calculate the regional center of each known category feature and regional centers of unknown class characteristics ; Will The clustering result of is regarded as an unknown category and regarded as a positive sample. The clustering results of are regarded as negative samples, then the feature loss function between known categories and unknown categories is for: , , , in, represents the feature similarity loss between positive samples, Represents the feature similarity loss between negative samples; express and The cosine similarity of represents a hyperparameter.

7. The semi-supervised remote sensing image semantic segmentation model training method according to claim 6, characterized in that: The loss function of the semantic segmentation model is: , in, represents the supervision loss function; represents the loss function of the auxiliary decoder; represents the consistency loss function; and Represents the weight of the corresponding loss; The auxiliary decoder is supervised by both labeled remote sensing image data and unlabeled remote sensing image data, and its loss function The calculation formula is: , , , in, Represents the loss function of labeled remote sensing image data in the auxiliary decoder, represents the loss function of unlabeled remote sensing image data in the auxiliary decoder; represents the predicted value, represents the true value; Consistency loss function The calculation formula is: , in, Represents the unlabeled remote sensing image data set participating in the training in the current training round, represents the parameters of the student model, Represents the segmentation map output grid The pixel address of Indicated in The segmentation predictions from the teacher module, Indicated in The segmentation predictions from the student model are Indicated in The confidence of the segmentation prediction from the teacher module.

8. A semi-supervised remote sensing image semantic segmentation model training device, characterized in that: include: A consistency score calculation module is used to calculate the consistency score of unlabeled remote sensing image data based on labeled remote sensing image data; The unlabeled data screening module is used to screen the unlabeled remote sensing image data according to the consistency score of the unlabeled remote sensing image data; The teacher prediction result acquisition module is used to input the filtered unlabeled remote sensing image data into the teacher module of the semantic segmentation model to obtain the teacher prediction result; A student prediction result acquisition module is used to input the labeled remote sensing image data and the screened unlabeled remote sensing image data into the student module of the semantic segmentation model, and supervise the student module through the teacher prediction result to obtain the student prediction result; A loss function construction module, used to construct a loss function of a semantic segmentation model based on the teacher prediction result and the student prediction result; The model parameter updating module is used to update the parameters of the semantic segmentation model through the loss function of the semantic segmentation model.

9. A computer device, characterized in that: include: Memory for storing computer programs; A processor, configured to execute the computer program to implement the steps of the semi-supervised remote sensing image semantic segmentation model training method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the semi-supervised remote sensing image semantic segmentation model training method described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method and device

    CN112132149A

  • Semi-supervised remote sensing image semantic segmentation method based on double consistency

    CN116416618A

  • Semi-supervised remote sensing image semantic segmentation method and system

    CN116597136A

  • Semi-supervised remote sensing image semantic segmentation method and system based on region contrast learning

    CN118710906A

  • Object detection model training method, detection method, apparatus, device and medium

    WO2024120157A1

Cited By

  • Coronary artery vulnerable plaque segmentation model training method and device based on multi-modal semi-supervised learning

    CN120876551A

  • Semi-supervised remote sensing image semantic segmentation system and method

    CN121564356A

  • Semi-supervised remote sensing image semantic segmentation system and method

    CN121564356B

  • Model training method and device for medical image segmentation, server and storage medium

    CN121600007A

  • Medical image segmentation model training methods, devices, servers, and storage media

    CN121600007B