Marine remote sensing coastline segmentation method based on text-guided semi-supervised pseudo-tag

By introducing a text-guided semi-supervised pseudo-label method in the coastline segmentation model, precise pseudo-labels are generated using text encoder and dual-span modal decoder, and the pseudo-label quality is optimized through uncertainty calibration modules, the problems of pseudo-label noise accumulation, insufficient multi-scale feature characterization and limited cross-region generalization capabilities in the prior art are solved, and the coastline segmentation effect with high precision and high adaptability is achieved.

CN120219407AActive Publication Date: 2025-06-27SANYA SCI & EDUCATION INNOVATION PARK WUHAN UNIV OF TECH

Patent Information

Application Number
CN202510694099.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

There are problems in the existing semi-supervised coastline segmentation technology such as pseudo-label noise accumulation, insufficient multi-scale feature characterization and limited cross-region generalization capabilities, resulting in insufficient coastline segmentation accuracy and adaptability.

Method used

Using a semi-supervised pseudo-label method based on text guidance, a segmentation model of a pseudo-label calibration module containing a shared image encoder, a text encoder, a dual-span modal decoder and an uncertainty-driven pseudo-label calibration module is carried out with and unsupervised learning training, optimize connection weights, generate accurate pseudo-labels, and improve segmentation efficiency and accuracy.

Benefits of technology

Effectively reduce pseudo-label noise, enhance the model's ability to identify complex coastline characteristics, improve the coastline segmentation accuracy and cross-regional generalization capabilities of the model, and ensure the high quality of coastline segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219407A_ABST
    Figure CN120219407A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent ocean and remote sensing image processing, and discloses an ocean remote sensing coastline segmentation method based on a text-guided semi-supervised pseudo tag, which comprises the following steps: collecting ocean remote sensing coastline image data, dividing into tag data and non-tag data, and constructing a text prompt; constructing a segmentation model comprising a shared image encoder, a text encoder, two decoders and a pseudo label calibration module, cooperatively extracting image features and text features by all the parts, generating a pseudo label and a prediction mask, and optimizing the pseudo label through uncertainty calibration; training the model in a supervised stage and an unsupervised stage, and adding loss function values of the two stages to a back propagation optimization model; and finally, based on the trained model, precise segmentation of the ocean remote sensing coastline image is realized. According to the method, the efficiency and the accuracy of cross-regional ocean remote sensing coastline segmentation are effectively improved by utilizing text guidance and semi-supervised pseudo labels, and the dependence on a large amount of labeled data is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of intelligent ocean and remote sensing image processing, and particularly relates to a method for segmenting ocean remote sensing coastlines based on text-guided semi-supervised pseudo-labels. Background Art

[0002] Precise segmentation of coastlines in remote sensing images is a core technical requirement in fields such as marine environmental monitoring, coastal zone resource management, and disaster warning. Traditional coastline segmentation methods mainly rely on fully supervised deep learning algorithms. By training segmentation networks with a large amount of manually labeled high-resolution remote sensing data, pixel-level coastline contour extraction is achieved. However, due to the dynamic complexity of coastline morphology affected by multiple factors such as tidal changes and human activities, and the interference factors such as cloud occlusion and uneven illumination in remote sensing images, the cost of manual labeling is high and the problem of inconsistent labeling is likely to occur.

[0003] Current mainstream solutions make full use of a large amount of unlabeled remote sensing data through a pseudo-label generation mechanism. Such methods usually adopt two independent decoders to generate pseudo-labels using labeled data, and gradually expand the training dataset through iterative optimization. However, existing methods have significant limitations in the coastline segmentation scenario. The accumulation of pseudo-label noise is likely to occur in the blurred boundary area between water and land in remote sensing images. Secondly, traditional feature-based methods are difficult to effectively capture the multi-scale problem features of the intertidal zone, resulting in insufficient adaptability of the model to the image. The existing technologies have the following disadvantages: (1) The problem of pseudo-label noise accumulation in existing semi-supervised segmentation technologies When traditional semi-supervised methods iteratively generate pseudo-labels, they lack an effective noise suppression mechanism for the blurred boundary between water and land in remote sensing images, resulting in the continuous accumulation of pseudo-label errors during the training process, significantly reducing the segmentation accuracy of the model for complex landforms; (2) Insufficient multi-scale feature representation in existing semi-supervised segmentation technologies Existing feature alignment methods do not fully consider the multi-scale characteristics of coastlines caused by the dynamic changes of tides, and it is difficult to simultaneously capture macroscopic contours and microscopic detail features, resulting in breaks and distortions in the segmentation results of low-resolution or occluded images; (3) Limited cross-regional generalization ability Current methods do not design an adaptive mechanism for the spectral feature differences of coastlines in different geographical environments, resulting in a significant decrease in performance when the model is migrated to a new region due to the domain shift problem.

[0004] Therefore, how to construct a robust pseudo-label mechanism and at the same time improve the discriminative ability of the model for complex coastline features has become the key challenge to improve the accuracy of remote sensing coastline segmentation. Summary of the Invention

[0005] In view of the existing technical defects in the background art, the present invention proposes a method for segmenting ocean remote sensing coastlines based on text-guided semi-supervised pseudo-labels. This method divides image data, constructs text prompts, builds a segmentation model including modules such as a shared image encoder and a text encoder, trains the model through supervised and unsupervised learning, optimizes the connection weights, and finally realizes the accurate segmentation of ocean remote sensing coastlines based on the trained model, effectively utilizing unlabeled data and improving the segmentation efficiency and accuracy.

[0006] The technical solution adopted by the present invention: A method for segmenting ocean remote sensing coastlines based on text-guided semi-supervised pseudo-labels, and the specific steps include: Step S1), collect ocean remote sensing coastline image data, randomly divide the labeled data and unlabeled data according to a ratio, construct an ocean remote sensing coastline image data set, and construct text prompts according to their respective categories; Step S2), construct an ocean remote sensing coastline segmentation model based on the semi-supervised pseudo-label method. The model structure includes four parts: a shared image encoder, a text encoder, two independent text-image cross-modal decoders, and an uncertainty-driven pseudo-label calibration module; (1) The shared image encoder extracts multi-scale remote sensing image features ; (2) The text encoder extracts text features of specific task prompts ; (3) The two independent text-image cross-modal decoders are a pseudo-label decoder and a prediction decoder respectively, which generate initial pseudo-labels and prediction masks ; (4) The uncertainty-driven pseudo-label calibration module perturbs the multi-scale remote sensing image features input to the pseudo-label decoder to make it have uncertainty, performs Monte Carlo sampling, calculates its uncertainty index and mean value, takes the mean value as the pseudo-label, and obtains the final pseudo-label through uncertainty calibration ; Step S3), train the ocean remote sensing coastline segmentation model, and the training process includes two stages: supervised learning and unsupervised learning; Supervised learning stage: Input the labeled data in the data set into the ocean remote sensing coastline segmentation model, and calculate the value of the supervised loss function between the final pseudo-label and the true label; Unsupervised learning stage: Input the unlabeled data in the data set into the ocean remote sensing coastline segmentation model, and calculate the value of the unsupervised loss function between the final pseudo-label and the prediction mask ; Add the supervised loss function value and the unsupervised loss function value to obtain the total loss function value, perform backpropagation, optimize the connection weights through the selected optimizer and corresponding parameters, and obtain the final ocean remote sensing coastline segmentation model after multiple rounds of training; Step S4) Based on the trained ocean remote sensing coastline segmentation model, input the ocean remote sensing coastline image data to be segmented, and output the related coastline segmentation image.

[0007] Preferably, in (1) of step S2, the shared image encoder extracts multi-scale remote sensing image features The process is as follows: Define the number of Blocks contained in the four layers of the ConvNeXt network as 3, 3, 9, and 3 respectively, with a total of 18 Blocks. Input the labeled data and unlabeled data into the ConvNeXt network. In the first Block, data processing operations are performed on the labeled data and unlabeled data to obtain image feature data. Then, perform a two-dimensional convolution downsampling with a convolution kernel size of 2 and a stride of 2 on the image feature data. The downsampled image feature data is used as the input for the next Block, and enter the next Block to continue the above repeated data processing operations until the last Block performs a skip connection, and then downsample to extract multi-scale remote sensing image features ; The specific process of performing data processing operations in the Block is as follows: Sa1) Extract image features from the data through preprocessing , for the image features Perform a two-dimensional convolution operation with a convolution kernel size of 7×7 and a padding of 3 to obtain convolution image features ; Sa2) Perform layer normalization on the convolution image features ; Sa3) Perform a two-dimensional convolution upsampling operation with a convolution kernel size of 1×1 on the convolution image features after layer normalization to increase its dimension by 4 times; Sa4) Perform GELU activation function processing on the convolution image features after upsampling operation Sa5) Perform a two-dimensional convolution downsampling operation with a convolution kernel size of 1×1 on the convolution image features after GELU activation function processing to reduce its dimension to the original value; Sa6) Perform a skip connection between the convolution image features after downsampling operation and the image features to obtain image feature data.

[0008]

[0008]

[0008] Preferably, in step (2) of step S2, the text encoder extracts the text features of the specific task prompt The specific process includes: The text prompt is converted into an initial text vector with a fixed length by the text encoder and then the initial text vector is input into the CXR-BERT model to extract text features , and the specific process is as follows: Sb1) Use the WordPiece tokenizer to split the text prompt into sub-word units and add special tokens: [CLS] and [SEP] at both ends of the sequence to obtain the Token sequence and the corresponding Token IDs; Sb2) According to the Token IDs, look up the table to obtain the word vectors of the corresponding sub-words to obtain the Token Embedding; Sb3) Learn a position vector for each position in the Token sequence so that the model can perceive the order of the tokens in the sequence to obtain the Position Embedding; Sb4) Add the obtained Token Embedding and Position Embedding to obtain the initial text vector ; Sb5) Processed by the multi-layer Transformer Encoder in the CXR-BERT model, which includes multi-head self-attention, residual connection, layer normalization, feed-forward network, second residual connection and layer normalization; Sb6) The output of the last layer of the Transformer Encoder The text vector at the [CLS] position is regarded as the summary representation of the whole sentence, and the text vectors at other positions are used as the context representations of each Token sequence, that is, the text features .

[0009] Further preferably, the specific process of the Transformer Encoder processing the initial text vector is as follows: Sc1) Divide the sub-space by performing a Linear linear transformation operation on the initial text vector to generate attention heads , and the calculation formula is as follows:

[0010] where, , n is the sequence length, is the model dimension, is the number of attention heads, is the a linear transformation that projects the initial text vector onto a subspace with dimension ; Sc2) Generate queries , keys , and values matrices through a Linear linear transformation, and then calculate the attention weights through an attention mechanism to focus on key information. The calculation formula is as follows: where

[0011] is a scaling factor to prevent the dot product value from being too large, is a normalization operation, is the transpose operation of with respect to , and is the attention weight of the th attention head; Sc3) Combine the multi-head self-attention weights of each subspace, that is, concatenate the attention results of all heads, transform through a projection matrix , and add to obtain the total attention weight . The calculation formula is as follows:

[0012] where is the output projection matrix, is the randomly zeroed partial activation value to prevent overfitting, is the concatenation operation; Sc4) Perform a residual connection and layer normalization on the total attention weight and the initial text vector to obtain the normalized text vector . The calculation formula is as follows:

[0013] where LayerNorm is layer normalization; Sc5) Process through a feed-forward neural network, that is, perform two fully connected transformations on , activate using , and add to obtain the feed-forward network output . The calculation formula is as follows:

[0014] where , are both fully connected weights, and ReLU is a non-linear activation function. is the bias vector of the first fully connected layer, is the bias vector of the second fully connected layer; Sc6) Quadratic residual connection and layer normalization, that is, and perform residual connection and then go through layer normalization to obtain the final output , and the calculation formula is as follows:

[0015] Preferably, in step (3) of step S2, the specific process of the pseudo-label decoder generating the initial pseudo-label includes: The pseudo-label decoder consists of 4 Blocks, and the specific process of each Block is as follows: Sd11) Map the text feature to through a fully connected layer, and its dimension is aligned with the dimension of the image feature. The calculation formula is as follows:

[0016] where Conv is a one-dimensional convolution with a kernel size of 1, is a learnable weight matrix, is the text feature after the fully connected layer; Sd12) Input the multi-scale remote sensing image feature into the self-attention layer, where let , , is , and the calculation formula is as follows:

[0017] where, is layer normalization, is multi-head self-attention, is the image feature after the attention layer; Sd13) Input the image feature after the attention layer and the text feature after the fully connected layer into the cross-modal attention layer, where Q is , and K and V are , and the calculation formula is as follows:

[0018] where, represents the weight parameter value, is the fused image feature after cross-modal fusion; Sd14) Take The features of the multi-scale remote sensing image of the previous input Input into the Unetr module for skip connection fusion, and the calculation formula is as follows:

[0019] Among them, is the feature of the multi-scale remote sensing image of the previous input, is the number of Blocks, is the output feature after skip connection fusion; Sd15) Perform a two-dimensional convolution with a kernel size of 1 on the output feature of the last Block, reduce its dimension to 1, and obtain the prediction result , and the calculation formula is as follows:

[0020] Sd16) Generate multiple pseudo-labels through the pseudo-label decoder , calculate the uncertainty index based on these pseudo-labels, and use the average value of these pseudo-labels as the initial pseudo-label , and the calculation formula is as follows:

[0021] Among them, represents the number of forward passes, is the total number of pseudo-labels, is from start, to end of the summation operation, is the time step.

[0022] Preferably, in (3) of step S2, the specific process of the prediction decoder generating the prediction mask is as follows: The prediction decoder is also composed of 4 Blocks, and the specific process of each Block is as follows: Sd21) Map the text feature to through a fully connected layer, and its dimension is aligned with the image feature dimension. The calculation formula is as follows:

[0023] Among them, Conv is a one-dimensional convolution with a kernel size of 1, is the learnable weight matrix, is the text feature after the fully connected layer; Sd22) Input the multi-scale remote sensing image feature into the self-attention layer, where let , , For , the calculation formula is as follows:

[0024] Wherein, is layer normalization, is multi-head self-attention, is the image feature after the attention layer; Sd23) Input the image feature after the attention layer and the text feature after the fully connected layer into the cross-modal attention layer, where Q is , K and V are , and the calculation formula is as follows:

[0025] Wherein, represents the weight parameter value, is the fused image feature after cross-modal fusion; Sd24) Input and the multi-scale remote sensing image feature of the previous input into the Unetr module for skip connection fusion, and the calculation formula is as follows:

[0026] Wherein, is the multi-scale remote sensing image feature of the previous input, is the number of Blocks, is the output feature after skip connection fusion; Sd25) Perform a two-dimensional convolution with a kernel size of 1 on the output feature of the last Block to reduce its dimension to 1 and obtain the prediction result , and the calculation formula is as follows:

[0027] Wherein, is the two-dimensional convolution operation with a kernel size of 1; Sd26) Perform a forward pass through the prediction decoder , and then obtain the final value through the function, which is the prediction mask , and the calculation formula is as follows:

[0028] Wherein, is the operation of converting to the probability form, that is, transforming the output of to make it conform to the probability distribution and output the probability value.

[0029] Preferably, in step (4) of step S2, the final pseudo-label is obtained through the uncertainty calibration module The process is as follows: Se1) Perform Dropout random perturbation on the multi-scale remote sensing image features input to the pseudo-label decoder to introduce uncertainty. The calculation formula is as follows:

[0030] Se2) Perform forward pass through the pseudo-label decoder to obtain different pseudo-labels , and the calculation formula is as follows:

[0031] wherein, represents the number of forward passes; Se3) Calculate the variance and mean for each pixel point of the different pseudo-labels . The calculation formula is as follows: ,

[0032] wherein, is the variance, representing the average estimate of pixel values, is the mean, measuring the uncertainty of pixels; Se4) Use the variance as the uncertainty index to calibrate the pseudo-label to obtain the final pseudo-label , and the calculation formula is as follows:

[0033] wherein, is element-wise multiplication, is the exponent.

[0034] Preferably, in step S3, the model training process is as follows: The total loss function value of the model is calculated by the sum of the binary cross-entropy and the loss function value. The specific process is as follows: Sf1) In the supervised learning stage, calculate the supervised loss function value by using the final pseudo-label and the true label. The calculation formula is as follows:

[0035] wherein, is the true label, is the binary cross-entropy, is the Dice loss function value; Sf2) In the unsupervised learning stage, through the final pseudo-label and the prediction mask calculate the unsupervised loss function value , and the calculation formula is as follows:

[0036] Sf3) Calculate the total loss function value of the model , and the calculation formula is as follows:

[0037] where and are the weight parameter values.

[0038] Compared with the prior art, the present invention proposes a method for segmenting the coastline of ocean remote sensing based on text-guided semi-supervised pseudo-labels, and the present invention has the following beneficial effects: (1) The segmentation model constructed based on the semi-supervised pseudo-label method adopts a combined architecture of a shared image encoder, a text encoder, two double cross-modal decoders, and an uncertainty-driven pseudo-label calibration module; the deep fusion of multi-scale image features and task-specific text features enables the model to capture the feature information of the coastline from different dimensions, effectively improving the perception and segmentation accuracy of the details of the coastline in complex ocean remote sensing images; the two decoders generate pseudo-labels and prediction masks respectively, with clear division of labor, further optimizing the model's processing ability for the coastline segmentation task; (2) Precise pseudo-label calibration to improve the reliability of segmentation; by perturbing the visual features and Monte Carlo sampling, quantifying and utilizing the uncertainty of the data, and taking the mean as the pseudo-label, the generated pseudo-label is more reliable; this calibration mechanism effectively reduces the noise of the pseudo-label, avoids the negative impact of incorrect pseudo-labels on model training, enhances the stability and accuracy of model prediction, and ensures the high quality of the coastline segmentation result; (3) The training strategy combining supervised learning and unsupervised learning enables the model to accurately learn using labeled data while learning more generalizable feature representations with the help of unlabeled data; this training method helps the model adapt to ocean remote sensing images in different regions and under different imaging conditions, significantly improving the generalization ability of the model, enabling it to perform well in cross-regional coastline segmentation tasks, and effectively coping with the challenges of complex and variable ocean environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is the flowchart of the method proposed by the present invention; Figure 2 is the framework diagram of the pseudo-label decoder and the prediction decoder proposed by the present invention; Figure 3 Block diagram of the uncertainty-driven pseudo-label calibration module proposed by the present invention. Specific implementation manners

[0040] Next, the technical solutions in the embodiments of the present application will be further clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. It should be noted that the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0041] In order to make the invention purpose, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings of the specification: In order to better understand the above objects, features and advantages of the present invention, the advantages of the present invention will be further illustrated by comparing with embodiments in combination with the drawings and specific implementation manners.

[0042] The present invention proposes a method for ocean remote sensing coastline segmentation based on text-guided semi-supervised pseudo-labels. The flowchart of this method is as Figure 1 shown. The steps of this method will be described in detail: Step S1), collect ocean remote sensing coastline image data, randomly divide the labeled data and unlabeled data according to three ratios of 1 / 16:1 / 8:1 / 4, construct an ocean remote sensing coastline image dataset, and construct text prompts according to their respective categories. The text prompt template is “This image depicts a {cls} region.” Replace cls with the region type; Step S2), construct an ocean remote sensing coastline segmentation model based on the semi-supervised pseudo-label method. The model structure includes four parts: a shared image encoder, a text encoder, two independent text-image cross-modal decoders, and an uncertainty-driven pseudo-label calibration module; (1) The shared image encoder extracts multi-scale remote sensing image features ; Specifically, in (1) of step S2, the process of the shared image encoder extracting multi-scale remote sensing image features is as follows: Define the number of Blocks in the four layers of the ConvNeXt network as 3, 3, 9, and 3 respectively. There are a total of 18 Blocks. Input the labeled data and unlabeled data into the ConvNeXt network. In the first Block, the labeled data and unlabeled data are processed to obtain image feature data. Then, the image feature data is downsampled by a two-dimensional convolution with a kernel size of 2 and a stride of 2. The downsampled image feature data is used as the input for the next Block, and the above-mentioned repeated data processing operations are continued until the last Block performs a skip connection, and then multi-scale remote sensing image features are extracted by downsampling. ; The specific process of data processing operations in the Block is as follows: Sa1) Extract image features from the data through preprocessing , and perform a two-dimensional convolution operation on the image features with a kernel size of 7×7 and a padding of 3 to obtain image features ; The local feature information in the image is extracted through the convolution operation. The padding of 3 is to keep the size of the feature map unchanged during the convolution process, so that the edge information of the image can also be effectively processed; Sa2) Perform layer normalization on the convolutional image features ; Layer normalization helps to accelerate the convergence speed of the model and can alleviate the problems of gradient disappearance or gradient explosion; Sa3) Perform a two-dimensional convolution upsampling operation with a kernel size of 1×1 on the convolutional image features after layer normalization to increase its dimension by 4 times; The 1×1 convolution can adjust the number of channels of the feature map without changing the spatial size (width and height) of the image; Increasing the number of channels to 4 times the original can increase the expression ability of the features; Sa4) Perform GELU activation function processing on the convolutional image features after the upsampling operation ; The activation function is used to introduce non-linearity into the neural network; Sa5) Perform a two-dimensional convolution downsampling operation with a kernel size of 1×1 on the convolutional image features after GELU activation function processing to reduce its dimension to the original value; After enriching the information by learning through the previous upsampling operation to expand the number of channels, the number of channels is restored to the original level through the downsampling operation to remove possible redundant information and further integrate the features; Sa6) Perform a skip connection on the convolutional image features after the downsampling operation and the original image features to obtain image feature data.

[0043] (2) The text encoder extracts the text features of the specific task prompt ; Specifically, in (2) of step S2, the text encoder extracts the text features of the specific task prompt The specific process is as follows: The text prompt is converted into an initial text vector with a fixed length by the text encoder Then the initial text vector is input into the CXR - BERT model to extract text features The specific process is as follows: Sb1) Use the WordPiece tokenizer to split the text prompt into sub - word units and add special tokens: [CLS] (classification token) and [SEP] (separator token) at both ends of the sequence to obtain the Token sequence and the corresponding Token IDs (sub - word numbers); The WordPiece tokenizer effectively processes out - of - vocabulary words (such as spelling mistakes or new words) to improve the accuracy of the text; [CLS] is used for aggregating the semantics of the whole sentence (such as classification tasks), and [SEP] is used to distinguish sentence pairs (if any); Sb2) According to the Token IDs, look up the table to get the word vectors of the corresponding sub - words to obtain TokenEmbedding; Sb3) Learn a position vector for each position in the Token sequence so that the model can perceive the order of tokens in the sequence to obtain Position Embedding; Sb4) Add the obtained Token Embedding and Position Embedding to get the initial text vector ; Sb5) Processed by multiple - layer Transformer Encoder, which includes multi - head self - attention (capturing dependencies between words), residual connection (preventing gradient vanishing), layer normalization (stabilizing training), feed - forward network (non - linear transformation), the second residual connection and layer normalization; Sb6) Consider the vector at the [CLS] position of the output of the last layer of the Transformer Encoder as the summary representation of the whole sentence (or sentence pair), and the vectors at other positions can be used as the context representation of each Token, that is, the text features .

[0044] More specifically, the specific process of the Transformer Encoder processing the text vector is as follows: Sc1) Divide the sub - space by performing a Linear linear transformation operation on the initial text vector to generate attention heads , which is used to capture different feature representations and enhance the model's ability to capture text information. The calculation formula is as follows:

[0045] Among them, , n is the sequence length, is the model dimension, h is the number of attention heads, is the th linear transformation, that is, projecting the initial text vector onto a subspace with dimension ; Sc2) Generates query , key , and value matrices by performing a Linear linear transformation on , and then calculates the attention weights through the attention mechanism to focus on key information. The calculation formula is as follows:

[0046] Among them, is the scaling factor to prevent the dot product value from being too large, is the normalization operation, is the transpose operation on , is the attention weight of the th attention head; Sc3) Combines the multi-head self-attention weights of each subspace, that is, concatenates the attention results of all heads, transforms through the projection matrix , and adds to obtain the total attention weight , fuses multi-head information, restores the model dimension, and enhances the generalization ability. The calculation formula is as follows:

[0047] Among them, is the output projection matrix, is the randomly zeroed partial activation value to prevent overfitting, is the concatenation operation; Sc4) Performs a residual connection and layer normalization on the attention weight and the text vector T to obtain the normalized text vector . The residual connection alleviates the vanishing gradient, and the layer normalization improves the training stability. The calculation formula is as follows:

[0048] Among them, LayerNorm is the layer normalization; Sc5) Processes through a feed-forward neural network, that is, on Perform a two-layer fully connected transformation and use for activation, and add to obtain the output of the feed-forward network , and the calculation formula is as follows:

[0049] where , are all fully connected weights, ReLU is a non-linear activation function, is the bias vector of the first-layer fully connected layer, is the bias vector of the second-layer fully connected layer; Sc6) Quadratic residual connection and layer normalization, that is, is connected with by residual connection, and then layer normalization is performed to obtain the final output , and the calculation formula is as follows:

[0050] (3) The two independent text-image cross-modal decoders are a pseudo-label decoder and a prediction decoder, which generate an initial pseudo-label and a prediction mask ; Specifically, in (3) of step S2, as Figure 2 shown, the specific process of the pseudo-label decoder generating the initial pseudo-label includes: The pseudo-label decoder consists of 4 Blocks, and the specific process of each Block is as follows: Sd11) Map the text feature to through a fully connected layer, and its dimension is aligned with the image feature dimension. The calculation formula is as follows:

[0051] where Conv is a one-dimensional convolution with a kernel size of 1, is a learnable weight matrix, is the text feature after the fully connected layer; Sd12) Input the multi-scale remote sensing image feature into the self-attention layer, where is set as . Capture the internal dependencies of the image features through the self-attention mechanism to enhance the expression ability of the image features. The calculation formula is as follows:

[0052] where is layer normalization, is multi-head self-attention, is the image feature after the attention layer; Sd13) The image features after the attention layer And the text features after the fully connected layer Input cross-modal attention layer, where Q is , K and V are , realize the cross-modal interaction between image features and text features, integrate image features with text information, and further enrich feature expression. The calculation formula is as follows:

[0053] in, represents the weight parameter value, is the fused image feature after cross-modal fusion; Sd14) Multi-scale remote sensing image features with the previous input Input the Unetr module for skip connection fusion, the calculation formula is as follows:

[0054] in, is the multi-scale remote sensing image feature of the previous input, is the number of Blocks, Output features after skip connection fusion; Sd15) Output features of the last layer of Block Perform a two-dimensional convolution with a convolution kernel size of 1 to reduce its dimension to 1 and obtain the prediction result , the calculation formula is as follows:

[0055] Sd16) through the pseudo-label decoder Generate multiple pseudo labels, calculate uncertainty indicators based on these pseudo labels, and take the average of these pseudo labels as the initial pseudo label , the calculation formula is as follows:

[0056] in, represents the number of forward passes, is the total number of pseudo labels, For Start to The summation operation ends, is the time step.

[0057] Specifically, in step S2 (3), the prediction decoder generates a prediction mask The specific process is as follows: The prediction decoder is also composed of 4 blocks, and the specific process of each block is as follows: Sd21) Map the text features through a fully connected layer to , whose dimension is aligned with the dimension of the image features. The calculation formula is as follows:

[0058] where Conv is a one-dimensional convolution with a kernel size of 1, is a learnable weight matrix, is the text feature after the fully connected layer; Sd22) Input the multi-scale remote sensing image features into the self-attention layer, where let be . Capture the internal dependencies of the image features through the self-attention mechanism to enhance the expression ability of the image features. The calculation formula is as follows:

[0059] where, is layer normalization, is multi-head self-attention, is the image feature after the attention layer; Sd23) Input the image feature after the attention layer and the text feature after the fully connected layer into the cross-modal attention layer, where Q is , and K and V are . Realize the cross-modal interaction between the image features and the text features, so that the image features fuse the text information and further enrich the feature expression. The calculation formula is as follows:

[0060] where, represents the weight parameter value, is the fused image feature after cross-modal fusion; Sd24) Input and the previous input multi-scale remote sensing image feature into the Unetr module for skip connection fusion. The calculation formula is as follows:

[0061] where, is the previous input multi-scale remote sensing image feature, is the number of blocks, is the output feature after skip connection fusion; Sd25) Perform a two-dimensional convolution with a kernel size of 1 on the output features of the last layer of the block, reduce its dimension to 1, and obtain the prediction result , and the calculation formula is as follows: ,

[0062] Sd26) Perform a forward pass through the prediction decoder , and then through the function to obtain the final value, which is the prediction mask , and the calculation formula is as follows:

[0063] where, is the operation of converting to probability form, that is, transforming the output of to make it conform to the probability distribution and output probability values.

[0064] (4) The uncertainty-driven pseudo-label calibration module perturbs the multi-scale remote sensing image features input to the pseudo-label decoder to make it have uncertainty, performs Monte Carlo sampling, calculates its uncertainty index and mean value, takes the mean value as the pseudo-label, and obtains the final pseudo-label through uncertainty calibration ; Specifically, in (4) of step S2, as Figure 3 shown, the process of obtaining the final pseudo-label through uncertainty calibration is as follows: Se1) Perform Dropout random perturbation on the multi-scale remote sensing image features input to the pseudo-label decoder to introduce uncertainty, and the calculation formula is as follows:

[0065] Se2) Perform a forward pass through the pseudo-label decoder to obtain different pseudo-labels , and the calculation formula is as follows:

[0066] where, represents the number of forward passes; Se3) Calculate the variance and mean value for each pixel point of the different pseudo-labels , and the calculation formula is as follows: ,

[0067] Among them, is the variance, representing the average estimate of pixel values, is the mean, measuring the uncertainty of pixels (the larger the variance, the higher the uncertainty); Se4) Use the variance as an uncertainty indicator to calibrate the pseudo - labels to obtain the final pseudo - labels , and the calculation formula is as follows:

[0068] Among them, is element - wise multiplication, is the exponent.

[0069] Step S3), train the ocean remote - sensing coastline segmentation model, and the training process includes two stages: supervised learning and unsupervised learning; Supervised learning stage: Input the labeled data in the dataset into the ocean remote - sensing coastline segmentation model, and calculate the value of the supervised loss function between the final pseudo - labels and the true labels; Unsupervised learning stage: Input the unlabeled data in the dataset into the ocean remote - sensing coastline segmentation model, and calculate the value of the unsupervised loss function between the final pseudo - labels and the predicted mask ; Add the value of the supervised loss function and the value of the unsupervised loss function to obtain the total loss function value, perform backpropagation, and optimize the connection weights through the selected optimizer and corresponding parameters. After multiple rounds of training, obtain the final ocean remote - sensing coastline segmentation model; Specifically, in step S3, the process of training the model is as follows: The process of model training is as follows: Use the sum of the binary cross - entropy and the loss function value to calculate the total loss function value of the model. The specific process is as follows: Sf1) In the supervised learning stage, calculate the value of the supervised loss function through the final pseudo - labels and the true labels. The calculation formula is as follows:

[0070] Among them, is the true label, is the binary cross - entropy, measuring the difference in probability distribution between the pseudo - labels and the true labels, focusing on the correctness of classification, is the Dice loss function value, measuring the overlap degree between the pseudo - labels and the true labels, being more robust to the problem of class imbalance and focusing on the accuracy of segmentation; Sf2) In the unsupervised learning stage, through the final pseudo - labels and the predicted mask Calculate the value of the unsupervised loss function , and the calculation formula is as follows:

[0071] Sf3) Calculate the total loss function value of the model , and the calculation formula is as follows:

[0072] Among them, and are the weight parameter values, which are used to adjust the contribution ratios of the supervised loss function value and the unsupervised loss function value to the total loss.

[0073] Step S4) Based on the trained ocean remote sensing coastline segmentation model, input the ocean remote sensing coastline image data to be segmented, and output the related coastline segmentation image.

[0074] The present invention proposes a core innovation mechanism, the text-guided decoding architecture, in the field of cross-domain coastline semantic segmentation under the semi-supervised pseudo-labeling method. Compared with the limitation of traditional semi-supervised coastline segmentation methods relying on a single visual representation, this architecture innovatively constructs a region-adaptive text-visual collaborative learning paradigm: at the stage of reconstructing features at all levels of the decoder, semantic cue vectors strongly related to geographical partitions are injected. Through hierarchical text semantic embedding and cross-modal fusion of multi-scale visual features, the decoupled enhancement of fine-grained region features is realized; this design enables the model to dynamically establish semantic alignment between the text description space and the pixel-level visual feature space during cross-region parallel training, effectively improving the joint representation learning ability of multi-region heterogeneous coastline elements, thereby breaking through the bottleneck of regional feature confusion existing in traditional methods in cross-domain generalization, and significantly enhancing the topological structure adaptability and edge segmentation accuracy of the model to unlabeled regions; Meanwhile, in order to improve the quality of the pseudo-labels generated by the pseudo-label decoder, an uncertainty-driven pseudo-label calibration module is further introduced to optimize the output reliability of the pseudo-label decoder in the semi-supervised framework; by performing Monte Carlo sampling on the decoder, a set of pseudo-labels with different probability distributions is generated, an uncertainty heat map is constructed based on the pixel-level prediction variance, and by minimizing the inter-class information entropy of high-uncertainty regions, the confidence threshold of the pseudo-labels is dynamically adjusted to suppress noise propagation; this joint optimization strategy not only realizes the probabilistic evaluation of the pseudo-label quality, but also alleviates the random deviation of single inference through the ensemble learning paradigm, significantly improving the topological consistency and boundary robustness of the pseudo-labels in cross-domain scenarios.

[0075] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.

[0076] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. An ocean remote sensing coastline segmentation method based on text-guided semi-supervised pseudo-labels, characterized in that Including: Step S1), collect ocean remote sensing coastline image data, randomly divide the labeled data and unlabeled data according to a ratio, construct an ocean remote sensing coastline image dataset, and construct text prompts according to their categories; Step S2), construct an ocean remote sensing coastline segmentation model based on the semi-supervised pseudo-label method, and the model structure includes four parts: a shared image encoder, a text encoder, two independent text-image cross-modal decoders, and an uncertainty-driven pseudo-label calibration module; (1)The shared image encoder extracts multi-scale remote sensing image features ; (2)The text encoder extracts the text features of the specific task prompt ; (3)The two independent text-image cross-modal decoders are a pseudo-label decoder and a prediction decoder respectively, which generate initial pseudo-labels and prediction masks ; (4) The uncertainty-driven pseudo-label calibration module perturbs the multi-scale remote sensing image features input to the pseudo-label decoder, making them have uncertainty, and performs Monte Carlo sampling, calculates its uncertainty index and mean value, takes the mean value as the pseudo-label, and obtains the final pseudo-label through uncertainty calibration ; performs Monte Carlo sampling, calculates its uncertainty index and mean value, takes the mean value as the pseudo-label, and obtains the final pseudo-label through uncertainty calibration ; Step S3), train the ocean remote sensing coastline segmentation model, and the training process includes two stages: supervised learning and unsupervised learning; Supervised learning stage: Input the labeled data in the dataset into the ocean remote sensing coastline segmentation model to calculate the final pseudo-label and the value of the supervised loss function between the true label; Unsupervised learning stage: Input the unlabeled data in the dataset into the ocean remote sensing coastline segmentation model to calculate the final pseudo-labels and the predicted masks to calculate the value of the unsupervised loss function between them; Add the supervised loss function value and the unsupervised loss function value to obtain the total loss function value, perform backpropagation, optimize the connection weights through the selected optimizer and corresponding parameters, and obtain the final ocean remote sensing coastline segmentation model after multiple rounds of training; Step S4), based on the trained ocean remote sensing coastline segmentation model, input the ocean remote sensing coastline image data to be segmented, and output the corresponding coastline segmentation image.

2. The method for segmenting ocean remote sensing shorelines based on text-guided semi-supervised pseudo-labels according to claim 1, wherein In step (1) of step S2, the shared image encoder extracts multi-scale remote sensing image features The process is as follows: The number of Blocks contained in the four layers of the ConvNeXt network is defined as 3, 3, 9, and 3 respectively, for a total of 18 Blocks. The labeled data and unlabeled data are input into the ConvNeXt network. In the first Block, the labeled data and unlabeled data are processed to obtain image feature data. Then, two-dimensional convolutional downsampling with a convolutional kernel size of 2 and a stride of 2 is performed on the image feature data. The downsampled image feature data is used as the input for the next Block, and the above-mentioned repeated data processing operations are continued in the next Block until the last Block performs skip connections, and multi-scale remote sensing image features are extracted by downsampling. ; The specific process of data processing operations in the said Block is as follows: Sa1) Preprocess the data to extract image features , for the image features perform a two-dimensional convolution operation with a convolution kernel size of 7×7 and a padding of 3 to obtain convolution image features ; Sa2) Perform layer normalization on the convolutional image features ; Sa3) The convolutional image features after layer normalization operation Perform a two-dimensional convolutional upsampling operation with a kernel size of 1×1 to increase its dimension by 4 times; Sa4) Perform GELU activation function processing on the convolutional image features after the dimension elevation operation ; Sa5) The convolutional image features after being processed by the GELU activation function Perform a two-dimensional convolutional dimensionality reduction operation with a convolutional kernel size of 1×1 to reduce its dimension to the original value; Sa6) The convolutional image features after the dimensionality reduction operation are skip-connected with the image features to obtain the image feature data.

3. The method for segmenting ocean remote sensing shorelines based on text-guided semi-supervised pseudo-labels according to claim 1, wherein In step (2) of step S2, the text encoder extracts the text features of the specific task prompt The specific process includes: Convert the text prompt into an initial text vector with a fixed length of through the text encoder, and then input the initial text vector into the CXR-BERT model to extract text features , and the specific process is as follows: Sb1) Use the WordPiece tokenizer to split the text prompt into sub-word units and add special tokens: [CLS] and [SEP] at both ends of the sequence to obtain the Token sequence and the corresponding Token IDs; Sb2) According to the Token IDs, look up the table to obtain the word vectors of the corresponding sub-words to obtain the Token Embedding; Sb3) Learn a position vector for each position in the Token sequence so that the model can perceive the order of tokens in the sequence to obtain the Position Embedding; Sb4) Add the obtained Token Embedding and Position Embedding to get the initial text vector ; Sb5) Processed by the multi-layer Transformer Encoder in the CXR-BERT model, which includes multi-head self-attention, residual connection, layer normalization, feed-forward network, the second residual connection and layer normalization; Sb6) Take the output of the last layer of the Transformer Encoder The text vector at the [CLS] position is regarded as the summary representation of the whole sentence, and the text vectors at other positions are used as the context representations of each Token sequence, that is, text features .

4. The method for segmenting marine remote sensing coastline based on text-guided semi-supervised pseudo-labels according to claim 3, characterized in that, The Transformer Encoder processes the initial text vector The specific process is as follows: Sc1) By performing a Linear linear transformation operation on the initial text vector to divide the subspaces, generating attention heads , the calculation formula is as follows: ; Among them, , where n is the sequence length, is the model dimension, is the number of attention heads, is the th linear transformation, that is, projecting the initial text vector onto a subspace of dimension . Sc2) Pairwise Generate queries through a Linear linear transformation , keys , values matrices, and then calculate the attention weights through the attention mechanism to focus on key information. The calculation formula is as follows: ; Among them, is a scaling factor to prevent the dot product value from being too large, is a normalization operation, is for the transpose operation of, is the attention weight of the (Sc3) Merge the multi-head self-attention weights of each subspace, that is, concatenate the attention results of all heads, and transform through the projection matrix and add to obtain the total attention weight The calculation formula is as follows: ; Among them, is the output projection matrix, is the randomly zeroed part of the activation values to prevent overfitting, is the concatenation operation; Sc4) Residual connect and layer normalize the total attention weight and the initial text vector to obtain the normalized text vector , and the calculation formula is as follows: ; Among them, LayerNorm is layer normalization; Sc5) Feedforward neural network processing, that is, for perform two-layer fully connected transformation, use for activation, and add to obtain the feedforward network output The calculation formula is as follows: ; Among them, , are all fully connected weights, is a non-linear activation function, is the bias vector of the first layer of fully connected layer, is the bias vector of the second layer of fully connected layer; Sc6) Secondary residual connection and layer normalization, that is, and perform residual connection and then go through layer normalization to obtain the final output , and the calculation formula is as follows: 。 5. The method for segmenting the coastline of marine remote sensing based on text-guided semi-supervised pseudo-labels according to claim 1, wherein In step (3) of step S2, the pseudo-label decoder generates an initial pseudo-label The specific process includes: The pseudo-label decoder consists of 4 Blocks, and the specific process of each Block is as follows: Sd11) Map the text features through a fully connected layer to , whose dimension is aligned with the dimension of the image features. The calculation formula is as follows: ; Among them, Conv is a one-dimensional convolution with a convolution kernel size of 1, is a learnable weight matrix, is the text feature after the fully connected layer; Sd12) Input the multi-scale remote sensing image features into the self-attention layer, where let , , be , and the calculation formula is as follows: ; Among them, is layer normalization, is multi-head self-attention, is the image feature after the attention layer; The image features after the attention layer and the text features after the fully connected layer are input into the cross-modal attention layer, where Q is , K and V are , and the calculation formula is as follows: ; Among them, represents the weight parameter value, is the fused image feature after cross-modal fusion; Sd14) Multi-scale remote sensing image features with the previous input Input the Unetr module for skip connection fusion, the calculation formula is as follows: ; Among them, is the multi-scale remote sensing image feature of the previous input, is the number of Blocks, is the output feature after skip connection fusion; Sd15) Perform a two-dimensional convolution with a kernel size of 1 on the output features of the last block, reduce its dimension to 1, and obtain the prediction result The calculation formula is as follows: , and the calculation formula is as follows: ; Sd16) Through the pseudo-label decoder Generate multiple pseudo-labels, calculate the uncertainty index based on these pseudo-labels, and use the average value of these pseudo-labels as the initial pseudo-label , and the calculation formula is as follows: ; Among them, represents the number of forward passes, is the total number of pseudo-labels, is from to the summation operation ending at, is the time step.

6. The method for segmenting ocean remote sensing shorelines based on text-guided semi-supervised pseudo-labels according to claim 1, wherein In step (3) of step S2, the prediction decoder generates a prediction mask The specific process is as follows: The prediction decoder also consists of 4 Blocks, and the specific process of each Block is as follows: Sd21) Map the text features through a fully connected layer to , whose dimension is aligned with the dimension of the image features. The calculation formula is as follows: ; Among them, Conv is a one-dimensional convolution with a convolution kernel size of 1, is a learnable weight matrix, is the text feature after the fully connected layer; Sd22) Input the multi-scale remote sensing image features into the self-attention layer, where let , , be , and the calculation formula is as follows: ; Among them, is layer normalization, is multi-head self-attention, is the image feature after the attention layer; The image features after the attention layer and the text features after the fully connected layer are input into the cross-modal attention layer, where Q is , and K and V are , and the calculation formula is as follows: ; Among them, represents the weight parameter value, is the fused image feature after cross-modal fusion; Sd24) Combine with the multi-scale remote sensing image features of the previous input and input them into the Unetr module for skip connection fusion. The calculation formula is as follows: ; Among them, is the multi-scale remote sensing image feature of the previous input, is the number of Blocks, is the output feature after skip connection fusion; Sd25) Perform a two-dimensional convolution with a kernel size of 1 on the output features of the last block, reducing its dimension to 1 to obtain the prediction result The calculation formula is as follows: ​ ; Among them, is a two-dimensional convolution operation with a convolution kernel size of 1; Sd26) Through the prediction decoder perform a forward pass, and then through The final value obtained by the function is the prediction mask , and the calculation formula is as follows: ; Among them, it is an operation to convert to a probability form, that is, the output of is transformed to conform to a probability distribution and a probability value is output.

7. The method for segmenting marine remote sensing coastline based on text-guided semi-supervised pseudo-labeling according to claim 1, characterized in that In step (4) of step S2, the final pseudo-label is obtained through the uncertainty calibration module The process is as follows: Se1) Perform Dropout random perturbation on the multi-scale remote sensing image features input to the pseudo-label decoder to introduce uncertainty. The calculation formula is as follows: ​ ; Se2) Perform a forward pass through the pseudo-label decoder to obtain different pseudo-labels , and the calculation formula is as follows: ; Among them, represents the number of forward passes; Se3) For different pseudo-labels calculate the variance and the mean pixel by pixel, and the calculation formula is as follows: , ; Among them, is the variance, representing the average estimate of pixel values, is the mean, measuring the uncertainty of pixels; Se4) Calibrate the pseudo-labels using variance as the uncertainty metric to obtain the final pseudo-labels , and the calculation formula is as follows: ; wherein, is element-wise multiplication, is the exponent.

8. The method for segmenting ocean remote sensing shorelines based on text-guided semi-supervised pseudo-labels according to claim 1, wherein In the said step S3, the model training process is as follows: Using binary cross-entropy and The total loss function value of the model is calculated by the sum of the loss function values. The specific process is as follows: Sf1) In the supervised learning stage, calculate the supervised loss function value through the final pseudo-label and the true label , and the calculation formula is as follows: ; wherein, is the true label, is the binary cross-entropy, is the value of the Dice loss function; Sf2) In the unsupervised learning stage, through the final pseudo-labels and the prediction mask calculate the value of the unsupervised loss function , and the calculation formula is as follows: ; Sf3) Calculate the total loss function value of the computational model , and the calculation formula is as follows: ; Among them, and are weight parameter values.

Citation Information

Patent Citations

  • Semi-supervised remote sensing image semantic segmentation method based on double consistency

    CN116416618A

  • Surface defect detection method based on semi-supervised training strategy

    CN118052770A

Cited By

  • Methane column concentration inversion method and system based on physically-driven variational auto-encoder

    CN121027002A

  • Remote sensing image building change detection system based on self-training and consistency learning

    CN121147220A

  • Marine ecological monitoring method based on ecological representation and target perception matching

    CN121834243A

  • Deep sea geology mapping method and system based on multi-modal large model

    CN122223151A

  • Text and space context guided remote sensing image cross-domain segmentation method and system

    CN122336311A