Graffiti supervision medical image segmentation method with edge detection module based on reliable pseudo label

The graffiti-supervised medical image segmentation method, which introduces reliable pseudo-labels and edge detection modules, solves the problems of pseudo-label reliability and lack of boundary information, thereby improving the accuracy and efficiency of medical image segmentation.

CN121033076APending Publication Date: 2025-11-28CHANGCHUN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511149332.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

In existing graffiti-supervised medical image segmentation methods, the reliability of pseudo-labels is difficult to guarantee, and there is a lack of complete target shape and boundary information, resulting in insufficient or unreliable supervision signals, which affects the model's learning performance.

Method used

An edge detection module based on reliable pseudo-labels is adopted, which combines an encoder, dual decoders, channel-space attention blocks and weighted loss functions. Reliable pseudo-labels are selected through reliable pseudo-label learning and edge supervision, thereby enhancing the model's ability to identify edge regions.

Benefits of technology

It significantly reduces noise interference from low-quality pseudo-labels and improves the model's discrimination accuracy and segmentation performance in boundary regions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033076A_ABST
    Figure CN121033076A_ABST
Patent Text Reader

Abstract

The invention provides a graffiti supervision medical image segmentation method with edge detection based on reliable pseudo labels, and the method comprises the steps: 1, selecting different public data sets with graffiti labels, dividing the public data sets into a training set, a test set and a verification set, and carrying out the preprocessing; secondly, a medical image segmentation model is created based on an encoder, double decoders, channel-space attention blocks, an edge detection module and a weighted loss function, and the channel-space attention modules are introduced into the encoder and the decoders respectively; the edge detection module is composed of three edge detection heads. And thirdly, training, verifying and testing the model through the graffiti labels in the training set. And 4, deploying the model with a good test effect in a server for a medical image automatic segmentation task. The method has the advantages that the high dependence of a medical image segmentation task on full-pixel labeling is reduced, and the segmentation precision of a medical image segmentation model on different medical images by utilizing graffiti labeling is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of medical image segmentation, and specifically, a medical image segmentation method based on reliable pseudo-label with an edge detection module is designed. The method is trained and tested on public medical image datasets, achieving better segmentation results. BACKGROUND

[0002] Medical image segmentation is a fundamental and critical task in medical image analysis, which aims to accurately separate pixels with specific anatomical structures or pathological regions from the background. Through the segmentation process, accurate positioning and quantification of organs, tissues or lesion regions can be achieved, providing reliable support for disease diagnosis, treatment planning and postoperative evaluation. Considering the significant time cost and manpower burden in the operation process of manual segmentation, and the accuracy largely depends on expert experience, automated segmentation technology is particularly critical. Deep neural networks (DNNs) have made great progress in medical image segmentation due to their excellent ability to learn complex patterns and features directly from data. However, such methods usually rely on a large number of pixel-level accurate annotations to achieve supervised learning, and the fine annotation of medical images not only consumes time and human resources, but also requires the support of professional medical knowledge, which is difficult to obtain on a large scale in clinical practice.

[0003] In order to reduce the time and cost of annotation, weakly supervised learning methods stand out. Weakly supervised methods only need to use lower-cost supervision signals such as image-level labels, bounding boxes, points, scribble labels, etc., which can significantly reduce the dependence on fine annotation while ensuring segmentation performance, thereby improving the scalability and practicality of the model. Among them, scribble annotation as a structured, sparse and more close-to-target region contour weak supervision form, not only has much lower annotation cost than full-pixel annotation, but also provides certain spatial and shape prior information. It can achieve a good balance between supervision intensity and annotation efficiency, so it has received more and more attention in medical image segmentation tasks in recent years.

[0004] Compared with full pixel annotation, scribble annotation only provides sparse supervision signal, therefore, how to effectively mine the potential structure information in unlabeled pixels is the key to improve the segmentation performance. In the existing scribble supervised segmentation methods, pseudo label learning is considered as a promising strategy, which effectively utilizes the unlabeled regions by assigning pseudo labels, and increases the number of labeled pixels available for training without additional human cost. But it also faces two problems. First, the reliability of pseudo labels is difficult to guarantee, and many existing methods rely on setting a confidence threshold to filter out low-confidence predictions, and only use high-confidence pseudo labels to train the model. But there are drawbacks: high threshold can ensure label quality, but it is easy to overlook a large number of potential effective areas, resulting in insufficient supervision signal; while low threshold may introduce unreliable labels, affecting model learning. In addition, compared with full annotation, scribble annotation often lacks complete target shape and boundary information, especially in medical image segmentation tasks, accurate boundary segmentation is crucial. SUMMARY

[0005] The purpose of the present application is to solve the high dependence on full pixel annotation in medical image segmentation tasks, and to propose a scribble supervised medical image segmentation based on reliable pseudo labels with edge detection module for accurate segmentation of different medical images using scribble annotation.

[0006] The technical scheme adopted by the present application is as follows:

[0007] The scribble supervised medical image segmentation method based on reliable pseudo labels with edge detection module has the following specific implementation steps:

[0008] Step 1, select different public datasets with scribble annotation, divide the dataset into training set, test set and validation set, and perform preprocessing operation on the dataset images;

[0009] Step 2, create a medical image segmentation model based on an encoder, a double decoder, a channel-space attention block, an edge detection module and a weighted loss function, wherein the encoder is composed of multiple cascaded convolution-downsampling modules, a spatial attention module (SAB) is introduced in the shallow layer, and a channel attention module (CAB) is introduced in the deep layer. For the double decoder, each decoder is composed of multiple cascaded convolution-upsampling units, each unit has two convolution blocks and an upsampling layer, and a lightweight channel-space attention module (CBAM) is introduced after each upsampling module. The edge detection module is mainly composed of three independent edge detection heads;

[0010] Step 3: Train the medical image segmentation model using the graffiti labels in the training set, verify the segmentation effect of the medical image segmentation model using the validation set, and test the validated medical image segmentation model using the test set.

[0011] Step 4: Deploy the tested and effective medical image segmentation model on the server and use this model to perform medical image segmentation tasks.

[0012] Furthermore, step 1 specifically includes:

[0013] The selected medical image segmentation datasets are ACDC and MSCMRseg heart datasets. During preprocessing, the images in the ACDC dataset are uniformly adjusted to 256×256 pixels, and the images in the MSCMRseg dataset are uniformly adjusted to 480×480 pixels.

[0014] Furthermore, in step 2, the encoder is used to downsample the dataset images layer by layer, gradually extracting medical image features. A shallow spatial attention module is used to enhance the model's focus on salient regions and edge structures, improving the ability to distinguish between foreground and background. A deep channel attention module is used to strengthen the model's response to channels with deep semantic information, making it more focused on semantic channels related to the segmentation target.

[0015] Furthermore, in step 2, a perturbation decoder is embedded within a general UNet network framework to construct a dual-decoder structure. This enhances branch diversity while effectively mitigating model instability caused by graffiti supervision. Additionally, each decoder embeds a channel-space attention module, considering both channel and spatial dimension weight adjustments, enabling the fused feature map to effectively suppress background noise and enhance foreground regions. The two decoders have different feature attention regions and prediction boundaries, making them complementary when handling blurred or uncertain regions.

[0016] Furthermore, in step 2, a partial cross-entropy function is used to supervise the learning of the bi-branch prediction results and the graffiti labels, ignoring unlabeled pixels in the graffiti annotations. The mean of the bi-branch prediction loss is taken as the basic graffiti supervision loss.

[0017] Further, in step 2, corresponding pseudo-labels are generated using the prediction results of the dual decoders. Simultaneously, the prediction confidence scores of the two decoders are calculated, and a corresponding reliability mask is generated based on the confidence scores to determine reliable pseudo-labels. To evaluate the reliability, the feature representations of the penultimate layer are extracted from both decoders, and pixel features of each category are selected by combining them with the pseudo-labels. Then, the mean value of each category of features is calculated to construct the corresponding class prototype. Finally, the pixel-level cosine similarity between the penultimate layer features and the class prototype is calculated, and the mean squared error between the dual-branch output and the cosine similarity is used to measure the reliability of the pseudo-labels. The loss L for reliable pseudo-label supervision is calculated. Rpl .

[0018] Furthermore, in step 2, the deep semantic features of the penultimate layer in the main decoder are fused with the low-level features of the shallow stage in the encoder. Then, the fused features are concatenated with the deep features in the auxiliary decoder and fed into the edge detection head. A channel-spatial attention module is added to enhance edge perception, generating an edge prediction map that fuses the shallow features. Simultaneously, pseudo-edge maps are generated using pseudo-labels, and these are compared with the fused edge prediction map for edge loss supervision, enhancing the model's ability to identify target edge regions. Then, the deep features of the two branches are used to generate corresponding edge maps through two independent edge detection heads, and the consistency loss between the two is calculated.

[0019] Furthermore, in step 2, the weighted loss function is obtained by multiplying the graffiti supervision loss, the reliable pseudo-label supervision loss, the edge supervision loss, and the edge consistency loss by their respective weights and then adding them together.

[0020] Furthermore, step 3 specifically includes:

[0021] The medical image segmentation model is trained using the graffiti labels in the training set. During the training process, the loss function, optimizer function, and learnable hyperparameters used by the medical image segmentation model are continuously optimized until the model's segmentation performance reaches its best.

[0022] The trained medical image segmentation model is validated using the validation set. If the validation results are unsatisfactory, training continues. If the validation results are good, the model is tested using the test set.

[0023] When validating the model's actual segmentation performance using a test set, training ends if the segmentation performance is good, and continues if the segmentation performance is poor.

[0024] Furthermore, step 4 specifically involves:

[0025] The tested medical image segmentation model is deployed to the server, and the calling interface and access permissions are set. Users input the preprocessed data to be segmented into the medical image segmentation model, which outputs segmentation results with detailed data annotations, completing the segmentation task.

[0026] The method proposed in this invention has the following main advantages:

[0027] Using a publicly available dataset with graffiti annotations, the dataset is divided into training, testing, and validation sets, and the images in the dataset undergo preprocessing. Then, a medical image segmentation model is created based on an encoder, dual decoders, channel-spatial attention blocks, an edge detection module, and a weighted loss function. The validated medical image segmentation model is tested using the training set. The tested and effective medical image segmentation model is deployed on a server and used for medical image segmentation tasks. A channel-spatial attention mechanism is integrated into the encoder and decoder to enhance the model's ability to model local and global information. A reliable pseudo-label learning mechanism is introduced to select more trustworthy pseudo-labels for supervision, significantly reducing noise interference from low-quality pseudo-labels. Furthermore, an edge supervision module is introduced, generating pseudo-edge labels by fusing dual-branch predictions and aligning them with edge information fused with shallow features. Consistency constraints are applied between the dual-branch edge predictions to improve the model's accuracy in boundary region discrimination. Attached Figure Description

[0028] Figure 1 This is a flowchart of the graffiti-supervised medical image segmentation method based on reliable pseudo-labels and an edge detection module, according to the present invention.

[0029] Figure 2 This is a general framework diagram of a graffiti-supervised medical image segmentation model based on reliable pseudo-labels and an edge detection module.

[0030] Figure 3 This is a structural diagram of the spatial attention module in the encoder.

[0031] Figure 4 This is a structural diagram of the channel attention module in the encoder.

[0032] Figure 5 This is a structural diagram of the channel-space attention module in the decoder. Detailed Implementation

[0033] The present invention will be further described below with reference to the accompanying drawings to enable those skilled in the art to better understand the invention. It should be noted that those skilled in the art can make some modifications to the present invention without departing from its core spirit, and these modifications all fall within the scope of protection of the present invention.

[0034] Please refer to Figures 1 to 5 As shown, a specific implementation of a graffiti-supervised medical image segmentation method based on reliable pseudo-labels and an edge detection module according to the present invention includes the following steps:

[0035] Step 1: Select different public datasets with graffiti annotations, divide the datasets into training, testing and validation sets, and preprocess the dataset images.

[0036] Step 2 involves creating a medical image segmentation model based on an encoder, dual decoders, channel-spatial attention blocks, an edge detection module, and a weighted loss function. The encoder consists of multiple cascaded convolutional-downsampling modules, with spatial attention modules introduced at shallow layers and channel attention modules at deeper layers. For the dual decoders, each decoder comprises multiple cascaded convolutional-upsampling units, each unit having two convolutional blocks and one upsampling layer. A channel-spatial attention module is introduced after each upsampling module. The edge detection module mainly consists of three independent edge detection heads.

[0037] Step 3: Train the medical image segmentation model using the graffiti labels in the training set, verify the segmentation effect of the medical image segmentation model using the validation set, and test the validated medical image segmentation model using the test set.

[0038] Step 4: Deploy the tested and effective medical image segmentation model on the server and use this model to perform medical image segmentation tasks.

[0039] Step 1 specifically involves:

[0040] The selected medical image segmentation datasets were the ACDC and MSCMRseg heart datasets. During preprocessing, images in the ACDC dataset were uniformly resized to 256×256 pixels, and images in the MSCMRseg dataset were uniformly resized to 480×480 pixels. In practice, due to the different dataset sizes, the two datasets were divided into training, validation, and test sets in ratios of 70:15:15 and 25:5:15, respectively.

[0041] In step 2, the encoder consists of multiple cascaded convolutional-downsampling modules, with a spatial attention module introduced in its shallow layer and a channel attention module introduced in its deep layer. The model framework is as follows: Figure 2 As shown.

[0042] The encoder is used to downsample the dataset images layer by layer, progressively extracting medical image features. For shallow features, a spatial attention module is introduced to enhance the model's focus on salient regions and edge structures, thereby improving the ability to distinguish between foreground and background. For deep features, a channel attention module is introduced to strengthen the model's channel response to deep semantic information, making it more focused on semantic channels related to the segmentation target. During downsampling, the resolution of the feature map is progressively reduced to 1 / 2, 1 / 4, 1 / 8, and 1 / 16 of the initial input image, while the number of feature channels increases proportionally layer by layer.

[0043] During the encoding stage, shallow features typically possess higher spatial resolution and richer edge details. To better extract information from salient regions, spatial attention modules are introduced, such as... Figure 3 As shown. First, for a given feature map... Calculate its average and max-pooling channel attention graphs separately, as shown in the following formulas:

[0044]

[0045] The convolution calculation, after concatenation along the channel dimension, is shown in the diagram below, and the formula is expressed as follows:

[0046]

[0047] Finally, its weighted feature map is obtained, and the formula is expressed as follows:

[0048] F s =F⊙M spatial ,

[0049] Where σ is the Sigmoid activation function, ⊙ represents element-wise multiplication, and Conv k×k It is a convolution operation with a kernel size of k (7×7). An attention map is generated based on the importance of pixel positions in the feature map, guiding the network to focus on the foreground region and boundary contours, thereby improving the sensitivity of the early features and the edge segmentation ability.

[0050] For deep features, which are rich in semantic information but have low spatial resolution, a channel attention module is introduced, such as... Figure 3 As shown, this can improve the network's selective focus on channels strongly related to the semantic target, suppress redundant features, and allow the model to focus on dimensions with stronger class discriminative power. For a given feature map... Global average pooling (GAP) and global max pooling (GMP) are performed along the channel dimension, as expressed by the following formulas:

[0051]

[0052] These features are then fed into a shared multilayer perceptron (MLP) to achieve feature compression and recovery in the channel attention mechanism, as expressed in the following formula:

[0053] M avg =f2ReLUf1(F avg M max =f2ReLU(f1(F max )),

[0054] Here, f1 and f2 are used to compress and restore the number of channels, respectively, and ReLU() is used as the activation function. Finally, the two attention maps are fused and channel attention weights are generated using the Sigmoid activation function, which are then applied to the original features. The formula is expressed as follows:

[0055]

[0056] In step 2, a perturbation decoder is embedded in a general UNet network framework to construct a dual-decoder structure. Each decoder consists of multiple cascaded convolutional-upsampling units, each unit having two convolutional blocks and one upsampling layer. A channel-space attention module is introduced after each upsampling module, such as... Figure 5 As shown, the weight adjustment of both channel and spatial dimensions is considered simultaneously, so that the fused feature map can effectively suppress background noise, enhance the foreground region, and maintain the integrity and detail consistency of the target structure in the process of restoring spatial resolution layer by layer.

[0057] In step 2, a partial cross-entropy function is used to supervise the learning of the bi-branch prediction results and the graffiti labels, ignoring unlabeled pixels in the graffiti annotations. The mean of the bi-branch prediction loss is taken as the basic graffiti supervision loss. The graffiti supervision loss L... SS The formula is expressed as follows:

[0058]

[0059] Where y1 is the main branch prediction and y2 is the auxiliary branch prediction. L pCE The partial formula for the cross-entropy function is expressed as follows:

[0060]

[0061] Where K is the set of strategies in graffiti annotation, Ω l The set of marker pixels in a doodle. and These are the predicted probabilities of the graffiti element and pixel i belonging to the k-th class, respectively.

[0062] In step 2, the prediction results from the dual decoders, y1 = D1(Encoder(X)) and y2 = D2(Encoder(X)), are used to generate corresponding pseudo-labels, which can be represented as follows: Simultaneously calculate the prediction confidence scores of both decoders. The confidence score can be expressed by the formula:

[0063]

[0064] Where ∈ (∈=10) -6 ) is a very small constant to avoid singularity. This represents the probability of the i-th class. The C2 counting method is the same as above. A corresponding reliability mask is generated based on the confidence score to determine reliable pseudo-labels. For each pixel (x, y), the reliability mask formula is expressed as follows:

[0065]

[0066] To assess reliability, feature representations from the penultimate layer are extracted from both decoders, and pixel features of each category are selected using pseudo-labels. Then, the mean value of each category's features is calculated, and the corresponding class prototype is constructed, expressed by the following formula:

[0067]

[0068] pf1 is the feature of the penultimate layer of D1. This represents the i-th prototype of D1. Finally, calculate pf1 and... Pixel-level cosine similarity between them: The reliability of pseudo-labels is measured by the mean square error between the bi-branch output and the cosine similarity, as expressed by the following formula:

[0069]

[0070] That is, the larger ω is, the more consistent the current pixel is with its predicted class prototype, and the more reliable the pseudo-label. The same method can also be used for decoder D2 to obtain ω2 and calculate the loss L for reliable pseudo-label supervision. Rpl The formula is expressed as follows:

[0071]

[0072] In step 2, due to the sparse annotation information in the graffiti supervision, the model struggles to accurately perceive the complete boundary of the target structure. An edge detection module is introduced into the network, utilizing multi-scale features to model and supervise edge information, guiding the model to focus on key edge regions. First, the deep semantic features of the penultimate layer in the main decoder are fused with the low-level features of the shallow stage in the encoder. Then, the fused features are concatenated with the deep features in the auxiliary decoder and fed into the edge detection head. A channel-spatial attention module is added to enhance edge perception, generating an edge prediction map E that fuses the shallow features. fused Simultaneously, pseudo-edge maps E are generated using pseudo-labels. pred The fused edge prediction map is then subjected to edge loss supervision using Dice loss to enhance the model's ability to identify target edge regions. The edge supervision loss function is denoted as L. esl The formula is expressed as follows:

[0073]

[0074] Where i represents the position of the i-th pixel in the image, H and W are the height and width of the image, respectively, and E pred E fused Let represent the predicted and fused edge prediction maps, respectively. σ(·) is the Sigmoid function, and ∈ is a minimal constant. Furthermore, to further improve the model's prediction consistency in edge regions, the deep features of the two branches are used to generate corresponding edge maps through two independent edge detection heads, and the consistency loss L between the two is calculated. ecl The formula is expressed as follows:

[0075]

[0076] Where E main E aux Let || represent the edge predictions of the main branch and the auxiliary branch, respectively, and |·| represent the absolute value, i.e., the L1 distance. This edge consistency mechanism not only enhances the information interaction between branches but also suppresses noise propagation caused by pseudo-label errors to a certain extent, thereby further improving the overall segmentation performance.

[0077] In step 2, the weighted loss function is obtained by multiplying the graffiti supervision loss, reliable pseudo-label supervision loss, edge supervision loss, and edge consistency loss by their respective weights and then summing them. The formula is expressed as follows:

[0078] L total =λ×L ss +β×L Rpl +γ1×L esl +γ2×L ecl ,

[0079] λ is set to 0.5, and β is set to a time-dependent Gaussian warming phenomenon function. In order to balance the role of edge supervision and edge consistency supervision in the training process, a warm-up strategy is adopted, in which the weight γ1 of the edge loss is linearly increased from 0.2 to 0.5, and the weight γ2 of the edge consistency loss is linearly increased from 0.1 to 0.3, gradually strengthening the edge constraints as the training progresses.

[0080] Step 3 specifically involves:

[0081] The medical image segmentation model is trained using the graffiti labels in the training set. During the training process, the loss function, optimizer function, and learnable hyperparameters used by the medical image segmentation model are continuously optimized until the model's segmentation performance reaches its best.

[0082] The trained medical image segmentation model is validated using the validation set. If the validation results are unsatisfactory, training continues. If the validation results are good, the model is tested using the test set.

[0083] When validating the model's actual segmentation performance using a test set, training ends if the segmentation performance is good, and continues if the segmentation performance is poor.

[0084] Step 4 specifically involves:

[0085] The tested medical image segmentation model is deployed to the server, and the calling interface and access permissions are set. Users input the preprocessed data to be segmented into the medical image segmentation model, which outputs segmentation results with detailed data annotations, completing the segmentation task.

[0086] In summary, the advantages of this invention are as follows:

[0087] Using a publicly available dataset with graffiti annotations, the dataset is divided into training, testing, and validation sets, and the images in the dataset undergo preprocessing. Then, a medical image segmentation model is created based on an encoder, dual decoders, channel-spatial attention blocks, an edge detection module, and a weighted loss function. The validated medical image segmentation model is tested using the training set. The tested and effective medical image segmentation model is deployed on a server and used for medical image segmentation tasks. A channel-spatial attention mechanism is integrated into the encoder and decoder to enhance the model's ability to model local and global information. A reliable pseudo-label learning mechanism is introduced to select more trustworthy pseudo-labels for supervision, significantly reducing noise interference from low-quality pseudo-labels. Furthermore, an edge supervision module is introduced, generating pseudo-edge labels by fusing dual-branch predictions and aligning them with edge information fused with shallow features. Consistency constraints are applied between the dual-branch edge predictions to improve the model's accuracy in boundary region discrimination.

Claims

1. A graffiti-supervised medical image segmentation method based on reliable pseudo-labels and edge detection, specifically implemented by the following steps: Step 1: Select different publicly available medical image datasets with graffiti tags, divide the datasets into training, testing, and validation sets, and perform certain preprocessing operations on the dataset images. Step 2: A medical image segmentation model is created based on the encoder, dual decoders, channel-spatial attention blocks, edge detection module, and weighted loss function. The encoder consists of multiple cascaded convolutional-downsampling modules, with a spatial attention block (SAB) introduced in its shallow layer and a channel attention block (CAB) introduced in its deep layer. For the dual decoders, each decoder consists of multiple cascaded convolutional-upsampling units, each unit having two convolutional blocks and one upsampling layer. A lightweight channel-spatial attention block (CBAM) is introduced after each upsampling module. The edge detection module mainly consists of three independent edge detection heads. Step 3: Train the medical image segmentation model using the graffiti labels in the training set, verify the segmentation effect of the medical image segmentation model using the validation set, and test the validated medical image segmentation model using the test set. Step 4: Deploy the tested and effective medical image segmentation model on the server and use this model to perform medical image segmentation tasks.

2. The graffiti-supervised medical image segmentation method based on reliable pseudo-labels and edge detection according to claim 1, characterized in that: The medical image segmentation datasets in step 1 are the ACDC and MSCMRseg heart datasets. Due to the inconsistent image sizes in the datasets, the images in the ACDC dataset were uniformly resized to 256×256, and the images in the MSCMRseg dataset were uniformly resized to 480×480.

3. The graffiti-supervised medical image segmentation method based on reliable pseudo-labels and edge detection according to claim 1, characterized in that: Using the encoder in step 2, the dataset images are downsampled layer by layer to gradually extract medical image features. The shallow spatial attention module is used to enhance the model's attention to salient regions and edge structures, improving the ability to distinguish between foreground and background. The deep channel attention module is used to strengthen the model's response to channels with deep semantic information, making it more focused on semantic channels related to the segmentation target.

4. The graffiti-supervised medical image segmentation method based on reliable pseudo-labels and edge detection according to claim 1, characterized in that: In step 2, a perturbation decoder is embedded within a general UNet network framework to construct a dual-decoder structure. This enhances branch diversity while effectively mitigating model instability caused by graffiti supervision. Furthermore, each decoder incorporates a channel-space attention module, considering weight adjustments for both channel and spatial dimensions, enabling the fused feature map to effectively suppress background noise and enhance foreground regions. The two decoders have different feature attention regions and prediction boundaries, making them complementary when handling blurred or uncertain regions.

5. The graffiti-supervised medical image segmentation method based on reliable pseudo-labels and edge detection according to claim 1, characterized in that: In step 2, a partial cross-entropy function is used to supervise the learning of the bi-branch prediction results and the graffiti labels, ignoring unlabeled pixels in the graffiti annotations. The mean of the bi-branch prediction loss is taken as the basic graffiti supervision loss.

6. The graffiti-supervised medical image segmentation method based on reliable pseudo-labels and edge detection according to claim 1, characterized in that: The prediction results from the dual decoders in step 2 are used to generate corresponding pseudo-labels. Simultaneously, the prediction confidence scores of the two decoders are calculated, and a corresponding reliability mask is generated based on these confidence scores to determine reliable pseudo-labels. To evaluate reliability, the penultimate layer feature representations are extracted from both decoders, and pixel features of each category are selected using the pseudo-labels. Then, the mean value of each category's features is calculated to construct the corresponding class prototype. Finally, the pixel-level cosine similarity between the penultimate layer features and the class prototype is calculated, and the mean squared error between the dual-branch output and the cosine similarity is used to measure the reliability of the pseudo-labels. A reliable pseudo-label supervision loss is calculated.

7. The graffiti-supervised medical image segmentation method based on reliable pseudo-labels and edge detection according to claim 1, characterized in that: In step 2, the deep semantic features of the penultimate layer in the main decoder are fused with the low-level features of the shallow stage in the encoder. Then, the fused features are concatenated with the deep features in the auxiliary decoder and fed into the edge detection head. A channel-spatial attention module is added to enhance edge perception, generating an edge prediction map that fuses the shallow features. Simultaneously, pseudo-edge maps are generated using pseudo-labels, and these are compared with the fused edge prediction map for edge loss supervision, enhancing the model's ability to identify target edge regions. Finally, the deep features from the two branches are used to generate corresponding edge maps through two independent edge detection heads, and the consistency loss between the two is calculated.

8. The graffiti-supervised medical image segmentation method based on reliable pseudo-labels and edge detection according to claim 1, characterized in that: In step 2, the weighted loss function is obtained by multiplying the graffiti supervision loss, the reliable pseudo-label supervision loss, the edge supervision loss, and the edge consistency loss by their respective weights and then adding them together.

9. The graffiti-supervised medical image segmentation method based on reliable pseudo-labels and edge detection according to claim 1, characterized in that: Step 3 is implemented as follows: The medical image segmentation model is trained using the graffiti labels in the training set. During the training process, the loss function, optimizer function, and learnable hyperparameters used by the medical image segmentation model are continuously optimized until the model's segmentation performance reaches its best. The trained medical image segmentation model is validated using the validation set. If the validation results are unsatisfactory, training continues. If the validation results are good, the model is tested using the test set. When validating the model's actual segmentation performance using a test set, training ends if the segmentation performance is good, and continues if the segmentation performance is poor.

10. The graffiti-supervised medical image segmentation method based on reliable pseudo-labels and edge detection according to claim 1, characterized in that: Step 4 is implemented as follows: The tested medical image segmentation model is deployed to the server, and the calling interface and access permissions are set. Users input the preprocessed data to be segmented into the medical image segmentation model, which outputs segmentation results with detailed data annotations, completing the segmentation task.