Meibomian gland segmentation method and system based on prompt enhancement and medium

By generating positive prompt images and global feature maps, combining multi-scale noise suppression and cross-dimensional shift feature modules, the meibomian gland glands are segmented layer by layer, solving the problems of adhesion and noise interference in meibomian gland segmentation, achieving efficient and accurate segmentation effect.

CN120279050APending Publication Date: 2025-07-08SUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510338574.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art is difficult to effectively segment the meibomian gland glands, especially in the case of adhesions, low contrast, severe noise interference and high computing resources, it is difficult to achieve accurate segmentation.

Method used

The meibomian gland segmentation method based on cue enhancement is adopted. By generating multiple positive cue images and global feature maps, the meibomian gland infrared images are segmented layer by layer, combining the multi-scale noise suppression module and the cross-dimensional shift feature module, dynamically update the segmentation results and gradually optimize the segmentation boundary.

Benefits of technology

显著提升了睑板腺腺体分割的精确性和鲁棒性,减少噪声干扰,适应复杂腺体形态变化,降低计算资源消耗,提高了分割效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279050A_ABST
    Figure CN120279050A_ABST
Patent Text Reader

Abstract

The invention relates to a meibomian gland segmentation method and system based on prompt enhancement and a medium, and the segmentation method comprises the following steps: collecting a meibomian gland infrared image, generating a plurality of positive prompt images according to the meibomian gland infrared image, each positive prompt image respectively comprising a part of meibomian gland, the same positive prompt image does not contain adjacent glands; extracting prompt features of each positive prompt image, extracting global features of the meibomian gland infrared image, and generating a global feature map; and according to the global feature map and the prompt features of each positive prompt image, realizing layer-by-layer segmentation of the infrared meibomian gland image. According to the method, the segmentation result can be refined step by step, the segmentation task of each layer depends on the prediction result of the previous layer, the prompt image is dynamically updated, the adhesion problem between adjacent glands is avoided, and the segmentation precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a meibomian gland segmentation method, system and medium based on prompt enhancement. Background Art

[0002] Traditional medical image segmentation methods, such as threshold-based methods, edge detection algorithms, and region growing algorithms, often rely on manually setting parameters and are sensitive to complex backgrounds and noises. In recent years, deep learning-based methods, especially convolutional neural networks (CNNs), have made significant progress in image segmentation tasks. However, when dealing with objects with similar features or prone to adhesion (such as adjacent meibomian glands), how to effectively segment meibomian glands remains a challenging problem, which is mainly reflected in the following aspects:

[0003] (1) The distance between glands is very close, and adjacent gland adhesion often occurs, which poses a high requirement for the accurate segmentation of gland boundaries; existing image segmentation technologies, including traditional image processing methods and deep learning-based methods, usually have difficulty coping with the above challenges: traditional methods rely on manually setting parameters, are sensitive to noises and image variations, and are difficult to adapt to the image features of different patients. Although deep learning methods have made significant progress in medical image segmentation, they still have limitations in dealing with adjacent gland adhesion and low contrast problems.

[0004] (2) The contrast of meibomian gland infrared images is low, there is a lot of noise, and the details are not obvious enough, which easily leads to the boundary between the meibomian gland structure and the surrounding tissues being blurred and difficult to distinguish. Although UNeXt uses a multi-layer perceptron (MLP) to process features, its local feature learning ability still mainly depends on the convolutional stage. This method is not sufficient to capture complex global information or long-distance dependencies in some cases. In addition, although the channel translation technology can enhance the learning ability of local dependencies, its effect may not be stable and robust enough when segmenting non-uniform or high-noise meibomian gland infrared images. Traditional semantic segmentation networks also fail to well alleviate the adhesion problem between adjacent glands, which is usually caused by reasons such as the low contrast, irregular strip structure of meibomian glands, and the small distance between glands. Although the skip connection design of UNeXt can fuse low-level and high-level features, it may lead to information redundancy or increased sensitivity to noises when dealing with multi-scale features, thus affecting the segmentation performance.

[0005] (3) The tarsal gland mainly presents a strip-shaped structure and there are significant scale differences. An infrared image of the tarsal gland may contain up to 40 glands. SAM and MedSAM rely on a large amount of labeled data during the training process, which increases the difficulty of data collection and processing. In addition, these models use heavyweight image encoders with a large parameter scale, and the input image size is fixed at 1024×1024, lacking flexibility. A large amount of computing resources and time are required during the training and deployment processes, making it difficult to meet the needs of real-time medical treatment. If all tarsal gland glands need to be completely segmented, the model may need to perform dozens of inference operations with high computational consumption. In addition, the inference speed of SAM and MedSAM also depends on the search for foreground points. When the foreground points are sparse, many fine-grained or small meaningful objects may be missed; when the foreground points are dense, the model needs to repeatedly run the mask decoder to generate a large number of candidate masks and select high-quality and non-overlapping masks, which transfers the computational bottleneck from image encoding to the mask generation and filtering stages.

[0006] (4) Spatiotemporal neural networks improve the segmentation performance by fusing spatiotemporal information, but their computational overhead is large. Especially when processing high-resolution videos or long video sequences, additional spatiotemporal information modeling may significantly reduce the inference speed. In addition, there may be redundant information between consecutive frames, especially in the case of small scene changes, which leads to information redundancy and increased sensitivity to noise. When the network processes consecutive frames, it may be too sensitive to the noise or small changes in the previous and subsequent frames, resulting in instability of the segmentation results. At the same time, since spatiotemporal neural networks need to balance the modeling of time information and spatial details, their ability to capture small targets or fine structures may be insufficient, thus limiting the improvement of segmentation performance. Summary of the Invention

[0007] Therefore, the technical problem to be solved by the present invention is to overcome the problems in the prior art such as low contrast, serious noise interference, and blurred gland boundaries in infrared images of the tarsal gland, and provides a method for segmenting tarsal gland glands based on prompt enhancement, which can effectively alleviate the problem of adjacent gland adhesion in tarsal gland gland segmentation. The segmentation process of the prompt image is regarded as a time series, and the prompt image and segmentation result of each layer depend on the result of the previous layer, gradually generating a complete segmentation result, enhancing the accuracy of the gland segmentation boundary.

[0008] To solve the above technical problems, the present invention provides a method for segmenting tarsal gland glands based on prompt enhancement, including the following steps:

[0009] S1: Collect infrared images of the tarsal gland, and generate multiple positive prompt images according to the infrared images of the tarsal gland. Each positive prompt image respectively contains some tarsal gland glands, and adjacent glands are not included in the same positive prompt image;

[0010] S2: Extract the prompt features for each of the positive prompt images, and at the same time extract the global features from the infrared meibomian gland image to generate a global feature map;

[0011] S3: Based on the global feature map and the prompt features of each positive prompt image, perform layer-by-layer segmentation of the infrared meibomian gland image; during the layer-by-layer segmentation process, the prediction result of the previous layer is dynamically updated to generate a negative prompt image for the next layer of segmentation, which is used to mark the segmented gland area; the positive prompt image of the next layer is adjusted according to the unsegmented gland area.

[0012] In an embodiment of the present invention, in step S2, the method for extracting the global features from the infrared meibomian gland image is as follows: Extract the global features of the infrared meibomian gland image through an image encoder to generate a global feature map, and suppress noise through a multi-scale noise suppression module and enhance the local perception ability of the global feature map through a cross-dimensional shift feature module during the process of generating the global feature map.

[0013] In an embodiment of the present invention, in step S2, the method for extracting the prompt features from the positive prompt image is as follows: Extract the prompt features for each positive prompt image through a prompt encoder, and enhance the alignment ability between the prompt features and the image features through the cross-dimensional shift feature module during the process of extracting the prompt features.

[0014] In an embodiment of the present invention, in step S3, the method for layer-by-layer segmentation of the infrared meibomian gland image is as follows: The method for layer-by-layer segmentation of the infrared meibomian gland image is as follows: Achieve layer-by-layer segmentation of the infrared meibomian gland image through a mask decoder according to the prompt features and the global feature map; and further suppress the noise features through a multi-scale noise suppression module during the layer-by-layer segmentation process, and enhance the cross-dimensional information interaction and the local perception ability through a cross-dimensional shift feature module.

[0015] In an embodiment of the present invention, the multi-scale noise suppression module includes:

[0016] A multi-scale feature extraction module that extracts multi-scale features from the input feature map using multiple convolutional kernels;

[0017] A global context fusion module that extracts the global information of the input feature map through global average pooling operation;

[0018] A residual shrinkage unit for generating a channel-level scaling factor and denoising the features through a soft thresholding mechanism;

[0019] The working principle of the multi-scale noise suppression module MSNS is:

[0020] F2 = Conv 1×1(GAP(F in ))

[0021] F2 = Conv 1×1 (GAP(F in ))

[0022] F3 = ReLU(BN(Conv 3×3 (ReLU(BN(Conv 3×3 (F in ))))))

[0023] F2 out = RSBU(F2)

[0024] F1 = ReLU(BN(Conv 3×3 (F in )))

[0025] F merge = Conv 1×1 (F2 out + F3 out ) + F1

[0026] F out = ReLU(BN(Conv 3×3 (F merge )))

[0027] Among them, F in represents the input feature map, Conv 3×3 represents a 3×3 convolution, BN represents batch normalization, ReLU represents the ReLU activation, Conv 1×1 represents a 1×1 convolution, GAP represents global average pooling, F1 and F3 represent the features extracted at different scales, F2 represents the global information after global average pooling, F2 out and F3 out represent the output features after passing through the residual shrinkage unit RSBU-CS, F merge represents the fused feature of F1, F2 out and F3 out , and F out represents the output feature of the multi-scale noise suppression module MSNS.

[0028] Among them, the internal working principle of the residual shrinkage unit RSBU-CS is as follows:

[0029] F abs = |F i |

[0030] θ = σ(FC(ReLU(BN(GAP(F abs )))))

[0031] F o = Sign(F i )·max(| F i| - θ, 0)

[0032] Wherein, F i represents the input feature, F abs represents the feature after taking the absolute value, FC represents the fully connected layer, θ represents the scaling factor, σ represents the Sigmoid activation function, Sign represents the Sign activation function, max(·) represents the maximum operation, and F o represents the output feature of the residual shrinkage unit.

[0033] In an embodiment of the present invention, the first - layer segmentation of the meibomian gland infrared image is used to predict a partial main body of the target gland. The subsequent segmentation depends on the prediction result of the previous layer, and the prediction result of the previous layer automatically generates a negative hint image for the next layer to avoid the network from repeatedly predicting the same gland;

[0034] Wherein, the process of layer - by - layer segmentation of the meibomian gland is as follows:

[0035] First - layer segmentation: Use the hint encoder of the first layer to process the hint image of the first layer, and generate the prediction result of the meibomian gland of the first layer through the mask decoder of the first layer;

[0036] Second - layer segmentation: Use the hint encoder of the second layer to process the hint image of the second layer, combine the prediction result of the meibomian gland of the first layer, and generate the prediction result of the meibomian gland of the second layer through the mask decoder of the second layer;

[0037] Subsequent - layer segmentation: Use the hint encoder of the current layer to process the hint image of the current layer, combine the prediction results of the meibomian gland of the previous several layers, and generate the prediction result of the meibomian gland of the current layer through the mask decoder of the current layer;

[0038] Integrate the results: Integrate the prediction results of each layer through the XOR operation to generate the overall prediction result of the meibomian gland.

[0039] In an embodiment of the present invention, the negative hint image processed by the hint encoder in the first - layer segmentation is a black image with all gray - scale values being 0.

[0040] In an embodiment of the present invention, in the decoding stage of the mask decoder, the down - sampled feature map of each layer is passed to the decoding path through skip connections.

[0041] Second, to solve the above - mentioned technical problems, the present invention provides a meibomian gland segmentation system based on hint enhancement, including:

[0042] An image acquisition module, used to acquire meibomian gland infrared images;

[0043] An image preprocessing module for preprocessing the acquired images;

[0044] An image encoder for extracting global features of the input image and generating a global feature map;

[0045] A prompt image generation module for generating multiple prompt images based on the input image;

[0046] There are three prompt encoders for processing the prompt images and extracting prompt features;

[0047] There are three mask decoders, each mask decoder is used in cooperation with a prompt encoder for combining the global feature map and the prompt features to gradually complete the segmentation of the target gland;

[0048] A segmentation result output module for outputting the final segmentation result.

[0049] In a third aspect, to solve the above technical problems, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method for meibomian gland segmentation based on prompt enhancement described in the first aspect.

[0050] The above technical solutions of the present invention have the following beneficial effects compared with the prior art:

[0051] (1) For the method for meibomian gland segmentation based on prompt enhancement described in the present invention, multiple prompt images are regarded as "video frames" with a time sequence, and the segmentation result is gradually optimized. This method effectively solves the adhesion problem between meibomian glands, significantly improves the accuracy of the segmentation boundary and gland morphology, and at the same time guides the segmentation network to more accurately complete the meibomian gland segmentation task by introducing prompt information. This mechanism combines the dynamic generation of prompt images and the step-by-step optimization strategy, significantly improving the accuracy and robustness of the segmentation.

[0052] (2) In the present invention, the multi-scale noise suppression module (MNSM) can not only suppress the accumulation of noise features and reduce the interference of noise on the segmentation result, but also enhance the multi-scale feature expression, improve the network's adaptability to complex gland morphologies, enhance the global perception ability of features, help the network better understand the overall structure of the gland, improve the segmentation accuracy, enhance the robustness of the network, enable it to better process images with low contrast and high noise, dynamically adapt to gland morphology changes, and improve the adaptability and generalization ability of the segmentation.

[0053] (3) The cross-dimensional shift feature module (CSFM) in the present invention enhances the information interaction ability between features by performing cross-dimensional shift operations in the spatial and channel dimensions and combining the window attention mechanism. This module can effectively capture the strip structure and morphological features of the meibomian glands, significantly improving the adaptability of the model to complex glandular distributions. Brief Description of the Drawings

[0054] To make the content of the present invention easier to understand clearly, the following further details the present invention according to specific embodiments of the present invention and in combination with the accompanying drawings, where

[0055] Figure 1 is a flowchart of the meibomian gland segmentation method based on prompt enhancement in the preferred embodiment of the present invention;

[0056] Figure 2 shows the overall architecture of the three-layer segmentation network based on prompt enhancement proposed by the present invention;

[0057] Figure 3 shows the specific structures of the image encoder, prompt encoder, and mask decoder;

[0058] Figure 4 shows the detailed structure of the multi-scale noise suppression module (MNSM);

[0059] Figure 5 details the structure of the cross-dimensional shift feature module (CSFM);

[0060] Figure 6 shows an infrared image of the meibomian gland in Dataset 1 and its corresponding gland gold standard.

[0061] Figure (a) is the original infrared image of the meibomian gland;

[0062] Figure (b) is the gland gold standard (white marks) manually annotated by professionals;

[0063] Figure 7 shows an infrared image of the meibomian gland in Dataset 2 and its corresponding gland gold standard;

[0064] Figure (a) is the original infrared image of the meibomian gland;

[0065] Figure (b) is the gland gold standard (white marks).

[0066] Figure 8 shows the segmentation visualization results of the three-layer segmentation network based on prompt enhancement proposed by the present invention and its comparison network on Dataset 1;

[0067] Figure 9Shows the segmentation visualization results of the three - layer segmentation network enhanced by prompts and its comparison network on Dataset 2;

[0068] Figure 10 Shows the magnified local segmentation visualization results of the three - layer segmentation network enhanced by prompts and its comparison network on Dataset 1;

[0069] Figure 11 Shows the magnified local segmentation visualization results of the three - layer segmentation network enhanced by prompts and its comparison network on Dataset 2;

[0070] Figure 12 For the comparison of segmentation results of different models on the meibomian glands;

[0071] Figure 13 For the ablation experiment results of the three - layer segmentation network enhanced by prompts. Detailed implementation manners

[0072] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and implement it, but the specific embodiments cited do not limit the present invention.

[0073] Refer to Figure 1 As shown, a method for meibomian gland segmentation based on prompt enhancement provided by the present invention includes the following steps:

[0074] S1: Collect infrared images of meibomian glands, and generate three positive prompt images according to the infrared images of meibomian glands. Each positive prompt image contains part of the meibomian glands, and adjacent glands are not included in the same positive prompt image;

[0075] S2: Extract prompt features for each positive prompt image through a prompt encoder respectively, and at the same time extract global features from the infrared image of meibomian glands through an image encoder to generate a global feature map;

[0076] S3: Receive the global feature map generated by the image encoder and the prompt features output by each prompt encoder through three mask decoders respectively to realize the layer - by - layer segmentation of the infrared image of meibomian glands; in the layer - by - layer segmentation process, the prediction result of the previous layer is dynamically updated to generate a negative prompt image for the next - layer segmentation to mark the segmented gland area; the positive prompt image of the next layer is adjusted according to the unsegmented gland area.

[0077] As Figure 2 shown, the meibomian gland segmentation method based on prompt enhancement in the present invention is implemented by using a three - layer segmentation network enhanced by prompts (TriPromptNet). The overall working principle of the three - layer segmentation network enhanced by prompts (TriPromptNet) can be expressed as:

[0078] E i = ImageEncoder(I o )

[0079] E p1 = PromptEncoder1(I p1 ), I p1 = Cat(Pos1, Neg1)

[0080] E p2 = PromptEncoder2(I p2 ), I p2 = Cat(Pos2, Neg1 + O1)

[0081] E p3 = PromptEncoder3(I p3 ), I p3 = Cat(Pos3, Neg1 + O1 + O2)

[0082] O1 = MaskDecoder1[Identity(E p1 ) + Identity(E i )]

[0083] O1 = MaskDecoder1[Identity(E p1 ) + Identity(E i )]

[0084] O3 = MaskDecoder3[Identity(E p3 ) + Identity(E i )]

[0085] O3 = MaskDecoder3[Identity(E p3 ) + Identity(E i )]

[0086] Among them, EmageEncoder represents the image encoder, PromptEncoder i represents the prompt encoder, the subscript number i represents the i-th layer encoder, MaskDecoder i represents the mask decoder, and the subscript number i is the same as the prompt encoder, I o represents the meibomian gland image, E i represents the set of image features after downsampling at each layer encoded by the image encoder, E p represents the set of prompt features, P osdenotes the positive prompt image, Neg1 denotes the negative prompt image of the first - layer prompt encoder (a black image with all gray - scale values being 0), and the image size is the same as that of P os C at denotes feature concatenation in the channel dimension, Identity represents the skip connection, XOR represents exclusive - or, O i represents the output probability map of the mask decoder of the i - th layer, σ represents the Sigmoid activation function, II(·) represents the indicator function (outputs 1 when the condition is true and 0 otherwise), and O A denotes the final segmentation result.

[0087] Among them, the image encoder and the mask decoder both include a multi - scale noise suppression module and a cross - dimensional shift feature module. Among them, the multi - scale noise suppression module is used to effectively suppress noise features and improve the ability of the segmentation model to capture gland boundaries and detailed features; the cross - dimensional shift feature module is used to enhance the interaction and fusion of features in the spatial and channel dimensions and improve the segmentation performance of the model for glandular strip structures.

[0088] Furthermore, the prompt encoder includes a cross - dimensional shift feature module, which is used to enhance the alignment ability between prompt features and image features and improve the segmentation accuracy.

[0089] Preferably, the multi - scale noise suppression module includes:

[0090] A multi - scale feature extraction module that extracts multi - scale features from the input feature map using multiple convolutional kernels;

[0091] A global context fusion module that extracts the global information of the input feature map through global average pooling operation;

[0092] A residual shrinkage unit that is used to generate a channel - level scaling factor and denoise the features through a soft - thresholding mechanism.

[0093] In this embodiment, the multi - scale noise suppression module combines multi - scale feature extraction with global context information, and by means of the soft - thresholding strategy of the residual shrinkage unit RSBU - CS, adaptively removes noise - related features while retaining significant features closely related to the tarsal gland structure.

[0094] This module can effectively capture features of different scales in the tarsal gland image, enhance the feature expression ability of the network; through the soft - thresholding mechanism, accurately suppress the interference of background noise on the segmentation result; at the same time, use global information for feature fusion to improve the robustness of the segmentation network to complex image details. The structure of the multi - scale noise suppression module MSNS is as shown in the appendix Figure 4 shown.

[0095] First, multi-scale features are extracted from the input feature map f through the multi-scale feature extraction module in Specifically, it is divided into two branches: The first branch extracts shallow features F1 through a 3×3 convolutional layer, combined with batch normalization and ReLU activation function, and the number of channels changes from 32 to 128;

[0096] The second branch extracts deeper features F3 through two consecutive 3×3 convolutional operations, and the number of channels changes successively to 32, 64, and 128. Next, through the global context fusion module, the global information F of the input features is extracted using global average pooling (GAP) global and the channel is compressed through a 1×1 convolution to generate the feature F2 to provide global feature support. In the feature denoising stage, the module applies the residual shrinkage unit RSBU-CS to F2 and F3 respectively. RSBU-CS calculates the absolute value of the input feature F i and performs global average pooling, combined with batch normalization, fully connected layer, and Sigmoid function, to generate a channel-level scaling factor θ. Subsequently, the soft thresholding mechanism is used to denoise the features, effectively removing noise-related features and retaining their significant feature information.

[0097] Finally, in the feature fusion stage, the denoised features F2 out and F3 out output by RSBU-CS are added to F1 after channel information interaction, feature recombination, and feature selection through 1×1 convolution to obtain the fused feature F merge . Then, through another 3×3 convolutional layer, combined with batch normalization and ReLU activation function, the fused feature is refined to generate the final output feature F out . The design of the entire module can extract and fuse features at different scales, while removing noise, improving the robustness and feature expression ability of the network.

[0098] The working principle of the multi-scale noise suppression module MSNS can be expressed as:

[0099] F2 = Conv 1×1 (GAP(F in ))

[0100] F2 = Conv 1×1 (GAP(F in ))

[0101] F3 = ReLU(BN(Conv 3×3 (ReLU(BN(Conv 3×3 (F in ))))))

[0102] F2 out=RSBU(F2)

[0103] F1 = ReLU(BN(Conv 3×3 (F in )))

[0104] F merge = Conv 1×1 (F2 out + F 3out ) + F1

[0105] F out = ReLU(BN(Conv 3×3 (F merge )))

[0106] Among them, F in represents the input feature map, Conv 3×3 represents a 3×3 convolution, BN represents batch normalization, ReLU represents the ReLU activation, Conv 1×1 represents a 1×1 convolution, GAP represents global average pooling, F1 and F3 represent the features extracted at different scales, F2 represents the global information after global average pooling, F2 out and F3 out represent the output features after passing through the residual shrinkage unit RSBU-CS, F merge represents the fused feature of F1, F2 out and F3 out , and F out represents the output feature of the multi-scale noise suppression module MSNS.

[0107] Among them, the internal working principle of the residual shrinkage unit RSBU-CS is as follows:

[0108] F abs = |F i |

[0109] θ = σ(FC(ReLU(BN(GAP(F abs )))))

[0110] F o = Sign(F i )·max(|F i | - θ, 0)

[0111] Among them, F i represents the input feature, F abs represents the feature after taking the absolute value, FC represents the fully connected layer, θ represents the scaling factor, σ represents the Sigmoid activation function, Sign represents the Sign activation function, max(·) represents the maximum operation, and F o represents the output feature of the residual shrinkage unit.

[0112] To address the problem of significant differences in the morphological characteristics of meibomian glands, the present invention proposes a Cross-Dimension Shift Feature Module (CSFM). The morphology of meibomian glands exhibits non-uniformity in spatial distribution and has specific characteristic patterns in the channel dimension. Therefore, the CSFM module realizes local perception and global fusion of features across the height, width, and channel dimensions through a shift operation similar to the window attention mechanism, enhancing the information interaction ability between features, effectively capturing the multi-scale morphological characteristics of meibomian glands, thereby improving the performance and diagnostic ability of the model. The structure of the cross-dimension shift feature module CSFM is as shown in the appendix Figure 5 as follows

[0113] At the beginning of the module process, input processing is first performed, and the input features are mapped to a new feature space through a 3×3 convolution operation and layer normalization, and are transformed into features with a fixed dimension through shape reshaping This can adjust the dimension arrangement of the feature map, enabling flexible implementation of operations across the height, width, and channel dimensions, and simultaneously explicitly representing the features of each dimension

[0114] In the Shifted MLP module, M in is divided into an odd number n of sub-channel blocks wherein each sub-block is shifted in different directions (height, width, channel), and the shift amount is where k = 0, 1,..., n - 1. The specific rules of the shift can be dynamically adjusted to ensure that the global feature distribution is not damaged. The shifted feature maps are merged through a concatenation operation to generate a new feature representation By shifting in different axial directions in each sub-block, the Shifted MLP module can effectively capture feature locality, dynamically simulate the characteristics of a random window, introduce feature perturbations, enhance the generalization ability and robustness of the network. However, there is no need to introduce explicit attention calculation, thereby improving the calculation efficiency

[0115] In this embodiment, the structure of the entire cross-dimension shift feature module CSFM has three branches

[0116] The first branch is to obtain spatial local features to strengthen the capture of the lateral features and directional distribution of the glands

[0117] The second branch introduces local interaction in the channel dimension to fuse the feature patterns captured in different channels, which helps to enhance the global feature expression of the gland morphology

[0118] The third branch is to perform a residual connection between the original features and the fused features of the aforementioned space and dimension. In previous discussions, the use of depthwise separable convolution (DWConv) can help encode the position information of features, and its use of fewer parameters improves efficiency. The GELU activation is used because its performance is better than that of ReLu.

[0119] The working principle of the cross-dimensional shift feature module (CSFM) can be expressed as:

[0120] P = Reshape(LN(Conv 3×3 (F in )))

[0121] M uidth = ShiftedMLP width (P)

[0122] M spatial = FC(ShiftedMLP Height (GELU(DWConv(FC(M width )))))

[0123] M channel = FC(ShiftedMLP channel (P))

[0124] F out = LN(M spatial + M channel ) + P

[0125] Among them, F in represents the input feature map, Conv 3×3 represents a 3×3 convolution, LN represents layer normalization, Reshape represents a shape reshaping operation, P represents the reshaped feature map, ShiftedMLP represents a shifted multi-layer perceptron module, DWConv represents depthwise separable convolution, GELU represents GELU activation, FC represents a fully connected layer, M spatial represents the obtained spatial local features, M channel represents the obtained channel features, and F out represents the module output features.

[0126] In this embodiment, the cross-dimensional shift feature module includes:

[0127] An input feature processing unit that reshapes the input feature map into a fixed dimension;

[0128] A shifted multi-layer perceptron that performs shift operations on the reshaped feature map in the spatial and channel dimensions;

[0129] A feature fusion unit that fuses the shifted features with the original features to generate an enhanced feature representation.

[0130] In this embodiment, the first layer of segmentation is used to predict a partial body of the target gland. Subsequent segmentations rely on the segmentation results of the previous layer. The prediction results of the previous layer automatically generate negative cues for the next layer to avoid the network from repeatedly predicting the same gland.

[0131] Among them, the process of layer-by-layer segmentation of the tarsal gland is specifically as follows:

[0132] First layer segmentation: The cue encoder of the first layer processes the cue image of the first layer, and the first layer of tarsal gland prediction results are generated through the mask decoder of the first layer.

[0133] Second layer segmentation: The cue encoder of the second layer processes the cue image of the second layer, combines the prediction results of the first layer of tarsal gland, and generates the prediction results of the second layer of tarsal gland through the mask decoder of the second layer.

[0134] Third layer segmentation: The cue encoder of the third layer processes the cue image of the third layer, combines the results of the previous two layers, and generates the prediction results of the third layer of tarsal gland through the mask decoder of the third layer.

[0135] Integrate the results: Integrate the prediction results of each layer through an exclusive OR operation to generate all the prediction results of the tarsal gland.

[0136] Among them, the negative cue image of the cue encoder of the first layer is a black image with all gray values being 0.

[0137] Compared with the traditional method of directly splicing features, the present invention processes the infrared image (query image) and the cue image of the tarsal gland respectively based on the three-layer segmentation network TriPromptNet with cue enhancement.

[0138] Specifically: The task of feature extraction of the query image is completed by the image encoder, and the learning difficulty of the query image is higher than that of the cue image; the cue image extracts features separately through the cue encoder, and the cue focus of each layer is different to ensure that the model can comprehensively capture the local features and global information of the gland.

[0139] In this embodiment, the input of the image encoder is the infrared image of the tarsal gland (query image), and the input image is a grayscale image. The encoding path includes five downsampling operations. The number of feature channels in each layer is the same as that of the UNeXt model, which are 16, 32, 128, 160, and 256 respectively. The design of the number of channels can minimize the number of model parameters and reduce the computational complexity while ensuring the feature extraction ability. The sizes of the feature maps after passing through each layer of the encoding path are 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image respectively. Each of the first two downsamplings includes a 3×3 convolution (stride of 1, padding of 1), batch normalization (BatchNormalization), ReLU activation, and max pooling operations. In the third layer, the multi-scale noise suppression module MNSM is used to extract more effective information in the shallow features and suppress useless noise features. After max pooling, the cross-dimension shift feature module (CSFM) is used to promote the interaction and fusion of features through shift operations across spatial and channel dimensions, thereby enhancing the segmentation ability of the model. Attached Figure 3 shows the overall structures of the image encoder and the prompt encoder.

[0140] To avoid the influence of the number of prompts on the model prediction results and inference speed, the present invention adopts an image-based prompt input method, including positive prompts and negative prompts. The prompt image has the same size as the original image, which is convenient for the model to locate the gland area and also supports the flexible use of patterns of different shapes and sizes as prompts.

[0141] Positive prompts are used to mark the gland area, and the present invention uniformly uses a 3×3 rectangular pixel representation. Negative prompts are automatically generated from the prediction results of the previous architecture to prompt the model to ignore the predicted targets, thus avoiding repeated predictions. The prompt image is a simple binary image with a fixed number of channels of 2 (one channel for positive prompts and the other channel for negative prompts). Similar to the dataset mask generation process, classification labels are assigned to the prompt image to distinguish different types of prompts.

[0142] Corresponding to the image encoding path, the prompt encoding path includes a total of five layers. Different from the image encoder, in the third downsampling, a 3×3 convolution (stride of 1, padding of 1), batch normalization, ReLU activation, and max pooling are used. This is because the input prompt image is a simple binary image and does not require excessive parameters for learning.

[0143] In the last two downsampling processes, the cross-dimension shift feature module CSFM is used to better align with the image features when shifting the features.

[0144] In the decoding stage, the downsampled feature maps of each layer are passed to the decoding path through skip connections. These skip connections can re-inject the high-level semantic information extracted in the encoder stage during the decoding stage, and at the same time make full use of the feature information of different levels, thereby improving the segmentation effect.

[0145] The decoding path is the "reverse" process of the image encoding path. The output features of each layer of the decoder are upsampled to the features of a fixed resolution through bilinear interpolation, and then the output features of this layer are fused with the features of the image encoder and the prompt encoder as the input of the next layer of the decoder. The overall structure of the mask encoder is as shown in the appendix Figure 3 as follows.

[0146] Embodiment 2

[0147] The present invention provides a meibomian gland segmentation system based on prompt enhancement, including:

[0148] An image acquisition module for acquiring infrared images of meibomian glands;

[0149] An image preprocessing module for preprocessing the acquired images;

[0150] An image encoder for extracting global features of the input image and generating global feature maps;

[0151] A prompt image generation module for generating multiple prompt images according to the input image;

[0152] There are three prompt encoders, which are used to process the prompt images and extract prompt features;

[0153] There are three mask decoders, and each mask decoder is used in cooperation with a prompt encoder to combine the global feature map and the prompt features, and gradually complete the segmentation of the target gland;

[0154] A segmentation result output module for outputting the final segmentation result.

[0155] Embodiment 3

[0156] The present invention provides a computer-readable storage medium, on which a computer program is stored. The computer program, when executed by a processor, implements the meibomian gland segmentation method based on prompt enhancement described in Embodiment 1.

[0157] The present invention verifies the meibomian gland segmentation method based on prompt enhancement through experiments. The specific process is as follows:

[0158] (1) Dataset

[0159] The present invention uses two datasets. Dataset 1 is collected from the instrument Keratograph 5M and contains 1,750 infrared images of meibomian glands from 471 patients. The image size is 1024×512. Under the supervision of senior ophthalmologists, trained professionals manually annotate the segmentation labels. Among them, there are 876 images of the upper eyelid and 874 images of the lower eyelid. The present invention divides the 1,750 images into 1,200 for the training set, 250 for the validation set, and 300 for the test set.

[0160] To ensure the independence of the data, four infrared images of meibomian glands of a single person will not appear in the same partition set at the same time. Figure 6 Figure 5 shows an infrared image of the meibomian gland in Dataset 1 and the corresponding gland gold standard.

[0161] Dataset 2 contains 505 infrared images of 300 subjects, with a size of 1024x1280. All the images are from 300 different eyes. In the present invention, the 505 images are divided into 355 for the training set, 50 for the validation set, and 100 for the test set. Figure 7 Figure 10 shows an infrared image of the meibomian gland in Dataset 2 and the corresponding gland gold standard.

[0162] In the experiment, several data augmentation methods are adopted, including horizontal flipping, image rotation within 30 degrees, random size cropping, random blurring, sharpening, and histogram equalization.

[0163] It is worth mentioning that in the prompt segmentation experiment, a method of randomly selecting prompts is used for network training. This is a powerful image enhancement technique. In this case, randomly selected points simulate different input conditions, prompting the model to adapt to a wider range of variations. This can be regarded as a kind of "perturbation" or "mutation" of the data. By diversifying the distribution of training samples, the model can more robustly handle different actual situations.

[0164] (2) Training and validation strategies of the model

[0165] The present invention uses the meibomian gland dataset to perform full-supervised learning on the model. The training process uses complete and accurate segmentation labels, and the model completes the segmentation task by learning the mapping relationship between the data and the labels.

[0166] During model training, for a single target gland, 1 - 11 positive prompt points (3×3 pixel grid) are randomly generated, and in the training, two situations are covered:

[0167] All gland segmentation: The three positive prompt images contain the prompts of all gland targets. After using the OpenCV library to extract the gland contours, the arrangement order of the glands is randomly shuffled to ensure that there are no positive prompts of adjacent glands in the same positive prompt image as much as possible;

[0168] Single gland segmentation: This segmentation situation is similar to SAM. Each time a prompt segmentation is performed, only a single target gland is segmented, and there is only a positive prompt for a single target on the prompt image.

[0169] Considering extreme cases, such as when the positive prompt point is not on the target, during training, when making positive prompt points, we randomly add positive prompts to positions without glands to simulate false prompt situations and enhance the robustness of the model.

[0170] The size of the image input to the model is 512×512. In the interactive verification, two trained medical staff and a professional doctor were invited to simulate real-world interactions for prompt-based segmentation. Each member was required to prompt the given pictures, and the users were required to prompt the images naturally without considering the model performance, and there was no limit to the number of prompts.

[0171] (3) Training settings and evaluation metrics of the model

[0172] The model proposed in the present invention was experimented on the public platform Pytorch (a deep learning framework) using an NVIDIA RTX 3090 graphics processor with 24GB of video memory. During training, the Adam optimization algorithm with an initial learning rate of 0.001 was adopted, and the weight decay parameter was set to 0.0001 to control the intensity of the regularization term, thereby helping to prevent overfitting. In addition, the Cosine Annealing Learning Rate Scheduler was used to set the minimum learning rate to 0.00001 to ensure that the model converges to the optimal solution. The training batch size was 8, and a total of 800 training epochs were carried out.

[0173] To comprehensively evaluate the segmentation performance of the proposed three-layer segmentation network TriPromptNet based on prompt enhancement for tarsal glands, the present invention uses the Dice Similarity Coefficient (DSC), Intersection over Union (IoU), and Variance of Target Count Difference (VTCD) as evaluation metrics. VTCD is an evaluation metric used for the first time to evaluate tarsal gland segmentation, and the formula is as follows:

[0174]

[0175] where n i represents the actual number of glands on the nth image, It represents the number of glands predicted on the nth image.

[0176] (4) Results

[0177] (a) Comparative experiments. To comprehensively evaluate the meibomian gland segmentation performance of the three-layer segmentation network TriPromptNet based on prompt enhancement proposed in the present invention, comparative experiments were conducted between TriPromptNet and multiple advanced semantic segmentation networks (TransUNet, UNet++, MA-Net, UNext, DeepLabV3++, LinkNet, U-Net, IF-Net, UNet 3+, BiO-Net) and prompt segmentation networks (SAM, MedASAM, MobileSAM, LiteMedSAM). Table 1 in the appendix shows the results of the comparative experiments.

[0178] In Figure 12 it, the three-layer segmentation network TriPromptNet based on prompt enhancement proposed in the present invention achieved the optimal experimental results in the meibomian gland segmentation task.

[0179] For dataset 1, TriPromptNet reached 81.29%, 68.57%, and 7.2856 in the DSC, IoU, and VTCD metrics respectively, far higher than other comparative network models.

[0180] For dataset 2, the DSC, IoU, and VTCD metrics of TriPromptNet were 83.54%, 71.77%, and 4.9800 respectively, all exceeding the performance of current mainstream segmentation networks, demonstrating the superiority of the method of the present invention.

[0181] TransUNet combines the self-attention mechanism of Transformer and the encoding-decoding structure of U-Net, and can handle long-distance dependencies, but its ability to finely segment the boundaries of meibomian glands is weak. MA-Net adaptively integrates local features and global dependencies through self-attention mechanism, multi-scale fusion attention block and position attention block to enhance feature extraction, but it fails to pay enough attention to boundary features, which limits its ability to process complex glandular morphology. UNet++ not only uses short connections, but also adds long connections to the network, and realizes multi-level fusion of features through nested U-Net structure, but introduces a large number of parameters, increases training costs, and is insufficient in boundary detection. DeepLabV3+ enhances the ability to capture multi-scale information through dilated convolution, but has limited attention to bar targets, and its segmentation accuracy is not as good as TriPromptNet. LinkNet simplifies the network through jump connections, but its description of the contour details of complex glands is insufficient. UNet 3+ enhances the use of contextual information through the combination of long and short jump connections, but it does not pay enough attention to the boundaries of glands, resulting in affected segmentation accuracy and the introduction of more parameters. UNext improves the traditional U-Net structure and transforms the convolution stage of the network into a tokenized MLP to reduce the model complexity and improve the feature extraction capability. Although the network shows a certain efficiency in processing strip-shaped gland targets, it lacks a detailed description of the gland boundary, which limits the segmentation accuracy in the segmentation task of complex morphological glands. BiO-Net optimizes the encoder-decoder structure by using recursive bidirectional connections, maps the decoded features back to the encoder, and recurs between the encoder and decoder, which can perform better feature refinement, but still cannot solve the adhesion problem of meibomian glands. IF-Net optimizes and fuses the information of the channel dimension and spatial dimension of the input features through the information fusion module, and alleviates the problem of background noise introduced by the jump connection and the loss of semantic information of the feature information during the jump connection process through the parallel path connection module, but it is not effective in solving the adhesion problem of the gland boundary.

[0182] In addition, the cue-based segmentation methods such as SAM and its variants (MedSAM, MobileSAM and LiteMedSAM) have improved the segmentation performance to a certain extent, and are better than the semantic segmentation network in terms of VTCD (variance of target number differences) indicators. In addition to LiteMedSAM, DSC and IoU are also better than the semantic segmentation network. However, compared with the method of the present invention, they are still lacking, perform poorly, and have difficulty in accurately capturing the details of the glandular structure. The reason why the DSC and IoU of LiteMedSAM have not been significantly improved may be that the input image size is fixed at 256×256, resulting in loss of details, limited feature expression capabilities, and unable to segment the fine boundaries of the meibomian glands.

[0183] In contrast, the TriPromptNet proposed in the present invention fully combines the prompt segmentation mechanism, significantly enhancing the network's attention to gland contours and strip regions. Its DSC and IoU reach 81.29% and 68.57% (Dataset 1) and 83.54% and 71.77% (Dataset 2) respectively, performing optimally among all methods. In addition, TriPromptNet also achieves the lowest values of 7.2856 and 4.9800 in the VTCD metric, indicating its significant advantages in maintaining the integrity of gland structure and reducing morphological variations.

[0184] (b) Ablation experiment

[0185] To verify the effectiveness of the multi-scale noise suppression module (MSNS), cross-dimensional shift feature module (CSFM), and multiple prompt-guided segmentation structure in meibomian gland segmentation, detailed ablation experiments were conducted. All experiments were trained and tested with the same hyperparameters and used consistent evaluation metrics to ensure the fairness and rationality of the experiments. The experimental design and results are as shown Figure 13 as follows.

[0186] The baseline network for Experiment ① uses the first three layers of UNext as the Baseline. Its structure is lightweight, with DSC and IoU being 69.24% and 53.51% respectively, while the FPS is 311.9 img / s, and the FLOPs and the number of parameters are 1360.53M and 0.09M respectively. This shows that the Baseline has a relatively high inference speed in lightweight design, but due to insufficient depth, its segmentation performance is average, especially in dealing with gland adhesions and morphological variations (VTCD is 57.3160).

[0187] After adding MSNS in Experiment ②, DSC and IoU are respectively increased to 72.36% and 57.06%, the FPS drops to 209.2 img / s, while the FLOPs and the number of parameters increase to 3109.01M and 0.25M. This shows that MSNS helps to improve the network's segmentation ability for gland targets by suppressing background noise and enhancing the expression of multi-scale features, but there is an increase in computational overhead. In particular, the significant optimization of MSNS for VTCD (from 57.3160 to 40.0601) further proves its ability to capture gland morphological variations.

[0188] In Experiment ③, after adding CDSF, DSC and IoU were further improved to 73.32% and 58.29% respectively, while the FPS was 128.9 img / s, and the FLOPs and the number of parameters increased to 2507.57M and 0.19M. Through the deep feature alignment mechanism, CSFM optimized the accuracy of the segmentation region, making the network handle the gland boundaries more accurately. At the same time, VTCD decreased significantly to 38.9833. This indicates that CSFM has a good effect on capturing the structural features and morphological variations of glands.

[0189] In Experiment ④, after combining the MSNS and CSFM modules, DSC and IoU were improved to 73.66% and 58.68% respectively, and VTCD further decreased to 39.2364. Although the performance was improved compared with Experiment ③, the FPS further decreased to 98.5 img / s, and the FLOPs and the number of parameters were 4256.05M and 1.92M respectively. This indicates that the synergistic effect of MSNS and CSFM can further enhance the network's comprehensive extraction ability for multi-scale and deep features, but it brings a higher computational cost.

[0190] In Experiment ⑤, after adding a single prompt guidance module, DSC and IoU were significantly improved to 83.47% and 71.70% respectively, the FPS decreased to 56.2 img / s, and the FLOPs and the number of parameters increased to 5644.99M and 2.71M, guiding the network to capture gland features more accurately and significantly improving the segmentation performance. Especially in terms of VTCD, it significantly decreased to 32.8614, indicating that the prompt mechanism effectively alleviated the problem of gland morphological variation.

[0191] In Experiment ⑥, based on Experiment ⑤, the number of prompt guidance modules was increased to two. DSC and IoU were improved to 80.72% and 67.78% respectively, the FPS was 56.2 img / s, the FLOPs was 8796.67M, and the number of parameters remained 2.71M. Compared with the single prompt module, multiple prompt guidance optimized the segmentation effect and further reduced VTCD to 14.3297. And in Experiment ⑦ (the method of the present invention), the number of prompt guidance modules was increased to three. DSC and IoU were improved to 81.29% and 68.57% respectively, the FPS was 31.3 img / s, the FLOPs was 10748.35M, and the number of parameters remained 2.71M, and VTCD decreased to 7.2856, indicating that the model has more advantages in dealing with gland adhesion and complex morphology. Although the computational speed decreased, considering the significant performance improvement, this design achieved a good balance between efficiency and effect.

[0192] In Experiments ⑧ to ⑨, a multi - level prompt structure with different depths was used. It can be seen that as the depth of the prompt - guiding structure increases (4×, 5×), the segmentation performance tends to saturate. The DSC and IoU of 4× and 5× not only decrease (for example, the DSC in Experiment ⑨ is 71.32%), but also the computational cost increases significantly. The FLOPs reach 15851.71M, and the FPS is only 20.2 img / s. This indicates that the number and depth of the prompt modules need to be appropriately controlled. Excessive guidance will instead introduce interference information, reduce performance, and increase the computational burden.

[0193] From the above experimental results, it can be seen that the multi - scale noise suppression module (MSNS), the cross - dimensional shift feature module (CSFM), and the prompt - guiding structure have all significantly improved the performance of gland segmentation. The 3 - layer prompt - guiding structure proposed in the present invention has reached 81.29% and 68.57% respectively in terms of DSC and IoU indicators, and has achieved a minimum value of 7.2856 in the VTCD indicator, with an FPS of 31.3 img / s, taking into account both performance and efficiency, and demonstrating excellent adaptability to the meibomian gland segmentation task.

[0194] Appendix Figures 7 - 11 shows the segmentation visualization results of the TriPromptNet proposed in the present invention and some comparison networks on different datasets;

[0195] Among them, in Appendix Figure 8 for each group of images, starting from the first row, from left to right are the original image, the gold standard, TransUNet, UNet++, UNext, DeepLabV3+, LinkNet, U - Net, IF - Net, UNet3+, BiO - Net, SAM, MedSAM, MobileSAM, LiteMedSAM, and the TriPromptNet of the present invention.

[0196] Appendix Figure 9 starting from the first row, from left to right are the original image, the gold standard, TransUNet, UNet++, UNext, DeepLabV3+, LinkNet, U - Net, IF - Net, UNet 3+, BiO - Net, SAM, MedSAM, MobileSAM, LiteMedSAM, and the TriPromptNet of the present invention;

[0197] Appendix Figure 10For each group of images, starting from the first row and from left to right, they are the original image, the gold standard, TransUNet, UNet++, UNext, DeepLabV3+, LinkNet, U-Net, IF-Net, UNet 3+, BiO-Net, SAM, MedSAM, MobileSAM, LiteMedSAM, and the TriPromptNet of the present invention;

[0198] Appendix Figure 11 Starting from the first row and from left to right, they are the original image, the gold standard, TransUNet, UNet++, UNext, DeepLabV3+, LinkNet, U-Net, IF-Net, UNet 3+, BiO-Net, SAM, MedSAM, MobileSAM, LiteMedSAM, and the TriPromptNet of the present invention.

[0199] Through the comparison of the above results, the TriPromptNet proposed by the present invention performs best on the gland boundary, and it can be seen that the method proposed by the present invention can effectively alleviate the gland adhesion situation and improve the segmentation performance.

[0200] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0201] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0202] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the function.

[0203] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the function.

[0204] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to exhaustively list all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.

Claims

1. A meibomian gland segmentation method based on prompt enhancement, characterized in that Including the following steps: S1: Collect infrared images of the tarsal glands, and generate multiple positive prompt images according to the infrared images of the tarsal glands. Each positive prompt image respectively contains partial tarsal gland glands, and adjacent glands are not included in the same positive prompt image; S2: Extract prompt features for each positive prompt image, and at the same time extract global features from the infrared image of the tarsal gland to generate a global feature map; S3: According to the global feature map and the prompt features of each positive prompt image, perform layer-by-layer segmentation on the infrared image of the tarsal gland; during the layer-by-layer segmentation process, the prediction result of the previous layer is dynamically updated to generate a negative prompt image for the next layer of segmentation, which is used to mark the segmented gland area; The positive prompt image of the next layer is adjusted according to the unsegmented gland area.

2. The method for meibomian gland segmentation based on prompt enhancement according to claim 1, wherein In step S2, the method for extracting global features from the infrared image of the tarsal gland is as follows: extract the global features of the infrared image of the tarsal gland through an image encoder to generate a global feature map, and suppress noise through a multi-scale noise suppression module and enhance the local perception ability of the global feature map through a cross-dimensional shift feature module during the process of generating the global feature map.

3. The method for meibomian gland segmentation based on prompt enhancement according to claim 2, wherein, In step S2, the method for extracting prompt features from the positive prompt image is as follows: extract prompt features for each positive prompt image through a prompt encoder, and enhance the alignment ability between the prompt features and the image features through the cross-dimensional shift feature module during the process of extracting the prompt features.

4. The method for meibomian gland segmentation based on prompt enhancement according to claim 3, wherein In step S3, the method for layer-by-layer segmentation of the infrared image of the tarsal gland is as follows: perform layer-by-layer segmentation on the infrared image of the tarsal gland through a mask decoder according to the prompt features and the global feature map; and further suppress noise features through a multi-scale noise suppression module during the layer-by-layer segmentation process, and enhance cross-dimensional information interaction and local perception ability through a cross-dimensional shift feature module.

5. The method for meibomian gland segmentation based on prompt enhancement according to claim 4, wherein The multi-scale noise suppression module includes: A multi-scale feature extraction module that extracts multi-scale features from the input feature map using multiple convolutional kernels; A global context fusion module that extracts the global information of the input feature map through global average pooling operation; A residual shrinkage unit for generating a channel-level scaling factor and denoising the features through a soft thresholding mechanism; The working principle of the multi-scale noise suppression module MSNS is: F2 = Conv 1×1 (GAP(F in )) F2 = Conv 1×1 (GAP(F in )) F3 = ReLU(BN(Conv 3×3 (ReLU(BN(Conv 3×3 (F in )))))) F2 out = RSBU(F2) F1 = ReLU(BN(Conv 3×3 (F in ))) F merge = Conv 1×1 (F2 out + F3 out ) + F1 F out = ReLU(BN(Conv 3×3 (F merge ))) Among them, F in represents the input feature map, Conv 3×3 represents a 3×3 convolution, BN represents batch normalization, ReLU represents ReLU activation, Conv 1×1 represents a 1×1 convolution, GAP represents global average pooling, F1 and F3 represent features of different scales extracted, F2 represents the global information after global average pooling, F2 out and F3 out represent the output features after passing through the residual shrinkage unit RSBU-CS, F merge represents F1, F2 out and F3 out 's fused feature, F out represents the output feature of the multi-scale noise suppression module MSNS. Among them, the internal working principle of the residual shrinkage unit RSBU-CS is: F abs = |F i | θ = σ(FC(ReLU(BN(GAP(F abs ))))) F o = Sign(F i )·max(|F i | - θ, 0) Among them, F i represents the input feature, F abs represents the feature after taking the absolute value, FC represents the fully connected layer, θ represents the scaling factor, σ represents the Sigmoid activation function, Sign represents the Sign activation function, max(·) represents the maximum value operation, F o represents the output feature of the residual shrinkage unit.

6. The meibomian gland segmentation method based on prompt enhancement according to claim 5, characterized in that , where the cross-dimensional shift feature module includes: An input feature processing module that reshapes the input feature map into a fixed dimension; A shifted multi-layer perceptron module that performs spatial and channel dimension shift operations on the reshaped feature map; A feature fusion module that fuses the shifted features with the original features to generate an enhanced feature representation; The working principle of the cross-dimensional shift feature module is: P = Reshape(LN(Conv 3×3 (F in ))) M width = Shifted MLP width (P) M spatial = FC(ShiftedMLPH eight (GELU(DWConv(FC(M width ))))) M channel = FC(Shifted MLP channel (P)) F out = LN(M spatial + M channel ) + P Among them, F in represents the input feature map, Conv 3×3 represents a 3×3 convolution, LN represents layer normalization, Reshape represents the shape reshaping operation, P represents the feature map after shape reshaping, ShiftedMLP represents the shifted multi-layer perceptron module, DWConv represents the depthwise separable convolution, GELU represents the GELU activation, FC represents the fully connected layer, M spatial represents the obtained spatial local features, M channel represents the obtained channel features, F out represents the module output feature.

7. The method for meibomian gland segmentation based on prompt enhancement according to claim 6, wherein: In step S3, the method for layer-by-layer segmentation of the tarsal gland glands is as follows: First layer segmentation: Process the prompt image of the first layer using the prompt encoder of the first layer, and generate the prediction result of the tarsal gland glands in the first layer through the mask decoder of the first layer; Second layer segmentation: Use the hint encoder of the second layer to process the hint image of the second layer, combine the prediction results of the meibomian gland in the previous layer, and generate the prediction results of the meibomian gland in the current layer through the mask decoder of the second layer; Subsequent layer segmentation: Use the hint encoder of the current layer to process the hint image of the current layer, combine the prediction results of the meibomian gland in the previous several layers, and generate the prediction results of the meibomian gland in the current layer through the mask decoder of the current layer; Integrate the results: Integrate the prediction results of each layer through an exclusive OR operation to generate all the prediction results of the meibomian gland.

8. The method for meibomian gland segmentation based on prompt enhancement according to claim 7, wherein In the first layer segmentation, the negative hint image processed by the hint encoder is a black image with all gray values being 0.

9. A tarsal gland segmentation system based on prompt enhancement, characterized in that, It includes: An image acquisition module for acquiring infrared images of the meibomian gland; An image preprocessing module for preprocessing the acquired images; An image encoder for extracting the global features of the input image and generating a global feature map; A hint image generation module for generating multiple hint images according to the input image; There are three hint encoders for processing hint images and extracting hint features; There are three mask decoders, and each mask decoder is used in cooperation with a hint encoder to gradually complete the segmentation of the target gland by combining the global feature map and the hint features; A segmentation result output module for outputting the final segmentation result.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the hint-enhanced meibomian gland segmentation method according to any one of claims 1 to 8.