Eyelid morphological parameter measurement method based on reconstruction-segmentation network of conjoined architecture

Through the improved TransUNet model TB-Net and the convolutional convolution of dynamic parameter, the problem of manual measurement error and insufficient generalization of segmentation model in eyelid morphology parameter measurement is solved, and high-precision eyelid morphology parameter detection is achieved.

CN120298368AActive Publication Date: 2025-07-11SHENZHEN EYE HOSPITAL

Patent Information

Application Number
CN202510415207.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing method of measuring eyelid morphological parameters relies on manual measurement, is greatly affected by subjective factors, and the generalization ability of the image segmentation model is insufficient, resulting in insufficient segmentation accuracy.

Method used

The reconstruction-segmentation network STB-Net based on the concatenated architecture is adopted, combined with the improved TransUNet model TB-Net and dynamic parameter convolution, the feature fusion capability of the decoder is enhanced through the local attention modulation module BLAM, and combined with the reconstruction task and the segmentation task, an adaptive convolution kernel is generated for ocular image segmentation.

Benefits of technology

The accuracy and segmentation accuracy of eyelid morphological parameters detection are improved, especially in the segmentation of complex structures and edge details, showing excellent robustness and anti-interference ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298368A_ABST
    Figure CN120298368A_ABST
Patent Text Reader

Abstract

The invention discloses an eyelid morphological parameter measurement method based on a reconstruction-segmentation network of a conjoined architecture, and relates to the field of image processing, and the method comprises the steps: obtaining an ocular surface image, and respectively constructing an ocular surface image classification data set and an ocular surface image segmentation data set; constructing a reconstruction-segmentation network of a conjoined architecture based on dynamic parameter convolution; the reconstruction-segmentation network comprises a reconstruction task part and a segmentation task part; training the reconstruction task part by using the ocular surface image classification data set; embedding the trained encoder of the reconstruction task part into the segmentation task part as a dynamic convolution module, and training a reconstruction-segmentation network by using the ocular surface image segmentation data set to obtain a trained reconstruction-segmentation network; and inputting a to-be-detected ocular surface image into the trained reconstruction-segmentation network, segmenting a cornea region and a palpebral fissure region, and measuring the morphological parameters of the eyelid by using a measurement module. According to the invention, the accuracy of image segmentation and eyelid morphological parameter detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and more specifically, to a method for measuring eyelid morphological parameters based on a reconstruction-segmentation network with a siamese architecture. Background Art

[0002] Eyelid morphological parameters have extremely high clinical value in the diagnosis and evaluation of ocular surface diseases, such as ophthalmic disease diagnosis, evaluation of the effects before and after ophthalmic surgery, etc.

[0003] However, traditional eyelid morphology measurement relies on doctors to manually use measurement tools to achieve, which is affected by subjective factors and has low efficiency. With the rapid development of computer image processing technology and deep learning technology, more and more research has begun to explore automated measurement methods for eyelid morphological parameters. For example, ResNet50, improved ResNet18, etc. are used as backbones or encoders combined with networks such as U-Net, or models such as improved DeepLabV3 are used to measure various eyelid morphological parameters such as the height of the medial canthus, the height of the tear river, and the palpebral fissure length by segmenting different eye regions. However, most existing research on image processing and deep learning technology improves the segmentation accuracy by upgrading different parts of U-Net, but often ignores the effective construction of long-range dependence information, and the datasets used are usually limited to segmentation datasets, restricting the generalization ability of the segmentation model and resulting in insufficient accuracy of ocular surface image segmentation.

[0004] Therefore, how to improve the accuracy of image segmentation and thus improve the accuracy of eyelid morphological parameter detection is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a method for measuring eyelid morphological parameters based on a reconstruction-segmentation network with a siamese architecture, proposes an improved TransUNet model TB-Net, and implements a reconstruction-segmentation network STB-Net with a siamese architecture based on dynamic parameter convolution using the TB-Net model for image segmentation. Finally, the segmented corneal and palpebral fissure regions are input into the measurement module to automatically obtain the required eyelid morphological parameters.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] The present invention discloses a method for measuring eyelid morphological parameters based on a reconstruction-segmentation network with a siamese architecture, and the steps are as follows:

[0008] Obtain an ocular surface image, and respectively construct an ocular surface image classification dataset and an ocular surface image segmentation dataset;

[0009] Construct a reconstruction-segmentation network based on a dynamic parameter convolution-based siamese architecture; the reconstruction-segmentation network includes a reconstruction task part and a segmentation task part, and both the reconstruction task part and the segmentation task part are constructed based on the TB-Net;

[0010] Use the ocular surface image classification dataset to perform unsupervised training on the reconstruction task part; embed the encoder of the trained reconstruction task part as a dynamic convolution module into the segmentation task part, and use the ocular surface image segmentation dataset to perform supervised training on the reconstruction-segmentation network to obtain a trained reconstruction-segmentation network;

[0011] Input the ocular surface image to be detected into the trained reconstruction-segmentation network, segment the corneal region and the palpebral fissure region, and use the measurement module to measure the eyelid morphological parameters.

[0012] Furthermore, the TB-Net includes an encoder and a decoder part; the encoder includes a first bottleneck module, a second bottleneck module, a third bottleneck module, and 12 Transformer Layer layers connected in sequence;

[0013] Extract features from the input image or feature map through a three-stage bottleneck module to obtain CNN feature maps with different resolutions;

[0014] Convert the CNN feature map output by the third bottleneck module into the feature format of the Transformer through a linear mapping, and encode the features through the 12 Transformer Layer layers to obtain the output feature map of the encoder.

[0015] Furthermore, the decoder includes a first deconvolution layer, a second deconvolution layer, a third deconvolution layer, and an output layer connected in sequence;

[0016] The output feature map of the encoder is converted in feature format as the input feature of the decoder and input into the first deconvolution layer;

[0017] The first deconvolution layer, the second deconvolution layer, and the third deconvolution layer respectively correspond to the third bottleneck module, the second bottleneck module, and the first bottleneck module; the deconvolution layer first performs an upsampling operation on the input feature map, then fuses the upsampled feature map with the CNN feature map output by the bottleneck module, then performs a convolution operation on the fused feature map, and finally outputs through the non-linear activation operation of the ReLu activation function;

[0018] The output layer performs a convolution operation and a non - linear activation operation on the feature map output by the third de - convolution layer and then outputs.

[0019] Furthermore, the decoder further includes a BLAM module, and the BLAM module includes a first BLAM layer, a second BLAM layer, a third BLAM layer, and a fourth BLAM layer connected in sequence; the first BLAM layer, the second BLAM layer, the third BLAM layer, and the fourth BLAM layer respectively correspond one - to - one with the first de - convolution layer, the second de - convolution layer, the third de - convolution layer, and the output layer;

[0020] The de - convolution layer uses a local channel attention mechanism to aggregate the input small - scale feature map and large - scale feature map and outputs a cross - layer fusion feature map; the small - scale feature map is the input feature map of the decoder or the feature map output by the previous BLAM layer, and the large - scale feature map is the feature map output by the de - convolution layer corresponding to the BLAM layer;

[0021] The outputs of all BLAM layers are concatenated along the channel dimension and then the final result is output.

[0022] Furthermore, the formula for aggregation in the BLAM layer is:

[0023] L(M)=σ(δ(B(PWConv2(δ(B(PWConv1(M)))))));

[0024]

[0025] Among them, L(M) represents the attention weight map; PWConv1 and PWConv2 respectively represent the first point - wise convolution and the second point - wise convolution; σ and δ respectively represent the Sigmoid function and the ReLU activation function, B represents the normalization operation; M, N, and Z respectively represent the low - level features of the large - scale feature map, the high - level features of the small - scale feature map, and the cross - layer fusion features; represents element - wise multiplication.

[0026] Furthermore, in the dynamic convolution module, the dynamic parameter convolution process is as follows:

[0027] Perform convolution, batch normalization, and ReLU non - linear combination operations on the feature map output by the encoder of the reconstruction task part, and then perform parameter deformation to determine the parameters of the dynamic convolution kernel;

[0028] Perform convolution, batch normalization, and ReLU non - linear combination operations on the feature map output by the encoder of the segmentation task part and use it as the convolution object;

[0029] Perform a convolution operation on the convolution object using the dynamic convolution kernel, and add the obtained feature map to the feature map output by the encoder of the segmentation task part as the final output feature map of the encoder of the segmentation task part.

[0030] Furthermore, the formula for the dynamic parameter convolution process is:

[0031] DPConv(X,Y) = Conv2d(γX,ψ(γ(Y)));

[0032] Where γ represents the combination of convolution, batch normalization, and ReLU non-linearity, X and Y respectively represent the feature maps output by the encoders of the segmentation task and the reconstruction task, ψ represents parameter deformation, and Conv2d represents two-dimensional convolution.

[0033] Furthermore, to determine the parameters of the dynamic convolution kernel by performing parameter deformation, specifically: perform parameter deformation ψ on the feature map output by the encoder of the reconstruction task to obtain a one-dimensional vector θ of size n n ; determine the expected number of channels of the target channel of the dynamic convolution kernel, the size of the convolution kernel, and the number of output channels respectively according to the n parameters of θ n For the semantic transformation, first use a 1×1 convolution

[0034] and adaptive average pooling P to adjust the channel dimension, and then use another 1×1 convolution to reshape the feature into θ and change the channel from z to n. The specific formula is: n

[0035]

[0036] Furthermore, the eyelid shape parameters are divided into three categories in total, including palpebral fissure height, palpebral fissure width, and palpebral fissure area. The palpebral fissure height includes the left palpebral fissure height, the central palpebral fissure height, and the right palpebral fissure height.

[0037] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method for measuring eyelid morphological parameters based on a reconstruction-segmentation network with a conjoined architecture, proposes an improved TransUNet model TB-Net, and constructs a reconstruction-segmentation network STB-Net with a conjoined architecture based on dynamic parameter convolution based on TB-Net for accurately segmenting the corneal and palpebral fissure regions in ocular surface images. Finally, the corresponding left palpebral fissure height, central palpebral fissure height, right palpebral fissure height, palpebral fissure width, and palpebral fissure area are measured according to the obtained segmentation results. The TB-Net model effectively solves the problem of missing fine and edge local information in ocular surface image segmentation and improves the segmentation accuracy by improving the decoder of TransUNet and introducing a bottom-up local attention modulation module BLAM; the STB-Net model adopts the SRSNetwork architecture, combines the reconstruction task and the segmentation task, and generates an adaptive convolution kernel using dynamic parameter convolution, thereby optimizing the segmentation effect on limited labeled data. The present invention segments ocular surface images based on a reconstruction-segmentation network with a conjoined architecture based on dynamic parameter convolution, improves the accuracy of image segmentation, and further improves the accuracy of eyelid morphological parameter detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0039] Figure 1 It is a schematic diagram of the overall process of the embodiment of the present invention.

[0040] Figure 2 It is a schematic diagram of the BLAM structure of the embodiment of the present invention.

[0041] Figure 3 It is a schematic diagram of the TB-Net structure of the embodiment of the present invention.

[0042] Figure 4 It is a schematic diagram of the principle of the dynamic convolution module of the embodiment of the present invention.

[0043] Figure 5(a) is a schematic diagram of the original image of the embodiment of the present invention.

[0044] Figure 5(b) is a schematic diagram of the palpebral fissure label of the embodiment of the present invention.

[0045] Figure 5(c) is a schematic diagram of the corneal label of the embodiment of the present invention.

[0046] Figure 6This is the test effect diagram of the reconstruction task of the embodiment of the present invention.

[0047] Figure 7 This is the comparison diagram of the corneal segmentation results of each model in the embodiment of the present invention.

[0048] Figure 8 This is the comparison diagram of the palpebral fissure segmentation results of each model in the embodiment of the present invention. Detailed implementation manners

[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] The embodiment of the present invention discloses a method for measuring eyelid morphological parameters based on a reconstruction-segmentation network with a siamese architecture, as Figure 1 shown, the steps are as follows:

[0051] Obtain an ocular surface image, and respectively construct an ocular surface image classification dataset and an ocular surface image segmentation dataset;

[0052] Construct a reconstruction-segmentation network with a siamese architecture based on dynamic parameter convolution; the reconstruction-segmentation network includes a reconstruction task part and a segmentation task part, and both the reconstruction task part and the segmentation task part are constructed based on TB-Net;

[0053] Use the ocular surface image classification dataset to perform unsupervised training and learning on the reconstruction task part; embed the encoder of the trained reconstruction task part as a dynamic convolution module into the segmentation task part, and use the ocular surface image segmentation dataset to perform supervised training and learning on the reconstruction-segmentation network to obtain a trained reconstruction-segmentation network;

[0054] Input the ocular surface image to be detected into the trained reconstruction-segmentation network, segment the corneal region and the palpebral fissure region, and use the measurement module to measure the eyelid morphological parameters.

[0055] In a specific embodiment, as Figure 3 shown, TB-Net includes an encoder and a decoder part; the encoder includes a first bottleneck module, a second bottleneck module, a third bottleneck module, and 12 Transformer Layer layers connected in sequence;

[0056] Perform feature extraction on the input image or feature map through a three-stage bottleneck module to obtain CNN feature maps with different resolutions;

[0057] The CNN feature map output by the third bottleneck module is converted into the feature format of the Transformer through a linear mapping, and the features are encoded through 12 Transformer Layer layers to obtain the output feature map of the encoder.

[0058] In a specific embodiment, the decoder includes a first deconvolution layer, a second deconvolution layer, a third deconvolution layer, and an output layer connected in sequence;

[0059] The output feature map of the encoder is converted in feature format and used as the input feature of the decoder, and is input into the first deconvolution layer;

[0060] The first deconvolution layer, the second deconvolution layer, and the third deconvolution layer correspond to the third bottleneck module, the second bottleneck module, and the first bottleneck module respectively; for the deconvolution layer, first perform an upsampling operation on the input feature map, then perform feature fusion on the upsampled feature map and the CNN feature map output by the bottleneck module, then perform a convolution operation on the fused feature map, and finally output through the non-linear activation operation of the ReLu activation function;

[0061] The output layer performs a convolution operation and a non-linear activation operation on the feature map output by the third deconvolution layer and then outputs.

[0062] Specifically, in the field of medical image segmentation, the U-Net network is a classic and widely used deep learning model. The U-Net network is a typical encoding-decoding structure, forming a U-shaped form. With the in-depth study of subsequent research, the U-Net network has gone through multiple development stages, forming different variants and derivative versions. Among them, U-Net++ is a typical improvement. It increases the feature expression and detail recovery ability of the model by introducing a multi-layer network structure in the skip connection, and further improves the segmentation accuracy. In recent years, variants of U-Net have also incorporated Transformer. By introducing the self-attention mechanism at the encoder level, it improves the shortcoming of the traditional convolutional neural network in constructing long-distance dependent information. The most representative one is TransUNet. However, in the feature transmission of the decoder itself in TransUNet, the fusion of high-level semantic information and low-level semantic information is not considered, resulting in local information limitations; and, in the feature extraction process, most convolutional networks learn high-level semantic features by gradually attenuating the size of the feature map, resulting in local information being easily submerged by the surrounding background features in the deep layer, which is not conducive to the local reproduction of the edges of the segmentation target.

[0063] Therefore, in the present invention, a Bottom-Up Local Attentional Modulation (BLAM) module is utilized to integrate the small-scale subtle features of low-level features into higher-level features at deeper levels. The low-level features are upsampled and then interact with the high-level features inside the BLAM to enhance the local information attention of the low-level features to the high-level features through dynamic weighting. Through the progressive enhancement of multiple BLAM modules layer by layer, TB-Net achieves efficient feature progressive optimization in the decoder, enabling features at different levels to have stronger pertinence and local attention ability during the fusion process. Each BLAM module dynamically adjusts the response of high-level semantic features to low-level detailed information, gradually reducing the over-reliance of the deep network on background information, thereby highlighting the saliency of the target region. This progressive enhancement mechanism ensures that the segmentation network can effectively capture the local features of the target while retaining spatial details, and improve the robustness and accuracy of segmentation in complex scenarios. Applying BLAM to the decoder module in the segmentation part can effectively enhance the decoder's attention to small and edge local information, which is helpful for more accurately restoring the boundaries of complex structures in medical image segmentation and improving the segmentation accuracy of small lesions and blurred boundaries.

[0064] In a specific embodiment, the decoder further includes a BLAM module, and the BLAM module includes a first BLAM layer, a second BLAM layer, a third BLAM layer, and a fourth BLAM layer connected in sequence; the first BLAM layer, the second BLAM layer, the third BLAM layer, and the fourth BLAM layer respectively correspond to the first transposed convolution layer, the second transposed convolution layer, the third transposed convolution layer, and the output layer one by one;

[0065] The transposed convolution layer utilizes a local channel attention mechanism to aggregate the input small-scale feature map and large-scale feature map and output a cross-layer fusion feature map; the small-scale feature map is the input feature map of the decoder or the feature map output by the previous BLAM layer, and the large-scale feature map is the feature map output by the transposed convolution layer corresponding to the BLAM layer;

[0066] After the outputs of all BLAM layers are concatenated by channels, the final result is output.

[0067] In a specific embodiment, the aggregation formula in the BLAM layer is:

[0068] L(M) = σ(δ(B(PWConv2(δ(B(PWConv1(M)))))));

[0069]

[0070] Among them, L(M) represents the attention weight map; PWConv1 and PWConv2 represent the first pointwise convolution and the second pointwise convolution respectively; σ and δ represent the Sigmoid function and the ReLU activation function respectively, and B represents the normalization operation; M, N, and Z respectively represent the low-level features of the large-scale feature map, the high-level features of the small-scale feature map, and the cross-layer fusion features. represents element-wise multiplication.

[0071] Specifically, as Figure 2 shown (in the figure, Point-wise Conv represents pointwise convolution), the local channel attention mechanism L aggregates the channel context features at each spatial position. The kernel sizes of PWConv1 and PWConv2 are C / 4×C×1×1 and C×C / 4×1×1 respectively, similar to the bottleneck structure. The attention weight map L(X) ∈ R C×H×W has the same shape as the input feature map, so it can highlight subtle details in an element-wise manner (spatially and cross-channel). The motivation of the BLAM module is to embed small-scale details into the high-level coarse feature map, which is achieved by dynamically weighting and modulating the high-level features under the guidance of the low-level features. Given X as the low-level feature and Y as the high-level feature, the cross-layer fusion feature Z ∈ R C×H×W , R C×H×W represents a three-dimensional real tensor, where C, H, and W represent the number of channels, height, and width respectively.

[0072] In a specific embodiment, a dynamic parameter convolution process is performed within the dynamic convolution module. Specifically:

[0073] Perform convolution, batch normalization, and ReLU non-linear combination operations on the feature map output by the encoder of the reconstruction task part, and then perform parameter deformation to determine the parameters of the dynamic convolution kernel;

[0074] Perform convolution, batch normalization, and ReLU non-linear combination operations on the feature map output by the encoder of the segmentation task part as the convolution object;

[0075] Use the dynamic convolution kernel to perform convolution operations on the convolution object, and the resulting feature map is added to the feature map output by the encoder of the segmentation task part as the final output feature map of the encoder of the segmentation task part.

[0076] Specifically, the present invention uses TB-Net for reconstruction tasks and segmentation tasks respectively, and the new model formed is named STB-Net. STB-Net uses unsupervised reconstruction tasks to learn high-level semantic information in the same type of medical image dataset without segmentation labels, providing a more reliable dynamic parameter convolution kernel for the segmentation task, thereby improving the performance and accuracy of the segmentation model. The reconstruction task and the segmentation task are associated to form a symmetric Siamese architecture, and the connection between the reconstruction task and the segmentation task is realized through the dynamic convolution module DPConv. The working principle of DPConv is as Figure 4 shown, where E r 、E s 、D s represent the encoder of the reconstruction model, the encoder of the segmentation model, and the decoder of the segmentation model respectively. The semantic features output by the upper half are transformed into convolution kernels through parameter deformation, and the feature maps generated by the lower half are used as convolution objects. The two perform convolution operations, and the output feature maps are added to the original feature maps and passed to the decoder, which is also the core of the entire network's operation. It should be noted that DPConv is only used at the end of the encoder to additionally generate new feature outputs and add them to the output of the original encoder, without changing the number of channels during the feature transmission process. This method similar to the residual connection can ensure that the features originally transmitted by the segmentation model remain valid and are further strengthened after the integration of new features.

[0077] In a specific embodiment, the formula for the dynamic parameter convolution process is:

[0078] DPConv(X,Y)=Conv2d(γX,ψ(γ(Y)));

[0079] where γ represents the combination of convolution, batch normalization, and ReLU non-linearity, X and Y represent the feature maps output by the encoders of the segmentation task and the reconstruction task respectively, ψ represents parameter deformation, and Conv2d represents two-dimensional convolution.

[0080] In a specific embodiment, parameter deformation is performed to determine the parameters of the dynamic convolution kernel. Specifically: according to the feature map output by the encoder of the reconstruction task, parameter deformation ψ is performed to obtain a one-dimensional vector θ n with a size of n; according to the n parameters of θ n , the expected number of channels of the target channel of the dynamic convolution kernel, the size of the convolution kernel, and the number of output channels are determined respectively;

[0081] Semantic transformation, first use a 1×1 convolution and adaptive average pooling P to adjust the channel dimension, and then use another 1×1 convolution to reshape the feature into θ n , and change the channel from z to n. The specific formula is:

[0082]

[0083] Specifically, to perform the DPConv operation, the convolutional kernel must have specific parameters: the expected number of target channels is denoted as M, the kernel size is denoted as K, and the output channels are denoted as N. The total number of parameters to be generated can be expressed as n = M × N × K 2 ; therefore, to generate n parameters, a one-dimensional vector θ of size n is selected for construction n ; to obtain θ n , the semantic feature Y ∈ R z =C×H×W needs to be converted into a one-dimensional vector θ n . However, if the average pooling operation is directly used to convert Y into θ n , it will cause the network not to converge.

[0084] In a specific embodiment, the eyelid shape parameters are divided into three categories in total, including the palpebral fissure height, palpebral fissure width, and palpebral fissure area. The palpebral fissure height includes the left palpebral fissure height, central palpebral fissure height, and right palpebral fissure height.

[0085] Specifically, the last step of the present invention is to measure three types of eyelid shape parameters based on the corneal segmentation result and palpebral fissure segmentation result obtained by the foregoing segmentation algorithm. In this embodiment, the resolution of the ocular surface image is 2974×1984. After professional ophthalmologists measure the height and width of all ocular surface images in the experiment, the average width is 14.65 cm and the average height is 9.77 cm. The pictures obtained by the instrument for taking ocular surface images are 4 times the actual size, so the actual average width is 3.6625 cm and the actual average height is 2.4425 cm. Then, on the original image, the width of each pixel is 0.012306 mm and the height is 0.012310 mm. Due to the Transformer layer, there are restricted conditions on the size of the input image. An overly large input size will cause insufficient video memory during training, and an overly small input size will result in a lower accuracy of the segmentation result. Therefore, the input image size is reduced to 384×256, and then the width of each pixel is 0.09537 mm and the height is 0.09541 mm. For each ocular surface image in the test set, record the three abscissas of the left boundary, right boundary, and center point of the corneal area in the corneal segmentation map. And find the number of target pixels in the columns corresponding to the three abscissas in the palpebral fissure segmentation map, and multiply by the corresponding pixel height to obtain the left palpebral fissure height, central palpebral fissure height, and right palpebral fissure height. For the palpebral fissure width, the corresponding palpebral fissure width can be obtained by calculating the difference between the leftmost and rightmost abscissas in the palpebral fissure segmentation map and then multiplying by the pixel width. For the palpebral fissure area, the corresponding palpebral fissure area can be directly obtained by calculating the number of all target pixels and then multiplying by the area of a single pixel.

[0086] In a specific embodiment, a specific example is used to illustrate the process of image dataset acquisition and processing, model training, evaluation, and calculation result evaluation and analysis. Specifically:

[0087] Since the algorithm of the present invention includes a reconstruction task and a segmentation task, the datasets involved include two types, namely the ocular surface image classification dataset and the ocular surface image segmentation dataset. These two types of images are captured using different instruments, so their actual sizes and image resolutions are different. However, the ocular surface image classification dataset also belongs to the ocular surface images and has extremely high reuse value. Therefore, it can be used together with the ocular surface image segmentation dataset of the present invention in the reconstruction task. And, STB-Net includes two parts: reconstruction and segmentation. Among them, reconstruction is unsupervised learning, and the non-matching data - classification dataset can be used for unsupervised training of the network to extract the features of the same type of data (ocular surface image data). The encoder obtained by the reconstruction task generates adaptive dynamic convolution parameters for the network of the segmentation task, which can indirectly make up for the accuracy defect caused by the insufficient segmentation dataset. For the same type of dataset here, the existing ocular surface classification dataset is used in the experimental process, but their classification labels are not used.

[0088] First, the ocular surface image segmentation dataset is provided by the Shenzhen Eye Hospital, with a total of 250 ocular surface images containing corneal and palpebral fissure labels, and the resolution is 2976×1984. All images maintain the same shooting conditions. Secondly, the ocular surface image classification dataset is provided by the Ocular Surface Disease Center of the Affiliated Eye Hospital of Nanjing Medical University, with a total of 2855 ocular surface disease images, and the resolution is 2976×1984, including three categories of images: normal ocular surface, ocular surface hemorrhage, and pterygium. All images have been desensitized and do not contain the privacy information of patients.

[0089] In the reconstruction task, the ocular surface image classification dataset and the measurement dataset are used together to implement the reconstruction task, with a total of 3105 ocular surface images. Since the reconstruction task is used to assist the segmentation task, it does not need to be tested and verified. Then, the dataset is divided into a training set and a validation set according to a ratio of 9:1, resulting in 2795 training set images and 310 validation set images. In the segmentation task, the ocular surface image segmentation dataset is divided into a training set, a validation set, and a test set according to a ratio of 7:1:2. The division process also follows the principle of stratified sampling, resulting in 174 training set images, 25 validation set images, and 51 test set images.

[0090] Secondly, since the width and height of the images are not the same, and the optimal input image size of the algorithm of the present invention is 384×384. In order to meet the input requirements of the network, directly stretching the height and width of the image to the same scale does not conform to the original proportion of the image, which easily leads to distortion of the target shape, and further affects the network's understanding and segmentation accuracy of the image. Therefore, the method of scaling by the same ratio is adopted, the image is scaled to the required scale according to the long side, and zeros are padded on both sides of the short side to the same size. In this way, the original proportion of the image can be retained, and the information loss caused by deformation can be avoided. The data augmentation methods used include random size stretching, random horizontal flipping, and random vertical flipping. The segmentation labels of the present invention are all completed under the guidance of doctors. The labeling software used is LabelMe, and the preprocessing method of the labels is the same as that of the original images. The ocular surface images and the corresponding labels are shown in Figures 5(a), 5(b), and 5(c).

[0091] In the specific training process, the reconstruction task trains the TB-Net. Since it is necessary to simultaneously restore the three RGB channels of the image and further evaluate the difference between the model prediction result and the true value, the Mean Squared Error (MSE) is used as the loss function. The specific formula is as follows:

[0092]

[0093] where N is the total number of pixels in the three channels, y i is the true value of the i-th sample, is the predicted value of the i-th sample.

[0094] In addition, Stochastic Gradient Descent is used as the optimizer, the momentum is set to 0.9, the weight decay is set to 0.0001, the learning rate is set to 0.001, the training batch size is set to 4, and the number of iterations is 200 times.

[0095] For the segmentation task, the STB-Net is trained, and the weighted sum of BCE Loss and Dice Loss is used as the loss function. The calculation process is as follows:

[0096]

[0097] BCEDiceLoss = αDiceLoss + βBCELoss;

[0098] where N represents the total number of pixels, y i is the true value of the i-th sample, is the predicted value of the i-th sample, and α and β represent the weights of the two loss functions, both of which are set to 0.5.

[0099] In addition, Adam is used as the optimizer, the weight decay is set to 0.0003, the learning rate is set to 0.0001, the training batch size is set to 8, the number of iterations is 300, and the cosine annealing learning rate adjustment strategy is adopted. This invention is completed under the PyTorch deep learning framework. Hardware configuration: Intel i7-14700KF 5.6GHZ, NVIDIA RTX4090 24GB. Software configuration: Windows11, python3.8, PyTorch2.3.1, OpenCV4.5.5.

[0100] To measure the segmentation performance of the model, the Dice coefficient, Intersection over Union (IoU), and Global Accuracy (GA) are used as the evaluation metrics of the model. The calculation formulas are shown as follows:

[0101]

[0102] In the formula, y represents the label pixel matrix, represents the predicted pixel matrix, TP represents the number of pixels correctly predicted as the target, TN represents the number of pixels correctly predicted as the background, FP represents the number of pixels wrongly predicted as the target, and FN represents the number of pixels wrongly predicted as the background; among them, the Dice coefficient is used to describe the similarity between the prediction result and the real result, the higher the similarity, the closer the value is to 1; IoU is used to measure the overlapping degree between the prediction result and the real result, and similarly, the closer the ratio is to 1, the higher the overlapping degree; GA represents the proportion of the pixel values correctly predicted by the model among all pixel points, and is an important indicator to measure the model's ability to distinguish the target and the background.

[0103] This invention also adopts two methods, linear regression analysis and Bland-Altman agreement analysis, to quantitatively evaluate the measurement results. In linear regression analysis, by fitting the regression line and calculating the correlation coefficient r 2 value, the relationship and fitting degree between the measured value and the actual value are quantified. The closer the r 2 value is to 1, the higher the degree of agreement between the system prediction value and the real value, and the better the measurement accuracy. In addition, Bland-Altman agreement analysis is used to further evaluate the agreement between the system and the doctor's measurement results, analyze the error distribution and Limits of Agreement (LoA) between the two, and intuitively reveal the deviation and reliability between the two. These two analysis methods complement each other. Linear regression focuses on the correlation and accuracy of the results, while agreement analysis focuses on the stability and consistency of the measurement results in practical applications, ensuring the reliability and effectiveness of the system in clinical applications.

[0104] The TB-Net constructed in the present invention is used for the reconstruction task. The average mean square error of the optimal reconstruction model obtained through training on the validation set is 0.00311. Some test effect diagrams are as shown in Figure 6 the following. It can be seen through observation that the reconstruction task generally shows good image restoration effects. Although there are spots in some local parts of some images and holes in some reflective areas, the overall reconstruction of the entire ocular surface area is relatively accurate. Especially in the restoration of the texture in the corneal area and the blood vessels and color in the palpebral fissure area, it shows high quality. This indicates that TB-Net has learned the ocular surface image features sufficiently in the reconstruction task, laying a good foundation for subsequent analysis and applications.

[0105] In the segmentation task of the present invention, in addition to the comparative analysis of TransUNet, TB-Net and STB-Net, two classic segmentation models, U-Net and U-Net++, are also selected for comparison. All models are tested on the same dataset and trained multiple times to reach the optimal state to measure the effectiveness of the method of the present invention. As shown in Table 1, for the segmentation of the palpebral fissure area, the proposed model STB-Net is superior to other models in terms of indicators, where Dice, GA, and IoU are 0.9875, 0.9955, and 0.9767 respectively. The three indicators of TB-Net have all improved compared with TransUNet, among which Dice has increased by 0.08%, GA has increased by 0.03%, and IoU has increased by 0.15%, indicating that it is effective to introduce BLAM into the decoder of TransUNet. STB-Net has increased Dice by 0.06%, GA by 0.02%, and IoU by 0.12% on the basis of TB-Net, which also indicates that it is effective to adopt the siamese architecture based on TB-Net.

[0106] Table 1 Evaluation indicators for palpebral fissure segmentation

[0107]

[0108]

[0109] Table 2 Evaluation indicators for corneal segmentation

[0110]

[0111] Similarly, as shown in Table 2, for the segmentation of the corneal region, the proposed model STB-Net also achieved the highest in all three metrics. Among them, GA tied for the highest with U-Net++ and TB-Net at 0.9978, Dice was 0.9891, and IoU was 0.9790. TB-Net had improvements in Dice, GA, and IoU compared to TransUNet. STB-Net had further improvements in Dice and IoU compared to TB-Net, except that GA remained unchanged. It can be seen that STB-Net can effectively segment the palpebral fissure region and the corneal region in ocular surface images.

[0112] In addition to measuring the segmentation performance of the model in ocular surface images through evaluation metrics, the present invention also intuitively demonstrates the advantages of the proposed model by showing the segmentation results of each model on ocular surface images. As Figure 8 shown, it can be observed that STB-Net can segment the palpebral fissure region more accurately. First, in terms of the segmentation details of the palpebral fissure angle, different degrees of missing or breakage occur in other types of models. For example, in the first group of segmentation results, STB-Net maximally restored the details at the palpebral fissure angle, while other models either did not consider it as part of the palpebral fissure or showed incomplete phenomena. Second, in terms of the segmentation details of polyps on the upper and lower eyelids, other types of models misjudged the polyps. For example, in the third group of segmentation results, STB-Net could distinguish the polyps and regarded them as the background, while other models regarded the polyps as part of the palpebral fissure. Similarly, as Figure 7 shown, STB-Net was also more accurate in corneal segmentation. For example, in the first group of segmentation results, STB-Net well segmented the cornea, while other models had holes and missing parts due to the influence of polyps. Another example is the second group of segmentation results. STB-Net was not affected by the eyelashes, while other models were affected and regarded the area that did not belong to the cornea as the target.

[0113] In summary, through the comparison of quantitative evaluation metrics and intuitive segmentation results, the present invention fully verified the advantages of STB-Net in the ocular surface image segmentation task. STB-Net not only had better segmentation accuracy of the palpebral fissure and cornea than other models, but also demonstrated excellent detail capture ability and anti-interference ability. Specifically, in the segmentation of complex structures (such as the palpebral fissure angle, polyps on the upper and lower eyelids, and eyelash regions), STB-Net could accurately restore the details, avoid common detail loss or misjudgment, effectively resist background interference, and ensure the integrity and accuracy of the segmentation results.

[0114] Furthermore, the segmentation results of STB-Net were measured and the relative errors were calculated with the true measurement values. Among them, the relative error of the left palpebral fissure height was 2%, the relative error of the central palpebral fissure height was 0.78%, the relative error of the right palpebral fissure height was 1.93%, the relative error of the palpebral fissure width was 1.31%, and the relative error of the palpebral fissure area was 0.98%. The measurement values of some samples are shown in Tables 3 and 4. In the tables, LH, MH, RH, PW, and PA represent the left palpebral fissure height, central palpebral fissure height, right palpebral fissure height, palpebral fissure width, and palpebral fissure area respectively. This further verifies that STB-Net can accurately segment the key areas in most samples and achieve precise measurement.

[0115] Table 3 True measurement values of some samples

[0116]

[0117] Table 4 Segmentation measurement values of STB-Net for some samples

[0118]

[0119]

[0120] The present invention adopts the Bland-Altman analysis method to evaluate the consistency between the measured values and the true values of each eyelid morphological parameter. In the analysis, the average value of the two groups of data was used as the horizontal axis and the difference was used as the vertical axis to generate a Bland-Altman test chart. Statistically, if the error is within the 95% consistency interval, the error is considered to be within the acceptable range. The results show that there are 3 sample points of the left palpebral fissure height exceeding the consistency interval, the same is true for the central palpebral fissure height with 3, 1 for the right palpebral fissure height, 4 for the palpebral fissure width, and 3 for the palpebral fissure area. This indicates that the method for measuring eyelid morphological parameters proposed in this study has a good agreement with the true values, demonstrating its potential value in practical applications.

[0121] In summary, the present invention proposes an improved TransUNet model, TB-Net, and constructs a reconstruction-segmentation network, STB-Net, based on a Siamese architecture with dynamic parameter convolution for accurately segmenting the corneal and palpebral fissure regions in ocular surface images and measuring the corresponding left palpebral fissure height, central palpebral fissure height, right palpebral fissure height, palpebral fissure width, and palpebral fissure area according to the obtained segmentation results. The TB-Net model effectively solves the problem of missing fine and edge local information in ocular surface image segmentation and improves the segmentation accuracy by improving the decoder of TransUNet and introducing a bottom-up local attention modulation module, BLAM. The STB-Net model adopts the SRSNetwork architecture, combines the reconstruction task with the segmentation task, and uses dynamic parameter convolution to generate adaptive convolution kernels, thereby optimizing the segmentation effect on limited labeled data.

[0122] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0123] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for measuring eyelid morphological parameters based on a reconstruction-segmentation network with a siamese architecture, characterized in that, The steps are as follows: Obtain ocular surface images, and respectively construct an ocular surface image classification dataset and an ocular surface image segmentation dataset; Construct a reconstruction-segmentation network based on a siamese architecture with dynamic parameter convolution; the reconstruction-segmentation network includes a reconstruction task part and a segmentation task part, and both the reconstruction task part and the segmentation task part are constructed based on TB-Net; Use the ocular surface image classification dataset to perform unsupervised training and learning on the reconstruction task part; Embed the encoder of the trained reconstruction task part into the segmentation task part as a dynamic convolution module, and use the ocular surface image segmentation dataset to perform supervised training and learning on the reconstruction-segmentation network to obtain a trained reconstruction-segmentation network; Input the ocular surface image to be detected into the trained reconstruction-segmentation network, segment the corneal region and the palpebral fissure region, and use the measurement module to measure the eyelid morphological parameters.

2. The method for measuring eyelid morphological parameters of a reconstruction-segmentation network based on a conjoined architecture according to claim 1, wherein The TB-Net includes an encoder and a decoder part; the encoder includes a first bottleneck module, a second bottleneck module, a third bottleneck module, and 12 Transformer Layer layers connected in sequence; Extract features from the input image or feature map through a three-stage bottleneck module to obtain CNN feature maps with different resolutions; Convert the CNN feature map output by the third bottleneck module into the feature format of Transformer through a linear mapping, and encode the features through the 12 Transformer Layer layers to obtain the output feature map of the encoder.

3. The eyelid morphology parameter measurement method based on the reconstruction-segmentation network of the Siamese architecture according to claim 2, characterized in that: The decoder includes a first deconvolution layer, a second deconvolution layer, a third deconvolution layer, and an output layer connected in sequence; The output feature map of the encoder is used as the input feature of the decoder after feature format conversion and input into the first deconvolution layer; The first deconvolution layer, the second deconvolution layer, and the third deconvolution layer respectively correspond to the third bottleneck module, the second bottleneck module, and the first bottleneck module; the deconvolution layer first performs an upsampling operation on the input feature map, then fuses the upsampled feature map with the CNN feature map output by the bottleneck module, then performs a convolution operation on the fused feature map, and finally outputs through the non-linear activation operation of the ReLu activation function; The output layer performs a convolution operation and a non-linear activation operation on the feature map output by the third deconvolution layer and then outputs.

4. The eyelid morphology parameter measurement method based on the reconstruction-segmentation network of the Siamese architecture according to claim 3, characterized in that: The decoder also includes a BLAM module, and the BLAM module includes a first BLAM layer, a second BLAM layer, a third BLAM layer, and a fourth BLAM layer connected in sequence; the first BLAM layer, the second BLAM layer, the third BLAM layer, and the fourth BLAM layer respectively correspond to the first deconvolution layer, the second deconvolution layer, the third deconvolution layer, and the output layer; The deconvolution layer utilizes a local channel attention mechanism to aggregate the input small-scale feature map and large-scale feature map and output a cross-layer fusion feature map; the small-scale feature map is the input feature map of the decoder or the feature map output by the previous BLAM layer, and the large-scale feature map is the feature map output by the deconvolution layer corresponding to the BLAM layer; After the outputs of all BLAM layers are concatenated along the channel dimension, the final result is output.

5. A method for measuring eyelid morphological parameters of a reconstruction-segmentation network based on a conjoined architecture according to claim 4, wherein, The aggregation formula in the BLAM layer is: L(M) = σ(δ(B(PWConv2(δ(B(PWConv1(M))))))); Among them, L(M) represents the attention weight map; PWConv1 and PWConv2 respectively represent the first pointwise convolution and the second pointwise convolution; σ and δ respectively represent the Sigmoid function and the ReLU activation function, and B represents the normalization operation; M, N, and Z respectively represent the low-level features of the large-scale feature map, the high-level features of the small-scale feature map, and the cross-layer fusion features; represents element-wise multiplication.

6. The method for measuring eyelid morphological parameters of a reconstruction-segmentation network based on a conjoined architecture according to claim 1, characterized in that, The dynamic parameter convolution process is performed within the dynamic convolution module, specifically: Perform convolution, batch normalization, and ReLU non-linear combination operations on the feature map output by the encoder in the reconstruction task part, and then perform parameter deformation to determine the parameters of the dynamic convolution kernel; Perform convolution, batch normalization, and ReLU non-linear combination operations on the feature map output by the encoder in the segmentation task part as the convolution object; Use the dynamic convolution kernel to perform a convolution operation on the convolution object, and the resulting feature map is added to the feature map output by the encoder in the segmentation task part as the final output feature map of the encoder in the segmentation task part.

7. A method for measuring eyelid morphological parameters of a reconstruction-segmentation network based on a conjoined architecture according to claim 6, characterized in that The formula for the dynamic parameter convolution process is: DPConv(X, Y) = Conv2d(γX, ψ(γ(Y))); Where γ represents the combination of convolution, batch normalization, and ReLU non-linearity, X and Y respectively represent the feature maps output by the encoders in the segmentation task and the reconstruction task, ψ represents parameter deformation, and Conv2d represents two-dimensional convolution.

8. The eyelid morphology parameter measurement method based on the reconstruction-segmentation network of the Siamese architecture according to claim 6, characterized in that: The parameters for determining the dynamic convolution kernel through parameter transformation are specifically as follows: based on the feature map output by the encoder of the reconstruction task, perform parameter transformation ψ to obtain a one-dimensional vector θ of size n. n ; Based on the n parameters of θ n respectively determine the expected number of target channels, the size of the convolution kernel, and the number of output channels of the dynamic convolution kernel. For the semantic transformation, first use a 1×1 convolution and adaptive average pooling P to adjust the channel dimension, and then use another 1×1 convolution to reshape the features into θ n , and change the channels from z to n. The specific formula is as follows:

9. The eyelid morphology parameter measurement method based on the reconstruction-segmentation network of the Siamese architecture according to claim 1, characterized in that: The eyelid shape parameters are divided into three categories in total, including palpebral fissure height, palpebral fissure width, and palpebral fissure area. The palpebral fissure height includes left palpebral fissure height, central palpebral fissure height, and right palpebral fissure height.

Citation Information

Patent Citations

  • Speaker identification method based on SRS-CL network

    CN116434759A

  • Eye fundus macular pore structure three-dimensional model reconstruction method and system based on layered multi-dimensional path coding

    CN118898681A

  • Dynamic multi-modal segmentation selection and fusion

    CN119156637A

  • Image segmentation method and device combining feature difference recognition and detail enhancement

    CN119339075A

  • Three-stage feature enhanced myocardial infarction positioning method

    CN119488296A

Cited By

  • Image extraction method and prediction method of lipid whisker morphological parameters

    CN121281054A