Ocular surface image segmentation method and device and ocular surface analyzer
By constructing a segmentation model of a multimodal hybrid attention module and a multi-decoder, the fixed requirements of the eye table image segmentation model in the prior art for shooting angle, imaging mode and image number are solved, and compatible segmentation of multi-type eye table images is realized, and the flexibility and accuracy of the segmentation task are improved.
Patent Information
- Application Number
- CN202510001274.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
AI Technical Summary
Existing eye table image segmentation models are usually designed to specifically segment a single eye table structure, and require input of fixed shooting angles, fixed imaging modes and fixed number of images, which are not compatible with multiple types of eye table image segmentation tasks.
A segmentation model is adopted based on multiple encoders, multimodal hybrid attention modules and multiple decoders. By obtaining eye table image sets with different shooting angles and modal types, the cross attention module and self-attention module are used to extract and reconstruct features to achieve accurate segmentation of different eye table structures.
The compatibility of different eye surface segmentation tasks is improved, and the eye surface images of different modes can be processed at the same shooting angle, and accurate segmentation results can still be achieved in the presence of modes or missing shooting angles.
Smart Images

Figure CN119942113A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an ocular surface image segmentation method, device and ocular surface analyzer. Background Art
[0002] Ocular surface diseases are one of the most common eye diseases, which can cause a variety of eye discomforts and even the risk of blindness. Common ocular surface diseases such as meibomian gland dysfunction, dry eyes and pterygium usually rely on the quantification of key biomarkers for diagnosis. However, it is difficult for ophthalmologists to obtain accurate quantitative indicators of different ocular surface structures from ocular surface images only through slit lamp examination. Therefore, ocular surface image segmentation technology based on image processing to accurately extract ocular surface structures has emerged.
[0003] However, existing ocular surface image segmentation models are generally designed as networks specifically for segmenting single ocular surface structures such as the tear meniscus or meibomian glands. They usually require the input of a fixed shooting angle, a fixed imaging mode (i.e., modality), and a fixed number of images. Otherwise, the accuracy of the segmentation results cannot be guaranteed, and even abnormal situations such as model crash may occur. As a result, the existing ocular surface image segmentation methods can only undertake a single and fixed ocular surface segmentation task and are not compatible with multiple types of ocular surface image segmentation tasks. Summary of the invention
[0004] The problem solved by the present invention is how to improve the compatibility with different ocular surface segmentation tasks.
[0005] In order to solve the above problems, the present invention provides a method for segmenting an ocular surface image, comprising:
[0006] Acquire at least one set of ocular surface image sets, wherein the ocular surface images in the set of ocular surface image sets correspond to the same shooting angle and different modality types;
[0007] Inputting each group of the ocular surface image sets into a preset segmentation model to obtain a segmentation result corresponding to each ocular surface image, wherein the segmentation model is constructed based on multiple encoders, a multimodal mixed attention module and multiple decoders, and the multimodal mixed attention module includes a cross attention module and a self-attention module;
[0008] Each of the encoders is used to extract features from at most one ocular surface image to obtain at most one target feature map;
[0009] When there is no modality missing in the ocular surface image set, the cross-attention module is used to simultaneously convert the target feature map corresponding to each of the ocular surface images in the ocular surface image set into a cross-attention feature;
[0010] When there is a modality missing in the ocular surface image set, the self-attention module is used to convert the target feature map corresponding to each ocular surface image in the ocular surface image set into a self-attention feature respectively;
[0011] Each of the decoders is used to reconstruct the features of at most one of the cross-attention features or the self-attention features to obtain at most one of the segmentation results.
[0012] Optionally, before inputting each group of the ocular surface image sets into a preset segmentation model to obtain a segmentation result corresponding to each ocular surface image, the method further includes:
[0013] Determining a modality type set corresponding to each group of the ocular surface image sets, and judging whether the modality type set matches a preset modality type set corresponding to the ocular surface image set;
[0014] If yes, then there is no modality missing in the ocular surface image set;
[0015] If not, there is a modality missing in the ocular surface image set.
[0016] Optionally, the encoder includes a multi-view capture module, a downsampling module and a residual module; the encoder is specifically used for:
[0017] The ocular surface image is converted into a plurality of feature maps of different perspectives by the multi-perspective capture module, and the ocular surface image and the plurality of feature maps are fused to obtain a multi-perspective fusion feature;
[0018] Downsampling the multi-view fusion feature by the downsampling module;
[0019] The residual module performs deep semantic extraction on the multi-view fusion features downsampled by the downsampling module to obtain the target feature map.
[0020] Optionally, converting the ocular surface image into feature maps of multiple different perspectives by the multi-perspective capture module includes:
[0021] Obtaining a plurality of duplicate images according to the ocular surface image, and cutting each of the duplicate images into a plurality of image blocks;
[0022] A perspective transformation operation is performed on each of the copy images to obtain multiple feature maps, wherein the perspective transformation operation includes rotating each of the image blocks in the same copy image by the same angle and then splicing them together, wherein the rotation angles corresponding to the image blocks in different copy images are different.
[0023] Optionally, the multi-view fusion feature satisfies:
[0024]
[0025] Among them, F fusion Represents the multi-view fusion feature; Conv 1×1 Represents 1×1 convolution; DConv 7×7 represents a 7×7 depthwise separable atrous convolution; cat represents a concatenation operation; F represents the ocular surface image; F1, F2 and F n represents a plurality of said feature maps.
[0026] Optionally, the preset modality type set includes a natural light modality and an infrared light modality; when there is no modality missing in the ocular surface image set, the cross-attention module is used to simultaneously convert the target feature map corresponding to each ocular surface image in the ocular surface image set into a cross-attention feature, including:
[0027] Projecting each of the target feature maps into an input vector respectively, wherein the input vector includes a query vector, a key vector, and a content vector;
[0028] The cross-attention feature is obtained based on each of the input vectors, and the cross-attention feature satisfies:
[0029]
[0030] Wherein, A represents the first ocular surface image corresponding to the natural light modality; B represents the second ocular surface image corresponding to the infrared light modality; V AB represents the cross attention feature corresponding to the first ocular surface image; V BA represents the cross-attention feature corresponding to the second eye surface image; Q, K and V represent the query vector, the key vector and the content vector respectively; d K represents the dimension of the key vector; FFN represents a two-layer perceptron with a GELU activation function; Norm represents layer normalization; and softamax represents a normalized exponential function.
[0031] Optionally, when there is a modality missing in the ocular surface image set, the self-attention module is used to convert the target feature map corresponding to each ocular surface image in the ocular surface image set into a self-attention feature, including:
[0032] Each of the target feature maps is projected as the input vector, and the self-attention feature is obtained based on each of the input vectors, and the self-attention feature satisfies:
[0033]
[0034] Wherein, C represents any one of the ocular surface images in the ocular surface image set; VCC Represents the self-attention feature corresponding to any one of the eye surface images.
[0035] Optionally, the multimodal mixed attention module also includes multiple convolutional modulation modules; the number of the encoders, the convolutional modulation modules and the decoders is the same, and each of the convolutional modulation modules is used to encode spatial information of at most one of the cross-attention features or the self-attention features.
[0036] In the present invention, for the task of segmenting the structure of the ocular surface image, different shooting angles can capture different ocular surface structures. In order to be compatible with the segmentation tasks of multiple ocular surface structures, the present invention obtains the ocular surface images in each group of ocular surface image sets corresponding to the same shooting angle, which is equivalent to classifying and acquiring the ocular surface images according to the shooting angle, so that the subsequent segmentation model can perform specific ocular surface structure segmentation for each shooting angle. In addition, the ocular surface images in a group of ocular surface image sets correspond to the same shooting angle and different modality types, which is convenient for judging whether there is a modality missing. On this basis, each group of ocular surface image sets is input into a preset segmentation model. The segmentation model is constructed based on multiple encoders, a multimodal mixed attention module and multiple decoders. Each encoder is used to extract features from at most one ocular surface image to obtain at most one target feature map. While ensuring that each input ocular surface image can realize independent feature extraction, it is also conducive to compatibility with different segmentation tasks corresponding to various situations such as modality missing or shooting angle missing. On this basis, when there is no modality missing in the eye surface image set, the cross-attention module can be used to simultaneously convert the target feature map corresponding to each eye surface image in the eye surface image set into a cross-attention feature, so as to realize the cross-modal feature fusion of multiple eye surface images of different modalities at the same shooting angle (i.e., the same eye surface image set), which is conducive to capturing the feature association between multiple eye surface images of different modal types. When there is a modality missing in the eye surface image set, the self-attention module can be used to convert the target feature map corresponding to each eye surface image in the eye surface image set into a self-attention feature, so as to realize the mining of the correlation between features at different positions in a single eye surface image, which is conducive to capturing the long-distance dependency in the image. On this basis, each decoder in the present invention is used to perform feature reconstruction on at most one cross-attention feature or self-attention feature, obtain at most one segmentation result, and realize the matching of multiple decoders with multiple encoders. Therefore, this embodiment can ensure that each eye surface image obtains accurate segmentation results regardless of segmentation tasks such as modality missing or simultaneous multi-modality processing at the same shooting angle, or modality missing at some shooting angles and cross-modal feature fusion is required for some shooting angles, thereby achieving compatibility with different segmentation tasks.
[0037] The present invention also provides an ocular surface image segmentation device, comprising:
[0038] An acquisition module, which is used to acquire at least one set of ocular surface image sets, wherein the ocular surface images in the set of ocular surface image sets correspond to the same shooting angle and different modality types;
[0039] A segmentation module, which is used to input each group of the ocular surface image sets into a preset segmentation model to obtain a segmentation result corresponding to each ocular surface image, wherein the segmentation model is constructed based on multiple encoders, a multimodal mixed attention module and multiple decoders, and the multimodal mixed attention module includes a cross attention module and a self-attention module;
[0040] Each of the encoders is used to extract features from at most one ocular surface image to obtain at most one target feature map;
[0041] When there is no modality missing in the ocular surface image set, the cross-attention module is used to simultaneously convert the target feature map corresponding to each of the ocular surface images in the ocular surface image set into a cross-attention feature;
[0042] When there is a modality missing in the ocular surface image set, the self-attention module is used to convert the target feature map corresponding to each ocular surface image in the ocular surface image set into a self-attention feature respectively;
[0043] Each of the decoders is used to reconstruct the features of at most one of the cross-attention features or the self-attention features to obtain at most one of the segmentation results.
[0044] The advantages of the ocular surface image segmentation device provided by the present invention and the ocular surface image segmentation method compared with the prior art are basically the same, and will not be repeated here.
[0045] The present invention also provides an ocular surface analyzer, comprising a computer-readable storage medium storing a computer program and a processor. When the computer program is read and executed by the processor, the ocular surface image segmentation method as described above is implemented.
[0046] The advantages of the ocular surface analyzer provided by the present invention and the ocular surface image segmentation method compared with the prior art are basically the same, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 Schematic diagram of the flow of the method for segmenting an ocular surface image according to an embodiment of the present invention;
[0048] Figure 2 is a schematic structural diagram of an encoder according to an embodiment of the present invention;
[0049] Figure 3 Schematic diagram of the structure of the residual module according to an embodiment of the present invention;
[0050] Figure 4 is a schematic structural diagram of a multi-view capturing module according to an embodiment of the present invention;
[0051] Figure 5 It is a structural schematic diagram of a segmentation model according to an embodiment of the present invention;
[0052] Figure 6 Schematic diagram of the structure of a decoder according to an embodiment of the present invention. DETAILED DESCRIPTION
[0053] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be interpreted as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not intended to limit the scope of protection of the present invention.
[0054] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0055] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments"; the term "optionally" means "optional embodiments". Relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0056] It should be noted that the modifications of "one" and "plurality" mentioned in the present invention are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0057] like Figure 1 As shown, an embodiment of the present invention provides a method for segmenting an ocular surface image, comprising:
[0058] S1: Acquire at least one set of ocular surface image sets, wherein the ocular surface images in a set of ocular surface image sets correspond to the same shooting angle and different modality types.
[0059] Specifically, a set of ocular surface image sets referred to in this embodiment represents an image set corresponding to a shooting angle and containing at least one ocular surface image. For example, multiple storage paths can be set up in advance, and each storage path corresponds to an ocular surface image of a shooting angle. In actual use, the operator can select the path to upload the image according to the shooting angle corresponding to the ocular surface image. After the upload is completed, the ocular surface images in each storage path correspond to an ocular surface image set. The shooting angles referred to in this embodiment may include facing the cornea and facing the meibomian glands, etc. The ocular surface images in each set of ocular surface image sets referred to in this embodiment correspond to different modality types. Different modality types represent images taken using different imaging modes. The imaging modes may include natural light imaging modes and infrared imaging modes, etc. It should be understood that for some shooting angles, the actual required imaging mode may be only one. For example, when extracting the ocular surface structures of the meibomian glands, glandular areas, and eyelid areas, only images of the modality facing the eyelids in the infrared imaging mode are required. At the same time, in this embodiment, the number of ocular surface image sets obtained each time is related to different segmentation task requirements, and may be all ocular surface image sets corresponding to all shooting angles, or may be partial ocular surface image sets corresponding to partial shooting angles.
[0060] In this embodiment, for the task of segmenting the structure of the ocular surface image, different shooting angles can capture different ocular surface structures. In order to be compatible with the segmentation tasks of multiple ocular surface structures, the ocular surface images in each group of ocular surface image sets obtained in this embodiment correspond to the same shooting angle, which is equivalent to classifying and acquiring the ocular surface images according to the shooting angle, so that the subsequent segmentation model can perform specific ocular surface structure segmentation for each shooting angle. In addition, the ocular surface images in a group of ocular surface image sets correspond to the same shooting angle and different modality types, which is conducive to reflecting the modality types covered by the ocular surface image set (such as the number of modality types in the ocular surface image set can be determined by the number of ocular surface images), which is convenient for determining whether there is modality missing in the ocular surface image set.
[0061] S2: input each set of ocular surface images into a preset segmentation model to obtain a segmentation result corresponding to each ocular surface image, wherein the segmentation model is constructed based on multiple encoders, a multimodal mixed attention module and multiple decoders, and the multimodal mixed attention module includes a cross attention module and a self-attention module;
[0062] Each encoder is used to extract features from at most one eye surface image to obtain at most one target feature map;
[0063] When there is no modality missing in the ocular surface image set, the cross-attention module is used to simultaneously convert the target feature map corresponding to each ocular surface image in the ocular surface image set into a cross-attention feature;
[0064] When there is a modality missing in the ocular surface image set, the self-attention module is used to convert the target feature map corresponding to each ocular surface image in the ocular surface image set into a self-attention feature;
[0065] Each decoder is used to reconstruct at most one cross-attention feature or self-attention feature to obtain at most one segmentation result.
[0066] Specifically, the present embodiment includes a plurality of encoders, each of which is used to extract features from at most one ocular surface image, wherein the number of encoders can be determined according to the number of shooting angles and the number of modal types corresponding to each shooting angle. For example, assuming that the shooting angles include two types, namely, facing the cornea and facing the eyelid, wherein the shooting angle facing the cornea can correspond to two modalities, namely, natural light imaging mode (recorded as image A) and infrared light imaging mode (recorded as image B), while the shooting angle facing the eyelid only corresponds to one infrared light imaging mode (recorded as image C). In actual use, at least three encoders need to be configured to ensure that when the input multiple ocular surface image sets include image A, image B, and image C, the target feature map corresponding to each ocular surface image can be extracted respectively. In actual use, due to different ocular surface segmentation tasks, the number of ocular surface images corresponding to each acquisition is also different. For example, based on image A, three ocular surface structures of pterygium, cornea, and pupil can be segmented, based on image B, two ocular surface structures of cornea and tear river can be segmented, and based on image C, three ocular surface structures of meibomian gland, gland area, and eyelid area can be segmented. If this segmentation task needs to obtain all the above 7 types of ocular surface structures, two ocular surface image sets need to be input. The first ocular surface image set includes image A and image B, and the second ocular surface image set only contains image C. At this time, each encoder is used to extract features from an ocular surface image. Assuming that this segmentation task only needs to segment the three ocular surface structures of the meibomian gland, glandular area, and eyelid area, only image C in the second ocular surface image set needs to be input. At this time, only one encoder is used to extract features from image C, and the other two encoders are idle. Therefore, in this embodiment, each encoder is used to extract features from at most one ocular surface image in a segmentation task. While ensuring that each input ocular surface image can achieve independent feature extraction, it is also conducive to compatible with different segmentation tasks corresponding to various situations such as missing modalities or missing shooting angles (that is, there is no ocular surface image in the ocular surface image set corresponding to a certain shooting angle).
[0067] In one embodiment, the multimodal hybrid attention module includes a cross-attention module and a self-attention module, wherein the cross-attention module referred to in this embodiment is used to simultaneously process multiple eye surface images in an eye surface image set, perform cross-modal feature fusion on multiple eye surface images of different modalities at the same shooting angle (i.e., the same eye surface image set), and capture the feature association between multiple eye surface images of different modal types at the same shooting angle. The self-attention module referred to in this embodiment is used to separately process each eye surface image in the eye surface image set, mine the correlation between features at different positions in a single eye surface image, and better capture the long-distance dependency in the image. After obtaining at least one set of eye surface image sets, it can be judged whether there is a modality missing for each set of eye surface image sets. For example, assuming that there are at most two modality types corresponding to each shooting angle, if the number of corresponding eye surface images in the eye surface image set is 2, it means that there is no modality missing, and if the number of corresponding eye surface images in the eye surface image set is 1, it means that there is a modality missing. It should be understood that for some shooting angles, there may be only one type of modality. Such ocular surface images can only be processed by the self-attention module, which can be classified as a case of modality missing. For the same ocular surface image set, if there is no modality missing, the cross-attention module can be used to simultaneously process each ocular surface image in the ocular surface image set. If there is a modality missing, the self-attention module can be used to process each ocular surface image in the ocular surface image set separately to avoid inaccurate segmentation results or model network collapse caused by modality missing. For multiple groups of ocular surface image sets, the segmentation model of this embodiment can also be compatible with different modal situations of different ocular surface image sets at the same time, that is, to achieve compatibility with multiple segmentation tasks.
[0068] In one embodiment, each decoder is used to reconstruct at most one cross-attention feature or self-attention feature, and matches with multiple encoders to achieve a separate output of a segmentation result corresponding to a single eye surface image. The segmentation result referred to in this embodiment may include a one-hot encoding map or a segmentation map. For example, each pixel in the segmentation map can be assigned a label to represent a type of eye surface structure.
[0069] In this embodiment, for the task of eye surface image structure segmentation, different shooting angles can capture different eye surface structures. In order to be compatible with the segmentation tasks of multiple eye surface structures, the eye surface images in each group of eye surface image sets obtained in this embodiment correspond to the same shooting angle, which is equivalent to classifying and acquiring the eye surface images according to the shooting angle, so that the subsequent segmentation model can perform specific eye surface structure segmentation for each shooting angle. In addition, the eye surface images in a group of eye surface image sets correspond to the same shooting angle and different modality types, which is conducive to accurately reflecting the situation of the modality types covered by the eye surface image set through the number of eye surface images, and is convenient for judging whether there is a modality missing. On this basis, each group of eye surface image sets is input into a preset segmentation model. The segmentation model is constructed based on multiple encoders, multimodal mixed attention modules and multiple decoders. Each encoder is used to extract features from at most one eye surface image to obtain at most one target feature map. While ensuring that each input eye surface image can realize independent feature extraction, it is also conducive to compatible with different segmentation tasks corresponding to various situations such as modality missing or shooting angle missing. On this basis, when there is no modality missing in the eye surface image set, the cross-attention module can be used to simultaneously convert the target feature map corresponding to each eye surface image in the eye surface image set into a cross-attention feature, so as to realize the cross-modal feature fusion of multiple eye surface images of different modalities at the same shooting angle (i.e., the same eye surface image set), which is conducive to capturing the feature association between multiple eye surface images of different modal types. When there is a modality missing in the eye surface image set, the self-attention module can be used to convert the target feature map corresponding to each eye surface image in the eye surface image set into a self-attention feature, so as to realize the mining of the correlation between features at different positions in a single eye surface image, which is conducive to capturing the long-distance dependency in the image. On this basis, each decoder in this embodiment is used to perform feature reconstruction on at most one cross-attention feature or self-attention feature, obtain at most one segmentation result, and realize the matching of multiple decoders with multiple encoders. Therefore, this embodiment can ensure that each eye surface image obtains accurate segmentation results regardless of segmentation tasks such as modality missing or simultaneous multi-modality processing at the same shooting angle, or modality missing at some shooting angles and cross-modal feature fusion is required for some shooting angles, thereby achieving compatibility with different segmentation tasks.
[0070] Optionally, before inputting each group of ocular surface image sets into a preset segmentation model to obtain a segmentation result corresponding to each ocular surface image, the method further includes:
[0071] Determining a modality type set corresponding to each group of ocular surface image sets, and judging whether the modality type set matches a preset modality type set corresponding to the ocular surface image set;
[0072] If so, there is no modality missing in the ocular surface image set;
[0073] If not, there is a modality missing in the ocular surface image set.
[0074] Specifically, after obtaining each group of ocular surface image sets, the ocular surface images in the same group of ocular surface image sets correspond to different modality types, and the modality type set corresponding to the group of ocular surface image sets can be determined according to the modality type corresponding to each ocular surface image. At the same time, according to the actual segmentation task requirements, all modality types required for the cross-attention module to perform cross-modal feature fusion can be determined in advance to obtain a preset modality type set. On this basis, it is determined whether each group of modality type sets matches the preset modality type set corresponding to the group of ocular surface image sets. For example, assuming that the preset modality type set also corresponds to modality A and modality B, if the modality type set corresponds to modality A and modality B, then the two are considered to match, indicating that there is no modality missing in the ocular surface image set. If the modality type set only contains modality A, then the two are considered to be mismatched, indicating that there is a modality missing.
[0075] Optionally, since each eye surface image in the same set of eye surface image sets corresponds to a different modality type, the number of eye surface images in each set of eye surface image sets can also be obtained to obtain the current modality number. When the current modality number is equal to the preset modality number corresponding to the eye surface image set, it can be considered that there is no modality missing; when the current modality number is less than the preset modality number, it can be considered that there is a modality missing.
[0076] In this embodiment, when the modality type set matches the preset modality type set corresponding to the ocular surface image set, it indicates that there is no modality missing in the ocular surface image set, which can ensure that the modality type and quantity of the ocular surface images subsequently input into the cross-attention module match the model structure, avoiding abnormal situations such as inaccurate segmentation results and network crashes due to the ocular surface images not meeting the requirements of the cross-attention modality for input content. At the same time, if the two do not match, it means that there is a modality missing, and the self-attention module can be used to process the ocular surface images in the ocular surface image set separately. In this way, based on whether there is a modality missing in the ocular surface image set, the most suitable processing method can be selected for each group of ocular surface image sets, which is conducive to compatibility with diverse segmentation tasks.
[0077] Optionally, the encoder includes a multi-view capture module, a downsampling module and a residual module; the encoder is specifically used for:
[0078] The ocular surface image is converted into feature maps of multiple different perspectives through a multi-perspective capture module, and the ocular surface image and multiple feature maps are fused to obtain multi-perspective fusion features;
[0079] Downsample the multi-view fusion features through the downsampling module;
[0080] The residual module is used to perform deep semantic extraction on the multi-view fusion features that have been downsampled by the downsampling module to obtain the target feature map.
[0081] Specifically, the multi-view capture module can convert the ocular surface image into feature maps of multiple different viewpoints. The different viewpoints referred to in this embodiment represent different viewpoints for the local area of the ocular surface image. For example, the ocular surface image can be orthogonally divided into four uniform blocks, and the angles of the blocks are changed and then reassembled into the original image size to obtain a feature map of different viewpoints. The original ocular surface image and the feature maps of multiple different viewpoints are fused to obtain multi-view fusion features, which is beneficial to improving the accuracy of image recognition. On this basis, the multi-view fusion features are downsampled by the downsampling module to increase the receptive field, help the segmentation model better capture the global information of the image, and improve the accuracy and robustness of the segmentation. The multi-view fusion features downsampled by the downsampling module are then deeply semantically extracted by the residual module.
[0082] Preferably, if Figure 2 As shown, in this embodiment, the encoder may include multiple multi-view capture modules, multiple downsampling modules, and multiple residual modules. Assuming that the input eye surface image size is H×W×3 (where H represents the original height of the eye surface image; W represents the original width of the eye surface image; 3 represents the number of channels), the eye surface image first passes through the first multi-view capture module to obtain the first multi-view fusion feature, and the first multi-view fusion feature is downsampled by the first downsampling module and encoded as a feature of H / 2×W / 2×dim (where dim represents the basic dimension of the network, and the basic dimension is equal to 64 in this embodiment). The first multi-view fusion feature is further fused by the second multi-view capture module, and the obtained second multi-view fusion feature is downsampled by the second downsampling module and encoded as a feature of H / 4×W / 4×2dim. On this basis, the first residual module performs deep semantic extraction on the downsampled second multi-view fusion feature to obtain a first intermediate feature map, and the first intermediate feature map is downsampled by the third downsampling module to obtain a feature encoded as H / 8×W / 8×4dim. The first intermediate feature map then passes through the second residual module and the fourth downsampling module to extract deeper dimensional features, and obtains the second intermediate feature map encoded as H / 16×W / 16×8dim. The second intermediate feature map is finally subjected to deep semantic extraction by the third residual module, and obtains the target feature map encoded as H / 16×W / 16×16dim. Therefore, as the network continues to deepen, the feature dimension continues to increase, so that the high-dimensional target feature map finally extracted has richer deep semantic information.
[0083] Alternatively, if Figure 3 As shown, the residual module can include 3×3 convolution, batch normalization and activation function (corresponding to Figure 3The RELU in is used to better learn the detailed information and local features of the image and further improve the accuracy of the segmentation results.
[0084] In this embodiment, the encoder first converts the ocular surface image into feature maps of multiple different perspectives through a multi-perspective capture module, and fuses the ocular surface image and multiple feature maps to obtain multi-perspective fusion features, which is conducive to encoding the local area of the ocular surface image from different perspectives, enhancing the segmentation model's ability to capture local features, and is conducive to alleviating the difficulty of feature extraction caused by blurred boundaries of the ocular surface image and light spot occlusion, and reducing the interference of noise on the final segmentation result. On this basis, the multi-perspective fusion features are downsampled and deep semantic information is extracted through the downsampling module and the residual module in turn, which is conducive to increasing the receptive field, helping the segmentation model to better capture the global information of the image and then extract high-level features, so that the target feature map can ensure the accuracy of subsequent ocular surface image segmentation.
[0085] Optionally, the ocular surface image is converted into feature maps of multiple different perspectives by a multi-perspective capture module, including:
[0086] Obtaining multiple copy images according to the ocular surface image, and cropping each copy image into multiple image blocks;
[0087] A perspective transformation operation is performed on each copy image to obtain multiple feature maps. The perspective transformation operation includes rotating each image block in the same copy image by the same angle and then splicing them together, wherein the rotation angles corresponding to the image blocks in different copy images are different.
[0088] Specifically, Figure 4 As shown, in this embodiment, the multi-view capture module may include a view transformation module, which is used to obtain multiple copy images based on the eye surface image, and perform a view transformation operation on each copy image to obtain multiple feature maps. For example, the eye surface image can be copied to obtain three copy images, which are respectively recorded as the first copy image, the second copy image, and the third copy image. On this basis, each copy image can be evenly cropped into multiple image blocks, such as orthogonal cropping into 16×16 image blocks ( Figure 4Only the example of cropping into 2×2 image blocks is shown in the figure), and each image block in the same copy image is rotated by the same angle through the perspective transformation module and then spliced to obtain multiple feature maps, wherein the rotation angles corresponding to the image blocks in different copy images are different. For example, each corresponding image block in the first copy image is rotated 90° clockwise and then spliced according to the original spatial position to obtain the first feature map; each corresponding image block in the second copy image is rotated 180° clockwise and then spliced according to the original spatial position to obtain the second feature map; each corresponding image block in the third copy image is rotated 270° clockwise and then spliced according to the original spatial position to obtain the third feature map. In this way, multiple feature maps of different perspectives can be obtained based on the original eye surface image.
[0089] In this embodiment, the perspective transformation operation is performed on each copy image respectively, and each image block in the same copy image is rotated by the same angle and then spliced together to obtain a feature map after the perspective change of the local area. The image blocks in different copy images have different corresponding rotation angles, which ensures the perspective richness of the final feature map and is beneficial to reducing the interference of problems such as blurred structural boundaries and light spot occlusion in the ocular surface image on the final segmentation result.
[0090] Optionally, the multi-view fusion feature satisfies:
[0091]
[0092] Among them, F fusion Represents multi-view fusion features; Conv 1×1 Represents 1×1 convolution; DConv 7×7 represents a 7×7 depthwise separable atrous convolution; cat represents a concatenation operation; F represents an ocular surface image; F1, F2, and F n Represents multiple feature maps.
[0093] Specifically, Figure 4 As shown, the multi-view capture module in this embodiment also includes a 7×7 depth-separable hole convolution and two 1×1 convolutions. After obtaining multiple feature maps of different viewpoints based on the original eye surface image, the original image can be spliced with multiple feature maps using a splicing operation to form a multi-view feature map. The multi-view feature map is spatially modeled, and the spatial relationship between different features is mined to obtain a spatial attention map. On this basis, the spatial attention map is subjected to a 1×1 convolution to mine the channel relationship between different features to obtain a spatial and channel attention map of the same size as the original image. Finally, the original eye surface image is dot-multiplied with the spatial and channel attention maps after a 1×1 convolution to obtain the final multi-view fusion feature.
[0094] In this embodiment, by constructing feature maps of multiple different perspectives, the local area of the ocular surface image is encoded from different perspectives. On this basis, the depth-separable dilated convolution is combined to further increase the perception range, and the 1×1 convolution is combined to deeply explore the spatial relationship and channel relationship between different features in the multi-perspective feature map, enrich the multi-perspective fusion feature information level, and further reduce the noise effect in the ocular surface image, which is conducive to improving the accuracy of the final segmentation result.
[0095] Optionally, the preset modality type set includes a natural light modality and an infrared light modality; when there is no modality missing in the ocular surface image set, the cross-attention module is used to simultaneously convert the target feature map corresponding to each ocular surface image in the ocular surface image set into a cross-attention feature, including:
[0096] Project each target feature map into an input vector respectively, the input vector includes a query vector, a key vector and a content vector;
[0097] Based on each input vector, a cross-attention feature is obtained, which satisfies:
[0098]
[0099] Where A represents the first ocular surface image corresponding to the natural light modality; B represents the second ocular surface image corresponding to the infrared light modality; V AB V represents the cross attention feature corresponding to the first eye image; BA represents the cross-attention feature corresponding to the second eye table image; Q, K and V represent the query vector, key vector and content vector respectively; d K represents the dimension of the key vector; FFN represents a two-layer perceptron with a GELU activation function; Norm represents layer normalization; softamax represents a normalized exponential function.
[0100] Specifically, in this embodiment, the preset modality type set includes natural light modality and infrared light modality. When the modality type set corresponding to a certain eye surface image set corresponds to natural light modality and infrared light modality, it means that there is no modality missing in the eye surface image set. At this time, the two encoders connected to the cross-attention module convert the eye surface images of the two modalities into two target feature maps respectively, and the cross-attention module can project each target feature map as an input vector. Figure 5As shown, the cross-attention module in the multimodal hybrid attention module may include multiple projection layers, and the number of projection layers is adapted to the number of corresponding connected encoders. In this embodiment, the cross-attention module includes two projection layers, which respectively project two target feature maps into corresponding input vectors. In the attention mechanism, the feature map can be projected into a query vector Q, a key vector K, and a content vector V, which is convenient for capturing the dependencies between different positions. In this embodiment, the first eye surface image A corresponding to the natural light modality is projected as Q A , K A and V A , the second ocular surface image B corresponding to the infrared light modality is projected as Q B , K B and V B :
[0101] Q A =F A W Q , K A =F A W K , V A =F A W V ;
[0102] Q B =F B W Q , K B =F B W K , V B =F B W V ;
[0103] Among them, F A represents the target feature map corresponding to the first eye image A; F B W represents the target feature map corresponding to the second eye surface image B; Q , W K and W V Represent the three weights of the projection layer respectively.
[0104] On this basis, the first eye surface image A and the second eye surface image B can be modeled with cross-attention, and they can be simultaneously converted into cross-modal cross-attention features. The cross-attention features satisfy:
[0105]
[0106] like Figure 5 As shown, the cross-intention module also includes a two-layer perceptron with a GELU activation function (corresponding to Figure 5 FFN in), a layer normalization (corresponding to Figure 5Norm) and a normalized exponential function (corresponding to Figure 5 In the above process, Q A With K B The result of matrix multiplication is the same as V A Perform matrix dot multiplication to obtain the first cross-modal feature map corresponding to the first eye surface image A; Q B With K A The result of matrix multiplication is the same as V B Perform matrix dot multiplication to obtain the second cross-modal feature map corresponding to the second eye surface image B. The first cross-modal feature map and the second cross-modal feature map are respectively subjected to layer normalization to extract the correlation information between the internal channels of the feature maps, and then the nonlinear features are extracted by a two-layer perceptron with a GELU activation function. Finally, the final cross-attention feature is obtained by a normalized exponential function, which fully explores the association between eye surface images of different modalities, and is conducive to using the complementary advantages of multiple modalities to further improve the accuracy of the segmentation results.
[0107] Preferably, if Figure 5 As shown, in order to make the cross-attention feature adapt to the input format of subsequent convolution, the cross-attention module in this embodiment also includes a deformation module. After obtaining the cross-attention feature, the deformation module can be used to perform a deformation operation on it to deform the vector into a feature map of H / 16×W / 16×16 dim.
[0108] Optionally, when there is a modality missing in the ocular surface image set, the self-attention module is used to convert the target feature map corresponding to each ocular surface image in the ocular surface image set into a self-attention feature, including:
[0109] Each target feature map is projected as an input vector, and a self-attention feature is obtained based on each input vector. The self-attention feature satisfies:
[0110]
[0111] Where C represents any ocular surface image in the ocular surface image set; V CC Represents the self-attention feature corresponding to any eye surface image.
[0112] Specifically, the preset modality type set corresponding to the ocular surface image set in this embodiment includes natural light modality and infrared light modality. When the modality type set corresponding to the ocular surface image set only includes natural light modality or infrared light modality, it indicates that there is a modality missing in the ocular surface image set. It should be understood that for some shooting angles, only one modality type is needed to complete the corresponding ocular surface structure segmentation, that is, the preset modality type set corresponding to the ocular surface image set only includes one modality type, and it can be assumed that there is a modality missing in this case (for example, a virtual modality type can be added to this preset modality type set so that it can be determined to be a modality missing). Optionally, the number of self-attention modules in this embodiment can be set according to actual needs. For example, assuming that the preset modality type set includes three modality types, when the current ocular surface image set includes only two modality types or one modality type, it belongs to the situation where there is a modality missing. At least two self-attention modules can be set accordingly, which is conducive to the parallel processing of multiple ocular surface images and improves processing efficiency.
[0113] In one embodiment, if Figure 5 As shown, each self-attention module includes a projection layer for projecting the target feature map into the corresponding input vector. For any eye surface image C in the eye surface image set with modality loss, the projection layer can project it into Q C , K C and V C :
[0114] Q C =F C W Q , K C =F C W K , V C =F C W V ;
[0115] Among them, F C represents the target feature map corresponding to the ocular surface image C; W Q , W K and W V Represent the three weights of the projection layer respectively.
[0116] On this basis, self-attention modeling can be performed on any eye surface image C in the eye surface image set, and the self-attention features corresponding to the eye surface image C can be obtained:
[0117]
[0118] like Figure 5 As shown, the self-attention module also includes a two-layer perceptron with a GELU activation function (corresponding to Figure 5FFN in), a layer normalization (corresponding to Figure 5 Norm) and a normalized exponential function (corresponding to Figure 5 In the above process, Q C With K C The result of matrix multiplication is the same as V C Perform matrix dot multiplication to obtain the self-attention feature map corresponding to any eye surface image C. The self-attention feature map is normalized to extract the correlation information between the internal channels of the feature map, and then the nonlinear features are extracted by a two-layer perceptron with a GELU activation function. Finally, the final self-attention feature is obtained by a normalized exponential function, which fully explores the association between information in the same eye surface image and between different positions, which is conducive to enhancing the contextual connection of the features.
[0119] Preferably, if Figure 5 As shown, in order to make the self-attention feature adapt to the input format of subsequent convolution, the self-attention module in this embodiment also includes a deformation module. After obtaining the self-attention feature, the deformation module can be used to perform a deformation operation on it, and the vector can be deformed into a feature map of H / 16×W / 16×16dim.
[0120] Optionally, the multimodal mixed attention module also includes multiple convolutional modulation modules; the number of encoders, convolutional modulation modules and decoders is the same, and each convolutional modulation module is used to encode spatial information of at most one cross-attention feature or self-attention feature.
[0121] In this embodiment, after the cross-attention feature or the self-attention feature is obtained, it can also be input into the convolution modulation module to further encode the spatial information and better capture the feature information. Figure 5 As shown, the convolutional modulation module in this embodiment includes a first linear layer (corresponding to Linear1 in the figure), a second linear layer (corresponding to Linear2 in the figure), a third linear layer (corresponding to Linear3 in the figure), and an 11×11 depth-separable dilated convolution (corresponding to DCnv in the figure). The output of the convolutional modulation module satisfies:
[0122] Z=W3(DConv 11×11 (W1X))⊙(W2X);
[0123] Among them, Z represents the output of the convolutional modulation module; X represents the input of the convolutional modulation module (i.e., cross-attention feature or self-attention feature); DConv 11×11 represents a 11×11 depthwise separable atrous convolution; ⊙ represents a Hadamard product operation (corresponding to Figure 5 W1, W2, and W3 represent the weights corresponding to the first linear layer, the second linear layer, and the third linear layer, respectively.
[0124] Preferably, in the segmentation model provided in this embodiment, weights are shared between each encoder, weights are shared between each projection layer, weights are shared between each two-layer perceptron, and weights are shared between the depthwise separable atrous convolutions in each convolutional modulation module, which is beneficial to reducing the amount of parameters that need to be trained in the model and improving the efficiency of parameter adjustment.
[0125] Alternatively, if Figure 6 As shown, the decoder in this embodiment includes multiple residual modules and multiple upsampling modules. The cross-attention features or self-attention features of the convolution modulation module are transformed into a feature map of H / 16×W / 16×16dim and input into the decoder. After passing through four groups of residual modules and upsampling modules, the size of the feature map is gradually restored from H / 16×W / 16×16dim to a feature map of H / 16×W / 16×8dim, a feature map of H / 8×W / 8×4dim, a feature map of H / 4×W / 4×2dim, and a feature map of H / 2×W / 2×dim. Finally, it is restored to a segmentation map of the same size as the original input eye surface image through a residual module, wherein each pixel in the segmentation map can be assigned a label, representing a type of eye surface structure, thereby completing the segmentation task of the eye surface image.
[0126] Another embodiment of the present invention further provides an ocular surface image segmentation device, comprising:
[0127] An acquisition module, which is used to acquire at least one set of ocular surface image sets, wherein the ocular surface images in a set of ocular surface image sets correspond to the same shooting angle and different modality types;
[0128] A segmentation module, which is used to input each set of ocular surface images into a preset segmentation model to obtain a segmentation result corresponding to each ocular surface image, wherein the segmentation model is constructed based on multiple encoders, a multimodal mixed attention module and multiple decoders, and the multimodal mixed attention module includes a cross attention module and a self-attention module;
[0129] Each encoder is used to extract features from at most one eye surface image to obtain at most one target feature map;
[0130] When there is no modality missing in the ocular surface image set, the cross-attention module is used to simultaneously convert the target feature map corresponding to each ocular surface image in the ocular surface image set into a cross-attention feature;
[0131] When there is a modality missing in the ocular surface image set, the self-attention module is used to convert the target feature map corresponding to each ocular surface image in the ocular surface image set into a self-attention feature;
[0132] Each decoder is used to reconstruct at most one cross-attention feature or self-attention feature to obtain at most one segmentation result.
[0133] The ocular surface image segmentation device and the ocular surface image segmentation method provided in this embodiment can produce substantially the same technical effects, which will not be described in detail herein.
[0134] Another embodiment of the present invention further provides an ocular surface analyzer, comprising a computer-readable storage medium storing a computer program and a processor. When the computer program is read and executed by the processor, the ocular surface image segmentation method as described above is implemented.
[0135] The ocular surface analyzer and the ocular surface image segmentation method provided in this embodiment can produce basically the same technical effects, which will not be described in detail here.
[0136] An electronic device that can be used as a server or client of the present invention will now be described, which is an example of a hardware device that can be applied to various aspects of the present invention. Electronic devices are intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples, and are not intended to limit the implementation of the present invention described and / or required herein.
[0137] The electronic device includes a computing unit, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) or a computer program loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for the operation of the device can also be stored. The computing unit, ROM and RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.
[0138] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.
[0139] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc. In the present application, the unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment of the present invention. In addition, each functional unit in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0140] Although the present invention is disclosed as above, the protection scope of the present invention is not limited thereto. Those skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will fall within the protection scope of the present invention.
Claims
1. A method for segmenting an ocular surface image, characterized in that: include: Acquire at least one set of ocular surface image sets, wherein the ocular surface images in the set of ocular surface image sets correspond to the same shooting angle and different modality types; Inputting each group of the ocular surface image sets into a preset segmentation model to obtain a segmentation result corresponding to each ocular surface image, wherein the segmentation model is constructed based on multiple encoders, a multimodal mixed attention module and multiple decoders, and the multimodal mixed attention module includes a cross attention module and a self-attention module; Each of the encoders is used to extract features from at most one ocular surface image to obtain at most one target feature map; When there is no modality missing in the ocular surface image set, the cross-attention module is used to simultaneously convert the target feature map corresponding to each of the ocular surface images in the ocular surface image set into a cross-attention feature; When there is a modality missing in the ocular surface image set, the self-attention module is used to convert the target feature map corresponding to each ocular surface image in the ocular surface image set into a self-attention feature respectively; Each of the decoders is used to reconstruct the features of at most one of the cross-attention features or the self-attention features to obtain at most one of the segmentation results.
2. The method for segmenting an ocular surface image according to claim 1, characterized in that: Before inputting each group of the ocular surface image sets into a preset segmentation model to obtain a segmentation result corresponding to each ocular surface image, the method further includes: Determining a modality type set corresponding to each group of the ocular surface image sets, and judging whether the modality type set matches a preset modality type set corresponding to the ocular surface image set; If yes, then there is no modality missing in the ocular surface image set; If not, there is a modality missing in the ocular surface image set.
3. The method for segmenting an ocular surface image according to claim 1, characterized in that: The encoder includes a multi-view capture module, a downsampling module and a residual module; the encoder is specifically used for: The ocular surface image is converted into a plurality of feature maps of different perspectives by the multi-perspective capture module, and the ocular surface image and the plurality of feature maps are fused to obtain a multi-perspective fusion feature; Downsampling the multi-view fusion feature by the downsampling module; The residual module performs deep semantic extraction on the multi-view fusion features downsampled by the downsampling module to obtain the target feature map.
4. The method for segmenting an ocular surface image according to claim 3, characterized in that: The step of converting the ocular surface image into feature maps of multiple different perspectives by the multi-perspective capture module includes: Obtaining a plurality of duplicate images according to the ocular surface image, and cutting each of the duplicate images into a plurality of image blocks; A perspective transformation operation is performed on each of the copy images to obtain multiple feature maps, wherein the perspective transformation operation includes rotating each of the image blocks in the same copy image by the same angle and then splicing them together, wherein the rotation angles corresponding to the image blocks in different copy images are different.
5. The method for segmenting an ocular surface image according to claim 3, characterized in that: The multi-view fusion feature satisfies: Among them, F fusion Represents the multi-view fusion feature; Conv 1×1 Represents 1×1 convolution; DConv 7×7 represents a 7×7 depthwise separable atrous convolution; cat represents a concatenation operation; F represents the ocular surface image; F1, F2 and F n represents a plurality of said feature maps.
6. The method for segmenting an ocular surface image according to claim 2, characterized in that: The preset modality type set includes a natural light modality and an infrared light modality; when there is no modality missing in the ocular surface image set, the cross-attention module is used to simultaneously convert the target feature map corresponding to each ocular surface image in the ocular surface image set into a cross-attention feature, including: Projecting each of the target feature maps into an input vector respectively, wherein the input vector includes a query vector, a key vector, and a content vector; The cross-attention feature is obtained based on each of the input vectors, and the cross-attention feature satisfies: Wherein, A represents the first ocular surface image corresponding to the natural light modality; B represents the second ocular surface image corresponding to the infrared light modality; V AB represents the cross attention feature corresponding to the first ocular surface image; V BA represents the cross-attention feature corresponding to the second eye surface image; Q, K and V represent the query vector, the key vector and the content vector respectively; d K represents the dimension of the key vector; FFN represents a two-layer perceptron with a GELU activation function; Norm represents layer normalization; and softamax represents a normalized exponential function.
7. The method for segmenting an ocular surface image according to claim 6, characterized in that: When there is a modality missing in the ocular surface image set, the self-attention module is used to convert the target feature map corresponding to each ocular surface image in the ocular surface image set into a self-attention feature, including: Each of the target feature maps is projected as the input vector, and the self-attention feature is obtained based on each of the input vectors, and the self-attention feature satisfies: Wherein, C represents any one of the ocular surface images in the ocular surface image set; V CC Represents the self-attention feature corresponding to any one of the eye surface images.
8. The ocular surface image segmentation method according to claim 1, characterized in that: The multimodal mixed attention module also includes multiple convolutional modulation modules; the number of the encoders, the convolutional modulation modules and the decoders is the same, and each of the convolutional modulation modules is used to encode spatial information of at most one of the cross-attention features or the self-attention features.
9. A device for segmenting an ocular surface image, characterized in that: include: An acquisition module, which is used to acquire at least one set of ocular surface image sets, wherein the ocular surface images in the set of ocular surface image sets correspond to the same shooting angle and different modality types; A segmentation module, which is used to input each group of the ocular surface image sets into a preset segmentation model to obtain a segmentation result corresponding to each ocular surface image, wherein the segmentation model is constructed based on multiple encoders, a multimodal mixed attention module and multiple decoders, and the multimodal mixed attention module includes a cross attention module and a self-attention module; Each of the encoders is used to extract features from at most one ocular surface image to obtain at most one target feature map; When there is no modality missing in the ocular surface image set, the cross-attention module is used to simultaneously convert the target feature map corresponding to each of the ocular surface images in the ocular surface image set into a cross-attention feature; When there is a modality missing in the ocular surface image set, the self-attention module is used to convert the target feature map corresponding to each ocular surface image in the ocular surface image set into a self-attention feature respectively; Each of the decoders is used to reconstruct the features of at most one of the cross-attention features or the self-attention features to obtain at most one of the segmentation results.
10. An ocular surface analyzer, characterized in that: The method comprises a computer-readable storage medium storing a computer program and a processor, and when the computer program is read and executed by the processor, the method for segmenting an ocular surface image according to any one of claims 1 to 8 is implemented.