Boundary Segmentation Method of Scleral Lens and Tear Lens in Anterior Segment OCT Images
By constructing an improved U-shaped network combining residual module and hollow convolution module, the time-consuming and subjective differences in scleral mirror and tear mirror segmentation in the anterior segment OCT image is solved, and high-precision and stable boundary segmentation is achieved, which is suitable for quantitative analysis of scleral mirrors.
Patent Information
- Application Number
- CN202111485159.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-07
AI Technical Summary
The prior art has problems such as time-consuming and labor-intensive, large subjective differences, noise sensitivity and insufficient small-objective segmentation capabilities when segmenting the OCT images of the anterior segment. In particular, the traditional U-shaped network has poor application effect in this field.
An improved U-shaped network is constructed, combining residual modules and hollow convolution modules to be used for the boundary segmentation of scleral mirrors and tear mirrors of OCT images of the anterior segment. Through multi-layer encoding paths and decoding paths, the weighted loss functions focal loss and dice are combined to solve the class imbalance problem and achieve accurate segmentation.
The segmentation accuracy and stability of the boundaries of scleral and tear mirrors in the anterior segment OCT images are improved, and are suitable for quantitative analysis, reducing subjective differences and noise influences, and improving segmentation performance.
Smart Images

Figure CN115330663B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and analysis, and particularly relates to a method for segmenting scleral lenses and tear lenses in anterior segment OCT images. Background Art
[0002] A scleral lens is a special rigid gas-permeable contact lens that does not touch the cornea and the corneoscleral limbus and is completely supported by the sclera and the conjunctival tissue above it. Its special design allows it to completely "vault" over the corneoscleral limbus, and the "tear lens" formed under the lens can stably cover a large range of irregular surfaces caused by corneal structure changes, artificially reshaping a regular optical surface and improving the eye's optical system. In recent years, the rapid development of highly gas-permeable new materials and personalized scleral lens design technologies has further promoted the development of modern scleral lenses in clinical practice and the industry, becoming one of the most concerned directions in this field in the past five years. Numerous studies and clinical practices have shown that scleral lens treatment is a safe and effective means for treating difficult refractive errors and has become the preferred non-surgical treatment method for irregular corneas in many countries. In addition, it has been reported that scleral lenses can, to a certain extent, delay the progression of keratoconus and postpone the need for transplantation surgery for patients.
[0003] The three-dimensional thickness distribution (thickness map) of the tear lens after the scleral lens is worn in the eye is one of the important indicators for evaluating the good fit of the scleral lens. Conventionally, in clinical practice, a slit lamp is used to obtain a certain cross-section, and the doctor subjectively evaluates the thickness of the tear lens, unable to obtain the three-dimensional and quantitative thickness distribution of the tear lens, which brings difficulties to the fitting evaluation. With the progress of optical coherence tomography (OCT) technology, many scholars have used OCT technology to obtain three-dimensional scanned tomographic images of the scleral lens worn in the eye to assist in the fitting of the scleral lens. However, currently, the thickness distribution of the tear lens is mostly obtained by manually outlining the image, which is not only time-consuming and laborious but also has subjective differences, severely restricting its clinical promotion.
[0004] Several non-machine learning algorithms for segmenting anterior segment OCT images have been developed in the past: methods such as using graph theory, fast active contours and polynomial fitting, Canny edge detection, Gaussian mixture models, and the combination of Haugh transform and Kalman filter. The disadvantage of these methods is that they rely on specific "ad hoc" rules, such as image grayscale, spatial texture, geometric shape, etc., resulting in poor segmentation performance and generalization ability of application scenarios in the presence of noise and / or artifacts.
[0005] At present, it has become a trend to use deep learning technology to solve the problem of medical image segmentation. Convolutional neural networks represented by U-type networks use symmetrical encoder and decoder structures and combine jump connections, which play an important role in medical image segmentation. However, the traditional U-type network has poor segmentation effect on small targets, and there are problems such as loss of internal data structure and spatial hierarchy information and inability to reconstruct small parts of structural information. So far, there has been no report on a method based on U-type networks for segmenting the scleral lens, tear lens and the entire cornea in anterior segment OCT images. Summary of the invention
[0006] In view of the shortcomings of the prior art, the object of the present invention is to provide a method for segmenting the scleral lens and the tear lens in anterior segment OCT images.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] A method for segmenting the boundaries of a scleral lens and a tear lens in an anterior segment OCT image comprises the following steps:
[0009] 1) constructing a segmentation network model, which includes an encoder module with a multi-layer encoding path and a corresponding decoder module with a multi-layer decoding path, each layer of the encoding path and the decoding path are provided with a convolution layer, each layer of the encoding path is provided with a residual module for downsampling to capture information and deepen the network depth, and a hole convolution module for increasing the convolution receptive field and extracting multi-scale information;
[0010] 2) Collect a large number of anterior segment OCT 3D images of the cornea after wearing scleral lenses in different conditions, and divide them into a training set and a validation set, input the training set into the above segmentation network model for training, and verify it with the validation set, and after multiple trainings, obtain the trained segmentation network model;
[0011] 3) Use the trained segmentation network model to predict the upper and lower boundaries of the scleral lens and tear lens in the anterior segment OCT 3D image to obtain the precise segmentation structure of the scleral lens and tear lens boundaries in the anterior segment OCT 3D image, calculate the thickness between the upper and lower boundaries of the tear lens, and obtain the 3D thickness distribution map of the tear lens through 3D reconstruction.
[0012] The encoding path has four layers, and each layer is connected by a convolutional layer with stride=2.
[0013] The encoding paths at different levels contain different numbers of residual modules.
[0014] The encoding paths at different levels contain atrous convolution modules with different atrous rates.
[0015] Each layer of the encoding path includes a main route and a branch route. The main route is provided with a convolutional layer and a residual module cascaded with the convolutional layer. The branch route is provided with a dilated convolutional module. After passing through the main route and the branch route, the image is spliced to obtain a feature map.
[0016] The three-dimensional OCT images of the anterior segment of the eye to be measured are respectively input into the main and branch routes of the first layer of the encoding path. The main route in the first layer of the encoding path passes through a 7×7 convolutional layer with 32 channels, and is processed by batch normalization and the activation function relu. Then it passes through a residual module cascaded with 3 convolutional layers with 32 channels and a stride of 1. The branch route inputs the three-dimensional OCT image of the anterior segment of the eye to be measured into a cascaded dilated convolutional module with a dilation rate of 7, and then is processed by the activation function relu. Finally, the two parts of the image passing through the main and branch routes are connected by splicing to obtain a feature map with a size of 512×512 and 64 channels.
[0017] The main route of the second layer of the encoding path takes the output of the first layer of the encoding path as input, and passes through a convolutional layer with 64 channels and a stride of 2 and a residual module cascaded with 3 convolutional layers with 64 channels and a stride of 1. The branch route takes the output of the branch route in the first layer of the encoding path as input, and passes through a dilated convolutional module cascaded with a dilation rate of 5. Subsequently, the output of the main route and the output of the branch route are spliced together to obtain a feature map with a size of 256×256 and 128 channels.
[0018] The main route of the third layer of the encoding path takes the output of the second layer of the encoding path as input, and passes through a convolutional layer with 128 channels and a stride of 2 and a residual module cascaded with 5 convolutional layers with 128 channels and a stride of 1. The branch route takes the output of the branch route in the second layer of the encoding path as input, and passes through a dilated convolutional module cascaded with a dilation rate of 3. Subsequently, the output of the main route and the output of the branch route are spliced together to obtain a feature map with a size of 128×128 and 256 channels.
[0019] The main route of the fourth layer of the encoding path takes the output of the third layer of the encoding path as input, and passes through a convolutional layer with 256 channels and a stride of 2 and a residual module cascaded with 7 convolutional layers with 256 channels and a stride of 1. The branch route takes the output of the branch route in the third layer of the encoding path as input, and passes through a dilated convolutional module cascaded with a dilation rate of 2. Subsequently, the main output and the branch output are spliced together to obtain a feature map with a size of 64×64 and 512 channels.
[0020] An intermediate layer is also provided between the encoding path and the decoding path. The intermediate layer includes a main route and a branch route. The main route of the intermediate layer takes the output of the fourth layer of the encoding path as input and passes through a convolutional layer with 512 channels, 3x3 kernel size, and stride = 2, and two residual modules cascaded with convolutional layers with 512 channels, 3x3 kernel size, and stride = 1. In the branch route, the output of the branch route in the fourth layer of the encoding path is taken as input and passes through an atrous convolution block cascaded with dilation rate of 1. Subsequently, the main output and the branch output are concatenated together to obtain a feature map with a size of 32x32 and 1024 channels.
[0021] The first layer of the decoding path performs transposed convolution on the output result of the intermediate layer to obtain a feature map with a size of 64x64 and 512 channels, and then concatenates it with the output result of the fourth layer of the encoding path to obtain a feature map with a size of 64x64 and 1024 channels;
[0022] The second layer of the decoding path first passes the output result of the first layer of the decoding path through two convolutional blocks with 512 channels and 3x3 kernel size, and then performs transposed convolution to obtain a feature map with a size of 128x128 and 256 channels, and then concatenates it with the output result of the third layer of the encoding path to obtain a feature map with a size of 128x128 and 512 channels;
[0023] The third layer of the decoding path first passes the output result of the second layer of the decoding path through two convolutional blocks with 256 channels and 3x3 kernel size, and then performs transposed convolution to obtain a feature map with a size of 256x256 and 128 channels, and then concatenates it with the output result of the third layer of the encoding path to obtain a feature map with a size of 256x256 and 256 channels;
[0024] The fourth layer of the decoding path first passes the output result of the third layer of the decoding path through two convolutional blocks with 128 channels and 3x3 kernel size, and then performs transposed convolution to obtain a feature map with a size of 512x512 and 64 channels, and then concatenates it with the output result of the third layer of the encoding path to obtain a feature map with a size of 512x512 and 128 channels;
[0025] Finally, the output result is first passed through two convolutional blocks with 64 channels and 3x3 kernel size, and then input into a 1x1 convolutional layer with a softmax activation function to obtain the probability map of the segmentation target in the entire image area.
[0026] After each passing through a convolutional block in the decoding path, the activation function relu is performed.
[0027] A weighted loss function combining focal loss and dice is used to solve the class imbalance problem caused by the small proportion of the scleral lens area and the tear lens area compared to the background area.
[0028] The beneficial effects of the present invention:
[0029] 1. An improved U-shaped segmentation network combining a residual module and a dilated convolution module has the advantages of high efficiency and high boundary segmentation accuracy. It is applicable to the segmentation of scleral lenses, tear lenses, and the entire ocular surface boundary in anterior segment OCT images, with good segmentation performance, laying a foundation for subsequent quantitative analysis.
[0030] 2. The residual module and the dilated convolution module are combined in the U-shaped network. A large number of anterior segment images after wearing scleral lenses of various corneas (normal corneas and irregular corneas) are selected in the training data, ensuring the stability of model application. Brief Description of the Drawings
[0031] Figure 1 It is a schematic diagram of the network structure in the embodiment of the present application.
[0032] Figures 2(a) and 2(b) are schematic diagrams of the structural design of the residual module and the dilated convolution module of the present invention.
[0033] Figure 3 It is a schematic diagram of the training data in the scleral lens and tear lens boundary segmentation algorithm for anterior segment OCT images of the present invention.
[0034] Figure 4 It is a schematic diagram of the gold standard in the scleral lens and tear lens boundary segmentation algorithm for anterior segment OCT images of the present invention.
[0035] Figure 5 It is a schematic diagram of the boundary segmentation result of the scleral lens and tear lens segmentation algorithm for anterior segment OCT images of the present invention.
[0036] Figure 6 It shows the PVD results of different methods.
[0037] Figure 7 It shows the PVD comparison results of different surfaces of different methods.
[0038] Figure 8 It shows the dice coefficient results of different methods. Detailed Embodiments
[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0040] The present invention provides a method for segmenting the boundaries of scleral lenses and tear lenses in anterior segment OCT images, which includes the following steps:
[0041] 1) Construct a segmentation network model, which includes an encoder module with multiple layers of encoding paths and a corresponding decoder module with multiple layers of decoding paths. A convolutional layer is provided between each layer of the encoding path and the decoding path. Each layer of the encoding path is provided with a residual module for downsampling to capture global information and local information, and a dilated convolution module for increasing the convolutional receptive field.
[0042] The network structure includes an encoder module, a residual module, a dilated convolution module, and a decoder module. The residual module and the dilated convolution module are arranged in the encoding paths of the encoder module and are used for downsampling to capture global information and local information. The skip connections in U-net and the residual blocks help information propagate from low-level features to high-level features and retain more detailed boundary information. In addition, the residual module, as a building unit, simplifies the training process and is stacked in a cascaded manner, which helps to extract coarse boundary contour features and edge detail features from the original image. By setting the dilated convolution module, the convolutional receptive field can be increased without increasing the computational amount, thereby improving the boundary detection accuracy and segmentation timeliness.
[0043] Among them, the encoding path has four layers, and each layer includes a residual module and a dilated convolution module. The layers are downsampled by a convolutional layer with stride = 2 to reduce the image resolution and expand the receptive field. The residual module is composed of two 3x3 convolutional layers, and different levels of the encoding path contain different numbers of residual modules, which are 3, 4, 6, 8, and 3 respectively. The dilated convolution module is composed of two cascaded dilated convolutions, and different levels of the encoding path contain dilated convolutions with different dilation rates, which are [7, 7], [5, 5], [3, 3], [2, 2], [1, 1] respectively.
[0044] Each layer of the encoding path includes a main route and a branch route. The three-dimensional OCT image of the anterior segment of the eye to be measured is input into the main and branch routes of the first layer of the encoding path respectively. The main route in the first layer of the encoding path passes through a 7×7 convolutional layer with 32 channels, and after batch normalization and activation function relu processing, it passes through a residual module cascaded by 3 convolutional layers with 32 channels and 3×3, stride = 1. The branch route inputs the three-dimensional OCT image of the anterior segment of the eye to be measured into a cascaded dilated convolution module with a dilation rate of 7, then performs activation function relu processing, and finally connects the two parts of the image passing through the main and branch routes by splicing to obtain a feature map with a size of 512×512 and 64 channels.
[0045] The main path of the second layer of the encoding path takes the output of the first layer of the encoding path as input, passing through a convolutional layer with 64 channels, 3×3 kernel size, and stride = 2, and 3 residual modules cascaded with convolutional layers with 64 channels, 3×3 kernel size, and stride = 1. The branch path takes the output of the branch path in the first layer of the encoding path as input, passing through a dilated convolutional module cascaded with a dilation rate of 5. Subsequently, the outputs of the main path and the branch path are concatenated together to obtain a feature map with a size of 256×256 and 128 channels;
[0046] The main path of the third layer of the encoding path takes the output of the second layer of the encoding path as input, passing through a convolutional layer with 128 channels, 3x3 kernel size, and stride = 2, and 5 residual modules cascaded with convolutional layers with 128 channels, 3x3 kernel size, and stride = 1. The branch path takes the output of the branch path in the second layer of the encoding path as input, passing through a dilated convolutional module cascaded with a dilation rate of 3. Subsequently, the outputs of the main path and the branch path are concatenated together to obtain a feature map with a size of 128x128 and 256 channels.
[0047] The main path of the fourth layer of the encoding path takes the output of the third layer of the encoding path as input, passing through a convolutional layer with 256 channels, 3x3 kernel size, and stride = 2, and 7 residual modules cascaded with convolutional layers with 256 channels, 3x3 kernel size, and stride = 1. The branch path takes the output of the branch path in the third layer of the encoding path as input, passing through a dilated convolutional module cascaded with a dilation rate of 2. Subsequently, the main output and the branch output are concatenated together to obtain a feature map with a size of 64x64 and 512 channels.
[0048] An intermediate layer is also provided between the encoding path and the decoding path. The intermediate layer includes a main path and a branch path. The main path of the intermediate layer takes the output of the fourth layer of the encoding path as input, passing through a convolutional layer with 512 channels, 3x3 kernel size, and stride = 2, and 2 residual modules cascaded with convolutional layers with 512 channels, 3x3 kernel size, and stride = 1. The branch path takes the output of the branch path in the fourth layer of the encoding path as input, passing through a dilated convolutional block cascaded with a dilation rate of 1. Subsequently, the main output and the branch output are concatenated together to obtain a feature map with a size of 32x32 and 1024 channels.
[0049] The first layer of the decoding path performs transposed convolution on the output result of the intermediate layer to obtain a feature map with a size of 64x64 and 512 channels, and then concatenates it with the output result of the fourth layer of the encoding path to obtain a feature map with a size of 64x64 and 1024 channels;
[0050] The second layer of the decoding path first passes the output result of the first layer of the decoding path through two convolutional blocks with 512 channels and a size of 3x3, and then performs a transposed convolution to obtain a feature map with a size of 128x128 and 256 channels. Subsequently, it is concatenated with the output result of the third layer of the encoding path to obtain a feature map with a size of 128x128 and 512 channels;
[0051] The third layer of the decoding path first passes the output result of the second layer of the decoding path through two convolutional blocks with 256 channels and a size of 3x3, and then performs a transposed convolution to obtain a feature map with a size of 256x256 and 128 channels. Subsequently, it is concatenated with the output result of the third layer of the encoding path to obtain a feature map with a size of 256x256 and 256 channels;
[0052] The fourth layer of the decoding path first passes the output result of the third layer of the decoding path through two convolutional blocks with 128 channels and a size of 3x3, and then performs a transposed convolution to obtain a feature map with a size of 512x512 and 64 channels. Subsequently, it is concatenated with the output result of the third layer of the encoding path to obtain a feature map with a size of 512x512 and 128 channels;
[0053] Finally, the output result is first passed through two convolutional blocks with 64 channels and a size of 3x3, and then input into a 1x1 convolutional layer with a softmax activation function to obtain the probability map of the segmented target in the entire image area.
[0054] In the problem of anterior segment OCT image segmentation, there is a class imbalance problem, that is, the proportion of the scleral lens area and the tear lens area is very small compared to the background area. This class imbalance problem will cause the neural network to perform well in the background area detection, but poorly in the scleral lens area and the tear lens area. The weighted loss function focal loss and dice are combined to solve this problem. The loss function includes emphasizing the weight of the pixels in the target area or introducing a scaling factor for the background area, which can effectively solve the problem of data type imbalance in the training process.
[0055] The dice formula is as follows:
[0056] Where pred is the set of predicted values, true is the set of true values, the numerator is the intersection between pred and true, and multiplying by 2 is because the denominator has double-counted the common elements between pred and true. The denominator is the union of pred and true.
[0057] The focal loss formula is as follows: FL(p t )=-α t (1-p t ) γ log(p t ),
[0058] Among them, γ is taken as 2 and α is taken as 0.25.
[0059] In addition, considering the imbalance of their values, the final loss function is FL + 0.2 Dice.
[0060] 2) Collect a large number of three-dimensional anterior segment OCT images after wearing scleral lenses under different conditions, and divide them into a training set, a validation set, and a test set. Input the training set into the above segmentation network model for training, and verify it through the validation set. After multiple trainings, obtain the trained segmentation network model, and test it through the test set to confirm the accuracy of the trained segmentation network model.
[0061] After collecting a large number of three-dimensional anterior segment OCT images, first preprocess the three-dimensional images. For example, obtain a total of 1217 pictures, among which 928 pictures are used as the training set, 229 pictures are used as the validation set, and 60 pictures are divided into two groups as the test set. The images are all labeled by professional physicians through the MIT Licensed LabelMe software. A total of four interfaces are labeled, namely: air - upper surface of the scleral lens, lower surface of the scleral lens - upper surface of the tear lens, lower surface of the tear lens - upper surface of the cornea, lower surface of the cornea - anterior chamber.
[0062] Before all the data is fed into the model for training, it will go through preprocessing. The preprocessing steps first obtain the coordinates of the four interfaces outlined by the LabelMe software, then fill each position of the label map with a gray value of 0, and then, in the form of upper closed and lower open, fill different colors through the coordinates of the four interfaces to finally obtain the regional label map. Finally, both the original image and the regional label map are scaled to 512×512 through the nearest neighbor interpolation method.
[0063] For the training of the model, it is written in Python and run and debugged under the Keras framework. The experimental hardware uses two NVIDIA NVIDA TITAN RTX graphics cards, with a total video memory capacity of 48GB, and the GPU is used to accelerate the model training. During the training process, the Adam optimizer with a learning rate of 0.0001 is used to optimize the weight parameters in the network. After each training of the data, it is verified on the validation set, and the model parameters with the highest segmentation performance on the validation set are saved. In addition, a learning rate scheduler is adopted, and the parameter factor in the ReduceLROnPlateau function is set to 0.8, and the patience is set to 5, that is, when the loss value during training does not decrease in five iterative processes, the learning rate decays to 0.8 times the original. An early stopper is used, and the parameter patience of the EarlyStopping function is set to 100, indicating that when the loss value does not decrease in 100 iterations, the training stops.
[0064] In terms of data augmentation, we perform operations such as randomly moving horizontally and vertically, randomly zooming in, and randomly flipping the image horizontally to augment the data.
[0065] 3) Use the trained segmentation network model to predict the upper and lower boundaries of the scleral lens and the tear lens for the three-dimensional OCT image of the anterior segment of the eye to be measured, obtain the precise segmentation structure of the boundaries of the scleral lens and the tear lens in the three-dimensional OCT image of the anterior segment of the eye, and calculate the thickness between the upper and lower boundaries of the tear lens. Through three-dimensional reconstruction, obtain the three-dimensional thickness distribution map of the tear lens.
[0066] In the comparative experiment, we adopt two evaluation indexes: the Dice coefficient and the pixel - value difference (hereinafter referred to as PVD). In the comparative experiment, the method of the present invention is compared with three other networks. Two of the methods are segmentation methods based on the U - net network, including the dilated convolution U - net and the residual convolution U - net, and the other one is the traditional U - net. Figure 6 List the PVD results of different methods. Figure 7 List the PVD comparison results of different surfaces of different methods. Figure 8 List the dice coefficient results of different methods.
[0067] For PVD, in the classification result, a total of four boundaries are obtained, namely the upper surface of the scleral lens, the lower surface of the scleral lens, the upper surface of the cornea, and the lower surface of the cornea. For an image, taking the lower left corner as the origin, the length as the X - axis, and the width as the Y - axis to establish a rectangular coordinate system. Among them, Y_true represents the y - coordinate of the true point at this x value, and Y_pred represents the y - coordinate of the predicted point at this x value. The pixel distance difference is equal to the absolute value of the difference between Y_true and Y_pred (represented by formula 1). Therefore, we use PVD - 0 to represent that the pixel distance difference is between 0 and 2 pixels (excluding 2 pixels), PVD - 1 to represent that the pixel distance difference is between 2 and 5 pixels (including 2 pixels, excluding 5 pixels), PVD - 2 to represent that the pixel distance difference is between 5 and 8 pixels (the same as above), and PVD - 3 to represent that the pixel distance difference is greater than or equal to 8 pixels.
[0068] Pixel distance difference = ∣Y 真实 -Y 预测 ∣ (1)
[0069] The dice formula is as follows:
[0070] Among them, pred is the set of predicted values, true is the set of true values. The numerator is the intersection between pred and true, and multiplying by 2 is because the denominator has double - counted the common elements between pred and true. The denominator is the union of pred and true.
[0071] The results are as Figures 6 to 8 shown.
[0072] From Figure 6 and Figure 7 it can be seen that in the comparative experiment, we reflect the reliability of the network by calculating the distance between the predicted pixel points and the real pixel points. On the basis of the traditional U-Net, the dilation module and the residual module are added respectively, but the prediction results are worse than those of the traditional U-Net. The U-Net-based network built by our method is better than the other three networks in all aspects except that the upper surface of the cornea is slightly higher than that of the U-Net.
[0073] From Figure 8 it can be seen that the dice coefficients of our method are 0.98209, 0.94744, and 0.97101 respectively. Compared with the DU-Net, the dice coefficients increase by 0.340%, 0.716%, and 0.264% respectively. Compared with the original U-net network, they increase by 0.473%, 0.395%, and 0.706% respectively. Compared with the RESU-Net network, they increase by 0.649%, 1.343%, and 0.431% respectively, indicating that the network of the present invention has higher performance.
[0074] The embodiments should not be regarded as limitations of the present invention, but any improvements made based on the spirit of the present invention should be within the protection scope of the present invention.
Claims
1. A method for segmenting the boundaries of scleral lenses and tear lenses in anterior segment OCT images, characterized in that: It includes the following steps: 1) Construct a segmentation network model, which includes an encoder module with multiple layers of encoding paths and a corresponding decoder module with multiple layers of decoding paths. A convolutional layer is provided between each layer of the encoding path and the decoding path. Each layer of the encoding path is provided with a residual module for downsampling to capture information and deepen the network depth, and a dilated convolution module for increasing the convolutional receptive field to extract multi-scale information; 2) Collect a large number of anterior segment OCT three-dimensional images after wearing a scleral lens under different conditions, and divide them into a training set and a validation set. Input the training set into the above segmentation network model for training, and verify it through the validation set. After multiple trainings, obtain the trained segmentation network model; 3) Use the trained segmentation network model to predict the upper and lower boundaries of the scleral lens and the tear lens for the anterior segment OCT three-dimensional image to be measured, obtain the precise segmentation structure of the boundaries of the scleral lens and the tear lens in the anterior segment OCT three-dimensional image, calculate the thickness between the upper and lower boundaries of the tear lens, and obtain the three-dimensional thickness distribution map of the tear lens through three-dimensional reconstruction; The encoding path has four layers, and each layer is connected by a convolutional layer with stride = 2. Different levels of the encoding path contain different numbers of residual modules, and different levels of the encoding path contain dilated convolution modules with different dilation rates. Each layer of the encoding path includes a main route and a branch route. The main route is provided with a convolutional layer and a residual module cascaded with the convolutional layer, and the branch route is provided with a dilated convolution module. The image passes through the main route and the branch route and then is spliced to obtain a feature map.
2. The method for segmenting the boundaries of the scleral lens and the tear lens in the anterior segment OCT image according to claim 1, wherein: The anterior segment OCT three-dimensional image to be measured is respectively input into the main and branch routes of the first layer of the encoding path. The main route in the first layer of the encoding path passes through a 7×7 convolutional layer with 32 channels, and is processed by batch normalization and the activation function relu, and then passes through 3 residual modules cascaded with 3×3 convolutional layers with 32 channels and stride = 1. The branch route inputs the anterior segment OCT three-dimensional image to be measured into a cascaded dilated convolution module with a dilation rate of 7, and then is processed by the activation function relu. Finally, the two parts of the image passing through the main and branch routes are connected by splicing to obtain a feature map with 512×512 and 64 channels; The main route of the second layer of the encoding path takes the output of the first layer of the encoding path as input, passes through a 3×3 convolutional layer with 64 channels and stride = 2 and 3 residual modules cascaded with 3×3 convolutional layers with 64 channels and stride = 1. The branch route takes the output of the branch route in the first layer of the encoding path as input, passes through a cascaded dilated convolution module with a dilation rate of 5, and then splices the output of the main line and the output of the branch line together to obtain a feature map with 256×256 and 128 channels; The third main path of the encoding path takes the output of the second layer of the encoding path as input, passing through a convolutional layer with 128 channels, 3x3 kernel size, and stride = 2, and 5 residual modules cascaded with convolutional layers of 128 channels, 3x3 kernel size, and stride = 1. In the branch path, the output of the branch path in the second layer of the encoding path is taken as input, passing through a dilated convolutional module composed of dilations with a rate of 3 cascaded. Subsequently, the outputs of the main path and the branch path are concatenated together to obtain a feature map of 128 x 128 with 256 channels; The fourth main path of the encoding path takes the output of the third layer of the encoding path as input, passing through a convolutional layer with 256 channels, 3x3 kernel size, and stride = 2, and 7 residual modules cascaded with convolutional layers of 256 channels, 3x3 kernel size, and stride = 1. In the branch path, the output of the branch path in the third layer of the encoding path is taken as input, passing through a dilated convolutional module composed of dilations with a rate of 2 cascaded. Subsequently, the outputs of the main path and the branch path are concatenated together to obtain a feature map of 64 x 64 with 512 channels.
3. The method for segmenting the boundaries of the scleral lens and the tear lens in the anterior segment OCT image according to claim 2, wherein: An intermediate layer is further provided between the encoding path and the decoding path. The intermediate layer includes a main path and a branch path. The main path of the intermediate layer takes the output of the fourth layer of the encoding path as input, passing through a convolutional layer with 512 channels, 3x3 kernel size, and stride = 2, and 2 residual modules cascaded with convolutional layers of 512 channels, 3x3 kernel size, and stride = 1. In the branch path, the output of the branch path in the fourth layer of the encoding path is taken as input, passing through a dilated convolutional block composed of dilations with a rate of 1 cascaded. Subsequently, the outputs of the main path and the branch path are concatenated together to obtain a feature map of 32 x 32 with 1024 channels.
4. The method for segmenting the boundaries of scleral lenses and tear lenses in anterior segment OCT images according to claim 3, characterized in that: The first layer of the decoding path performs transposed convolution on the output result of the intermediate layer to obtain a feature map of 64 x 64 with 512 channels, and then concatenates it with the output result of the fourth layer of the encoding path to obtain a feature map of 64 x 64 with 1024 channels; The second layer of the decoding path first passes the output result of the first layer of the decoding path through two convolutional blocks with 512 channels and 3x3 kernel size, and then performs transposed convolution to obtain a feature map of 128 x 128 with 256 channels. Subsequently, it is concatenated with the output result of the third layer of the encoding path to obtain a feature map of 128 x 128 with 512 channels; The third layer of the decoding path first passes the output result of the second layer of the decoding path through two convolutional blocks with 256 channels and 3x3 kernel size, and then performs transposed convolution to obtain a feature map of 256 x 256 with 128 channels. Subsequently, it is concatenated with the output result of the third layer of the encoding path to obtain a feature map of 256 x 256 with 256 channels; The fourth layer of the decoding path first passes the output result of the third layer of the decoding path through two convolutional blocks with 128 channels and a size of 3x3, and then performs transposed convolution to obtain a feature map with a size of 512 x 512 and 64 channels. Subsequently, it is concatenated with the output result of the third layer of the encoding path to obtain a feature map with a size of 512 x 512 and 128 channels; Finally, the output result first passes through two convolutional blocks with 64 channels and a size of 3 x 3, and then is input into a 1×1 convolutional layer with a softmax activation function to obtain the probability map of the segmented target in the entire image region.
5. The method for segmenting the boundaries of a scleral lens and a tear lens in an anterior segment OCT image according to claim 4, wherein: After each passing through a convolutional block in the decoding path, the activation function relu is performed.
6. The method for segmenting the boundaries of a scleral lens and a tear lens in an anterior segment OCT image according to claim 1, wherein: The weighted loss function combining focal loss and dice is adopted to solve the class imbalance problem caused by the small proportion of the scleral lens region and the tear lens region compared to the background region.
Citation Information
Patent Citations
A choroidal segmentation method for OCT images based on improved U-net network
CN109509178A
Melanoma segmentation method based on cavity convolution and multi-scale fusion
CN112446890A