Clinical feature-based model training method and image processing method
By introducing the CATLOSs loss function and combining thickness and curvature features, the deep learning network model was optimized, which solved the problem of insufficient segmentation accuracy of the retina and choroid in patients with high myopia, and achieved higher segmentation accuracy and feature preservation.
Patent Information
- Application Number
- CN202511316436.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-09-16
AI Technical Summary
Existing methods for segmenting the retina and choroid are insufficient in terms of segmentation precision and thickness characteristics in patients with high myopia. Traditional methods are time-consuming and difficult to adapt to complex pathological conditions, resulting in a decline in segmentation performance.
A comprehensive loss function, CATLouss, was designed, which combines loss components of thickness and curvature features. The segmentation of the retina and choroid is optimized through a deep learning network model. The model is trained using an OCT image training set, and loss functions of single-layer thickness and multi-layer overall curvature are introduced to improve segmentation accuracy.
In the complex pathological conditions of patients with high myopia, it significantly improves the segmentation accuracy and feature preservation ability of the retina and choroid, and is applicable to ophthalmic image segmentation for different types of people.
Smart Images

Figure CN120823461A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent medical technology, and in particular to a model training method and an image processing method based on clinical characteristics. Background Art
[0002] The morphology and function of the retinal layers are important biomarkers for many ophthalmic diseases. For example, age-related macular degeneration (AMD) can be assessed by the thickness of the retinal pigment epithelium (RPE). Changes in retinal thickness and morphology are closely related to the progression of multiple diseases.
[0003] The thickness of the retina and choroid plays an important role in identifying and monitoring a variety of ophthalmic and neurodegenerative diseases. Traditional segmentation of the retina and choroid involves a two-step process: retinal segmentation and choroidal segmentation. Manual features and topological knowledge are used to calculate features such as thickness and curvature of each segmented layer. Due to the time-consuming process, only a few studies have used small sample sizes. Although deep learning-based methods have achieved good results in retinal segmentation in healthy individuals, patients with high myopia exhibit significant morphological changes in the retina and choroid. High myopia causes localized thinning of the choroid layer, and as myopia progresses to high myopia, the layer structure degenerates, with significant reductions in RPE thickness and blurred interlaminar boundaries. Furthermore, high myopia can lead to localized bulging of the posterior sclera, resulting in stretched and thinned retina and choroid. Accurate segmentation and analysis is a very challenging task. Summary of the Invention
[0004] The embodiments of the present application provide a model training method, an image processing method, and a computing device based on clinical characteristics, which can accurately segment the retinal layer and the choroid layer of eye detection images of different types of people.
[0005] In a first aspect, an embodiment of the present application provides a model training method based on clinical features, the method comprising: constructing multiple training sets; each training set corresponds to different types of ophthalmic images, each training set includes at least a first type of samples, the input data of the first type of samples includes a first image obtained by preprocessing the ophthalmic image, the label data of the first type of samples includes a label image obtained by segmenting multiple target structures in the first image using image segmentation technology, the label image includes a hierarchical representation of each target structure; the multiple target structures include a choroid layer and different types of retinal layers; the different types include normal people, high myopia patients and Alzheimer's patients; using multiple training sets, a deep learning network model is trained to obtain a target model; the loss function of the deep learning network model includes a loss component of the target clinical feature, the characteristic value of the target clinical feature is calculated based on the hierarchical representation, and the target clinical feature includes the thickness of a single target structure and the overall curvature of multiple target structures.
[0006] In this way, a target model that can accurately segment various types of ophthalmic images can be obtained, thereby achieving a higher-precision image segmentation effect.
[0007] In a second aspect, an embodiment of the present application provides an image processing method based on clinical features, the method comprising: inputting an ophthalmic target image into a target model, outputting a segmented image of the ophthalmic target image, and characteristic values of target clinical features; the target model is pre-trained based on a loss function including a loss component of the target clinical features; wherein the segmented image includes a hierarchical representation of each of multiple target structures in the ophthalmic target image, the target clinical features include the thickness of a single target structure, and the overall curvature of multiple target structures; the characteristic values of the target clinical features are calculated based on the hierarchical representation; the multiple target structures include a choroid layer and different types of retinal layers.
[0008] In a third aspect, an embodiment of the present application provides a computing device, including:
[0009] Multiple memories for storing programs;
[0010] Multiple processors are used to execute programs stored in the memory. When the programs stored in the memory are executed, the processors are used to execute the method described in the first aspect or any possible implementation of the first aspect, or the implementation of the second aspect.
[0011] In a fourth aspect, an embodiment of the present application provides a computer storage medium, in which instructions are stored. When the instructions are executed on a computer, the computer executes the method described in the first aspect or any possible implementation of the first aspect, or the implementation of the second aspect.
[0012] In a fifth aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the method described in the first aspect or any possible implementation of the first aspect, or the implementation of the second aspect.
[0013] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0015] Figure 1 A schematic diagram of a model training based on clinical characteristics provided in an embodiment of the present application; Figure 2 Flowchart of the model training method based on clinical characteristics provided in the embodiment of the present application; Figure 3 This is an example diagram of the segmentation of the retina and choroid in the OCT image of a patient with high myopia provided in an embodiment of the present application; Figure 4 This is an example diagram of the segmentation of the retina and choroid in OCT images under different training sets provided in the embodiments of the present application; Figure 5 A comparison chart of the segmentation results using CATLoss and traditional loss functions provided in the embodiment of this application; Figure 6 A schematic diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.
[0017] In the description of the embodiments of this application, any embodiment or design scheme using "exemplary," "for example," or "for example" should not be understood as being more preferred or advantageous than other embodiments or designs. Rather, the use of words such as "exemplary," "for example," or "for example" is intended to present the relevant concepts in a concrete manner.
[0018] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly identifying the technical features being referred to. Thus, a feature specified as "first" or "second" may explicitly or implicitly include one or more of such features. The terms "include," "comprising," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized. "Multiple" can mean one or more, where "multiple" means two or more.
[0019] For example, optical coherence tomography (OCT) images are used in ophthalmology to diagnose and monitor retinal diseases such as AMD, diabetic retinopathy, and glaucoma. OCT images are an imaging technology used to obtain high-resolution cross-sectional images of biological tissues and are particularly suitable for the field of ophthalmology.
[0020] In OCT images, different retinal or choroidal layers are typically marked with colored lines, such as the nerve fiber layer (NFL), retinal pigment epithelium (RPE), and inner nuclear layer (INL). OCT images display the cross-sectional structure of the retina and choroid, allowing observation of the thickness, shape, and possible abnormalities of each layer, such as edema, atrophy, or thickening. By analyzing structural changes in these layers, physicians can assess disease progression and treatment effectiveness.
[0021] OCT images provide detailed layered information of the retina and choroid through color coding and legends, which is very important for the diagnosis and research of ophthalmic diseases.
[0022] Existing retinal segmentation methods typically rely on traditional image segmentation techniques. However, the severe elongation of the axial length and thinning of the retinal layers in patients with high myopia significantly reduce the segmentation performance of these methods in this situation, resulting in insufficient segmentation accuracy and thickness feature accuracy. Therefore, it is necessary to create an automated tool to segment the choroid and calculate indicators to help ophthalmologists more objectively evaluate choroidal pathological changes and assist researchers in conducting studies with larger sample sizes.
[0023] The embodiment of the present application designs a comprehensive loss function for multi-layer segmentation of choroid and retina in ophthalmic images. By explicitly introducing the loss components of target clinical features such as single-layer thickness and multi-layer overall curvature, the shortcomings of traditional segmentation algorithms on complex pathological images are optimized. This method incorporates the upper and lower boundary distance errors of the thickness feature and the geometric deviation of the curvature feature into the loss calculation, and combines it with the traditional pixel-level cross entropy loss to achieve higher segmentation accuracy and feature preservation. It is particularly suitable for ophthalmic image segmentation of patients with high myopia, and effectively improves the accuracy of multi-layer segmentation in complex pathological conditions.
[0024] For example, Figure 1 A schematic diagram of the model training based on clinical characteristics provided in an embodiment of the present application is shown in FIG.
[0025] like Figure 1 As shown in the figure, the loss function CATLoss of the deep learning network model (such as the UNet model) includes Thickness Loss and Curvature Loss, where Thickness Loss is used to represent the loss component of the single-layer thickness of the choroid and retina in the OCT image, and Curvature Loss is used to represent the loss component of the multi-layer overall curvature of the choroid and retina in the OCT image.
[0026] Before model training, a training set is constructed. The samples in the training set include OCT images and their corresponding true label images.
[0027] Figure 1 The upper left OCT image shows the internal structure of the choroid and retina. Figure 1 In the lower left ground truth image, you can see the choroid layer and different types of retinal layers, which are displayed in different grayscales in the image.
[0028] During the model training phase, sample data is fed into the deep learning network model. The segmentation network within the deep learning network propagates the input data layer by layer, from the input layer to the output layer. Each layer performs specific transformations on the data until the final layer generates output results. These results include segmentation maps of the OCT image, as well as tabular information showing single-layer thickness and multi-layer overall curvature.
[0029] Figure 1 The upper right image shows the segmentation map output by the model, which shows the thickness changes of the choroid layer and different types of retinal layers, which can be used to analyze the changes in layer thickness at different locations.
[0030] Figure 1 The lower right table lists the thickness and global curvature measurements of each layer output by the model. The layers in the table include: retinal nerve fiber layer (RNFL), ganglion cell layer (GCL), inner nuclear layer (INL), outer plexiform layer (OPL), outer segment (OS), anterior retinal layer (PRE), middle zone (MEZ), and choroid (CHOROID). The table also lists the thickness and global curvature values of each layer.
[0031] Next, the model compares the segmentation map with the true label map and calculates the value of the loss function. This value reflects the current performance level of the model. The smaller the loss, the closer the model's prediction result is to the true value.
[0032] Figure 1The upper middle image shows the segmentation results for the choroid and different retinal layers. The results include color-coded retinal layers, such as the retinal nerve fiber layer (RNFL) and retinal pigment epithelium (RPE). This color coding helps distinguish and analyze the structure of each layer and is used to calculate the thickness loss component of each layer.
[0033] Figure 1 The lower middle figure shows an overall view of the choroid layer and different types of retinal layers. This view provides a more intuitive representation of the loss component Curvature Loss used to calculate the multi-layer global curvature.
[0034] The model then calculates the gradient of the loss function with respect to each layer's parameters, layer by layer, from the output layer to the input layer, to determine how to adjust the parameters to reduce the loss. The gradient shows the rate of change of the loss function in parameter space and is used to guide parameter updates.
[0035] Then, an optimization algorithm (such as gradient descent, Adam algorithm, etc.) is used to update the model parameters based on the calculated gradients. The optimization algorithm determines the strategy for updating the parameters to minimize the loss function.
[0036] Repeat the above process of forward propagation, loss calculation, backpropagation and parameter update until the performance of the model reaches the expected level or the preset number of iterations is completed to obtain the target model.
[0037] Therefore, the deep learning network model uses the loss function designed above to learn, and continuously adjusts the model parameters through iterative optimization to minimize the loss function, thereby improving the prediction accuracy of the model.
[0038] Based on the above content, the model training method and image processing method based on clinical features proposed in this application are introduced in detail.
[0039] For example, Figure 2 The flowchart of the model training method based on clinical characteristics provided in the embodiment of the present application is shown. The method mainly includes the following execution steps:
[0040] Step S301: Construct multiple training sets. Each training set corresponds to a different type of ophthalmic image. Each training set includes at least a first type of sample. The input data for the first type of sample includes a first image obtained by preprocessing the ophthalmic image. The label data for the first type of sample includes a label image obtained by segmenting multiple target structures in the first image using image segmentation techniques. The label image includes a hierarchical representation of each target structure. The multiple target structures include the choroid layer and different types of retinal layers. The different types include normal people, patients with high myopia, and patients with Alzheimer's disease.
[0041] In one embodiment, ophthalmic images of different types of people are collected. The ophthalmic images may be OCT images, fundus photography, or fluorescein fundus angiography images. Ophthalmic images include multiple target structures, including the choroid layer and different types of retinal layers. Taking OCT images as an example, the choroid layer and different types of retinal layers are included. Different types of people include normal people, patients with high myopia, and patients with Alzheimer's disease. Ophthalmic images of different types of people can reflect the differences in the structure and thickness of the choroid and retina of different people. These differences can be used to train the model to improve the model's segmentation ability of the choroid layer and different types of retinal layers.
[0042] Each training set includes at least first-category samples. The input data for these samples includes a first image obtained by preprocessing an ophthalmic image. The label data for these samples includes a label image obtained by segmenting multiple target structures in the first image using image segmentation techniques. The label image includes a hierarchical representation of each target structure.
[0043] Preprocessing involves scaling ophthalmic images to a target size, ensuring that the images fed into the model have a uniform size. The target size is the desired size of the image after scaling. This size is typically a specific pixel value pair, such as 256x256, indicating that the scaled image will have a width of 256 pixels and a height of 256 pixels. By choosing an appropriate scaling algorithm for the scaling operation, such as bilinear or bicubic interpolation, image quality can be maintained during resizing. Scaling the images can meet the model's input data size requirements and improve the computational efficiency of model training while reducing memory usage.
[0044] Preprocessing also involves randomly performing at least one of the following target operations on the scaled ophthalmic images: horizontal flipping, rotation ±10°, cropping ±15%, and axial scaling ±10%, among other data augmentation operations. These data augmentation methods can generate image variants of varying sizes or shapes, thereby increasing the diversity of the training data. This helps improve the model's generalization capabilities, enabling it to better recognize and classify images of varying sizes, shapes, angles, and scales. Data augmentation can also be combined with other techniques, such as color transformation and noise addition, to further increase data diversity.
[0045] Different target structures in the first image obtained by preprocessing the ophthalmic image can be segmented manually or using image segmentation techniques (such as threshold-based or region-based methods, or deep learning-based methods) to obtain a label image corresponding to the first image. The label image includes a layered representation of the choroid and different types of retinas.
[0046] For example, the hierarchical representation can be Figure 1 The grayscale label shown in the real label map represents different layers with different grayscale values.
[0047] For example, there is an OCT image of ophthalmology. The goal of the segmentation task is to distinguish the multiple layers of the choroid and retina. The segmentation results may include the following layered representations:
[0048] Background: The grayscale value is 0.
[0049] Retinal nerve fiber layer: grayscale value is 1.
[0050] Rod and cone layer: grayscale value is 2.
[0051] Choroid layer: grayscale value is 3.
[0052] This grayscale label image can be directly used to train deep learning models, helping the model learn features at different layers.
[0053] In addition, the hierarchical representation can also be boundaries, contours or pixel-level labels, etc., which will not be repeated here.
[0054] Each training set may also include a second type of sample. The input data for the second type of sample consists of a second image obtained by randomly cropping the first image. The label data for the second type of sample consists of a label image obtained by segmenting multiple target structures in the second image using image segmentation techniques. The proportion of the second type of sample in the training set can range from 15% to 25%, for example, 20%.
[0055] For example, randomly cropping an image to retain 50% of the original image area involves randomly selecting a region from the first image that is half the size of the entire image and then cropping this region as the second image. This operation can help improve the model's adaptability and generalization capabilities to changes in image size.
[0056] For each training set, it can also include a third type of sample. The input data of the third type of sample includes a third image obtained by randomly distorting the first image according to a preset intensity. The label data of the third type of sample includes a label image obtained by segmenting multiple target structures in the third image using image segmentation technology. The proportion of the third type of sample in the training set is 5% to 15%, for example, it can be 10%.
[0057] For example, a 16x16 grid is used with a distortion strength of 2 to randomly transform the first image to obtain the third image. This transformation can increase the robustness of the model, enabling it to better generalize to different image deformations.
[0058] Step S302: Train the deep learning network model using multiple training sets to obtain a target model. The loss function of the deep learning network model includes a loss component for the target clinical features, the eigenvalues of which are calculated based on the hierarchical representation. The target clinical features include the thickness of a single target structure and the overall curvature of multiple target structures.
[0059] In one embodiment, a loss function is designed for a deep learning network model to be trained, the loss function including a loss component for target clinical features, including the thickness of a single choroid or retinal layer, and the overall curvature of the choroid and retinal layers.
[0060] The deep learning network model, based on an encoder-decoder architecture, is used to extract features of the retina and choroid and generate segmentation results. The encoder extracts multi-scale features, while the decoder gradually restores these features to their original resolution to produce an accurate segmentation image. Skip connections are used to fuse information between the encoder and decoder, improving the accuracy of segmentation details.
[0061] For example, the loss function is expressed as:
[0062] (1)
[0063] In formula (1), is the loss function, The cross entropy loss of any pixel in the sample input image (such as the first image) is, is the thickness loss of a single layer, is the overall curvature loss of the multilayer, is the hyperparameter of the sum of the thickness losses for each single layer, is a hyperparameter of the overall curvature loss. α and β are used to balance the contribution of thickness and curvature loss, respectively, and a 1:1 weight ratio can be used.
[0064] Specifically, the thickness loss function The thickness of the different layers obtained by segmentation is directly encoded into the end-to-end loss function as a partial weight in the overall loss function. The thickness is the distance between the upper and lower boundaries of each layer. For each pair of adjacent boundaries, the average distance between their pixels is calculated and compared with the true thickness. This distance is used as one of the weights in the loss function during model training and is adjusted in real time to ensure the thickness of the segmented choroid or retina layer is accurate.
[0065] Curvature loss function The curvature loss function is used to measure the accuracy of the overall curvature of multiple layers, ensuring that the segmented layers maintain their normal geometric shape. By explicitly incorporating curvature features into the loss function, the model can learn and optimize the curvature during training, which is particularly important for accurately capturing pathological changes in patients with high myopia.
[0066] Overall loss function :Combining the thickness of each layer and the overall curvature loss, as well as traditional pixel-level loss functions (such as cross entropy loss ) to improve the overall segmentation effect. The overall loss function aims to strike a balance between pixel-level accuracy and feature preservation, making the segmentation results both accurate and clinically meaningful.
[0067] For example, since the segmentation of the choroid and retinal layers is a multi-class segmentation task, the final segmentation result will consist of N+1 regions, denoted as ,in Represents the background of the image, and the remaining areas correspond to the various retinal layers and choroid layers. A single layer thickness loss can be used to train the model.
[0068] This thickness loss is used as a constraint during model training to maintain the true anatomical relationship between layers. Related research has shown that incorporating specific clinical features of medical images into model training can significantly improve the medical usability of segmentation results.
[0069] The thickness loss of a single layer is expressed as:
[0070] (2)
[0071] In formula (2), is the predicted image output by the deep learning network model for the input image, For label images. is a constant, which is introduced to prevent equation (2) from performing division by zero. It is usually set according to experience. .
[0072] It can be further expressed as:
[0073] (3)
[0074] In formula (3), is the thickness loss of the mth layer, is the hierarchical representation of the mth layer in the predicted image output by the deep learning network model for the input image, is the hierarchical representation of the mth layer in the true label image, 1≤m≤N, m∈Z, and N is the number of multiple layers.
[0075] Based on The calculated thickness of the mth layer, Based on The calculated thickness of the mth layer, (x, y) is the coordinate of any pixel in the input image, For any pixel to be segmented The probability in For any pixel to be segmented The probability in or The value of is 0 or 1, and there is only one or is 1.
[0076] For example, a curvature loss function is designed , which measures the global curvature characteristics of multiple layers to ensure that the segmented layers maintain their normal geometry. The global curvature is converted from the coordinates of the segmented image. Global curvature is important in determining how light is focused on the retina, which affects vision and perception. Understanding this curvature is crucial for diagnosing and treating various eye diseases, including retinal detachment, age-related macular degeneration, and diabetic retinopathy. Abnormal curvature can lead to refractive errors such as myopia or hyperopia. It also plays a role in the development and fitting of contact lenses and in the planning of refractive surgery.
[0077] This curvature can be described mathematically using differential geometry, which is to calculate the curvature of a curve (or surface) at a given point. For a curve on a plane, the curvature at a point is defined as the inverse of the radius of the osculating circle (R): .
[0078] For the segmented image, its curvature It can be expressed as:
[0079] (4)
[0080] In formula (4), the function is the fitting of multiple layered contours, and represent the first and second order derivatives respectively.
[0081] By explicitly incorporating curvature features into the loss function, the model is able to learn and optimize the curvature of the choroid and retina during training, which is particularly important for accurately capturing pathological changes, especially in patients with high myopia.
[0082] Furthermore, the overall curvature loss is expressed as:
[0083] (5)
[0084] In formula (5), Based on The calculated overall curvature is Based on The calculated overall curvature, M, is or The marginal fitting function The number of valid points on the x-axis.
[0085] Among them, the predicted curvature It can be expressed as:
[0086] (6)
[0087] True curvature It can be expressed as:
[0088] (7)
[0089] For example, the cross entropy loss of any pixel (x, y) in the input image is It can be expressed as:
[0090] (8)
[0091] In formula (8), For any pixel (x,y) to be segmented into The probability of the jth layer in is the predicted probability. is the probability that any pixel (x, y) is segmented into the jth layer in G, that is, the true probability. N is the number of layers. or The value of is 0 or 1, and there is only one or is 1.
[0092] Therefore, the overall loss function shown in formula (1) can be designed for the deep learning network model .
[0093] After designing the loss function for the deep learning network model based on the above, the deep learning network model is trained using multiple training sets to obtain the target model.
[0094] Exemplarily, each training set is input into the deep learning network model for training in turn, or the order of samples in multiple training sets is shuffled and then input into the deep learning network model for training to obtain the target model.
[0095] Exemplary training and optimization procedures include: Using a diverse OCT image training set consisting of choroidal and retinal images from healthy individuals, patients with high myopia, and patients with Alzheimer's disease. The Adam optimizer is used with a learning rate of 0.0001 and a batch size of 16. Cross-validation is used to assess model robustness, and data augmentation techniques (such as random flipping and rotation) are used to increase data diversity.
[0096] Specifically, the thickness is the distance between the upper and lower boundaries of each layer. For any sample data, the input image is input to the model, and the segmentation network of the model is used to segment it and predict the thickness of the mth layer in S. and the overall curvature of S .
[0097] First, the segmentation network is used to predict the probability that any pixel (x, y) in the input image belongs to the mth layer , thereby predicting the boundary positions of each layer in S.
[0098] Secondly, for the mth layer, calculate the absolute value of the y coordinates of the two pixels with the same x coordinates on its upper and lower boundaries, and then average the absolute values corresponding to all x coordinates of the layer to get the thickness of the layer .
[0099] Similarly, based on the real label map Calculated and .
[0100] Then use formula (3) to calculate the thickness loss of the mth layer.
[0101] Furthermore, S is converted into a binary image , apply Canny edge detection to the binary image B to obtain edge information , using a quadratic polynomial Perform curve fitting to obtain the fitting edge . Use formula (6) to calculate the fitting edge The curvature of each x in .
[0102] Similarly, based on the real label map The curvature is calculated using formula (7): .
[0103] Then use formula (5) to calculate the overall curvature loss.
[0104] Furthermore, the cross entropy loss is calculated using formula (8): .
[0105] Finally, the overall loss is calculated using formula (1) .
[0106] Therefore, by adding thickness and curvature losses to the loss function, the model not only performs segmentation at the pixel level, but also retains important structural information, especially for complex images of patients with high myopia, ensuring segmentation accuracy and clinical significance.
[0107] While curvature loss primarily aims to improve segmentation results by focusing on the overall curvature of multiple layers in the segmented image, it may not accurately capture errors caused by spatial displacement or other geometric deformations in the image. Therefore, combining it with other loss functions (such as cross-entropy loss that can capture displacement or shape changes) may be helpful. The CATLoss comprehensive loss function of this invention strikes a balance between pixel-level accuracy and feature preservation, resulting in segmentation results that are both accurate and clinically meaningful.
[0108] In summary, the embodiment of the present application designs a comprehensive loss function for multi-layer segmentation of choroid and retina in ophthalmic images. By explicitly introducing the loss components of target clinical features such as single-layer thickness and multi-layer overall curvature, the shortcomings of traditional segmentation algorithms on complex pathological images are optimized. This method incorporates the upper and lower boundary distance errors of the thickness feature and the geometric deviation of the curvature feature into the loss calculation, and combines it with the traditional pixel-level cross entropy loss to achieve higher segmentation accuracy and feature preservation. It is particularly suitable for OCT image segmentation of patients with high myopia, and effectively improves the accuracy of multi-layer segmentation in complex pathological conditions.
[0109] For example, an embodiment of the present application provides a clinical feature-based image processing method, comprising: inputting an ophthalmic target image into a target model, and outputting a segmented image of the ophthalmic target image and a feature value of a target clinical feature. The target model is pre-trained based on a loss function that includes a loss component for the target clinical feature.
[0110] For example, the ophthalmic target image is an OCT image, which is the input image of the target model. The target model is based on Figure 2 The method shown is pre-trained.
[0111] The segmented image includes a hierarchical representation of each of the multiple target structures in the ophthalmic target image. The target clinical features include the thickness of individual target structures and the overall curvature of the multiple target structures. Feature values of the target clinical features are calculated based on the hierarchical representation. The multiple target structures include the choroid layer and different retinal layers.
[0112] For example, the feature values of the segmented image and the target clinical feature are as follows: Figure 1 Model output shown. Multiple target structures include the choroid layer and different types of retinal layers.
[0113] For example, Figure 3An example diagram of the segmentation of the retina and choroid in an OCT image of a high myopia patient provided in an embodiment of the present application is shown.
[0114] Table 1 shows the performance comparison of different loss functions on the choroid and retinal layers of the UNet model in the Alzheimer's disease patient training set. Performance parameters include precision (PRE), recall, intersection over union (IoU), and DICE coefficient.
[0115] Table 1
[0116] Figure 3 Table 1 shows the segmentation results and specific parameter performance of the Alzheimer's disease patient training set using CATLoss and other loss methods. CATLoss achieved the best precision of 0.92, 0.892, 0.897, and 0.884 at the CHR, GCIPL, RNFL, and OS layers, respectively. In terms of recall, CATLoss achieved the highest scores of 0.92, 0.911, 0.911, and 0.884 at the CHR, GCIPL, RNFL, and OS layers, respectively. Regarding the Dice coefficient, CATLoss achieved maximum values of 0.92, 0.901, 0.858, and 0.882 at the CHR, GCIPL, RNFL, and OS layers, respectively. Regarding the IoU metric, CATLoss performed best, achieving scores of 0.834, 0.826, 0.838, and 0.715 at the choroid, GCIPL, RNFL, and OS layers, respectively.
[0117] Figure 3 The meaning of the relevant English letters:
[0118] IMAGE: ophthalmological images,
[0119] GT: Real group,
[0120] DICE: Dice coefficient loss,
[0121] CAT-LOSS: clinically characterized enhanced thickness and curvature loss,
[0122] CE: Cross Entropy Loss.
[0123] It can be found that CATLoss performs best in segmentation under high curvature retina.
[0124] For example, Figure 4 Shown are example diagrams of segmentation of the retina and choroid in OCT images under different training sets provided in the embodiments of the present application.
[0125] Table 2 shows the comparison of retinal segmentation performance of different loss functions (CATLoss, NLL-Loss, DICE-Loss, CE-Loss) under different training sets (normal people AD, high myopia patients CN, Alzheimer's disease patients HM).
[0126] Table 2
[0127] Figure 4 Table 2 shows the retinal layer segmentation results across three training sets using the UNet architecture as the backbone.
[0128] The CATLoss function demonstrated robustness and consistency across all training sets, achieving significantly high precision, recall, and Dice scores in both inner and outer retinal layers. Specifically, in the AD training set, CATLoss achieved Dice scores of 0.913 and 0.907 for HFL and CHR, respectively, indicating comparable balanced performance to NLL-Loss. While NLL-Loss achieved a similar Dice score of 0.928 in HFL, its performance in other retinal layers was slightly lower. In the CN training set, CATLoss demonstrated excellent performance, achieving the highest Dice scores of 0.933 and 0.91 in the CHR and GCIPL layers, respectively, thereby surpassing the performance of Dice-Loss and CE-Loss, which exhibited lower Dice values in these specific layers.
[0129] Figure 4 The meaning of the relevant English letters: EN: Normal people, HM: Patients with high myopia, AD: Alzheimer's disease patients, GT: Real group, DICE: Dice coefficient loss, CAT-LOSS: clinically characterized enhanced thickness and curvature loss, CE: Cross Entropy Loss,
[0130] NLL: NLL loss.
[0131] In the HM training set, CATLoss achieves Dice scores of 0.914 and 0.873 on the CHR and OS layers, respectively, which are comparable to Dice-Loss, especially the score of 0.914 on the CHR layer, although Dice-Loss generally performs worse on inner layers (such as INL).
[0132] CATLoss consistently provides high accuracy and balanced segmentation across different layers and training sets, outperforming NLL-Loss, Dice-Loss, and CE-LOSS in terms of overall adaptability and accuracy, especially in the more complex outer layers.
[0133] It can be found that CATLoss performs best in complex pathological conditions.
[0134] For example, Figure 5 A comparison chart of the segmentation results using CATLoss and traditional loss functions provided in an embodiment of the present application is shown.
[0135] Figure 5 The meaning of the relevant English letters: CAT-Loss: CAT loss, Loss: loss, HM: Patients with high myopia, GCL: Ganglion cell layer, CE-GCL: CE loss in the ganglion cell layer, DICE-GCL: DICE loss in the ganglion cell layer, CAT-GCL: CAT loss in the ganglion cell layer, HFL: Henle fiber layer, CE-HFL: CE loss of Henle fiber layer, DICE-HFL: DICE loss in the fiber layer of Henle, CAT-HFL: CAT loss in the fiber layer of Henle, OS: (photoreceptor) outer segment, CE-OS: CE loss of the outer segment (photoreceptor), DICE-OS: DICE loss in the outer segments of photoreceptors. CAT-OS: CAT loss in the outer segments of photoreceptors. CHOROID: choroid, CE-CHOROID: CE loss of the choroid, DICE-CHOROID: DICE loss of the choroid, CAT-CHOROID: Loss of CAT in the choroid.
[0136] It can be found that compared with traditional loss functions, CATLoss performs best in feature preservation.
[0137] It is understandable that the size of the sequence number of each step in the above embodiment does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. In addition, in some possible implementations, the steps in the above embodiment can be selectively executed according to actual conditions, and can be partially executed or fully executed, which is not limited here. In addition, all or part of any features in the above embodiment can be freely and arbitrarily combined without contradiction. The combined technical solution is also within the scope of this application.
[0138] For example, the present application embodiment further provides a computing device 1000. Figure 6 As shown, computing device 1000 includes a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. Processor 1004, memory 1006, and communication interface 1008 communicate with each other via bus 1002. It should be understood that the present application does not limit the number of processors and memories in computing device 1000.
[0139] The bus 1002 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 The bus 1004 is shown as a single line, but does not indicate that there is only one bus or only one type of bus. The bus 1004 may include a path for transmitting information between various components of the computing device 1000 (eg, the memory 1006, the processor 1004, and the communication interface 1008).
[0140] The processor 1004 may include any one or more processors such as a central processing unit, a graphics processing unit (GPU), a microprocessor (MP), a digital signal processor (DSP), a baseboard management controller, and the like.
[0141] The memory 1006 may include volatile memory, such as random access memory (RAM). The processor 1004 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0142] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1000 or a cluster of multiple computing devices 1000 and other devices or communication networks.
[0143] The computing device 1000 includes a network device internally, or the computing device 1000 is connected to multiple network devices externally. The internal network device communicates with the processor 1004, the memory 1006 and the communication interface 1008 through the bus 1002, and the external network device communicates with the computing device 1000 through interfaces such as Ethernet, Fibre Channel, and InfiniBand.
[0144] The memory 1006 stores executable program codes / instructions, and the processor 1004 executes the executable program codes / instructions to implement Figure 2 In other words, the memory 1006 stores a program / instruction for executing all or part of the steps in the method of the above embodiment.
[0145] An embodiment of the present application provides a computing device, including: a memory and a processor; the memory and the processor are coupled; the memory is used to store programs; the processor is used to execute the programs stored in the memory, and when the programs stored in the memory are executed, the processor is used to execute the methods in the above embodiments.
[0146] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.
[0147] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the method in the above embodiment.
[0148] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an ASIC.
[0149] In the above embodiments, all or part of the embodiments can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that integrates one or more available media. Available media can include magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
[0150] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.
Claims
1. A model training method based on clinical characteristics, characterized in that: The method comprises: Constructing multiple training sets; each training set corresponds to a different type of ophthalmic image, and each training set includes at least a first type of sample, wherein the input data of the first type of sample includes a first image obtained by preprocessing the ophthalmic image, and the label data of the first type of sample includes a label image obtained by segmenting multiple target structures in the first image using image segmentation technology, and the label image includes a hierarchical representation of each target structure; the multiple target structures include a choroid layer and different types of retinal layers; the different types include normal people, patients with high myopia, and patients with Alzheimer's disease; The deep learning network model is trained using the multiple training sets to obtain a target model; the loss function of the deep learning network model includes a loss component of the target clinical feature, the characteristic value of the target clinical feature is calculated based on the hierarchical representation, and the target clinical feature includes the thickness of a single target structure and the overall curvature of multiple target structures.
2. The method according to claim 1, characterized in that The loss function includes the thickness loss of a single target structure and the overall curvature loss of multiple target structures; The loss function is expressed as follows: , in, is the loss function, is the cross entropy loss of any pixel in the first image, is the thickness loss of the single target structure, is the overall curvature loss, is a hyperparameter of the sum of thickness losses for each target structure, is the hyperparameter of the overall curvature loss.
3. The method according to claim 2, characterized in that The thickness loss of the single target structure is expressed as: , in, is the thickness loss of the mth target structure, is a hierarchical representation of the mth target structure in the predicted image output by the deep learning network model for the first image, is the hierarchical representation of the mth target structure in the label image, 1≤m≤N, m∈Z, N is the number of multiple target structures; Based on The calculated thickness of the mth target structure, Based on The calculated thickness of the mth target structure, (x, y) is the coordinate of any pixel in the first image, For any pixel to be segmented into The probability in For any pixel to be segmented into The probability in and The value of is 0 or 1, and there is only one or is 1, is a constant.
4. The method according to claim 2, characterized in that The overall curvature loss is expressed as: , in, Based on The overall curvature calculated is, Based on The overall curvature calculated is, is a predicted image output by the deep learning network model for the first image, is the label image, M is or The marginal fitting function The number of valid points on the x-axis.
5. The method according to claim 1, wherein The training set also includes a second type of samples, the input data of the second type of samples includes a second image obtained by randomly cropping the first image, and the label data of the second type of samples includes a label image obtained by segmenting multiple target structures in the second image using image segmentation technology; The proportion of the second category of samples in the training set is 15% to 25%.
6. The method according to claim 1, characterized in that The training set also includes a third type of samples, the input data of the third type of samples including a third image obtained by randomly distorting the first image according to a preset intensity, and the label data of the third type of samples including a label image obtained by segmenting multiple target structures in the third image using an image segmentation technique; The proportion of the third category of samples in the training set is 5% to 15%.
7. The method according to claim 1, characterized in that The method of training the deep learning network model using multiple training sets includes: Input each training set into the deep learning network model for training in turn; or The samples in multiple training sets are shuffled in order and then input into the deep learning network model for training.
8. The method according to claim 1, characterized in that The preprocessing includes scaling the ophthalmic image to a target size, and randomly performing at least one of the following target operations on the scaled ophthalmic image, wherein the target operations include: horizontal flipping, rotation of ±10°, cropping of ±15%, and axial scaling of ±10%.
9. An image processing method based on clinical features, characterized in that: The method comprises: Inputting an ophthalmic target image into a target model and outputting a segmented image of the ophthalmic target image and a feature value of a target clinical feature; the target model is pre-trained based on a loss function including a loss component of the target clinical feature; The segmented image includes a hierarchical representation of each of the multiple target structures in the ophthalmic target image, the target clinical feature includes the thickness of a single target structure and the overall curvature of the multiple target structures; the characteristic value of the target clinical feature is calculated based on the hierarchical representation; the multiple target structures include a choroid layer and different types of retinal layers.
10. A computing device, characterized in that include: Multiple memories for storing programs; Multiple processors are used to execute the program stored in the memory. When the program stored in the memory is executed, the processors are used to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Vascular feature-based retinal vessel segmentation method and device, and storage medium
CN113379741A
Retinal exudate full-automatic segmentation method based on curvature loss function
CN114170150A
Sclera lens OCT (Optical Coherence Tomography) image tear layer segmentation model, method and equipment based on deep learning
CN117392156A
Method and system for calculating gray level, hollowing degree and thickness of cell wall
CN120451973A
System and Method for Early Detection of Diabetic Retinopathy Using Optical Coherence Tomography
US20110275931A1