Eye image intelligent analysis method and system and electronic equipment
By combining the lightweight semantic segmentation network KiU-Net2D with a pupil localization and diagnosis model of at least two levels of ConvNeXt, the problems of image segmentation accuracy and manual diagnosis in cataract diagnosis are solved, and high-precision, automated graded diagnosis of ocular lesions is achieved.
Patent Information
- Application Number
- CN202511383675.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2025-10-31
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing cataract diagnosis methods, image segmentation models have poor segmentation accuracy and cannot effectively handle irrelevant noise such as light spots. Furthermore, manual diagnosis requires a high level of experience, resulting in low diagnostic efficiency and poor accuracy.
A pupil localization model is established using the lightweight semantic segmentation network KiU-Net2D, and a diagnostic model is established by combining it with at least two levels of ConvNeXt model. Automatic hierarchical diagnosis is achieved through the coupling of the pupil localization model and the diagnostic model, thereby improving the accuracy of pupil boundary localization and diagnosis.
It improves the precision of pupil region image cropping and the accuracy of diagnostic results, realizes an end-to-end automated diagnostic process from image input to hierarchical diagnosis, enhances diagnostic efficiency and accuracy, and has high clinical practical value.
Smart Images

Figure CN120877359A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of eye disease treatment technology, and in particular to an intelligent image analysis method, system and electronic device for the eye. Background Technology
[0002] Cataracts are a disease in which the lens of the eye becomes cloudy, leading to visual impairment. Clinically, the visual observation of cataracts mainly relies on symptoms and externally visible signs, including blurred vision, abnormal pupil color (grayish-white or pale yellow), and light sensitivity. Cataract diagnosis can also be achieved using ocular image analysis techniques.
[0003] In existing technologies, cataract diagnosis methods based on ocular image analysis typically employ optical coherence tomography (OCT) images or slit-lamp images. Image segmentation models are used to segment the OCT or slit-lamp images, and then doctors manually diagnose the extent of cataract lesions by combining the segmented image features (e.g., abnormal pupil color or degree of lens opacity). However, existing technologies have the following problems: Firstly, existing image segmentation models have poor segmentation accuracy for OCT images of the eye and cannot effectively handle irrelevant noise such as light spots, resulting in low image segmentation accuracy and affecting the accuracy of analysis results. Secondly, manual diagnosis requires a high level of experience. Manual diagnosis necessitates operators accurately understanding the characteristics of different lesion severity levels, such as the lesion area, location, and density. This results in low diagnostic efficiency and a high risk of misdiagnosis. Summary of the Invention
[0004] To address the aforementioned issues, this invention provides an intelligent image analysis method, system, and electronic device for the eye. By coupling a pupil localization model and a diagnostic model, it achieves automatic hierarchical diagnosis of eye lesions, solving the problems of poor segmentation accuracy of existing cataract diagnosis methods, high experience requirements for human diagnosis, low diagnostic efficiency, and poor diagnostic accuracy.
[0005] According to one aspect of the present invention, an intelligent analysis method for eye images is provided, comprising: acquiring an eye image and performing image preprocessing on the eye image; acquiring a pre-trained pupil localization model and a diagnostic model; wherein the pupil localization model is established based on a KiU-Net2D segmentation model; the diagnostic model is established based on at least two levels of a ConvNeXt model; extracting pupil region images from the eye image based on the pupil localization model; classifying and diagnosing the pupil region images based on the diagnostic model, and outputting diagnostic results for eye lesions.
[0006] Optionally, the diagnostic model includes a first diagnostic sub-model and a second diagnostic sub-model, wherein the lesion level output by the first diagnostic sub-model is lower than the lesion level output by the second diagnostic sub-model; the step of classifying and diagnosing the pupil region image based on the diagnostic model and outputting the ocular lesion diagnosis result includes: importing the pupil region image into the first diagnostic sub-model and obtaining the preliminary classification prediction output by the first diagnostic sub-model; when the preliminary classification prediction is a first category, using the first category as the classification diagnosis result; when the preliminary classification prediction is a second category, importing the pupil region image into the second diagnostic sub-model and performing a secondary diagnosis; wherein the lesion level corresponding to the first category is lower than the lesion level corresponding to the second category.
[0007] Optionally, during the model pre-training stage, the first diagnostic sub-model trains the first ConvNeXt model using the first category and the second category; the second diagnostic sub-model trains the second ConvNeXt model using the subcategories of the second category; wherein the lesion level corresponding to the first category is lower than the lesion level corresponding to any of the subcategories.
[0008] Optionally, the step of extracting the pupil region image from the eye image based on the pupil localization model includes: importing the preprocessed eye image into the pupil localization model and obtaining the probability map output by the pupil localization model; performing threshold segmentation on the probability map and selecting the largest connected component region as the pupil prediction region based on connected component analysis; and cropping the eye image based on the pupil prediction region to obtain the pupil region image.
[0009] Optionally, acquiring the eye image and performing image preprocessing on the eye image includes: performing grayscale conversion and image resampling processing on the eye image.
[0010] Optionally, during the model pre-training stage, the pupil localization model uses an eye image processed by a first data augmentation strategy and its corresponding refined mask to train the KiU-Net2D segmentation model; wherein, the refined mask is a mask established by binarization annotation after coarse segmentation using SAM; the first data augmentation strategy includes at least one of the following: geometric transformation processing, color enhancement processing, spatial enhancement processing, and data normalization processing.
[0011] Optionally, during the model pre-training phase, the diagnostic model uses pupil region images processed by the second data augmentation strategy and their corresponding lesion grading labels to perform at least two stages of grading training on the ConvNeXt model; wherein, the lesion grading labels include at least three lesion levels; the second data augmentation strategy includes at least one of the following: geometric transformation processing, color enhancement processing, spatial enhancement processing, occlusion enhancement processing, noise enhancement processing, and data normalization processing.
[0012] Optionally, during the model pre-training stage, the loss function of the pupil localization model is constructed using a weighted combination of the Dice loss function and the binary cross-entropy loss function; the loss function of the diagnostic model is constructed using the Focal loss function with a class balance factor.
[0013] According to another aspect of the present invention, an intelligent analysis system for eye images is provided, comprising: an image acquisition module for acquiring eye images and performing image preprocessing on the eye images; a model acquisition module for acquiring a pre-trained pupil localization model and a diagnostic model; wherein the pupil localization model is established based on a KiU-Net2D segmentation model; the diagnostic model is established based on at least two levels of a ConvNeXt model; an image segmentation module for extracting pupil region images from the eye images based on the pupil localization model; and a diagnostic module for classifying and diagnosing the pupil region images based on the diagnostic model and outputting diagnostic results for eye lesions.
[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the above-described intelligent eye image analysis method.
[0015] The technical solution of the present invention has the following technical effects: Firstly, a pupil localization model is established by using the lightweight semantic segmentation network KiU-Net2D. It has a dual-path structure. The main branch uses the U-Net structure to extract semantic context information, while the auxiliary branch uses the Ki-Net structure to enhance the edge detail perception capability. The two branches achieve complementary fusion of low-level and high-level features through a bidirectional fusion module, which can improve the localization accuracy of pupil boundaries, improve the cropping accuracy of pupil region images, and reduce the impact of background region images on disease diagnosis results.
[0016] Secondly, a diagnostic model is established by introducing at least two levels of ConvNeXt models. Using the ConvNeXt model as the backbone network, it features large convolutional kernels, an improved residual structure, and a Transformer-style regularization strategy. This effectively captures details such as blurriness, texture corruption, and light spots, achieving both high accuracy and computational efficiency, making it suitable for mobile diagnostic scenarios. Through at least two levels of hierarchical training and diagnostic strategies, the diagnostic model's ability to distinguish between blurred boundary categories (such as the difference between moderate and severe lesions) is improved, thereby enhancing the accuracy of diagnostic results.
[0017] Third, automatic grading diagnosis of ocular lesions is achieved by coupling the pupil localization model and the diagnostic model. The output of the pupil localization model is used as the cropping basis for the input of the diagnostic model. The segmentation accuracy of KiU-Net2D determines the effectiveness and center alignment of the subsequent grading images. The two models are independent during the training phase and form a highly coupled closed-loop system during the application phase. Ultimately, an end-to-end automatic diagnostic process from image input to grading diagnosis output is realized, improving the diagnostic efficiency and accuracy of ocular diseases. It has high clinical practical value and deployability, and solves the problems of poor segmentation accuracy of existing cataract diagnosis methods, high experience requirements for human diagnosis, low diagnostic efficiency, and poor diagnostic accuracy.
[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A flowchart of an intelligent eye image analysis method provided in an embodiment of the present invention; Figure 2 A flowchart of another intelligent eye image analysis method provided in an embodiment of the present invention; Figure 3 A flowchart illustrating another intelligent eye image analysis method provided in this embodiment of the invention; Figure 4 This is a schematic diagram of the probability map output by a pupil localization model provided in an embodiment of the present invention; Figure 5 A schematic diagram of an eye image provided in an embodiment of the present invention; Figure 6 A schematic diagram of an eye image and its rough mask provided for an embodiment of the present invention; Figure 7 A schematic diagram of an eye image and its refined mask provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of an intelligent eye image analysis system provided in an embodiment of the present invention; Figure 9 A schematic diagram of the structure of an electronic device for implementing the intelligent eye image analysis method of this invention. Detailed Implementation
[0021] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0022] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0023] Figure 1 This is a flowchart of an intelligent eye image analysis method provided by an embodiment of the present invention. This embodiment is applicable to the application scenario of graded diagnosis of cataract lesions. The method can be executed by an intelligent eye image analysis system, which can be implemented in hardware and / or software. The intelligent eye image analysis system can be configured in an eye treatment device or other types of electronic devices.
[0024] like Figure 1 As shown, the intelligent eye image analysis method of the present invention includes the following steps: S1: Acquire the eye image and perform image preprocessing on the eye image.
[0025] In this embodiment, the eye image can be understood as a slit-lamp image of the eye of the patient to be diagnosed, acquired using a handheld cataract screening device. Typically, the acquired eye images are stored in .jpg or .png format.
[0026] Optionally, an eye image is acquired, and image preprocessing is performed on the eye image, including grayscale conversion and image resampling. Image resampling can be understood as changing the image resolution through interpolation or downsampling. For example, in the diagnosis and treatment of eye diseases, the slit-lamp image of the eye to be processed can be grayscaled and its size adjusted to a uniform size, such as an image with a resolution of 1024×1024.
[0027] S2: Obtain the pre-trained pupil localization model and diagnostic model.
[0028] The pupil localization model is based on the KiU-Net2D segmentation model; the diagnostic model is based on at least two levels of ConvNeXt models. In this embodiment, ConvNeXt models of different levels can be trained using data sources of different lesion levels.
[0029] Specifically, during the model training phase, the training input data for the pupil localization model consists of pre-processed and labeled slit-lamp images of the eye and masks used to crop the pupil region image. All eye images acquired by the slit-lamp device are in color format. To ensure training consistency and model generalization ability, all eye images acquired by the slit-lamp device undergo the following preprocessing: First, grayscale conversion: Since pupil region segmentation mainly relies on brightness information, converting color images to grayscale reduces color interference and improves the sensitivity of the segmentation model to the pupil region. Second, image resampling: All images and their corresponding pupil masks are adjusted to a uniform size, for example, images with a resolution of 1024×1024, ensuring consistent network input dimensions and facilitating batch training and network structure alignment. During the model training phase, the pre-processed data also needs to be labeled. After coarse segmentation using the Segment Anything Model (SAM), binary labeling is performed by professionals (e.g., labeling the background region as 0 and the pupil region as 1) to supervise model training.
[0030] The training input data for the diagnostic model consists of RGB color images with pre-localized and cropped pupil regions, along with corresponding cataract grading labels annotated by professional ophthalmologists (e.g., divided into three levels: none or mild, moderate, and severe, corresponding to category labels 0, 1, and 2, respectively). All input images are resized to a uniform size, such as 352×352 resolution, to accommodate the input requirements of the deep neural network structure.
[0031] In the diagnosis and treatment of eye diseases, the pre-trained pupil localization model can be used for image cropping, and then the cropped image can be input into the pre-trained diagnostic model, which will then perform a graded diagnosis.
[0032] S3: Extract pupil region images from eye images based on pupil localization models.
[0033] The pupil region image can be understood as an image cropped from the preprocessed eye image based on a mask corresponding to the pupil region.
[0034] S4: Classify and diagnose pupil region images based on the diagnostic model, and output the diagnostic results of eye lesions.
[0035] In this context, the diagnostic results for ocular lesions can be understood as the lesion grade or category labels corresponding to different lesion grades. In this embodiment, ocular lesions include, but are not limited to, cataract lesions.
[0036] Specifically, in the diagnosis and treatment of eye diseases, a single slit-lamp image of the eye, preprocessed (e.g., grayscale conversion and image resampling adjustment), is input into a pre-trained pupil localization model. The output data of the pupil localization model (i.e., the pupil region image obtained by pupil region masking) is used as the input data for the diagnostic model. The high segmentation accuracy of the pupil localization model determines the effectiveness and center alignment of the subsequent graded diagnosis. The pupil localization model and the diagnostic model operate independently during the model pre-training stage, but form a highly coupled closed-loop system during the diagnosis and treatment of eye diseases, ultimately achieving an end-to-end automated diagnostic process from image input to graded diagnostic output.
[0037] Therefore, the technical solution of this invention establishes a pupil localization model by using the lightweight semantic segmentation network KiU-Net2D. It has a dual-path structure. The main branch uses the U-Net structure to extract semantic context information, while the auxiliary branch uses the Ki-Net structure to enhance the edge detail perception capability. The two branches achieve complementary fusion of low-level features and high-level features through a bidirectional fusion module, which can improve the localization accuracy of pupil boundaries, improve the cropping accuracy of pupil region images, and reduce the impact of image background areas on disease diagnosis results.
[0038] A diagnostic model is built by introducing at least two levels of ConvNeXt models. Using the ConvNeXt model as the backbone network, it features large convolutional kernels, an improved residual structure, and a Transformer-style regularization strategy. This effectively captures details such as blurriness, texture corruption, and light spots, achieving both high accuracy and computational efficiency, making it suitable for mobile diagnostic scenarios. Through at least two levels of hierarchical training and diagnostic strategies, the model's ability to distinguish between blurred boundary categories (such as the difference between moderate and severe lesions) is improved, thereby enhancing the accuracy of diagnostic results.
[0039] By coupling the pupil localization model and the diagnostic model, automatic hierarchical diagnosis of ocular lesions is achieved, which improves the diagnostic efficiency and accuracy of ocular diseases. It has high clinical practical value and deployability, and solves the problems of poor segmentation accuracy of the image segmentation model used in existing cataract diagnosis methods, as well as the high experience requirements of human diagnosis, resulting in low diagnostic efficiency and poor diagnostic accuracy.
[0040] In some optional embodiments, the diagnostic model includes a first diagnostic sub-model and a second diagnostic sub-model, wherein the lesion grade output by the first diagnostic sub-model is lower than the lesion grade output by the second diagnostic sub-model. In this embodiment, the first and second diagnostic sub-models are ConvNeXt models trained using pupil region images of different lesion grades.
[0041] Optionally, during the model pre-training phase, the first diagnostic sub-model trains the first ConvNeXt model using the first category and the second category; the second diagnostic sub-model trains the second ConvNeXt model using the subcategories of the second category; wherein, the lesion level corresponding to the first category is lower than the lesion level corresponding to any subcategory.
[0042] For example, taking lesion grades including mild or no lesion, moderate lesion, and severe lesion as an example, a two-stage grading strategy is used for training during the pre-training phase of the diagnostic model: In the first stage of model training, the original three-class classification task is simplified to a two-class classification task. "Mild lesions" and "no lesions" are merged into the first category, and "moderate lesions" and "severe lesions" are merged into the second category. The ConvNeXt model is trained in this stage to obtain a coarse classification model, namely the first diagnostic sub-model, which focuses on identifying the presence of significant lesions. In the second stage of model training, the second category is further subdivided into "moderate lesion" and "severe lesion". The ConvNeXt model is trained using the category labels corresponding to "moderate lesion" and "severe lesion" and pupil area images to obtain a secondary classification model, namely the second diagnostic sub-model. This second diagnostic sub-model is used to identify whether a significant lesion belongs to "severe" or "moderate", and finally achieves a three-class classification diagnosis.
[0043] Figure 2 A flowchart of another intelligent eye image analysis method provided in an embodiment of the present invention is shown. Figure 1 Based on the illustrated embodiment, a specific implementation method for classification diagnosis based on a diagnostic model is shown as an example. See also Figure 2 As shown, in step S4 above, the pupil region image is classified and diagnosed based on the diagnostic model, and the diagnostic results of the eye lesion are output, specifically including: S401: Import the pupil region image into the first diagnostic sub-model and obtain the preliminary classification prediction output by the first diagnostic sub-model.
[0044] S402: When the initial classification prediction is Category I, Category I shall be used as the classification diagnosis result.
[0045] S403: When the initial classification prediction is the second category, import the pupil region image into the second diagnostic sub-model to perform a secondary diagnosis.
[0046] Among them, the lesion grade corresponding to the first category is lower than that corresponding to the second category.
[0047] For example, taking the first category as "mild or no lesion" and the second category as "moderate lesion" and "severe lesion" as an example, in the diagnosis and treatment of eye diseases, the diagnostic model receives the pupil region image automatically cropped and output by the pupil localization model, and inputs the pupil region image into the trained first diagnostic sub-model for preliminary classification. If the preliminary classification prediction result is the first category (mild or no lesion), the current classification label is directly output; if the preliminary classification prediction result is the second category (moderate or severe lesion), the pupil region image is imported into the second diagnostic sub-model for further classification, and the final classification label (moderate or severe) is output. For example, the classification label for mild or no lesion can be recorded as 0, the classification label for moderate lesion can be recorded as 1, and the classification label for severe lesion can be recorded as 2.
[0048] Figure 3 This flowchart illustrates another intelligent eye image analysis method provided in this embodiment of the invention, demonstrating a specific implementation of extracting pupil region images based on a pupil localization model. See also... Figure 3 As shown, when performing step S3 above, the pupil region image is extracted from the eye image based on the pupil localization model, specifically including: S301: Import the preprocessed eye image into the pupil localization model and obtain the probability map output by the pupil localization model.
[0049] The probability map can be understood as the pupil localization model (i.e., the KiU-Net2D model) predicting the target category of each pixel.
[0050] For example, Figure 4 This is a schematic diagram of the probability map output by a pupil localization model provided in an embodiment of the present invention.
[0051] S302: Perform threshold segmentation on the probability map and select the largest connected region as the pupil prediction region based on connected component analysis.
[0052] In binary images, a connected component can be understood as a group of pixels belonging to the same category and spatially interconnected. The largest connected component refers to the foreground region with the largest area (number of pixels). In medical images or object detection and segmentation results, the main body region often contains small, isolated noise. For example, in pupil localization scenarios, small noise points may accompany the segmentation of the pupil. By analyzing the largest connected component, only the largest connected component can be retained, thereby removing small noise patches and ultimately cropping out the pupil prediction region.
[0053] S303: Cropping the eye image based on the pupil prediction region to obtain the pupil region image.
[0054] In this embodiment, cropping the eye image based on the pupil prediction region to obtain the pupil region image includes: inverting the size of the mask corresponding to the pupil prediction region based on the initial eye image size (i.e., the size of the eye image before performing image preprocessing), and cropping the eye image based on the inverted mask to obtain the pupil region image.
[0055] Specifically, the slit-lamp image of the eye to be processed is converted to grayscale and its size is adjusted to a uniform size, for example, a resolution of 1024×1024. The preprocessed slit-lamp image is then fed into a pre-trained KiU-Net2D model, i.e., the pupil localization model. The pupil localization model outputs a probability map of size 1024×1024. Further, the output probability map is subjected to threshold segmentation (e.g., binarization segmentation) and connected component analysis is applied to extract the region with the largest area as the final pupil prediction region. Finally, the mask size corresponding to the pupil prediction region is inversely scaled based on the initial eye image size (i.e., the size of the eye image before image preprocessing), and the eye image is cropped based on the inversely scaled mask to obtain the pupil region image.
[0056] Optionally, during the model pre-training stage, the pupil localization model uses an eye image processed by the first data augmentation strategy and its corresponding refined mask to train the KiU-Net2D segmentation model; wherein, the refined mask is a mask established by binarization annotation after coarse segmentation using SAM; the first data augmentation strategy includes at least one of the following: performing geometric transformation processing on the eye image; performing color enhancement processing on the eye image; performing spatial enhancement processing on the eye image; and performing data normalization processing on the eye image based on the data mean and standard deviation.
[0057] Figure 5 A schematic diagram of an eye image provided in an embodiment of the present invention; Figure 6 A schematic diagram of an eye image and its rough mask provided for an embodiment of the present invention; Figure 7This is a schematic diagram of an eye image and its refined mask provided in an embodiment of the present invention.
[0058] Specifically, see Figures 5 to 7 As shown, during the training phase of the pupil localization model, the SegmentAnything Model (SAM) is first used to coarsely segment the eye image and create a coarse mask. Then, a labeling tool (e.g., a 3D Slicer) is used for binarization annotation, deleting pixels outside the pupil region from the eye image. If the coarse segmentation is broken due to a slit lamp, it needs to be repaired until a refined mask covering the pupil region is generated. It should be noted that the edges of the refined mask do not need to perfectly fit the pupil. Further, after cropping the refined mask and removing impurity areas, only the mask corresponding to the pupil region is retained, resulting in the eye image / mask pair used to train the pupil localization model.
[0059] The first data augmentation strategy is an augmentation strategy for the training data of the pupil localization model, which specifically includes: performing geometric transformation processing on the eye image based on horizontal flipping and / or vertical flipping; performing color enhancement processing on the eye image based on random adjustment of brightness and contrast; performing spatial enhancement processing on the eye image based on affine transformation; and performing data normalization processing on the eye image based on the data mean and standard deviation.
[0060] Optionally, during the model pre-training phase, the diagnostic model uses pupil region images processed by the second data augmentation strategy and their corresponding lesion grading labels to perform at least two stages of grading training on the ConvNeXt model; wherein, the lesion grading labels include at least three lesion levels; the second data augmentation strategy includes at least one of the following: performing geometric transformation processing on the pupil region image; performing color enhancement processing on the pupil region image; performing spatial enhancement processing on the pupil region image; performing occlusion enhancement processing on the pupil region image; performing noise enhancement processing on the pupil region image; and performing data normalization processing on the pupil region image based on the data mean and standard deviation.
[0061] Specifically, the lesion grading labels include at least the category labels corresponding to no or mild lesions, the category labels corresponding to moderate lesions, and the category labels corresponding to severe lesions.
[0062] The second data augmentation strategy is an augmentation strategy for the model training data of the diagnostic model. Specifically, it includes: performing spatial-level augmentation processing on the pupil region image based on at least one of affine transformation, elastic transformation, and optical distortion; performing color augmentation processing on the pupil region image based on at least one of brightness, contrast saturation, hue random perturbation, and random gamma correction; performing occlusion augmentation processing on the pupil region image based on randomly erasing small areas and / or randomly occluding into a grid; performing noise augmentation processing on the pupil region image based on at least one of simulated motion blur and multiplicative noise; performing geometric transformation processing on the pupil region image based on at least one of horizontal flip, vertical flip, and random 90° multiple rotation; and normalizing the pupil region image to a common distribution based on the data mean and standard deviation.
[0063] Optionally, during the model pre-training stage, the loss function of the pupil localization model is constructed using a weighted combination of the Dice loss function and the binary cross-entropy loss function; the loss function of the diagnostic model is constructed using the Focal loss function with a class balance factor.
[0064] Specifically, the loss function design idea of the pupil localization model is as follows: the Dice loss function is introduced to enhance the optimization of target region overlap, and the binary cross-entropy (BCE) loss function is introduced to ensure pixel-level classification accuracy.
[0065] The Dice coefficient was originally used to measure the similarity between two sets. For example, the Dice coefficient can be defined as: In the application scenario of eye image segmentation, A represents the predicted set of foreground pixels; B represents the true set of foreground pixels. The Dice value ranges from [0,1], and the closer it is to 1, the closer the prediction is to the true value.
[0066] The mathematical expression for the Dice coefficient in a segmentation scenario is shown in Formula 1: (Formula 1) in, This represents the model prediction probability for the i-th pixel. ; This represents the true label of the i-th pixel. N represents the total number of pixels.
[0067] The Dice loss function is typically defined as 1 - the Dice coefficient, and its mathematical expression is shown in Formula 2: (Formula 2) in: This represents the model prediction probability for the i-th pixel. ; This represents the true label of the i-th pixel. N represents the total number of pixels; This indicates the denominator adjustment parameter. It is usually set to a very small number to prevent the denominator from being 0.
[0068] The binary cross-entropy loss function (BCE loss) is a commonly used loss function in binary classification problems, used to measure the difference between the target value and the predicted value. In binary classification problems, it is employed... This represents the real label (where 0 = background, 1 = foreground); This represents the probability predicted by the model. If the BCE for a single sample is defined as: If this is extended to multiple pixels, the mathematical expression for the loss function in image segmentation is as shown in Formula 3: (Formula 3) Where N represents the total number of pixels; This represents the model prediction probability for the i-th pixel. ; This represents the true label of the i-th pixel. .
[0069] In summary, the pupil localization model adopts Form serves as the loss function during model training. These represent the weighting coefficients of the two loss functions. , The specific value can be adjusted according to actual training needs, and there is no limit to its value.
[0070] The design idea for the loss function of the diagnostic model is as follows: In many classification or segmentation tasks, there is class imbalance (some classes have many samples, while others have few, causing the model to favor predicting the majority class) or difficulty imbalance (most "easy-to-classify" samples dominate the gradient, while the contribution of "difficult-to-classify" samples is submerged). In cataract grading tasks, there is a situation where the amount of data for no or mild / moderate cataracts is greater than that for severe cataracts. Focal Loss addresses the problem of difficulty imbalance, while Class-Balanced Loss addresses the problem of class imbalance. Therefore, CB-FocalLoss is a combination of Focal Loss and Class-Balanced Loss.
[0071] Specifically, assuming the total number of classes in the training set is C, and the number of samples in class c is... Then the number of valid samples for this category is defined as ,in, It is a hyperparameter. The closer the value is to 1, the greater the weight of the rare category. The category balance factor, or category balance weight, can be expressed as... It should be noted that rarer categories have larger category balance weights.
[0072] The mathematical expression for the standard Focal loss function is shown in Formula 4: (Formula 4) in, Predict the probability for the model (if the label is 1, then) If the label is 0, then ); This represents a regulation factor used to control the degree of suppression for easily classified samples. .
[0073] The Focal loss function with a class balance factor can be implemented by adjusting the class balance weights. By incorporating the standard Focal loss function, a loss function is constructed as shown in Equation 5: (Formula 5) in, This represents the balance factor for category c; This represents the standard Focal loss function part, used to reduce the weight of easy samples; This represents the cross-entropy component, used to measure the difference between the prediction and the actual value.
[0074] In summary, when constructing the loss function for the diagnostic model, a class balancing factor, i.e., class balancing weights, is introduced on top of the standard Focal loss function. It also solves the problems of class imbalance and sample difficulty imbalance during model training, making it particularly suitable for application scenarios with few samples and class imbalance, such as medical imaging and object detection.
[0075] Based on the same inventive concept as the above embodiments, this invention also provides an intelligent eye image analysis system. This system can execute the intelligent eye image analysis method provided in any embodiment of this invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0076] Figure 8 This is a schematic diagram of the structure of an intelligent eye image analysis system provided in an embodiment of the present invention. Figure 8As shown, the system includes: an image acquisition module 101, used to acquire eye images and perform image preprocessing on the eye images; a model acquisition module 102, used to acquire a pre-trained pupil localization model and a diagnostic model; wherein, the pupil localization model is established based on the KiU-Net2D segmentation model; the diagnostic model is established based on at least two levels of ConvNeXt model; an image segmentation module 103, used to extract pupil region images from eye images based on the pupil localization model; and a diagnostic module 104, used to classify and diagnose pupil region images based on the diagnostic model and output the diagnostic results of eye lesions.
[0077] In some optional embodiments, the diagnostic model includes a first diagnostic sub-model and a second diagnostic sub-model, wherein the lesion level output by the first diagnostic sub-model is lower than the lesion level output by the second diagnostic sub-model. In this embodiment, the diagnostic module 104 is configured to: import the pupil region image into the first diagnostic sub-model and obtain the preliminary classification prediction output by the first diagnostic sub-model; when the preliminary classification prediction is a first category, use the first category as the classification diagnosis result; when the preliminary classification prediction is a second category, import the pupil region image into the second diagnostic sub-model and perform a secondary diagnosis; wherein the lesion level corresponding to the first category is lower than the lesion level corresponding to the second category.
[0078] In some optional embodiments, during the model pre-training stage, the first diagnostic sub-model trains the first ConvNeXt model using a first category and a second category; the second diagnostic sub-model trains the second ConvNeXt model using subcategories of the second category; wherein the lesion level corresponding to the first category is lower than the lesion level corresponding to any subcategory.
[0079] In some optional embodiments, the image segmentation module 103 is configured to: import the preprocessed eye image into the pupil localization model and obtain the probability map output by the pupil localization model; perform threshold segmentation on the probability map and select the largest connected component region as the pupil prediction region based on connected component analysis; and crop the eye image based on the pupil prediction region to obtain the pupil region image.
[0080] In some optional embodiments, the image acquisition module 101 is configured to perform grayscale conversion and image resampling processing on the eye image.
[0081] Optionally, the pupil localization model uses an eye image processed by the first data augmentation strategy and its corresponding refined mask to train the KiU-Net2D segmentation model; wherein, the refined mask is a mask established by binarization annotation after coarse segmentation using SAM; the first data augmentation strategy includes at least one of the following: geometric transformation processing, color enhancement processing, spatial enhancement processing, and data normalization processing.
[0082] Optionally, the diagnostic model uses pupil region images processed by the second data augmentation strategy and their corresponding lesion grading labels to perform at least two stages of grading training on the ConvNeXt model; wherein the lesion grading labels include at least three lesion levels; the second data augmentation strategy includes at least one of the following: geometric transformation processing, color enhancement processing, spatial enhancement processing, occlusion enhancement processing, noise enhancement processing, and data normalization processing.
[0083] Optionally, the loss function of the pupil localization model is constructed using a weighted combination of the Dice loss function and the binary cross-entropy loss function; the loss function of the diagnostic model is constructed using the Focal loss function with a class balance factor.
[0084] Based on any of the above embodiments, the present invention also provides an electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to execute the intelligent eye image analysis method provided in any of the above embodiments.
[0085] Figure 9 This is a schematic diagram of the structure of an electronic device for implementing the intelligent eye image analysis method of this invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0086] like Figure 9 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0087] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0088] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the intelligent eye image analysis method described above.
[0089] In some embodiments, the above-described intelligent eye image analysis method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the intelligent eye image analysis method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the above-described intelligent eye image analysis method by any other suitable means (e.g., by means of firmware).
[0090] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0091] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0092] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0093] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0094] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0095] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system, addressing the shortcomings of traditional physical hosts and VPS servers, such as high management difficulty and weak business scalability.
[0096] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0097] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for intelligent analysis of eye images, characterized in that, include: Acquire an eye image and perform image preprocessing on the eye image; Obtain a pre-trained pupil localization model and a diagnostic model; wherein the pupil localization model is based on the KiU-Net2D segmentation model; and the diagnostic model is based on at least two levels of the ConvNeXt model. Based on the pupil localization model, the pupil region image is extracted from the eye image; The diagnostic model is used to classify and diagnose the pupil region image and output the diagnostic results of eye lesions.
2. The method according to claim 1, characterized in that, The diagnostic model includes a first diagnostic sub-model and a second diagnostic sub-model, wherein the lesion level output by the first diagnostic sub-model is lower than the lesion level output by the second diagnostic sub-model. The process of classifying and diagnosing the pupil region image based on the diagnostic model and outputting the diagnostic results of eye lesions includes: The pupil region image is imported into the first diagnostic sub-model, and the preliminary classification prediction output by the first diagnostic sub-model is obtained. When the initial classification prediction is the first category, the first category is used as the classification diagnosis result; When the initial classification prediction is the second category, the pupil region image is imported into the second diagnostic sub-model to perform a secondary diagnosis; The lesion grade corresponding to the first category is lower than that corresponding to the second category.
3. The method according to claim 2, characterized in that, During the model pre-training phase, the first diagnostic sub-model trains the first ConvNeXt model using the first category and the second category; The second diagnostic sub-model is trained on the second ConvNeXt model using the subcategories of the second category; Wherein, the lesion level corresponding to the first category is lower than the lesion level corresponding to any of the subcategories.
4. The method according to claim 1, characterized in that, The step of extracting the pupil region image from the eye image based on the pupil localization model includes: The preprocessed eye image is imported into the pupil localization model, and the probability map output by the pupil localization model is obtained. The probability map is segmented by a threshold, and the region with the largest connected component is selected as the pupil prediction region based on connected component analysis. The eye image is cropped based on the pupil prediction region to obtain the pupil region image.
5. The method according to claim 1, characterized in that, The process of acquiring an eye image and performing image preprocessing on the eye image includes: The eye image is subjected to grayscale conversion and image resampling.
6. The method according to any one of claims 1-5, characterized in that, During the model pre-training stage, the pupil localization model uses an eye image processed by the first data augmentation strategy and its corresponding refined mask to train the KiU-Net2D segmentation model. The refined mask is a mask created by performing coarse segmentation using SAM and then labeling it with binary data. The first data augmentation strategy includes at least one of the following: geometric transformation processing, color enhancement processing, spatial enhancement processing, and data normalization processing.
7. The method according to any one of claims 1-5, characterized in that, During the model pre-training phase, the diagnostic model uses pupil region images processed by the second data augmentation strategy and their corresponding lesion grading labels to perform at least two phases of grading training on the ConvNeXt model. The lesion grading label includes at least three lesion levels; The second data augmentation strategy includes at least one of the following: geometric transformation processing, color enhancement processing, spatial enhancement processing, occlusion enhancement processing, noise enhancement processing, and data normalization processing.
8. The method according to any one of claims 1-5, characterized in that, During the model pre-training phase, the loss function of the pupil localization model is constructed using a weighted combination of the Dice loss function and the binary cross-entropy loss function. The loss function of the diagnostic model is constructed using the Focal loss function with a class balance factor.
9. An intelligent image analysis system for the eye, characterized in that, include: An image acquisition module is used to acquire an eye image and perform image preprocessing on the eye image; The model acquisition module is used to acquire pre-trained pupil localization and diagnostic models; wherein, the pupil localization model is based on the KiU-Net2D segmentation model; and the diagnostic model is based on at least two levels of ConvNeXt model. The image segmentation module is used to extract the pupil region image from the eye image based on the pupil localization model; The diagnostic module is used to classify and diagnose the pupil region image based on the diagnostic model and output the diagnostic results of eye lesions.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the intelligent eye image analysis method according to any one of claims 1-8.
Citation Information
Patent Citations
Three-dimensional medical image processing device and method
CN109754394A
Medical image segmentation method and system based on AS-UNet
CN113012172A
Mandibular impacted wisdom tooth image processing method and system
CN115170531A
Cataract grading method, device and equipment and storage medium
CN116258698A
Video image segmentation and classification network and tongue movement ultrasonic video segmentation and classification method
CN118314337A