Image processing method and system for intelligent screening of congenital cataract
Through portable devices and deep learning technology, cataracts in infants and young children are automatically identified, solving the problem of high misdiagnosis rate in primary medical institutions, and achieving efficient and accurate cataract screening.
Patent Information
- Application Number
- CN202510499581.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art congenital cataract screening in primary medical institutions depends on doctor experience, with high misdiagnosis and missed diagnosis rates, poor portability of traditional equipment, and difficult to adapt to the needs of large-scale screening.
A portable pediatric ophthalmology examiner is used to combine Otsu method, OpenCV image repair algorithm and YOLOv8 model to perform image preprocessing and deep learning processing, automatically identify pupil areas and determine cataracts, and reduce manual intervention.
It improves the accuracy and stability of screening, reduces the burden on doctors, and realizes structured diagnostic results output, which is suitable for use in grassroots environments.
Smart Images

Figure CN120412927A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cataract diagnosis, and more specifically, to an image processing method and system for intelligent screening of congenital cataracts. Background Art
[0002] Since the critical period of human visual development is concentrated within 7 years after birth, if congenital cataracts are not screened and intervened in a timely manner during this period, it is extremely likely to cause irreversible permanent visual impairment. Therefore, early identification and intervention are crucial for ensuring children's visual health.
[0003] Currently, the commonly used screening methods in clinics include red reflex examination and slit lamp microscopy examination. The red reflex examination mostly observes the pupil reflex through a direct ophthalmoscope to judge whether there is a light axis occlusion lesion, while the slit lamp image can directly show the anterior segment structural abnormalities. Although the above methods have certain effects in clinical diagnosis, there are still the following prominent problems: both the red reflex judgment and the slit lamp image analysis rely on the clinical experience and observation skills of ophthalmologists. Especially in primary medical institutions, the lack of professional personnel easily leads to uneven screening quality, with relatively high misdiagnosis and missed diagnosis rates. There are problems such as great operation difficulty, low cooperation degree, and poor equipment portability when using traditional slit lamp equipment to collect images of infants and young children, making it difficult to meet the application requirements of large-scale screening or the primary environment. Summary of the Invention
[0004] The purpose of the present invention is to provide an image processing method and system for intelligent screening of congenital cataracts to solve the problems raised in the above background art.
[0005] Technical Solution: An image processing method for intelligent screening of congenital cataracts, the method includes the following specific steps:
[0006] S1. Acquisition step: Use a portable pediatric ophthalmic examination instrument (with an image sensor built-in with a resolution not lower than 1920×1080) to collect anterior segment images and fundus red reflex images of infants and children, and store them in the form of original image data (in JPEG or RAW format);
[0007] S2. Preprocessing step:
[0008] S2.1. Apply the Otsu method to the original image for global segmentation to filter out background noise;
[0009] S2.2. Use an image inpainting algorithm implemented based on OpenCV to perform pixel interpolation correction on local overexposed and reflective areas in the image;
[0010] S2.3. Adopt an edge detection algorithm to accurately define the boundary of the pupil area in the image and output the area coordinates to obtain a preprocessed image;
[0011] S3. Annotation step: On the preprocessed image in S2, calibrate the pupil area and interference areas in a manual box selection manner through an HTML5-based annotation platform, record the vertex coordinates of each calibrated area, and output an annotation file in JSON format to obtain the annotated image;
[0012] S4. Deep learning processing step:
[0013] S4.1. Input the image data after annotation in S3 into the YOLOv8 model that has been supervised trained and whose parameters have been fine-tuned, use the convolutional neural network to extract the features of the target area in the image, and output the detection data including the target bounding box coordinates and the confidence of the candidate area;
[0014] S4.2. Adopt the softmax classification algorithm to determine the category of the detected area, determine it as "congenital cataract" or "normal", and generate a classification result with specific confidence values;
[0015] S5. Result output step: Generate an extracted area image by using image cropping and local enhancement techniques based on the detection data, organize the target boundary and classification results into an XML format report, and store this report in a predetermined database;
[0016] S6. Model update step: Periodically extract new image data from the database, retrain the YOLOv8 and CNN models by using the gradient descent and incremental learning algorithms, and automatically update the model weight file after the model is retrained.
[0017] Preferably, in S2.1, the Otsu method is used to determine the global threshold.
[0018] Preferably, the image repair algorithm in S2.2 is based on local mean brightness adjustment and bilinear interpolation techniques.
[0019] Preferably, the YOLOv8 model is fine-tuned by using transfer learning technology on the basis of the pre-trained model, and the detection data is output in the form of bounding box coordinates (upper left and lower right coordinates) and confidence percentages.
[0020] Preferably, the model update step is based on the incremental learning algorithm of gradient descent, and the model performance improvement is confirmed through preset evaluation metrics (such as accuracy and F1 value) after each retraining.
[0021] An image processing system for intelligent screening of congenital cataract, comprising:
[0022] Image acquisition unit: A portable pediatric ophthalmic examination instrument with an image sensor with a resolution of not less than 1920×1080 built in, used to collect anterior segment and fundus red light reflection images and store them in digital format;
[0023] Image preprocessing unit: equipped with an adaptive threshold segmentation module, an OpenCV-based local image restoration module, and a region growing edge detection module to preprocess the collected image and output a preprocessed image containing the pupil area coordinates;
[0024] Annotation unit: Integrates an online polygon annotation tool, supports manual selection of pupils and interference areas on pre-processed images, and generates annotation data in JSON format;
[0025] Data management unit: uses SQL database to store collected images, pre-processed data, annotation files and inspection reports, and supports data output in XML format.
[0026] Preferably, the adaptive threshold segmentation module of the image preprocessing unit adopts the Otsu method and outputs binary image data in the processing result.
[0027] Preferably, the data management unit supports the output of detection reports based on the XML protocol and uses data encryption technology to ensure the security of image data.
[0028] Compared with the prior art, the advantages of the present invention are:
[0029] (1) By introducing the Otsu method for background removal and OpenCV-based image restoration (local brightness adjustment and bilinear interpolation) and other preprocessing algorithms, we can effectively solve interference factors such as overexposure, reflection, and eyelash occlusion in the image, making the system more robust to low-quality images and improving the accuracy and stability of the overall model.
[0030] (2) The pupil area is accurately extracted using the YOLOv8 target detection model, and the CNN model is combined to complete feature recognition and classification, effectively reducing the reliance on manual judgment of doctors and alleviating the workload of professionals.
[0031] (3) Through the Softmax classification algorithm, a structured classification label and a specific confidence value are output for each detection area, making the judgment basis quantifiable, enhancing clinical reference, and facilitating subsequent decision-making and referral. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 This is an overall flow chart of an image processing method for intelligent screening of congenital cataracts according to the present invention. DETAILED DESCRIPTION
[0033] Example
[0034] An image processing method for intelligent screening of congenital cataracts, the method comprising the following specific steps:
[0035] S1. Acquisition steps: Use a portable pediatric ophthalmological examination device (with a built-in image sensor with a resolution of not less than 1920×1080) to collect anterior segment images and fundus red reflex images of infants and children, and store them in raw image data (in JPEG or RAW format);
[0036] S2. Preprocessing steps:
[0037] S2.1. Apply Otsu's method to perform global segmentation on the original image and filter out background noise;
[0038] Steps:
[0039] 1. Calculate the image histogram and probability distribution: Assume that the image grayscale ranges from 0 to L-1 (usually L=256), where n i Represents the number of pixels with gray value i, the total number of pixels is N, then the probability of each gray level is p i for
[0040] p i =n i / N,i=0,1,…,L-1
[0041] 2. Divide foreground and background: Assuming a certain value t, divide the image into two parts:
[0042] Background: grayscale value range [0, t]
[0043] Foreground: grayscale value range [t+1, L-1]
[0044] Calculate the weights of the two parts:
[0045] Background Weight:
[0046]
[0047] Foreground Weight:
[0048]
[0049] 3. Calculate the mean of the two categories
[0050] Background mean:
[0051]
[0052] Outlook mean:
[0053]
[0054] 4. Find the between-class variance: The between-class variance is defined as:
[0055] σ - B 2f(t) = w0(t)·w1(t)·[μ0(t)-μ1(t)] 2 ;
[0056] 5. Select the optimal threshold: Traverse t = 0, 1, …, L-1 and select the threshold t that maximizes σ - B 2 (t). * .
[0057] t ★ = argmax (t=0…L-1) {σ - B 2 (t)}.
[0058] 6. Generate a binary image: Use the threshold t * to perform global segmentation on the original image: For the pixel l(x, y), if I(x,y) > t * , then set it as the foreground (or set it to 1); otherwise set it as the background (or set it to 0).
[0059] S2.2. Use the image inpainting algorithm implemented based on OpenCV to perform pixel interpolation correction on the locally overexposed and specular regions in the image;
[0060] Objective: Correct the regions in the image caused by local overexposure or specular reflection to make the local brightness of the image more balanced and facilitate subsequent processing.
[0061] (1) Local mean brightness adjustment
[0062] 1. Determine the overexposed or specular region
[0063] Use the threshold method or local statistics (for example, judge that the average brightness within a certain window exceeds the preset upper limit T h ) to determine the abnormal region. Let the window size be k×k (such as 3×3 or 5×5). For each pixel l(x, y), calculate the local mean μ local :
[0064]
[0065] 2. Judge the correction condition: If μ local ≥T h , then it is considered that there is an overexposure / specular reflection problem in this local region and needs to be corrected.
[0066] 3. Local mean adjustment: For the target pixel, calculate the adjustment factor α (for example, according to the set target brightness T target ), and update the brightness:
[0067] I adj (x, y) = I(x, y)·(T target / μlocal ).
[0068] I adj (x, y) is the brightness of the (x, y) coordinates. This method ensures that the overall brightness of the region approaches T target closer.
[0069] (2) Bilinear interpolation technique: When the information of some pixels is missing due to overexposure or needs to be re - estimated, bilinear interpolation is used to fill in new values. Let the coordinates of the point to be interpolated be (x, y), its upper - left neighborhood coordinates be (x1, y1) and the lower - right neighborhood coordinates be (x2, y2). The bilinear interpolation formula is:
[0070] I interp (x, y)=I(x1, y1)·(x2 - x)·(y2 - y)+I(x2, y1)·(x - x1)·(y2 - y)+I(x1, y2)·(x2 - x)·(y - y1)+I(x2, y2)·(x - x1)·(y - y1)
[0071] where I represents the pixel value of the image at a certain pixel point;
[0072] S2.3. Adopt an edge - detection algorithm to accurately define the boundary of the pupil region in the image and output the region coordinates to obtain a pre - processed image; The image - restoration algorithm in S2.2 is based on local mean brightness adjustment and bilinear interpolation technology.
[0073] Step description:
[0074] 1. Grayscale conversion: If the restored image is a color image, it is first converted to a grayscale image G(x, y) to simplify the calculation.
[0075] 2. Apply an edge - detection algorithm: Common methods include the Sobel operator, Canny edge detection, etc. Here, take the Sobel operator as an example:
[0076] Sobel convolution kernel in the X direction:
[0077] K x [-1, 0, 1 - 2, 0, 2 - 1, 0, 1]
[0078] Sobel convolution kernel in the Y direction:
[0079] K v =[-1, -2, -1 0, 0, 0 1, 2, 1]
[0080] Calculate the gradients of the image in the X and Y directions:
[0081] G x (x, y)=(K x *G)(x, y), Gγ (x, y) = (K γ *G)(x, y).
[0082] The gradient magnitude is:
[0083]
[0084] 3. Thresholding and Binarization: Threshold the gradient magnitude image (either a fixed threshold or an adaptive method can be used) to obtain a binary edge image E(x, y), where the edge pixel values are set to 1 and the rest are set to 0.
[0085] 4. Edge Connected Region Analysis: Perform connected component analysis on the binary edge image (such as using the findContours function in OpenCV) to extract all independent boundary curves.
[0086] 5. Pupil Region Extraction: Combine prior knowledge (such as the pupil region shape being close to an ellipse and located in the center of the image, etc.) to filter out the connected regions that meet the characteristics. Calculate geometric parameters (minimum enclosing ellipse of the boundary, area, circularity, etc.) for the candidate regions, and select the regions that meet the threshold conditions as the pupil region.
[0087] 6. Output Region Coordinates: For the extracted pupil boundary, output the coordinates of its contour points or the boundary coordinates of the minimum enclosing rectangle / ellipse. The storage format can be an array, JSON, or XML.
[0088] S3. Annotation Step: On the preprocessed image in S2, calibrate the pupil region and interference regions in a manual box selection manner through an HTML5-based annotation platform, record the vertex coordinates of each calibrated region, and output a JSON-formatted annotation file to obtain the annotated image;
[0089] S4. Deep Learning Processing Step:
[0090] S4.1. Input the image data annotated in S3 into the YOLOv8 model that has been supervised-trained and whose parameters have been fine-tuned, use the convolutional neural network to extract the features of the target region in the image, and output the detection data including the target bounding box coordinates and the confidence of the candidate region; among them, the YOLOv8 model is fine-tuned using transfer learning technology based on the pre-trained model, and the detection data is output in the form of bounding box coordinates (upper left and lower right coordinates) and confidence percentage.
[0091] S4. Select the Softmax classification algorithm to perform class determination on the detected regions, determine whether it is "congenital cataract" or "normal", and generate a classification result with specific confidence values;
[0092] S4.2: Specific operation steps of the Softmax classification algorithm
[0093] 1. Feature extraction and classifier input: For each candidate region detected by the YOLOv8 model, extract its corresponding feature vector f{z} = [z1, z2], where:
[0094] z1 represents the score that the region belongs to the category of "congenital cataract";
[0095] z2 represents the score that the region belongs to the category of "normal".
[0096] These scores are usually calculated through the forward propagation of a convolutional neural network.
[0097] 2. Apply the Softmax function to calculate the probability distribution: Input the above score vector f{z} into the Softmax function to calculate the probability of each category:
[0098]
[0099] where, represents the exponential operation on the score z i and the denominator is the sum of the exponentials of all category scores, ensuring that the output probability values are between 0 and 1, and the sum of the probabilities of all categories is 1.
[0100] 3. Category determination and confidence output: According to the calculated probability distribution, select the category with the highest probability as the final determination result: If {Softmax}(z1) > {Softmax}(z2), then it is determined as "congenital cataract"; otherwise, it is determined as "normal".
[0101] At the same time, output the probability value of the corresponding category as the confidence. For example, if {Softmax}(z1) = 0.92, it means that the region is determined as "congenital cataract" with a confidence of 92%.
[0102] 4. Structured representation of classification results: Structurally represent the classification results of each candidate region in the following form: Category label: "congenital cataract" or "normal";
[0103] Confidence: The probability value of the corresponding category, ranging from 0 to 1.
[0104] For example: Category label: "congenital cataract"; Confidence: 0.92.
[0105] This structured result can be used for subsequent diagnostic report generation or further medical analysis. [[ID=4,2]]
[0106] S5. Result output step: Generate the extracted region image by using image cropping and local enhancement techniques based on the detection data, organize the target boundary and classification results into an XML format report, and store the report in a predetermined database;
[0107] S6. Model Update Step: Periodically extract newly added image data from the database, and retrain the YOLOv8 and CNN models using the gradient descent and incremental learning algorithms. After the model retraining, the model weight file is automatically updated. The model update step is based on the incremental learning algorithm of gradient descent, and after each retraining, the model performance improvement is confirmed through preset evaluation metrics (such as accuracy and F1 value).
[0108] An image processing system for intelligent screening of congenital cataracts, comprising:
[0109] Image Acquisition Unit: A portable pediatric ophthalmic examination instrument with an image sensor built-in with a resolution of not less than 1920×1080, used to collect anterior segment and fundus red reflex images and store them in digital format;
[0110] Image Preprocessing Unit: Configured with an adaptive threshold segmentation module, a local image repair module based on OpenCV, and a region growing edge detection module, used to preprocess the collected images and output preprocessed images containing pupil region coordinates;
[0111] Annotation Unit: Integrated with an online polygon annotation tool, supporting the determination of pupil and interference regions in the preprocessed images in a manual box selection manner and generating annotation data in JSON format;
[0112] Data Management Unit: Uses an SQL database to store collected images, preprocessed data, annotation files, and detection reports, and supports data output in XML format.
[0113] The adaptive threshold segmentation module of the image preprocessing unit uses the Otsu method and outputs binary image data in the processing result.
[0114] The data management unit supports the output of detection reports based on the XML protocol and uses data encryption technology to ensure the security of image data.
[0115] The above shows and describes the basic principles, main features, and advantages of the present invention; those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions in the specification are only preferred examples of the present invention and do not limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed; the scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. An image processing method for congenital cataract intelligent screening, characterized in that, It includes the following specific steps: S1. Acquisition step: Use a portable pediatric ophthalmic examination instrument to acquire anterior segment images and fundus red light reflex images of infants and children, and store them as original image data; S2. Preprocessing step: S2.
1. Apply the Otsu method to the original image for global segmentation to filter out background noise; S2.
2. Use an image inpainting algorithm implemented based on OpenCV to perform pixel interpolation correction on local overexposed and reflective areas in the image; S2.
3. Adopt an edge detection algorithm to accurately define the boundary of the pupil area in the image, output the area coordinates, and obtain a preprocessed image; S3. Annotation step: On the preprocessed image in S2, calibrate the pupil area and interference areas in a manual box selection manner through an HTML5-based annotation platform, record the vertex coordinates of each calibrated area, and output a JSON format annotation file to obtain an annotated image; S4. Deep learning processing step: S4.
1. Input the image data of the annotated image in S3 into the YOLOv8 model that has been supervised trained and whose parameters have been fine-tuned, use a convolutional neural network to extract the feature of the target area in the image, and output detection data including the target bounding box coordinates and the confidence of the candidate area; S4.
2. Adopt a softmax classification algorithm to perform class determination on the detected area, determine it as "congenital cataract" or "normal", and generate a classification result with specific confidence values; S5. Result output step: Generate an extracted area image by using image cropping and local enhancement techniques based on the detection data, organize the target boundary and classification result into an XML format report, and store the report in a predetermined database; S6. Model update step: Periodically extract new image data from the database, and retrain the YOLOv8 and CNN models using the gradient descent and incremental learning algorithms, and automatically update the model weight file after the model is retrained.
2. The image processing method for congenital cataract intelligent screening according to claim 1, wherein: In S2.1, the Otsu method is used to determine the global threshold.
3. The image processing method for congenital cataract intelligent screening according to claim 1, characterized in that: The image inpainting algorithm in S2.2 is based on local mean brightness adjustment and bilinear interpolation techniques.
4. An image processing method for congenital cataract intelligent screening according to claim 1, characterized in that: Among them, the YOLOv8 model is fine-tuned using transfer learning technology based on a pre-trained model, and the detection data is output in the form of bounding box coordinates and confidence percentages.
5. The image processing method for congenital cataract intelligent screening according to claim 1, characterized in that: The model update step is based on the incremental learning algorithm of gradient descent, and the model performance improvement is confirmed through preset evaluation indicators after each retraining.
6. An image processing system for intelligent screening of congenital cataracts, characterized in that, The system includes: Image acquisition unit: A portable pediatric ophthalmic examination instrument with an image sensor with a resolution of not less than 1920×1080 built-in, used to acquire anterior segment and fundus red light reflex images and store them in digital format; Image preprocessing unit: Configured with an adaptive threshold segmentation module, a local image inpainting module based on OpenCV, and a region growing edge detection module, used to preprocess the acquired image and output a preprocessed image containing pupil area coordinates; Annotation unit: Integrated with an online polygon annotation tool, supporting the determination of pupil and interference areas in a manual box selection manner on the preprocessed image and generating annotation data in JSON format; Data Management Unit: It uses an SQL database to store the acquired images, preprocessed data, annotation files, and detection reports, and supports the output of data in XML format.
7. An image processing system for congenital cataract intelligent screening according to claim 6, characterized in that: The adaptive threshold segmentation module of the Image Preprocessing Unit adopts the Otsu method and outputs binary image data in the processing results.
8. An image processing system for congenital cataract intelligent screening according to claim 6, characterized in that: The Data Management Unit supports the output of detection reports based on the XML protocol and uses data encryption technology to ensure the security of image data.