Method and device for training a detection model for detecting abnormal areas in fundus images
By dividing ultra-wide-angle retinal images into blocks and using global and local classifiers with supervised and semi-supervised loss functions, the method improves anomaly detection accuracy and reliability in ultra-wide-angle retinal images, addressing the challenges of large pixel size and limited annotation data.
Patent Information
- Application Number
- CN202310961124.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-01
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-08-01
AI Technical Summary
When detecting abnormal areas, the existing ultra-wide-angle fundus images have problems such as large amount of information, high labeling cost and scarce labeling data, which leads to inaccurate detection results and it is difficult to accurately identify and locate in the absence of abnormal areas labeling.
By blocking ultra-wide-angle fundus images, intermediate features are extracted using feature extractors, and fully supervised and semi-supervised loss functions are calculated in combination with global and local classifiers, the detection model is trained, the detailed information loss is reduced, and the detection accuracy is improved.
In the absence of abnormal area annotation, the abnormal area in the ultra-wide-angle fundus image is accurately identified and positioned, which improves the reliability and accuracy of the detection model and reduces the labeling cost.
Smart Images

Figure CN116934730B_ABST
Abstract
Description
Technical Field
[0001] The present application generally relates to the field of artificial intelligence technology. More specifically, the present application relates to a method, device, and computer-readable storage medium for training a detection model for detecting abnormal regions in fundus images. Further, the present application also relates to a method, device, and computer-readable storage medium for detecting abnormal regions in fundus images. Background Art
[0002] Disclosed by the World Health Organization ("WHO"), more than 2.2 billion people suffer from visual impairment or blindness, and more than 1 billion of them can effectively prevent the deterioration of the condition through early prevention or effective treatment. Therefore, early detection and correct treatment of eye diseases are of great significance for preventing visual impairment. With the continuous development of artificial intelligence, artificial intelligence, especially machine learning and deep learning, is expected to provide better solutions for prevention and treatment.
[0003] In recent years, ultra-widefield fundus imaging (as shown in the left figure in Figure 1 ) has been gradually applied to the diagnosis and treatment of ophthalmic diseases. This technology can take fundus photos without mydriasis, and compared with the color fundus images taken by traditional color fundus cameras (as shown in the right figure in Figure 1 ), ultra-widefield fundus imaging can obtain a wider retinal field of view (for example, the field of view can be as high as 200°). Thus, the ultra-widefield fundus imaging technology significantly reduces the complexity of the imaging process for identifying fundus diseases, improves the visible range of fundus photos, and provides more valuable information and data for clinical diagnosis and treatment planning. Currently, many research works have effectively utilized deep learning technology on ultra-widefield fundus photos to explore the feasibility of automated diagnosis of some specific lesions. However, the existing methods often target the entire image processing, and the ultra-widefield images have large pixels, which will ignore some small abnormal details, resulting in inaccurate detection results. In addition, most of the existing research works focus on the abnormal classification in the image, and less research is done on the specific location of the abnormal region. On the contrary, compared with traditional color fundus images, ultra-widefield images have a large amount of information, the cost of annotating the detection boxes of abnormal regions is high, and there is less data of ultra-widefield fundus images that are annotated and of high quality, which is not conducive to the training of the model.
[0004] In view of this, there is an urgent need to provide a solution for training a detection model for detecting abnormal regions in fundus images, so as to accurately identify and locate the abnormal regions in ultra-widefield fundus images in the case of lack or few annotations of abnormal regions. Summary of the Invention
[0005] To at least solve one or more of the above-mentioned technical problems, the present application proposes a solution for training a detection model for detecting abnormal regions in fundus images in multiple aspects.
[0006] In a first aspect, the present application provides a method for training a detection model for detecting abnormal regions in fundus images, wherein the detection model at least includes a feature extractor and a classifier, and the method includes: obtaining ultra-wide field of view fundus images for training the detection model, wherein the ultra-wide field of view fundus images include ultra-wide field of view fundus images without labeled detection frames and ultra-wide field of view fundus images with labeled detection frames; cropping the ultra-wide field of view fundus images without labeled detection frames and the ultra-wide field of view fundus images with labeled detection frames to obtain respective multiple image patches; using the trained feature extractor to perform feature operations related to abnormal region detection on the multiple image patches to obtain respective multiple intermediate features; inputting the multiple intermediate features into a first classifier to perform global classification operations and into a second classifier to perform local classification operations, and calculating a fully supervised loss function and a semi-supervised loss function based on whether there is a labeled detection frame; and training the detection model for detecting abnormal regions in fundus images by using the fully supervised loss function and the semi-supervised loss function.
[0007] In one embodiment, the ultra-wide field of view fundus images with labeled detection frames include ultra-wide field of view fundus images with labeled detection frames and labeled abnormal classifications, and the ultra-wide field of view fundus images without labeled detection frames include ultra-wide field of view fundus images with labeled abnormal classifications but without labeled detection frames and ultra-wide field of view fundus images without labeled detection frames and without labeled abnormal classifications.
[0008] In another embodiment, the trained feature extractor is obtained by the following operations: performing data augmentation operations on the multiple image patches to obtain data-augmented image patches; inputting the data-augmented image patches and the multiple image patches into the feature extractor for feature operations, and calculating a contrast loss function; and training the feature extractor by using the contrast loss function.
[0009] In yet another embodiment, a fully connected module is further connected to the trained feature extractor, and using the trained feature extractor to perform feature operations related to abnormal region detection on the multiple image patches to obtain respective multiple intermediate features includes: using the trained feature extractor to perform feature operations related to abnormal region detection on the multiple image patches, and performing a fully connected operation on the feature operation results via the fully connected module to obtain respective multiple intermediate features.
[0010] In yet another embodiment, inputting the plurality of intermediate features into a first classifier to perform a global classification operation and into a second classifier to perform a local classification operation, and calculating a fully supervised loss function and a semi-supervised loss function based on whether a detection box is labeled includes: inputting the plurality of intermediate features into the first classifier to perform a global classification operation, and calculating a first fully supervised loss function based on the labeled detection box and a first semi-supervised loss function based on the unlabeled detection box; and inputting the plurality of intermediate features into the second classifier to perform a local classification operation, and calculating a second fully supervised loss function based on the labeled detection box and a second semi-supervised loss function based on the unlabeled detection box.
[0011] In yet another embodiment, wherein the first classifier and the second classifier include fully connected layers, and the first classifier further includes an attention module, inputting the plurality of intermediate features into the first classifier to perform a global classification operation and inputting the plurality of intermediate features into the second classifier to perform a local classification operation includes: inputting the plurality of intermediate features into the attention module in the first classifier to perform an attention operation, and performing a fully connected operation on the attention operation result via the fully connected layer in the first classifier to perform a global classification operation; and inputting the plurality of intermediate features into the fully connected layer in the second classifier to perform a fully connected operation to perform a local classification operation.
[0012] In yet another embodiment, training a detection model for detecting an abnormal area in a fundus image by using the fully supervised loss function and the semi-supervised loss function includes: calculating a first loss sum of the first fully supervised loss function and the first semi-supervised loss function; calculating a second loss sum of the second fully supervised loss function and the second semi-supervised loss function; and training a detection model for detecting an abnormal area in a fundus image based on the first loss sum and the second loss sum.
[0013] In yet another embodiment, the method further includes: visualizing weights in the attention module to determine weights corresponding to each of the image patches.
[0014] In a second aspect, the present application provides a device for training a detection model for detecting an abnormal area in a fundus image, including: a processor; a memory storing program instructions for training a detection model for detecting an abnormal area in a fundus image, and when the program instructions are executed by the processor, enabling the device to implement multiple embodiments in the foregoing first aspect.
[0015] In a third aspect, the present application further provides a method for detecting abnormal regions in fundus images, including: obtaining an ultra-wide field of view fundus image to be detected; inputting the ultra-wide field of view fundus image into a detection model trained according to multiple embodiments in the foregoing first aspect for detection, so as to output a detection result of the abnormal region in the fundus image.
[0016] In a fourth aspect, the present application further provides a device for detecting abnormal regions in fundus images, including: a processor; a memory storing program instructions for detecting abnormal regions in fundus images, and when the program instructions are executed by the processor, the device implements the embodiments in the foregoing third aspect.
[0017] In a fifth aspect, the present application further provides a computer-readable storage medium storing computer-readable instructions for training a detection model for detecting abnormal regions in fundus images and for detecting abnormal regions in fundus images. When the computer-readable instructions are executed by one or more processors, the embodiments in the foregoing first aspect and the embodiments in the foregoing third aspect are implemented.
[0018] Through the solution for training a detection model for detecting abnormal regions in fundus images provided as above, the embodiments of the present application obtain ultra-wide field of view fundus images with unlabeled detection frames and labeled detection frames, and crop them into multiple image blocks, and extract respective multiple intermediate features through a trained feature extractor. Then, by respectively inputting the multiple intermediate features into the first and second classifiers to perform global classification operations and local classification operations, and calculating corresponding fully supervised loss functions and semi-supervised loss functions, the detection model is trained. Based on this, the embodiments of the present application perform block processing on the ultra-wide field of view fundus image to reduce the loss of detailed information, making the information of the feature results richer.
[0019] Furthermore, the embodiments of the present application combine global and local classification, and train the detection model by calculating fully supervised and semi-supervised loss functions based on labeled detection frames and unlabeled detection frames, which can make the classification results more accurate and ensure the reliability of the detection model even in the case of lack of abnormal region labels. Subsequently, by using the detection model trained in the present application, accurate identification and positioning of abnormal regions in ultra-wide field of view fundus images can be achieved. In addition, the embodiments of the present application also visualize the weights in the attention module to view the image blocks where the abnormal regions are located, bringing great convenience to the detection of abnormal regions in clinical practice. Description of the Drawings
[0020] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present application will become readily understandable. In the drawings, several embodiments of the present application are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, where:
[0021] Figure 1 is an exemplary schematic diagram showing an ultra-wide-angle fundus image and a normal color fundus image;
[0022] Figure 2 is an exemplary flowchart showing a method for training a detection model for detecting abnormal regions in a fundus image according to an embodiment of the present application;
[0023] Figure 3 is an exemplary overall schematic diagram showing a detection model for detecting abnormal regions in a fundus image according to an embodiment of the present application;
[0024] Figure 4 is an exemplary schematic diagram showing an attention operation according to an embodiment of the present application;
[0025] Figure 5 is an exemplary schematic diagram showing visualization of weights in an attention module according to an embodiment of the present application;
[0026] Figure 6 is an exemplary flowchart showing a method for detecting abnormal regions in a fundus image according to an embodiment of the present application; and
[0027] Figure 7 is an exemplary structural block diagram showing a device for training a detection model for detecting abnormal regions in a fundus image and for detecting abnormal regions in a fundus image according to an embodiment of the present application. Detailed Embodiments
[0028] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the embodiments described in this specification are only some embodiments provided by the present application for the convenience of clear understanding of the solution and for compliance with legal requirements, rather than all embodiments that can implement the present application. All other embodiments obtained by those skilled in the art based on the embodiments disclosed in this specification without making creative efforts belong to the scope of protection of the present application.
[0029] Figure 1 is an exemplary schematic diagram showing an ultra-wide-angle fundus image and a normal color fundus image. As Figure 1 shown in the left figure in Figure 1The right middle figure shows a normal color fundus image. According to the description of the background technology, compared with the normal color fundus image, the imaging technology of the ultra-wide field of view fundus image is relatively simple, and a wider range of the retina can be obtained, which can provide more valuable information and data for clinical diagnosis and treatment planning. However, the existing methods usually process the entire image (i.e., a piece of image), which will result in more details being lost and inaccurate detection results.
[0030] In addition, the existing methods have less research on the specific location of the abnormal area, and often mark the detection box of the abnormal area in the ultra-wide field of view fundus image as the training label. However, due to the wide field of view and large amount of information of the ultra-wide field of view fundus image, marking the detection box of the abnormal area in the image will consume a large amount of human resources and time costs, and the public datasets of ultra-wide field of view fundus images with high-quality and large-scale pixel-level annotations are scarce. For small lesions, they become imperceptible after being scaled and input into the neural network model, so it brings certain difficulties to model training.
[0031] Based on this, the present application provides a solution for training a detection model for detecting abnormal areas in fundus images. By performing block processing on the ultra-wide field of view fundus images with unmarked detection boxes and marked detection boxes, the loss of details is reduced and the training difficulty of the detection model is lowered. Further, by respectively inputting multiple intermediate features extracted by the feature extractor into the first and second classifiers to perform global classification operations and local classification operations, and calculating the corresponding fully supervised loss function and semi-supervised loss function. Thus, through the combination of global and local classification, the classification result is more accurate, and the reliability of the detection model can be ensured even in the case of lack of abnormal area annotation. Subsequently, through the trained detection model, the abnormal areas in the ultra-wide field of view fundus images can be accurately identified and located.
[0032] The following will be combined with Figures 2 - 5 to describe in detail the solution for training a detection model for detecting abnormal areas in fundus images according to the embodiments of the present application.
[0033] Figure 2 is an exemplary flowchart showing a method 200 for training a detection model for detecting abnormal areas in fundus images according to an embodiment of the present application. As Figure 2As shown in [Figure 0], at step 201, ultra-widefield fundus images for training a detection model are obtained. In one implementation scenario, the aforementioned detection model can be, for example, a ResNet-50 model, and this detection model can at least include a feature extractor and a classifier. In some embodiments, the aforementioned feature extractor can be composed of, for example, multiple convolutional layers, Dropout layers, activation functions, and Batch Normalization layers, and the aforementioned classifier can include, for example, fully connected layers. In one embodiment, the aforementioned ultra-widefield fundus images can be acquired, for example, by an ultra-widefield fundus camera, and the ultra-widefield fundus images include ultra-widefield fundus images without labeled detection boxes and ultra-widefield fundus images with labeled detection boxes. Among them, the ultra-widefield fundus images with labeled detection boxes include ultra-widefield fundus images with labeled detection boxes and labeled abnormal classifications, and the ultra-widefield fundus images without labeled detection boxes include ultra-widefield fundus images with labeled abnormal classifications but without labeled detection boxes and ultra-widefield fundus images without labeled detection boxes and without labeled abnormal classifications.
[0034] In other words, the ultra-widefield fundus images with labeled detection boxes in this application are ultra-widefield fundus images that are both labeled with detection boxes and labeled with abnormal classifications. The ultra-widefield fundus images without labeled detection boxes include ultra-widefield fundus images that are only labeled with abnormal classifications but not with detection boxes and ultra-widefield fundus images without any labels (i.e., neither labeled with detection boxes nor labeled with abnormal classifications). That is, the ultra-widefield fundus images in this application include three categories: those that are both labeled with detection boxes and labeled with abnormal classifications, those that are only labeled with abnormal classifications, and those without any labels. It can be understood that the aforementioned labeled detection boxes refer to marking the specific location of the abnormal area in the ultra-widefield fundus image with, for example, a rectangular box, and the aforementioned labeled abnormal classification refers to marking whether there is an abnormal area in the ultra-widefield fundus image with, for example, "1" or "0". For example, when there is an abnormal area, it is labeled "1", and when there is no abnormal area, it is labeled "0". In an actual application scenario, the abnormal area in the embodiments of this application can be, for example, the lesion area in the ultra-widefield fundus image.
[0035] Based on the obtained ultra-widefield fundus images, at step 202, the ultra-widefield fundus images without labeled detection boxes and the ultra-widefield fundus images with labeled detection boxes are cropped to obtain multiple image patches for each. That is, the aforementioned three categories of ultra-widefield fundus images are each cropped from the whole image into several image patches (for example, Figure 3 as shown in [Figure 0], each category of ultra-widefield fundus image is cropped into four image patches). In some embodiments, before cropping each category of ultra-widefield fundus image into image patches, the ultra-widefield fundus images can be first subjected to data cleaning, that is, removing ultra-widefield fundus images with poor quality.
[0036] Next, at step 203, the trained feature extractor is used to perform feature operations related to abnormal area detection on multiple image patches to obtain respective multiple intermediate features. In one embodiment, the trained feature extractor is obtained through the following operations: performing data augmentation operations on multiple image patches to obtain data-augmented image patches, inputting the data-augmented image patches and multiple image patches into the feature extractor for feature operations, and calculating a contrast loss function to train the feature extractor using the contrast loss function to obtain the trained feature extractor. In some implementation scenarios, the aforementioned data augmentation operations may include but are not limited to horizontal flipping of images, image scale scaling, and / or changing image colors, etc.
[0037] Based on the data-augmented image patches and the original (without data augmentation) multiple image patches, the data-augmented image patches and the original multiple image patches are used as training samples to train the feature extractor. Specifically, the feature extractor is used to extract corresponding feature results respectively to calculate the contrast loss function, and then based on this contrast loss function, the feature extractor is trained forward and backward to obtain the trained feature extractor. In one implementation scenario, the contrast loss function can be calculated by the following formula:
[0038]
[0039] where, v i represents the feature result extracted by the feature extractor from the original ultra-widefield fundus image, represents the feature result extracted by the feature extractor after the corresponding image has undergone one data augmentation, represents other samples in the same batch of training. For example, taking v i as an example, which represents the feature result extracted by the feature extractor from the ultra-widefield fundus image with both detection boxes and abnormal classifications labeled, represents the feature result extracted by the feature extractor after the ultra-widefield fundus image with both detection boxes and abnormal classifications labeled has undergone one data augmentation, represents the ultra-widefield fundus image with only abnormal classifications labeled and / or the ultra-widefield fundus image without labels in the same batch of training. Further, T represents a hyperparameter (its value is, for example, 2 - 5), and exp represents the exponential function. In this scenario, the contrast loss function can learn the features of the ultra-widefield fundus image without labels by pulling in the distances of the same sample in the feature space after different data augmentations and pulling away the distances from other samples in the feature space.
[0040] After obtaining the above-trained feature extractor, feature operations related to abnormal area detection can be performed on multiple image patches to obtain respective multiple intermediate features. In one embodiment, the trained feature extractor is further connected with a fully connected module. By using the trained feature extractor to perform feature operations related to abnormal area detection on multiple image patches and performing a fully connected operation on the results of the feature operations via the fully connected module, respective multiple intermediate features are obtained. That is to say, in the embodiment of the present application, multiple image patches corresponding to various ultra-widefield fundus images are first input into the feature extractor to extract features, and then the extracted features are input into the fully connected module for a fully connected operation to obtain respective multiple intermediate features corresponding to various ultra-widefield fundus images.
[0041] Based on the multiple intermediate features obtained above, at step 204, the multiple intermediate features are respectively input into a first classifier to perform a global classification operation and input into a second classifier to perform a local classification operation, and a full supervision loss function and a semi-supervision loss function are calculated based on whether a detection box is labeled. In one embodiment, the multiple intermediate features are input into the first classifier to perform a global classification operation, and a first full supervision loss function is calculated based on the labeled detection box and a first semi-supervision loss function is calculated based on the unlabeled detection box, and the multiple intermediate features are input into the second classifier to perform a local classification operation, and a second full supervision loss function is calculated based on the labeled detection box and a second semi-supervision loss function is calculated based on the unlabeled detection box.
[0042] In one embodiment, both the aforementioned first classifier and the second classifier include a fully connected layer, and the first classifier further includes an attention module. Among them, for the first classifier, the multiple intermediate features are input into the attention module in the first classifier to perform an attention operation, and a fully connected operation is performed on the result of the attention operation via the fully connected layer in the first classifier to perform a global classification operation. For the second classifier, the multiple intermediate features are input into the fully connected layer in the second classifier to perform a fully connected operation to perform a local classification operation.
[0043] That is to say, the embodiment of the present application uses two classifiers to form two classification branches. One classification branch performs a global (i.e., for the whole image) classification operation via the attention module and the fully connected layer, and the other classification branch performs a local (i.e., for the local image) classification operation via the fully connected layer. The respective multiple intermediate features corresponding to various ultra-widefield fundus images are respectively input into the two classifiers, and corresponding loss functions are calculated based on whether a detection box is labeled in each classification branch. Specifically, for the labeled detection box, the first and second full supervision loss functions can be calculated based on the labeled detection box (i.e., the true label) and the classification results of the first and second classifiers; for the unlabeled detection box, the first and second semi-supervision loss functions can be calculated based on the pseudo label and the classification results of the first and second classifiers.
[0044] In an implementation scenario, for the labeled detection boxes, the first and second fully supervised loss functions can be calculated based on the following formula:
[0045]
[0046] where L Xi represents the first (or second) fully supervised loss function, p i represents the true label, represents the classification result of multiple intermediate features of the ultra-wide-angle fundus image of the labeled detection box after passing through the first (or second) classifier, and i represents the classification category.
[0047] In another implementation scenario, for the unlabeled detection boxes, the first and second semi-supervised loss functions can be calculated based on the following formula:
[0048]
[0049] where L Ui represents the first (or second) semi-supervised loss function, represents the pseudo-label, represents the classification result of multiple intermediate features of the ultra-wide-angle fundus image of the unlabeled detection box (including only labeled abnormal classification and no label) after passing through the first (or second) classifier, and i represents the classification category. In some implementation scenarios, when the value of the foregoing in category i is the maximum value among all categories and exceeds a pre-set threshold (for example, 0.5), it can be used as the foregoing for loss calculation. That is, the foregoing pseudo-label can select the one with the largest value among all categories and exceeding the pre-set threshold
[0050] Furthermore, at step 205, the detection model for detecting the abnormal area in the fundus image is trained using the fully supervised loss function and the semi-supervised loss function. In one embodiment, the first loss sum of the first fully supervised loss function and the first semi-supervised loss function and the second loss sum of the second fully supervised loss function and the second semi-supervised loss function can be calculated, and then the detection model for detecting the abnormal area in the fundus image is trained based on the first loss sum and the second loss sum. Specifically, the first loss sum is obtained by adding the first fully supervised loss function and the first semi-supervised loss function, the second loss sum is obtained by adding the second fully supervised loss function and the second semi-supervised loss function, and then the first loss sum and the second loss sum are subjected to an addition operation or a weighted sum operation to obtain the final total loss, so as to forward and backward train the detection model based on the final total loss to realize the training of the detection model for detecting the abnormal area in the fundus image.
[0051] As described above, in the embodiment of the present application, by obtaining ultra-wide-angle fundus images of unlabeled detection frames and labeled detection frames, and cropping them into multiple image patches, a trained feature extractor is used to extract multiple intermediate features of various ultra-wide-angle fundus images. Then, the multiple intermediate features of various ultra-wide-angle fundus images are input into two classification branches formed by the first and second classifiers to perform global classification operations and local classification operations respectively, and real labels and pseudo-labels are set based on whether there is a labeled detection frame, and the first and second fully supervised loss functions and the first and second semi-supervised loss functions are calculated according to the classification results of the corresponding classifiers for the multiple intermediate features of various ultra-wide-angle fundus images. Further, according to the sum of the losses of the first and second fully supervised loss functions and the sum of the losses of the first and second semi-supervised loss functions, an addition or weighted sum operation is performed to obtain the final total loss, so as to train the detection model for detecting abnormal areas in fundus images.
[0052] Based on this, in the embodiment of the present application, by cropping the overall image into multiple image patches and performing feature extraction operations on each image patch, the loss of detailed information can be reduced, making the information of the feature results more complete and rich. Further, in the embodiment of the present application, two classifiers are also set up to form two classification branches to perform global classification operations and local classification operations, and the corresponding fully supervised loss function and semi-supervised loss function are calculated. Thus, by combining global and local classification, the classification result is more accurate, and the reliability of the detection model can be ensured even in the case of lack of abnormal area annotation. Based on the trained detection model, accurate identification and positioning of abnormal areas in ultra-wide-angle fundus images can be achieved.
[0053] Figure 3 is an exemplary overall schematic diagram showing the training of a detection model for detecting abnormal areas in fundus images according to an embodiment of the present application. It should be understood that Figure 3 the training process shown is a specific embodiment of the above Figure 1 training method 100, so the above description about Figure 1 also applies to Figure 3 .
[0054] As Figure 3The ultra-wide-angle fundus images obtained are shown on the upper left side of the middle. It includes the ultra-wide-angle fundus image 301 with a detection box marked, and the ultra-wide-angle fundus image without a detection box marked (including the ultra-wide-angle fundus image 302 with abnormal classification marked but without a detection box marked and the ultra-wide-angle fundus image 303 without a detection box marked and without abnormal classification marked). That is, the obtained ultra-wide-angle fundus images include three types of ultra-wide-angle fundus images: the ultra-wide-angle fundus image 301 with both a detection box and an abnormal classification marked, the ultra-wide-angle fundus image 302 with only an abnormal classification marked, and the ultra-wide-angle fundus image 303 without any marks. In one embodiment, the above three types of ultra-wide-angle fundus images can be first subjected to data cleaning to remove the ultra-wide-angle fundus images with poor quality. Then, the above three types of ultra-wide-angle fundus images are cropped to obtain multiple image patches for each of them. In an exemplary scenario, assuming the above three types of ultra-wide-angle fundus imaging are denoted as S, and they are equally divided into multiple image patches ( "patch"), for example, denoted as s1, s2,..., s n , where n represents the number of image patches.
[0055] According to the foregoing, data augmentation operations (such as horizontal flipping of the image, image scale scaling, and / or changing the image color, etc.) can be performed on the multiple image patches of the above three types of ultra-wide-angle fundus images respectively to obtain the image patches after data augmentation. By using the image patches after data augmentation and the original multiple image patches as training samples 304 and inputting them into the feature extractor 305 for feature extraction, the corresponding feature extraction results can be obtained. For example, the feature result v extracted from the original ultra-wide-angle fundus image by the feature extractor is obtained i and the feature result extracted from the corresponding image after one data augmentation by the feature extractor Then, based on the above formula (1), the contrast loss function is calculated. By training the feature extractor forward and backward, the trained feature extractor 305 can be obtained. In one implementation scenario, the foregoing feature extractor 305 can be composed of, for example, multiple convolutional layers, Dropout layers, activation functions, and Batch Normalization layers, and a fully connected module 306 can be connected after the feature extractor 305.
[0056] In the implementation scenario, after the foregoing feature extractor performs feature operations related to abnormal area detection on multiple image patches s1, s2,..., s n , multiple initial features F can be obtained, where the multiple initial features F can be denoted as F = {f1, f2,..., f n}. It can be understood that after feature extraction, each image patch corresponds to one extracted feature. For example, the image patch s1 corresponds to the initial feature f1. Similarly, the image patch s2 corresponds to the initial feature f2, and the image patch s n corresponds to the initial feature f nNext, input multiple initial features into the fully connected module 306 for fully connected operations to obtain multiple corresponding intermediate features L (“Logits”) = {l1, l2,..., l n}. Among them, for the image with the detection box marked, the multiple intermediate features obtained correspondingly are denoted as L X and L X = {l X1 , l X2 ,..., l Xn}; for the image without the detection box marked, the multiple intermediate features obtained correspondingly are denoted as L U and L U = {l U1 , l U2 ,..., l Un}. Below, taking the example of cropping various ultra-widefield fundus images into four image patches will be described.
[0057] As Figure 3 shown in the lower middle, input multiple image patches (for example, four image patches are shown in the figure) 301-1 of the above-mentioned ultra-widefield fundus image 301 that is both marked with a detection box and an abnormal classification, multiple image patches 302-1 of the ultra-widefield fundus image 302 that is only marked with an abnormal classification, and multiple image patches 303-1 of the ultra-widefield fundus image 303 without any markings into the trained feature extractor 305 to perform feature operations related to abnormal area detection, and perform fully connected operations on the feature operation results via the fully connected module 306 to obtain their respective multiple intermediate features 301-2, multiple intermediate features 302-2, and multiple intermediate features 303-2.
[0058] Based on the multiple intermediate features 301-2 to 303-2 obtained above, input the multiple intermediate features 301-2 to 303-2 into the first classifier and the second classifier respectively. The first classifier includes an attention module 307 and a fully connected layer 308, and the second classifier includes a fully connected layer 309. For the first classifier, input the multiple intermediate features 301-2 to 303-2 into the attention module 307 in the first classifier to perform attention operations, and perform fully connected operations on the attention operation results via the fully connected layer 308 in the first classifier to obtain their respective corresponding classification results to perform global classification operations. For the second classifier, input the multiple intermediate features 301-2 to 303-2 into the fully connected layer 309 in the second classifier to perform fully connected operations to obtain their respective corresponding classification results to perform local classification operations.
[0059] In an implementation scenario, the above-mentioned attention module 307 may also include a fully-connected layer. Specifically, when performing the attention operation, an attention mechanism is added to each of the multiple intermediate features by using the attention module to obtain a new intermediate feature corresponding to each intermediate feature. Then, a product operation is performed on the new intermediate feature and the corresponding intermediate feature to obtain the corresponding attention operation result. The foregoing attention operation will be described in detail later in conjunction with Figure 4 Describe the foregoing attention operation in detail.
[0060] Furthermore, for the ultra-wide-angle fundus image 301 that is labeled with both detection boxes and abnormal classifications, based on the classification results of the true label and the corresponding multiple intermediate features 301-2 via the first and second classifiers, and according to the above formula (2), the first fully supervised loss function 310 and the second fully supervised loss function 311 are calculated. For the ultra-wide-angle fundus image 302 that is only labeled with abnormal classifications and the unlabeled ultra-wide-angle fundus image 303, based on the classification results of the pseudo-label and the respective corresponding multiple intermediate features 302-2 and multiple intermediate features 303-2 via the first and second classifiers, and according to the above formula (3), the first semi-supervised loss function 312 and the second semi-supervised loss function 313 are calculated. After obtaining the foregoing first fully supervised loss function 310 and second fully supervised loss function 311, and first semi-supervised loss function 312 and second semi-supervised loss function 313, an addition or weighted sum operation is performed based on the sum of the losses of the first and second fully supervised loss functions and the sum of the losses of the first and second semi-supervised loss functions to obtain the final total loss, so as to train the detection model for detecting abnormal areas in the fundus image. In some embodiments, the detection model may be, for example, a ResNet-50 model.
[0061] Figure 4 It is an exemplary schematic diagram showing the attention operation according to an embodiment of the present application. As Figure 4 shown, taking one of the multiple intermediate features 301-2 as an example, the intermediate feature 301-2 is input into the fully-connected layer in the attention module 307 to output a new intermediate feature 401. Then, a product operation (such as shown in the figure shown) is performed on the new intermediate feature 401 and the intermediate feature 301-2 to obtain the corresponding attention operation result 402. Similarly, by performing the foregoing operation on other intermediate features among the multiple intermediate features 301-2, the corresponding attention operation results can be obtained respectively. By performing the foregoing operation on each intermediate feature among the multiple intermediate features 302-2 and multiple intermediate features 303-2, the corresponding attention operation results can also be obtained respectively.
[0062] In one embodiment, the embodiments of the present application can also visualize the weights in the attention module to determine the weights corresponding to each image patch. Based on this, the probability value of the existence of an abnormal area in each image patch can be viewed to determine the image patch where the abnormal area is located and the position in the image patch where the abnormal area is located, thus bringing great convenience to the detection of abnormal areas in clinical practice. For example Figure 5 as shown
[0063] Figure 5 FIG. is an exemplary schematic diagram showing the visualization of the weights in the attention module according to an embodiment of the present application. As Figure 5 exemplarily shown in, there are four image patches, and the corresponding weights are shown below each image patch, that is, the probability value of the existence of an abnormal area in each image patch. For example Figure 5 the weight w1 of the image patch shown in the upper left in is 0.9, Figure 5 the weight w2 of the image patch shown in the upper right in is 0.04, Figure 5 the weight w3 of the image patch shown in the lower left in is 0.05, Figure 5 the weight w4 of the image patch shown in the lower right in is 0.01. According to the corresponding weights, it can be known that Figure 5 the probability of the existence of an abnormal area in the image patch shown in the upper left in is relatively large
[0064] Figure 6 FIG. is an exemplary flowchart of a method 600 for detecting an abnormal area in a fundus image according to an embodiment of the present application. As Figure 6 shown in, at step 601, an ultra-wide field of view fundus image to be detected is acquired. In one embodiment, the ultra-wide field of view fundus image can be acquired by, for example, an ultra-wide field of view fundus camera. Based on the acquired ultra-wide field of view fundus image, at step 620, the ultra-wide field of view fundus image is input into a trained detection model for detection to output the detection result of the abnormal area in the fundus image. By using the trained detection model of the present application for detection, accurate identification and positioning of the abnormal area in the ultra-wide field of view fundus image can be achieved. That is, the output detection result can include the classification result of the abnormal area and can also include the specific position of the abnormal area
[0065] Figure 7 FIG. is an exemplary structural block diagram of a device 700 for training a detection model for detecting an abnormal area in a fundus image and for detecting an abnormal area in a fundus image according to an embodiment of the present application. It can be understood that the device implementing the solution of the present application can be a single device (such as a computing device) or a multi-functional device including various peripheral devices
[0066] As Figure 7As shown in the figure, the device of the present application may include a central processor or central processing unit ("CPU") 711, which may be a general-purpose CPU, a dedicated CPU, or other information processing and program execution units. Further, the device 700 may also include a mass storage 712 and a read-only memory ("ROM") 713, where the mass storage 712 may be configured to store various types of data, including various ultra-widefield fundus images, algorithm data, intermediate results, and various programs required to run the device 700. The ROM 713 may be configured to store data and instructions for power-on self-test of the device 700, initialization of each functional module in the system, drivers for basic input / output of the system, and data and instructions required to boot the operating system.
[0067] Optionally, the device 700 may also include other hardware platforms or components, such as the shown tensor processing unit ("TPU") 714, graphics processing unit ("GPU") 715, field-programmable gate array ("FPGA") 716, and machine learning unit ("MLU") 717. It can be understood that although multiple hardware platforms or components are shown in the device 700, they are merely exemplary rather than restrictive, and those skilled in the art can add or remove corresponding hardware according to actual needs. For example, the device 700 may only include a CPU, related storage devices, and interface devices to implement the method of the present application for training a detection model for detecting abnormal regions in fundus images and the method for detecting abnormal regions in fundus images.
[0068] In some embodiments, to facilitate the transfer and interaction of data with an external network, the device 700 of the present application further includes a communication interface 718, so that it can be connected to a local area network / wireless local area network ("LAN / WLAN") 705 through the communication interface 718, and then can be connected to a local server 706 or connected to the Internet ("Internet") 707 through the LAN / WLAN. Alternatively or additionally, the device 700 of the present application may also be directly connected to the Internet or a cellular network based on wireless communication technology through the communication interface 718, such as based on the 3rd generation ("3G"), 4th generation ("4G"), or 5th generation ("5G") wireless communication technology. In some application scenarios, the device 700 of the present application may also access a server 708 and a database 709 of an external network as needed to obtain various known algorithms, data, and modules, and may remotely store various data, such as various types of data or instructions for presenting, for example, ultra-widefield fundus images.
[0069] The peripheral devices of device 700 may include a display device 702, an input device 703, and a data transmission interface 704. In one embodiment, the display device 702 may include, for example, one or more speakers and / or one or more visual displays configured to provide voice prompts and / or image / video displays for the training of the detection model for detecting abnormal regions in fundus images and the detection of abnormal regions in fundus images in the present application. The input device 703 may include, for example, a keyboard, a mouse, a microphone, a gesture capture camera, and other input buttons or controls configured to receive input of audio data and / or user instructions. The data transmission interface 704 may include, for example, a serial interface, a parallel interface, or a Universal Serial Bus interface ("USB"), a Small Computer System Interface ("SCSI"), Serial ATA, FireWire, PCI Express, and a High-Definition Multimedia Interface ("HDMI"), etc., configured for data transmission and interaction with other devices or systems. According to the solution of the present application, the data transmission interface 704 may receive ultra-widefield fundus images collected by an ultra-widefield fundus camera and transmit to device 700 ultra-widefield fundus images or various other types of data or results.
[0070] The above-mentioned CPU 711, mass storage 712, ROM 713, TPU 714, GPU 715, FPGA 716, MLU 717, and communication interface 718 of device 700 of the present application may be interconnected with each other through a bus 719 and achieve data interaction with the peripheral devices through this bus. In one embodiment, through this bus 719, the CPU 711 may control other hardware components and their peripheral devices in device 700.
[0071] As described above in conjunction with Figure 7 devices that can be used to execute the method for training the detection model for detecting abnormal regions in fundus images and for detecting abnormal regions in fundus images of the present application have been described. It should be understood that the device structure or architecture here is merely exemplary, and the implementation manner and implementation entity of the present application are not limited by it, but may be changed without departing from the spirit of the present application.
[0072] According to the above description in conjunction with the drawings, those skilled in the art can also understand that the embodiments of the present application can also be implemented by software programs. Thus, the present application also provides a computer-readable storage medium. This computer-readable storage medium can be used to implement the methods for training the detection model for detecting abnormal regions in fundus images and for detecting abnormal regions in fundus images described in the present application in conjunction with the Figure 2 and Figure 6 drawings.
[0073] It should be noted that although the operations of the method of the present application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the order of the steps depicted in the flowchart can be changed. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.
[0074] It should be understood that when terms such as "first", "second", "third", and "fourth" are used in the claims, the specification, and the accompanying drawings of the present application, they are only used to distinguish different objects and not to describe a specific order. The terms "comprising" and "including" used in the specification and claims of the present application indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0075] It should also be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and claims of the present application, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms. It should be further understood that the term "and / or" used in the specification and claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0076] Although the embodiments of the present application are as above, the above content is only an example adopted for the convenience of understanding the present application and is not intended to limit the scope and application scenarios of the present application. Any person skilled in the art within the technical field of the present application can make any modifications and changes in the form of implementation and details without departing from the spirit and scope disclosed by the present application. However, the scope of patent protection of the present application shall still be subject to the scope defined by the appended claims.
Claims
1. A method for training a detection model for detecting abnormal regions in fundus images, wherein the detection model at least includes a feature extractor and a classifier, and the method includes: Obtaining ultra-wide field of view fundus images for training the detection model, wherein the ultra-wide field of view fundus images include ultra-wide field of view fundus images without labeled detection frames and ultra-wide field of view fundus images with labeled detection frames; Cropping the ultra-wide field of view fundus images without labeled detection frames and the ultra-wide field of view fundus images with labeled detection frames to obtain respective multiple image patches; Using the trained feature extractor to perform feature operations related to abnormal region detection on the multiple image patches to obtain respective multiple intermediate features; Inputting the multiple intermediate features into a first classifier to perform global classification operations and into a second classifier to perform local classification operations, and calculating a full-supervised loss function and a semi-supervised loss function based on whether there are labeled detection frames; and Training the detection model for detecting abnormal regions in fundus images by using the full-supervised loss function and the semi-supervised loss function.
2. The method according to claim 1, wherein the ultra-wide field of view fundus images with labeled detection frames include ultra-wide field of view fundus images with labeled detection frames and labeled abnormal classifications, and the ultra-wide field of view fundus images without labeled detection frames include ultra-wide field of view fundus images with labeled abnormal classifications but without labeled detection frames and ultra-wide field of view fundus images without labeled detection frames and without labeled abnormal classifications.
3. The method according to claim 1, wherein the trained feature extractor is obtained through the following operations: Performing data augmentation operations on the multiple image patches to obtain data-augmented image patches; Inputting the data-augmented image patches and the multiple image patches into the feature extractor to perform feature operations, and calculating a contrast loss function; and Training the feature extractor by using the contrast loss function.
4. The method according to claim 3, wherein a fully-connected module is further connected to the trained feature extractor, and using the trained feature extractor to perform feature operations related to abnormal region detection on the multiple image patches to obtain respective multiple intermediate features includes: Using the trained feature extractor to perform feature operations related to abnormal region detection on the multiple image patches, and performing a fully-connected operation on the feature operation results via the fully-connected module to obtain respective multiple intermediate features.
5. The method according to claim 1, wherein inputting the multiple intermediate features into a first classifier to perform global classification operations and into a second classifier to perform local classification operations, and calculating a full-supervised loss function and a semi-supervised loss function based on whether there are labeled detection frames includes: Inputting the multiple intermediate features into the first classifier to perform global classification operations, and calculating a first full-supervised loss function based on the corresponding labeled detection frames and a first semi-supervised loss function based on the corresponding unlabeled detection frames; And Input the multiple intermediate features into the second classifier to perform local classification operations, and calculate a second fully supervised loss function corresponding to the labeled detection boxes and a second semi-supervised loss function corresponding to the unlabeled detection boxes.
6. The method according to claim 5, wherein the first classifier and the second classifier include fully connected layers, and the first classifier further includes an attention module. Inputting the multiple intermediate features into the first classifier to perform global classification operations and inputting the multiple intermediate features into the second classifier to perform local classification operations includes: Input the multiple intermediate features into the attention module in the first classifier to perform attention operations, and perform a fully connected operation on the result of the attention operation via the fully connected layer in the first classifier to perform global classification operations; And Input the multiple intermediate features into the fully connected layer in the second classifier to perform a fully connected operation to perform local classification operations.
7. The method according to claim 5, wherein training a detection model for detecting abnormal regions in fundus images using the fully supervised loss function and the semi-supervised loss function includes: Calculate a first total loss of the first fully supervised loss function and the first semi-supervised loss function; Calculate a second total loss of the second fully supervised loss function and the second semi-supervised loss function; And Train a detection model for detecting abnormal regions in fundus images based on the first total loss and the second total loss.
8. The method according to claim 6, further comprising: Visualize the weights in the attention module to determine the weights corresponding to each of the image patches.
9. An apparatus for training a detection model for detecting abnormal regions in fundus images, comprising: A processor; A memory storing program instructions for training a detection model for detecting abnormal regions in fundus images, which when executed by the processor cause the apparatus to implement the method according to any one of claims 1-8.
10. A method for detecting abnormal regions in fundus images, comprising: Obtain an ultra-wide field of view fundus image to be detected; Input the ultra-wide field of view fundus image into a detection model trained by the method according to any one of claims 1-8 for detection to output a detection result of the abnormal region in the fundus image.
11. An apparatus for detecting abnormal regions in fundus images, comprising: A processor; A memory storing program instructions for detecting abnormal regions in fundus images, which when executed by the processor cause the apparatus to implement the method according to claim 10.
12. A computer-readable storage medium storing computer-readable instructions for training a detection model for detecting abnormal regions in fundus images and for detecting abnormal regions in fundus images, which when executed by one or more processors implement the method according to any one of claims 1-8 and the method according to claim 10.
Citation Information
Patent Citations
Abnormality recognition method and device based on eye fundus image, equipment and storage medium
CN110210286A
Training method and training device for glaucoma recognition based on multiple features
CN116343008A