Method and product for training ultra-wide-angle fundus image processing model
Through deep learning technology that combines pre-training and fine-tuning, the dependence on large-scale labeled data in ultra-wide-angle fundus image processing is solved, the model's generalization and multi-tasking capabilities are achieved, and the accuracy and efficiency of eye disease detection and systemic disease screening are improved.
Patent Information
- Application Number
- CN202510701773.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-10-17
AI Technical Summary
Existing deep learning models rely on large-scale, precisely labeled data in ultra-wide-angle fundus image processing, which makes training difficult and it is difficult to fully mine the information of multi-task and multi-modal data.
By combining pre-trained visual language models with self-supervised learning and multimodal information, a large number of unlabeled ultra-wide-angle fundus images are used for pre-training. In the fine-tuning stage, a small number of precisely labeled images are used, and joint training is performed with downstream task models to achieve model generalization and feature extraction.
A model with strong generalization ability is trained using small-scale annotated data, which can automatically extract and analyze the complex features of ultra-wide-angle fundus images, improve the accuracy of eye disease detection and support systemic disease screening, and support multi-task processing such as disease detection and lesion classification.
Smart Images

Figure CN120808096A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application generally relates to the technical field of computer vision. More specifically, the present application relates to a method, an electronic device and a computer-readable storage medium for training an ultra-widefield fundus image processing model. Further, the present application also relates to a method, an electronic device and a computer-readable storage medium for generating an ultra-widefield fundus image. Further, the present application also relates to a method, an electronic device and a computer-readable storage medium for processing an ultra-widefield fundus image. BACKGROUND
[0002] Ultra-Widefield Fundus Imaging is an important breakthrough in the field of ophthalmic medical imaging. It can capture up to 200 degrees of retinal range in a single imaging, far exceeding the 30-50 degree field of view limitation of traditional fundus photography. This technology realizes comprehensive observation of the structure from the fovea to the peripheral retina through the imaging characteristics of large field of view and high resolution, significantly improving the diagnostic efficiency of eye diseases. Ultra-widefield fundus images contain a variety of biomarkers such as microaneurysms, hemorrhages, exudates, neovascularization, retinal breaks and detachment, and pigment abnormalities, providing important evidence for early detection, diagnosis and prognosis evaluation of eye diseases and systemic diseases.
[0003] With the rapid development of deep learning technology in the field of medical image diagnosis, the method of fundus image analysis based on convolutional neural network has achieved remarkable results in the automatic detection and classification of diseases such as diabetic retinopathy, macular degeneration and glaucoma. However, the existing deep learning model faces the following technical bottlenecks: first, the model training relies on large-scale labeled data, while the ultra-widefield fundus image dataset is relatively scarce, especially the lack of large-scale samples with accurate labeling, which limits the training of deep learning model and affects the generalization ability of the model; second, the existing model has limitations in processing multi-task and multi-modal data, making it difficult to fully exploit the rich information of ultra-widefield fundus images.
[0004] Therefore, it is urgent to provide a scheme for training an ultra-widefield fundus image processing model, so as to reduce the dependence on large-scale accurate labeled data, train a model with strong generalization ability, and realize automatic extraction and analysis of complex features of ultra-widefield fundus images, and improve the adaptability and accuracy of the model in processing information in a wide retinal area. SUMMARY
[0005] In order to at least solve one or more of the above-mentioned technical problems, the present application provides a scheme for processing an ultra-widefield fundus image in the following aspects.
[0006] In a first aspect, the present application provides a method for training an ultra-wide angle fundus image processing model, comprising: obtaining a first ultra-wide angle fundus image, multi-modal information paired with the first ultra-wide angle fundus image, a second ultra-wide angle fundus image, and annotation information of the second ultra-wide angle fundus image; inputting the first ultra-wide angle fundus image and the multi-modal information as training data into a visual language model for pre-training to obtain a pre-trained visual language model; combining the pre-trained visual language model with a downstream task model to obtain the ultra-wide angle fundus image processing model; inputting the second ultra-wide angle fundus image and the annotation information as training data into the ultra-wide angle fundus image processing model for fine-tuning to obtain a final ultra-wide angle fundus image processing model.
[0007] In some embodiments, the multi-modal information includes image description and / or diagnostic report, the image description includes image quality and patient information, and the diagnostic report contains lesion information and diagnostic opinion; the annotation information includes one or more of quality grading, disease type, disease grading, lesion site, lesion type, lesion location, and optic cup and disc segmentation mask.
[0008] In some embodiments, before inputting the first ultra-wide angle fundus image as training data into the visual language model and before inputting the second ultra-wide angle fundus image as training data into the ultra-wide angle fundus image processing model, the method further comprises: performing regularization processing on the green channel and the red channel of the first ultra-wide angle fundus image and the second ultra-wide angle fundus image.
[0009] In some embodiments, inputting the first ultra-wide angle fundus image and the multi-modal information as training data into the visual language model for pre-training comprises: inputting the first ultra-wide angle fundus image into an image encoder of the visual language model for feature extraction to obtain image features; inputting the multi-modal information into a text encoder of the visual language model for feature extraction to obtain text features; and pre-training the visual language model using a contrastive learning method based on the image features and the text features.
[0010] In some embodiments, the downstream task model in the super-wide-angle fundus image processing model is a decoder or a classifier, and the fine-tuning of the super-wide-angle fundus image processing model by inputting the second super-wide-angle fundus image and the annotation information as training data comprises: inputting part of the training data into the super-wide-angle fundus image processing model, jointly training the pre-trained visual language model and the downstream task model to obtain a jointly trained visual language model and downstream task model; freezing the weights of the image encoder and the text encoder in the jointly trained visual language model, and inputting the other part of the training data into the super-wide-angle fundus image processing model to train the jointly trained downstream task model.
[0011] In a second aspect, the present application provides a method for generating a super-wide-angle fundus image, comprising: obtaining an image generation target in text form; inputting the image generation target into the super-wide-angle fundus image processing model trained according to the method and multiple embodiments thereof in the first aspect to perform an image generation operation, so as to output a super-wide-angle fundus image meeting the image generation target, wherein the downstream task model in the super-wide-angle fundus image processing model is a decoder.
[0012] In a third aspect, the present application provides an electronic device, comprising: a processor; and
[0013] a memory storing program instructions for training a super-wide-angle fundus image processing model for generating a super-wide-angle fundus image, which, when executed by the processor, causes the implementation of the method according to the first aspect or the second aspect and multiple embodiments thereof.
[0014] In a fourth aspect, the present application provides a device for processing a super-wide-angle fundus image, comprising: a processor; and a memory storing program instructions for processing a super-wide-angle fundus image, which, when executed by the processor, causes the device to perform the following operations:
[0015] obtaining a super-wide-angle fundus image to be processed; inputting the super-wide-angle fundus image to be processed into the super-wide-angle fundus image processing model trained according to the method and multiple embodiments thereof in the first aspect to perform an image processing operation, so as to output an image processing result, wherein the downstream task model in the super-wide-angle fundus image processing model is a decoder or a classifier.
[0016] In a fifth aspect, the present application provides a computer-readable storage medium having stored thereon program instructions for training an ultra-wide-angle fundus image processing model, for generating an ultra-wide-angle fundus image, or for processing an ultra-wide-angle fundus image, which, when executed by a processor, implements the operations of the method of the first aspect, the second aspect, and the plurality of embodiments thereof, or the device of the third aspect.
[0017] The method for training an ultra-wide-angle fundus image processing model as provided above reduces the dependence of model training on large-scale accurately labeled data by using a large number of unlabeled ultra-wide-angle fundus images in the pre-training stage and a small amount of accurately labeled ultra-wide-angle fundus images in the fine-tuning stage. Compared with existing training methods that rely on large-scale labeled data, the scheme of the present application can train a model with strong generalization ability even in a small-scale labeled data scenario, effectively solving the problem of data scarcity.
[0018] In addition, by combining self-supervised learning, text-image multi-modal joint pre-training, and fine-tuning, the visual language model can learn the alignment relationship between images and text, better understand and generate text descriptions related to images, and obtain a visual language model with strong feature representation capability. Furthermore, through the combination of pre-training and fine-tuning, the model can fully adapt to the complexity of ultra-wide-angle fundus images, automatically extract and analyze rich information covering a wider range of retinal areas, which not only improves the detection accuracy of eye diseases such as diabetic retinopathy, macular degeneration, and glaucoma, but also provides support for early screening of systemic diseases.
[0019] Thereafter, in the application stage of the ultra-wide-angle fundus image processing model, based on the trained ultra-wide-angle fundus image processing model, only the text form of the image generation target needs to be input into the trained ultra-wide-angle fundus image processing model, and the ultra-wide-angle fundus image satisfying the image generation target can be obtained, effectively alleviating the problem of relative scarcity of ultra-wide-angle fundus image datasets and providing a new approach for data expansion and model optimization. In addition, the trained ultra-wide-angle fundus image processing model has multi-task processing capability and can simultaneously support disease detection, lesion classification, lesion positioning, image segmentation, and image quality analysis, etc. Only by inputting the ultra-wide-angle fundus image to be processed into the trained ultra-wide-angle fundus image processing model, the corresponding image processing result can be quickly obtained, realizing end-to-end prediction, improving the processing efficiency of high-resolution, large-size ultra-wide-angle fundus images, and providing strong technical support for the automation analysis and clinical application of eye diseases. BRIEF DESCRIPTION OF DRAWINGS
[0020] The above and other objects, features and advantages of the exemplary embodiments of the present application will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present application are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0021] Figure 1 A schematic diagram showing an ultra-wide-angle fundus image according to an embodiment of the present application is shown;
[0022] Figure 2 An exemplary flowchart of a method for training an ultra-wide-angle fundus image processing model according to an embodiment of the present application is shown;
[0023] Figure 3 An exemplary flow chart of a method for generating an ultra-wide-angle fundus image according to an embodiment of the present application is shown;
[0024] Figure 4 An exemplary flowchart of an operation for processing an ultra-wide-angle fundus image according to an embodiment of the present application is shown;
[0025] Figure 5 An exemplary structural block diagram of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0027] It should be understood that the terms "include" and "comprising" used in the description and claims of this application indicate the presence of described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof.
[0028] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this specification and claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should also be further understood that the term "and / or" as used in this specification and claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.
[0029] As used in the specification and claims, the term “if’ can be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” depending on the context. Similarly, the phrase “if it is determined” or “if [a described condition or event] is detected” can be construed to mean “upon determining” or “in response to determining” or “upon detecting [the described condition or event]” or “in response to detecting [the described condition or event],” depending on the context.
[0030] For ease of understanding, before the technical solutions in the embodiments of the present application are clearly and completely described, the technical terms involved in the present application are introduced in detail.
[0031] The base model generally refers to a model pre-trained on large-scale data, which can learn general feature representations. The base model can be applied to various fields and tasks, including natural language processing, computer vision, speech recognition, etc. For example, BERT (Bidirectional Encoder Representations from Transformers) is a base model in the field of natural language processing, and ViT (Vision Transformer) is a base model in the field of computer vision. In addition, the architecture of the base model can be diverse, such as convolutional neural network (CNN), Transformer, etc., depending on the needs of the task and the characteristics of the data.
[0032] The aforementioned ViT is a base model that applies the Transformer architecture to the field of computer vision, which performs well in tasks such as image classification, object detection, and semantic analysis. Traditional convolutional neural networks (CNNs) perform well in image recognition tasks, while ViT innovatively leverages the successful experience of Transformers in natural language processing. Specifically, ViT divides the input image into fixed-size image blocks (patches), then flattens and maps each image block to a high-dimensional embedding vector as the input sequence of the Transformer encoder. Through the self-attention mechanism, ViT can capture the relationships and dependencies between different regions in the image in a global range. Experiments show that ViT, when trained on large-scale datasets, can achieve or even exceed the performance of traditional CNNs in tasks such as image classification. In addition, as a base model, ViT can be pre-trained on large-scale image datasets, and then fine-tuned on specific tasks.
[0033] For ease of understanding, the input stage, intermediate process, and output stage of the ViT model are described in detail as follows.
[0034] Input stage: The input of the ViT model is a two-dimensional image, which is first divided into fixed-size image patches. For example, an image of size 224×224 is divided into 16×16 image patches, resulting in a total of 196 image patches. Each image patch is flattened into a one-dimensional vector and then mapped to a fixed-dimensional embedding space through a linear transformation (fully connected layer) to form an embedded representation of the image patch. In order to preserve position information, a learnable position embedding is added to each image patch. In addition, a special classification token ([CLS]Token) is added to the beginning of the input sequence, and its embedded representation will be used for the final classification task.
[0035] Intermediate Process: The processed input sequence is fed into the Transformer encoder, which typically consists of multiple stacked encoder layers. Each encoder layer includes a multi-head self-attention mechanism and a feedforward neural network. The multi-head self-attention mechanism globally calculates correlations between embeddings, capturing long-range dependencies between image patches and learning global image features. The feedforward neural network applies a nonlinear transformation to the embeddings at each position, further enhancing the model's expressive power. Residual connections and layer normalization are used after each sublayer to stabilize training and mitigate the vanishing gradient problem.
[0036] Output stage: After processing by the Transformer encoder, an embedded representation of the classification token ([CLS]Token) at the beginning of the sequence is extracted. This embedded representation passes through a classification head (typically a fully connected layer), which outputs a probability distribution corresponding to each category, thus completing the image classification task. Ultimately, the model determines the category of the input image based on the output probability distribution, achieving image recognition and classification.
[0037] A visual language model is a multimodal model that integrates visual and linguistic information. It can understand and generate text descriptions related to images, or understand the content of an image based on the text description. A visual language model typically consists of a visual encoder, a language encoder, and a fusion module for fusing visual and linguistic features. The visual encoder is used to extract image features, such as ViT or other visual models; the language encoder is used to process and understand text features, such as BERT, GPT (Generative Pre-trained Transformer), or a custom Transformer decoder; and the fusion module is used to fuse image features and text features. Common fusion methods include cross-attention mechanism, splicing, and contrastive learning.
[0038] Next, combine Figure 1 The ultra-wide-angle fundus image of the embodiment of the present application is introduced. Figure 1 FIG. 1 shows a schematic diagram of an ultra-wide-angle fundus image according to an embodiment of the present application.Figure 1 As shown, the ultra-widefield fundus image shows a variety of biomarkers and lesion sites, such as the Optic Disc, Macula, Vein, and Artery. In addition, the image also shows some lesions related to diabetic retinopathy, such as Retinal Detachment, Retinal Hemorrhages, Exudates, Hemangioma, and New Vessels Elsewhere, etc. For the sake of brevity, other possible lesions shown in the ultra-widefield fundus image are not listed one by one.
[0039] As can be known from the foregoing description, due to the limited field of view, ordinary fundus imaging may miss lesions in the peripheral retina, such as retinal breaks, peripheral degeneration, etc., and thus is suitable for general ophthalmic examination and diagnosis of common diseases. In contrast, ultra-widefield imaging can comprehensively capture information of the peripheral retina and is more advantageous in complex cases and when a comprehensive assessment of the retinal condition is needed. Through comprehensive examination of a large range of retina, ultra-widefield imaging can early detect the effects of systemic diseases such as diabetes and hypertension on the eye, achieve early intervention, and prevent the occurrence of serious complications. In addition, in the process of obtaining the scheme of the present application, the inventors found that the ultra-widefield fundus image is a projection of the fundus on a two-dimensional plane captured by a camera. Unlike other eye scanning techniques such as optical coherence tomography and angiography, the ultra-widefield fundus image can be obtained in a non-invasive and cost-effective manner, and is more suitable for large-scale screening.
[0040] The inventors also found that, in combination with machine learning and deep learning techniques, the ultra-widefield fundus image can fully play its rich retinal information. By training a large number of ultra-widefield fundus images, the model can learn the characteristics of various fundus structures and lesions, thereby developing intelligent diagnostic devices. These devices can automatically identify and analyze lesions in the fundus image, such as retinal detachment, retinal hemorrhage, exudates, hemangioma, and new vessels, lesions related to diabetic retinopathy, and other possible lesions. This not only improves the efficiency of diagnosis, but also significantly improves the accuracy of diagnosis, providing strong support for early detection and intervention of eye diseases.
[0041] The specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0042] Figure 2An exemplary flow chart of a method 200 for training an ultra-wide-angle fundus image processing model according to an embodiment of the present application is shown. It is understood that the method 200 can be executed by any suitable device with data processing capabilities, including but not limited to a terminal device, a processor, and a server.
[0043] like Figure 2 As shown, at step S201, method 200 can obtain a first ultra-wide-angle fundus image, multimodal information paired with the first ultra-wide-angle fundus image, a second ultra-wide-angle fundus image, and annotation information of the second ultra-wide-angle fundus image. At step S202, method 200 can input the first ultra-wide-angle fundus image and the multimodal information as training data into a visual language model for pre-training to obtain a pre-trained visual language model. Then, at step S203, method 200 can combine the pre-trained visual language model with the downstream task model to obtain an ultra-wide-angle fundus image processing model. At step S204, method 200 can input the second ultra-wide-angle fundus image and the annotation information as training data into the ultra-wide-angle fundus image processing model for fine-tuning to obtain the final ultra-wide-angle fundus image processing model.
[0044] In step S201, the first ultra-wide-angle fundus image obtained is a component of training data used to pre-train the visual language model, while the second ultra-wide-angle fundus image is a component of training data used to fine-tune the ultra-wide-angle fundus image processing model. In an embodiment of the present application, the difference between the first ultra-wide-angle fundus image and the second ultra-wide-angle fundus image is that the first ultra-wide-angle fundus image has no annotated information, while the second ultra-wide-angle fundus image has accurate annotated information.
[0045] The multimodal information paired with the first ultra-wide-angle fundus image may include an image description and / or a diagnostic report. The image description may further include image quality and patient information. In practical applications, the image description may be a description of the image quality and patient information in natural language. Specifically, the image quality may describe the clarity, contrast, brightness, noise level, etc. of the image, and indicate whether the image has problems such as blur, overexposure, underexposure, excessive noise, etc. The patient information may include items such as the patient's age, gender, left eye / right eye, and medical history. The aforementioned diagnostic report may be an eye disease diagnosis report given by a doctor or a machine, which may include lesion information and diagnostic opinions.
[0046] Further, the lesion information can include a lesion site, and a lesion type and a lesion position of the lesion site, etc. The lesion site refers to a region or tissue organ of the eye where pathological changes occur. For example, the lesion site can be a relatively wide area such as the entire retina, the optic disc, the macular area, etc. The lesion is a local structure with specific pathological characteristics in the lesion site. In the fundus examination, if it is found that there are hemorrhagic spots, exudation spots, cotton wool spots or drusen on the retina, these are specific lesions. The lesion has clear characteristics such as shape, size, color, etc., and can more accurately reflect the specific situation of the disease. For example, in diabetic retinopathy, the lesion site can be the entire retina, but the lesion is the specific pathological changes such as the scattered microaneurysms, hemorrhagic spots, exudation spots, etc.
[0047] The diagnostic opinion clearly gives the diagnostic result based on the ultra-wide-angle fundus image, which includes whether there is an eye disease, and in the case of an eye disease, the grading of the eye disease, for example, in the case of the eye disease being diabetic retinopathy, the diagnostic result can also include the diabetic retinopathy grading.
[0048] The annotation information of the second ultra-wide-angle fundus image can include one or more of quality grading, disease type, disease grading, lesion site, lesion type, lesion position, and optic cup and disc segmentation mask. The quality grading is used to identify the usability of the image for deep learning model training and for clinical diagnosis, which can be any one of excellent, usable, and unusable. The disease type is used to identify the specific eye disease suffered by the patient, and the disease grading is used to identify the level of the eye disease.
[0049] As an example, when the disease classification is diabetic retinopathy, the disease grading can be any one of grade 0, grade 1, grade 2, grade 3, and grade 4. Here, grade 0 indicates no diabetic retinopathy, grade 1 indicates mild non-proliferative diabetic retinopathy, grade 2 indicates moderate non-proliferative diabetic retinopathy, grade 3 indicates severe non-proliferative diabetic retinopathy, and grade 4 indicates proliferative diabetic retinopathy. In particular, when the patient does not suffer from an eye disease, the disease type of the second ultra-wide-angle fundus image can be annotated as 0, and the corresponding disease grading is also annotated as 0.
[0050] The lesion site is used to identify the region or tissue organ of the second ultra-wide-angle fundus image where pathological changes occur, the lesion type is used to identify whether there are hemorrhage, exudation, cotton wool spots and / or drusen in the second ultra-wide-angle fundus image. The lesion position is used to identify the position of the lesion in the second ultra-wide-angle fundus image, and can be identified by the bounding box coordinates of the lesion. The optic cup and disc segmentation mask is a binary image corresponding to the size of the second ultra-wide-angle fundus image, a total of two, corresponding to the optic cup and disc regions in the image respectively.
[0051] In the embodiments of the present application, the first and second ultra-wide-angle fundus images can be dual-channel pseudo-color images (including only green and red channels) or three-channel pseudo-color images (including green, red and yellow channels). In the green channel of the ultra-wide-angle fundus image, the contrast between the retinal blood vessels and the background is the highest, which makes the blood vessel structure clearer and is conducive to subsequent blood vessel segmentation and analysis. The red channel can better highlight the lesion areas on the retina, such as hemorrhagic spots, etc. Thus, the red channel focuses more on capturing lesions and abnormalities on the retina, while the green channel is more conducive to the extraction and analysis of blood vessel structures. In contrast, the yellow channel usually has no significant advantage in the fundus image, and its information content is relatively small, and in some cases it can introduce noise or interference. Therefore, in actual processing, the yellow channel is usually ignored or directly discarded to reduce the computational complexity and improve the processing efficiency.
[0052] Based on this, before inputting the first ultra-wide-angle fundus image as training data into the visual language model and before inputting the second ultra-wide-angle fundus image as training data into the ultra-wide-angle fundus image processing model, the green and red channels of the first and second ultra-wide-angle fundus images can be subjected to regularization processing to unify the image scale and reduce the overfitting possibility of the model. The regularization processing can adjust the pixel values of the green and red channels to a specific range or distribution, which helps to unify the scales of different images and reduce the pixel value differences caused by different image acquisition conditions (such as light intensity, contrast, etc.), so that the model is more robust to changes in image brightness, contrast, etc. during the training process. In addition, the green and red channels of the ultra-wide-angle fundus image can contain some redundant information, and the regularization processing can remove or weaken these redundant features, so that the model pays more attention to the key information in the image, thereby reducing the risk of overfitting of the model to the training data.
[0053] In actual operation, in addition to the above-mentioned regularization processing, before inputting the first ultra-wide-angle fundus image as training data into the visual language model and before inputting the second ultra-wide-angle fundus image as training data into the ultra-wide-angle fundus image processing model, the first and second ultra-wide-angle fundus images can also be subjected to brightness adjustment to improve the brightness and contrast of the images, enhance the details and edge information of the images, and make the lesion features in the images more prominent and clear, which helps the model to more accurately identify and classify lesions. In actual application, any one of gamma correction, Laplacian enhancement and adaptive histogram equalization can be used to adjust the brightness of the first and second ultra-wide-angle fundus images.
[0054] At the step S202, the first ultra-wide-angle fundus image and the multi-modal information are input into the visual language model for pre-training. Specifically, the first ultra-wide-angle fundus image is input into an image encoder of the visual language model for feature extraction to obtain image features; the multi-modal information is input into a text encoder of the visual language model for feature extraction to obtain text features; and the visual language model is pre-trained based on the image features and the text features by using a contrastive learning manner. In the embodiments of the present application, the visual language model can be a CLIP (Contrastive Language-Image Pre-Training) model.
[0055] As a typical visual language model, the core mechanism of CLIP (Contrastive Language-Image Pre-Training) is to map image features and text features to a shared semantic feature space through contrastive learning, thereby realizing cross-modal semantic alignment and joint understanding.
[0056] In terms of structural design, CLIP adopts a dual-encoder architecture: the visual encoder is usually based on the VisionTransformer (ViT) model, which divides the input image into fixed-size image patches, and inputs these patch sequences together with position encodings into the Transformer to extract the global semantic features of the image through multiple layers of self-attention mechanism; the language encoder is based on the Transformer architecture, and common implementations include BERT variants or customized Transformer encoders, whose processing flow is as follows: first, the input text is tokenized and tagged to generate a token sequence containing position information, and then the context dependency of the text is captured through multiple Transformer layers to finally output an embedding vector representing the semantic of the text.
[0057] In actual operation, when the visual language model is pre-trained based on image features and text features using a contrastive learning manner, the model tries to pull the matching image features and text features closer in the common semantic space, while pushing the non-matching image features and text features further apart. Specifically, a contrastive loss function (such as InfoNCE loss) can be used to optimize the model parameters. The loss function calculates the similarity between image features and text features, and updates the model by maximizing the similarity between similar samples and minimizing the similarity between non-similar samples. The pre-trained CLIP model can efficiently handle image-text matching tasks, supporting zero-shot image classification, cross-modal retrieval, image-text generation, and other downstream tasks, and its architecture design and training strategy lay an important foundation for the development of multi-modal models.
[0058] The target of the downstream task model at the foregoing step S203 is a specific downstream task, such as disease detection, lesion classification, lesion positioning, image generation, image segmentation, and image quality analysis. In some embodiments of the present application, the number of the foregoing downstream task models is one or more, the tasks supported by each downstream task model are different from each other, and each downstream task model has its own task-specific layer, such as a classification layer or a regression layer.
[0059] The output of the downstream task model for implementing the disease detection is whether the eye disease is present, and in the case of the eye disease, the disease type and the disease classification of the eye disease. The output of the downstream task model for implementing the lesion classification task is the lesion type, such as one or more of hemorrhage, exudation, cotton wool patch, and drusen. The output of the downstream task model for implementing the lesion positioning task is the lesion position, such as the bounding box coordinates of one or more of hemorrhage, exudation, cotton wool patch, and drusen.
[0060] In addition, the output of the downstream task model for implementing the image generation task is the ultra-wide field fundus image meeting the image generation target. The output of the downstream task model for implementing the image segmentation task is two binary images, respectively corresponding to the cup and disc regions in the image. The output of the downstream task model for implementing the image quality analysis task is the quality classification, such as any one of excellent, usable, and unusable.
[0061] In the embodiments of the present application, by jointly using the pre-trained ultra-wide field fundus image model and one or more downstream task models, simultaneous fine-tuning can be achieved, so that the feature representation of the pre-trained ultra-wide field fundus image model can be shared among different tasks, which helps to improve the generalization ability and data efficiency of the model. In actual operation, in order to balance the learning difficulty and convergence speed among different tasks, the task weights of the downstream task models can be dynamically adjusted to promote the balanced convergence of the downstream task models.
[0062] In actual application, in order to meet the different downstream task requirements described above, the downstream task model can be a decoder or a classifier. The decoder is suitable for tasks such as lesion positioning, image generation, and image segmentation, and can realize generation and positioning by decoding image features into specific coordinates, masks, or images. The classifier is suitable for tasks such as disease detection, lesion classification, and image quality analysis, and can realize classification by mapping image features to predefined type labels.
[0063] At the foregoing step S203, when inputting the second ultra-wide-angle fundus image and the annotation information as training data into the ultra-wide-angle fundus image processing model for fine-tuning, part of the training data can be first input into the ultra-wide-angle fundus image processing model, and the pre-trained visual language model and the downstream task model are jointly trained, so as to simultaneously optimize the parameters of the pre-trained visual language model and the downstream task model, so that the model can learn the feature representation of the image and the text and the task-specific prediction ability at the same time, until the model converges, and the jointly trained visual language model and the downstream task model can be obtained.
[0064] Next, the weights of the image encoder and the text encoder in the jointly trained visual language model can be frozen, and the other part of the training data is input into the ultra-wide-angle fundus image processing model to train the jointly trained downstream task model. During the training process, the parameters of the downstream task model are adjusted according to the specific task target and data, so that the downstream task model can better complete the downstream task, until the training is completed, and the final ultra-wide-angle fundus image processing model can be obtained.
[0065] The above Figure 2 The method for training the ultra-wide-angle fundus image processing model is described, which reduces the dependence of model training on large-scale accurate annotation data by using a large amount of unannotated ultra-wide-angle fundus images in the pre-training stage and a small amount of accurate annotated ultra-wide-angle fundus images in the fine-tuning stage. Compared with the existing training method that depends on large-scale annotation data, the scheme of the present application can train a model with strong generalization ability even in a small-scale annotation data scenario, effectively solving the problem of data scarcity.
[0066] In addition, by combining self-supervised learning, text-image multi-modal joint pre-training and fine-tuning and other deep learning technologies, the visual language model can learn the alignment relationship between the image and the text, so as to better understand and generate text descriptions related to the image, and obtain a visual language model with strong feature representation capability. Furthermore, through the combination of pre-training and fine-tuning, the model can fully adapt to the complexity of the ultra-wide-angle fundus image, automatically extract and analyze rich information covering a wider retinal area, which not only improves the detection accuracy of the model for diseases such as diabetic retinopathy, macular degeneration and glaucoma, but also provides support for early screening of systemic diseases.
[0067] Subsequently, by integrating the trained ultra-wide-angle fundus image processing model into medical equipment and software, intelligent fundus image analysis equipment can be developed. This device can realize the automated interpretation and diagnosis of ultra-wide-angle fundus images, significantly improving the accuracy and efficiency of medical diagnosis. In addition, the model can also be applied to portable medical devices or mobile applications to support eye screening and diagnosis in telemedicine and primary healthcare institutions. At the same time, with the model's multi-task learning capabilities, the model can also be extended to other ophthalmic image analysis fields. In this way, it can provide comprehensive early detection of eye diseases and personalized treatment plans, bringing broader and deeper application value to ophthalmic medical care.
[0068] Combined with the above Figures 1 to 2 A method for training an ultra-wide-angle fundus image processing model is described. Accordingly, the present application also provides a method for generating an ultra-wide-angle fundus image. Figure 3 , which shows an exemplary flow chart of a method 300 for generating an ultra-wide-angle fundus image according to some embodiments of the present application. It is understood that the method 300 can be executed by any appropriate device with data processing capabilities, including but not limited to a processor, a terminal device, and a server.
[0069] like Figure 3 As shown, at step S301, method 300 can obtain an image generation target in text form. The image generation target usually uses natural language to describe the characteristics of the ultra-wide-angle fundus image to be generated. The characteristics of the image may include but are not limited to left eye / right eye, eye structure, image quality, disease type and disease grade, presence or absence of obvious lesions, lesion site, lesion type, lesion location, and the area where the optic cup and optic disc are located. It can be understood that the clearer and more detailed the description of the image generation target, the more realistic the ultra-wide-angle fundus image generated by the ultra-wide-angle fundus image processing model based on it. In actual applications, those skilled in the art can customize the image generation target according to actual data requirements or model training needs.
[0070] Then, at step S302, method 300 can input the image generation target into the ultra-wide-angle fundus image processing model trained according to the method for training the ultra-wide-angle fundus image processing model described in the previous embodiment to perform an image generation operation to output an ultra-wide-angle fundus image that meets the image generation target. Here, the downstream task model in the ultra-wide-angle fundus image processing model is a decoder.
[0071] Combination of the above Figure 3The method for generating the ultra-wide-angle fundus image is described, based on the trained ultra-wide-angle fundus image processing model, only the image generation target in the form of text is input into the trained ultra-wide-angle fundus image processing model, the ultra-wide-angle fundus image meeting the image generation target can be obtained, the problem of relatively insufficient ultra-wide-angle fundus image dataset is effectively alleviated, and a new way is provided for data expansion and model optimization.
[0072] Next, in combination with Figure 4 An electronic device 400 provided by the embodiment of the application is exemplarily introduced. As shown in the figure, the electronic device 400 of the embodiment of the application can include a processor 401, a memory 402 and a communication bus 403. Figure 4
[0073] In the process of the specific embodiment, the processor 401 can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing image processing device (DSPD), a programmable logic image processing device (PLD), a field programmable gate array (FPGA), a CPU, a controller, a microcontroller, and a microprocessor. It can be understood that, for different devices, the electronic device for realizing the function of the processor can also be other electronic devices, and the embodiment is not limited in detail.
[0074] In the embodiment of the application, the communication bus 403 is used to realize the connection and communication between the processor 401 and the memory 402; the memory 402 stores program instructions for training the ultra-wide-angle fundus image processing model or for generating the ultra-wide-angle fundus image; and the processor 401 executes the program instructions stored in the memory 402 to realize the method for training the ultra-wide-angle fundus image processing model described in combination with Figures 1 to 2 the method for generating the ultra-wide-angle fundus image described in combination with Figure 3 the method for training the ultra-wide-angle fundus image processing model.
[0075] The above describes the method for training the ultra-wide-angle fundus image processing model, or the method for generating the ultra-wide-angle fundus image described in combination with Figure 4 This disclosure describes an electronic device capable of executing program instructions for training an ultra-wide-angle fundus image processing model or generating ultra-wide-angle fundus images according to the present application. It should be understood that the device structure or architecture described herein is merely exemplary, and the implementation and implementation of the present application are not limited thereto and may be modified without departing from the spirit of the present application.
[0076] As described above, the trained ultra-wide-angle fundus image processing model has multi-tasking capabilities and can simultaneously support multiple tasks such as disease detection, lesion classification, lesion localization, image segmentation, and image quality analysis, providing strong technical support for the automated analysis and clinical application of eye diseases.
[0077] Next, combine Figure 5 The following describes devices for processing ultra-wide-angle fundus images according to other embodiments of the present application. The device for processing ultra-wide-angle fundus images has the same device structure or architecture as the device 400 described above. The difference is that the device includes a memory storing program instructions for processing ultra-wide-angle fundus images. When the program instructions are executed by the processor, the device performs the following operation 500 for processing the ultra-wide-angle fundus image:
[0078] like Figure 5 As shown, at step S501, an ultra-wide-angle fundus image to be processed can be acquired. Next, at step S502, the ultra-wide-angle fundus image to be processed can be input into an ultra-wide-angle fundus image processing model trained according to the method for training an ultra-wide-angle fundus image processing model described in the previous embodiment for image processing, thereby outputting an image processing result. The downstream task model in the ultra-wide-angle fundus image processing model is a decoder or a classifier. In practical applications, the image processing result may include, but is not limited to, one or more of quality grading, disease type, disease grade, lesion type, lesion location, and optic cup and disc segmentation mask.
[0079] In actual operation, the acquired to-be-processed ultra-wide-angle fundus image can be a dual-channel pseudo-color image or a three-channel pseudo-color image. In order to improve the accuracy of the image processing result, the to-be-processed ultra-wide-angle fundus image can be preprocessed before being input into the model. The preprocessing manner can include but is not limited to normalizing the green channel and the red channel of the to-be-processed ultra-wide-angle fundus image and adjusting the brightness of the to-be-processed ultra-wide-angle fundus image. Through the normalization, the redundant information in the green channel and the red channel can be removed or weakened, so that the model pays more attention to the key information in the image; through the brightness adjustment, the brightness and contrast of the image can be improved, the details and edge information of the image can be enhanced, and the lesion features in the image can be more prominent and clear, which helps the model to more accurately identify and classify the lesion, thereby improving the accuracy of the image processing result.
[0080] The above description is combined with Figure 5 The device for processing the ultra-wide-angle fundus image is described, the trained ultra-wide-angle fundus image processing model has multi-task processing capability and can simultaneously support multiple tasks such as disease detection, lesion classification, lesion positioning, image segmentation and image quality analysis. When analyzing and clinically applying the eye diseases, the to-be-processed ultra-wide-angle fundus image only needs to be input into the trained ultra-wide-angle fundus image processing model, and the image processing result, i.e., the eye disease analysis result, can be obtained, realizing end-to-end prediction, improving the processing efficiency of the high-resolution and large-size ultra-wide-angle fundus image, and providing strong technical support for the automatic analysis and clinical application of the eye diseases.
[0081] It can be understood that the description of each embodiment of the present disclosure emphasizes the differences between each embodiment, and the same or corresponding parts can be referred to each other. For the purpose of brevity, the present disclosure will not be repeated.
[0082] According to the above description combined with the drawings, those skilled in the art can also understand that the embodiments of the present application can also be implemented by software programs. Therefore, the present application also provides a computer readable storage medium. The computer readable storage medium stores program instructions for training an ultra-wide-angle fundus image processing model, generating an ultra-wide-angle fundus image, or processing an ultra-wide-angle fundus image, which can be used to implement the methods described above for training an ultra-wide-angle fundus image processing model, generating an ultra-wide-angle fundus image, or processing an ultra-wide-angle fundus image. Figures 1 to 2 The described method for training an ultra-wide-angle fundus image processing model, the described method for generating an ultra-wide-angle fundus image, or the described device for processing an ultra-wide-angle fundus image can be implemented by the program instructions stored in the computer readable storage medium. Figure 3 The described method for generating an ultra-wide-angle fundus image, or the described device for processing an ultra-wide-angle fundus image can be implemented by the program instructions stored in the computer readable storage medium. Figure 5 The described device for processing an ultra-wide-angle fundus image implements the operations described above.
[0083] It should be noted that, although the operations of the methods of the present application are described in a particular, sequential order for convenient presentation, some operations can be performed concurrently, out of order, or parallel to one another. Furthermore, some operations can be optional. Additionally, it should be noted that some operations can be performed multiple times, by one or more components, and / or can be performed by one component while one or more other components are performing other operations. Additionally or alternatively, some operations can be performed in a different order, combined, omitted, and / or divided into multiple operations.
[0084] While several embodiments of the application have been shown and described herein, it is to be understood that all the terms used herein are descriptive and not limiting, and that many changes, modifications, and substitutions can be made by one having ordinary skill in the art without departing from the idea and scope of the application. It is to be understood that in some instances some features of this application can be employed without a corresponding use of other features, and it is further understood that features described herein can be used in any combination. It is the intent, therefore, to be limited only by the scope of the claims and the patent laws.
Claims
1. A method for training an ultra-wide-angle fundus image processing model, comprising: Acquire a first ultra-wide-angle fundus image, multimodal information paired with the first ultra-wide-angle fundus image, a second ultra-wide-angle fundus image, and annotation information of the second ultra-wide-angle fundus image; Inputting the first ultra-wide-angle fundus image and the multimodal information as training data into a visual language model for pre-training to obtain a pre-trained visual language model; Combining the pre-trained visual language model with a downstream task model to obtain the ultra-wide-angle fundus image processing model; The second ultra-wide-angle fundus image and the annotation information are input as training data into the ultra-wide-angle fundus image processing model for fine-tuning to obtain a final ultra-wide-angle fundus image processing model.
2. The method according to claim 1, wherein The multimodal information includes an image description and / or a diagnostic report, wherein the image description includes image quality and patient information, and the diagnostic report contains lesion information and diagnostic opinions; the annotation information includes one or more of quality grade, disease type, disease grade, lesion location, lesion type, lesion location, and optic cup and disc segmentation mask.
3. The method according to claim 1, before inputting the first ultra-wide-angle fundus image as training data into a visual language model and before inputting the second ultra-wide-angle fundus image as training data into the ultra-wide-angle fundus image processing model, the method further comprises: Regularization processing is performed on the green channel and the red channel of the first ultra-wide-angle fundus image and the second ultra-wide-angle fundus image.
4. The method according to any one of claims 1 to 3, wherein: Inputting the first ultra-wide-angle fundus image and the multimodal information as training data into a visual language model for pre-training includes: Inputting the first ultra-wide-angle fundus image into the image encoder of the visual language model to perform feature extraction to obtain image features; Inputting the multimodal information into a text encoder of the visual language model to perform feature extraction to obtain text features; and Based on the image features and the text features, the visual language model is pre-trained using a contrastive learning approach.
5. The method according to claim 4, wherein The downstream task model in the ultra-wide-angle fundus image processing model is a decoder or a classifier, and the second ultra-wide-angle fundus image and the annotation information are input as training data into the ultra-wide-angle fundus image processing model for fine-tuning, including: Inputting part of the training data into the ultra-wide-angle fundus image processing model, and jointly training the pre-trained visual language model and the downstream task model to obtain a jointly trained visual language model and downstream task model; Freeze the weights of the image encoder and the text encoder in the jointly trained visual language model, and input the remaining data in the training data into the ultra-wide-angle fundus image processing model to train the jointly trained downstream task model.
6. A method for generating an ultra-wide-angle fundus image, comprising: Get the image generation target in text form; The image generation target is input into the ultra-wide-angle fundus image processing model trained according to the method described in any one of claims 1 to 5 to perform an image generation operation to output an ultra-wide-angle fundus image that meets the image generation target, wherein the downstream task model in the ultra-wide-angle fundus image processing model is a decoder.
7. An electronic device comprising: processor; as well as A memory storing program instructions for training an ultra-wide-angle fundus image processing model and for generating an ultra-wide-angle fundus image, wherein when the program instructions are executed by a processor, the method according to any one of claims 1 to 5 or the method according to claim 6 is implemented.
8. A device for processing ultra-wide-angle fundus images, comprising: processor; and a memory storing program instructions for processing an ultra-wide-angle fundus image, wherein when the program instructions are executed by the processor, the device implements the following operations: Acquiring an ultra-wide-angle fundus image to be processed; The ultra-wide-angle fundus image to be processed is input into an ultra-wide-angle fundus image processing model trained according to the method according to any one of claims 1 to 5 for image processing to output an image processing result, wherein the downstream task model in the ultra-wide-angle fundus image processing model is a decoder or a classifier.
9. A computer-readable storage medium storing program instructions for training an ultra-wide-angle fundus image processing model, for generating an ultra-wide-angle fundus image, or for processing an ultra-wide-angle fundus image, wherein when the program instructions are executed by a processor, the operations implemented by the method according to any one of claims 1 to 5, the method according to claim 6, or the device according to claim 8 are implemented.