Eye fundus disease detection model training method and device based on multi-view feature fusion

CN122737656APending Publication Date: 2026-09-11NING BO EYE HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610888400.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0004]然而,使用ImageNet数据集进行预训练,ImageNet数据集包含的是大量自然图像,这与医疗图像存在显著差异,这将导致预训练模型性能不佳,部分类型无法识别,不能很好地适应眼底疾病的识别任务

Benefits of technology

[0019]与现有技术相比,本发明提供的基于多视角特征融合的眼底疾病检测模型训练方法及装置,通过采用基于多个无标签的超广角眼底图像确定的第一数据集,对初始图像重建模型进行自监督训练,得到训练好的目标图像重建模型,能够有效利用眼科医疗废弃影像数据,采用自监督学习技术,使目标图像重建模型学习眼底疾病通用表征,实现变废为宝;通过分类层和目标图像重建模型中包含多视角特征融合层的编码器,构建初始眼底疾病检测模型,能够通过使用学习眼底疾病通用表征的目标图像重建模型中的编码器,结合分类层构成初始眼底疾病检测模型;采用基于开源数据集和/或至少一个有标签的超广角眼底图像确定的基于第二数据集,对初始眼底疾病检测模型进行监督训练,得到训练好的目标眼底疾病检测模型,能够显著提升目标眼底疾病检测模型对眼底疾病检测的准确性和疾病智能诊断性能,进而有利于促进临床应用。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122737656A_ABST
    Figure CN122737656A_ABST
Patent Text Reader

Abstract

The application provides a fundus disease detection model training method and device based on multi-view feature fusion, and belongs to the technical field of artificial intelligence. The method comprises the following steps: performing self-supervised training on an initial image reconstruction model based on a first data set to obtain a target image reconstruction model, wherein the first data set is determined based on multiple unlabeled ultra-wide-angle fundus images; constructing an initial fundus disease detection model based on a classification layer and an encoder in the target image reconstruction model; and performing supervised training on the initial fundus disease detection model based on a second data set to obtain a trained target fundus disease detection model, wherein the second data set is determined based on an open source data set and / or at least one labeled ultra-wide-angle fundus image. Through the application, medical waste image data of ophthalmology can be effectively utilized, the target image reconstruction model can learn general representations of fundus diseases, and waste can be turned into treasure; and the accuracy of the target fundus disease detection model for fundus disease detection and the intelligent disease diagnosis performance can be significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a training method and apparatus for a fundus disease detection model based on multi-view feature fusion. Background Technology

[0002] Artificial intelligence (AI) can significantly improve the level of disease detection, progression monitoring, and personalized treatment in the field of ophthalmology.

[0003] However, a key challenge in developing robust AI models is their reliance on large-scale, precisely labeled datasets, which requires ophthalmologists to invest significant time and expertise in annotation. Therefore, current intelligent diagnostic models in ophthalmology typically use ultra-large-scale visual datasets (ImageNet) for pre-training, followed by fine-tuning (retraining) using eye images. In addition, optimization strategies such as label smoothing, label weights, and normalization are employed to adjust for class imbalances and improve model performance.

[0004] However, using the ImageNet dataset for pre-training, which contains a large number of natural images that differ significantly from medical images, leads to poor performance of the pre-trained model, failing to recognize some types of images and making it unsuitable for tasks involving the identification of fundus diseases. Therefore, an effective solution is urgently needed to address these issues. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides a training method and apparatus for a fundus disease detection model based on multi-view feature fusion.

[0006] In a first aspect, the present invention provides a method for training a fundus disease detection model based on multi-view feature fusion, comprising: The initial image reconstruction model is self-supervised and trained based on the first dataset to obtain a trained target image reconstruction model. The first dataset is determined based on multiple unlabeled ultra-wide-angle fundus images. Based on the classification layer and the encoder in the target image reconstruction model, an initial fundus disease detection model is constructed. The encoder includes a multi-view feature fusion layer, which is used for multi-view feature extraction and fusion. Based on the second dataset, the initial fundus disease detection model is trained under supervision to obtain a trained target fundus disease detection model. The second dataset is determined based on an open-source dataset and / or at least one labeled ultra-wide-angle fundus image.

[0007] In some embodiments, the step of performing self-supervised training on the initial image reconstruction model based on the first dataset to obtain a trained target image reconstruction model includes: For any first sample image in the first dataset, random occlusion of local content is performed on the first sample image to obtain a second sample image; The second sample image is input into the encoder of the initial image reconstruction model. The encoder performs cascaded encoding on each encoding layer to obtain the encoding features output by each encoding layer. The multi-view feature fusion layer then extracts and fuses the encoding features output by each encoding layer to obtain the visible features corresponding to the unmasked content in the second sample image. Based on the visible features and the occlusion content corresponding to the second sample image, the comprehensive features corresponding to the second sample image are determined. The integrated features are input into the decoder of the initial image reconstruction model to reconstruct the image, thereby obtaining the reconstructed image; Based on the first sample image and the reconstructed image, the model parameters of the initial image reconstruction model are adjusted to obtain the adjusted initial image reconstruction model; Continue training the adjusted initial image reconstruction model until the first training stopping condition is met, and obtain the trained target image reconstruction model.

[0008] In some embodiments, the multi-view feature fusion layer includes a multi-view feature extraction layer, a feature fusion layer, and a spatial transformation layer; The multi-view feature fusion layer extracts and fuses the encoded features output by each encoding layer to obtain visible features corresponding to the unoccluded content in the second sample image, including: The encoded features output by each of the encoding layers are input into the multi-view feature extraction layer to extract multi-view features, resulting in N target features, where N is less than or equal to the number of encoding layers contained in the encoder. Each of the target features is input into the feature fusion layer for feature fusion to obtain fused features; The fused features are input into the spatial transformation layer for spatial transformation to obtain the visible features corresponding to the second sample image.

[0009] In some embodiments, the encoder comprises M coding layers, where M is a positive integer greater than 1; The step of inputting the second sample image into the encoder of the initial image reconstruction model, and performing concatenated encoding through each encoding layer in the encoder to obtain the encoded features output by each encoding layer includes: The second sample image is input into the first coding layer for encoding processing to obtain the encoded features output by the first coding layer; The encoded features output from the first encoding layer are input into the second encoding layer for encoding processing to obtain the encoded features output from the second encoding layer; The encoded features output from the second encoding layer are input into the third encoding layer for encoding processing to obtain the encoded features output from the third encoding layer. This process continues until the encoded features output by the Mth encoding layer are obtained.

[0010] In some embodiments, determining the comprehensive features corresponding to the second sample image based on the visible features and the occlusion content corresponding to the second sample image includes: Determine the first location information of the occluded content and the second location information of the unoccluded content in the second sample image; Based on the first location information and the second location information, the visible features are concatenated with the empty features corresponding to the occluded content to obtain the comprehensive features corresponding to the second sample image.

[0011] In some embodiments, constructing an initial fundus disease detection model based on the classification layer and the encoder in the target image reconstruction model includes: Using the encoder in the target image reconstruction model as a feature extractor and a multilayer perceptron as a classification layer, the feature extractor and the classification layer are connected in series to obtain the initial fundus disease detection model.

[0012] In some embodiments, the step of supervising the training of the initial fundus disease detection model based on a second dataset to obtain a trained target fundus disease detection model includes: Based on the set ratio and the second dataset, the target training set, target validation set, and target test set are determined. Based on the target training set, the initial fundus disease detection model is subjected to multiple rounds of supervised training to obtain backup fundus disease detection models for each round. Based on the target validation set, each of the backup fundus disease detection models is validated to obtain the detection indicators of each backup fundus disease detection model; The standby fundus disease detection model that meets the testing conditions is determined as the target fundus disease detection model; After supervising the training of the initial fundus disease detection model based on the second dataset to obtain the trained target fundus disease detection model, the process further includes: Based on the target test set, the target fundus disease detection model is tested, and the target fundus disease detection model that passes the test is used for fundus disease detection.

[0013] In some embodiments, determining the target training set, target validation set, and target test set based on a set ratio and the second dataset includes: The second dataset is divided into an initial training set, an initial validation set, and an initial test set according to the set ratio. Data augmentation processing is performed on the initial training set, the initial validation set, and the initial test set to obtain the target training set, the target validation set, and the target test set.

[0014] In some embodiments, before performing self-supervised training on the initial image reconstruction model based on the first dataset to obtain the trained target image reconstruction model, the method further includes: Collect the aforementioned multiple unlabeled ultra-wide-angle fundus images; Using cubic difference, the plurality of unlabeled ultra-wide-angle fundus images are adjusted to a first set size, and the adjusted plurality of unlabeled ultra-wide-angle fundus images are converted to a set image format to obtain the first dataset, and / or, each first sample image in the first dataset is adjusted to a second set size; The step of performing self-supervised training on the initial image reconstruction model based on the first dataset to obtain a trained target image reconstruction model includes: Based on the first dataset or the first dataset after resizing, the initial image reconstruction model is trained under self-supervised conditions to obtain the target image reconstruction model.

[0015] Secondly, the present invention also provides a training device for a fundus disease detection model based on multi-view feature fusion, comprising: The self-supervised training module is configured to perform self-supervised training on the initial image reconstruction model based on the first dataset to obtain the trained target image reconstruction model. The first dataset is determined based on multiple unlabeled ultra-wide-angle fundus images. The building module is configured to construct an initial fundus disease detection model based on the classification layer and the encoder in the target image reconstruction model. The encoder includes a multi-view feature fusion layer, which is used for multi-view feature extraction and fusion. The supervised training module is configured to supervise the training of the initial fundus disease detection model based on a second dataset to obtain a trained target fundus disease detection model. The second dataset is determined based on an open-source dataset and / or at least one labeled ultra-wide-angle fundus image.

[0016] Thirdly, the present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement the training method for a fundus disease detection model based on multi-view feature fusion as described in the first aspect above.

[0017] Fourthly, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the training method for a fundus disease detection model based on multi-view feature fusion as described in the first aspect above.

[0018] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the training method for a fundus disease detection model based on multi-view feature fusion described in the first aspect above.

[0019] Compared with existing technologies, the method and apparatus for training a fundus disease detection model based on multi-view feature fusion provided by this invention, through the use of a first dataset determined based on multiple unlabeled ultra-wide-angle fundus images, performs self-supervised training on an initial image reconstruction model to obtain a trained target image reconstruction model. This effectively utilizes discarded ophthalmic medical image data, and the self-supervised learning technique enables the target image reconstruction model to learn general representations of fundus diseases, turning waste into treasure. An initial fundus disease detection model is constructed through a classification layer and an encoder containing a multi-view feature fusion layer in the target image reconstruction model. This allows the initial fundus disease detection model to be formed by combining the encoder in the target image reconstruction model, which learns general representations of fundus diseases, with the classification layer. A second dataset, determined based on an open-source dataset and / or at least one labeled ultra-wide-angle fundus image, is used to supervise the training of the initial fundus disease detection model, resulting in a trained target fundus disease detection model. This significantly improves the accuracy and intelligent diagnostic performance of the target fundus disease detection model, thereby promoting clinical applications.

[0020] Details of one or more embodiments of the present invention are set forth in the following drawings and description to make other features, objects and advantages of the invention more readily apparent. Attached Figure Description

[0021] The accompanying drawings, which are provided to further illustrate the invention, constitute a part of this invention. Those skilled in the art will recognize that other drawings can be derived from these drawings without any inventive effort. The illustrative embodiments and descriptions of the invention are used to explain the invention and do not constitute an undue limitation thereof.

[0022] Figure 1 This is a flowchart illustrating the training method for a fundus disease detection model based on multi-view feature fusion provided by the present invention.

[0023] Figure 2 This is a schematic diagram of the structure of the initial image reconstruction model provided by the present invention.

[0024] Figure 3 This is a schematic diagram of the structure of the initial fundus disease detection model provided by the present invention.

[0025] Figure 4 This is a schematic diagram of the structure of the fundus disease detection model training device based on multi-view feature fusion provided by the present invention.

[0026] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0027] To more clearly understand the objectives, technical solutions, and advantages of this invention, the technical solutions of this invention will be clearly and completely described and explained below with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are only some embodiments of this invention, not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0028] Unless otherwise defined, the technical or scientific terms used in this invention shall have the general meaning understood by one of ordinary skill in the art to which this invention pertains. Words such as “a,” “an,” “an,” “the,” “the,” and “these” used in this invention do not indicate quantitative limitation and may be singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this invention are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include steps or modules (units) not listed, or may include other steps or modules (units) inherent to these processes, methods, products, or devices. The terms “connected,” “linked,” and “coupled” used in this invention are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. “A plurality” used in this invention refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. Normally, the character " / " indicates that the objects before and after it are in an "or" relationship. The terms "first," "second," "third," etc., used in this invention are merely for distinguishing similar objects and do not represent a specific ordering of the objects.

[0029] First, a brief description of the relevant content involved in this invention will be given.

[0030] Given the scarcity of medical professionals and the labor-intensive nature of manual annotation, only a small portion of available medical image data is currently annotated, while the vast majority of data remains unusable due to a lack of annotation. This results in low accuracy of artificial intelligence models, limiting the development of AI and hindering its clinical application.

[0031] With the development of ophthalmology in society, people are paying more and more attention to eye health. Ophthalmology clinics have accumulated a large amount of unlabeled data, which was previously called waste data and was not used properly.

[0032] With the rapid development of artificial intelligence and the rapid iteration of self-supervised algorithms in recent years, new ideas have been provided for the rational use of these discarded data. Based on this, this invention provides a training method for a fundus disease detection model based on multi-view feature fusion. It effectively utilizes "discarded" ophthalmic medical image data to improve the recognition performance of intelligent diagnostic models. Employing self-supervised learning technology, it trains an image reconstruction model using a large number of unlabeled ultra-wide-angle fundus images to learn general representations of fundus diseases, thus "turning waste into treasure." Using the encoder in the target image reconstruction model that has learned general representations of fundus diseases, combined with a classification layer, an initial fundus disease detection model is constructed. Supervised training of this initial model yields the final target fundus disease detection model. This does not affect the use of optimization strategies such as label smoothing, label weights, and normalization. Moreover, by using unlabeled data for training, the accuracy of model detection and the intelligent diagnostic performance of diseases can be significantly improved, thereby promoting clinical applications.

[0033] The following is combined with Figures 1 to 4 This invention describes a training method and apparatus for a fundus disease detection model based on multi-view feature fusion.

[0034] Figure 1 This is a flowchart illustrating the training method for a fundus disease detection model based on multi-view feature fusion provided by the present invention, as shown below. Figure 1 As shown, the training method for the fundus disease detection model based on multi-view feature fusion includes the following steps: Step 101: Perform self-supervised training on the initial image reconstruction model based on the first dataset to obtain a trained target image reconstruction model. The first dataset is determined based on multiple unlabeled ultra-wide-angle fundus images.

[0035] Specifically, the initial image reconstruction model is an untrained image reconstruction model. This model can be an improved version of the traditional Masked Autoencoders (MAE) algorithm, with a multi-view feature fusion layer added to the encoder. Alternatively, the image reconstruction model can be any other currently available or future self-supervised learning algorithm. Ultra-wide-angle fundus images are images that acquire a large area (200 degrees and above) of fundus information in a single scan using advanced laser scanning technology. These images clearly show the outermost edge (peripheral part) of the retina, effectively avoiding the omission of lesions.

[0036] In practical applications, with user authorization, a large number (e.g., 1 million) of unlabeled ultra-wide-angle fundus images can be collected from cooperating medical institutions. Then, for each unlabeled ultra-wide-angle fundus image, a first sample image corresponding to that unlabeled ultra-wide-angle fundus image is created; the first sample images corresponding to each unlabeled ultra-wide-angle fundus image form a first dataset.

[0037] In addition, the first dataset can also be constructed based on medical images of other modalities (such as ocular surface images and optical coherence tomography (OCT) images): for each other modality of medical image, a first sample image corresponding to that medical image is created and added to the first dataset.

[0038] Furthermore, the initial image reconstruction model can be self-supervised and trained based on the first dataset to obtain the target image reconstruction model.

[0039] It should be noted that when training the initial image reconstruction model, image enhancement techniques can be used to enhance the images on the first dataset to improve the performance of the image reconstruction model. These image enhancement techniques can include at least one of the following: image normalization, random horizontal flipping of the image, and image resizing (e.g., resizing the image to 224 pixels * 224 pixels).

[0040] In addition, the initial image reconstruction model and the target image reconstruction model contain an encoder and a decoder. The embedding size of the encoder can be set to 1024, and the embedding size of the decoder can be set to 512.

[0041] Self-supervised training of the initial image reconstruction model can employ batch training, with the batch size set according to requirements, such as 1536. The number of training epochs for the initial image reconstruction model can also be set based on needs, such as 800 epochs. Furthermore, a set number of epochs in the self-supervised training can be used to warm up the learning rate, which can be set according to requirements, such as using the first 40 epochs to warm up the learning rate (from 0 to 1.5 × 10⁻⁶). -4 ).

[0042] Step 102: Based on the classification layer and the encoder in the target image reconstruction model, construct an initial fundus disease detection model. The encoder contains a multi-view feature fusion layer, which is used for multi-view feature extraction and fusion.

[0043] Specifically, the classification layer can be a feedforward neural network with a normalized exponential function (softmax), or it can be a logistic regression (LR), a support vector machine (SVM), or a multi-layer perceptron (MLP), etc.

[0044] Specifically, multi-view feature fusion involves extracting features at different depths of the encoder (such as 3 / 6 / 9 / 12 layers) and fusing them through average pooling and linear layers to generate a multi-view feature representation that contains shallow, medium, and deep semantic information.

[0045] In practical applications, the encoder and classification layer in the target image reconstruction model can be combined to obtain an initial fundus disease detection model.

[0046] Step 103: Based on the second dataset, supervised training is performed on the initial fundus disease detection model to obtain a trained target fundus disease detection model. The second dataset is determined based on an open-source dataset and / or at least one labeled ultra-wide-angle fundus image.

[0047] Specifically, the open-source dataset can be at least one of the Deep Diabetic Retinopathy Image Dataset (DeepDRid) and datasets provided by the Tumor Online Prognostic Platform (TOPP).

[0048] Specifically, at least one labeled ultra-wide-angle fundus image can be a labeled ultra-wide-angle fundus image collected from a medical institution with user authorization. It should be noted that at least one labeled ultra-wide-angle fundus image needs to be anonymized, i.e., privacy information needs to be removed.

[0049] In practical applications, for each image in an open-source dataset and / or at least one labeled ultra-wide-angle fundus image, a third sample image can be created, and the label of that image can be used as the label of the third sample image. Each third sample image and its label form a second dataset.

[0050] Furthermore, based on the second dataset, the initial fundus disease detection model is trained under supervision to obtain the target fundus disease detection model.

[0051] For example, the supervised training process for an initial fundus disease detection model can be as follows: For any third sample image in the second dataset, the third sample image is input into the initial fundus disease detection model to detect fundus diseases and obtain the detection result; Based on the detection results and the labels of the third sample image, the loss value is determined; Based on the loss value, the model parameters of the initial fundus disease detection model are adjusted to obtain the adjusted initial fundus disease detection model; Continue training the adjusted initial fundus disease detection model until the second training stop condition is met, and obtain the trained target fundus disease detection model.

[0052] The second training stopping condition can be at least one of the following: the loss value is less than a loss threshold, the rate of change of the loss value is less than a loss rate of change threshold, or the number of iterations of self-supervised training reaches an iteration threshold. The loss function used to determine the loss value can be set based on requirements, such as using a classification task loss function.

[0053] It should be noted that supervised training of the initial fundus disease detection model can employ batch training, with the batch size set according to requirements, such as 32. The training epochs for supervised training of the initial fundus disease detection model can also be set based on requirements, such as 100. Furthermore, stochastic gradient descent can be used as the optimizer in supervised training, with the initial learning rate and weight decay set according to requirements, such as an initial learning rate of 0.01 and a weight decay of 1×10⁻⁶. -4 .

[0054] The present invention provides a training method for a fundus disease detection model based on multi-view feature fusion. It employs a first dataset determined from multiple unlabeled ultra-wide-angle fundus images to perform self-supervised training on an initial image reconstruction model, resulting in a trained target image reconstruction model. This method effectively utilizes discarded ophthalmic medical image data. By employing self-supervised learning, the target image reconstruction model learns general representations of fundus diseases, turning waste into treasure. An initial fundus disease detection model is constructed using a classification layer and an encoder containing a multi-view feature fusion layer within the target image reconstruction model. This allows the initial fundus disease detection model to be formed by combining the encoder in the target image reconstruction model, which learns general representations of fundus diseases, with the classification layer. A second dataset, determined from an open-source dataset and / or at least one labeled ultra-wide-angle fundus image, is then used to supervise the training of the initial fundus disease detection model, resulting in a trained target fundus disease detection model. This significantly improves the accuracy and intelligent diagnostic performance of the target fundus disease detection model, thereby facilitating clinical applications.

[0055] In some embodiments, before performing self-supervised training on the initial image reconstruction model based on the first dataset to obtain the trained target image reconstruction model, the method further includes: Collect the aforementioned multiple unlabeled ultra-wide-angle fundus images; Using cubic difference, the plurality of unlabeled ultra-wide-angle fundus images are adjusted to a first set size, and the adjusted plurality of unlabeled ultra-wide-angle fundus images are converted to a set image format to obtain the first dataset, and / or, each first sample image in the first dataset is adjusted to a second set size.

[0056] In practical applications, with user authorization, a large number of unlabeled ultrawide-angle fundus images can be collected from cooperating medical institutions. Then, to improve the speed of self-supervised training, cubic interpolation can be used to uniformly adjust all the collected unlabeled ultrawide-angle fundus images to a first set size, such as 512 pixels × 512 pixels, and all the resized unlabeled ultrawide-angle fundus images are converted to a set image format, such as the lossless format of Portable Network Graphics (PNG), thus obtaining the first dataset.

[0057] Furthermore, the initial image reconstruction model can be self-supervised trained based on the first dataset or the first dataset after resizing to obtain the target image reconstruction model.

[0058] In this embodiment of the invention, by setting the ultra-wide-angle fundus images in the first dataset to a uniform size and a uniform format, it is beneficial to improve the rate of self-supervised training, that is, to improve the acquisition efficiency of the target image reconstruction model, and thus improve the acquisition efficiency of the target fundus disease detection model determined based on the target image reconstruction model.

[0059] In some embodiments, before performing self-supervised training on the initial image reconstruction model based on the first dataset to obtain the trained target image reconstruction model, the method further includes: Adjust each first sample image in the first dataset to a second set size; The step of performing self-supervised training on the initial image reconstruction model based on the first dataset to obtain a trained target image reconstruction model includes: Based on the first dataset after resizing, the initial image reconstruction model is trained under self-supervised conditions to obtain the target image reconstruction model.

[0060] In practical applications, to improve the speed of self-supervised training, the size of the first sample images in the first dataset can be made more suitable for the requirements of the image reconstruction model. Therefore, the first dataset can be preprocessed by adjusting each first sample image in the first dataset to a second predetermined size, such as 224 pixels × 224 pixels. Here, the second predetermined size is the input size of the image reconstruction model. Furthermore, based on the resized first dataset, the initial image reconstruction model can be self-supervised trained to obtain the target image reconstruction model.

[0061] In some embodiments, before performing self-supervised training on the initial image reconstruction model based on the first dataset to obtain the trained target image reconstruction model, the method further includes: Collect the aforementioned multiple unlabeled ultra-wide-angle fundus images; Using cubic difference, the plurality of unlabeled ultra-wide-angle fundus images are adjusted to a first set size, and the adjusted plurality of unlabeled ultra-wide-angle fundus images are converted into a set image format to obtain the first dataset; Adjust each first sample image in the first dataset to a second set size; Based on the first dataset after resizing, the initial image reconstruction model is trained under self-supervised conditions to obtain the target image reconstruction model.

[0062] In practical applications, with user authorization, a large number of unlabeled ultra-wide-angle fundus images can be collected from collaborating medical institutions. Then, to improve the speed of self-supervised training, cubic interpolation can be used to uniformly resize all collected unlabeled ultra-wide-angle fundus images to a first predetermined size, such as 512 pixels × 512 pixels. All resized unlabeled ultra-wide-angle fundus images are then converted to a predetermined image format, such as the lossless format of Portable Network Graphics (PNG), thus obtaining the first dataset. The first predetermined size can be a general size used in image processing.

[0063] Furthermore, to improve the speed of self-supervised training, the size of the first sample images in the first dataset can be made more suitable for the image reconstruction model. Therefore, the first dataset can be preprocessed by adjusting each first sample image in the first dataset to a second predetermined size. This second predetermined size is the input size of the image reconstruction model. Then, based on the resized first dataset, the initial image reconstruction model is self-supervised trained to obtain the target image reconstruction model.

[0064] Thus, by setting the ultra-wide-angle fundus images in the first dataset to a uniform size and format, the first dataset becomes applicable to most image processing scenarios, improving its applicability. By adjusting each first sample image in the first dataset to a second set size, the self-supervised training rate is improved, i.e., the acquisition efficiency of the target image reconstruction model is improved, thereby improving the acquisition efficiency of the target fundus disease detection model determined based on the target image reconstruction model.

[0065] In some embodiments, the step of performing self-supervised training on the initial image reconstruction model based on the first dataset to obtain a trained target image reconstruction model includes: For any first sample image in the first dataset, random occlusion of local content is performed on the first sample image to obtain a second sample image; The second sample image is input into the encoder of the initial image reconstruction model. The encoder performs cascaded encoding on each encoding layer to obtain the encoding features output by each encoding layer. The multi-view feature fusion layer then extracts and fuses the encoding features output by each encoding layer to obtain the visible features corresponding to the unmasked content in the second sample image. Based on the visible features and the occlusion content corresponding to the second sample image, the comprehensive features corresponding to the second sample image are determined. The integrated features are input into the decoder of the initial image reconstruction model to reconstruct the image, thereby obtaining the reconstructed image; Based on the first sample image and the reconstructed image, the model parameters of the initial image reconstruction model are adjusted to obtain the adjusted initial image reconstruction model; Continue training the adjusted initial image reconstruction model until the first training stopping condition is met, and obtain the trained target image reconstruction model.

[0066] See Figure 2 , Figure 2 This is a schematic diagram of the structure of an initial image reconstruction model provided by the present invention: the initial image reconstruction model includes an encoder and a decoder. The encoder adopts a multi-layer coding layer structure, wherein the coding layer can be a Transformer, and the Transformer can be a vision Transformer. Furthermore, based on the characteristics of the eye, the encoder of the initial image reconstruction model adds a multi-view feature fusion layer on the basis of the traditional image reconstruction model to extract features of the image from different viewpoints and obtain a more fine-grained feature representation. The decoder can adopt an 8-layer Transformer structure for reconstructing the original image.

[0067] In practical applications, for each first sample image in the first dataset, the first sample image can be divided into blocks, that is, the first sample image is divided into M1×M2 image blocks, where M1 and M2 are both positive integers.

[0068] Next, a predetermined number of image blocks in the first sample image are randomly masked to obtain a second sample image corresponding to the first sample image. The predetermined number can be M3, where M3 is a positive integer; it can also be a value greater than 0 and less than 1, such as 75%, meaning 75% of the image blocks in the first sample image are randomly masked. The predetermined number can be set based on requirements, and this invention does not impose any limitations on it.

[0069] Then, the second sample image can be input into the encoder of the initial image reconstruction model. The encoder includes at least two cascaded coding layers and a multi-view feature fusion layer: first, the second sample image is cascaded and encoded by the at least two cascaded coding layers to obtain the encoded features output by each coding layer; then, the multi-view feature fusion layer performs feature extraction, fusion, and transformation processing on the encoded features output by each coding layer to obtain the visible features corresponding to the unmasked content in the second sample image.

[0070] For example, the size of the first sample image is 224 pixels × 224 pixels. The first sample image is divided into blocks, each block being 16 pixels × 16 pixels, that is, the first sample image is divided into (224 / 16) × (224 / 16) = 14 × 14 blocks. Then, 75% of the block is randomly occluded, that is, 14 * 14 * 75% = 147 blocks are randomly occluded. The visible 25% (14 * 14 * 25% = 49) of the block (the second sample image) is used as the input to the encoder. After the encoder, the features of the 25% visible block are obtained, which are the visible features.

[0071] Based on the visible features obtained, the visible features are concatenated and fused with the invisible features corresponding to the occluded content in the second sample image to obtain the comprehensive features corresponding to the second sample image. These comprehensive features are then input into the decoder of the initial image reconstruction model, where the decoder performs image reconstruction based on the comprehensive features to obtain the reconstructed image. The invisible features can be represented virtually, i.e., a learnable vector of the same dimension (e.g., 1*1024) is constructed to represent the occluded content. For example, the invisible features may include the location of the occluded content but not the image features of the occluded content itself.

[0072] Furthermore, based on the first sample image and the reconstructed image, the difference value is determined, and the model parameters of the initial image reconstruction model are adjusted based on the difference value to obtain the adjusted initial image reconstruction model. The loss function used to determine the difference value can be set according to requirements, such as a structural similarity loss function or a pixel-level loss function.

[0073] Following the steps described above, continue training the adjusted initial image reconstruction model until the first training stopping condition is met, thereby obtaining the trained target image reconstruction model. The first training stopping condition can be at least one of the following: the difference value is less than a difference threshold, the rate of change of the difference value is less than a difference rate of change threshold, or the number of iterations in supervised training reaches an iteration limit.

[0074] In this embodiment of the invention, by randomly occluding local content of a first sample image, visible content (unoccluded content) and invisible content (occluded content) are formed. Features of the visible content are then extracted and fused based on at least two cascaded coding layers to form visible features. Subsequently, image reconstruction is performed based on the comprehensive features formed by the visible features and the invisible features corresponding to the invisible content. This reconstructs the image to differentiate it from the original first sample image, allowing for adjustment of model parameters. This enables efficient self-supervised training and fully utilizes unlabeled data (unlabeled ultra-wide-angle fundus images), considered medical waste data, to construct a target image reconstruction model and, consequently, a target fundus disease detection model, thereby improving the detection performance of the intelligent fundus disease detection model. Furthermore, by adding a multi-view feature fusion layer based on the characteristics of eye tasks, a multi-view feature fusion layer is added to the traditional image reconstruction model to extract features from images under different viewpoints, obtaining finer-grained feature representations and making the extraction of eye lesion features more accurate and effective.

[0075] In some embodiments, the encoder comprises M coding layers, where M is a positive integer greater than 1; The step of inputting the second sample image into the encoder of the initial image reconstruction model, and performing concatenated encoding through each encoding layer in the encoder to obtain the encoded features output by each encoding layer includes: The second sample image is input into the first coding layer for encoding processing to obtain the encoded features output by the first coding layer; The encoded features output from the first encoding layer are input into the second encoding layer for encoding processing to obtain the encoded features output from the second encoding layer; The encoded features output from the second encoding layer are input into the third encoding layer for encoding processing to obtain the encoded features output from the third encoding layer. This process continues until the encoded features output by the Mth encoding layer are obtained.

[0076] In practical applications, see Figure 2 The encoder of the initial image reconstruction model adopts an M-layer coding layer structure. M can be set according to requirements. Figure 2M is 12. Based on this, the concatenated coding process of each coding layer can be as follows: The second sample image is input into the first coding layer for coding processing to obtain the coding features output by the first coding layer; the coding features output by the first coding layer are input into the second coding layer for coding processing to obtain the coding features output by the second coding layer; the coding features output by the second coding layer are input into the third coding layer for coding processing to obtain the coding features output by the third coding layer; the coding features output by the third coding layer are input into the fourth coding layer for coding processing to obtain the coding features output by the fourth coding layer; the coding features output by the fourth coding layer are input into the fifth coding layer for coding processing to obtain the coding features output by the fifth coding layer; the coding features output by the fifth coding layer are input into the sixth coding layer for coding processing to obtain the coding features output by the sixth coding layer. The coding features are processed as follows: the coding features output from the 6th coding layer are input into the 7th coding layer for encoding, resulting in the coding features output from the 7th coding layer; the coding features output from the 7th coding layer are input into the 8th coding layer for encoding, resulting in the coding features output from the 8th coding layer; the coding features output from the 8th coding layer are input into the 9th coding layer for encoding, resulting in the coding features output from the 9th coding layer; the coding features output from the 9th coding layer are input into the 10th coding layer for encoding, resulting in the coding features output from the 10th coding layer; the coding features output from the 10th coding layer are input into the 11th coding layer for encoding, resulting in the coding features output from the 11th coding layer; the coding features output from the 11th coding layer are input into the 12th coding layer for encoding, resulting in the coding features output from the 12th coding layer. This yields the coding features output from each coding layer. In this way, the coding features of the second sample image are obtained from 12 perspectives, which improves the richness of the coding features and thus improves the accuracy of image reconstruction.

[0077] In some embodiments, the multi-view feature fusion layer includes a multi-view feature extraction layer, a feature fusion layer, and a spatial transformation layer; The multi-view feature fusion layer extracts and fuses the encoded features output by each encoding layer to obtain visible features corresponding to the unoccluded content in the second sample image, including: The encoded features output by each of the encoding layers are input into the multi-view feature extraction layer to extract multi-view features, resulting in N target features, where N is less than or equal to the number of encoding layers contained in the encoder. Each of the target features is input into the feature fusion layer for feature fusion to obtain fused features; The fused features are input into the spatial transformation layer for spatial transformation to obtain the visible features corresponding to the second sample image.

[0078] See Figure 2 The multi-view feature fusion layer includes a multi-view feature extraction layer, a feature fusion layer, and a spatial transformation layer.

[0079] In practical applications, to reduce data processing volume while ensuring feature quality, a multi-view feature extraction layer can be used to select N coded features, i.e., target features, from M coded features. Here, N is a positive integer less than or equal to M, which can be set according to requirements; feature extraction can be random or it can be fixed by extracting certain coded features output by the coding layers.

[0080] For example, M is 12 and N is 4. Feature extraction adopts a skipping method to extract the coding features of the 3rd coding layer, the 6th coding layer, the 9th coding layer and the 12th coding layer as target features.

[0081] Based on the N target features obtained, the feature fusion layer can use average pooling to fuse the extracted N target features to obtain fused features; further, the spatial transformation layer can use a fully connected (Linear) layer to perform spatial transformation on the fused features (e.g., convert to 1024 dimensions) to obtain visible features.

[0082] In this embodiment of the invention, feature extraction can reduce the amount of feature processing while ensuring the accuracy of visible features, thereby improving the efficiency of image reconstruction; and average pooling and spatial transformation operations further improve the reliability of visible features.

[0083] In some embodiments, determining the comprehensive features corresponding to the second sample image based on the visible features and the occlusion content corresponding to the second sample image includes: Determine the first location information of the occluded content and the second location information of the unoccluded content in the second sample image; Based on the first location information and the second location information, the visible features are concatenated with the empty features corresponding to the occluded content to obtain the comprehensive features corresponding to the second sample image.

[0084] In practical applications, the first location information of the occluded content (the location information of each occluded image block) and the second location information of the unoccluded content (the location information of each unoccluded image block) in the second sample image can be determined. Then, according to the first and second location information, the visible features are concatenated with the empty features (empty image features) corresponding to the occluded content to obtain the comprehensive features.

[0085] For example, the first sample image is divided into 14×14 image blocks, and 75% of the image blocks (147 image blocks) are randomly occluded to obtain the second sample image. After the second sample image is encoded, the features of the 25% visible blocks are obtained, which are the visible features, and the visible features have a dimension of 1024. The invisible features of the unoccluded content are represented virtually, that is, a learnable vector of the same dimension (1*1024) is constructed to represent the invisible blocks, that is, an empty feature of 1024 dimensions. Then, the 25% visible block features (i.e., visible features, with a total dimension of 49*1024, one feature vector per image block) and the 75% invisible block features (i.e., empty features, containing 147 1*1024 features, that is, the total dimension of the invisible block features is 147*1024, one feature vector per image block) are used together as input and concatenated in the original order (first position information and second position information) to serve as the input to the decoder.

[0086] It should be noted that the processing flow of the traditional image reconstruction model is as follows: the image is cut into patches (such as 16×16), and then 75% of the patches are randomly occluded. The encoder only processes the visible 25% of the patches, and the decoder receives the visible + virtual token (invisible block) and reconstructs the original image. In the traditional image reconstruction model, only the output of the last layer of the encoder's 12 coding layers is used for reconstruction.

[0087] The target image reconstruction model provided by this invention essentially adds a multi-view feature fusion layer inside the encoder of a traditional image reconstruction model to obtain more fine-grained and stable ultra-wide-angle fundus features.

[0088] For example, taking an encoder containing 12 coding layers and a multi-view feature fusion layer extracting 4 target features as an example, the image reconstruction process of the target image reconstruction model is explained: The multi-view feature fusion layer uses the skip method to extract the coding features output by the 3rd, 6th, 9th and 12th coding layers as target features. The coding features output by the 3rd, 6th, 9th and 12th coding layers represent different depths of the network, that is, different semantic perspectives.

[0089] Among them, the coding features output by the third coding layer represent the low-level texture, brightness, and color details of shallow vision; the coding features output by the sixth coding layer represent the intermediate structure and shape of mid-level vision; the coding features output by the ninth coding layer represent the background and local lesion relationships of deep semantics; and the coding features output by the twelfth coding layer represent the global semantics / final representation of high-level vision.

[0090] Next, the multi-view feature fusion layer uses Average Pooling to fuse multi-view features (all target features) to obtain fused features. In this way, information from different perspectives can be unified and integrated into a more stable and comprehensive patch representation. Then, a Linear transformation layer is used to obtain visible features, ensuring that the visible features are consistent with the input dimension of the decoder.

[0091] Furthermore, the visible features (the true features of 25% of the patches) and the invisible blocks (the learnable vectors of 75% of the patches) are concatenated in the original patch order to obtain comprehensive features, which are then fed into the decoder for image reconstruction, thus realizing image reconstruction based on multi-view feature fusion.

[0092] As can be seen, multi-view feature fusion can extract features such as small lesions (superficial features), vascular network relationships (medium-layer features), and overall lesion patterns (deep features) to address the complex structure of ultra-wide-angle fundus images. This can compensate for the problem that single-layer features are insufficient to fully express information, thereby significantly improving the noise resistance, fine-grained lesion expression accuracy, global and local consistency, and cross-dataset robustness of fundus disease detection models.

[0093] In this embodiment of the invention, by concatenating visible features with empty features corresponding to occluded content based on location information, the accuracy of the comprehensive features is ensured, which is beneficial to improving the accuracy of image reconstruction.

[0094] In some embodiments, constructing an initial fundus disease detection model based on the classification layer and the encoder in the target image reconstruction model includes: Using the encoder in the target image reconstruction model as a feature extractor and a multilayer perceptron as a classification layer, the feature extractor and the classification layer are connected in series to obtain the initial fundus disease detection model.

[0095] In the interdisciplinary field of artificial intelligence and medicine, disease detection and identification mostly employ image classification algorithms. Medical image information serves as input, and a feature extractor converts this input information into feature information (usually feature vectors) before classification. This feature extractor can also be called an encoder. Therefore, this invention uses the encoder of the target image reconstruction model as the feature extractor.

[0096] Specifically, see Figure 3 , Figure 3This is a schematic diagram of the structure of the initial fundus disease detection model provided by the present invention: the encoder of the target image reconstruction model can be used as a feature extractor (for receiving a third sample image and extracting features from the third sample image), and a multilayer perceptron can be used as a classification layer (for classifying and detecting the features output by the feature extractor and outputting the detection results) to construct an ultra-wide-angle image eye disease recognition network, i.e., the initial fundus disease detection model.

[0097] In this embodiment of the invention, since the target image reconstruction model makes full use of unlabeled data that is considered medical waste data, the feature extraction accuracy of the decoder in the target image reconstruction model is improved. Therefore, the encoder in the trained target image reconstruction model is used as the feature extractor of the initial fundus disease detection model, which can improve the detection performance of the initial fundus disease detection model and thus improve the detection performance of the target fundus disease detection model.

[0098] In some embodiments, the step of supervising the training of the initial fundus disease detection model based on a second dataset to obtain a trained target fundus disease detection model includes: Based on the set ratio and the second dataset, the target training set, target validation set, and target test set are determined. Based on the target training set, the initial fundus disease detection model is subjected to multiple rounds of supervised training to obtain backup fundus disease detection models for each round. Based on the target validation set, each of the backup fundus disease detection models is validated to obtain the detection indicators of each backup fundus disease detection model; The standby fundus disease detection model that meets the testing conditions for the detection indicators is determined as the target fundus disease detection model.

[0099] In practical applications, the second dataset can be divided into a target training set, a target validation set, and a target test set according to a set ratio to ensure that the same image in the second dataset can only appear in the same subset.

[0100] Furthermore, based on the target training set, the initial fundus disease detection model can undergo a first round of supervised training to obtain a backup fundus disease detection model corresponding to the first round. Then, based on the target training set, the backup fundus disease detection model corresponding to the first round can undergo a second round of supervised training to obtain a backup fundus disease detection model corresponding to the second round. Next, based on the target training set, the backup fundus disease detection model corresponding to the second round can undergo a third round of supervised training to obtain a backup fundus disease detection model corresponding to the third round, and so on, until the backup fundus disease detection model corresponding to the last round is obtained.

[0101] Based on the backup fundus disease detection model corresponding to each round, the target validation set can be input into each backup fundus disease detection model for testing to obtain the detection index of each backup fundus disease detection model. The detection index may include at least one of accuracy and detection rate.

[0102] Then, the backup fundus disease detection model with the best detection index is used as the target fundus disease detection model, that is, the backup fundus disease detection model whose detection index meets the testing conditions is used as the target fundus disease detection model.

[0103] For example, the second dataset is uniformly divided into a 7:1.5:1.5 ratio to obtain a target training set, a target validation set, and a target test set, ensuring that the same image in the second dataset can only appear in the same subset. Then, the initial fundus disease detection model is subjected to multiple rounds of supervised training to obtain backup fundus disease detection models for each round. The backup fundus disease detection model that performs best on the target validation set is selected as the target fundus disease detection model.

[0104] In this embodiment of the invention, by dividing the target training set and performing multiple rounds of supervised training on the initial fundus disease detection model, multiple backup fundus disease detection models can be obtained. The backup fundus disease detection model with the best detection performance is selected from the backup fundus disease detection models through the target validation set as the target fundus disease detection model, which to a certain extent ensures the detection accuracy and reliability of the target fundus disease detection model.

[0105] In some embodiments, after supervising the training of the initial fundus disease detection model based on the second dataset to obtain a trained target fundus disease detection model, the method further includes: Based on the target test set, the target fundus disease detection model is tested, and the target fundus disease detection model that passes the test is used for fundus disease detection.

[0106] In practical applications, to further ensure that the detection performance of the target fundus disease detection model meets the standards, the target fundus disease detection model can be tested using a target test set. If the test is passed, the target fundus disease detection model can be used to detect fundus diseases; if the test is failed, the target fundus disease detection model can be retrained until the retrained target fundus disease detection model passes the test.

[0107] In some embodiments, determining the target training set, target validation set, and target test set based on a set ratio and the second dataset includes: The second dataset is divided into an initial training set, an initial validation set, and an initial test set according to the set ratio. Data augmentation processing is performed on the initial training set, the initial validation set, and the initial test set to obtain the target training set, the target validation set, and the target test set.

[0108] In practical applications, in order to ensure the amount of data and the effectiveness of training, validation and testing, the second dataset can be divided into an initial training set, an initial validation set and an initial test set according to a set ratio. Then, data augmentation strategies, such as image normalization, random horizontal image flipping and image resizing, are used to augment the initial training set, initial validation set and initial test set respectively to obtain the target training set, target validation set and target test set.

[0109] The target fundus disease detection model trained by this invention has been validated on ultra-wide-angle fundus image disease detection tasks: In disease detection tasks such as diabetic retinopathy (DR), lattice degeneration (LD), retinal breaks (RB), retinal detachment (RD), retinitis pigmentosa (RP), artery occlusion (AO), retinal vein occlusion (RVO), age-related macular degeneration (AMD), macular hole (MH), and pathological myopia (PM), compared with the model initialized using ImageNet, the target fundus disease detection model trained by this invention has significantly improved the area under the curve (AUC) and accuracy (ACC) in disease detection tasks.

[0110] This invention also provides a training device for a fundus disease detection model based on multi-view feature fusion. This device is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. The terms "module," "unit," "subunit," etc., used below refer to combinations of software and / or hardware that perform a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0111] The following describes the training device for a fundus disease detection model based on multi-view feature fusion provided by the present invention. The training device for a fundus disease detection model based on multi-view feature fusion described below can be referred to in correspondence with the training method for a fundus disease detection model based on multi-view feature fusion described above.

[0112] Figure 4 This is a schematic diagram of the structure of the fundus disease detection model training device based on multi-view feature fusion provided by the present invention, as shown below. Figure 4 As shown, the training device for the fundus disease detection model based on multi-view feature fusion includes: The self-supervised training module 401 is configured to perform self-supervised training on the initial image reconstruction model based on the first dataset to obtain the trained target image reconstruction model. The first dataset is determined based on multiple unlabeled ultra-wide-angle fundus images. The construction module 402 is configured to construct an initial fundus disease detection model based on the classification layer and the encoder in the target image reconstruction model. The encoder includes a multi-view feature fusion layer, which is used for multi-view feature extraction and fusion. The supervised training module 403 is configured to perform supervised training on the initial fundus disease detection model based on a second dataset to obtain a trained target fundus disease detection model. The second dataset is determined based on an open-source dataset and / or at least one labeled ultra-wide-angle fundus image.

[0113] The present invention provides a training device for a fundus disease detection model based on multi-view feature fusion. It uses a first dataset determined from multiple unlabeled ultra-wide-angle fundus images to perform self-supervised training on an initial image reconstruction model, resulting in a trained target image reconstruction model. This device effectively utilizes discarded ophthalmic medical image data and employs self-supervised learning technology to enable the target image reconstruction model to learn general representations of fundus diseases, thus turning waste into treasure. An initial fundus disease detection model is constructed using a classification layer and an encoder containing a multi-view feature fusion layer within the target image reconstruction model. This allows the initial fundus disease detection model to be formed by combining the encoder in the target image reconstruction model, which learns general representations of fundus diseases, with the classification layer. A second dataset, determined from an open-source dataset and / or at least one labeled ultra-wide-angle fundus image, is then used to supervise the training of the initial fundus disease detection model, resulting in a trained target fundus disease detection model. This significantly improves the accuracy and intelligent diagnostic performance of the target fundus disease detection model, thereby facilitating clinical applications.

[0114] In some embodiments, the self-supervised training module 401 is specifically configured as follows: For any first sample image in the first dataset, random occlusion of local content is performed on the first sample image to obtain a second sample image; The second sample image is input into the encoder of the initial image reconstruction model. The encoder performs cascaded encoding on each encoding layer to obtain the encoding features output by each encoding layer. The multi-view feature fusion layer then extracts and fuses the encoding features output by each encoding layer to obtain the visible features corresponding to the unmasked content in the second sample image. Based on the visible features and the occlusion content corresponding to the second sample image, the comprehensive features corresponding to the second sample image are determined. The integrated features are input into the decoder of the initial image reconstruction model to reconstruct the image, thereby obtaining the reconstructed image; Based on the first sample image and the reconstructed image, the model parameters of the initial image reconstruction model are adjusted to obtain the adjusted initial image reconstruction model; Continue training the adjusted initial image reconstruction model until the first training stopping condition is met, and obtain the trained target image reconstruction model.

[0115] In some embodiments, the encoder comprises M coding layers, where M is a positive integer greater than 1; The self-supervised training module 401 is specifically configured as follows: The second sample image is input into the first coding layer for encoding processing to obtain the encoded features output by the first coding layer; The encoded features output from the first encoding layer are input into the second encoding layer for encoding processing to obtain the encoded features output from the second encoding layer; The encoded features output from the second encoding layer are input into the third encoding layer for encoding processing to obtain the encoded features output from the third encoding layer. This process continues until the encoded features output by the Mth encoding layer are obtained.

[0116] In some embodiments, the multi-view feature fusion layer includes a multi-view feature extraction layer, a feature fusion layer, and a spatial transformation layer; The self-supervised training module 401 is specifically configured as follows: The encoded features output by each of the encoding layers are input into the multi-view feature extraction layer to extract multi-view features, resulting in N target features, where N is less than or equal to the number of encoding layers contained in the encoder. Each of the target features is input into the feature fusion layer for feature fusion to obtain fused features; The fused features are input into the spatial transformation layer for spatial transformation to obtain the visible features corresponding to the second sample image.

[0117] In some embodiments, the self-supervised training module 401 is specifically configured as follows: Determine the first location information of the occluded content and the second location information of the unoccluded content in the second sample image; Based on the first location information and the second location information, the visible features are concatenated with the empty features corresponding to the occluded content to obtain the comprehensive features corresponding to the second sample image.

[0118] In some embodiments, the construction module 402 is specifically configured as follows: Using the encoder in the target image reconstruction model as a feature extractor and a multilayer perceptron as a classification layer, the feature extractor and the classification layer are connected in series to obtain the initial fundus disease detection model.

[0119] In some embodiments, the supervised training module 403 is specifically configured as follows: Based on the set ratio and the second dataset, the target training set, target validation set, and target test set are determined. Based on the target training set, the initial fundus disease detection model is subjected to multiple rounds of supervised training to obtain backup fundus disease detection models for each round. Based on the target validation set, each of the backup fundus disease detection models is validated to obtain the detection indicators of each backup fundus disease detection model; The standby fundus disease detection model that meets the testing conditions is determined as the target fundus disease detection model; Based on the target test set, the target fundus disease detection model is tested, and the target fundus disease detection model that passes the test is used for fundus disease detection.

[0120] In some embodiments, the supervised training module 403 is specifically configured as follows: The second dataset is divided into an initial training set, an initial validation set, and an initial test set according to the set ratio. Data augmentation processing is performed on the initial training set, the initial validation set, and the initial test set to obtain the target training set, the target validation set, and the target test set.

[0121] In some embodiments, the training device for the fundus disease detection model based on multi-view feature fusion further includes an adjustment module configured to: Collect the aforementioned multiple unlabeled ultra-wide-angle fundus images; Using cubic difference, the plurality of unlabeled ultra-wide-angle fundus images are adjusted to a first set size, and the adjusted plurality of unlabeled ultra-wide-angle fundus images are converted to a set image format to obtain the first dataset, and / or, each first sample image in the first dataset is adjusted to a second set size; The self-supervised training module 401 is specifically configured as follows: Based on the first dataset or the first dataset after resizing, the initial image reconstruction model is trained under self-supervised conditions to obtain the target image reconstruction model.

[0122] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination.

[0123] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 5 As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a training method for a fundus disease detection model based on multi-view feature fusion. This method includes: performing self-supervised training on an initial image reconstruction model based on a first dataset to obtain a trained target image reconstruction model, wherein the first dataset is determined based on multiple unlabeled ultra-wide-angle fundus images; constructing an initial fundus disease detection model based on a classification layer and an encoder in the target image reconstruction model, wherein the encoder includes a multi-view feature fusion layer for multi-view feature extraction and fusion; and performing supervised training on the initial fundus disease detection model based on a second dataset to obtain a trained target fundus disease detection model, wherein the second dataset is determined based on an open-source dataset and / or at least one labeled ultra-wide-angle fundus image.

[0124] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0125] The present invention also provides an electronic device including a memory and a processor, the memory storing a computer program and the processor being configured to run the computer program to perform the steps in any of the above method embodiments.

[0126] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0127] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program: The initial image reconstruction model is self-supervised and trained based on the first dataset to obtain a trained target image reconstruction model. The first dataset is determined based on multiple unlabeled ultra-wide-angle fundus images. Based on the classification layer and the encoder in the target image reconstruction model, an initial fundus disease detection model is constructed. The encoder includes a multi-view feature fusion layer, which is used for multi-view feature extraction and fusion. Based on the second dataset, the initial fundus disease detection model is trained under supervision to obtain a trained target fundus disease detection model. The second dataset is determined based on an open-source dataset and / or at least one labeled ultra-wide-angle fundus image.

[0128] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated in this embodiment.

[0129] On the other hand, in conjunction with the fundus disease detection model training method based on multi-view feature fusion provided in the above embodiments, the present invention also provides a non-transitory computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above embodiments of the fundus disease detection model training method based on multi-view feature fusion.

[0130] In another aspect, in conjunction with the fundus disease detection model training method based on multi-view feature fusion provided in the above embodiments, the present invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the fundus disease detection model training method based on multi-view feature fusion.

[0131] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0132] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A training method for a fundus disease detection model based on multi-view feature fusion, characterized in that, include: The initial image reconstruction model is self-supervised and trained based on the first dataset to obtain a trained target image reconstruction model. The first dataset is determined based on multiple unlabeled ultra-wide-angle fundus images. Based on the classification layer and the encoder in the target image reconstruction model, an initial fundus disease detection model is constructed. The encoder includes a multi-view feature fusion layer, which is used for multi-view feature extraction and fusion. Based on the second dataset, the initial fundus disease detection model is trained under supervision to obtain a trained target fundus disease detection model. The second dataset is determined based on an open-source dataset and / or at least one labeled ultra-wide-angle fundus image.

2. The method for training a fundus disease detection model based on multi-view feature fusion according to claim 1, characterized in that, The step of performing self-supervised training on the initial image reconstruction model based on the first dataset to obtain a trained target image reconstruction model includes: For any first sample image in the first dataset, random occlusion of local content is performed on the first sample image to obtain a second sample image; The second sample image is input into the encoder of the initial image reconstruction model. The encoder performs cascaded encoding on each encoding layer to obtain the encoding features output by each encoding layer. The multi-view feature fusion layer then extracts and fuses the encoding features output by each encoding layer to obtain the visible features corresponding to the unmasked content in the second sample image. Based on the visible features and the occlusion content corresponding to the second sample image, the comprehensive features corresponding to the second sample image are determined. The integrated features are input into the decoder of the initial image reconstruction model to reconstruct the image, thereby obtaining the reconstructed image; Based on the first sample image and the reconstructed image, the model parameters of the initial image reconstruction model are adjusted to obtain the adjusted initial image reconstruction model; Continue training the adjusted initial image reconstruction model until the first training stopping condition is met, and obtain the trained target image reconstruction model.

3. The method for training a fundus disease detection model based on multi-view feature fusion according to claim 2, characterized in that, The multi-view feature fusion layer includes a multi-view feature extraction layer, a feature fusion layer, and a spatial transformation layer; The multi-view feature fusion layer extracts and fuses the encoded features output by each encoding layer to obtain visible features corresponding to the unoccluded content in the second sample image, including: The encoded features output by each of the encoding layers are input into the multi-view feature extraction layer to extract multi-view features, resulting in N target features, where N is less than or equal to the number of encoding layers contained in the encoder. Each of the target features is input into the feature fusion layer for feature fusion to obtain fused features; The fused features are input into the spatial transformation layer for spatial transformation to obtain the visible features corresponding to the second sample image.

4. The method for training a fundus disease detection model based on multi-view feature fusion according to claim 2 or 3, characterized in that, The encoder comprises M coding layers, where M is a positive integer greater than 1; The step of inputting the second sample image into the encoder of the initial image reconstruction model, and performing concatenated encoding through each encoding layer in the encoder to obtain the encoded features output by each encoding layer includes: The second sample image is input into the first coding layer for encoding processing to obtain the encoded features output by the first coding layer; The encoded features output from the first encoding layer are input into the second encoding layer for encoding processing to obtain the encoded features output from the second encoding layer; The encoded features output from the second encoding layer are input into the third encoding layer for encoding processing to obtain the encoded features output from the third encoding layer. This process continues until the encoded features output by the Mth encoding layer are obtained.

5. The method for training a fundus disease detection model based on multi-view feature fusion according to claim 2 or 3, characterized in that, The step of determining the comprehensive features corresponding to the second sample image based on the visible features and the occlusion content corresponding to the second sample image includes: Determine the first location information of the occluded content and the second location information of the unoccluded content in the second sample image; Based on the first location information and the second location information, the visible features are concatenated with the empty features corresponding to the occluded content to obtain the comprehensive features corresponding to the second sample image.

6. The method for training a fundus disease detection model based on multi-view feature fusion according to any one of claims 1-3, characterized in that, The initial fundus disease detection model is constructed based on the classification layer and the encoder in the target image reconstruction model, including: Using the encoder in the target image reconstruction model as a feature extractor and a multilayer perceptron as a classification layer, the feature extractor and the classification layer are connected in series to obtain the initial fundus disease detection model.

7. The method for training a fundus disease detection model based on multi-view feature fusion according to any one of claims 1-3, characterized in that, The step of supervising the training of the initial fundus disease detection model based on the second dataset to obtain a trained target fundus disease detection model includes: Based on the set ratio and the second dataset, the target training set, target validation set, and target test set are determined. Based on the target training set, the initial fundus disease detection model is subjected to multiple rounds of supervised training to obtain backup fundus disease detection models for each round. Based on the target validation set, each of the backup fundus disease detection models is validated to obtain the detection indicators of each backup fundus disease detection model; The standby fundus disease detection model that meets the testing conditions is determined as the target fundus disease detection model; After supervising the training of the initial fundus disease detection model based on the second dataset to obtain the trained target fundus disease detection model, the process further includes: Based on the target test set, the target fundus disease detection model is tested, and the target fundus disease detection model that passes the test is used for fundus disease detection.

8. The method for training a fundus disease detection model based on multi-view feature fusion according to claim 7, characterized in that, The determination of the target training set, target validation set, and target test set based on a set ratio and the second dataset includes: The second dataset is divided into an initial training set, an initial validation set, and an initial test set according to the set ratio. Data augmentation processing is performed on the initial training set, the initial validation set, and the initial test set to obtain the target training set, the target validation set, and the target test set.

9. The method for training a fundus disease detection model based on multi-view feature fusion according to any one of claims 1-3, characterized in that, Before the step of performing self-supervised training on the initial image reconstruction model based on the first dataset to obtain the trained target image reconstruction model, the following steps are also included: Collect the aforementioned multiple unlabeled ultra-wide-angle fundus images; Using cubic difference, the plurality of unlabeled ultra-wide-angle fundus images are adjusted to a first set size, and the adjusted plurality of unlabeled ultra-wide-angle fundus images are converted to a set image format to obtain the first dataset, and / or, each first sample image in the first dataset is adjusted to a second set size; The step of performing self-supervised training on the initial image reconstruction model based on the first dataset to obtain a trained target image reconstruction model includes: Based on the first dataset or the first dataset after resizing, the initial image reconstruction model is trained under self-supervised conditions to obtain the target image reconstruction model.

10. A training device for a fundus disease detection model based on multi-view feature fusion, characterized in that, include: The self-supervised training module is configured to perform self-supervised training on the initial image reconstruction model based on the first dataset to obtain the trained target image reconstruction model. The first dataset is determined based on multiple unlabeled ultra-wide-angle fundus images. The building module is configured to construct an initial fundus disease detection model based on the classification layer and the encoder in the target image reconstruction model. The encoder includes a multi-view feature fusion layer, which is used for multi-view feature extraction and fusion. The supervised training module is configured to supervise the training of the initial fundus disease detection model based on a second dataset to obtain a trained target fundus disease detection model. The second dataset is determined based on an open-source dataset and / or at least one labeled ultra-wide-angle fundus image.