Establishing and training method and device for fundus image multi-task model

Through the coordinated fusion of two-stage training strategies and visual encoder and visual projector, the problem of the lack of multimodal task system in fundus disease assisted diagnosis is solved, and more accurate fundus image multi-task model training is achieved, which improves the model's performance in fundus disease diagnosis.

CN120297420APending Publication Date: 2025-07-11BEIHANG UNIV
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510447878.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art lacks a multimodal task system and evaluation benchmark in the auxiliary diagnosis of fundus diseases. The data set covers a single and limited scale, making it difficult to conduct comprehensive and accurate model training and evaluation. The multimodal large language model has challenges in capturing local and global visual features.

Method used

Using a two-stage training strategy, multiple open-source fundus image data sets are collected, structured annotations and synthetic text descriptions are integrated, fine-grained pathological details and advanced diagnostic semantics are combined with large language models to generate predictive text.

Benefits of technology

The performance of multimodal large language model in fundus image-assisted diagnostic tasks is improved, and can more comprehensively and accurately reflect the visual performance of fundus diseases and support multi-task learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297420A_ABST
    Figure CN120297420A_ABST
Patent Text Reader

Abstract

The invention provides a construction and training method and device for an eye fundus image multi-task model, and belongs to the field of image processing, and the method comprises the steps: S1, collecting and sorting a public eye fundus image data set, constructing an image text pair according to a real label, and carrying out the two-stage training of a multi-modal large language model, the multi-mode large language model comprises a visual encoder, a visual projector and a large language model; s2, inputting the image data # imgabs0 # into a visual encoder in the trained multi-modal large language model to obtain an enhanced visual feature # imgabs1 #, and extracting a visual feature # imgabs3 # from the # imgabs2 # through a visual projector; and S3, embedding the text input # imgabs4 # to obtain a text feature # imgabs5 #, splicing the text feature # imgabs5 # with the visual feature # imgabs6 #, and inputting the spliced text feature # imgabs5 # and the visual feature # imgabs6 # into a large language model to generate a prediction text A. According to the method, a wide range of fundus image data is collected for training, multilevel lesion features in the fundus image are fully utilized, and the performance of the model for executing a fundus disease auxiliary diagnosis task can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and particularly relates to a method and device for constructing and training a multi-task model for fundus images. Background Art

[0002] Fundus images can directly reflect the health status of the retina structure, optic nerve and blood vessels, and have irreplaceable clinical value in the early screening of blinding eye diseases such as diabetic retinopathy, glaucoma, and macular degeneration. According to statistics from the World Health Organization, approximately 2.2 billion people worldwide have visual impairments, and at least half of them can achieve early intervention through fundus screening. Computer-aided diagnosis technology based on fundus images can improve the efficiency of clinical diagnosis and provide reliable remote diagnosis and treatment support for areas with scarce medical resources. Fundus image-assisted diagnosis mainly includes disease classification, lesion localization, local diagnosis, and image description. Among them, the disease classification task requires identifying diseases in fundus images. The lesion localization task involves providing detection frames for specific lesions. The local lesion diagnosis task requires lesion diagnosis based on fundus images and the provided detection frames, which is the reverse process of lesion localization. The image description task aims to analyze the visual features of the image and diagnose potential diseases. These four tasks have been widely studied in the field of medical image computing, simulating the clinical diagnosis process: the disease classification task identifies potential fundus diseases, equivalent to the preliminary screening process. The fundus lesion localization and local lesion diagnosis tasks simulate the detailed analysis of lesions. The fundus image description task is similar to the process of generating a diagnosis report.

[0003] In recent years, Multimodal Large Language Models (MLLMs) have demonstrated powerful cross-modal understanding capabilities in visual language tasks in the general field. However, in the field of medical imaging, especially in the application of fundus disease-assisted diagnosis, there are still two challenges:

[0004] Currently, there is a lack of a multi-modal task system and evaluation benchmark for fundus disease-assisted diagnosis. There are a wide variety of fundus diseases with complex pathological characteristics, while the existing datasets cover only a single disease and have a limited data scale, making it difficult for researchers to conduct comprehensive and accurate model training and evaluation. On the other hand, the existing datasets often focus on the optimization of a single task and cannot cover the multi-task requirements in fundus disease-assisted diagnosis, such as image classification, lesion detection, and clinical report generation. In addition, the existing datasets only provide fundus images and corresponding simple annotations, lacking high-quality image-text pairs required for training multi-modal large language models.

[0005] The visual features of fundus diseases exhibit complexity and diversity. For example, microaneurysms and hemorrhages in diabetic retinopathy typically present as small-scale pathological features; the optic disc and optic cup structures in glaucoma show more extensive global morphological changes; and the lesion areas in macular degeneration may include local exudates and global deformations in the macular region. These multi-scale visual features require the model to be able to capture both local and global information simultaneously. However, existing work has mainly focused on data design and paid little attention to the model's ability to extract medical features, making it challenging for multi-modal large language models in the field of fundus image analysis, which requires simultaneous attention to local and global features, for diagnosis. Summary of the Invention

[0006] To solve the above technical problems, the present invention provides a method for constructing and training a multi-task model for fundus images, including the following steps:

[0007] Step S1: Collect and organize publicly available fundus image datasets, construct image-text pairs according to real annotations, and perform two-stage training on the multi-modal large language model, where the multi-modal large language model includes: a visual encoder, a visual projector, and a large language model;

[0008] Step S2: Input the image data into the visual encoder of the trained multi-modal large language model to obtain the outputs of different visual layers, and then synergistically fuse fine-grained pathological details and advanced diagnostic semantics through a multi-layer feature extraction mechanism to obtain enhanced visual features , and pass through the visual projector to extract visual features ;

[0009] Step S3: Input the text for embedding to obtain text features , concatenate with the visual features and input into the large language model to generate prediction text A.

[0010] Advantageous Effects:

[0011] 1. The method disclosed in the present invention collects and integrates multiple open-source fundus image datasets, constructs a fundus image-assisted diagnosis dataset, and integrates structured annotations and synthetic text descriptions to support powerful multi-task learning.

[0012] 2. The present invention adopts a two-stage training strategy. In the first stage, the visual-language alignment ability of the visual projector is trained on a pseudo-label dataset, and in the second stage, the visual projector and the large language model are trained with data constructed by real labels. This training method introduces relevant knowledge of fundus images and simultaneously improves the performance of the multi-modal large language model in fundus image-assisted diagnosis tasks.

[0013] 3. The method disclosed in the present invention integrates visual encoding, including the detailed lesion features of the low-level visual layer and the semantic abstraction features of the high-level layer, which coincides with the complexity and diversity of the visual features of fundus diseases. By fusing these two levels of features, the model can more comprehensively and accurately reflect the visual manifestations of fundus diseases, thereby improving the performance of the fundus image assisted diagnosis task. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a schematic flowchart of the construction and training method of a multi-task model for fundus images according to the present invention;

[0015] Figure 2 It is a schematic diagram of the prediction text generation process of the multi-modal large language model;

[0016] Figure 3 It is a structural block diagram of the construction and training device of a multi-task model for fundus images according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0018] Embodiment 1

[0019] As Figure 1 shown, a construction and training method of a multi-task model for fundus images provided by an embodiment of the present invention includes the following steps:

[0020] Step S1: Collect and organize the publicly available fundus image dataset, construct image-text pairs according to the true annotations, and perform two-stage training on the multi-modal large language model, where the multi-modal large language model includes: a visual encoder, a visual projector, and a large language model;

[0021] Step S2: Input the image data into the visual encoder in the trained multi-modal large language model to obtain the outputs of different visual layers, and then collaboratively fuse the fine-grained pathological details and high-level diagnostic semantics through a multi-layer feature extraction mechanism to obtain enhanced visual features , and extract visual features after passing through the visual projector;

[0022] Step S3: Embed the text to obtain text features , and combine with the visual features After splicing, input it into the large language model to generate the predicted text A.

[0023] In one embodiment, the above step S1: collect and organize the publicly available fundus image dataset, construct image-text pairs according to the real annotations, and perform two-stage training on the multi-modal large language model. The multi-modal large language model includes: a visual encoder, a visual projector, and a large language model. Specifically, it includes:

[0024] Step S11: Collect and reorganize 11 publicly available datasets. The datasets are divided into classification datasets, detection datasets, and image description datasets;

[0025] Step S12: Integrate all the collected fundus images, design prompts, require the model to describe the visual features of the images and diagnose possible fundus diseases, input the images and prompts into the Qwen-VL multi-modal large language model to generate pseudo-labels for each image, and construct the fundus images and pseudo-labels in the form of a dialogue to form the training data for the first stage;

[0026] Step S13: Construct classification dialogue data, merge low-frequency categories into the "other" category, design multiple-choice dialogue templates and true / false dialogue templates for the classification task, and construct classification task dialogue data according to the classification labels and the two dialogue templates;

[0027] When constructing the classification task data in the embodiment of the present invention, first obtain all disease categories in the current dataset, construct multiple judgment questions for a fundus image and all categories to judge whether the disease exists in the fundus image; for multiple-choice questions, only one question is constructed for each image, and one or more diseases existing in the image are required to be given;

[0028] Step S14: Construct detection dialogue data, convert the segmentation masks of the collected detection datasets into the form of detection boxes, and normalize the coordinates of the detection boxes to the range of [0, 1000], design dialogue templates for the fundus lesion detection task and the local lesion diagnosis task, and construct data for the two tasks according to the detection boxes, lesion types, and templates;

[0029] Step S15: Construct image description dialogue data, input the keywords provided by the image description dataset into Qwen-VL to guide the generation of fundus image descriptions, design image description dialogue templates, and construct image description dialogue data according to the fundus image descriptions generated by Qwen-VL and the dialogue templates;

[0030] Step S16: Freeze the visual encoder and the large language model of the model, and use the pseudo-labels generated in step S12 to train the visual projector for medical semantic alignment;

[0031] Step S17: Freeze the visual encoder of the model, and use the multi-task data generated in Steps S13 - S15 to train the visual projector and the large language model, so as to improve the performance of the multi-modal large language model in different fundus image assisted diagnosis tasks.

[0032] In the embodiment of the present invention, when training the model, the generated multi-task data is randomly divided into a training set and a validation set according to a ratio of 4:1, and the training set data is used to train the multi-modal large language model.

[0033] In one embodiment, the above-mentioned Step S2: The image data is input into the visual encoder of the trained multi-modal large language model to obtain the outputs of different visual layers, and then the fine-grained pathological details and high-level diagnostic semantics are synergistically fused through a multi-layer feature extraction mechanism to obtain enhanced visual features , and through the visual projector, the visual features are extracted , specifically including:

[0034] Step S21: Preprocess the input image , uniformly convert it to the RGB format, randomly apply JPEG degradation, then use the bicubic interpolation method to adjust the image to a fixed size, and finally perform normalization processing on the image to obtain the preprocessed input image ;

[0035] In this project, the torchvision.transforms library is used for data preprocessing, and each channel is normalized using the mean mean = [0.485, 0.456, 0.406] and the standard deviation std = [0.229, 0.224, 0.225] to ensure that the image features conform to the expected distribution range and improve the training effect of the model.

[0036] Step S22: Input into the visual encoder, and sequentially pass through the embedding layer and several visual layers to extract the outputs of different visual layers , as shown in formula (1):

[0037] (1)

[0038] wherein, represents the visual encoder;

[0039] Step S23: According to the hierarchical similarity analysis, the multi-layer visual features are strategically divided into shallow layers and deep layers, and the means are calculated for the shallow layers and the deep layers respectively to obtain the shallow visual representation and the deep visual representation , as shown in formulas (2) - (3):

[0040] (2)

[0041] (3)

[0042] Among them, and respectively represent the starting layer number and the ending layer number of the shallow features, and respectively represent the starting layer number and the ending layer number of the deep features, represents the output of the visual encoder at the -th layer;

[0043] In the embodiment of the present invention, hierarchical similarity analysis is performed and it is found that in the visual encoder of the multimodal large language model, the cosine similarity of the first several layers and the last several layers is relatively high, indicating that these two parts extract different information respectively. Previous studies have shown that the shallow layer of the visual encoder mainly captures fine-grained local patterns, including edges, textures, and geometric details, while the deep visual layer gradually integrates details to form a globally represented semantics-rich. This step enhances the visual features by aggregating the shallow layer output to extract the detailed features in the image and aggregating the deep features to extract the semantic information in the image.

[0044] Step S24: The output of the last layer of the visual encoder retains the complete semantic context and is fused with the hierarchical features through residual summation to obtain enhanced visual features , as shown in formula (4):

[0045] (4)

[0046] Step S25: Input into the visual projector. After pixel rearrangement and downsampling, the dimension is compressed through a multi-layer perceptron to obtain the extracted visual features , as shown in formula (5):

[0047] (5)

[0048] Among them, represents the multi-layer perceptron, represents the pixel rearrangement operation.

[0049] In this step, pixel rearrangement is used to convert the high-dimensional feature map into a lower spatial resolution to achieve data compression, while avoiding losing too much spatial information in traditional downsampling. The multi-layer perceptron further maps the features after pixel rearrangement into features that the large language model can understand.

[0050] Figure 2The left part shows the schematic structure of the visual encoder and the visual projector. By extracting the hierarchical features of the visual encoder and synergistically fusing fine-grained pathological details (such as microaneurysms, hemorrhages) with high-level diagnostic semantics, it overcomes the limitation of traditional multi-modal large language models that overly focus on semantic features and neglect clinical key details.

[0051] In one embodiment, the above step S3: Input the text to perform embedding to obtain text features , and after splicing with the visual features, input them into the large language model to generate the predicted text A, which specifically includes:

[0052] Step S31: Insert special placeholder tags to mark the positions of visual features in the text , and input it into the Tokenizer to convert it into a text feature vector ;

[0053] Step S32: Fill the visual features into the positions reserved in the text feature vector to obtain the multi-modal input features , and input the multi-modal features into the large language model to generate the predicted text A, and then complete the auxiliary diagnosis task of fundus diseases.

[0054] Figure 2 The right part shows the generation process of text features. Figure 2 The overall shows the schematic diagram of the predicted text generation process of the multi-modal large language model.

[0055] Embodiment 2

[0056] As Figure 3 shown, the embodiment of the present invention provides a device for constructing and training a multi-task model for fundus images, including the following modules:

[0057] The multi-modal large language model training module 41 is used to collect and organize the publicly available fundus image dataset, construct image-text pairs according to the real annotations, and perform two-stage training on the multi-modal large language model. The multi-modal large language model includes: a visual encoder, a visual projector, and a large language model; this module includes the following steps:

[0058] Step S11: Collect and reorganize the publicly available dataset. The dataset is divided into a classification dataset, a detection dataset, and an image description dataset;

[0059] Step S12: Integrate all the collected fundus images, design prompts that require the model to describe the visual features of the images and diagnose possible fundus diseases, input the images and prompts into the Qwen-VL multimodal large language model to generate pseudo-labels for each image, and construct the fundus images and pseudo-labels in the form of a dialogue to form the training data for the first stage;

[0060] Step S13: Construct classification dialogue data, merge low-frequency categories into the "other" category, design multiple-choice dialogue templates and true / false dialogue templates for the classification task, and construct classification task dialogue data according to the classification labels and the two dialogue templates;

[0061] Step S14: Construct detection dialogue data, convert the segmentation masks of the collected detection data set into the form of detection boxes, and normalize the coordinates of the detection boxes to the range of [0, 1000]. Design dialogue templates for the fundus lesion detection task and the local lesion diagnosis task, and construct data for the two tasks according to the detection boxes, lesion types, and templates;

[0062] Step S15: Construct image description dialogue data, input the keywords provided by the image description data set into Qwen-VL to guide the generation of fundus image descriptions, design an image description dialogue template, and construct image description dialogue data according to the fundus image descriptions generated by Qwen-VL and the dialogue template;

[0063] Step S16: Freeze the visual encoder and the large language model of the model, and use the pseudo-labels generated in Step S12 to train the visual projector for medical semantic alignment;

[0064] Step S17: Freeze the visual encoder of the model, and use the multi-task data generated in Steps S13 - S15 to train the visual projector and the large language model;

[0065] The visual feature extraction module 42 is used to input the image data into the visual encoder in the trained multimodal large language model to obtain the outputs of different visual layers, and then use a multi-layer feature extraction mechanism to synergistically fuse fine-grained pathological details and high-level diagnostic semantics to obtain enhanced visual features , and extract visual features through the visual projector ; This module includes the following steps:

[0066] Step S21: Preprocess the input image by uniformly converting it to the RGB format, randomly applying JPEG degradation, then using bicubic interpolation to adjust the image to a fixed size, and finally standardizing the image to obtain the preprocessed input image ;

[0067] Step S22: Input visual encoder, Through the embedding layer and several visual layers in turn, extract the output of different visual layers , as shown in formula (1):

[0068] (1)

[0069] in, represents the visual encoder;

[0070] Step S23: Based on the hierarchical similarity analysis, the multi-layer visual features are strategically divided into shallow and deep layers, and the mean of the shallow and deep layers is calculated to obtain the shallow visual representation. and deep visual representation , as shown in formulas (2) to (3):

[0071] (2)

[0072] (3)

[0073] in, and Respectively represent the start and end layer numbers of shallow features, and Respectively represent the start and end layer numbers of the deep features, Indicates The visual encoder output of the layer;

[0074] Step S24: The last layer output of the visual encoder The complete semantic context is retained and fused with the hierarchical features through residual summation to obtain enhanced visual features. , as shown in formula (4):

[0075] (4)

[0076] Step S25: Input the visual projector, after pixel rearrangement and downsampling, the multi-layer perceptron performs dimensional compression to extract visual features. , as shown in formula (5):

[0077] (5)

[0078] in, represents a multi-layer perceptron, Represents a pixel rearrangement operation.

[0079] Prediction module 43, used to input text Embedding to get text features , after being spliced with visual features , input it into a large language model to generate predicted text A. The prediction module includes the following steps:

[0080] Step S31: Insert special placeholder identifiers at the positions of visual features in the text , input it into the Tokenizer, and convert it into a text feature vector ;

[0081] Step S32: Fill the visual features into the positions reserved in the text feature vector to obtain multi-modal input features . The multi-modal features are input into the large language model to generate predicted text A, and thus complete the auxiliary diagnosis task of fundus diseases.

[0082] An electronic device includes: one or more processors; a memory for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement a method and apparatus for constructing and training a multi-task model for fundus images.

[0083] A computer-readable storage medium stores executable instructions thereon. When the instructions are executed by a processor, the processor is caused to implement a method and apparatus for constructing and training a multi-task model for fundus images.

[0084] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for constructing and training a multi-task model for fundus images, characterized in that, It includes the following steps: Step S1: Collect and organize the publicly available fundus image dataset, construct image-text pairs according to the true annotations, and perform two-stage training on the multimodal large language model, where the multimodal large language model includes: a visual encoder, a visual projector, and a large language model; Step S2: The image data is input into the visual encoder of the trained multi-modal large language model to obtain the outputs of different visual layers, and then the fine-grained pathological details and high-level diagnostic semantics are synergistically fused through a multi-layer feature extraction mechanism to obtain enhanced visual features , and through the visual projector, the visual features are extracted ; Step S3: Input the text Perform embedding to obtain text features , and concatenate with the visual features Then input the concatenated result into the large language model to generate the predicted text A.

2. The construction and training method for the multi-task model of fundus images according to claim 1, characterized in that, The step S1: Collect and organize the publicly available fundus image dataset, construct image-text pairs according to the true annotations, and perform two-stage training on the multimodal large language model, where the multimodal large language model includes: a visual encoder, a visual projector, and a large language model, specifically includes: Step S11: Collect and re-organize the publicly available dataset, where the dataset is divided into a classification dataset, a detection dataset, and an image description dataset; Step S12: Integrate all the collected fundus images, design prompts, require the model to describe the visual features of the images and diagnose possible fundus diseases, input the images and prompts into the Qwen-VL multimodal large language model to generate pseudo-labels for each image, and construct the fundus images and pseudo-labels in a dialogue form to form the training data for the first stage; Step S13: Construct classification dialogue data, merge the low-frequency categories into the "other" category, design multiple-choice dialogue templates and true / false dialogue templates for the classification task, and construct classification task dialogue data according to the classification labels and the two dialogue templates; Step S14: Construct detection dialogue data, convert the segmentation masks of the collected detection dataset into the form of detection boxes, and normalize the coordinates of the detection boxes to the range of [0, 1000], design dialogue templates for the fundus lesion detection task and the local lesion diagnosis task, and construct the data for the two tasks according to the detection boxes, lesion types, and templates; Step S15: Construct image description dialogue data, input the keywords provided by the image description dataset into Qwen-VL to guide the generation of fundus image descriptions, design image description dialogue templates, and construct image description dialogue data according to the fundus image descriptions generated by Qwen-VL and the dialogue templates; Step S16: Freeze the visual encoder and the large language model of the model, and use the pseudo-labels generated in step S12 to train the visual projector for medical semantic alignment; Step S17: Freeze the visual encoder of the model, and use the multi-task data generated in steps S13~S15 to train the visual projector and the large language model.

3. The construction and training method of the multi-task model for fundus images according to claim 2, characterized in that, The said step S2: The image data is input into the visual encoder of the trained multi-modal large language model to obtain the outputs of different visual layers, and then the fine-grained pathological details and high-level diagnostic semantics are synergistically fused through a multi-layer feature extraction mechanism to obtain enhanced visual features , and the visual features are extracted through a visual projector , specifically including: Step S21: Preprocess the input image by converting it to the RGB format uniformly, randomly applying JPEG degradation, adjusting the image to a fixed size using the bicubic interpolation method, and finally normalizing the image to obtain the preprocessed input image ; Step S22: Take as the input to the visual encoder, and sequentially pass it through the embedding layer and several visual layers to extract the outputs of different visual layers , as shown in formula (1): (1) Among them, represents a visual encoder; Step S23: According to the hierarchical similarity analysis, the multi-layer visual features are strategically divided into shallow and deep layers, and the means are calculated for the shallow and deep layers respectively to obtain the shallow visual representation and the deep visual representation , as shown in formulas (2) to (3): (2) (3) Among them, and respectively represent the starting layer number and the ending layer number of the shallow features, and respectively represent the starting layer number and the ending layer number of the deep features, represents the output of the visual encoder of the th layer; Step S24: Output of the last layer of the visual encoder retains the complete semantic context, and fuses it with the hierarchical features through residual summation to obtain enhanced visual features , as shown in Equation (4): (4) Step S25: Input into the visual projector. After pixel rearrangement and downsampling, dimension compression is performed through a multi-layer perceptron to obtain the extracted visual features , as shown in Equation (5): (5) Among them, represents a multi-layer perceptron, represents a pixel rearrangement operation.

4. The method for constructing and training a multi-task model for fundus images according to claim 3, wherein The said step S3: Input the text Perform embedding to obtain text features , and splice with visual features Then input the spliced result into the large language model to generate the predicted text A, specifically including: Step S31: Insert a special placeholder in the text to identify the position of the visual feature, input the Tokenizer, and convert it into a text feature vector ; Step S32: Fill the visual features into the positions reserved in the text feature vector to obtain the multi-modal input features , and input the multi-modal features into the large language model to generate the predicted text A.

5. An apparatus for constructing and training a multi-task model for fundus images, characterized in that, It includes the following modules: The multimodal large language model training module is used to collect and organize the publicly available fundus image dataset, construct image-text pairs according to the true annotations, and perform two-stage training on the multimodal large language model, where the multimodal large language model includes: a visual encoder, a visual projector, and a large language model; A visual feature extraction module for image data Input it into the visual encoder of the trained multi-modal large language model to obtain the outputs of different visual layers, and then use a multi-layer feature extraction mechanism to synergistically fuse fine-grained pathological details and high-level diagnostic semantics to obtain enhanced visual features , and Extract visual features through a visual projector ; Prediction module, for inputting text to perform embedding to obtain text features , and concatenate with visual features and then input the concatenated result into a large language model to generate prediction text A.

6. The construction and training device for the multi-task model of fundus images according to claim 5, characterized in that, The multimodal large language model training module includes the following steps: Step S11: Collect and re-organize the publicly available dataset, where the dataset is divided into a classification dataset, a detection dataset, and an image description dataset; Step S12: Integrate all the collected fundus images, design prompts that require the model to describe the visual features of the images and diagnose possible fundus diseases, input the images and prompts into the Qwen-VL multi-modal large language model to generate pseudo-labels for each image, and construct the fundus images and pseudo-labels in the form of a dialogue to form the training data for the first stage; Step S13: Construct classification dialogue data, merge low-frequency categories into the "other" category, design multiple-choice dialogue templates and true / false dialogue templates for the classification task, and construct classification task dialogue data according to the classification labels and the two dialogue templates; Step S14: Construct detection dialogue data, convert the segmentation masks of the collected detection data set into the form of detection boxes, and normalize the coordinates of the detection boxes to the range of [0, 1000]. Design dialogue templates for the fundus lesion detection task and the local lesion diagnosis task, and construct data for the two tasks according to the detection boxes, lesion types, and templates; Step S15: Construct image description dialogue data, input the keywords provided by the image description data set into Qwen-VL to guide the generation of fundus image descriptions, design an image description dialogue template, and construct image description dialogue data according to the fundus image descriptions generated by Qwen-VL and the dialogue template; Step S16: Freeze the visual encoder and the large language model of the model, and use the pseudo-labels generated in Step S12 to train the visual projector for medical semantic alignment; Step S17: Freeze the visual encoder of the model, and use the multi-task data generated in Steps S13 - S15 to train the visual projector and the large language model.

7. The construction and training device for the multi-task model of fundus images according to claim 5, characterized in that, The visual feature extraction module includes the following steps: Step S21: Preprocess the input image by converting it to the RGB format uniformly, randomly applying JPEG degradation, then using the bicubic interpolation method to resize the image to a fixed size, and finally normalizing the image to obtain the preprocessed input image ; Step S22: Take as the input to the vision encoder, and successively pass it through the embedding layer and several vision layers to extract the outputs of different vision layers , as shown in formula (1): (1) Among them, represents a visual encoder; Step S23: According to the hierarchical similarity analysis, the multi-layer visual features are strategically divided into shallow and deep layers, and the means are calculated for the shallow and deep layers respectively to obtain the shallow visual representation and the deep visual representation , as shown in formulas (2) to (3): (2) (3) Among them, and represent the start layer number and the end layer number of the shallow features respectively, and represent the start layer number and the end layer number of the deep features respectively, represents the output of the visual encoder of the th layer; Step S24: Output of the last layer of the visual encoder retains the complete semantic context, and fuses it with the hierarchical features through residual summation to obtain enhanced visual features , as shown in formula (4): (4) Step S25: Place into the input visual projector. After pixel rearrangement and downsampling, perform dimensionality compression through a multi-layer perceptron to obtain the extracted visual features , as shown in formula (5): (5) Among them, represents a multi-layer perceptron, represents a pixel rearrangement operation.

8. The construction and training device for the multi-task model of fundus images according to claim 5, characterized in that The prediction module includes the following steps: Step S31: Insert a special placeholder in the text to identify the visual feature position, input the Tokenizer, and convert it into a text feature vector ; Step S32: Fill the visual features into the positions reserved in the text feature vector to obtain the multi-modal input features , and input the multi-modal features into the large language model to generate the predicted text A.

9. An electronic device, characterized in that, Including: One or more processors; A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that, Stored thereon are executable instructions that, when executed by a processor, cause the processor to implement the method according to any one of claims 1 to 4.

Citation Information

Cited By

  • Multimodal automatic driving training method based on DeepSeek training framework

    CN120910477A

  • A multi-modal automatic driving training method based on a DeepSeek training framework

    CN120910477B

  • Retina image classification method and system based on multi-modal incremental learning

    CN120954078A

  • A retinal image classification method and system based on multi-modal incremental learning

    CN120954078B

  • Layered vision Token compression method for multi-modal large language model

    CN121350590A