Small-sample digestive tract disease diagnosis method and device based on medical attribute driving
The medical attribute description of the endoscopic image was obtained through visual question and answer and cross-modal prompt optimization was solved, which solved the limitations and category imbalance of the whole sample method in the diagnosis of digestive tract disease, and achieved high-precision and explainable few-sample diagnosis.
Patent Information
- Application Number
- CN202510490788.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art relies on full-sample methods in the diagnosis of digestive tract diseases, which has high cost, cumbersome labeling and difficult to cope with category imbalance. In addition, traditional CNNs have limitations in long-distance dependence and global information capture, which is difficult to meet the needs of high-precision diagnosis.
The medical attribute description of the endoscopic image is obtained through visual question-and-answer methods, combined with the learnable visual and attribute prompts, the AVF module is used to bridge the visual and attribute information, optimize the loss function to improve the model performance, and use the cross-modal prompt optimization method for diagnosis.
It significantly improves the diagnostic accuracy and interpretability of digestive tract diseases, especially in the learning scenario of few samples. By introducing text information and cross-modal prompt optimization, the semantic understanding and classification capabilities of the model are enhanced.
Smart Images

Figure CN120412978A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image classification and diagnosis, and particularly to a few-shot digestive disease diagnosis method and device driven by medical attributes. Background Art
[0002] Digestive diseases are one of the common health problems worldwide, and early diagnosis is crucial for improving patient survival rate and prognosis. According to the statistics of the World Health Organization (WHO), the incidence of digestive diseases (such as gastric cancer, colorectal cancer, etc.) has been increasing year by year, especially in developing countries, which has brought great pressure to the medical system. As the gold standard for the diagnosis of digestive diseases, endoscopy can provide intuitive lesion image information, but its diagnostic process highly depends on the experience and professional knowledge of doctors, and has limitations of subjectivity and labor intensity. In recent years, image classification technology based on deep learning has made remarkable progress in endoscopic image analysis. In particular, convolutional neural networks (CNNs) can automatically extract features and reduce the complexity of manual design, and are widely used in this field. However, with the increase in task complexity, CNNs have limitations in capturing long-range dependencies and global information, and it is difficult to meet the requirements of high-precision diagnosis.
[0003] However, traditional full-sample-based digestive endoscopy image classification methods rely on large-scale labeled data. In practical applications, the acquisition and annotation of endoscopic images face many challenges. First, the acquisition of endoscopic images depends on professional medical equipment and clinical examinations, with high costs. Second, the accurate annotation of images requires the participation of experienced gastroenterologists, and the process is cumbersome and time-consuming. In addition, there are significant differences in the incidence of different diseases. For example, the imaging data of early canceration is much less than that of common lesions, resulting in the problem of class imbalance. These factors jointly limit the application of full-sample-based methods and prompt researchers to explore more efficient few-shot learning strategies. Although existing research has made some progress in the field of few-shot learning, the few-shot classification research for digestive endoscopy images is still in its infancy. In addition, existing methods mainly rely on single-modal image data, and the knowledge that the model can extract is relatively limited, making it difficult to handle complex and changeable clinical scenarios.
[0004] In view of this, the present application aims to provide a few-shot digestive disease diagnosis method and device driven by medical attributes, which can improve the accuracy and comprehensiveness of early diagnosis of digestive diseases by combining the visual features of endoscopic images with medical attribute descriptions, and provide a solution that takes into account both performance and interpretability for medical image classification in few-shot learning scenarios. Summary of the Invention
[0005] To solve the above problems, the present invention provides a few-shot digestive disease diagnosis method and device based on medical attribute drive, which generates based on attribute description, optimizes cross-modal prompts, fully integrates the complementary advantages of endoscopic images and medical attributes, and improves the prediction diagnosis accuracy and comprehensiveness of digestive diseases.
[0006] The technical solution adopted by the present invention is as follows:
[0007] A few-shot digestive disease diagnosis method based on medical attribute drive, comprising the steps of:
[0008] Ask a large language model through visual question answering to obtain the medical attributes contained in the endoscopic image;
[0009] Use the attribute description generated in the attribute description generation step for prompt optimization, introduce learnable visual prompts and attribute prompts, and use the AVF module to bridge the visual prompts and attribute prompts to inject medical attribute knowledge into the visual prompts;
[0010] Through the loss function optimization module, optimize the prediction accuracy of each disease label in the classification task, and update the model parameters through the backpropagation algorithm to improve the overall classification performance of the model and achieve disease diagnosis prediction.
[0011] Further, the learnable visual prompt is based on the CLIP model, and the CLIP model consists of a two-tower structure composed of a visual encoder ViT-B / 16 and a text encoder transformer.
[0012] Further, in the visual question answering method, a customized prompt template and a few-shot picture set are used to ask questions to the large language model GPT-4.
[0013] Further, the prompt template is standardized as an interrogation structure including four attributes: color, shape, texture, and boundary, and is interrogated for a certain lesion category to obtain standardized semantic information about different attributes of a certain category.
[0014] Further, the few-shot picture set is instances with different sample numbers for each class, including 4, 8, and 16, which are used as the picture input in visual question answering to prompt the large language model to obtain visual information about attributes in the image features.
[0015] Further, the medical attributes include color, shape, texture, and boundary, which are used to provide structured semantic guidance for the model.
[0016] Further, the visual prompt adopts a learnable prompt with a depth of 12 and a length of 16.
[0017] Furthermore, the attribute prompts adopt a diversified design, specifically set to have 4 different attribute prompts corresponding to different attribute descriptions for prompt learning. The depth of each attribute prompt is 1, and the length is 16.
[0018] Furthermore, the AVF module is based on the attention mechanism, adaptively focuses on the key parts of the attribute prompts and visual features, and realizes the efficient fusion of attribute semantics and visual features, thereby enhancing the model's semantic understanding and utilization of medical images.
[0019] Based on the same inventive concept, the present application also provides a few-shot digestive tract disease diagnosis device driven by medical attributes, including a processor and a memory for storing processor-executable instructions. When the processor executes the instructions, the steps of the above diagnosis method are implemented.
[0020] The beneficial effects of the present invention are as follows:
[0021] The few-shot digestive tract disease diagnosis method driven by medical attributes provided by the present invention includes: attribute description generation, which is used to ask a large language model through a visual question and answer method to obtain the medical attributes contained in the endoscopic image; cross-modal prompt optimization, which is used to optimize the prompts by using the attribute descriptions generated by the attribute description generation module. By introducing learnable visual prompts and attribute prompts, and using the AVF module to bridge the visual prompts and attribute prompts, medical attribute knowledge is injected into the visual prompts; the loss function optimization module is used to optimize the prediction accuracy of each disease label in the multi-label classification task through the cross-entropy loss function, and update the model parameters through the backpropagation algorithm to improve the overall classification performance of the model. This diagnosis method generates medical attribute descriptions of endoscopic images through a large language model, and uses learnable visual prompts and attribute prompts for cross-modal prompt optimization. During the prompt optimization process, the AVF module realizes the efficient fusion of attribute semantics and visual features, thereby enhancing the model's semantic understanding and utilization of medical images. Compared with the traditional classification method that only relies on single image data, the present invention significantly improves the accuracy and interpretability of digestive tract disease classification by introducing attribute description generation based on a large language model, introducing text information outside the image, and performing cross-modal prompt optimization at the same time, providing a solution that takes into account both performance and semantic understanding for endoscopic image analysis in the few-shot learning scenario. Description of the Drawings
[0022] Figure 1 It is the overall system diagram of the few-shot digestive tract disease diagnosis system based on medical attributes in the embodiments of the present invention;
[0023] Figure 2The following is a comparison result of the accuracy of the method in the embodiments of the present invention and other advanced methods (CoOp(CSC), CoOp, CoCoOp, Tip-Adapter, Tip-Adapter-F, CLIP Adapter, MapLe, Linear-probe CLIP, PromptSRC, PLOT) under different few-shot settings. Among them, Figure a is the Kvasir dataset, Figure b is the Hyper-Kvasir dataset, Figure c is the Nerthus dataset, and Figure d is the LIMUC dataset;
[0024] Figure 3 The following is a comparison result of the F1 score of the method in the embodiments of the present invention and other advanced methods (CoOp(CSC), CoOp, CoCoOp, Tip-Adapter, Tip-Adapter-F, CLIP Adapter, MapLe, Linear-probe CLIP, PromptSRC, PLOT) under different few-shot settings. Among them, Figure a is the Kvasir dataset, Figure b is the Hyper-Kvasir dataset, Figure c is the Nerthus dataset, and Figure d is the LIMUC dataset;
[0025] Figure 4 The following are the visualization effect diagrams of the method in the embodiments of the present invention for the original, final, and four (color, shape, texture, and boundary) attributes. Detailed implementation manners
[0026] To facilitate the understanding of the present invention, the present invention will be described more comprehensively through embodiments below. The following are the preferred embodiments of the present invention. However, the present invention can be implemented in various different forms and is not limited to the embodiments described herein. All other implementation manners obtained by modifying the technical solutions of the present invention or making equivalent replacements without creative results are within the protection scope of the present invention.
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention. The terms used in the specification of the present invention are only for describing specific embodiments and are not intended to limit the present invention.
[0028] The numerical values disclosed in the embodiments of the present invention are approximate values, not definite values. Within the allowable range of errors or experimental conditions, all values within the error range are included and are not limited to the specific numerical values disclosed in the embodiments of the present invention.
[0029] As Figure 1 shown, this embodiment provides a few-shot digestive disease diagnosis method based on medical attribute drive, including:
[0030] Attribute description generation: Ask the large language model through Visual Question Answering (VQA) to obtain the medical attributes contained in the endoscopic images.
[0031] Specifically, the purpose of attribute description generation is to address the limitations of the CLIP model's understanding of medical class names. Considering the characteristics of gastrointestinal endoscopic images, this step mimics the clinical diagnosis process of doctors and standardizes the diagnostic basis into four key attributes: color, shape, texture, and boundary. To avoid the appearance of medical class names while ensuring that the generated descriptions can be understood by the CLIP model;
[0032] The specific implementation steps are as follows:
[0033] 1. Input: Prompt GPT-4 using the following prompt template:
[0034] Question:
[0035] Please provide visual features regarding color, shape, texture, and boundary of this set of [CLASS] image.
[0036] Please list in bullet points and explain in plain words that CLIP understands. Avoid using words such as [CLASS].
[0037] Please provide the visual features of this set of image in the following format:
[0038] color: "one sentence describing the color"
[0039] shape: "one sentence describing the shape"
[0040] texture: "one sentence describing the texture"
[0041] boundary: "one sentence describing the boundary"
[0042] Among them, [CLASS] needs to be replaced with the class name of the target dataset (e.g., polyps). As Figure 1 As shown in the upper part of Figure 1 , in this step, the above prompt template is used as the Question of Visual Question Answering (VQA), and the set of images corresponding to the few-shot number (4-shot, 8-shot, or 16-shot) of the corresponding class is used as the Vision input of VQA. By asking GPT-4, a standardized response divided by attributes will be obtained. Thanks to the powerful reasoning ability of GPT-4, the generated response can accurately combine image and text information, providing a logical explanation related to each attribute of each class. These attribute descriptions not only enhance the CLIP model's understanding and classification ability of medical symptoms, but also improve the interpretability of the model, helping doctors and researchers understand the model's decision-making process.
[0043] 3. Output: The GPT-4 model outputs a standardized response divided by attributes, in the following format:
[0044] color: "Sentence describing the color"
[0045] shape: "Sentence describing the shape"
[0046] texture: "Sentence describing the texture"
[0047] boundary: "Sentence describing the boundary"
[0048] The generated attribute descriptions will be used as the input for cross-modal prompt optimization to enhance the semantic understanding and classification ability of the CLIP model for digestive endoscopy lesions.
[0049] Cross-modal prompt optimization uses the attribute descriptions generated in the attribute description generation step for prompt optimization;
[0050] Specifically, by introducing learnable visual prompts and attribute prompts, and using the AVF module to bridge the visual prompts and attribute prompts, medical attribute knowledge is injected into the visual prompts;
[0051] The specific implementation steps are as follows:
[0052] Initialize prompts: Initialize visual prompts and attribute prompts
[0053] Construct attribute-enhanced visual prompts: Pass p v and p a through a linear layer with weights and respectively to obtain and
[0054] Q v = p v W q , K a = p a W k .
[0055] Calculate the inner product of Q v and K a to obtain the similarity matrix Use this matrix to weight-sum and fuse the attribute information of p a and add the fusion result to Q v using residual connection to get the fusion hint
[0056]
[0057] Pass through an adjustment layer with weights to adjust the dimension of p av to obtain the attribute-enhanced visual hint for fine-tuning the visual encoder
[0058] Text embedding: After tokenizing the attribute description, concatenate it with the corresponding p a to obtain the input embedding of the text encoder i.e., the text embedding;
[0059] Visual embedding: After dividing the image into patches and performing linear mapping on the image, obtain the input embedding, and concatenate the input embedding with the corresponding p av ' and an additional class token c0 to get the visual embedding of the first layer [c0, P0, p av '[0]];
[0060] Output: The two branches of vision and text independently encode the visual and text embeddings layer by layer to obtain the final visual feature v and text feature t;
[0061] [c i , P i , __] = I i ([c i-1 , P i-1 , p av '[i - 1]]) i = 1, 2,..., N.
[0062] v = Proj v (c N )
[0063] W i = T i (W i-1 ) i = 1, 2,..., N.
[0064] t = Proj t (c N )
[0065] Calculate the similarity: Calculate the cosine similarity between the visual feature v and the set of attribute features {t color , t shape , t texture , t boundary}.
[0066]
[0067] Calculate the output probability: Given the class label y ∈ {1, 2, …, C} of a dataset, the probability that the image x is classified as class y C can be expressed as:
[0068]
[0069] Backpropagation optimization: Optimize the attribute-enhanced visual cue p CE by minimizing the cross-entropy loss L av ', the attribute cue p a and the AVF module:
[0070]
[0071] During the experiment, the system used 4 digestive endoscopy datasets: the Hyper-Kvasir dataset, the Kvasir dataset, the Nerthus dataset, and the LIMUC dataset. The Hyper-Kvasir dataset contains 10,662 annotated images, covering 23 categories, divided into anatomical landmarks, mucosal field quality, pathological findings, and therapeutic interventions, covering the upper and lower digestive tracts; the Kvasir dataset contains 8,000 high-quality endoscopic images, divided into 8 categories, covering anatomical landmarks, pathological abnormalities, and endoscopic procedures; the Nerthus dataset contains 21 videos (5,525 frames), focusing on the colon region, and defining four quality rating categories based on BBPS; the LIMUC dataset contains 11,276 colonoscopy images, divided into four categories (MES0-3) according to the Mayo Endoscopic Score (MES) for ulcerative colitis assessment. To ensure a sufficient number of samples, the Hyper-Kvasir dataset was screened, and 7 categories with a small number of samples were discarded, and finally 16 categories were retained for the experiment. This processing ensured that the sample size of each category was sufficient to support the few-shot learning task, while avoiding the impact of data imbalance on the model performance.
[0072] Specifically, see Table 1 below for the medical class names of each dataset used by the system;
[0073] Table 1
[0074]
[0075] As Figure 2 and Figure 3 shown, it is a comparative line chart of the system method and other advanced few-shot methods based on different evaluation metrics under different sample instances.
[0076] See Figure 2 , on the Kvasir dataset, the method ADPL provided in this embodiment is significantly better than other methods under three sample instance settings. It is worth noting that the overall performance of the prompt learning-based methods is better than that of the adapter-based methods (such as CLIPAdapter, Tip-Adapter, and Tip-Adapter-F). This phenomenon may be attributed to the fact that prompt learning can better adapt to the knowledge distribution of the pre-trained model, thus reducing the optimization difficulty under few-shot conditions. In addition, the introduced dual-branch prompt is more advantageous than a single text prompt. For example, PromptSRC, PLOT, and MaPLe are better than CoOp and CoCoOp in terms of accuracy. This is mainly because the dual-branch prompt learning not only guides the model to focus on semantic information through text prompts but also captures key visual features in the image through visual prompts, thus achieving more robust feature alignment and knowledge transfer under few-shot conditions. The same trend is also observed in terms of the F1 score. ADPL shows a higher score compared to other methods, demonstrating its consistent performance (see Figure 3 ).
[0077] Under the three sample instance settings, ADPL improves by 1.56%, 6.63%, and 4.55% respectively compared to the second-best Maple. Especially when the number of samples is 8 and 16, the performance improvement of ADPL is particularly significant. This phenomenon is mainly attributed to the high degree of fit between the attribute description and various categories in the dataset, especially in the characterization of disease categories such as intestinal polyps. The attribute description can accurately depict the category features from multiple perspectives. Combining the dual-branch prompt learning and AVF to promote the prompt fusion between modalities further enhances the effect of prompt learning. As the number of image samples increases, the alignment degree between the attribute description and the image content is further improved, thus significantly enhancing the performance of the model.
[0078] Compared with the worst CLIPAdapter, ADPL shows a more significant performance gap, especially in the setting with a sample size of 4, where the performance gap reaches 34.6%. The performance limitation of CLIPAdapter mainly stems from the following two aspects: First, the hard prompt template of CLIPAdapter lacks text prompt learning, which limits its ability to learn abstract semantics outside the dictionary, especially in the learning of medical domain-related knowledge. Second, CLIPAdapter uses ResNet50 based on the CNN architecture as the visual encoder, while the method provided in this application is based on ViT-B / 16 of the Transformer architecture. In medical image classification tasks, the Transformer architecture can better represent the complex structures and detailed features in medical images due to its stronger ability to capture global context information, thus significantly improving the performance of the model.
[0079] Next, analyze the Hyper-Kvasir dataset. The Hyper-Kvasir and Kvasir datasets have many similarities in content, both covering a variety of gastrointestinal-related diseases and normal gastrointestinal images, providing extensive support for gastrointestinal image analysis. The difference is that Hyper-Kvasir is a refined classification version of Kvasir, with more disease categories and more detailed labels, suitable for higher-precision fine-grained disease classification tasks. For example, the Kvasir dataset contains 8 general classification categories, such as "ulcerative colitis", "esophagitis", etc.; while Hyper-Kvasir further subdivides these categories into more refined subcategories, such as "ulcerative colitis grade 1", "ulcerative colitis grade 2", "esophagitis_a", "esophagitis_b-d", etc.
[0080] See Figure 2, on the Hyper-Kvasir dataset, ADPL is also competitive under different few-shot instance settings, and the performance improvement trends of all models are similar to those on the Kvasir dataset. In addition, the performance of CoCoOp, CLIPAdapter, and Tip-Adapter is even lower than that of the unimodal Linear-probe CLIP. This is mainly because CLIPAdapter and Tip-Adapter fail to utilize text prompt learning. Although CoCoOp introduces prompt learning and alleviates the performance bottleneck of CoOp in unknown class generalization by adding image instances, this method is prone to causing the model to overfit on the base classes, resulting in poor performance in the few-shot classification scenario. Specifically, the lack of effective text prompt learning makes it difficult for these methods to fully leverage the advantages of CLIP in cross-modal learning, leading to sub-optimal performance.
[0081] Next, analyze the Nerthus dataset. Different from the aforementioned Kvasir and HyperKvasir datasets, the Nerthus dataset involves the classification task of gastrointestinal cleanliness. See Figure 2 and Figure 3 the lower left part of. Overall, the performance differences of all models in terms of accuracy and F1 score are similar to the results of the Kvasir and HyperKvasir datasets. However, under the three sample sizes, ADPL improves the accuracy and F1 score by 11.4%, 1.3%, 2.17% and 3.72%, 2.04%, 1.3% respectively compared to the two-branch prompt MaPLe; in addition, compared to the single-branch prompt CoOp, the accuracy and F1 score improvements of ADPL are 8.2%, 15.96%, 9.17% and 9.61%, 10.97%, 9.5% respectively. It can be seen from the data that the improvement in the F1 score is relatively stable.
[0082] Accuracy measures the proportion of correct predictions, while the F1 score is a balanced metric that takes into account both precision and recall. Although accuracy can provide an overview of the overall performance, in the context of imbalanced class distributions, the F1 score is particularly important because it focuses more on the model's performance in correctly identifying positive class samples. Given the slight sample imbalance in the Nerthus dataset, the F1 score is further introduced for analysis to more accurately evaluate the performance of the models.
[0083] Experiments show that CLIPAdapter still performs the worst. When the number of samples is 16, there is a huge gap of 61.72% between it and ADPL. This phenomenon may be related to the degree of task refinement. In the fine-grained classification task of a single class, the model's cognitive limitations regarding medical class names are more significant. Especially when the subtle differences between classes are represented only by numerical scores (such as 0, 1, 2, 3), these numerical differences essentially lack significant visual distinctiveness and are difficult to effectively distinguish through the explicit features of images. As the number of samples increases, the model may overfit the noise in the data during training. Especially when the visual features between different scoring levels are extremely similar, the model may wrongly rely on some minor and irrelevant features for discrimination, resulting in a decline in the generalization ability on the test set.
[0084] Next, analyze the LIMUC dataset. Similar to the Nerthus dataset, the LIMUC dataset belongs to the fine-grained classification task of a single class, specifically the classification task of the severity of ulcerative colitis. There is a significant class imbalance problem in this dataset. Among them, the proportion of normal ulcerative colitis samples in the test set is as high as 54%, far higher than the other three abnormal classes.
[0085] See Figure 3 , on the LIMUC dataset, from the perspective of F1 score, the ADPL method is significantly better than other methods. Especially in the scenario of a very small number of samples 4, ADPL has achieved a significant improvement of 4.61% compared to CoOp(CSC), while CoOp(CSC) has also improved by 2.58% compared to CoOp. This result indicates that in the case of an extremely low number of samples, class-specific prompts are more effective than non-class-specific prompts, indirectly verifying the importance of prompt diversity in few-shot classification tasks. However, the number of prompt parameters of CoOp(CSC) that linearly grows based on the number of classes may cause overfitting problems due to the mismatch between the scale of prompt parameters and the limited number of training samples. In contrast, ADPL only uses 4 different attribute prompts to represent all classes. This not only fixes the number of prompt parameters but also introduces prompt diversity, especially more prominent in datasets with a larger number of classes.
[0086] In addition, ADPL uses attribute descriptions to replace class names. This method provides richer knowledge information for the prompts, thus further enhancing the diversity of the prompts. We also observed that Tip-Adapter had a slight performance decline of 0.04% when the number of samples was from 8 to 16. Considering that this method is a non-training method, its performance saturation phenomenon may be due to the limited image metric features captured during the random augmentation process.
[0087] To further verify the superiority of this method, a method based on Gradient-weighted Class Activation Mapping (Grad-CAM) was adopted. Two samples were selected from each category to conduct visual analysis on the inputs of the visual branch and the text branch respectively.
[0088] The specific implementation steps are as follows:
[0089] Visual branch visualization: By generating a Class Activation Map (CAM), the influence degree of different regions in the image on the model decision-making is visualized, and the attention heatmap corresponding to each attribute is further displayed to intuitively reflect the attention area of the model in the image feature extraction process.
[0090] Text branch visualization: Screen the top 10 attention tokens of the text input after model optimization, and for the learnable vectors contained therein, find the words closest to them through the nearest neighbor search method to reveal the semantic preference and key information capture ability of the model in text feature learning.
[0091] As Figure 4 shown, the generated class activation map reveals the attention distribution of the model when processing different types of tissues. For the dyed resection margins, the final result of the model fits the residual polyp tissue after resection. In addition, the model also has different focuses on different attributes: for the color attribute, the model mainly focuses on the dyed area; for the shape attribute, the model focuses on the dyed incision position of the resected polyp; in terms of the texture attribute, the model also focuses on the incision area, but compared with the shape attribute, the attention distribution of the texture attribute is more diffuse, and the boundary part shows a stronger attention to the large-area dyed area. This detailed attention distribution indicates that the model can effectively extract key information from different tissue features, thereby improving the classification accuracy. Further analyzing the attention tokens of the text branch, as shown in Table 2, we find that the model can extract tokens highly relevant to visual information from the attribute description. For "dyed resection margins", the model not only focuses on the blue color change, but also pays attention to the rough (incision) or smooth (near the incision) texture and the blue or sharp (clear) boundary, which shows its sensitivity to common signs in clinical medicine. In addition, the model also shows a relatively high attention to some general descriptive terms (such as regions, hues, shapes, surfaces, textures, surface, edges, etc.). It is worth noting that the learnable vectors appearing in the top 10 attention tokens are all non-Latin characters, which may reflect that the model captures some abstract semantic information during the learning process.
[0092] The above results indicate that this system can not only extract key features from the visual modality, but also capture semantic information highly relevant to visual features through the text modality, thus realizing the effective fusion and collaborative decision-making of multimodal information, and significantly improving the accuracy and robustness of the classification task.
[0093] Table 2
[0094]
[0095]
[0096]
[0097]
[0098]
[0099] In summary, the few-shot digestive disease diagnosis method (ADPL) based on medical attribute driving provided by the present invention includes: attribute description generation, which is used to ask a large language model in the way of visual question answering (VQA) to obtain the medical attributes contained in the endoscopic image; cross-modal prompt optimization, which is used to optimize the prompt by using the attribute description generated by the attribute description generation module. By introducing learnable visual prompts and attribute prompts, and using the AVF module to bridge the visual prompt and the attribute prompt, medical attribute knowledge is injected into the visual prompt; the loss function optimization module is used to optimize the prediction accuracy of each disease label in the multi-label classification task through the cross-entropy loss function, and update the model parameters through the backpropagation algorithm to improve the overall classification performance of the model. This system generates medical attribute descriptions of endoscopic images through a large language model, and uses learnable visual prompts and attribute prompts for cross-modal prompt optimization. During the prompt optimization process, the AVF module is used to achieve the efficient fusion of attribute semantics and visual features, thereby enhancing the model's semantic understanding and utilization of medical images. Compared with the traditional classification method that only relies on single image data, the present invention significantly improves the accuracy and interpretability of digestive disease classification by introducing attribute description generation based on a large language model, introducing text information other than images, and performing cross-modal prompt optimization at the same time, providing a solution that takes into account both performance and semantic understanding for endoscopic image analysis in the few-shot learning scenario.
[0100] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. A few-shot digestive disease diagnosis method driven by medical attributes, characterized in that, Including the steps: Ask the large language model through visual question answering to obtain the medical attributes contained in the endoscopic image; Use the attribute description generated in the attribute description generation step to optimize the prompt, introduce learnable visual prompts and attribute prompts, and use the AVF module to bridge the visual prompt and the attribute prompt to inject medical attribute knowledge into the visual prompt; Through the loss function optimization module, optimize the prediction accuracy of each disease label in the classification task, and update the model parameters through the backpropagation algorithm to improve the overall classification performance of the model and achieve disease diagnosis and prediction.
2. The few-shot digestive disease diagnosis method based on medical attribute driving according to claim 1, characterized in that, The learnable visual prompt is based on the CLIP model, and the CLIP model consists of a two-tower structure composed of a visual encoder ViT-B / 16 and a text encoder transformer.
3. The few-shot digestive disease diagnosis method based on medical attribute drive according to claim 1, characterized in that In the visual question answering method, a customized prompt template and a few-shot image set are used to ask questions to the large language model GPT-4.
4. The few-shot digestive disease diagnosis method based on medical attribute drive according to claim 3, wherein The prompt template is specified as an interrogation structure including four attributes: color, shape, texture, and boundary, and is interrogated for a certain lesion category to obtain normalized semantic information about different attributes of a certain category.
5. The few-shot digestive disease diagnosis method based on medical attribute drive according to claim 3, characterized in that, The few-shot image set is instances with different numbers of samples for each class, including 4, 8, and 16, which are used as the image input in visual question answering to prompt the large language model to obtain visual information about attributes in the image features.
6. The few-shot digestive disease diagnosis method based on medical attribute drive according to claim 1, characterized in that The medical attributes include color, shape, texture, and boundary, which are used to provide structured semantic guidance for the model.
7. The method for diagnosing rare-sample digestive tract diseases driven by medical attributes according to claim 1, wherein The visual prompt uses a learnable prompt with a depth of 12 and a length of 16.
8. The method for diagnosing rare-sample digestive tract diseases driven by medical attributes according to claim 1, wherein The attribute prompt adopts a diversified design, specifically set that 4 different attribute prompts correspond to different attribute descriptions for prompt learning, and the depth of each attribute prompt is 1 and the length is 16.
9. The few-shot digestive disease diagnosis method based on medical attribute drive according to claim 1, characterized in that, The AVF module is based on the attention mechanism, adaptively focuses on the key parts of the attribute prompt and the visual features, and realizes the efficient fusion of the attribute semantics and the visual features, thereby enhancing the model's semantic understanding and utilization of medical images.
10. A few-shot digestive disease diagnosis device driven by medical attributes, characterized in that, Including a processor and a memory for storing processor-executable instructions, and when the processor executes the instructions, the steps of the method according to any one of claims 1-9 are implemented.
Citation Information
Patent Citations
Large model prompt learning method based on category attribute knowledge enhancement
CN116994098A
Transform-based gastrointestinal endoscopic image classification and segmentation method
CN117830631A
Chest medical image multi-label intelligent diagnosis algorithm based on multi-modal comparative learning
CN118136239A
Cervical panoramic image few-sample classification method based on visual guidance and language prompt
CN118230052A
Zero sample image classification method and system based on prompt guidance
CN118691899A
Cited By
Multi-modal large model deployment method and system for space governance
CN121010861A