Alzheimer's disease image classification method and system based on multi-modal large language model

By combining a multimodal large language model with cross-modal learning and diversity-driven token pruning techniques, the problems of high computational cost and low transparency in existing technologies are solved, achieving efficient and accurate Alzheimer's disease image classification and interpretability analysis.

CN121834441APending Publication Date: 2026-04-10SANYA HOSPITAL OF TRADITIONAL CHINESE MEDICINE +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing deep learning-based Alzheimer's disease assessment methods suffer from high computational costs, low efficiency, poor decision-making transparency, and failure to effectively integrate multimodal data, which limits their clinical application.

Method used

Employing a multimodal large language model, MRI image features are extracted through cross-modal contrastive learning. Combined with diversity-driven token pruning and visual-semantic pre-fusion techniques, structured radiology reports are generated and deeply integrated with clinical text, outputting high-precision and interpretable image classification results.

Benefits of technology

It achieves efficient and accurate image classification for Alzheimer's disease, provides reliable auxiliary analysis and decision support, and improves the transparency of the model and the efficiency of multimodal data utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834441A_ABST
    Figure CN121834441A_ABST
Patent Text Reader

Abstract

The invention discloses an Alzheimer's disease image classification method and system based on a multi-modal large language model, and the method comprises the steps: firstly obtaining a brain MRI image and clinical text data of a patient suffering from the Alzheimer's disease, and extracting MRI image features through a three-dimensional image encoder, thereby obtaining an initial visual token sequence; secondly, the initial visual token sequence is compressed to form visual embedding, and a medical multi-modal large language model is utilized to generate a radiology report according to an MRI image and splice the radiology report with clinical text data to form text input; and then fusing the initial visual token sequence with the embedding representation of the text input to generate text embedding. And finally, splicing visual embedding and text embedding to form multi-modal representation, inputting the multi-modal representation into the large language model, and outputting an Alzheimer disease image classification result and interpretability information. According to the method, the pathological features and semantic description of the patient are analyzed in combination with the multi-modal large language model, and a classification result with high accuracy and clinical interpretability is given.
Need to check novelty before this filing date? Find Prior Art