基于视觉和文本语义对齐的舌象多任务系统、方法、终端及介质
The tongue image multi-task system, which aligns visual and textual semantics, solves the problems of insufficient fine-grained feature capture and poor text adaptability in tongue diagnosis. It achieves accurate alignment and multi-task sharing between tongue image and labeled text, thereby improving the intelligence and standardization of TCM tongue diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI NAT GRP HEALTH TECH CO LTD
- Filing Date
- 2026-03-13
- Publication Date
- 2026-07-17
AI Technical Summary
Existing visual and text alignment technologies for tongue diagnosis suffer from insufficient fine-grained feature capture and poor text adaptation, resulting in insufficient accuracy in tongue diagnosis.
A tongue image multi-task system based on visual and text semantic alignment is adopted, including a data preprocessing module, a multi-functional architecture module, a multi-task joint fine-tuning module, and a general embedding output module. Through an image encoder, a text encoder, and a cross-modal alignment unit, the tongue image and label text are converted into general features, and then jointly fine-tuned and projected to the same dimension vector space to achieve alignment of visual features and text features and multi-task sharing.
It improves the intelligence and standardization of tongue diagnosis, and can accurately capture the microscopic pathological features of the tongue and the deep semantic relationship between text tags, thereby enhancing the accuracy and reliability of TCM tongue diagnosis.
Smart Images

Figure CN122417352A_ABST