基于视觉和文本语义对齐的舌象多任务系统、方法、终端及介质

The tongue image multi-task system, which aligns visual and textual semantics, solves the problems of insufficient fine-grained feature capture and poor text adaptability in tongue diagnosis. It achieves accurate alignment and multi-task sharing between tongue image and labeled text, thereby improving the intelligence and standardization of TCM tongue diagnosis.

CN122417352APending Publication Date: 2026-07-17SHANGHAI NAT GRP HEALTH TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI NAT GRP HEALTH TECH CO LTD
Filing Date
2026-03-13
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing visual and text alignment technologies for tongue diagnosis suffer from insufficient fine-grained feature capture and poor text adaptation, resulting in insufficient accuracy in tongue diagnosis.

Method used

A tongue image multi-task system based on visual and text semantic alignment is adopted, including a data preprocessing module, a multi-functional architecture module, a multi-task joint fine-tuning module, and a general embedding output module. Through an image encoder, a text encoder, and a cross-modal alignment unit, the tongue image and label text are converted into general features, and then jointly fine-tuned and projected to the same dimension vector space to achieve alignment of visual features and text features and multi-task sharing.

Benefits of technology

It improves the intelligence and standardization of tongue diagnosis, and can accurately capture the microscopic pathological features of the tongue and the deep semantic relationship between text tags, thereby enhancing the accuracy and reliability of TCM tongue diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122417352A_ABST
    Figure CN122417352A_ABST
Patent Text Reader

Abstract

本发明提供基于视觉和文本语义对齐的舌象多任务系统、方法、终端及介质,包括:数据预处理模块,用于对原始舌象图像及标签文本进行预处理,以生成舌象图像与标签文本相配对的数据集;多功能架构模块,用于输出舌象图像的视觉特征和标签文本的动态文本特征;多任务联合微调模块,用于对所述视觉特征和动态文本特征进行联合微调;通用嵌入输出模块,用于将对齐后的所述视觉特征和动态文本特征投射成同一维度向量,以形成目标舌象多任务共享的通用嵌入空间;多任务预测测试模块,用于将被投射成同一维度向量的视觉特征和动态文本特征,与通用嵌入空间进行相似度匹配,并基于相似度匹配结果对所述舌象多任务进行预测。
Need to check novelty before this filing date? Find Prior Art