Multi-modal large model-based traditional Chinese medicine tongue diagnosis medicine drink recommendation method

By extracting tongue features using BLIP-2 and Q-Former structures, and combining them with a large-scale language model and a medicinal decoction knowledge base, a complete link from tongue image recognition to medicinal decoction recommendation in traditional Chinese medicine tongue diagnosis was realized. This solved the shortcomings of traditional Chinese medicine tongue diagnosis in multimodal fusion and cross-modal information alignment, and improved diagnostic accuracy and the efficiency of personalized medicinal decoction recommendation.

CN120913769AActive Publication Date: 2025-11-07GUILIN UNIV OF ELECTRONIC TECH +1

Patent Information

Application Number
CN202511060346.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-07
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

In clinical applications, TCM tongue diagnosis suffers from several drawbacks. Diagnostic results are influenced by factors such as the doctor's experience, light, and temperature. It lacks standardized quantitative analysis and a complete closed-loop system from tongue diagnosis to herbal recommendations. In particular, it has shortcomings in multimodal fusion and cross-modal information alignment, which hinders the implementation of personalized health interventions.

Method used

The BLIP-2 multimodal coding framework, combined with an improved Q-Former structure and BridgeToken, is used to extract tongue image features, generate tongue image symptom descriptions through a large language model, and make personalized medicine recommendations using a medicine knowledge base, thus realizing a complete chain of "tongue image recognition - health reasoning - medicine recommendation".

Benefits of technology

It significantly improved the detection rate and F1 score of tongue features, shortened the inference time, enhanced the accuracy and interpretability of diagnosis, enabled the rapid deployment and traceability of personalized herbal recommendations, and enhanced the professionalism and reliability of TCM health intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913769A_ABST
    Figure CN120913769A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large health and artificial intelligence, and particularly provides a traditional Chinese medicine tongue diagnosis medicine drink recommendation method based on a multi-modal large model, and the method comprises the steps: employing BLIP-2 as a multi-modal coding framework to extract a complete query vector group according to a preprocessed tongue picture image; constructing a visual semantic vector based on the complete query vector group; a large language model is adopted, and tongue picture disease description is generated based on the visual semantic vector; constructing a medicinal drink knowledge base, and retrieving tongue picture disease description and user self-description based on the medicinal drink knowledge base to obtain final candidate medicinal drinks; and based on a large language model, according to the tongue picture disease description and the final candidate medicine drink, generating a recommendation scheme in combination with a target template. By means of the cross-modal representation capability of the BLIP-2 model, user text description and tongue picture image features are subjected to combined modeling, personalized health state description is generated based on the language model, and a more accurate semantic basis is provided for subsequent medicine beverage recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the fields of health and artificial intelligence, in particular to a traditional Chinese medicine tongue diagnosis medicine and drink recommendation method based on a multi-modal large model. BACKGROUND

[0002] Traditional Chinese medicine tongue diagnosis is one of the key components of the traditional Chinese medicine diagnosis and treatment system. It can assist in judging the internal organ function state, the rise and fall of qi and blood, and the nature of the disease by observing the color, shape, texture, dryness and wetness, size, position change of the tongue, and the characteristics of the tongue fur. According to the theory of traditional Chinese medicine, the tongue is considered to be the external reflection of the heart, and the physiological functions of the five zang organs and six fu organs depend on the nourishment of stomach qi. Different dietary intake and physical conditions can be reflected on the tongue. Therefore, the tongue can directly reveal the running state of the zang and fu organs. Among them, the tongue texture mainly reflects the filling degree of the zang and fu essence, and the tongue fur can represent the nature of the pathogenic factor and the strength of the stomach qi. Therefore, based on the detailed analysis of the tongue, combined with the four diagnostic methods of observation, listening, questioning and palpation, doctors can more comprehensively understand the patient's internal health status, and provide an important basis for clinical diagnosis and treatment of diseases. However, there are still some limitations in the clinical application of traditional Chinese medicine tongue diagnosis in China at present. The diagnosis results are affected by many factors such as the personal clinical experience of the doctor, subjective judgment, dietary habits of the patient, and the light and temperature of the treatment environment. There is a lack of unified and standardized quantitative analysis indicators and diagnosis system. This uncertainty limits the promotion and application of tongue diagnosis in the intelligent development of traditional Chinese medicine.

[0003] Although the current intelligent research of traditional Chinese medicine tongue diagnosis has achieved preliminary results, it still mainly stays at the level of traditional image processing and health status recognition, and has not yet effectively constructed a complete "tongue image recognition-symptom reasoning-medicine and drink recommendation" closed-loop system. Especially in the aspects of multi-modal fusion, cross-modal information alignment and fine-grained tongue feature analysis, there are still obvious shortcomings, which makes it difficult to land the personalized health intervention scheme. First, the existing tongue image analysis methods are mostly based on shallow visual models or general computer vision architectures, which are limited by light changes, individual differences and shooting quality. The extraction of fine-grained features such as color, cracks, tooth marks and fur thickness in the tongue image has the problem of insufficient robustness, which affects the accurate reasoning of the health status. Second, most traditional Chinese medicine tongue diagnosis systems use a single modal reasoning path, which cannot effectively fuse the joint semantics of user self-reported symptoms and tongue image, thereby limiting the accuracy and explainability of traditional Chinese medicine syndrome identification. In addition, existing patents and researches still mainly focus on tongue image processing, diagnosis result analysis or interrogation systems, and lack of landing application of tongue diagnosis results in the health intervention level, especially in the medicine and drink recommendation which is still in the blank stage in the aspect of personalized regulation of traditional Chinese medicine. SUMMARY

[0004] In view of this, the present application proposes a traditional Chinese medicine tongue diagnosis medicine and drink recommendation method based on a multi-modal large model, aiming to improve the diagnostic accuracy of traditional Chinese medicine tongue diagnosis, standardize the clinical diagnosis and treatment path, and provide feasible suggestions for promoting the development of traditional Chinese medicine and integrated traditional Chinese and Western medicine, so as to solve the problems existing in the prior art.

[0005] To achieve the above-mentioned purpose, the present application proposes a traditional Chinese medicine tongue diagnosis medicine and drink recommendation method based on a multi-modal large model, comprising:

[0006] According to the pre-processed tongue image, BLIP-2 is used as a multi-modal encoding framework to extract a complete query vector group;

[0007] Based on the complete query vector group, a visual semantic vector is constructed;

[0008] A large language model is used to generate a tongue image disease description based on the visual semantic vector;

[0009] A medicine and drink knowledge base is constructed, and the tongue image disease description and user self-reports are retrieved based on the medicine and drink knowledge base to obtain final candidate medicines and drinks;

[0010] Based on the large language model, a recommended scheme is generated according to the tongue image disease description and the final candidate medicines and drinks combined with a target template.

[0011] Further, the process of extracting tongue image features using BLIP-2 as a multi-modal encoding framework includes:

[0012] Based on the visual encoder, a feature sequence is generated from the processed tongue image;

[0013] Based on the improved Q-Former structure, information extraction and global aggregation are performed on the feature sequence to obtain a complete query vector group.

[0014] Further, the improved Q-Former structure queries the visual features in the multi-head cross-attention mechanism based on the query vector group, which is represented as follows:

[0015]

[0016] wherein, represents the query vector group, is the BridgeToken, is the number of learnable QueryToken, represents the query vector.

[0017] Further, the improved Q-Former structure includes introducing BridgeToken in the Q-Former network;

[0018] In the cross-attention stage, the feature sequence is weighted based on the query vector group to generate a new query vector, key information extraction is performed on the feature sequence based on the new query vector, and weights are generated based on the BridgeToken, and the discriminative regions in the feature sequence are dynamically converged based on the weights;

[0019] In the self-attention stage, the new query vector is weighted to obtain a final query vector, and local detailed information is absorbed through the BridgeToken;

[0020] Based on the final query vector, a complete query vector group is constructed.

[0021] Further, the process of constructing a visual semantic vector based on the tongue image features includes:

[0022] The improved Q-Former structure is trained for visual-linguistic representation based on the contrast learning and discriminative pairing relationship method, and visual-linguistic generation training is performed based on the language generation task of the large language model, and the visual semantic vector is obtained according to the hidden state of the BridgeToken in the complete query vector group.

[0023] Further, the process of generating a tongue symptom description based on the constructed visual semantic vector includes:

[0024] The visual semantic vector is linearly mapped to the hidden dimension of the large language model based on the weight matrix to construct an output tensor after mapping;

[0025] A visual-linguistic prompt is constructed based on the output tensor after mapping;

[0026] The visual-linguistic prompt is input into the OPT model and the Flan-T5 model, respectively;

[0027] The tongue symptom description is generated based on the output of the OPT model and the Flan-T5 model.

[0028] Further, the medicine knowledge base includes the names, materials, methods, effects, indications, and sources of several medicines.

[0029] Further, the process of retrieving the tongue symptom description and the user's self-description based on the medicine knowledge base includes:

[0030] The candidate medicine is obtained by performing nearest neighbor retrieval on the tongue symptom description and the user's self-description and the vectors in the medicine knowledge base.

[0031] The final candidate medicine is determined by performing Boolean filtering on the candidate medicine based on the indications and effects fields in the medicine knowledge base and the keywords of the tongue symptom description.

[0032] Further, the preprocessing process of the tongue image includes:

[0033] Boundary detection is performed on the tongue image based on a target detection model, the tongue image is cropped according to the detection result, and the cropped image is standardized.

[0034] Compared with the prior art, the present application has the following advantages:

[0035] The present application adopts BLIP-2Q-Former+BridgeToken, and the model adaptively aggregates global and local visual information in cross-attention, so that the detection rate and F1 value of fine-grained features such as tongue coating thickness, color, cracks and tooth marks are significantly higher than those of traditional U-Net / ViT and other simple visual networks, meeting the clinical demand for high-resolution diagnostic signals.

[0036] The present application directly maps the visual PromptToken to OPT or Flan-T5 to generate a structured tongue disease description in one forward inference; experimental data show that the CIDEr and METEOR indicators are improved by more than 10% compared with the public baseline model, and the inference time is shortened by about 30%, which is suitable for mobile and cloud deployment;

[0037] The present application converts the disease description into a vector retrieval query, combines the Faiss semantic index to quickly locate the medicine and drink items, and then generates a personalized conditioning plan containing

efficacy

material

source

[0038] The method of the present application only needs single interaction from tongue image input to medicine and drink scheme output, that is, to realize the complete link of "tongue image recognition -> health reasoning -> medicine and drink recommendation", which greatly improves the user experience and clinical application efficiency.

[0039] The present application adopts physical separation based on visual-linguistic backbone, knowledge base retrieval and generation model, and any module upgrade does not affect the rest, when new medicine and drink or updated medical cases are added, only the item needs to be imported and the vector index needs to be reconstructed, without retraining the backbone network.

[0040] The present application automatically presents the compatibility principle and literature source of the recommended text, which is convenient for doctors to review and users to verify; the knowledge base can be managed hierarchically according to regulations to realize safe and controllable intelligent Chinese medicine health services, which is superior to the best known art in terms of tongue feature analysis accuracy, cross-modal generation consistency, recommendation traceability and system maintainability, etc., and provides a more professional and reliable solution for intelligent and personalized medicine and drink intervention of Chinese tongue diagnosis. BRIEF DESCRIPTION OF DRAWINGS

[0041] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of preferred embodiments, and are not meant to limit the present application. In the drawings:

[0042] Figure 1 The overall implementation flowchart of the traditional Chinese medicine tongue diagnosis medicine recommendation method based on a multi-modal large model proposed in the present application is shown in the figure.

[0043] Figure 2 The improved Q-Former model structure diagram in the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0044] Exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be accurately conveyed to those skilled in the art. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0045] The present embodiment proposes a traditional Chinese medicine tongue diagnosis medicine recommendation method based on a multi-modal large model, as shown in the figure, which includes: Figure 1

[0046] Step 1: Tongue image acquisition and preprocessing

[0047] In order to obtain clear and diagnostic tongue image, the present embodiment first requires the user to upload the front tongue photo through the applet using the mobile device (such as a smart phone). Since the original image usually contains too much irrelevant area (such as face, background), after the system receives the image, it will call the lightweight target detection model YOLO to accurately locate and crop the tongue area. YOLO uses a multi-scale feature pyramid structure and a convolutional network to detect possible tongue targets in the image in real time and outputs the corresponding bounding box. The system automatically crops the original image according to the bounding box coordinates and extracts the effective tongue area as the input image for subsequent visual modeling.

[0048] ​The "tongue body automatic positioning and cropping" process realized by YOLO not only effectively removes background interference, facial occlusion and other non-key information, but also significantly improves the focusing ability of subsequent visual feature extraction on fine-grained features such as tongue fur thickness, color, cracks, and tooth marks. At the same time, this method takes into account user experience, allowing users to obtain clear and feature-rich tongue image even in non-professional shooting environments, enhancing the system's versatility and robustness.

[0049] The cropped image will be standardized (including normalization, size scaling, and channel arrangement adjustment) to adapt to the image input requirements of the multi-modal visual coding model BLIP-2, ensuring that the model can fully capture the key structure and texture information of the tongue image in the encoding stage.

[0050] Step two: tongue image feature extraction

[0051] This embodiment uses BLIP-2 (Bootstrapping Language-Image Pre-training v2) as the multi-modal encoding framework for tongue image, and introduces BridgeToken in its Q-Former to further enhance the visual semantic aggregation ability, as shown in the structure of Figure 2 BLIP-2 decouples the frozen visual encoder (Vision Transformer, hereinafter referred to as ViT) from the trainable query transformer (Q-Former), allowing the model to accurately capture fine-grained features of the tongue image while maintaining efficient training, and seamlessly integrate with large language models (OPT or Flan-T5) in the future.

[0052] Step 2.1 Visual encoding

[0053] The normalized and size-adjusted tongue image is first input into the ViT visual encoder. ViT divides the image into fixed-size patches and maps them into a high-dimensional vector sequence

[0054]

[0055] where is used to represent the global semantics of the entire tongue image, represents the local features of the th image block, is the number of image patches output by the visual encoder, hidden feature dimension (ViT and Q-Former are unified to 768).

[0056] Step 2.2 Q-Former and BridgeToken aggregation

[0057] Subsequently, the feature sequence is fed into a trainable Q-Former. The Q-Former structure is shown in Figure 2 , which consists of a set of learnable query vectors:

[0058]

[0059] where is the BridgeToken, is the number of learnable QueryTokens. Under the multi-head cross-attention mechanism, the Q-Former uses to perform information extraction and global aggregation on .

[0060] (1) Cross-attention stage: from image to query

[0061] In the first cross-attention layer, the query matrix is , and the keys and values come from the visual features . The weight calculation and feature aggregation formula are:

[0062]

[0063] where is a trainable linear mapping. The bridge vector b establishes attention connections with all patches in this process, obtaining a set of weights , dynamically converging the discriminative regions of the entire tongue image.

[0064] (2) Self-attention stage: information exchange between queries

[0065] The updated query enters the same layer self-attention:

[0066]

[0067] In this stage, the BridgeToken broadcasts its global semantic to the remaining QueryTokens, and at the same time absorbs their local detailed information, forming a "global-local" fusion representation.

[0068] Step 2.3 Output and LLM interface

[0069] After layers of alternating stacking (cross-attention + self-attention + feedforward network), we get

[0070]

[0071]

[0072] where​ is the complete query matrix output by Q-Former, is the hidden state of the corresponding BridgeToken in Q-Former, as the visual global semantic vector. BridgeToken is specially learned in the Forward process to capture and align the fine-grained features of tongue thickness, color, cracks, and tooth marks, etc. It has stronger cross-modal information transmission ability than the original BLIP-2 in experiments.

[0073] Step 2.4 Cross-modal pre-training target

[0074] BLIP-2 still adopts a “two-stage” pre-training strategy:

[0075] Stage 1: Visual-linguistic representation learning (ITC+ITM)

[0076] • ITC (Image-Text Contrastive) contrastive learning, which brings paired images and texts closer and pushes non-paired samples further apart.

[0077] • ITM (Image-Text Matching) discriminates the pairing relationship and strengthens the consistency discrimination ability of images and texts.

[0078] Stage 2: Visual-to-linguistic generation learning

[0079] After linear mapping, the Q-Former output is input as PromptEmbedding into the frozen large language model (OPT or Flan-T5), which further refines the mapping of visual semantics to language space through the language generation task.

[0080] After completing the above pre-training, the obtained visual semantic vector as formula (6) (i.e. taking the hidden state corresponding to BridgeToken) will be projected into the input space of LLM, providing high-quality multi-modal semantic basis for subsequent health status description generation and RAG+T5 medicine recommendation.

[0081] Step three: Tongue image description generation

[0082] After completing 2 tongue image feature extraction, the visual semantic vector output by Q-Former (including the BridgeToken row) needs to be seamlessly connected with the large language model (LLM) to generate structured and medically accurate tongue disease descriptions. This invention supports both OPT (Decoder-Only) and Flan-T5 (Encoder-Decoder) LLMs; the access methods of the two are as follows.

[0083] Step 3.1 Visual Prompt Construction

[0084] 1. Linear Mapping

[0085]

[0086] is the weight matrix of the single-layer linear mapping, is the bias vector of the mapping, and the function is to map from the Q-Former hidden dimension (=768) to the hidden dimension used by the back-end large language model (OPT or Flan-T5) , so as to realize the matching of tensor shape and numerical range.

[0087] 2. Visual Prompt Token Set

[0088] Take the first rows (including the BridgeToken row) to form

[0089]

[0090] This as a visual-linguistic unified prompt (VisualPrompt), carries the global tongue image semantics.

[0091] Step 3.2 Input the visual-linguistic unified prompt into the OPT model and the Flan-T5 model respectively

[0092] (1) OPT path (Decoder-Only)

[0093] OPT self-attention only depends on single sequence input. Therefore, the is directly concatenated to the front of the user text embedding:

[0094]

[0095] where is the embedding of the user input text Token, such as the first word "tongue", the second word "picture", etc.; in the OPT path, it is directly concatenated after , is the number of visual Prompt Token, and is the number of user text Token, and the multi-head self-attention of OPT is executed at the decoding step :

[0096]

[0097] where Contains visual Prompt and historical word vectors. Since visual Token and text Token share the same attention matrix, the model can adaptively retrieve the global representation and local details of the tongue image when generating each word, ensuring that the disease description is consistent with the image features.

[0098] (2) Flan-T5 path (Encoder-Decoder)

[0099] Flan-T5 divides the visual and textual information into two parts:

[0100] 1. Visual side (Encoder Memory)

[0101] Directly output as an encoder;

[0102] 2. Text side (Decoder Query)

[0103] User input symptom text is embedded as .

[0104] At the first decoding step, perform standard cross-attention:

[0105]

[0106] where is the current decoding hidden state, is the projection matrix, is the scaling coefficient, the global tongue image semantics carried by the BridgeToken row can be retrieved at each time step, supplementing the details of the local patch Query, improving the generation accuracy and interpretability.

[0107] Step 3.3 Generate output and examples

[0108] Through the above mechanism, the LLM recursively generates a tongue medical description, such as: "tongue pale red, thin white moss slightly greasy, tooth marks obvious, can see the tendency of spleen deficiency and dampness, accompanied by fatigue and weakness." The output will be used as the query statement for RAG retrieval, matched with the medicinal knowledge base, and finally return a personalized conditioning scheme.

[0109] Step four: RAG retrieval and medicinal drink recommendation

[0110] After obtaining the structured disease text through "Step three tongue description generation", the present application outputs personalized traditional Chinese medicine and medicinal drink schemes for the user through retrieval-augmented generation (RAG). The process consists of three steps: retrieval, filtering, and generation.

[0111] ​Step 4.1 Knowledge base and index construction

[0112] To interface with the Chinese medicine drinking scene, this embodiment prepares an Excel format of medicine drinking knowledge base, which collates classics such as “Prescriptions” and “Medical Classics” and modern clinical medical records, covering more than a dozen basic medicine drinking items .

[0113] Vectorization uses sentence-BERT(zh), which is a BERT-based sentence vector encoding framework that specifically maps complete sentences or short texts into fixed-length dense vectors, so that synonymous sentences are close to each other in semantic space. The “efficacy + applicable syndrome” field is encoded into a 768-dimensional vector; stored in the Faiss index (L2&IVF), which is an open source vector retrieval library published by Meta AI, used for efficient search of nearest neighbors in large-scale vector sets. L2 is the retrieval distance used as a similarity measure. IVF (Inverted File) is an inverted file nearest neighbor index structure that first clusters vectors into several “center” buckets, and then only performs accurate distance calculation within the nearest bucket, which can significantly reduce search time and memory. The retrieval interface supports top-k similarity query and threshold filtering (default ).

[0114] Step 4.2 Retrieval and candidate item filtering

[0115] (1) Query construction

[0116]

[0117] Example: “Tongue coating yellow and greasy, stomach and intestines hot; patient reports dry mouth, fatigue.”

[0118] (2) Vector matching

[0119] Take and do nearest neighbor search with knowledge base vectors; return .

[0120] (3) Rule filtering

[0121] According to the keywords (such as “clearing heat and dampness”) in the “applicable syndrome / efficacy” field and the LLM description, a Boolean filter is performed to obtain the final candidate .

[0122] Step 4.3 Generate personalized recommendations using Flan-T5

[0123] The same Flan-T5 as the tongue description (already fine-tuned on a small number of “disease ↔ medicine drinking instructions” alignment examples) is selected as the generator.

[0124] Encoder input:

[0125] [DESCR] Yellow and greasy tongue fur … bitter taste in the mouth;

[0126] [CAND] 1) Coix seed and red bean soup | Effect = clearing heat and removing dampness; …

[0127] 2) Sanren decoction | Effect = promoting qi movement; …

[0128] Decoder target template:

[0129] It is recommended to use { recipe name}, { material compatibility principle}, which is suitable for { applicable syndrome description}.

[0130] Method: { steps}. Source: { source}.

[0131] The performance of the Bridge-BLIP model proposed in the application and the original BLIP-2 model is compared, and the results are shown in Tables 1, 2 and 3:

[0132] Table 1

[0133]

[0134] Table 2

[0135]

[0136] Table 3

[0137]

[0138] The experimental results show that the Bridge-BLIP proposed in the application keeps or slightly exceeds the original BLIP-2 performance in the public cross-domain retrieval and description task, and achieves good performance (zero-shot CIDEr 122.1, few-shot CIDEr 143.8) in the tongue diagnosis description task, verifying the effectiveness and stability of the mechanism in the medical fine-grained scene.

[0139] It can be understood that the present application proposes to introduce a Bridge Token mechanism in the BLIP-2 multimodal large model structure, which strengthens the transmission and alignment ability of visual features in the language understanding process through a learnable global aggregation token, effectively improves the perception accuracy of small pathological features in the tongue image and the cross-modal expression ability. Compared with the original BLIP-2, the introduction of Bridge Token realizes efficient mapping of image semantic information to the language model, effectively enhances the aggregation and transmission ability of image information in the language modeling process, and significantly improves the perception and expression effect of fine-grained features of tongue image (such as tooth marks, cracks, moss color, etc.). The model not only realizes the cross-modal alignment and comprehensive reasoning of tongue image and user symptom description, but also connects the classical medicine knowledge base through the integration of RAG (retrieval augmented generation) technology, realizes personalized medicine recommendation based on the diagnosis result, and forms an intelligent closed loop of "tongue image recognition-health status analysis-regulation scheme generation". Compared with the previous scheme of providing only diagnosis information, the present application not only improves the accuracy and interpretability of the tongue diagnosis intelligent system, but also expands its practical application scene in traditional Chinese medicine health intervention and regulation suggestion, fills the key blank of current traditional Chinese medicine tongue diagnosis in intelligent recommendation, and provides a valuable path for the deep integration of traditional medicine and artificial intelligence.

[0140] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application and not to limit it, although the present application has been described in detail with reference to the above examples, those skilled in the art should understand that: the specific embodiments of the present application can still be modified or replaced, without departing from the spirit and scope of the present application. Any modification or equivalent replacement, which should be covered within the protection scope of the claims of the present application.

Claims

1. A multi-modal large model-based traditional Chinese medicine tongue diagnosis medicine recommendation method, characterized in that, The method comprises the following steps: According to the pretreated tongue image, BLIP-2 is used as a multi-modal encoding framework to extract a complete query vector group; Based on the complete query vector group, a visual semantic vector is constructed; A large language model is used to generate a tongue disease description based on the visual semantic vector; A drug knowledge base is constructed, and the tongue disease description and user self-reports are retrieved based on the drug knowledge base to obtain the final candidate drug; Based on the large language model, the tongue disease description and the final candidate drug are combined with the target template to generate a recommended scheme.

2. The multi-modal large model-based traditional Chinese medicine tongue diagnosis medicine recommendation method according to claim 1, characterized in that, The process of extracting tongue image features using BLIP-2 as a multi-modal encoding framework includes: Based on the visual encoder, a feature sequence is generated from the processed tongue image; Based on the improved Q-Former structure, information extraction and global aggregation are performed on the feature sequence to obtain a complete query vector group.

3. The multi-modal large model-based traditional Chinese medicine tongue diagnosis medicine recommendation method according to claim 2, characterized in that, The improved Q-Former structure queries the visual features in the multi-head cross-attention mechanism based on the query vector group, which is represented by the following formula: , wherein, denotes a set of query vectors, is a BridgeToken, is the number of learnable QueryTokens, denotes a query vector.

4. The multi-modal large model-based traditional Chinese medicine tongue diagnosis medicine recommendation method according to claim 2, characterized in that, The improved Q-Former structure includes introducing BridgeToken in the Q-Former network; In the cross-attention stage, the feature sequence is weighted calculated based on the query vector group to generate a new query vector, and the feature sequence is extracted based on the new query vector. At the same time, BridgeToken generates a weight, and the discriminative region in the feature sequence is dynamically converged based on the weight; In the self-attention stage, the new query vector is weighted calculated to obtain the final query vector, and the local detail information is absorbed through BridgeToken; Based on the final query vector, a complete query vector group is constructed.

5. The multi-modal large model-based traditional Chinese medicine tongue diagnosis medicine recommendation method according to claim 2, characterized in that, The process of constructing a visual semantic vector based on the tongue image features includes: Based on the contrast learning and discriminative pairing relationship method, the improved Q-Former structure is trained for visual-linguistic representation, and the visual-linguistic generation training is performed based on the language generation task of the large language model. The visual semantic vector is obtained based on the hidden state of BridgeToken in the complete query vector group.

6. The multi-modal large model-based traditional Chinese medicine tongue diagnosis medicine recommendation method according to claim 1, characterized in that, The process of generating a tongue disease description based on the constructed visual semantic vector includes: Based on the weight matrix, the visual semantic vector is linearly mapped to the hidden dimension of the large language model to construct a mapped output tensor; Based on the mapped output tensor, a visual-linguistic prompt is constructed; The visual-linguistic prompt is input into the OPT model and the Flan-T5 model respectively; The tongue disease description is generated based on the output of the OPT model and the Flan-T5 model.

7. The multi-modal large model-based traditional Chinese medicine tongue diagnosis medicine recommendation method according to claim 1, characterized in that, The drug knowledge base includes the names, materials, methods, effects, indications, and sources of several drugs.

8. The multi-modal large model-based traditional Chinese medicine tongue diagnosis medicine recommendation method according to claim 1, characterized in that, The process of retrieving the tongue disease description and user self-reports based on the drug knowledge base includes: According to the tongue disease description and user self-reports, the nearest neighbor retrieval is performed with the vectors in the drug knowledge base to obtain candidate drugs; According to the indications and efficacy fields in the medicine knowledge base, and in combination with the keywords of the tongue disease description, the candidate medicine is subjected to Boolean screening to determine the final candidate medicine.

9. The multi-modal large model-based traditional Chinese medicine tongue diagnosis medicine recommendation method according to claim 1, characterized in that, The preprocessing process of the tongue image includes: Boundary detection is performed on the tongue image based on a target detection model, the tongue image is cropped according to the detection result, and the cropped image is subjected to standardization processing.

Citation Information

Patent Citations

  • Traditional Chinese medicine prescription compatibility method and device based on large model and mapping knowledge domain

    CN119149754A

  • Traditional Chinese medicine tongue diagnosis and prescription recommendation system based on multi-modal feature fusion

    CN120089345A

Cited By

  • Herbal medicine recommendation method and system based on semantic driving and tongue picture completion

    CN121117287A