Traditional Chinese medicine multi-mode LLM training method and application thereof

Through the integrated training method of multimodal data, a multimodal large language model of traditional Chinese medicine was generated, which solved the problem of insufficient processing capabilities of the existing traditional Chinese medicine knowledge question-and-answer method for complex classical Chinese and diverse language expressions, and achieved more accurate and personalized Chinese medicine diagnosis and teaching.

CN119941636AActive Publication Date: 2025-05-06GUANGDONG HONG KONG MACAO GREATER BAY AREA PRECISION MEDICINE RESEARCH INSTITUTE (GUANGZHOU)

Patent Information

Application Number
CN202411905586.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-06
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

The existing TCM knowledge Q&A methods are difficult to effectively deal with complex classical Chinese and diverse language expressions, and rely on a single data source, and cannot fully integrate multiple data types, resulting in limited semantic understanding and knowledge extraction capabilities.

Method used

A multimodal large language model (LLM) training method of traditional Chinese medicine is proposed. By obtaining labeled multimodal medical data, including tongue coating images, pulse images and text data, the visual feature extraction model, timing feature extraction model and text feature extraction model are trained respectively. Combining the cross-modal self-attention mechanism and the cross-modal attention mechanism, fusion features are generated and integrated into the large language model.

Benefits of technology

It significantly improves the multimodal reasoning ability and decision-making accuracy of the model, can more comprehensively integrate multiple data types, and provides more accurate and personalized Chinese medicine diagnosis and teaching suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941636A_ABST
    Figure CN119941636A_ABST
Patent Text Reader

Abstract

The invention provides a traditional Chinese medicine multi-mode LLM training method and application thereof, and relates to the field of biology and artificial intelligence crossing. The method comprises the following steps: acquiring multi-modal medical data with labels; training a visual and time sequence feature extraction model to generate tongue coating and pulse features; performing feature combination on the tongue coating and pulse features through a cross-modal self-attention mechanism model; training the text feature extraction model to generate text features; inputting the combined features and the text features into a cross-modal attention mechanism model to generate fusion features; and finally, integrating each model into a large language model, and generating a traditional Chinese medicine multi-modal LLM (Logical Language Model). According to the training method, the cooperative expression ability of the multi-modal information is improved, the deep semantic association between the features is enhanced, and the efficient fusion of the multi-modal information of the traditional Chinese medicine is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the intersection of biology and artificial intelligence, and specifically, to a multimodal LLM training method for traditional Chinese medicine and its application. Background Art

[0002] Traditional Chinese medicine is a treasure of the Chinese nation, recording the experience and knowledge of the Chinese people in fighting diseases and pursuing health for thousands of years. Creating a feasible TCM knowledge question-and-answer method can not only serve practitioners in the field of TCM, but also facilitate ordinary users to obtain and understand TCM knowledge. At present, most researchers have created TCM knowledge question-and-answer methods based on text information. This type of method has some obvious shortcomings. For example, for some ancient TCM documents, it cannot cope well with the complex classical Chinese, vernacular and other diversified language expressions, resulting in limited semantic understanding and knowledge extraction capabilities. At the same time, the existing question-and-answer methods rely on a single data source and cannot fully integrate multiple data types (such as tongue coating images, pulse conditions, etc.), resulting in limited judgment accuracy and personalized requirements. Summary of the invention

[0003] The present application aims to solve at least one of the existing problems. To this end, the present application proposes a training method for a multimodal large language model (LLM) of traditional Chinese medicine, which is used to integrate multimodal data for training, effectively improving the multimodal reasoning ability of the model and the final decision accuracy.

[0004] Specifically, this application provides the following technical solutions:

[0005] In the first aspect of the present application, the present application proposes a multimodal LLM training method for traditional Chinese medicine, the method comprising: S1, obtaining labeled multimodal medical data of traditional Chinese medicine, the multimodal medical data comprising: at least one of: tongue coating image, pulse image or and text data, the text data comprising: at least one of: patient information, main symptoms or diagnosis results; S2, using the labeled tongue coating image and pulse image to respectively train a visual feature extraction model and a temporal feature extraction model to generate tongue coating visual features and pulse temporal features; S3, using the tongue coating visual features and pulse temporal features to train a cross-modal self-attention mechanism model to generate joint features; S4, using text data to train a text feature extraction model to generate text features; S5, inputting the joint features and text features into a cross-modal attention mechanism model to generate fusion features; S6, integrating the trained visual feature extraction model, temporal feature extraction model, cross-modal self-attention mechanism model, text feature extraction model and cross-modal attention mechanism model into a large language model to generate a traditional Chinese medicine auxiliary diagnosis model.

[0006] The model training method of this application fully mines the key information in tongue coating images, pulse images and text data by training visual feature extraction models, temporal feature extraction models and text feature extraction models respectively; further improves the collaborative expression ability of multimodal information by fusing tongue coating visual features and pulse temporal features through a cross-modal self-attention mechanism model; and enhances the deep semantic association between features by fusing joint features and text features through a cross-modal attention mechanism model. Finally, all models are integrated into a large language model to achieve efficient fusion of multimodal information of traditional Chinese medicine.

[0007] In the second aspect of the present application, the present application proposes the application of TCM multimodal LLM in knowledge extraction and relationship extraction of multi-source data, and the TCM multimodal LLM is trained based on the TCM multimodal LLM training method of the first aspect.

[0008] The model trained based on the aforementioned TCM multimodal LLM training method is used for knowledge extraction and relationship extraction, which can effectively improve the comprehensiveness of knowledge extraction and the accuracy of relationship extraction.

[0009] In the third aspect of the present application, the present application proposes the application of a TCM multimodal LLM in patient diagnosis, wherein the TCM multimodal LLM is obtained by training based on the TCM multimodal LLM training method of the first aspect.

[0010] The model trained based on the aforementioned TCM multimodal LLM training method is used for patient diagnosis. It can fully integrate multimodal data such as tongue coating images, pulse images and text data, and achieve comprehensive analysis and accurate diagnosis of the patient's health status through cross-modal attention mechanism and deep reasoning ability of large language model. The model uses the joint information of visual, temporal and text features, which can not only efficiently capture the potential correlation in multimodal data, but also provide personalized disease diagnosis and treatment recommendations based on specific knowledge in the field of TCM, significantly improving the accuracy, timeliness and intelligence of diagnosis, thereby reducing the workload of doctors and providing patients with better medical services.

[0011] In the fourth aspect of the present application, the present application proposes the application of a TCM multimodal LLM in TCM teaching tasks, wherein the TCM multimodal LLM is obtained through training based on the TCM multimodal LLM training method of the first aspect.

[0012] The model trained based on the aforementioned TCM multimodal LLM training method is used for TCM teaching tasks. It can integrate multimodal information such as tongue coating images, pulse images and related text data, and comprehensively present TCM theory and clinical knowledge through cross-modal attention mechanisms and the knowledge expression capabilities of large language models. This model can not only provide an intuitive demonstration of multimodal data, but also generate detailed teaching explanations and case analyses through deep semantic understanding, helping learners deepen their understanding of TCM theory and practice. At the same time, the model can enhance the interactivity and practicality of teaching through interactive Q&A and simulated diagnosis scenarios, providing intelligent, personalized and diversified learning tools for TCM education, thereby optimizing teaching effects and learning experience.

[0013] In a fifth aspect of the present application, the present application proposes a computing device. According to an embodiment of the present application, the computing device includes: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program to implement the traditional Chinese medicine multimodal LLM training method as described in the first aspect of the present application.

[0014] The aforementioned computing device automatically executes the aforementioned multimodal LLM training method of traditional Chinese medicine through computer instructions, thereby achieving efficient automation and effectively improving training efficiency.

[0015] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 It is a flowchart of the multimodal LLM training method of traditional Chinese medicine provided by some examples of this application;

[0018] Figure 2 It is a schematic diagram of a multimodal LLM training system for traditional Chinese medicine provided by some examples of this application;

[0019] Figure 3 It is a schematic diagram of an electronic device provided in some examples of this application. DETAILED DESCRIPTION

[0020] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0021] In this article, unless otherwise specified, the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described here can be implemented in a sequence other than those illustrated or described here. In the present application, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. In the description of the present application, unless otherwise specified, "multiple" refers to two or more than two.

[0022] This application proposes a TCM multimodal LLM training method, a TCM multimodal LLM training system, an application of TCM multimodal LLM in knowledge extraction and relationship extraction of multi-source data, an application of TCM multimodal LLM in patient diagnosis, an application of TCM multimodal LLM in TCM teaching tasks, and an electronic device. The following describes each of them in detail:

[0023] Multimodal LLM training method for traditional Chinese medicine

[0024] In one aspect of the present application, the present application proposes a multimodal LLM training method for traditional Chinese medicine, referring to Figure 1 , the method comprising:

[0025] S1, obtaining labeled multimodal medical data of traditional Chinese medicine, wherein the multimodal medical data includes: at least one of a tongue coating image, a pulse image, and text data, wherein the text data includes: at least one of patient information, main complaint symptoms, or diagnosis results;

[0026] In some examples of this application, the training data of tongue coating images can be obtained from public databases, such as http: / / gitcode.com / open-source-toolkit / 7542e / blob / main / LICENSE, https: / / github.com / BioHit / TongeImageDataset, or collected by medical imaging equipment. The image data at least includes information such as tongue color, the condition of the keratinized epithelium (tongue coating) of the tongue, and the condition of the tongue papillae.

[0027] The label of the aforementioned tongue coating image can be selected from the theory of tongue diagnosis in traditional Chinese medicine.

[0028] In some examples of the present application, the pulse image is selected from the pulse data of a one-minute electrocardiogram, which can be from a public database, such as MIT-BIH Arrhythmia Database (https: / / www.physionet.org / content / mitdb / 1.0.0 / ), PhysioNet-2017 (https: / / physionet.org / content / challenge-2017 / 1.0.0 / ) INCART (http: / / physionet.org / content / incartdb / 1.0.0 / ), CPSC-2018 (http: / / 2018.icbeb.org / Challenge.html), or it can be collected by medical imaging equipment.

[0029] The label of the aforementioned pulse image can be selected from the traditional Chinese medicine pulse interpretation.

[0030] In some examples of this application, text data includes patient information (such as age, gender, medical history, etc.), main symptoms (such as headache, fatigue, stomachache, etc.), and the diagnosis or preliminary evaluation given by the doctor. The text data can be selected from the TCM-SD Chinese Medicine Dialectics Dataset (https: / / tianchi.aliyun.com / dataset / 139034) and the Chinese Medicine Literature Question Generation Dataset (https: / / tianchi.aliyun.com / dataset / 86895, https: / / github.com / Mengqi97 / chinese-medical-dataset).

[0031] To ensure data quality, each piece of multimodal medical data is professionally labeled, including the corresponding disease category, symptom characteristics and diagnostic basis, to ensure that the multimodal features can be accurately mapped with actual medical knowledge during model training. In addition, the collection of multimodal data in this application covers different populations and disease types to enhance the generalization ability of the model.

[0032] S2, using the labeled tongue coating image and pulse image, respectively train the visual feature extraction model and the temporal feature extraction model to generate the tongue coating visual features and the pulse temporal features;

[0033] In step S2, the labeled tongue coating images and pulse images will be used to train the visual feature extraction model and the temporal feature extraction model, respectively.

[0034] In some examples of the present application, for data inputs of different modalities, modality-specific loss functions are used to independently train corresponding sub-models.

[0035] The tongue coating image is combined with the corresponding TCM tongue coating theory as annotated data and input into the visual feature extraction model - Vision Transformer (ViT). The ViT model embeds the image into a high-dimensional space through learning so as to extract visual features with deep semantics from it. During the training process, the cross entropy loss function is used as the loss function of the visual module to ensure that the model can establish a correct mapping between the image input and the annotated tongue coating theory knowledge.

[0036] The pulse image is combined with the corresponding pulse label optimized ECG signal as the labeled data and input into the time series feature extraction model - Timesformer. Timesformer can capture the time series dynamic information in the pulse image and extract the time series features of the pulse image. During the training process, the dynamic time warping (DTW) is combined with the loss function of the classification task to optimize the performance of the model. DTW loss can help the model align and compare the variability of ECG signals, such as acceleration or deceleration of the heart rate, making the model more robust when extracting pulse features. The calculation formula of DTW is as follows:

[0037]

[0038] Among them, X = {x1, x2, ..., xn} and Y = {y1, y2, ..., ym} are two time series, π is the path of the time series, d(x i ,y j ) is the Euclidean distance of sequence elements;

[0039] is the prediction sequence of the model, Y is the true label sequence, and the DTW loss is expressed as:

[0040] The loss function of the cross entropy for the classification task is as follows:

[0041]

[0042] Where C is the number of categories (divided into 24 categories according to pulse description) i is the true label (one-hot representation), is the probability predicted by the model.

[0043] The two loss functions are summarized as follows: Among them, α (0.2~0.5) and β (0.5~0.8) are weight parameters.

[0044] The difference between two time series is measured by Euclidean distance, and the matching degree of the two signals is calculated based on the optimal alignment path, which helps the model learn how to deal with the temporal changes in ECG signals.

[0045] In addition, the cross entropy loss function of the classification task will be used for pulse classification to ensure that the model can accurately predict the pulse category based on the features extracted from the pulse image.

[0046] S3, using the tongue coating visual features and pulse time series features, training a cross-modal self-attention mechanism model to generate joint features;

[0047] In step S3, the tongue coating visual features and pulse timing features are mapped to the same feature space after passing through independent feature extraction models (visual feature extraction model and timing feature extraction model). In this step, by calculating the similarity between the visual features and the timing features, the self-attention mechanism model is used to weightedly fuse the features of these two modalities. Through such alignment, the model can learn the relationship between visual and timing features and perform effective fusion, thereby improving the multimodal reasoning ability of the model. However, due to the different sources and properties of visual features and timing features, their feature spaces are also different. Therefore, before inputting these features into the cross-modal self-attention mechanism model, feature embedding is first required. For visual features, the Vision Transformer (ViT) is used for embedding. The ViT model captures the global semantic information in the image through the self-attention mechanism, converts the original pixel-level information into a feature vector with deep semantics, embeds the high-dimensional image data into a unified vector space, and generates a feature vector of unified dimension. For temporal features, the ECG time series data learns the dynamic changes of the time series data through the self-attention mechanism of Timesformer, so that it can take into account the contextual information of the time series when processing the pulse image, thereby extracting the temporal features of the pulse and embedding the data into a unified vector space. Through these processes, the visual and temporal features are transformed into a vector space of the same dimension, so that they can be effectively aligned in subsequent cross-modal fusion.

[0048] The visual features and temporal features are input into the Transformer-based cross-modal self-attention mechanism model, and feature fusion is achieved through weighted learning. The model captures the global relationship between visual features and temporal features by calculating the similarity between them, and combines the information of the two modalities in a weighted manner. To ensure stable multimodal training, the model gradually unfreezes the weights of the single-modal modules in the first two steps during training, allowing the model to gradually adapt to the feature interactions between different modalities. The entire model is optimized by using a joint loss function. During training, not only the feature extraction quality of each modality is focused on, but also the interaction features between modalities are optimized to ensure the effective fusion of visual features and temporal features in the cross-modal space. Finally, these joint features are further optimized through residual connections and multi-layer perceptrons to form a cross-modal fusion representation.

[0049] S4, using text data to train a text feature extraction model and generate text features;

[0050] After the text data is processed by natural language processing technology, it is input into the text feature extraction model (such as LLaMA, GPT-4, BERT or its variants, etc.). These models will embed the text data into a high-dimensional vector space to generate a text feature vector with rich semantics. In this process, the vocabulary and sentence structure of the text are converted into a fixed-dimensional vector representation. These embedded vectors capture the important information and contextual relationships in the text.

[0051] Compared with visual and temporal features, text features usually have more complex expressions in semantic space, so in multimodal learning, text features need to be aligned with other modal features (such as tongue coating visual features and pulse temporal features). Through cross-modal alignment, text features are converted to a space that matches visual and temporal features, enabling them to be effectively fused with other modal features. This process ensures the similarity and semantic consistency between modalities through specific alignment algorithms, such as co-training or alignment methods based on attention mechanisms.

[0052] S5, inputting the joint feature and the text feature into a cross-modal attention mechanism model to generate a fusion feature;

[0053] In step S5, the joint features and text features are fused and trained through a cross-modal attention mechanism, aiming to achieve effective interaction between the two modalities. Specifically, the text features are first mapped to the joint feature space, so that the model can learn the shared representation of cross-modal features. In this process, the model calculates the similarity weights between the text features and the joint features through the dot product attention mechanism, and then performs weighted fusion based on the calculated weights to generate fused features.

[0054] In order to optimize the performance of the entire model, this step adopts an end-to-end joint training strategy. During the training process, multi-cross-modal alignment loss, classification loss, and generation task loss are used to jointly optimize the model. The cross-modal alignment loss ensures the consistency of text features and joint features in the shared space, while the classification loss and generation task loss help the model to accurately generate results in specific tasks (such as diagnosis, treatment, educational reasoning, etc.). During the training process, the parameters of all modalities are gradually updated to ensure that the model can achieve the best performance in multimodal interaction, so that the network can learn to balance the features of different modalities and ultimately generate accurate diagnosis or educational reasoning results.

[0055] S6, integrates the trained visual feature extraction model, temporal feature extraction model, cross-modal self-attention mechanism model, text feature extraction model and cross-modal attention mechanism model into the large language model to generate a traditional Chinese medicine auxiliary diagnosis model.

[0056] In step S6, in order to further optimize the semantic consistency of cross-modal features and the fusion effect of the model, contrastive learning is used to improve feature alignment capabilities. Specifically, during the training process, the visual features, temporal features, and text features in the multimodal data are treated as positive samples, that is, the corresponding visual, temporal, and text features come from different data modalities (modal features) of the same instance or the same patient. At the same time, non-corresponding visual features, temporal features, and text features are used as negative samples, that is, data modalities (modal features) from different instances or different patients are paired.

[0057] Through this method of constructing positive and negative samples, the goal of the model is to maximize the similarity between positive samples and minimize the similarity between negative samples. In this way, the model can not only learn the feature representation within the modality, but also effectively enhance the feature alignment ability between different modalities through cross-modal contrastive learning. Contrastive learning helps ensure that visual features, temporal features, and text features remain consistent in the same shared space, thereby improving the cross-modal fusion ability of the model.

[0058] This contrastive learning process is combined with a joint loss function (including classification loss, regression loss, and generation task loss) to jointly optimize the performance of the entire model in a weighted manner. During the model training process, the parameters of all modalities are gradually updated to achieve end-to-end optimization. After training, the model is able to achieve optimal performance in multimodal interaction and can accurately generate diagnostic or educational reasoning results. This cross-modal feature alignment strategy significantly improves the model's multimodal reasoning capabilities and the accuracy of the final decision.

[0059] The TCM multimodal LLM trained based on the above method is particularly suitable for extracting complex text information of TCM, and can efficiently and accurately identify, analyze and extract information from complex types of TCM literature and ancient books. By combing the source text information, a list of relationships such as prescriptions, corresponding symptoms, contained ingredients and ingredient effects can be compiled. The knowledge graph technology can be further introduced to structure the TCM knowledge system and enhance the system's own reasoning ability.

[0060] Traditional Chinese Medicine Multimodal LLM Training System

[0061] On the other hand, this application proposes a multimodal LLM training system for traditional Chinese medicine, referring to Figure 2 The system includes: a multimodal medical data acquisition module 100, a tongue coating visual feature and pulse time series feature extraction module 200, a feature combination module 300, a text feature extraction module 400, a feature fusion module 500 and an integration module 600. Among them,

[0062] Module 100 is used to obtain labeled multimodal medical data of traditional Chinese medicine, wherein the multimodal medical data includes: at least one of tongue coating image, pulse image or and text data, wherein the text data includes: at least one of patient information, main complaint symptoms or diagnosis results.

[0063] Module 200 is used to use the labeled tongue coating images and pulse images to train the visual feature extraction model and the temporal feature extraction model respectively, so as to generate the tongue coating visual features and the pulse temporal features.

[0064] Module 300 is used to train a cross-modal self-attention mechanism model using the tongue coating visual features and pulse timing features to generate joint features.

[0065] Module 400 is used to train a text feature extraction model using text data and generate text features.

[0066] Module 500 is used to input the joint features and text features into a cross-modal attention mechanism model to generate fusion features.

[0067] The 600 module is used to integrate the trained visual feature extraction model, temporal feature extraction model, cross-modal self-attention mechanism model, text feature extraction model and cross-modal attention mechanism model into the large language model to generate a TCM auxiliary diagnosis model.

[0068] In some examples of the present application, the aforementioned modules are connected, and the aforementioned connection includes a physical connection or a network connection.

[0069] Those skilled in the art will appreciate that the features and advantages described above for the multimodal LLM training method for traditional Chinese medicine are also applicable to the above-mentioned system and will not be elaborated here.

[0070] It should be understood that the system embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here. Specifically, Figure 2 The system shown can execute the above-mentioned embodiment of the traditional Chinese medicine multimodal LLM training method, and the operations and / or functions performed by each module in the system correspond to the method embodiment, which will not be repeated here for the sake of brevity.

[0071] The system of the embodiment of the present application is described above from the perspective of the functional module in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in hardware form, can be implemented by instructions in software form, and can also be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software form instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to perform, or a combination of hardware and software modules in the decoding processor to perform. Optionally, the software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory, and completes the steps in the above method embodiment in conjunction with its hardware.

[0072] Application of Traditional Chinese Medicine Multimodal LLM in Knowledge Extraction and Relation Extraction from Multi-source Data

[0073] On the other hand, the present application proposes the application of TCM multimodal LLM in knowledge extraction and relationship extraction of multi-source data, and the TCM multimodal LLM is trained based on the aforementioned TCM multimodal LLM training method.

[0074] The TCM multimodal LLM uses the structured data in the database to perform fine-tuning training on the basis of the pre-trained model, and uses the formatted dictionary form: {"source data": "structured data"} as training data to allow the TCM multimodal LLM to perform supervised fine-tuning training. The fine-tuned TCM multimodal LLM can be used as the TCM multimodal knowledge extraction LLM, which aims to extract knowledge from newly added unformatted data, extract important and key information from text, ECG electrocardiogram data, and tongue coating image data, and classify and plan these data separately, and store them in a unified format as a knowledge database for model training iteration.

[0075] Application of multimodal LLM of traditional Chinese medicine in patient diagnosis

[0076] On the other hand, the present application proposes the application of TCM multimodal LLM in patient diagnosis, and the TCM multimodal LLM is trained based on the aforementioned TCM multimodal LLM training method.

[0077] The TCM multimodal LLM uses structured data in the database and performs fine-tuning training on the basis of the pre-trained model. It uses the formatted dictionary form: {"symptom description + ECG + tongue coating image": "diagnosis result"} as training data to allow the TCM multimodal LLM to perform supervised fine-tuning training. The fine-tuned TCM multimodal LLM can be used as the TCM multimodal QA consultation LLM, which performs the traditional "look, listen, ask, and feel" in a modern way, using the patient's symptom description, electrocardiogram, and tongue coating image data as input, and outputs the patient's diagnosis and treatment results and treatment opinions in the form of natural language output in an end-to-end manner.

[0078] Application of multimodal LLM in TCM teaching tasks

[0079] On the other hand, the present application proposes the application of TCM multimodal LLM in TCM teaching tasks, and the TCM multimodal LLM is obtained through training based on the aforementioned TCM multimodal LLM training method. TCM multimodal LLM uses structured data in the database and performs fine-tuning training on the basis of the pre-trained model, using the formatted dictionary form: {"symptom description + ECG + tongue coating image": "feature description + corresponding diagnosis result"} as training data, so that TCM multimodal LLM can be supervised fine-tuned. The fine-tuned TCM multimodal LLM can be used as a TCM multimodal teaching LLM. By inputting the symptom description, electrocardiogram and tongue coating image of the "simulated patient", the feature analysis of the corresponding data and the diagnosis results of the corresponding features are output in an end-to-end manner, and TCM diagnosis and treatment teaching is carried out in a more detailed description method.

[0080] Computing equipment

[0081] In another aspect, the present application provides a computing device, based on which the aforementioned multimodal LLM training method of traditional Chinese medicine is executed.

[0082] The so-called electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing devices may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are intended to be examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0083] like Figure 3As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 502 or a computer program loaded from a storage unit 508 to a RAM (Random Access Memory) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.

[0084] A number of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0085] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as a multimodal LLM training method for traditional Chinese medicine. For example, in some embodiments, the multimodal LLM training method for traditional Chinese medicine may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute the aforementioned traditional Chinese medicine multimodal LLM training method by any other appropriate means (eg, by means of firmware).

[0086] In the present application, the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, which can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program may be printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or processing in other suitable ways as necessary, and then stored in a computer memory. The various computer-readable storage media described herein may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.

[0087] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0088] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

[0089] In addition, each functional unit in each embodiment of the present invention may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0090] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0091] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0092] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in the field can change, modify, replace and modify the above embodiments within the scope of the present application without departing from the principles and purpose of the present application.

Claims

1. A multimodal LLM training method for traditional Chinese medicine, characterized in that: include: S1, obtaining labeled multimodal medical data of traditional Chinese medicine, wherein the multimodal medical data includes: at least one of a tongue coating image, a pulse image, and text data, wherein the text data includes: at least one of patient information, main complaint symptoms, or diagnosis results; S2, using the labeled tongue coating image and pulse image, respectively train the visual feature extraction model and the temporal feature extraction model to generate the tongue coating visual features and the pulse temporal features; S3, using the tongue coating visual features and pulse time series features, training a cross-modal self-attention mechanism model to generate joint features; S4, using text data to train a text feature extraction model and generate text features; S5, inputting the joint feature and the text feature into a cross-modal attention mechanism model to generate a fusion feature; S6, integrates the trained visual feature extraction model, temporal feature extraction model, cross-modal self-attention mechanism model, text feature extraction model and cross-modal attention mechanism model into the large language model to generate a traditional Chinese medicine auxiliary diagnosis model.

2. The method according to claim 1, characterized in that The visual feature extraction model is selected from VisionTransformer or its variants; Optionally, the time series feature extraction model Timesformer or its variants; Optionally, the text feature extraction model is selected from LLaMA, GPT-4, BERT or their variants.

3. The method according to claim 2, characterized in that The S3 step includes: Embed visual features and temporal features; Calculate self-attention distribution based on visual features and temporal features to capture the global relationship between multi-modalities; The calculation results are optimized through residual connection and multi-layer perceptron to generate joint features.

4. The method according to claim 3, characterized in that The S5 step includes: The point-wise attention mechanism is used to calculate the similarity weights between text features and joint features; The text features and joint features are weighted and superimposed according to the weights to generate fusion features.

5. The method according to claim 4, characterized in that The model is optimized by the following steps: For data inputs of different modalities, the corresponding sub-models are trained independently using modality-specific loss functions; For the fused multimodal features, a joint loss function, including classification loss, regression loss and generation loss, is used to jointly optimize the entire model.

6. The method according to claim 5, characterized in that During the training process, contrastive learning is used to optimize the semantic consistency of cross-modal features, including: The corresponding visual features, temporal features and text features in the multimodal data are taken as positive samples; Use the non-corresponding visual features, temporal features, and text features as negative samples; The feature alignment ability of the model is improved by maximizing the similarity of positive samples and minimizing the similarity of negative samples.

7. Application of traditional Chinese medicine multimodal LLM in knowledge extraction and relationship extraction of multi-source data, wherein the traditional Chinese medicine multimodal LLM is trained based on the method described in any one of claims 1-6.

8. Application of a traditional Chinese medicine multimodal LLM in patient diagnosis, wherein the traditional Chinese medicine multimodal LLM is trained based on the method described in any one of claims 1-6.

9. Application of a multimodal LLM of traditional Chinese medicine in teaching tasks of traditional Chinese medicine, wherein the multimodal LLM of traditional Chinese medicine is obtained by training based on the method described in any one of claims 1 to 6.

10. A computing device, characterized in that: include: Processor and memory; The memory is used to store computer programs; The processor is used to execute the computer program to implement the traditional Chinese medicine multimodal LLM training method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Interrogation method and device for assisting traditional Chinese medicine, equipment and storage medium

    CN119132519A

  • Cancer auxiliary diagnosis and treatment method based on vision-language large model

    CN119170257A

  • Non-cuff blood pressure measuring apparatus based on multi-stage multi-modal learning and method

    WO2023185873A1

Cited By

  • Traditional Chinese medicine rehabilitation diagnosis system based on multi-modal knowledge graph and large language model

    CN120221058A

  • Multi-mode synchronous monitoring method and system based on traditional Chinese medicine pulse condition and microcirculation

    CN120458529A