Traditional Chinese medicine syndrome differentiation aided decision-making method based on cross-modal attention Transform
By integrating multimodal data through cross-modal attention Transformer, a multi-level diagnostic model is constructed, which solves the problems of insufficient data modalities and isolated diagnostic methods in traditional Chinese medicine (TCM) diagnostic methods. This achieves the systematicness and accuracy of TCM diagnostic methods and enhances the ability to capture and understand diagnostic and treatment information.
Patent Information
- Application Number
- CN202511066339.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-21
AI Technical Summary
Existing TCM diagnostic methods suffer from insufficient data modalities, isolated diagnostic methods, and unscientific use of attention, leading to one-sidedness and inaccuracy in diagnosis.
We employ a cross-modal attention Transformer-based approach to integrate multimodal data such as visual, auditory, linguistic, and pulse diagnosis. We acquire data from the four diagnostic methods of traditional Chinese medicine through image processing, sound recognition, and sensor technology, and construct a multi-level diagnostic model. We then use a multi-head intensive self-attention mechanism and a multilayer perceptron for comprehensive diagnosis.
It enriches the content of medical records, improves the ability to capture and understand diagnostic and treatment information, realizes the systematicness and accuracy of TCM syndrome differentiation methods, and reflects the complexity and systematicness of TCM holistic view and syndrome differentiation thinking.
Smart Images

Figure CN120995234A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of traditional Chinese medicine and software technology, and specifically relates to a traditional Chinese medicine syndrome differentiation auxiliary decision-making method based on a cross-modal attention Transformer. BACKGROUND
[0002] There are various methods of traditional Chinese medicine syndrome differentiation, each with different focuses. Common differentiation methods include eight-category differentiation, six-meridian differentiation, and triple-heat differentiation. Eight-category differentiation divides the patient's symptoms into four categories: exterior and interior, cold and heat, deficiency and excess, and yang and yin, to guide clinical treatment. In the eight categories, the other six categories can be summarized as two categories of yin and yang, that is, the exterior, heat, and excess are yang, and the interior, cold, and deficiency are yin. Therefore, yin and yang are the general categories in the eight categories. Six-meridian differentiation is a classification method for differentiating and classifying the pathogenic wind-cold evil in the disease process, and is the differentiation principle of colds. Triple-heat differentiation and Weiqi Yingxue differentiation are mainly applicable to the differentiation of warm diseases. Triple-heat differentiation is divided into upper, middle, and lower heat diseases, including lung, heart, and pericardium diseases, spleen, stomach, and large intestine diseases, and liver, kidney, and bladder diseases, respectively. Triple-heat differentiation focuses on dividing the three major parts of the body invaded by warm diseases.
[0003] For a specific syndrome, it is composed of syndrome architecture and disease site identification. Syndrome architecture is the overall framework, which integrates etiology, disease location, disease nature, disease progression, and other factors to form a complete image of the syndrome. However, the syndrome architecture is not clear enough in explaining the specific location of the disease or the stage of the syndrome, as well as the target of the treatment goal. Therefore, disease site identification will be performed again to provide a positioning basis for the syndrome architecture, helping to clarify the scope of the disease and the treatment target. Through the above hierarchical differentiation mode, the syndrome architecture and disease site identification corresponding to the patient's syndrome are obtained, and more accurate patient syndrome information is obtained.
[0004] As the core of traditional Chinese medicine diagnosis and treatment, traditional Chinese medicine syndrome differentiation relies on four examinations (inspection, auscultation and olfaction, inquiry, and palpation) to comprehensively obtain the patient's disease information. Zhou (Analysis of the lesion characteristics of chloasma faciei by looking at the face with the syndrome differentiation of zang-fu organs as the core) et al. observed the color, shape, and boundary of chloasma patches, combined with local and systemic differentiation of the face, supplemented objective information collected by modern equipment, used zang-fu organ differentiation method for lesion analysis and treatment guidance, and explored the objective and quantitative method of facial inspection for chloasma treatment in traditional Chinese medicine. Shi (Research on lung cancer risk warning model based on tongue feature logistic regression) et al. used TFDA-1 digital tongue diagnosis instrument to capture tongue images, used feature extraction technology to obtain index tongue images. The statistical features and correlations of the tongue were analyzed, and six machine learning algorithms were used to construct lung cancer prediction models based on different data sets.
[0005] Modern TCM research gradually introduces multi-modal data fusion technology in order to improve the accuracy and objectivity of syndrome differentiation. Zhao (Construction of a deep learning multi-modal fusion TCM syndrome element differentiation model for type 2 diabetes) et al. based on deep fully connected neural network, U2-Net and ResNet34 network to construct a symptom differentiation model (S-Model) based on tongue map data and symptom data, a tongue map differentiation model (T-Model), and a multi-modal fusion differentiation model (TS-Model) with both as common input. Yang (Construction of a multi-modal feature fusion syndrome diagnosis model for coronary heart disease based on tongue image objectification) et al. collected patient tongue images, used deep learning algorithms to construct a tongue region boundary detection and segmentation model, established a tongue segmentation model, constructed a multi-modal feature fusion diagnosis model for coronary heart disease in traditional Chinese and Western medicine, and analyzed the distribution of TCM syndrome elements and syndrome element combinations to construct a coronary heart disease syndrome diagnosis model.
[0006] Each method of TCM syndrome differentiation is interdependent. The existing work mainly has the following problems: ① The data modalities used in syndrome differentiation are insufficient. The existing methods may not be comprehensive due to various factors such as time constraints, insufficient patient cooperation, and insufficient diagnosis and treatment conditions. At the same time, in TCM syndrome differentiation, tongue and pulse are two important modalities, but patient tongue and pulse data cannot be fully obtained from medical records. ② The syndrome differentiation methods used are isolated. In existing methods, there is a tendency to use only one syndrome differentiation method, ignoring the comprehensive application of other methods. This isolated syndrome differentiation method may lead to one-sided diagnosis and make it difficult to accurately grasp the overall condition of the patient. ③ In terms of attention use, existing solutions do not consider the differences in attention in specific syndrome differentiation methods, and simply use a single and simple attention mechanism, lacking simulation of the complex diagnosis process of TCM. SUMMARY
[0007] To solve the above problems, the present invention proposes a TCM syndrome differentiation based on cross-modal attention Transformer, which integrates visual, auditory, linguistic and pulse data, aiming to explore a more scientific and systematic TCM syndrome differentiation method. Specifically, visual diagnosis obtains information such as facial color and tongue fur through image processing technology; auditory diagnosis obtains sound information such as breathing and cough through sound recognition technology; interview diagnosis analyzes the patient's complaints and symptom descriptions through natural language processing technology; and palpation diagnosis collects pulse data through sensor technology. Through the fusion of multi-modal data, a multi-level syndrome differentiation model is constructed to generate more reasonable syndrome differentiation results.
[0008] The technical solution of the present invention is:
[0009] A TCM syndrome differentiation auxiliary decision-making method based on cross-modal attention Transformer, comprising the following steps:
[0010] S1. Obtain data from the four diagnostic methods of Traditional Chinese Medicine, including obtaining tongue and facial images through observation, obtaining sound images through auscultation, obtaining medical record text data through inquiry, and obtaining pulse data through palpation; among them, tongue and facial images are image data, and sound and pulse data are waveform data.
[0011] S2. Enhanced medical record text data is obtained based on pulse and tongue data. The specific method is as follows:
[0012] The pulse and tongue data are identified by a neural network to obtain the predicted text categories corresponding to the pulse and tongue data. Then, the text data is concatenated with the medical record text data obtained from the consultation to obtain enhanced medical record text data.
[0013] S3. Perform data preprocessing, including:
[0014] Image data, namely tongue image data and facial image data, are first encoded by OpenFace network and ResNet network respectively, and then fused using bidirectional gating units to obtain representation vectors of tongue image data and facial image data.
[0015] Waveform data, namely acoustic image data and pulse image data, are first encoded by the Wav2Vec2-XLSR network and the Low-level prosodic network, respectively, and then fused using a bidirectional gating unit to obtain the representation vectors of acoustic image data and pulse image data.
[0016] For text data, the enhanced medical record text data obtained based on S2 is used to obtain sentence-level representation vectors, i.e., the representation vectors of text data, using the BERT model;
[0017] This yields five modalities: representation vectors for tongue image data, facial image data, acoustic image data, pulse image data, and text data.
[0018] S4. Multimodal dialectics based on the obtained representation vectors, including two stages:
[0019] The first stage involves preliminary diagnosis, namely, the Eight Principles of Diagnosis, the specific method of which is as follows:
[0020] First, we perform the differentiation of the eight principles of syndrome differentiation: emptiness and fullness, cold and heat, and exterior and interior. For each type of syndrome differentiation, we use one of the five modes of the input vector to generate the key vector and the value vector, and the other four modes to generate the query vector. After passing through the multi-head dense self-attention mechanism, we generate five hidden layer result vectors corresponding to different modes, thus obtaining the hidden layer result vectors of the differentiation of emptiness and fullness, cold and heat, and exterior and interior.
[0021] Then the yin and yang differentiation of eight principles is performed, specifically, the hidden layer result vectors of one of the five modalities in the input vector and the corresponding modalities of the yin and yang differentiation of eight principles are spliced to generate the key vector and the value vector, and the vectors of the other four modalities are used to generate the query vector, after the multi-head dense self-attention mechanism, the final result vector of five corresponding different modalities is generated.
[0022] In the second stage, according to the results of the pre-differentiation, that is, according to the hidden layer result vector of the cold and heat differentiation in the pre-differentiation and the final result vector of the yin and yang differentiation, the warm and cold diseases are differentiated, and it is judged whether it is warm or cold, and according to the judgment result, the re-differentiation is performed, and the specific method of the re-differentiation is:
[0023] If the judgment result is cold, the cold disease differentiation is performed, and the six meridian differentiation is adopted, the potential representation generation of the six meridian differentiation is performed first, the hidden layer result vector of the cold disease differentiation of five modalities is generated, and then the final result generation of the six meridian differentiation is performed, the final result vector of five modalities is generated.
[0024] If the judgment result is warm, the warm disease differentiation is performed, and the triple energizer differentiation is adopted, the potential representation generation of the triple energizer differentiation is performed first, the hidden layer result vector of the warm disease differentiation of five modalities is generated, and then the final result generation of the triple energizer differentiation is performed, the final result vector of five modalities is generated.
[0025] S5, the results of each differentiation module are independently output by a separate classifier, the five modal final result vectors output by the final result generation module are fused, and finally a multilayer perceptron is used for classification to generate the differentiation result of traditional Chinese medicine, and the final differentiation output result is the differentiation result of the eight principle differentiation and the cold and warm disease differentiation result of the patient, and the six meridian differentiation or the triple energizer differentiation result based on the cold and warm disease differentiation result of the patient.
[0026] Further, in S2, the pulse data and tongue data are recognized by a neural network respectively, and the specific method for obtaining the predicted text category corresponding to the pulse data and tongue data is:
[0027] The pulse input feature sequence obtained from the pulse data is defined as The tongue input feature sequence obtained from the tongue data is First, a convolutional neural network is trained to obtain a text translation vector, which is represented as:
[0028] ,
[0029] ,
[0030] Among them, represents the pulse translation vector, is a one-dimensional multi-scale convolution operation, is the weight of pulse convolution at each scale, is the number of pulse multi-scale convolution kernels, represents the tongue translation vector, is the weight of tongue convolution at each scale, is the number of tongue multi-scale convolution kernels;
[0031] Then the pulse prediction and tongue prediction are performed:
[0032] ,
[0033] ,
[0034] wherein, belongs to a set {Y}, which represents all possible real pulses, represents the number of all possible pulse classification categories, is the text encoding vector corresponding to the real pulse, represents the current pulse prediction category, and represents the direct inner product calculation of the vector; belongs to a set {Y}, which represents all possible real pulses, represents the number of all possible pulse classification categories, is the text encoding vector corresponding to the real pulse, represents the current pulse prediction category, and represents the direct inner product calculation of the vector; represents the current first prediction result is true probability; Finally, the enhanced medical record text data is represented as:
[0035]
[0036] ,
[0037] wherein, is the original medical record text data, is the enhanced text generated by the medical record text data enhancement.
[0038] Further, in S3, the specific processing method of the image data is:
[0039] ,
[0040] wherein, represents the initial input data, the superscript of represents the type of input data, and defines represents the tongue data and face data, and respectively represent the vectors encoded using the OpenFace network and the ResNet network, respectively represent the vectors encoded using the OpenFace network and the ResNet network, and ;
[0041] The specific processing method of the waveform data is:
[0042] ,
[0043] wherein, the superscript of represents the type of input data, and defines represents pulse data and sound data, and respectively represent the vectors encoded using the Wav2Vec2-XLSR network and the Low-level prosodic network, respectively represent the vectors encoded using the Wav2Vec2-XLSR network and the Low-level prosodic network, and ;
[0044] The specific processing method of the text data is:
[0045] ,
[0046] wherein, represents the BERT model, represents the enhanced medical text;
[0047] The representation vectors of the four diagnostic data are finally obtained as wherein and are obtained from , and are obtained from .
[0048] Further, in S4, the specific method of the deficiency-excess, cold-heat and exterior-interior differentiation in the pre-differentiation is:
[0049] For deficiency-excess differentiation, the processing method is:
[0050] ,
[0051] wherein, , respectively represent tongue, sound, text, pulse and face; for , there are , the hidden layer result vector representing the virtual-real differentiation, the query vector, the key vector and the value vector, and the concatenation of vectors, the learnable weight matrix. The multi-head dense self-attention embodies the uniqueness of the eight-constitution differentiation; thus, there are five hidden layer result vectors corresponding to different modalities .
[0052] The same method is used for cold-heat differentiation and exterior-interior differentiation, and the hidden layer result vectors of cold-heat differentiation and exterior-interior differentiation are defined as and .
[0053] The method for yin-yang differentiation of eight-constitution differentiation is:
[0054] The hidden layer result vectors of virtual-real differentiation, cold-heat differentiation, and exterior-interior differentiation are concatenated to obtain , which is then processed:
[0055] ,
[0056] wherein, the query vector, the key vector and the value vector, the final result vector of yin-yang differentiation of eight-constitution differentiation; five final result vectors corresponding to different modalities are generated .
[0057] Further, in S4, according to the cold-heat differentiation hidden layer result vector in the pre-differentiation and the final result vector of yin-yang differentiation, the specific method for warm-heat and cold-heat differentiation is:
[0058] The cold-heat differentiation hidden layer result vector and the final result vector of yin-yang differentiation are fused through a bilinear transformation mechanism to perform warm-heat and cold-heat differentiation of the patient:
[0059] ,
[0060] wherein, is the specific result of warm-heat and cold-heat differentiation, indicating whether the patient is a warm-heat syndrome or a cold-heat syndrome, , is the fusion vector of a certain modality in the warm-heat and cold-heat differentiation stage, is the multi-layer perceptron used in the warm-heat and cold-heat differentiation stage;
[0061] For cold-heat differentiation, the specific method of six-meridian differentiation is:
[0062] ,
[0063] wherein, represents the six-meridian syndrome differentiation, represents the query vector, represents the key vector and the value vector, and finally generates the hidden layer result vector of the five modalities , which represents the hidden layer result vector of the latent representation generation stage of the six-meridian syndrome differentiation, represents the multi-head sparse self-attention embodying the uniqueness of the six-meridian syndrome differentiation;
[0064] In the final result generation stage of the six-meridian syndrome differentiation, by taking the representation vector of the four diagnostic data and the hidden layer result vector generated by itself in the last step , the QKV is generated for attention fusion:
[0065] ,
[0066] Finally, the final result vector of the five modalities is generated, which represents the final result vector of the final result generation stage of the six-meridian syndrome differentiation.
[0067] For warm diseases, the specific method of tri-jiao syndrome differentiation is:
[0068] ,
[0069] wherein, represents the tri-jiao syndrome differentiation, represents the query vector, represents the key vector and the value vector, and finally generates the hidden layer result vector of the five modalities , which represents the hidden layer result vector of the latent representation generation stage of the tri-jiao syndrome differentiation, represents the multi-head local self-attention embodying the uniqueness of the tri-jiao syndrome differentiation;
[0070] The final result generation stage of the tri-jiao syndrome differentiation is:
[0071] ,
[0072] Finally, the final result vector of the five modalities is generated, which represents the final result vector of the final result generation stage of the tri-jiao syndrome differentiation.
[0073] Further, in S5, the eight-meridian syndrome differentiation result output is represented as:
[0074] ,
[0075] wherein, is the result of the eight-meridian syndrome differentiation, indicating the specific syndrome of the patient in the eight-meridian syndrome differentiation, This represents a multilayer perceptron, and [,] represents the concatenation of vectors;
[0076] The other diagnostic results are derived in the same way as the results of the Eight Principles of Diagnosis.
[0077] The beneficial effects of this invention are as follows:
[0078] 1) A multimodal text enhancement module is adopted to combine traditional medical record texts with pulse diagnosis and tongue diagnosis data. Through the interaction and supplementation of data, the expression mode and information content of medical records are enriched and optimized, thereby enhancing the system's ability to capture and understand diagnostic and treatment information.
[0079] 2) This paper innovatively proposes a multimodal hierarchical dialectical logic architecture, which hierarchically integrates various TCM diagnostic methods and achieves comprehensive application of different diagnostic methods through the organic connection of multimodal information flow. This architecture ensures the coordination and consistency of various diagnostic methods in different diagnostic and treatment contexts, fully reflecting the complexity and systematic nature of TCM's holistic view and dialectical thinking.
[0080] 3) An independent attention selection mechanism is introduced for each diagnostic module, configuring a dedicated attention module for each diagnostic method, thereby achieving relative independence and flexible adaptation between diagnostic modules. Through this mechanism, each diagnostic module can independently focus on its own relevant information elements, further strengthening the characteristics and independence of different diagnostic methods, and improving the diagnostic accuracy and treatment precision of the system. Attached Figure Description
[0081] Figure 1 This is a flowchart of the method of the present invention.
[0082] Figure 2 This is a model architecture diagram of the present invention.
[0083] Figure 3 This is a schematic diagram of the vector encoding for the dialectical module. Detailed Implementation
[0084] The technical principles and solutions of the present invention will now be described in detail with reference to the accompanying drawings:
[0085] like Figure 1 As shown, the overall process of the model, after acquiring the data from the four diagnostic methods of traditional Chinese medicine, mainly includes steps such as medical record text data enhancement, data preprocessing, and multimodal dialectics, as follows. Figure 2 The diagram shows the specific architecture of the model, as follows:
[0086] Enhancement of medical record text data:
[0087] (1) Text enhancement based on pulse translation
[0088] After several times of convolution training by pulse input data, the text translation vector of pulse is generated:
[0089]
[0090] wherein is the one-dimensional multi-scale convolution operation described above. represents the pulse translation vector. is the input feature sequence of pulse. is the weight of pulse convolution under each scale, is the number of multi-scale pulse convolution kernels.
[0091] The pulse prediction is as follows:
[0092]
[0093] wherein belongs to a set { }, which represents all possible real pulses. represents the number of all possible pulse classification categories. is the text encoding vector corresponding to the real pulse. represents the current th pulse prediction result probability of being true. represents the current pulse prediction category. The pulse prediction category can be fine weak knot generation, pulse almost, fine number, sinking string slip. represents the direct inner product calculation of vector.
[0094] (2) Text enhancement based on tongue translation
[0095] After several times of convolution training by tongue input data, the text translation vector of tongue is generated:
[0096]
[0097] wherein is the two-dimensional multi-scale convolution operation described above. represents the tongue translation vector. is the input feature sequence of tongue. is the weight of tongue convolution under each scale, is the number of multi-scale tongue convolution kernels.
[0098] The following formula is used to predict the category of the translation vector:
[0099]
[0100] wherein belongs to a set { }, which represents all possible real tongue. represents the number of all possible tongue image classification categories. represents the text encoding vector corresponding to the real tongue image. represents the current first tongue image prediction result. represents the current tongue image prediction category. The tongue image prediction category can be white fur, pale tongue, purple dark, or pale and tender tongue, etc. represents the direct inner product calculation of the vector.
[0101] (3) Generating enhanced medical record text data
[0102] Finally, the predicted patient tongue image and pulse image are organized through a template and spliced with the original medical record text data to generate enhanced medical record text data .
[0103] For example, the preliminary prediction of the patient's tongue image and pulse image in the medical record text data enhancement is tongue purple dark and fine weak knot.
[0104]
[0105]
[0106] wherein is the original medical record text, is the enhanced text generated by the medical record text data enhancement.
[0107] Data preprocessing:
[0108] For the input traditional Chinese four diagnostic data, it is divided into text data, image data, and waveform data, and the following method is used for preliminary encoding of the data.
[0109] (1) Image data processing.
[0110] For image data, first encode it using the OpenFace network and the ResNet network, and use the bidirectional gate unit (BI-GRU) for fusion to obtain the final feature vector of the image data:
[0111]
[0112] wherein represents the initial input data, represents the tongue image data and the face image data. and respectively represent the vectors encoded using the OpenFace network and the ResNet network. represents the bidirectional gate unit (BI-GRU) for image fusion.
[0113] (2) Waveform data processing.
[0114] For waveform data, similarly, first encode using Wav2Vec2-XLSR network and Low-level prosodic network, and then fuse using bidirectional gated unit (BI-GRU) to obtain the final representation vector of waveform data:
[0115]
[0116] wherein Input represents the initial input data. represents pulse data and sound data. and respectively represent vectors encoded using Wav2Vec2-XLSR network and Low-level prosodic network. Then, fuse using bidirectional gated unit (BI-GRU). represents bidirectional gated unit (BI-GRU) for graph fusion.
[0117] (3) Text data processing.
[0118] For text data, use BERT model to generate text embedding to enhance the semantic expression of medical records, thereby obtaining sentence-level representation vector:
[0119]
[0120] wherein represents BERT model, represents enhanced medical record text.
[0121] Multimodal syndrome differentiation stage:
[0122] After obtaining the final representation vector of four diagnoses (tongue, sound, text, pulse, face), multimodal syndrome differentiation will be performed, which is divided into three stages: pre-differentiation, warm-heat cold differentiation, and re-differentiation. Among them, the pre-differentiation stage uses eight-class differentiation; in the re-differentiation stage, warm-heat disease uses triple-jiao differentiation, and cold disease uses six-meridian differentiation. Each specific differentiation method will be divided into two sub-steps. For eight-class differentiation, the first step is to perform deficiency-excess, cold-heat, and interior-exterior differentiation to generate a hidden layer result vector, and the second step is to perform yin-yang differentiation to generate a final result vector; for triple-jiao differentiation and six-meridian differentiation, the first step is to perform latent representation generation to generate a hidden layer result vector, and the second step is to perform final result generation to generate a final result vector. The first sub-step result vector will be sent to the second sub-step, together with other related vectors to generate a joint input vector for the second sub-step, capturing the characteristics of each stage of differentiation, providing more sufficient auxiliary information for the result vector generation of the second sub-step, and outputting more accurate differentiation results.
[0123] (1) Pre-differentiation
[0124] This stage will carry out eight-class differentiation. First, respectively, the virtual and real, cold and heat, and inside and outside differentiation. Relative to the traditional Transformer architecture, the self-attention mechanism is modified, and the virtual and real differentiation is taken as an example:
[0125]
[0126] Among them , represent tongue, voice, text, pulse, and face. For , there are , represent the hidden layer result vector of the virtual and real differentiation of eight-class differentiation. represents the query vector. represents the key vector and value vector. [] represents the splicing of the vector. represents the learnable weight matrix. represents the multi-head dense self-attention embodying the uniqueness of eight-class differentiation.
[0127] The above operation uses one of the four diagnostic vectors to generate the key vector and value vector while using the other four vectors to generate the query vector. Thus, after the multi-head dense self-attention mechanism, five kinds of hidden layer result vectors corresponding to different modalities are generated .
[0128] Then, the hidden layer result vectors of the cold and heat differentiation and the inside and outside differentiation of eight-class differentiation are obtained respectively by using the same method as above .
[0129] Then, the yin and yang differentiation of eight-class differentiation is carried out. First, by splicing the hidden layer result vectors of the virtual and real, cold and heat, and inside and outside differentiation, we obtain , and the self-attention mechanism of the traditional Transformer architecture is modified.
[0130]
[0131] represents the query vector. represents the key vector and value vector. represents the final result vector of the yin and yang differentiation of eight-class differentiation.
[0132] The above operation splices one of the four diagnostic vectors and the hidden layer result vectors of the virtual and real, cold and heat, and inside and outside differentiation of eight-class differentiation to generate the key vector and value vector while using the other four modalities to generate the query vector. Thus, after the multi-head dense self-attention mechanism, five kinds of final result vectors corresponding to different modalities are generated Specific syndrome module vector organization and coding details, such as Figure 3 As shown in the figure, eight steels are taken as an example, and the structure is similar for other syndrome modules.
[0133] (2) Warm fever syndrome
[0134] After the end of the pre-diagnosis, the preliminary warm fever judgment of the patient is made through the results of the pre-diagnosis. In the eight syndromes of the pre-diagnosis, there is a step of cold and heat syndrome differentiation. Then in the specific syndrome, there are complex conditions such as deficiency and excess mixed, cold and heat mixed, and internal and external diseases. The patient's disease type cannot be determined by cold and heat syndrome differentiation. Here, the cold and heat syndrome hidden layer vector and the final result vector of yin and yang syndrome differentiation are fused through the bilinear transformation mechanism to make the warm fever syndrome of the patient. Specifically
[0135]
[0136] Among them is the specific result of warm fever syndrome differentiation, indicating whether the patient is a warm fever or a cold fever, , represent tongue, voice, text, pulse and face, respectively. is the corresponding weight matrix. is the fusion vector of a certain modality in the warm fever syndrome differentiation stage, which is used for the classification of the multilayer perceptron afterwards. is the multilayer perceptron used in the warm fever syndrome differentiation stage. Based on the classification result of , we will select the subsequent syndrome differentiation method for more accurate description of the patient's syndrome.
[0137] (3) Re-diagnosis
[0138] After the above warm fever syndrome differentiation, the patient's disease will be divided into two categories: cold fever and warm fever, and different treatment methods will be used. At the same time, in order to reflect the overall guiding role of the pre-diagnosis, i.e. the eight syndromes, in the diagnosis, the input vector of the potential representation generation stage in the re-diagnosis stage will be enhanced by introducing the final result vector of the yin and yang syndrome of the eight syndromes , to enhance the logical connection between the diagnosis systems.
[0139] 1. For cold fever syndrome differentiation, six meridian syndrome differentiation is used. First, the potential representation generation stage of six meridian syndrome differentiation is carried out, and at this stage, the self-attention mechanism of the traditional Transformer architecture is modified:
[0140]
[0141] Among them represents the six meridian syndrome differentiation. represents the query vector. represent the key and value vectors. Finally, the hidden layer result vectors of the five modalities are generated represent the hidden layer result vectors of the potential representation generation stage of the six meridian syndrome differentiation. represent the multi-head sparse self-attention embodying the uniqueness of the six meridian syndrome differentiation. The six meridian syndrome differentiation has sparse correlation, which reduces redundant information interference and strengthens local meridian feature modeling.
[0142] After that, the final result generation stage of the six meridian syndrome differentiation is carried out, and the QKV is generated by combining the representation vectors of the four diagnostic data and the hidden layer result vectors generated in the previous step , attention fusion is performed, and the self-attention mechanism of the traditional Transformer architecture is modified.
[0143]
[0144] wherein represents the six meridian syndrome differentiation. represents the query vector. represent the key and value vectors. Finally, the hidden layer result vectors of the five modalities are generated represent the final result vectors of the final result generation stage of the six meridian syndrome differentiation.
[0145] 2. For warm disease syndrome differentiation, three-energies syndrome differentiation is used. First, the potential representation generation stage is carried out, and the self-attention mechanism of the traditional Transformer architecture is modified:
[0146]
[0147] wherein represents the three-energies syndrome differentiation. represents the query vector. represent the key and value vectors. Finally, the hidden layer result vectors of the five modalities are generated represent the hidden layer result vectors of the potential representation generation stage of the three-energies syndrome differentiation. represent the multi-head local self-attention embodying the uniqueness of the three-energies syndrome differentiation. The three-energies syndrome differentiation involves the transmission of changes in traditional Chinese medicine theory, and there is strong local correlation between each level. Therefore, using local attention can effectively capture the details and changes between these levels.
[0148] After that, the final result generation of the three-energies syndrome differentiation is carried out, and the self-attention mechanism of the traditional Transformer architecture is modified.
[0149]
[0150] wherein represents the three-energies syndrome differentiation. represents the query vector. represents the key vector and the value vector. Finally, the final result vector of the five modalities is generated , which represents the final result vector of the final result generation stage of the triple energizer syndrome.
[0151] The result output is:
[0152] Finally, through a separate classifier, the results of each syndrome module are independently output. By fusing the five modalities final result vectors output by the final result generation module, a multilayer perceptron is used for classification to generate the syndrome result of traditional Chinese medicine. The final syndrome output result is the syndrome result of the eight syndromes and the syndrome result of the six meridians (if it is a cold disease) or the triple energizer syndrome (if it is a warm disease). Take the eight syndromes result output generation as an example. By splicing the five modalities final result generation, a multilayer perceptron is used for classification:
[0153]
[0154] wherein is the specific classification result of the eight syndromes. The eight syndromes classification result is either a yin syndrome or a yang syndrome. represents a multilayer perceptron, and represents the splicing of vectors.
[0155] The syndrome result derivation of the triple energizer syndrome or the six meridian syndrome is similar. For patients with warm disease, the triple energizer syndrome classification result may be wind-heat invading the defensive aspect, damp-heat obstruction in the middle-jiao, etc. For patients with cold disease, the six meridian syndrome classification result may be sun wind, sun blood accumulation, too cold, too hot, etc.
[0156] Finally, the overall loss of the model is as follows:
[0157]
[0158] wherein is the loss of the cold and warm disease differentiation stage:
[0159]
[0160] represents whether the prediction is a warm disease or a cold disease, taking the value of 0 or 1.
[0161] is the total loss of the syndrome module:
[0162]
[0163] represents the number of all syndrome modules, i.e. . represents the number of possible syndrome result categories under a specific syndrome, represents the model parameters, represents the regularization coefficient, represents the authenticity (0 or 1) of the predicted result under the category, represents the total number of all trainable parameters in the model.
[0164] The scheme of the present application adopts a multi-modal text enhancement module to enhance medical record texts using tongue and pulse data. A multi-modal hierarchical syndrome differentiation logic architecture is constructed, various syndrome differentiation methods are used and are hierarchically organized according to TCM theory, and each syndrome differentiation method is divided into two stages. Cross-modal fusion attention is adopted, an independent syndrome differentiation attention mechanism is designed, the differences in TCM theory of various syndrome differentiation schemes are considered, and different attention mechanisms are used to reflect the differences in syndrome differentiation.
Claims
1. A TCM diagnostic auxiliary decision-making method based on cross-modal attention Transformer, characterized in that, Includes the following steps: S1. Obtain data from the four diagnostic methods of Traditional Chinese Medicine, including obtaining tongue and facial images through observation, obtaining sound images through auscultation, obtaining medical record text data through inquiry, and obtaining pulse data through palpation; among them, tongue and facial images are image data, and sound and pulse data are waveform data. S2. Enhanced medical record text data is obtained based on pulse and tongue data. The specific method is as follows: The pulse and tongue data are identified by a neural network to obtain the predicted text categories corresponding to the pulse and tongue data. Then, the text data is concatenated with the medical record text data obtained from the consultation to obtain enhanced medical record text data. S3. Perform data preprocessing, including: Image data, namely tongue image data and facial image data, are first encoded by OpenFace network and ResNet network respectively, and then fused using bidirectional gating units to obtain representation vectors of tongue image data and facial image data. Waveform data, namely acoustic image data and pulse image data, are first encoded by the Wav2Vec2-XLSR network and the Low-level prosodic network, respectively, and then fused using a bidirectional gating unit to obtain the representation vectors of acoustic image data and pulse image data. For text data, the enhanced medical record text data obtained based on S2 is used to obtain sentence-level representation vectors, i.e., the representation vectors of text data, using the BERT model; Thus, five modalities are obtained: representation vectors for tongue image data, facial image data, acoustic image data, pulse image data, and text data. S4. Multimodal dialectics based on the obtained representation vectors, including two stages: The first stage involves preliminary diagnosis, namely, the Eight Principles of Diagnosis, the specific method of which is as follows: First, we perform the differentiation of the eight principles of syndrome differentiation: emptiness and fullness, cold and heat, and exterior and interior. For each type of syndrome differentiation, we use one of the five modes of the input vector to generate the key vector and the value vector, and the other four modes to generate the query vector. After passing through the multi-head dense self-attention mechanism, we generate five hidden layer result vectors corresponding to different modes, thus obtaining the hidden layer result vectors of the differentiation of emptiness and fullness, cold and heat, and exterior and interior. Then, the Yin-Yang differentiation of the Eight Principles of Differentiation is performed. Specifically, the hidden layer result vectors of one of the five modes of the input vector and the corresponding modes of the Eight Principles of Differentiation (Virtual and Real, Cold and Heat, Exterior and Interior) are concatenated to generate key vectors and value vectors. At the same time, the vectors of the other four modes are used to generate query vectors. After passing through the multi-head dense self-attention mechanism, a total of five final result vectors corresponding to different modes are generated. The second stage involves differentiating between febrile and typhoid fever based on the results of the preliminary diagnosis, specifically the hidden layer result vector of the cold-heat differentiation and the final result vector of the yin-yang differentiation. This process determines whether the condition is febrile or typhoid fever. Further diagnosis is then conducted based on the results of this determination. The specific method for this further diagnosis is as follows: If the diagnosis is typhoid fever, the typhoid fever syndrome differentiation is performed using the Six Channels differentiation method. First, the latent representation of the Six Channels differentiation is generated, producing a hidden layer result vector of the typhoid fever syndrome differentiation in five modalities. Then, the final result of the Six Channels differentiation is generated, producing a final result vector of the five modalities. If the diagnosis is febrile disease, the febrile disease syndrome differentiation is performed using the Sanjiao syndrome differentiation. First, the potential representation of the Sanjiao syndrome differentiation is generated, producing a hidden layer result vector of febrile disease syndrome differentiation in five modalities. Then, the final result of the Sanjiao syndrome differentiation is generated, producing a final result vector of five modalities. S5. Using a separate classifier, the results of each diagnostic module are output independently. The five modal final result vectors output by the final result generation module are fused together, and finally, a multilayer perceptron is used for classification to generate the TCM diagnostic results. The final diagnostic output results are the diagnostic results of the Eight Principles, the patient's typhoid fever and febrile disease diagnosis results, and the comprehensive diagnostic results of the Six Channels or Triple Energizer based on the patient's typhoid fever and febrile disease diagnosis results.
2. The TCM syndrome differentiation auxiliary decision-making method based on cross-modal attention Transformer according to claim 1, characterized in that, In S2, the specific method for using a neural network to identify pulse and tongue data and obtain the predicted text categories corresponding to the pulse and tongue data is as follows: Define the pulse input feature sequence obtained from pulse data as follows: The tongue image input feature sequence obtained from the tongue image data is as follows: First, a convolutional neural network is used for training to obtain text translation vectors, which are represented as follows: , , in, Represents the pulse translation vector. This is a one-dimensional multi-scale convolution operation. The weights for pulse convolution at each scale are given. The number of multi-scale convolution kernels for pulse diagnosis. This represents the tongue image translation vector. The weights for tongue image convolution at various scales are given. The number of multi-scale convolution kernels for the tongue image; Then, pulse and tongue diagnosis are performed: , , in, Belongs to a set { This set represents all possible true pulse signs. This indicates the number of all possible pulse diagnosis categories. This is the text encoding vector corresponding to the actual pulse. This indicates the current pulse prediction category, and <.> indicates the inner product calculation between vectors; Belongs to a set { This set represents all possible true tongue images. This indicates the number of all possible tongue image classification categories. This is the text encoding vector corresponding to the actual tongue image. This indicates the current tongue appearance prediction category. Indicates the current number The probability that a prediction is true; The resulting enhanced medical record text data Represented as: , in, The original medical record text data, Enhanced text generated from medical record text data.
3. The TCM syndrome differentiation auxiliary decision-making method based on cross-modal attention Transformer according to claim 2, characterized in that, In S3, the specific processing method for image data is as follows: , in, This represents the initial input data. superscript Indicates the type of input data, defines This represents tongue image data and facial image data. and These represent vectors encoded using the OpenFace network and the ResNet network, respectively. This represents a bidirectional gating unit for image fusion, which fuses the output vectors from two coding networks to ultimately obtain... and ; The specific processing method for waveform data is as follows: , in, superscript Indicates the type of input data, defines This represents pulse data and sound data. and These represent vectors encoded using the Wav2Vec2-XLSR network and the Low-level prosodic network, respectively. This represents a bidirectional gated unit for waveform fusion, which fuses the output vectors from two coding networks to ultimately obtain... and ; The specific processing method for text data is as follows: , in, Represents the BERT model. This indicates an enhanced medical record text; The final representation vectors of the four diagnostic methods data are as follows: ,in and Depend on get, and Depend on get.
4. The TCM syndrome differentiation auxiliary decision-making method based on cross-modal attention Transformer according to claim 3, characterized in that, In S4, the specific methods for differentiating deficiency / excess, cold / heat, and exterior / interior syndromes during the preliminary diagnosis are as follows: The approach to the dialectic of reality and illusion is as follows: , in, , These respectively represent tongue appearance, sound appearance, text appearance, pulse appearance, and facial appearance; for ,have , This represents the hidden layer result vector of the dialectic between reality and illusion. Represents the query vector. Represents the key vector and value vector, and [,] indicates the concatenation of vectors. Represents the learnable weight matrix; This indicates that the multi-head dense self-attention embodies the uniqueness of the Eight Principles of Dialectics; thus generating five hidden layer result vectors corresponding to different modalities. ; Using the same method for cold-heat differentiation and exterior-interior differentiation, the hidden layer result vectors of cold-heat differentiation and exterior-interior differentiation are defined and obtained. and ; The method of Yin-Yang differentiation in the Eight Principles of Differentiation is as follows: By concatenating the hidden layer result vectors of the differentiation of emptiness and fullness, cold and heat, and exterior and interior, we obtain Then proceed with the processing: , in, Represents the query vector. Represents key vectors and value vectors. This represents the final result vector of Yin-Yang differentiation in the Eight Principles of Differentiation; it generates five final result vectors corresponding to different modalities. .
5. A TCM diagnostic auxiliary decision-making method based on cross-modal attention Transformer according to claim 4, characterized in that, In S4, based on the hidden layer result vector of cold-heat differentiation in the pre-diagnosis and the final result vector of yin-yang differentiation, the specific method for determining whether it is a febrile disease or a typhoid fever is as follows: By fusing the hidden layer result vector of cold-heat syndrome differentiation and the final result vector of Yin-Yang syndrome differentiation through a bilinear transformation mechanism, the patient's febrile-heat typhoid fever syndrome can be diagnosed: , in, The specific results for differentiating between febrile and typhoid fever indicate whether the patient has a febrile disease or a typhoid fever. , This represents the fusion vector of a specific mode in the differentiation stage of febrile diseases. It is a multilayer sensor used in the diagnosis stage of febrile typhoid fever; The specific method of differentiation of syndromes using the Six Channels theory for typhoid fever is as follows: , in, This indicates the differentiation of syndromes according to the Six Classics. Represents the query vector. Representing key vectors and value vectors, the final result vectors of the hidden layer for five modalities are generated. It represents the hidden layer result vector of the latent representation generation stage of the Six Classics Dialectics. This indicates that the sparse, multi-headed approach reflects the uniqueness of the Six Classics of Differentiation; In the final stage of generating the results of the Six Channels Differentiation, the representation vectors of the four diagnostic methods are used... and the hidden layer result vector generated in the previous step Generate QKV and perform attention fusion: , Finally, a final result vector of five modalities is generated. , representing the final result vector of the final result generation stage of the Six Classics Differentiation; For febrile diseases, the specific method of differentiation of syndromes using the Triple Energizer is as follows: , in, This indicates the differentiation of syndromes based on the three jiaos (triple burner). Represents the query vector. Representing key vectors and value vectors, the final result vectors of the hidden layer for five modalities are generated. It represents the hidden layer result vector of the latent representation generation stage of the Sanjiao syndrome differentiation. This indicates that the uniqueness of the Triple Burner syndrome differentiation is reflected in the localized self-attention of multiple heads; The final stage of generating the results of the Triple Energizer Differentiation is as follows: , Finally, a final result vector of five modalities is generated. , represents the final result vector in the final result generation stage of the Sanjiao differentiation.
6. A TCM diagnostic auxiliary decision-making method based on cross-modal attention Transformer as described in claim 5, characterized in that, In S5, the output of the Eight Principles of Differentiation is represented as follows: , in, The results of the Eight Principles Differentiation indicate the specific symptoms of the patient according to the Eight Principles Differentiation. This represents a multilayer perceptron, and [,] represents the concatenation of vectors; The other diagnostic results are derived in the same way as the results of the Eight Principles of Diagnosis.
Citation Information
Cited By
Dialectical analysis method and system for traditional Chinese medicine auscultation and diagnosis, electronic equipment and medium
CN122157706A
Traditional Chinese medicine auscultation and diagnosis syndrome differentiation analysis method, system, electronic device and medium
CN122157706B