Traditional Chinese medicine tongue diagnosis and prescription recommendation system based on multi-modal feature fusion
By introducing a multimodal feature fusion method in the traditional Chinese medicine tongue diagnosis technology, combining tongue coating images and patient condition text data, a high-integrated multimodal physique identification model was established, which solved the problems of insufficient physical identification accuracy and lack of prescription grouping and conditioning recommendations in the existing technology, and achieved personalized traditional Chinese medicine tea drink formula grouping recommendations.
Patent Information
- Application Number
- CN202510565510.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing traditional Chinese medicine tongue diagnosis technology has insufficient accuracy in physical fitness identification, and lacks expansion in formula and conditioning recommendations, so it is impossible to effectively provide personalized traditional Chinese medicine tea drink formula recommendations.
A Chinese medicine tongue diagnosis and prescription recommendation system based on the fusion of multimodal features is designed. By analyzing the tongue coating image and patient's condition text data, a highly integrated multimodal physique identification model is established, and a physical condition information is quickly and accurately extracted, physical condition information is speculated, and a patient's TCM disease name and syndrome are provided, and personalized Chinese medicine tea drink formula recommendations are provided.
It realizes the efficient integration of tongue object data and living habit data, improves the accuracy of physical fitness identification and the degree of personalization of the recommendations of the prescription, and can have a more comprehensive understanding of the individual's health status and living habits, thereby providing more accurate Chinese medicine diagnosis and treatment suggestions.
Smart Images

Figure CN120089345A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent medicine, and particularly relates to a traditional Chinese medicine tongue diagnosis and formula recommendation system based on multi-modal feature fusion. Background Art
[0002] Tongue diagnosis is an important method in traditional Chinese medicine diagnostics. By observing the color, shape, moisture level of the tongue, and the changes in the tongue coating, it is used to assist in the diagnosis and differentiation of diseases. As one of the main contents of inspection by observation, tongue diagnosis has the characteristics of being intuitive, simple, and non-invasive, and is an important basis for syndrome differentiation and treatment in traditional Chinese medicine clinical practice. The changes in tongue manifestations are rapid and clear, which is the most sensitive external reaction to the changes in the condition, providing an intuitive diagnostic basis for doctors. Tongue manifestations can relatively objectively reflect the internal conditions of the human body, and are of great significance for the early detection of diseases and the dynamic monitoring of the condition. However, traditional tongue diagnosis relies on doctors' experience and subjective judgment, which is based on doctors' theoretical learning and practical accumulation. Such a diagnostic method lacks the defect of objective quantitative indicators. At present, there are already some mature and feasible solutions in the field of intelligent medicine for the constitution identification method based on tongue diagnosis. For example, image segmentation is performed on the tongue image to obtain the respective segmentation images of the five-zang body surface regions, and a classification network obtained through machine learning is used to classify the segmentation images to obtain the classification results. To improve the accuracy of classification, adding a feature fusion part in model training is a current trend. In the construction of tongue image-related models, the currently visible feature fusion is the feature fusion of multi-view tongue images, that is, the feature extraction is extended from the extraction after tongue surface image segmentation to the extraction of features from tongue images of different views, and a usable model for constitution identification based on tongue images is trained.
[0003] Therefore, due to the characteristics of clear tongue image features and obvious changes with the physical condition, the constitution discrimination model built based on tongue images has a high utilization value in the conditioning of sub-healthy states (that is, judging the constitution characteristics of the person based on tongue image information and giving corresponding health interventions). However, an important part of the health intervention means in traditional Chinese medicine is to use the theory of the nature and flavor of different herbs and their meridians tropism, combined with an individual's constitution and health status, to select suitable herbs for decocting or brewing for drinking, in order to achieve the purpose of conditioning the body and preventing diseases. This requires the model to output a usable conditioning formula after completing constitution discrimination based on tongue image information. However, the existing model construction schemes have not considered this point.
[0004] Currently, the existing model construction and training schemes related to tongue diagnosis have the following defects: First: The accuracy of constitution discrimination is generally average, and the data types used for discrimination are single. For example, in actual operation, the existing tongue diagnosis technical schemes basically only discriminate based on tongue image data, and do not comprehensively analyze other information of the patient such as living habit data. There are a large number of traditional Chinese medicine disease names and syndrome types, and the descriptions of each disease are detailed and diverse, and comprehensive analysis is required to achieve a better constitution discrimination effect.
[0005] Second: There is a lack of expansion in the aspect of formula conditioning recommendation. The current tongue diagnosis technical solution finally aims to identify the constitution of patients, but does not involve subsequent corresponding formula conditioning recommendations. Traditional Chinese medicine theory is complex, with numerous syndromes, and the corresponding conditioning methods are also rich and diverse. Adding formula conditioning recommendations based on constitution identification is an expansion direction with application value.
[0006] Therefore, it is very important to design a multi-modal feature fusion-based traditional Chinese medicine tongue diagnosis and formula recommendation system that can analyze tongue coating image and patient's condition text data, establish a highly integrated multi-modal constitution identification model, and use artificial intelligence and deep learning technologies to quickly and accurately extract constitution information, infer the traditional Chinese medicine disease name and traditional Chinese medicine syndrome of patients, and provide personalized traditional Chinese medicine tea formula suggestions. Summary of the Invention
[0007] The present invention aims to overcome the problems in the prior art that the existing traditional Chinese medicine tongue diagnosis technology has general accuracy in constitution identification and lacks expansion in the aspect of formula conditioning recommendation. A multi-modal feature fusion-based traditional Chinese medicine tongue diagnosis and formula recommendation system is provided, which can analyze tongue coating image and patient's condition text data, establish a highly integrated multi-modal constitution identification model, and use artificial intelligence and deep learning technologies to quickly and accurately extract constitution information, infer the traditional Chinese medicine disease name and traditional Chinese medicine syndrome of patients, and provide personalized traditional Chinese medicine tea formula suggestions.
[0008] To achieve the above invention object, the present invention adopts the following technical solutions: A multi-modal feature fusion-based traditional Chinese medicine tongue diagnosis and formula recommendation system, comprising: A data preprocessing module for preprocessing tongue image data and text data; A feature extraction module for extracting image features and text features from the preprocessed tongue image data and text data; A multi-modal feature fusion module for effectively fusing the extracted image features and text features by using the self-attention mechanism as the core fusion strategy to generate joint features; A multi-task learning module for completing parallel learning of multiple tasks based on the representation of the joint features; the tasks include disease name judgment and diagnosis, syndrome judgment, and formula conditioning recommendation; A model training and optimization module for designing training methods, loss functions, and optimization algorithms to enable the multi-task neural network model to achieve the goals of disease name judgment, syndrome judgment, and formula conditioning recommendation after processing the joint features.
[0009] Preferably, in the data preprocessing module, the process of preprocessing the tongue image data is specifically as follows: S11. Clean and filter the tongue image data, removing unclear, blurred, overly dark or bright images. At the same time, perform data augmentation on the cleaned and filtered tongue image data. Finally, perform normalization and standardization operations on the data-augmented tongue image data. S12. Perform a color space conversion from RGB to HSV on the tongue image data processed in step S11. S13. Perform threshold segmentation on the tongue image data after color space conversion. S14. Perform morphological operations on the tongue image data after threshold segmentation to complete image correction and smoothing, and remove noise or holes. The morphological operations include opening operation and closing operation, etc. S15. Through the DeepLabV3+ network model, perform segmentation on the tongue image data processed by morphological operations and output the tongue body region image.
[0010] Preferably, in the data preprocessing module, the process of preprocessing text data is as follows: S16. Remove the noise in the text data, that is, remove irrelevant characters; unify the text format, unify the handling of upper and lower case, full-width and half-width characters; for long Chinese and English texts, perform word segmentation processing accordingly, that is, split the original continuous text into independent words or sub-word units. S17. Perform standardization processing on the text data, including two steps: removing stop words and lemmatization; the lemmatization is for English texts, that is, restoring English words to their corresponding basic forms.
[0011] S18. Output the standard text, which is a vocabulary list containing all text vocabulary; each vocabulary in the vocabulary list represents a feature in the text.
[0012] Preferably, in the feature extraction module, the process of extracting image features is as follows: S21. Use color histogram to statistically analyze the main color components in the tongue image to reflect the color characteristics of the tongue coating; at the same time, extract the average color value of the tongue body region to quantify the overall color characteristics of the tongue image. S22. Extract the texture features in the tongue image by calculating the gray-level co-occurrence matrix GLCM; at the same time, use the local binary pattern LBP to calculate the texture features of the local area to describe the texture changes in different parts of the tongue image; calculate GLCM and LBP in different pixel windows for multi-scale analysis to extract texture features at different scales. S23. For the extraction of the shape features of the tongue image, use the feature pyramid network FPN for multi-scale feature extraction to capture the shape information of the tongue image. S24. Use the self-attention mechanism to dynamically adjust the weights of color, texture, and shape features; then concatenate the extracted color, texture, and shape features to form a comprehensive feature vector.
[0013] Preferably, in the feature extraction module, the process of extracting text features is as follows: S25. Use the TF-IDF method to capture the importance of each word in the document; the TF-IDF method includes calculating the term frequency TF and the inverse document frequency IDF. The formula for calculating the term frequency TF is as follows: ; where is the number of times the word t appears in the document d; is the total number of words in the document d; The formula for calculating the inverse document frequency IDF is as follows: ; where N is the total number of documents; is the number of documents containing the word t; 1 is used as a smoothing term to avoid a denominator of 0; Combining the calculation results of the term frequency TF and the inverse document frequency IDF, calculate the TF-IDF value of each word, that is, the term frequency feature. The specific formula is as follows: ; where TF(t, d) is the term frequency of the word t in the document d; IDF(T) is the inverse document frequency of the word t; S26. Use the pre-trained Word2Vec model to further process the words; for each word, the Word2Vec model will generate a corresponding word vector; the word vector is a dense vector of a fixed dimension that retains the semantic features of the corresponding word. S27. Concatenate the term frequency feature and the semantic feature extracted in step S25 and step S26 to obtain a text feature vector.
[0014] Preferably, in the multi-modal feature fusion module, the process of effectively fusing the extracted image features and text features specifically includes the following steps: S31. Map the feature representations of the image and text to the same vector space; set the output feature dimension to , then convert the image features to a -dimensional vector through a fully connected layer; at the same time, use the ReLU activation function to ensure the non-linear expression of the features. S32. Similarly, use a fully connected layer to map the text features to the same dimension as the image features ; At the same time, when aligning text features, maintain the dependency relationship of the context; S33. Fuse the image features and text features through the self-attention mechanism. The specific process is as follows: Take the image features and the text features as the inputs of the self-attention mechanism respectively to generate a query matrix Q, a key matrix K, and a value matrix V; The image features are mapped to , , through a fully connected layer; Similarly, map the text features to , , ; Calculate the similarity between the queries and keys of the image features and text features to obtain the self-attention weights. The specific process is as follows: For the self-attention weights of the image features: ; For the self-attention weights of the text features: ; Among them, is the dimension of the key matrix; According to the calculated self-attention weights, perform weighted summation on the value matrices V of the image features and text features to obtain the final fused features; The weighted fused features contain the mutual information between the image features and text features and are used to reflect the dependency relationship between the image features and text features. The specific process is as follows: The calculation formula for weighting the image features is as follows: ; The calculation formula for weighting the text features is as follows: ; Obtain the final fused features by concatenating or weighted averaging the weighted image features and text features : ; Among them, is a trainable weight parameter used to control the relative contributions of the image and text features during the fusion process.
[0015] Preferably, the multi-task learning module adopts an architecture combining shared parameters and task-specific parameters, specifically including a shared feature layer and task-specific sub-networks; The shared feature layer is used to process the fused features generated from the multi-modal feature fusion module and pass the fused features to each task-specific sub-network; The input of the shared feature layer is the fused features ; A shared fully-connected layer is used to reduce the dimension of the generated fused features to ensure that the fused features are applicable to different tasks in the multi-task learning module; Subsequently, batch normalization is performed on the fused features with reduced dimension; Then, the ReLU activation function is used to enhance the non-linear expression ability of the fused features; Finally, shared features are generated .
[0016] Preferably, the task-specific sub-networks include a disease name judgment sub-network, a symptom judgment sub-network, and a prescription conditioning recommendation sub-network; The task of the disease name judgment sub-network is to predict the types of diseases suffered by the patient; The input of the disease name judgment sub-network is the shared features , and the shared features are processed through several layers of fully-connected layers, and the softmax activation function is used to generate the class probability distribution of the diseases ; The generated class probability distribution represents the prediction probability of each disease name; The disease name judgment sub-network selects the class with the highest probability as the final disease name judgment result; The symptom judgment sub-network is used to identify and judge the symptoms shown by the patient; The symptom judgment sub-network generates a multi-label prediction of symptom diagnosis by processing the shared features and the disease name judgment result, representing the probability of each symptom; The symptom judgment sub-network sequentially selects several symptoms with the highest probability from high to low as the final symptom diagnosis result of the patient; The prescription conditioning recommendation sub-network is used to recommend suitable traditional Chinese medicine prescriptions and conditioning methods according to the patient's disease name and symptoms; The input of the prescription conditioning recommendation sub-network is the shared features , the disease name judgment result, and the symptom diagnosis result; After being processed by several layers of fully-connected layers and the Softmax activation function, the prescription conditioning recommendation sub-network generates a prescription recommendation result , and the prescription recommendation result includes the recommended traditional Chinese medicine prescriptions and specific conditioning method suggestions.
[0017] Preferably, in the model training and optimization module, Focal Loss is used as the loss function; The Focal Loss function is defined as follows: ; Among them, is the probability of predicting the correct category; is the coefficient for adjusting the category weights, used to control the imbalance between positive and negative samples; is the adjustment parameter for adjusting the weights of easy and difficult samples. When Focal Loss is equivalent to the cross-entropy loss; Each task-specific sub-network is trained using Focal Loss; the total loss function is set as the weighted sum of the Focal Losses of each task: ; wherein, , , are hyperparameters for balancing the losses of each task, and ; is the loss function of the disease name judgment sub-network; is the loss function of the symptom judgment sub-network; is the loss function of the prescription conditioning recommendation sub-network.
[0018] Preferably, in the model training and optimization module, the Adam optimization algorithm is adopted, and the specific optimization formula is as follows: ; wherein, is the first-order momentum at the t-th iteration, regarded as the exponentially weighted moving average of the gradient, representing the trend of historical gradients; is the momentum estimate value of the previous iteration; is the first-order momentum decay rate, used to control the influence degree of historical gradients on the current gradient; is the gradient value at the current time step t; ; wherein, is the second-order momentum at the t-th iteration, that is, the exponentially weighted moving average of the gradient square, reflecting the variance of gradient changes; is the second-order momentum estimate value of the previous iteration; is the second-order momentum decay rate, used to control the influence of the historical information of the gradient square on the current estimate; is the square of the gradient value at the current time step, used to estimate the variance of the gradient; ; wherein, is the model parameter value at the t-th iteration, is the model parameter value of the previous iteration, is the learning rate, is a constant added for numerical stability; The cosine annealing scheduler is used to dynamically adjust the learning rate; during training, the learning rate gradually decreases according to the following formula, presenting a cosine curve shape: ; where and are the minimum learning rate and the maximum learning rate respectively; is the current training iteration number, is the total number of iterations; In the model training and optimization module, the early stopping strategy and the resume training from breakpoint strategy are adopted for the training method.
[0019] Compared with the prior art, the beneficial effects of the present invention are: (1) By analyzing tongue image data and living habit data, the present invention establishes a highly integrated multi-modal constitution identification model; using artificial intelligence and deep learning technologies, it can quickly and accurately extract constitution information, infer the traditional Chinese medicine disease name and traditional Chinese medicine syndrome of the patient, and provide personalized traditional Chinese medicine tea prescription suggestions; (2) In the model training of the present invention, tongue coating image data and living habit text data are used, covering both physiological and behavioral aspects. After comprehensive processing, the traditional Chinese medicine constitution identification result is obtained based on multi-modal data fusion; the tongue coating image reflects the physiological state of an individual, while the patient's condition text provides detailed information about the condition. The combination of the two can help to more comprehensively understand the individual's health status and living habits; the present invention utilizes the complementarity and comprehensiveness of different modal data, enabling the model to understand the data from different perspectives, which is helpful for more accurate constitution classification and output of corresponding prescription recommendations. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a framework diagram of a traditional Chinese medicine tongue diagnosis and prescription recommendation system based on multi-modal feature fusion according to the present invention; Figure 2 is a flow chart of the preprocessing process of tongue image data in the present invention; Figure 3 is a flow chart of the preprocessing process of text data in the present invention; Figure 4 is a flow chart of the image feature extraction process in the present invention; Figure 5 is a flow chart of the text feature extraction process in the present invention; Figure 6 is a flow chart of the traditional Chinese medicine tongue diagnosis and prescription recommendation system based on multi-modal feature fusion provided by the embodiment of the present invention in actual application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] To more clearly illustrate the embodiments of the present invention, the following will describe the specific implementation manners of the present invention with reference to the accompanying drawings. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, and other implementation manners can also be obtained.
[0022] As Figure 1 shown, the present invention provides a traditional Chinese medicine tongue diagnosis and formula recommendation system based on multi-modal feature fusion, including: A data preprocessing module for preprocessing tongue image data and text data; A feature extraction module for extracting image features and text features from the preprocessed tongue image data and text data; A multi-modal feature fusion module for effectively fusing the extracted image features and text features by using the self-attention mechanism as the core fusion strategy to generate joint features; A multi-task learning module for completing parallel learning of multiple tasks based on the representation of the joint features; the tasks include disease name judgment and diagnosis, syndrome judgment, and formula conditioning recommendation; A model training and optimization module for designing a training method, a loss function, and an optimization algorithm, so that the multi-task neural network model can complete the objectives of disease name judgment, syndrome judgment, and formula conditioning recommendation after processing the joint features.
[0023] Further, the data preprocessing module mainly preprocesses tongue data and text data: Among them, the preprocessing of tongue data, as Figure 2 shown, the specific process is as follows: The preprocessing of tongue data first performs data cleaning and screening, removing low-quality images in the data set, and removing unclear, blurred, too dark or too bright images to ensure the accuracy of subsequent processing. Subsequently, all tongue images are adjusted to a unified size of 256x256 for input into the deep learning model.
[0024] Data augmentation of tongue images is performed by randomly rotating the tongue images, which can avoid overfitting of the model to the tongue direction; at the same time, the images are randomly sheared or scaled to simulate tongue data at different shooting distances and angles; and the color information of the images is enhanced by adjusting brightness, contrast, saturation, etc., to improve the model's recognition ability for different tongue coating colors; in addition, Gaussian noise is added to improve the model's anti-interference ability.
[0025] Finally, normalization and standardization operations are performed. The pixel normalization operation scales the pixel values of the image to the range of 0 to 1 or -1 to 1. This step aims to accelerate the convergence process of the model and improve training stability. Subsequently, a standardization method is adopted, which calculates the mean and standard deviation based on the training set and standardizes the image data accordingly.
[0026] The above steps have performed preliminary processing on the tongue image. To improve the accuracy of subsequent tongue segmentation, further operations on the image are required.
[0027] The currently obtained image is an RGB image. In contrast, the HSV color space separates hue, saturation, and value, making it more intuitive when dealing with colors. Therefore, the RGB image is converted to the HSV color space to extract color and brightness features for distinguishing the tongue from the background. The specific principle is as follows: 1. Normalize the RGB values: Scale the R, G, and B values to between 0 and 1.
[0028] 2. Calculate the maximum and minimum values of the three scales of R, G, and B: ; ; ; 3. Calculate the hue (H): If , then H = 0; If , then ; If , then ; If , then ; 4. Calculate the saturation S: If , S = 0; Otherwise, ; 5. Calculate the value V: V = .
[0029] Subsequently, threshold segmentation is performed on the image data after color space conversion. The role of threshold segmentation is to simplify the image data. By converting the complex image into a binary image, it highlights the regions of interest (such as the tongue body) and suppresses the background. This can assist subsequent image processing steps, such as feature extraction, shape analysis, and classification, improving the processing efficiency and accuracy. The basic principle is as follows: The image is divided into foreground and background by setting a threshold T. For each pixel I(x, y): If I(x, y) > T, then this pixel is classified as the foreground (the tongue body region).
[0030] If I(x, y) ≤ T, then this pixel is classified as the background.
[0031] Given that a series of image processing tasks have been carried out previously, the global threshold (Otsu) method is used here to automatically select the optimal brightness feature threshold by maximizing the between-class variance, and the tongue body region and the background are binarized. The specific formula is: ; where T is the threshold to be determined, used to divide the image into foreground and background, is the between-class variance, used to measure the segmentation effect, and are the number of pixels in the background and foreground respectively, N is the total number of pixels, and are the average gray values of the background and foreground respectively.
[0032] At this time, the tongue image has been converted into a binary image. Inevitably, there are small noises and holes in the image. At this time, morphological operations (such as opening operation, closing operation, etc.) are applied to correct and smooth the image, and remove small noises or holes.
[0033] For images with more small interferences, an opening operation is performed on them. A binary image is eroded with a structuring element (such as a small rectangle or circle) to remove small objects and noises.
[0034] ; An expansion operation is applied to the eroded image to restore the structure of the tongue body region.
[0035] ; For images with missing parts in the tongue body region, a closing operation is used to fill holes and connect broken parts. First, the binary image is expanded to fill small holes.
[0036] ; An erosion operation is applied to the expanded image to remove redundant pixels.
[0037] ; After completing the above operations, an image segmentation operation needs to be performed. Here, the DeepLabV3+ algorithm is used to segment the image. This algorithm can handle tongue image of different scales well and is suitable for tongue image segmentation in complex scenarios. The input image size of this algorithm is H×W×C, where H is the image height, W is the image width, and C is the number of channels. A backbone network (such as ResNet or Xception) is used for feature extraction. This network includes multiple convolutional layers, batch normalization layers (BatchNormalization), and activation functions (usually ReLU). In the feature extraction network, dilated convolutions are used to extract features with a dilation rate r, enhancing the model's ability to capture context information. The formula is: ; where y[i] is the output feature, x[i] is the input feature, and w[k] is the convolutional kernel.
[0038] In addition, the ASPP module of this algorithm extracts features through dilated convolutions with multiple different dilation rates to achieve multi-scale feature fusion. Specifically, it is expressed as: ; where r1, r2, r3, and r4 are different dilation rates.
[0039] DeepLabV3+ also introduces a decoder structure to restore image details. The decoder enhances the details and edge information of the segmentation result by gradually upsampling (i.e., boosting the low-resolution feature map to a high-resolution feature map) and combining the feature map from the encoder. This process makes the segmentation result more refined and accurate. The upsampling operation process can be expressed as: ; ; where is the resolution feature, is the upsampled feature, is the feature from the encoder. represents "fused feature", which is the combination of the upsampled feature and the feature from the encoder .
[0040] Finally, the feature map processed by the decoder passes through a convolutional layer to convert it into an output image of the same size as the input image. The output is the segmentation result after passing through the Softmax activation function, and each pixel point of the output image corresponds to a class probability distribution.
[0041] The segmentation results obtained by the above operations have uneven quality. In this case, morphological operations are used again to improve the segmentation results. For images with many small interferences, open operations are performed to remove noise. For images with missing tongue areas, closed operations are used to fill holes and broken connections.
[0042] After that, the processed segmentation mask is combined with the original image to remove the background and highlight the tongue features. The specific operation is: Combine the segmentation mask with the original image: multiply the processed binary mask with the original image, and retain the tongue area corresponding to the mask. Then, set the background area outside the mask to black (or transparent) to highlight the tongue.
[0043] In addition, text data preprocessing, such as Figure 3 As shown, the specific process is as follows: Hospital medical record documents are used here as the source of text features, but this type of data has many problems, such as containing irrelevant fields such as the social unified credit code, etc. Therefore, it is necessary to clean the text data. The preprocessing of text data in the present invention includes two major steps: text cleaning and data normalization.
[0044] In terms of text cleaning, the present invention first removes the noise in the text data, that is, removes irrelevant characters, such as special symbols and punctuation marks, or standardizes them. Then the text format is unified, and the processing of uppercase and lowercase, full-width and half-width characters is unified. For long texts in Chinese and English, word segmentation is required to split the original continuous text into independent words or sub-word units. This is the basis for subsequent feature extraction. Through word segmentation, the long text is disassembled into an ordered sequence of words, from which the model can learn the dependency relationship and contextual semantics between words. Jieba is used for Chinese text and NLTK is used for English text.
[0045] When the Jieba model is used, a prefix dictionary is introduced. Based on the characteristics of TCM consultation, custom medical professional vocabulary is added to improve the word segmentation effect. The prefix dictionary allows the model to give priority to custom words to avoid them being segmented incorrectly.
[0046] Data normalization refers to the standardization of text data to make it more consistent in format, structure and content, which is convenient for subsequent model training and feature extraction. The normalization of the processed data mainly includes two steps: removing stop words and restoring the word form. Stop words refer to words that appear frequently in the text but do not contribute much semantically to the task itself. The existence of stop words will cause the model to be disturbed by irrelevant information during feature extraction. Therefore, common words in the text that do not contribute to classification are removed, such as common Chinese words such as "的" and "了", or English words such as "the" and "is". By removing these high-frequency, low-value words, the redundancy of the feature space can be reduced and the model's sensitivity to useful information can be enhanced.
[0047] Subsequently, the data after removing the stop words is subjected to morphological restoration. This processing step is for English text, restoring the words to their basic forms (such as restoring "running" to "run"), ensuring that the same words with different morphologies can be normalized into the same feature. This patent uses the morphological restoration method WordNet to simplify the word form.
[0048] For the preprocessed text data, a vocabulary containing all text words is constructed. Each word in the vocabulary represents a feature in the text.
[0049] Furthermore, the feature extraction module is mainly used to extract image features and text features.
[0050] Among them, tongue image feature extraction, such as Figure 4 As shown, the specific process is as follows: Color feature extraction for tongue image data. The present invention uses color histogram to count the main color components (such as red, white, yellow, etc.) in the tongue image to reflect the color characteristics of the tongue coating; in addition, the average color value of the tongue body area is extracted to quantify the overall color characteristics of the tongue image.
[0051] For the texture feature extraction of tongue images, the present invention extracts the texture features (such as roughness, contrast, energy, homogeneity, etc.) in the tongue image by calculating the gray level co-occurrence matrix (GLCM). The local binary pattern (LBP) is used to calculate the texture features of the local area and describe the texture changes of different parts of the tongue image. The GLCM and LBP are calculated in different pixel windows, and multi-scale analysis is performed to extract texture features of different scales.
[0052] For the shape feature extraction of tongue image, the Feature Pyramid Network (FPN) is adopted for multi-scale feature extraction to capture the detailed shape information of the tongue image. Subsequently, the geometric features of the tongue are extracted as shape descriptors, including the area, perimeter, convex hull area of the tongue, and the dimensions of the circumscribed rectangle, which jointly reflect the shape characteristics of the tongue. Further, by extracting and analyzing the contour after tongue image segmentation, the edge shape features of the tongue are calculated, specifically involving parameters such as curvature and convexity. In addition, the Hu moment invariants are also calculated to extract the shape invariant features of the tongue body region, ensuring the stability and accuracy of shape description.
[0053] After extracting the above three types of tongue image features, these features need to be fused to obtain an image feature vector that can be used for subsequent model training. The self-attention mechanism is used before feature concatenation to dynamically adjust the weights of different features. Subsequently, the extracted color, texture, and shape features are concatenated to form a comprehensive feature vector.
[0054] In the design of the model structure, this feature extraction model uses FPN as the basic feature extractor, and uses the FPN network to extract multi-level feature pyramids, especially for the fine-grained information of shape features. GLCM and LBP are used as texture feature extractors, which are combined with the multi-level features of FPN to form a rich texture expression. After passing through the fully connected layer, the extracted tongue image feature vector is generated.
[0055] In addition, for text feature extraction, as Figure 5 shown, the specific process is as follows: Traditional feature extraction methods rely on word frequency statistics, but are easily affected by common words. The present invention uses the TF-IDF method to improve the feature evaluation effect. TF-IDF is an enhanced representation based on the bag-of-words model. By representing the importance of vocabulary with TF-IDF weights, the discriminability of features is enhanced. This method includes calculating the term frequency (TF) and calculating the inverse document frequency (IDF). If a word appears frequently in a document but rarely in other documents, its TF-IDF value will be very high, indicating that it has high discriminability and importance for this document. The specific formula is as follows: Calculating the term frequency (TF): For each vocabulary in each document, calculate its term frequency value, which represents the number of times a certain vocabulary appears in the document. The larger this value is, the more important the word is in the document: ; where, is the number of times the vocabulary t appears in the document d; is the total number of vocabularies in the document d.
[0056] Calculate the Inverse Document Frequency (IDF): For each term in the vocabulary, calculate its inverse document frequency to measure the distribution of the term across all documents. When a term appears in many documents, its IDF value will be small; conversely, if the term appears in only a few documents, the IDF value will be large: ; where N is the total number of documents; is the number of documents containing term t; 1 is used as a smoothing term to avoid a zero denominator.
[0057] Finally, combine the above calculation results of term frequency and inverse document frequency to calculate the TF-IDF value for each term. The TF-IDF value is used to measure the importance of a term in the current document, and its calculation formula is as follows: ; where TF(t, d) is the term frequency of term t in document d; IDF(t) is the inverse document frequency of term t.
[0058] In the above operations of the present invention, the TF-IDF statistical method based on term frequency is used to capture the importance of words in a document, but the order or context of the words is not considered. Therefore, TF-IDF cannot reflect the semantic similarity of words. However, during subsequent model training, the model needs to perform classification tasks. Therefore, it is still necessary to focus on the semantic understanding of the text. To meet this requirement, the present invention uses Word2Vec to make a useful supplement to the text features. This model can map each word to a continuous vector of a fixed dimension while preserving the semantic similarity between words, and is widely used in natural language processing tasks.
[0059] Process the above preprocessed text data using a pre-trained Word2Vec model. For each term, the Word2Vec model will generate a corresponding word vector. The word vector is a dense vector of a fixed dimension that preserves the semantic features of the term. For each term in the text, the corresponding word embedding vector can be obtained by looking it up in the model.
[0060] TF-IDF generates term frequency features, which are suitable for capturing the importance of words in a document. Word2Vec generates semantic features, which capture the context relationship between words. Concatenating the features of the two can combine statistical information and semantic information to provide a more comprehensive feature representation for downstream tasks (such as classification, clustering, etc.).
[0061] When performing this operation, the present invention first concatenates two feature matrices column - by - column to form a larger feature vector. Since the value ranges of TF - IDF and Word2Vec are different, the concatenated features can be normalized, for example, using StandardScaler to scale the features to the same range. If the dimension of the concatenated features is relatively high, PCA is used for dimensionality reduction.
[0062] Furthermore, the multi - modal feature fusion module in the present invention adopts the self - attention mechanism as the core fusion strategy to effectively fuse the features extracted from medical images and diagnostic texts. The self - attention mechanism can dynamically calculate the correlation between image and text features, thereby achieving a more refined and context - related fusion and enhancing the model's comprehensive expression ability for multi - modal features. The specific operation process of the multi - modal feature fusion module is as follows: 1. Feature alignment: Due to the different dimensions and natures of medical image features and text features, it is necessary to align the two types of features. The goal of feature alignment is to map the feature representations of images and texts into the same vector space, thereby providing a basis for subsequent fusion and prediction tasks.
[0063] The features obtained from the image feature extraction module are usually multi - dimensional vectors. To align with text features, a fully - connected layer is used to map the image features into vectors of a fixed dimension. Assuming the output feature dimension is , the image features are converted into vectors of dimensions through the fully - connected layer.
[0064] Meanwhile, the ReLU activation function is used to ensure the non - linear expression of features.
[0065] After the text features obtained from the text feature extraction module are encoded by BERT or other Transformers, they usually already have high semantic expression ability. To be consistent with image features, a fully - connected layer is also used to map the text features to the same dimension as the image features .
[0066] When aligning text features, the context - dependent relationship is maintained to ensure that key medical terms and symptom descriptions in the text can interact effectively with image features through the attention mechanism.
[0067] 2. Feature fusion strategy: The introduction of the self - attention mechanism enables features of different modalities to dynamically capture the correlation between each other during fusion and fuse them with weights according to the requirements of specific tasks. This mechanism automatically selects the features most meaningful for the current task to be strengthened by calculating the cross - correlation between image features and text features.
[0068] The core of the self-attention mechanism is to selectively enhance important features by calculating the similarity between "queries" (Q), "keys" (K), and "values" (V). The specific process is as follows: ; Among them, Attentiom(Q, K, V) represents the calculation function for Q, K, and V. Q is the query matrix, generally derived from image or text features. K is the key matrix, representing the content to be matched in the features, and V is the value matrix, representing the feature information to be focused on. is the dimension of the key matrix to ensure scaling balance.
[0069] In the present invention, the steps of fusing image features and text features using the self-attention mechanism are as follows: Take the image features and text features as the inputs of the self-attention mechanism respectively to generate the query matrix Q, the key matrix K, and the value matrix V; The image features are mapped to , , through a fully connected layer; similarly, the text features are mapped to , , ; By calculating the similarity between the queries and keys of the image features and text features, the self-attention weights are obtained. The specific process is as follows: For the self-attention weights of the image features: ; For the self-attention weights of the text features: ; Among them, is the dimension of the key matrix; According to the calculated self-attention weights, weighted summation is performed on the value matrices V of the image features and text features to obtain the final fused features; the weighted fused features contain the mutual information between the image features and text features, which is used to reflect the dependence relationship between the image features and text features. The specific process is as follows: The calculation formula for weighting the image features is as follows: ; The calculation formula for weighting the text features is as follows: ; The final fused features are obtained by concatenating or weighted averaging the weighted image features and text features : ; Among them, is a trainable weight parameter used to control the relative contributions of image and text features during the fusion process.
[0070] 3. Post-processing of the fused features: After the fused features are generated, further processing is required to ensure their suitability for different tasks in the multi-task learning module. Since the dimension of the fused features is relatively high, the present invention uses a fully connected layer to reduce the dimension of the features, so as to reduce redundant information and ensure the computational efficiency of the model.
[0071] Batch normalization is performed on the dimension-reduced fused features to maintain the stability of the features in different batches; subsequently, the ReLU activation function is used to enhance the non-linear expression ability. The feature representation after fusion and post-processing will be used as the input of the multi-task learning module of the present invention for subsequent tasks such as disease name judgment, symptom judgment, and prescription conditioning recommendation.
[0072] Furthermore, the multi-task learning module of the present invention is used to complete parallel learning of multiple tasks based on the joint feature representation generated by the multi-modal feature fusion module, including disease name judgment and diagnosis, symptom judgment, and prescription conditioning recommendation. The design of the multi-task learning module is based on a shared and task-specific sub-network structure, which can process multiple tasks simultaneously and make full use of the shared information between different tasks, thereby improving the overall performance of the model.
[0073] The multi-task learning module of the present invention includes the following three tasks: disease name judgment task, symptom judgment task, and prescription conditioning recommendation task.
[0074] The disease name judgment task diagnoses and outputs the corresponding disease name based on the input medical images and text data; the symptom judgment task diagnoses and outputs the corresponding symptoms based on the input medical images and text data; the prescription conditioning recommendation task provides personalized traditional Chinese medicine prescription and conditioning method suggestions for patients according to the disease name judgment result and the symptom judgment result.
[0075] The multi-task learning module adopts an architecture that combines shared parameters and task-specific parameters, including a shared feature layer and task-specific sub-networks.
[0076] The shared feature layer is used to process the joint features generated from the multi-modal feature fusion module and pass these features to each task sub-network. The design of this layer is to extract common features useful for all tasks and reduce redundant calculations between tasks.
[0077] The input of this layer is the fused features , which is usually a high-dimensional feature vector and contains the joint information of images and texts. The joint features are further processed through several shared fully connected layers and activation functions (ReLU). Batch Normalization is applied after each fully connected layer to accelerate the training of the model and improve its generalization ability. After processing, a shared feature representation is generated. , which contains the latent information important for multiple tasks.
[0078] The task-specific sub-networks include the disease name judgment sub-network, the symptom judgment sub-network, and the prescription and conditioning recommendation sub-network. The task of the disease name judgment sub-network is to predict the type of disease the patient has; the input of the disease name judgment sub-network is the shared feature. , and the shared feature is processed through several fully connected layers. and the softmax activation function is used to generate the class probability distribution of the disease. The generated class probability distribution. represents the prediction probability of each disease name; the disease name judgment sub-network selects the class with the highest probability as the final disease name judgment result. The symptom judgment sub-network is used to identify and judge the symptoms shown by the patient (for example, the symptom groups involved in syndrome differentiation and treatment in traditional Chinese medicine theory); the symptom judgment sub-network processes the shared feature. and the disease name judgment result to generate a multi-label prediction for symptom diagnosis. represents the probability of each symptom; the symptom judgment sub-network selects several symptoms with the highest probability from high to low as the final symptom diagnosis result of the patient. The prescription and conditioning recommendation sub-network is used to recommend suitable traditional Chinese medicine prescriptions and conditioning methods according to the patient's disease name and symptoms; the input of the prescription and conditioning recommendation sub-network is the shared feature. , the disease name judgment result, and the symptom diagnosis result; after being processed through several fully connected layers and the Softmax activation function, the prescription and conditioning recommendation sub-network generates a prescription recommendation result. The prescription recommendation result. includes the recommended traditional Chinese medicine prescriptions and specific conditioning method suggestions.
[0079] Furthermore, the model training and optimization module of the present invention ensures that the multi-task neural network can achieve the goals of disease name judgment, symptom judgment, and prescription and conditioning recommendation after processing multi-modal medical image and text data by designing appropriate training methods, loss functions, and optimization algorithms. In this module, the Focal Loss is adopted as the main loss function to specifically solve the class imbalance problem and enhance the learning ability of the model for minority classes.
[0080] Define the Focal Loss function as follows: ; where is the probability of the correctly predicted class; is the coefficient for adjusting the class weights, used to control the imbalance between positive and negative samples; is the adjustment parameter for adjusting the weights of easy and hard samples. When , the Focal Loss is equivalent to the cross-entropy loss; Each task-specific sub-network is trained using the Focal Loss; set the total loss function as the weighted sum of the Focal Losses of each task: ; where , , are hyperparameters for balancing the losses of each task and need to be set optimally according to the actual situation (for these three weight parameters, the present invention adopts a normalization constraint, i.e., , to ensure that the losses of each task are on the same scale, so that the loss value of a certain task will not be much larger than that of other tasks, causing the model to tend to minimize this loss); is the loss function of the disease name judgment sub-network; is the loss function of the symptom judgment sub-network; is the loss function of the prescription conditioning recommendation sub-network.
[0081] The optimization algorithm used in model training has a great impact on the stability and efficiency of the training process. The multi-task neural network in the present invention adopts the Adam optimization algorithm. This algorithm is a widely used optimization algorithm that can automatically adjust the learning rate and adapt to different gradient change situations. Its optimization formula is as follows: ; where is the first-order momentum at the t-th iteration, which can be regarded as the exponentially weighted moving average of the gradient and represents the trend of historical gradients. is the momentum estimate value of the previous iteration. is the first-order momentum decay rate (hyperparameter, usually taking the value = 0.9), used to control the influence degree of historical gradients on the current gradient. is the gradient value at the current time step t (the current gradient of the model parameters).
[0082] ; where is the second-order momentum at the t-th iteration, that is, the exponentially weighted moving average of the gradient square, reflecting the variance of gradient changes. is the second-order momentum estimate value for the previous iteration. is the second-order momentum decay rate (a hyperparameter, often taking the value = 0.999), which is used to control the influence of the historical information of the gradient square on the current estimate. is the square of the gradient value at the current time step, which is used to estimate the variance of the gradient.
[0083] ; where, is the model parameter value at the t-th iteration, is the model parameter value of the previous iteration, is the learning rate, is a constant added for numerical stability.
[0084] To improve the efficiency and effect of model training, a cosine annealing scheduler is used to dynamically adjust the learning rate; during the training process, the learning rate gradually decreases according to the following formula, presenting a cosine curve shape: ; where, and are the minimum learning rate and the maximum learning rate respectively; is the current training iteration number, is the total iteration number.
[0085] To prevent overfitting, an early stopping strategy is adopted during the training process. When the performance of the validation set no longer improves, the model training will automatically stop to avoid overfitting of the model on the training set.
[0086] To improve the flexibility and fault tolerance of training, a resume training strategy from breakpoint is adopted. During the training process, the checkpoints of the model are saved regularly, including the model parameters and the optimizer state. If the training is interrupted due to external reasons, the model can continue training from the nearest checkpoint without starting from scratch.
[0087] To further improve the generalization ability of the model, the Dropout method is adopted in the present invention to randomly set the outputs of some neurons in the network to zero, aiming to reduce the over-dependence of the model on individual neurons, thereby enhancing the generalization performance of the model. At the same time, the L2 regularization technique is used to introduce the sum of squares of weights into the loss function as a penalty mechanism for overly large weight values, effectively preventing the phenomenon of overfitting during the training process of the model.
[0088] Based on the technical solution of the present invention, the following case scenario is used to illustrate the implementation process of the present invention in practical applications. The specific application implementation plan is as follows: Such as Figure 6As shown, taking a patient with typical tongue image features and daily living habit data as an example, this paper elaborates on how to implement the process of constitution identification and formula recommendation through this technical solution.
[0089] 1. Data collection Collect the tongue image and living habit questionnaire information of the patient.
[0090] Use a high-definition camera to take pictures of the patient's tongue, ensuring that the obtained images have uniform light distribution, no significant blurring, and the image resolution is set to 512x512 pixels to ensure image clarity and analysis accuracy.
[0091] Collect the patient's conditions through a detailed questionnaire in the mobile device. This questionnaire covers the patient's specific physical conditions, discomfort symptoms, etc. The questionnaire examples are as follows: Q1: Is the defecation cycle long, less than three times a week? Q2: Is defecation unsmooth and laborious? Q3: Do you have a feeling of abdominal distension and pain before defecation? Q4: Is the stool dry and hard? Q5: Do you have a history of traditional Chinese medicine allergy? 2. Data preprocessing 2.1 Image data preprocessing First, perform image cleaning and enhancement processing. This step aims to eliminate unclear and low-quality tongue images, and perform standardization operations on the remaining images to ensure that all image sizes are unified to 256x256 pixels. Subsequently, a series of data enhancement operations, including rotation, shearing, and color enhancement, are performed on the standardized images, aiming to improve the model's recognition ability and robustness for different tongue images by increasing the diversity of image data.
[0092] Next, perform color space conversion and segmentation processing. Convert the tongue image from the RGB color space to the HSV color space to better separate color information. Subsequently, use the Otsu method to automatically determine the segmentation threshold, segment the tongue body area, and obtain a binary segmentation result. On this basis, use morphological processing methods to perform opening and closing operations to further remove the noise in the segmentation result, so as to accurately define the tongue body area and provide high-quality image data for subsequent analysis.
[0093] 2.2 Text data preprocessing Clean and segment the questionnaire information. This step aims to remove irrelevant characters and noise words in the text to improve data quality. Subsequently, a word segmentation algorithm (such as jieba) is used to perform detailed word segmentation on long texts for subsequent analysis. To improve the accuracy of word segmentation, a medical-specific dictionary is specially introduced as an aid to ensure that medical-related terms can be accurately segmented.
[0094] Immediately afterwards, perform data normalization. This step includes removing stop words in the text, which usually do not contribute substantially to the meaning of the text but may affect the effect of subsequent analysis. At the same time, for English words in the text, perform lemmatization to convert them to the basic form (such as removing tense, plural, etc. changes) to ensure that all text features remain consistent in subsequent processing, facilitating accurate identification and analysis by the model.
[0095] 3. Feature Extraction 3.1 Tongue Image Feature Extraction For the collected tongue images. First, statistically analyze the color information, focusing on the distribution of main colors such as red and white. By calculating the color histogram, calculate the composition and distribution of colors in the image to obtain color features.
[0096] Use GLCM to extract various texture features including roughness, contrast, energy, and homogeneity by analyzing the spatial relationship between pixels in the image. And LBP generates binary patterns that can reflect the local texture information of the image by comparing the gray values of each pixel with those of its neighboring pixels. Combine the two methods to obtain texture features.
[0097] Through the contour extraction algorithm, accurately depict the edge contour of the tongue body. On this basis, calculate the shape features such as the area, perimeter, and edge curvature of the tongue body to quantify the geometric structure of the tongue and obtain shape features.
[0098] After obtaining the three types of features, perform weighted fusion on them to obtain comprehensive tongue image features.
[0099] 3.2. Text Feature Extraction Calculate TF-IDF features based on the text data to obtain the weight distribution of important keywords. Use the pre-trained Word2Vec model to convert the text information into word vector representations to obtain the semantic relationships between words. Scale the feature vectors obtained from the two processes to the same range and concatenate them to obtain text features.
[0100] 4. Multimodal Feature Fusion Input the extracted tongue image features and lifestyle text features into the self-attention module. Through the self-attention mechanism, perform weighted fusion on the image and text features, enabling the model to dynamically focus on relevant features and output the fused feature vector.
[0101] 5. Model Classification Input the fused feature vectors into a multi-task neural network model to sequentially complete the following tasks: Disease Name Judgment Task: Identify the traditional Chinese medicine disease name of the patient and generate a probability distribution of disease name predictions.
[0102] Syndrome Judgment Task: Perform syndrome judgment on the input data and output multi-label syndrome results.
[0103] Prescription Conditioning Recommendation: Combine the patient's constitution and disease name, and generate suitable traditional Chinese medicine tea drinks or medicated diet plans through the prescription recommendation network.
[0104] 6. Result Output The model output results include the patient's traditional Chinese medicine constitution type, disease name judgment result, syndrome judgment result, and personalized prescription. The final output format is as follows: Disease Name Judgment: Qi-stagnation Constipation Syndrome Judgment: Spleen Deficiency and Qi Stagnation, Gastrointestinal Stagnation Prescription Recommendation: Zhizhu Tongbianyin Exocarpium Citri Grandis 6g, Fructus Aurantii 6g, Semen Arecae 6g, Radix Linderae 6g, Rhizoma Atractylodis Macrocephalae 10g, Semen Cassiae 6g.
[0105] The present invention aims to solve the defects existing in the digitalization process of traditional Chinese medicine tongue diagnosis technology. At the same time, it builds a bridge between traditional Chinese medicine tongue diagnosis and prescription recommendation, expands the functional characteristics of traditional constitution discrimination and disease diagnosis, and focuses on "preventive treatment of disease". That is, by utilizing the characteristics that the tongue image can quickly feedback the changes of the human body condition and has prominent signs, the tongue image is used as the main data source for constitution discrimination, and then supplemented with the disease description data of the patient collected in the form of questionnaires, etc. The constitution and syndromes of the patient are comprehensively analyzed, an effective feature engineering is established, and an efficient and accurate prediction model is established to accurately predict traditional Chinese medicine disease names, traditional Chinese medicine syndromes, and traditional Chinese medicine conditioning methods.
[0106] From the perspective of improving the efficiency of constitution discrimination, the present invention uses tongue coating image data and lifestyle text data in model training, covering both physiological and behavioral aspects. After comprehensive processing, the traditional Chinese medicine constitution identification result is obtained based on multi-modal data fusion. The tongue coating image reflects the physiological state of an individual, while the patient's disease text provides detailed information about the disease. The combination of the two can comprehensively understand the individual's health status and lifestyle. The present invention utilizes the complementarity and comprehensiveness of different modal data, enabling the model to understand the data from different perspectives, which helps to more accurately classify the constitution and output the corresponding prescription recommendation.
[0107] The above description only elaborates in detail on the preferred embodiments and principles of the present invention. For those of ordinary skill in the art, according to the idea provided by the present invention, there will be changes in the specific implementation manners, and these changes should also be regarded as the protection scope of the present invention.
Claims
1. A TCM tongue diagnosis and prescription recommendation system based on multimodal feature fusion, characterized in that: include: A data preprocessing module, used for preprocessing tongue image data and text data; A feature extraction module is used to extract image features and text features from the pre-processed tongue image data and text data; The multimodal feature fusion module uses the self-attention mechanism as the core fusion strategy to effectively fuse the extracted image features with the text features to generate joint features; A multi-task learning module is used to complete parallel learning of multiple tasks based on the representation of the joint features; the tasks include disease name judgment and diagnosis, symptom judgment and prescription conditioning recommendation; The model training and optimization module is used to design training methods, loss functions and optimization algorithms, so that the multi-task neural network model can complete the goals of disease name judgment, symptom judgment and prescription conditioning recommendation after processing joint features.
2. The TCM tongue diagnosis and prescription recommendation system based on multimodal feature fusion according to claim 1, characterized in that: In the data preprocessing module, the process of preprocessing the tongue image data is as follows: S11, cleaning and screening the tongue image data to remove unclear, blurred, dark or bright images; at the same time, performing data enhancement processing on the cleaned and screened tongue image data; finally, normalizing and standardizing the data-enhanced tongue image data; S12, performing color space conversion processing on the tongue image data processed in step S11 to convert RGB into HSV; S13, performing threshold segmentation on the tongue image data after color space conversion; S14, performing morphological operation processing on the tongue image data after the threshold segmentation processing, completing image correction and smoothing processing, and removing noise or holes; the morphological operation includes opening operation and closing operation, etc.; S15, through the DeepLabV3+ network model, the tongue image data processed by morphological operations is segmented and the tongue area image is output.
3. The TCM tongue diagnosis and prescription recommendation system based on multimodal feature fusion according to claim 2 is characterized in that: In the data preprocessing module, the process of preprocessing text data is as follows: S16, remove noise from text data, that is, remove irrelevant characters; unify text format, unify uppercase and lowercase, full-width and half-width character processing; for long texts in Chinese and English, perform word segmentation processing, that is, split the original continuous text into independent words or sub-word units; S17, standardizing the text data, including two steps: removing stop words and restoring the word form; the word form restoration is for English text, that is, restoring English words to corresponding basic forms; S18, outputting a standard text, wherein the standard text is a vocabulary containing all text words; each word in the vocabulary represents a feature in the text.
4. The TCM tongue diagnosis and prescription recommendation system based on multimodal feature fusion according to claim 3 is characterized in that: In the feature extraction module, the image feature extraction process is as follows: S21, using a color histogram to count the main color components in the tongue image, so as to reflect the color characteristics of the tongue coating; at the same time, extracting the average color value of the tongue body area, so as to quantify the overall color characteristics of the tongue image; S22, extracting the texture features in the tongue image by calculating the gray level co-occurrence matrix GLCM; at the same time, using the local binary pattern LBP to calculate the texture features of the local area, which is used to describe the texture changes of different parts of the tongue image; calculating the GLCM and LBP in different pixel windows, and performing multi-scale analysis to extract texture features of different scales; S23, for shape feature extraction of tongue images, feature pyramid network FPN is used to perform multi-scale feature extraction to capture the shape information of the tongue image; S24, uses the self-attention mechanism to dynamically adjust the weights of color, texture and shape features; then the extracted color, texture and shape features are concatenated to form a comprehensive feature vector.
5. The TCM tongue diagnosis and prescription recommendation system based on multimodal feature fusion according to claim 4 is characterized in that: In the feature extraction module, the text feature extraction process is as follows: S25, using the TF-IDF method to capture the importance of each word in the document; the TF-IDF method includes calculating the word frequency TF and calculating the inverse document frequency IDF; The word frequency TF calculation formula is as follows: ; in, is the number of times word t appears in document d; is the total number of words in document d; The inverse document frequency IDF calculation formula is as follows: ; Where N is the total number of documents; is the number of documents containing word t; 1 is used as a smoothing term to avoid the denominator being 0; Combine the calculation results of term frequency TF and inverse document frequency IDF to calculate the TF-IDF value of each word, that is, the term frequency feature. The specific formula is as follows: ; Where TF(t,d) is the word frequency of word t in document d; IDF(T) is the inverse document frequency of word t; S26, further processing the vocabulary using a pre-trained Word2Vec model; for each vocabulary, the Word2Vec model generates a corresponding word vector; the word vector is a dense vector of a fixed dimension that retains the semantic features of the corresponding vocabulary; S27, concatenating the word frequency features and semantic features extracted in step S25 and step S26 to obtain a text feature vector.
6. The TCM tongue diagnosis and prescription recommendation system based on multimodal feature fusion according to claim 5 is characterized in that: In the multimodal feature fusion module, the process of effectively fusing the extracted image features with the text features specifically includes the following steps: S31, map the feature representations of the image and text to the same vector space; set the output feature dimension to , the image features are converted into Dimensional vector; at the same time, the nonlinear expression of features is guaranteed by using the ReLU activation function; S32, also uses a fully connected layer to map text features to the same dimension as image features ; At the same time, the contextual dependencies are maintained when aligning text features; S33, the image features and text features are integrated through the self-attention mechanism. The specific process is as follows: The image features and text features As the input of the self-attention mechanism, the query matrix Q, key matrix K and value matrix V are generated respectively; Image features The image features are mapped to , , ; Similarly, the text features Mapping , , ; The self-attention weight is obtained by calculating the similarity between the query and the key of the image features and text features. The specific process is: Self-attention weights for image features: ; Self-attention weights for text features: ; in, is the dimension of the key matrix; According to the calculated self-attention weights, the value matrix V of the image features and text features is weighted and summed to obtain the final fusion feature; the weighted fusion feature contains the mutual information between the image features and the text features, which is used to reflect the dependency relationship between the image features and the text features. The specific process is: The calculation formula for image feature weighting is as follows: ; The calculation formula for text feature weighting is as follows: ; The final fusion feature is obtained by concatenating or weighted averaging the weighted image features and text features. : ; in, is a trainable weight parameter used to control the relative contribution of image and text features in the fusion process.
7. The TCM tongue diagnosis and prescription recommendation system based on multimodal feature fusion according to claim 6, characterized in that: The multi-task learning module adopts an architecture that combines shared parameters and task-specific parameters, specifically including a shared feature layer and a task-specific sub-network; The shared feature layer is used to process the fused features generated from the multimodal feature fusion module and pass the fused features to each task-specific sub-network; The input of the shared feature layer is the fused feature ; Use a shared fully connected layer to reduce the dimension of the generated fusion features to ensure that the fusion features are suitable for different tasks in the multi-task learning module; Subsequently, the fusion features after dimension reduction are batch normalized; and the ReLU activation function is used to enhance the nonlinear expression ability of the fusion features; Finally, shared features are generated .
8. The TCM tongue diagnosis and prescription recommendation system based on multimodal feature fusion according to claim 7 is characterized in that: The task-specific subnetwork includes a disease name judgment subnetwork, a symptom judgment subnetwork, and a prescription conditioning recommendation subnetwork; The task of the disease name judgment subnetwork is to predict the type of disease the patient suffers from; the input of the disease name judgment subnetwork is the shared feature , through several layers of fully connected layers to share features Process and use the softmax activation function to generate the disease category probability distribution ; Generated category probability distribution , represents the predicted probability of each disease name; the disease name judgment subnetwork selects the category with the highest probability as the final disease name judgment result; The symptom judgment subnetwork is used to identify and judge the symptoms exhibited by the patient; the symptom judgment subnetwork processes the shared features and disease name judgment results to generate multi-label predictions for symptom diagnosis , represents the probability of each symptom; the symptom judgment subnetwork selects several symptoms with the highest probability from high to low in turn as the final symptom diagnosis result of the patient; The prescription and conditioning recommendation subnetwork is used to recommend appropriate TCM prescriptions and conditioning methods according to the patient's disease name and symptoms; the input of the prescription and conditioning recommendation subnetwork is the shared feature , disease name judgment results and symptom diagnosis results; the prescription conditioning recommendation sub-network is processed by several layers of fully connected layers and Softmax activation function to generate a prescription recommendation result , the formula recommendation result Contains recommended TCM formulas and specific conditioning method suggestions.
9. The TCM tongue diagnosis and prescription recommendation system based on multimodal feature fusion according to claim 8, characterized in that: In the model training and optimization module, Focal Loss is used as the loss function; The Focal Loss function is defined as follows: ; in, is the probability of predicting the correct category; It is the coefficient for adjusting the category weight to control the imbalance of positive and negative samples; is the adjustment parameter for adjusting the weight of difficult and easy samples. When , Focal Loss is equivalent to cross entropy loss; Each task-specific subnetwork is trained using Focal Loss; the total loss function is set to be the weighted sum of the Focal Loss of each task: ; in, , , is a hyperparameter used to balance the loss of each task, and ; is the loss function of the disease name judgment sub-network; is the loss function of the symptom judgment sub-network; Loss function for the prescription conditioning recommendation subnetwork.
10. The TCM tongue diagnosis and prescription recommendation system based on multimodal feature fusion according to claim 9, characterized in that: In the model training and optimization module, the Adam optimization algorithm is used. The specific optimization formula is as follows: ; in, is the first-order momentum of the t-th iteration, which is regarded as the exponentially weighted moving average of the gradient and represents the trend of the historical gradient; is the momentum estimate of the previous iteration; is the first-order momentum decay rate, which is used to control the influence of historical gradient on current gradient; is the gradient value of the current time step t; ; in, is the second-order momentum of the t-th iteration, that is, the exponentially weighted moving average of the square of the gradient, reflecting the variance of the gradient change; is the second-order momentum estimate of the previous iteration; is the second-order momentum decay rate, which is used to control the impact of historical information of the squared gradient on the current estimate; is the square of the gradient value at the current time step, used to estimate the variance of the gradient; ; in, is the model parameter value of the tth iteration, is the model parameter value of the previous iteration, is the learning rate, is a constant added for numerical stability; The cosine annealing scheduler is used to dynamically adjust the learning rate. During the training process, the learning rate gradually decreases according to the following formula, showing a cosine curve shape: ; in, and are the minimum learning rate and the maximum learning rate respectively; is the current training iteration number, is the total number of iterations; In the model training and optimization module, the training method adopts the early stopping strategy and the breakpoint continuation training strategy.
Citation Information
Patent Citations
Tongue diagnosis detection method based on AI semantic segmentation and image recognition
CN117541574A
Traditional Chinese medicine disease classification method and system based on tongue picture and text information fusion
CN117975101A
Dynamic correlation enhancement retrieval generation system and method driven by intelligent knowledge graph
CN118839021A
Multi-modal model traditional Chinese medicine tongue diagnosis method and system
CN118942638A
Visual question and answer method based on multi-modal feature fusion and model thereof
CN119832535A
Cited By
Visual text feature fused image sensitive information automatic covering method
CN120374419A
An automatic masking method for image sensitive information based on visual-text feature fusion
CN120374419B
Self-adaptive partition storage method for multi-modal file
CN120670392A
Multi-modal fusion traditional Chinese medicine physique intelligent evaluation system
CN120853900A
Multi-modal large model-based traditional Chinese medicine tongue diagnosis medicine drink recommendation method
CN120913769A