An intelligent evaluation system fusing tongue appearance, interrogation and environment, and a medium

CN122599053APending Publication Date: 2026-08-18GUANGDONG HOSPITAL OF TRADITIONAL CHINESE MEDICINE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610793483.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

然而对节气要素的利用多为静态的知识推送,缺乏对节气因素对健康状况实时影响的结构化量化建模,无法反映环境变化对个体体质或症状的动态作用

Benefits of technology

本发明提供了一种融合舌象、问诊及环境的智能评估系统、介质,依次设置舌象数据处理、动态问诊、环境信息构建及多源数据融合模块,使舌象、症状信息和节气环境因素能够在同一体系中被逐层关联、深度融合。舌象特征用于提供初步体征基础,动态问诊模块根据该基础补充与之相关的症状信息,环境模块进一步引入节气及相关因素,输出更为完整、连续的多维度状况评估结果,实现三者之间关联分析的健康评估技术方案,提升评估的准确性、针对性与时效性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122599053A_ABST
    Figure CN122599053A_ABST
Patent Text Reader

Abstract

The application discloses a smart evaluation system fusing tongue appearance, inquiry and environment, and comprises a tongue appearance data processing module, a dynamic inquiry module and an environment data processing module.The tongue appearance data processing module is used for acquiring tongue appearance data, segmenting and extracting features of the tongue appearance data, and obtaining tongue appearance features.The dynamic inquiry module is used for dynamically generating an inquiry question sequence according to the tongue appearance features, and extracting inquiry features based on response data of the inquiry question sequence.The environment data processing module is used for constructing environment seasonal feature based on the inquiry features and environment data.A multi-source data fusion module is used for cross-modal deep fusion of the tongue appearance features, the inquiry features and the environment seasonal feature according to the influence of the tongue appearance data on the inquiry question sequence and the influence of the inquiry features on the environment weight adjustment, and for generating a multi-dimensional condition evaluation result.A result output module is used for generating a physical condition evaluation report based on the multi-dimensional condition evaluation result.The application fuses multi-source data of tongue diagnosis, inquiry and environment, realizes health evaluation of correlation analysis among the three, and improves the accuracy, pertinence and timeliness of the evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent traditional Chinese medicine technology, and in particular to an intelligent assessment system and medium that integrates tongue diagnosis, medical history taking, and environmental factors. Background Technology

[0002] The traditional Chinese medicine diagnostic system centers on "inspection, auscultation and olfaction, inquiry, and palpation," with tongue diagnosis and inquiry being crucial means of obtaining physiological and pathological information about the human body. Changes in the color, shape, and coating of the tongue can reflect the imbalance of qi and blood in the body's internal organs; inquiry, through communication with the patient, obtains information such as symptoms, subjective feelings, and lifestyle habits that are not easily obtained through objective testing. Therefore, the combination of tongue diagnosis and inquiry has always been an important basis for evaluating an individual's health status in traditional Chinese medicine clinical practice.

[0003] With the development of digital health management, some intelligent assessment methods based on tongue image analysis or reasoning based on consultation rules have emerged. However, most of them still remain at the level of processing single data types. For example, some technologies only analyze the color or texture of the tongue image under visible light, making it difficult to capture deeper physiological information; other systems can conduct consultations according to preset rules, but the content of the consultation has a low correlation with the previous diagnostic results, often exhibiting fixed procedures and a lack of specificity, making it difficult to effectively simulate the diagnostic approach of traditional Chinese medicine that guides consultation through observation.

[0004] Furthermore, traditional Chinese medicine emphasizes the harmony between humanity and nature, believing that the natural environment, especially the changes in the solar terms, has a significant impact on the human body's condition. However, the utilization of solar term elements is mostly static knowledge dissemination, lacking structured and quantitative modeling of the real-time impact of solar term factors on health, and thus failing to reflect the dynamic effects of environmental changes on individual constitution or symptoms. Summary of the Invention

[0005] Based on the shortcomings of the existing technology, the present invention provides an intelligent assessment system and medium that integrates tongue image, medical history and environment. It integrates multi-source data from tongue diagnosis, medical history and environment, and realizes the correlation analysis between the three to improve the accuracy, pertinence and timeliness of the assessment.

[0006] To address the aforementioned technical problems, the first aspect of this invention discloses an intelligent assessment system integrating tongue imagery, medical history taking, and environmental factors, comprising: The tongue image data processing module acquires tongue image data, segments and extracts features from the tongue image data to obtain tongue image features; The dynamic consultation module dynamically generates a sequence of consultation questions based on the tongue image features, and extracts consultation features based on the response data of the consultation question sequence; The environmental data processing module is used to construct environmental seasonal features based on diagnostic features and environmental data; The multi-source data fusion module performs cross-modal deep fusion of tongue image features, consultation features, and environmental seasonal features based on the influence of tongue image data on the sequence of consultation questions and the adjustment of environmental weights by consultation features, generating multi-dimensional condition assessment results. The results output module generates a physical condition assessment report based on the multi-dimensional condition assessment results.

[0007] In some implementations, the dynamic consultation module adopts a traditional Chinese medicine consultation model based on the MoE architecture. It is trained using a proximal strategy optimization algorithm based on an objective function and a reward function to generate a sequence of consultation questions. The traditional Chinese medicine consultation model introduces traditional Chinese medicine diagnosis and treatment logic constraints, including following the principle of observation before consultation and avoiding invalid questions across syndrome types.

[0008] In some implementations, the multi-source data fusion module includes a cross-modal fusion Transformer network; the multi-source data fusion module maps tongue image features, consultation features, and environmental and seasonal features to a unified latent feature space to obtain image codes, consultation codes, and environmental codes; Construct cross-modal feature sequences that include tongue image encoding, medical history encoding, and environment encoding, and add modality embedding or location embedding for different modalities; The cross-modal feature sequence is input into the cross-modal fusion Transformer network, and multi-head attention calculation is performed. Tongue image encoding is used as query vector, consultation encoding and environment encoding are used as key-value pairs, and environment encoding is used as context bias input, so that bidirectional or multidirectional associations are established between tongue image, consultation and environmental seasonal features. Based on the quality of multi-source data, the contributions of the tongue image modality, the consultation modality, and the environment modality are dynamically weighted to output multi-dimensional status assessment results; the quality of multi-source data includes the completeness of tongue image data, the completeness of consultation responses, and / or the completeness of environment data.

[0009] In some implementations, the multi-source data fusion module further includes an uncertainty quantification unit and an interpretability unit; the uncertainty quantification unit maintains the Dropout layer active state during the inference phase, performs multiple random forward propagations on the same input to obtain a set of prediction results, and calculates the confidence and uncertainty of the cross-modal fusion Transformer network output based on the statistical distribution of the prediction result set. The interpretability unit includes a SHAP interpretation subunit for generating global feature contribution and a LIME interpretation subunit for constructing a local linear approximation model for a single prediction, which interpretably presents the results of multi-dimensional status evaluation of tongue features, consultation features and environmental and seasonal features.

[0010] In some implementations, the environmental data includes solar terms, Five Elements and Six Qi, and meteorological and geographical information; The environmental data processing module adopts a spatiotemporal dynamic heterogeneous graph neural network. The heterogeneous graph includes graph nodes and graph edge sets. The graph nodes include visceral nodes, meridian nodes, emotional nodes, and pathogenic nodes. The graph edge sets include the five elements' generating and restraining relationships, exterior-interior relationships, and tonifying and purging relationships. The spatiotemporal dynamic heterogeneous graph neural network takes the Five Elements and Six Qi and meteorological and geographical information as input to update the graph node attributes and graph edge weights; based on the spatiotemporal dynamic heterogeneous graph neural network, it performs multi-step propagation on the heterogeneous graph to obtain the activation state of each graph node under the current environmental conditions, and uses the activation state as the output to generate environmental solar term features.

[0011] In some embodiments, the tongue image data includes visible light images and thermal imaging images. The tongue image data processing module uses the YOLO target detection model to identify the tongue body and tongue coating in the visible light images. It extracts the color, texture, and morphological features of the visible light images and the temperature distribution and gradient features of the thermal imaging images through a visual state space model. It then combines the tongue body, tongue coating, color, texture, morphological features, and temperature distribution and gradient features to generate tongue image features.

[0012] In some implementations, the objective function of the dynamic consultation module is: LCLIP ( i )=E^ t [min( rt ( i ) A ^ t ,clip( rt ( i ),1 ,1+ ) A ^ t )] in, rt ( i () represents the ratio of the old to the new strategies. A ^ t It is the estimation of the advantage function. Update the constraint coefficients for the strategy; The reward function is:

[0013] Where Δ H ( St This represents the reduction in information entropy resulting from the consultation question. Fuser It is the user's implicit / explicit feedback rating. w 1, w 2 represents the weighting coefficient.

[0014] In some implementations, a federated learning module is also included for training the model locally on the user terminal and performing differential privacy gradient pruning and noise injection processing when uploading multi-source data.

[0015] In some implementations, the physical condition assessment report includes a five-dimensional dynamic conditioning report and an interactive interpretability report; the five-dimensional dynamic conditioning report includes dimensions of diet, daily life, exercise, emotions, and acupoints, and provides pre-adjustment suggestions based on the seasonal changes over a preset time period; The interactive interpretability report is used to demonstrate the impact of different combinations of features on the physical condition assessment report.

[0016] Secondly, a computer storage medium is disclosed, on which a computer program is stored, which, when executed by a processor, implements an intelligent assessment system integrating tongue image, medical history and environment as described in any of the above.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention provides an intelligent assessment system and medium that integrates tongue appearance, medical history, and environmental factors. It sequentially sets up modules for tongue appearance data processing, dynamic medical history taking, environmental information construction, and multi-source data fusion, enabling tongue appearance, symptom information, and seasonal environmental factors to be progressively correlated and deeply integrated within the same system. Tongue appearance features provide a preliminary physical examination basis; the dynamic medical history taking module supplements this basis with related symptom information; and the environmental module further incorporates seasonal factors and related elements, outputting a more complete and continuous multi-dimensional condition assessment result. This health assessment technology solution achieves correlation analysis among the three elements, improving the accuracy, relevance, and timeliness of the assessment. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the structure of an intelligent assessment system that integrates tongue image, medical history and environment provided by the present invention. Detailed Implementation

[0019] To better understand and implement this invention, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0020] The terms “comprising” and “having” and any variations thereof in this invention are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such processes, methods, products or devices.

[0021] This invention discloses an intelligent assessment system that integrates tongue image, medical history, and environmental data. By acquiring, organizing, and processing multiple inputs such as tongue image, medical history, and environmental data, tongue image features, medical history, and environmental seasonal features are formed respectively. Through data fusion, the system achieves correlation analysis and comprehensive judgment of the three types of features, thereby generating targeted physical condition assessment results. This system realizes effective correlation between information from different sources to meet the needs of multi-dimensional intelligent health status assessment.

[0022] like Figure 1 As shown, this system includes a tongue image data processing module, a dynamic consultation module, an environmental data processing module, a multi-source data fusion module, and a result output module. The tongue image data processing module acquires tongue image data, segments and extracts features from it to obtain tongue image features. The dynamic consultation module dynamically generates a sequence of consultation questions based on the tongue image features and extracts consultation features from the response data of the consultation question sequence. The environmental data processing module constructs environmental and seasonal features based on the consultation features and environmental data. The multi-source data fusion module performs cross-modal deep fusion of tongue image features, consultation features, and environmental and seasonal features based on the influence of tongue image data on the consultation question sequence and the adjustment of the environmental weights of the consultation features, generating multi-dimensional condition assessment results. The result output module generates a physical condition assessment report based on the multi-dimensional condition assessment results.

[0023] Specifically, the system can acquire or capture tongue image data via a client device, including visible light images and thermal images. Tongue image acquisition can be performed by the device's built-in or external image acquisition components. Visible light images reflect the color, texture, and morphology of the tongue and its coating, while thermal images reflect the temperature distribution and temperature gradient changes on the tongue surface. By calling the device's integrated or external sensor components through the client, both visible light tongue images (RGB) and long-wave infrared thermal images (8–14 μm) can be acquired simultaneously. To ensure stable image quality, the acquisition interface provides posture guidance, distance prompts, and light detection functions, and activates soft light compensation when lighting is insufficient. Furthermore, a prompt to rest for a period of time before thermal imaging acquisition allows the tongue surface temperature to reach a relatively stable state, facilitating subsequent temperature feature extraction.

[0024] After acquiring the tongue image, the tongue image data processing module performs region recognition on the visible light tongue image. This embodiment uses the YOLO object detection model to segment the tongue image, inputting natural language prompts in addition to predefined categories, enabling the system to identify the tongue body, tongue coating, and any potential abnormal areas. In this embodiment, the YOLO object detection model is the YOLO-World model, the input image size is 640×640, the text prompt embedding dimension is 512, the confidence threshold is set to 0.5, and the non-maximum suppression (NMS) IoU threshold is set to 0.3. By sharing the embedding space matching relationship between visual features and text features, the system can locate and identify tongue features described in natural language, such as "cracks," "edge teeth marks," and "ecchymosis." The visible light image provides structured region input for subsequent feature extraction by obtaining the tongue body region, tongue coating region, and any potential abnormal areas.

[0025] After completing region segmentation, this embodiment extracts deep features from the image based on the Vision-Mamba (ViM) state-space model. This model adopts the form of a state-space model (SSM), and its mathematical representation is as follows: ht=Ah t 1+Bx t y t =Ch t +Dx t Where, x t The input sequence is y. t It is the output sequence, h t This refers to the hidden state. The visual state space model flattens the image into a sequence and processes it using an efficient parallel scanning algorithm, enabling it to capture global contextual information with linear complexity. When processing visible light images, the RGB branch takes the segmented visible light tongue image as input and extracts features such as color changes, texture, and morphology of the tongue body and coating. Color changes include reddish, purplish, and pale tones; textures include dryness / wetness, roughness, and greasy coating features; and morphology includes features such as teeth marks, cracks, size, and shape. The RGB branch uses the ViM-tiny version, with a hidden layer dimension of 768, a state space dimension of 16, a dilation factor of 2, an attention window size of 8×8, and the activation function SiLU.

[0026] Simultaneously, the visual state space model processes thermal imaging images through a thermal imaging branch. This branch takes a temperature matrix as input and extracts the overall temperature distribution of the tongue surface, areas with excessively high or low temperatures, temperature gradient changes, and structured patterns of temperature anomalies. The two branches are fused at the rear of the network through a cross-attention mechanism, making surface information such as color and texture complementary to temperature field information. Finally, a multi-layer perceptron (MLP) outputs a multi-dimensional tongue image feature to characterize the overall state of the tongue. This tongue image feature includes overall tongue features, tongue coating characteristics, thermal imaging features, and structural features of abnormal areas. The resulting multi-dimensional tongue image feature corresponds to tongue physiological and pathological indicators, facilitating subsequent comprehensive judgment based on medical history and environmental models.

[0027] The YOLO-World model and the ViM model are jointly fine-tuned on the dataset, with the loss function being a weighted sum of the segmentation loss (DiceLoss + Focal Loss) and the feature extraction loss (Cross-Entropy Loss).

[0028] After obtaining tongue image features, the dynamic consultation module dynamically generates a sequence of consultation questions based on the clues provided by these features. During the consultation process, the system can infer the possible syndrome types based on tongue image clues, ensuring targeted consultation, reducing redundant questions, and improving consultation efficiency. User answers to the consultation questions are structured to form consultation feature vectors. The dynamic consultation module incorporates a Traditional Chinese Medicine (TCM) consultation model based on the Mixture-of-Experts (MoE) architecture. This model is based on TCM knowledge and uses extensive classical TCM literature, real medical cases, and modern research data to construct training corpora. The model possesses capabilities in areas such as detailed symptom inquiry, lifestyle habit inquiry, emotional state analysis, and review of past medical history. When generating consultation questions, the system automatically selects the most suitable expert network to participate in reasoning based on tongue image features and user answers, improving the rationality and professionalism of the consultation question sequence.

[0029] To achieve continuous interaction with users and dynamically adjust the consultation direction, this embodiment employs deep reinforcement learning technology. In this embodiment, the proximal policy optimization (PPO) algorithm is used for deep learning to update the consultation strategy while maintaining training stability. The core optimization objective of PPO can be represented as a pruning objective function, which is: LCLIP ( i )=E^ t [min( rt ( i ) A ^ t ,clip( rt ( i ),1 ,1+ ) A ^ t )] in, It is the ratio of the old and new strategies. A ^ t It is the estimation of the advantage function. This sets the limit coefficient for policy updates. This objective function helps avoid instability caused by excessively large policy updates.

[0030] To improve the quality of consultations, the questions generated in each round and the user feedback are used to construct a reward signal. The reward function is:

[0031] Where, Δ H ( St The amount of information entropy reduction brought about by this round of question and answer is used to measure the effectiveness of the question; Fuser It is the user's implicit / explicit feedback rating. w 1, w 2 represents the weighting coefficient. Through the reward function, the TCM consultation model tends to ask questions with higher information gain and better user experience.

[0032] To optimize the naturalness of the consultation content and the doctor-patient communication experience, the dynamic consultation module further employs the Direct Preference Optimization (DPO) method based on human preferences to train the consultation model. This method directly optimizes the model based on the preference data of TCM doctors for question-answer pairs, resulting in questions that are more in line with the style of TCM clinical consultation and are more organized and empathetic.

[0033] Furthermore, to ensure that the question sequence generation process conforms to the TCM diagnostic and treatment system, this embodiment incorporates diagnostic and treatment logic constraints into the question strategy generation process. These constraints include: following the observation-questioning sequence; determining the initial direction based on tongue characteristics before organizing questions around that direction to reduce irrelevant questions; and avoiding ineffective follow-up questions across syndrome types. If the tongue appearance shows characteristics such as coldness, deficiency, or dampness, the questions will revolve around that syndrome type, avoiding questions about irrelevant syndromes and improving consultation efficiency. With these constraints added, the consultation path more closely aligns with TCM clinical diagnostic and treatment habits, enhancing the effectiveness of information collection.

[0034] After the user completes their response, the system will process the user's response data in a structured manner, transforming it into diagnostic features that include dimensions such as symptom intensity, duration, triggering factors, lifestyle habits, and emotional state.

[0035] The environmental data includes solar terms, Five Elements and Six Qi, and meteorological and geographical information. The environmental data processing module employs a spatiotemporal dynamic heterogeneous graph neural network, where graph nodes include organ nodes, meridian nodes, emotional nodes, and pathogenic factor nodes. The spatiotemporal dynamic heterogeneous graph neural network takes the Five Elements and Six Qi and meteorological and geographical information as input and outputs the activation state of the graph nodes under the influence of the current environmental data to generate environmental solar term features. Specifically, the solar term information represents the current twenty-four solar terms and their corresponding climate characteristics; the Five Elements and Six Qi data includes elements related to natural changes such as excess, deficiency, primary Qi, and secondary Qi; and the meteorological and geographical information includes real-time meteorological parameters such as temperature, humidity, and air pressure in the user's location.

[0036] To reflect the systematic connections between the internal organs, meridians, emotions, and pathogenic factors in Traditional Chinese Medicine (TCM) theory, this embodiment constructs a heterogeneous graph G=(V,E), where the graph nodes V include nodes representing the internal organs, meridians, emotions, and pathogenic factors. The internal organ nodes include the heart, liver, spleen, lungs, and kidneys; the meridian nodes include the twelve meridians; the emotion nodes include joy, anger, pensiveness, worry, and fear; and the pathogenic factor nodes include wind, cold, summer heat, dampness, dryness, and fire. The graph edge set E includes relationships such as the five elements (Wu Xing) generating and controlling relationships, exterior-interior relationships, and tonification / reduction relationships. The five elements generating and controlling relationships include the liver (wood) generating the heart (fire) and the kidney (water) controlling the heart (fire); the exterior-interior relationships are exemplified by the lungs and large intestine being exterior-interior pairs. Both graph edges and nodes can carry attributes describing their strength, weight, and dynamic changes.

[0037] The aforementioned environmental data is encoded into a time-varying vector, which is the environmental input vector. f env(t) Node attributes and edge weights will change at each time step t according to f env(t) Update the system to enable it to simulate the effects of the natural environment on the human body. For example, the organ node v z Attribute update hvz ( t =MLP([ hvzstatic , fenv ( t )]).in, hvzstatic The static attributes of the viscera nodes are represented by MLP, which stands for Multilayer Perceptron Network. This update method enables responses to changes in the solar terms and the imbalances of the Five Elements and Six Qi.

[0038] The spatiotemporal dynamic heterogeneous graph neural network takes the Five Elements and Six Qi and meteorological and geographical information as input to update the graph node attributes and graph edge weights. Based on the spatiotemporal dynamic heterogeneous graph neural network, it performs multi-step propagation on the heterogeneous graph to obtain the activation state of each graph node under the current environmental conditions. The activation state is used as the output to generate environmental solar term features. Specifically, this embodiment realizes spatiotemporal modeling by combining a spatiotemporal dynamic heterogeneous graph neural network, that is, a graph neural network (GNN) with a time recurrent unit (GRU). Information propagation is carried out through a multi-layer graph convolutional network to aggregate neighbor node information. In the spatial dimension, at each time step, graph convolution is performed on the node state; in the temporal dimension, a gated recurrent unit (GRU) is used to capture the evolution trend of the node state, and finally obtains:

[0039] Simultaneously, it captures the dynamic changes of the environment over time and the spatial dependencies between nodes. After multiple layers of spatiotemporal propagation, the activation states of various nodes at the current time will tend to stabilize. In this embodiment, the final node embeddings are combined according to certain rules to form an environmental influence feature vector. Finally, the spatiotemporal dynamic heterogeneous graph neural network outputs the stable state embeddings of all nodes at the current moment, which are then spliced ​​together to form environmental seasonal features.

[0040] The processing of environmental data quantifies traditional Chinese medicine theories such as solar terms and the Five Elements and Six Qi into computable dynamic features. Through the dynamic updating of graph nodes and edges, the system reflects the comprehensive impact of the environment on the internal organs, meridians, emotions, and pathogenic factors. The resulting environmental solar term features are highly correlated with the human body's condition, thus improving the overall evaluation performance of the system.

[0041] After obtaining tongue appearance features, consultation features, and environmental / seasonal features, the multi-source data fusion module performs cross-modal deep fusion of these features based on the influence of tongue appearance data on the consultation question sequence and the adjustment of consultation features on environmental weights, generating multi-dimensional condition assessment results. Specifically, the three types of features are projected onto the same latent feature space through linear mapping or a feedforward network to obtain image encoding, consultation encoding, and environmental encoding. Modality embedding and position embedding are added to each encoding to construct a cross-modal feature sequence, preserving modality differences and sequence structure information. Each encoding contains a composite feature composed of a content vector and a modality embedding, used to distinguish different information sources.

[0042] The cross-modal feature sequence is input into a cross-modal fusion Transformer network to perform multi-head attention computation, establishing bidirectional or multi-directional associations between tongue image, medical history, and environmental / seasonal features. For any query vector Q, key vector K, and value vector V, the multi-head attention computation is as follows:

[0043] To ensure that the consultation questions revolve around tongue abnormalities, this embodiment uses the tongue image code as the query vector and the consultation code and environment code as key-value pairs, establishing the following association through an attention mechanism:

[0044] By observing the abnormal features of the tongue, we can proactively focus on the symptoms most closely related to it during consultation. The degree to which consultation features are affected by environmental features can be dynamically reflected by attention scores, thus establishing a two-way or multi-way correlation between tongue appearance, consultation, and environmental and seasonal features.

[0045] By using environmental encoding as contextual bias input to a cross-modal fusion Transformer network, the changes in different solar terms and the Five Elements and Six Qi can influence the attention distribution during tongue appearance and patient consultation. For example, during solar terms with predominant damp-heat, attention to the yellow, greasy tongue coating—damp-heat symptom link will increase; while in environments with predominantly cold and cool conditions, confidence in symptoms such as thirst and irritability will decrease. Through this mechanism, environmental data can control the weight relationships between different modalities within the model in a structured manner.

[0046] To enhance the robustness of the fusion, this embodiment dynamically weights the contributions of the tongue image modality, medical history modality, and environmental modality based on the quality of the multi-source data. The quality of the multi-source data includes the completeness of the tongue image data, the completeness of the medical history responses, and / or the completeness of the environmental data. The sum of the confidence scores of the three modalities is 1, and the final fusion feature is obtained by recalculating based on the confidence scores. Through confidence score calculation, adaptive features are provided in the event of missing or insufficient data, improving the overall evaluation stability.

[0047] To enhance the reliability of the evaluation results, this application includes an uncertainty quantification (MC-Dropout) unit in the multi-source data fusion module. This unit calculates the stability and cognitive uncertainty of the prediction results during the model inference phase. By maintaining the random deactivation mechanism of the Dropout layer during inference, the model can generate a series of different prediction values ​​under the same input, thereby forming a prediction distribution and estimating model uncertainty. During inference, Dropout is kept active, and model parameters are randomly sampled. The set of prediction results obtained through multiple samplings can be used to analyze the volatility of the model output.

[0048] For the input feature vector, i.e., the fused features, T forward propagations are performed with Dropout enabled to obtain a prediction set. The final prediction value is the mean of all results, and the prediction uncertainty, i.e., the variance, is calculated. The larger the variance, the lower the confidence of the multi-source data fusion module in the prediction; when the uncertainty exceeds a preset threshold, the user can be prompted to supplement the consultation information or re-acquire the tongue image. The uncertainty information can be written into the multi-dimensional condition assessment results along with the evaluation results, providing transparency for users or physicians.

[0049] To enhance the interpretability of the evaluation results, this embodiment introduces an interpretability unit after the multimodal fusion model. The interpretability unit includes a SHAP interpretation subunit and a LIME interpretation subunit, calculating the overall contribution of features to the final evaluation based on the SHAP method. By mapping TCM tongue appearance features, consultation symptom features, and environmental and seasonal factors to an additive feature contribution space, the system can obtain the role of different features in the overall prediction. In this way, users can understand the comprehensive influence of multidimensional factors such as color, texture, tongue coating thickness, symptom severity, and seasonal predominance on the current judgment. To interpret a specific prediction, this embodiment constructs a local linear approximation based on the LIME model, fitting the local decision boundary near the current sample by slightly perturbing the input features. Through the LIME model, users can clearly see which features drive a particular constitution result or syndrome judgment, and how the evaluation result will change if certain features change. Since tongue appearance, consultation, and environmental and seasonal factors are all trimodal inputs, the strength of their correlations can be simultaneously displayed during the interpretation process, making the model's reasoning logic more transparent.

[0050] Traditional deep learning models are often limited to fixed, single-output methods. This embodiment, however, introduces a random sampling mechanism during the inference phase, enabling the model to express the robustness of predictions in a high-dimensional space after feature fusion, and to determine whether the current assessment is affected by modality loss, image quality degradation, or incomplete consultation information. Traditional SHAP and LIME methods are mostly used for interpreting single-modal image or language models. This embodiment, however, demonstrates the correlation paths between tongue visual features, consultation semantic features, and environmental temporal features, allowing the model's decisions to be mapped to the diagnostic logic of "observation-inquiry-seasonal environment" in Traditional Chinese Medicine. In this way, a cross-modal perspective is used to explain the correlation between tongue abnormalities and specific symptoms, and how environmental imbalances can shift the weight of syndrome differentiation, facilitating quick understanding of multi-dimensional condition assessment results for both doctors and users.

[0051] After obtaining the dimensional condition assessment results, the result output module generates a physical condition assessment report based on the multi-dimensional condition assessment results. This embodiment's physical condition assessment report provides a judgment of the current constitution or syndrome, and, combined with seasonal changes, living scenarios, and individual characteristics, generates dynamically adjustable health management suggestions for users, meeting the needs of individualized and dynamic conditioning in traditional Chinese medicine assessment. The physical condition assessment report includes a five-dimensional integrated dynamic conditioning report and an interactive interpretable report; the five-dimensional integrated dynamic conditioning report includes dimensions of diet, daily life, exercise, emotions, and acupoints, and provides pre-adjustment suggestions based on seasonal changes over a preset time period. For example, in a season with predominantly damp heat, dietary suggestions emphasize light and refreshing foods; in the exercise dimension, users are encouraged to reduce excessive energy-consuming exercises; and in the acupoint dimension, guidance on acupoints that help regulate and clear heat is added. Seasonal trends not only participate in the assessment process but also in the pre-adjustment of the conditioning plan.

[0052] The interactive interpretability report is used to demonstrate the impact of different feature combinations on the physical condition assessment results. Based on the contribution analysis of SHAP and LIME interpretations to cross-modal features, the system displays the strength of the influence of various tongue features, consultation symptoms, and environmental factors on the results in the current assessment. Users can interactively adjust some features on the interface, such as sliding features like "fatigue level," "mouth stickiness," and "sleep quality" to different levels. The assessment output will be recalculated in real time in the background, and the changing trend of the predicted results will be displayed visually. This interactive approach helps users understand the relationships between different factors, such as how the thickness of the tongue coating and seasonal humidity jointly affect the judgment of a certain syndrome, or how changes in certain symptoms can lead to a shift in the overall physical condition assessment.

[0053] During deployment, this system employs a cross-device federated learning model to ensure continuous model updates and protect user privacy. A federated learning module is configured to train the model locally on the user's terminal and perform differential privacy gradient pruning and noise injection when uploading multi-source data. A central server maintains global model parameters and coordinates the training process across clients; client applications distributed across user mobile terminals perform several rounds of training locally using user-collected data to calculate model gradients or parameter updates. In this way, the system can continuously absorb feature information from multi-source samples without directly uploading raw user data, maintaining the model's generalization performance.

[0054] To further reduce the risk of inferring user privacy information from gradients, this embodiment sets up a differential privacy mechanism before the client uploads parameter updates, performing two-step processing on the gradient. First, norm clipping is performed on each local gradient to limit the influence of a single sample on the update result; then, Gaussian noise is superimposed on the clipped gradient, so that the contribution of any single user cannot be effectively identified during the global aggregation process.

[0055] Furthermore, this embodiment employs a microservices approach to construct the overall system architecture. Different functional modules are broken down into independent service units, such as tongue image processing, diagnosis generation, environmental feature processing, and fusion assessment. Each service is encapsulated in containers and automated deployment, scheduling, and scaling management are achieved through Kubernetes, maintaining stability and scalability under high concurrency and varying business loads. In addition, the distributed structure of microservices facilitates modular updates and maintenance, allowing for flexible iteration without affecting the overall service.

[0056] In one implementation, the system can be deployed in a cloud-edge-device manner to support model inference, data processing, and federated learning tasks. On the device side, the system can run on a smartphone or health all-in-one device with the application installed. The smartphone can be equipped with a tongue image acquisition device and long-wave infrared thermal imaging to ensure the quality of tongue image acquisition.

[0057] The edge layer can be comprised of primary healthcare institutions and health check centers, deploying corresponding edge servers. Edge nodes are responsible for model inference tasks within their respective areas and undertake intermediate aggregation work in federated learning, effectively reducing cloud load and request latency. In the cloud, the system is deployed in a private or hybrid cloud environment that meets medical data security requirements. The cloud platform provides high-performance computing nodes for model training and inference, and uses a distributed storage structure to manage model parameters, graph structure data, and environmental time-series data. Backend functional modules all run in a containerized manner and are automatically orchestrated and scheduled by Kubernetes; the API gateway is used to uniformly handle external access, authentication, and traffic control.

[0058] It should be noted that this system constructs a training dataset based on clinical samples from multiple sources, and trains and debugs the YOLO object detection model, the visual state space model, the TCM consultation model of the dynamic consultation module, the spatiotemporal dynamic heterogeneous graph neural network, and the cross-modal fusion Transformer network respectively. To ensure the reliability of model training and the standardization of data sources, we collaborated with affiliated tertiary TCM hospitals and collected over 3,000 high-quality, multimodal case data sets under ethical review and informed consent conditions. The case data includes RGB-T images of the tongue taken under standard lighting conditions, consultation texts recorded by TCM physicians and subsequently structured, time and geographical information at the time of collection, and constitution types and syndrome diagnosis labels independently annotated by multiple physicians. Constitution labeling follows the national standard "Classification and Determination of TCM Constitutions," and syndrome diagnosis is based on the corresponding TCM clinical guidelines for the respective diseases.

[0059] To ensure the model adapts to different imaging conditions and text features, this embodiment preprocesses and enhances the image and text data: before inputting the image data into the model, white balance adjustment, reflective area processing, and various spatial transformations are performed to enhance its robustness; the text data undergoes cleaning, entity recognition, and relation extraction to construct a TCM knowledge graph, providing a structured semantic foundation for the consultation model. The processed dataset is divided into training, validation, and test sets in approximately a 7:2:1 ratio, with the test set used only for final performance validation.

[0060] During the model training phase, the tongue image data processing module uses YOLO-World and Vision-Mamba networks to jointly train the tongue body region, tongue coating region, and multi-dimensional features of the tongue image. The loss function consists of segmentation loss and feature classification loss to simultaneously optimize segmentation accuracy and feature representation ability. The dynamic consultation module is based on a TCM consultation model with a multi-expert structure, continuously pre-trained on a large-scale TCM corpus, and fine-tuned for instructions and preference alignment on structured consultation data. The reinforcement learning part optimizes the strategy in a simulated environment to improve the consultation sequence generation strategy. The environmental data processing module is trained using historical environmental data and health status labels to capture the temporal correlation between solar terms, the Five Elements and Six Qi, and health status. In terms of the overall fusion model, this embodiment adopts an end-to-end training approach, focusing on optimizing the cross-modal fusion Transformer structure and multi-task output layer based on the fixed feature extraction module. The training objective consists of multi-task losses, including cross-entropy loss for constitution classification, mean squared error loss for risk scoring, and binary cross-entropy loss for syndrome prediction. The balance of different prediction objectives is achieved through task weights.

[0061] To evaluate the system's performance in multimodal scenarios, multiple comparative experiments were designed during training. The comparative models included a unimodal visual model based solely on tongue images, a text model based solely on medical history text, a simple multimodal model that directly concatenates tongue images with medical history features, and a reproduced existing multimodal TCM diagnostic model. All models were tested under the same data partitioning and evaluation conditions. Evaluation metrics included accuracy, precision, recall, F1 score, and AUC. Experimental results are shown in Table 1.

[0062] Experimental results show that the model in this embodiment performs well in all major evaluation metrics, and has complete technical implementation in terms of interpretability and privacy protection. Compared with the comparison model, it has a more obvious advantage in overall functional completeness.

[0063] This invention provides an intelligent assessment system that integrates tongue image analysis, medical history taking, and environmental data. It includes modules for tongue image data processing, dynamic medical history taking, environmental data processing, and multi-source data fusion, enabling the gradual association and integration of health-related information from different sources within the same system. Tongue image features provide preliminary physical examination data, the medical history taking module supplements symptom information, and the environmental module further incorporates external factors such as seasonal changes and environmental conditions, allowing the assessment process to comprehensively reflect multidimensional changes in an individual's condition. The assessment results obtained after cross-modal fusion are relatively complete and continuous, and the final report can be presented in a structured format, helping to provide users with more valuable analytical conclusions in health management scenarios.

[0064] Based on the same inventive concept, the present invention also provides a computer device, comprising: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed as described above in the steps of an intelligent assessment system integrating tongue image, medical history and environment.

[0065] The processing methods for computer devices can be referred to the description of the methods above, and will not be repeated here.

[0066] This application also provides a non-transitory machine-readable storage medium storing an executable program, which, when run by a microprocessor, causes the processor to execute the method provided in the above embodiments.

[0067] This invention discloses a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform the described methods.

[0068] This invention discloses a computer program product including a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform the described method.

[0069] The embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0070] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0071] Finally, it should be noted that the embodiments disclosed in this invention are merely preferred embodiments of this invention and are only used to illustrate the technical solutions of this invention, not to limit it. Although this invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this invention.

Claims

1. An intelligent assessment system integrating tongue imagery, medical history taking, and environmental factors, characterized in that: include: The tongue image data processing module acquires tongue image data, segments and extracts features from the tongue image data to obtain tongue image features; The dynamic consultation module dynamically generates a sequence of consultation questions based on the tongue image features, and extracts consultation features based on the response data of the consultation question sequence; The environmental data processing module is used to construct environmental seasonal features based on diagnostic features and environmental data; The multi-source data fusion module performs cross-modal deep fusion of tongue image features, consultation features, and environmental seasonal features based on the influence of tongue image data on the sequence of consultation questions and the adjustment of environmental weights by consultation features, generating multi-dimensional condition assessment results. The results output module generates a physical condition assessment report based on the multi-dimensional condition assessment results.

2. The intelligent assessment system integrating tongue imagery, medical history, and environmental factors according to claim 1, characterized in that, The dynamic consultation module adopts a TCM consultation model based on the MoE architecture. It is trained based on the objective function and reward function through a proximal strategy optimization algorithm to generate a sequence of consultation questions. The TCM consultation model introduces TCM diagnosis and treatment logic constraints, including following the principle of observation before consultation and avoiding invalid questions across syndrome types.

3. The intelligent assessment system integrating tongue imagery, medical history, and environmental factors according to claim 2, characterized in that, The multi-source data fusion module includes a cross-modal fusion Transformer network; the multi-source data fusion module maps tongue image features, consultation features and environmental and seasonal features to a unified latent feature space to obtain image coding, consultation coding and environmental coding; Construct cross-modal feature sequences that include tongue image encoding, medical history encoding, and environment encoding, and add modality embedding or location embedding for different modalities; The cross-modal feature sequence is input into the cross-modal fusion Transformer network, and multi-head attention calculation is performed. Tongue image encoding is used as query vector, consultation encoding and environment encoding are used as key-value pairs, and environment encoding is used as context bias input, so that bidirectional or multidirectional associations are established between tongue image, consultation and environmental seasonal features. Based on the quality of multi-source data, the contributions of the tongue image modality, the consultation modality, and the environment modality are dynamically weighted to output multi-dimensional status assessment results; the quality of multi-source data includes the completeness of tongue image data, the completeness of consultation responses, and / or the completeness of environment data.

4. The intelligent assessment system integrating tongue imagery, medical history, and environmental factors according to claim 3, characterized in that, The multi-source data fusion module also includes an uncertainty quantification unit and an interpretability unit; the uncertainty quantification unit maintains the Dropout layer in an active state during the inference phase, performs multiple random forward propagations on the same input to obtain a set of prediction results, and calculates the confidence and uncertainty of the cross-modal fusion Transformer network output based on the statistical distribution of the prediction result set. The interpretability unit includes a SHAP interpretation subunit for generating global feature contribution and a LIME interpretation subunit for constructing a local linear approximation model for a single prediction, which interpretably presents the results of multi-dimensional status evaluation of tongue features, consultation features and environmental and seasonal features.

5. The intelligent assessment system integrating tongue imagery, medical history, and environmental factors according to claim 2, characterized in that, The environmental data includes solar terms, Five Elements and Six Qi, and meteorological and geographical information; The environmental data processing module adopts a spatiotemporal dynamic heterogeneous graph neural network. The heterogeneous graph includes graph nodes and graph edge sets. The graph nodes include visceral nodes, meridian nodes, emotional nodes, and pathogenic nodes. The graph edge sets include the five elements' generating and restraining relationships, exterior-interior relationships, and tonifying and purging relationships. The spatiotemporal dynamic heterogeneous graph neural network takes the Five Elements and Six Qi and meteorological and geographical information as input to update the graph node attributes and graph edge weights; based on the spatiotemporal dynamic heterogeneous graph neural network, it performs multi-step propagation on the heterogeneous graph to obtain the activation state of each graph node under the current environmental conditions, and uses the activation state as the output to generate environmental solar term features.

6. The intelligent assessment system integrating tongue imagery, medical history, and environment as described in claim 3, characterized in that, The tongue image data includes visible light images and thermal imaging images. The tongue image data processing module uses the YOLO target detection model to identify the tongue body and tongue coating in the visible light images. It extracts the color, texture, and morphological features of the visible light images and the temperature distribution and gradient features of the thermal imaging images through the visual state space model. It then combines the tongue body, tongue coating, color, texture, morphological features, and temperature distribution and gradient features to generate tongue image features.

7. The intelligent assessment system integrating tongue imagery, medical history, and environmental factors according to claim 2, characterized in that, The objective function of the dynamic consultation module is: LCLIP ( θ )=E^ t [my( rt ( θ ) A ^ t ,clip( rt ( θ ),1 ,1+ ) A ^ t )] in, rt ( θ () represents the ratio of the old to the new strategies. A ^ t It is the estimation of the advantage function. Update the constraint coefficients for the strategy; The reward function is: Where Δ H ( St This represents the reduction in information entropy resulting from the consultation question. Fuser It is the user's implicit / explicit feedback rating. w 1, w 2 represents the weighting coefficient.

8. The intelligent assessment system integrating tongue imagery, medical history, and environment as described in claim 7, characterized in that, It also includes a federated learning module for training the model locally on the user terminal and performing differential privacy gradient pruning and noise injection processing when uploading multi-source data.

9. The intelligent assessment system integrating tongue imagery, medical history, and environment as described in claim 7, characterized in that, The physical condition assessment report includes a five-dimensional dynamic conditioning report and an interactive interpretable report; the five-dimensional dynamic conditioning report includes dimensions of diet, daily life, exercise, emotions, and acupoints, and provides pre-adjustment suggestions based on the seasonal changes within a preset time period; The interactive interpretability report is used to demonstrate the impact of different combinations of features on the physical condition assessment report.

10. A computer storage medium, characterized in that, It stores a computer program, which, when executed by a processor, implements the steps of an intelligent assessment system that integrates tongue image, medical history, and environment as described in any one of claims 1-9.