Traditional Chinese medicine tongue picture multispectral intelligent diagnosis system

The TCM tongue image multispectral intelligent diagnostic system, which combines multispectral imaging and deep learning technology, solves the problems of subjectivity and insufficient information acquisition in traditional TCM tongue diagnosis. It achieves accurate quantification and intelligent diagnosis of deep physiological and biochemical information of the tongue, and provides personalized TCM intervention suggestions.

CN121817796APending Publication Date: 2026-04-10刘永芳
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
刘永芳
Filing Date
2025-12-01
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Traditional Chinese medicine tongue diagnosis relies on personal experience, is highly subjective, and is difficult to standardize in different environments. It also cannot fully obtain the deep physiological and biochemical information of the tongue, and cannot achieve multi-dimensional and intelligent diagnosis of TCM syndromes.

Method used

The TCM tongue image multispectral intelligent diagnostic system includes modules for data acquisition and preprocessing, adaptive spectral super-resolution reconstruction, biological tissue optical inversion, multidimensional feature fusion analysis, TCM syndrome inference decision and result interpretation and visualization interaction. It acquires rich spectral information through multispectral imaging technology and performs intelligent diagnosis by combining deep learning and TCM knowledge graph.

Benefits of technology

It enables precise quantification of the microscopic physiological state of the tongue, improves the comprehensiveness and accuracy of diagnosis, provides personalized TCM intervention suggestions, overcomes the subjectivity and limitations of traditional tongue diagnosis, and provides technical support for the modernization, objectification and standardization of TCM.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121817796A_ABST
    Figure CN121817796A_ABST
Patent Text Reader

Abstract

The invention discloses a traditional Chinese medicine tongue picture multispectral intelligent diagnosis system, and relates to the technical field of computer vision. Comprising a data acquisition and preprocessing interface, a self-adaptive spectrum super-resolution reconstruction module, a biological tissue optical inversion module, a multi-dimensional feature fusion analysis module, a traditional Chinese medicine syndrome deduction decision module and a result interpretation and visual interaction module, the system is deployed on a cloud server or a high-performance edge computing terminal, and is connected with a multispectral imaging module through a USB, GigE or wireless mode so as to access a multispectral acquisition device, so that integrated operation of multispectral acquisition and system diagnosis is realized. According to the invention, through introduction of the multispectral imaging technology, rich spectral information far beyond a common visible light image can be obtained from the tongue body, a data basis is provided for subsequent inversion of deep physiological and biochemical indexes, and the problem that traditional tongue diagnosis cannot deeply explore the microscopic physiological state of the tongue body is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a multispectral intelligent diagnostic system for tongue appearance in traditional Chinese medicine. Background Technology

[0002] Tongue diagnosis in Traditional Chinese Medicine (TCM) is an important part of "inspection" in TCM. It involves observing the color, shape, and coating of the tongue to determine the health status and nature of diseases. Traditional tongue diagnosis relies heavily on the personal experience of TCM doctors and suffers from problems such as strong subjectivity, low standardization, and susceptibility to interference from factors such as ambient light. In particular, with the accelerated pace of modern life, people have higher demands for the convenience, objectivity, and accuracy of TCM diagnosis and treatment.

[0003] Existing TCM tongue diagnosis methods mainly focus on the acquisition and analysis of ordinary visible light images. Although they can identify the thickness and color changes of the tongue coating, as well as macroscopic features such as cracks and teeth marks on the tongue, they are difficult to obtain deep physiological and biochemical information inside the tongue, such as microcirculation, tissue water content, and blood oxygen saturation. Traditional methods are sensitive to light conditions and are difficult to standardize under different diagnostic environments. Currently, there is no mature technical solution that can comprehensively and objectively acquire the deep microscopic physiological and biochemical information of the tongue and combine it with macroscopic morphological features to conduct multi-dimensional and intelligent TCM syndrome diagnosis, thereby achieving objectification, quantification, and intelligentization of traditional TCM tongue diagnosis. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a multispectral intelligent diagnostic system for tongue appearance in Traditional Chinese Medicine.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a multispectral intelligent diagnostic system for tongue appearance in traditional Chinese medicine, comprising a data acquisition and preprocessing interface, an adaptive spectral super-resolution reconstruction module, a biological tissue optical inversion module, a multidimensional feature fusion analysis module, a traditional Chinese medicine syndrome deduction and decision-making module, and a result interpretation and visualization interaction module. Each module interacts through a data bus. The system is deployed on a cloud server or a high-performance edge computing terminal and connects to a multispectral imaging module via USB, GigE, or wireless means to access a multispectral acquisition device, thereby realizing the integrated operation of multispectral acquisition and system diagnosis.

[0006] As a further description of the above technical solution: The data acquisition and preprocessing interface connects to a multispectral acquisition device, reads raw RAW image data containing several discrete narrow bands, and uses a semantic segmentation network to classify the raw RAW image data pixel by pixel, identifying and segmenting the oral cavity background, lips, teeth, and tongue, generating a tongue region mask, and outputting tongue region data after removing pixels from non-tongue regions. Using standard whiteboard data and dark current data acquired simultaneously by the acquisition device during shooting, the raw RAW image data of the tongue region is corrected. By subtracting dark current noise from the RAW image data and comparing it with the whiteboard data, the pixel grayscale values ​​of the image are converted into spectral reflectance values, eliminating the influence of uneven light source illumination and sensor response characteristics, and outputting a reflectance image.

[0007] As a further description of the above technical solution: The adaptive spectral super-resolution reconstruction module receives pre-processed discrete narrow-band spectral reflectance images, loads a pre-trained spectral reconstruction Transformer network, and uses a multi-head self-attention mechanism to analyze the correlation between different discrete bands in the spectral reflectance image, as well as the texture context information in the image space. It performs feature extraction and nonlinear interpolation through residual dense blocks to extract spatial features and complete the reconstruction from discrete narrow bands to continuous spectrum. The image data is expanded into a continuous hyperspectral data cube covering the entire visible to near-infrared band. During the reconstruction process, the spectral absorption peaks and reflection valleys corresponding to the tongue tissue are restored, and the spectral gaps between discrete bands are filled, so that each pixel corresponds to a continuous high-resolution spectral curve.

[0008] As a further description of the above technical solution: The biological tissue optical inversion module is based on the multilayer modified Kubelka-Munk model. It constructs a photon transmission simulation environment for tongue tissue, defining the tongue as a layered structure containing a surface liquid film layer, an epithelial layer, and a lamina propria. Standard absorption and scattering characteristics of biochemical substances such as hemoglobin, melanin, and water are set, and inversion constraints consistent with the spectral response of tongue tissue are established. The differential evolution algorithm is used to solve the parameters of the hyperspectral data cube in reverse. Taking the pixels of the hyperspectral data cube as the optimization object, the physiological parameters in the multilayer modified Kubelka-Munk model are iteratively adjusted to generate theoretical spectra. The error between the theoretical spectrum and the reconstructed spectrum of the corresponding pixels is calculated, and the physiological parameters are determined when the error reaches the minimum.

[0009] As a further description of the above technical solution: Physiological parameters include total hemoglobin concentration, blood oxygen saturation, tissue water content, scattering parameters, and melanin and bilirubin concentrations. The inverted physiological parameters are mapped back to a two-dimensional image space to generate physiological and biochemical feature maps. These maps include blood perfusion maps, body fluid distribution maps, tongue coating structure maps, and pigment distribution maps. The blood perfusion map records total hemoglobin concentration and blood oxygen saturation, reflecting the congestion or ischemia of the tongue. The body fluid distribution map records tissue water content, reflecting the moisture level of the tongue. The tongue coating structure map records scattering parameters, reflecting the thickness and fineness of the tongue coating. The pigment distribution map records melanin and bilirubin concentrations, reflecting the distribution of ecchymosis or jaundice.

[0010] As a further description of the above technical solution: The multidimensional feature fusion analysis module performs dual-stream feature extraction of spatial texture flow and physiological feature flow. Spatial texture flow consists of macroscopic morphological features, while physiological feature flow consists of microscopic physiological features. Through a spatial texture flow network based on the EfficientNet architecture, crack, tooth mark, dot, and old / young morphological texture features are extracted from the synthesized RGB image. Through a physiological feature flow network based on a three-dimensional convolutional neural network, the spatial distribution features of physiological parameters in different reflex zones of the tongue are extracted from the physiological and biochemical feature map. A multimodal cross-attention mechanism is used to align and fuse the spatial texture flow and physiological feature flow. Using spatial texture features as a query guide, the corresponding explanatory basis is searched in the physiological feature map, and the fused features are output and mapped into a high-dimensional latent space feature vector.

[0011] As a further description of the above technical solution: The TCM syndrome inference and decision-making module loads a pre-built TCM knowledge graph and maps the high-dimensional feature vectors output by the multi-dimensional feature fusion analysis module to tongue feature nodes in the graph. Using a graph attention network, message passing and probability propagation are performed in the knowledge graph, so that the weights of the tongue feature nodes are passed along the graph edge relationships to the pathogenesis nodes and converge to the syndrome nodes, forming the syndrome activation probability. Based on a preset dynamic threshold, syndrome nodes whose syndrome activation probabilities reach the threshold are selected as syndrome diagnosis results, and multi-label output is supported.

[0012] As a further description of the above technical solution: The results interpretation and visualization interaction module visualizes and textualizes the diagnostic process and results. It includes a multi-level visualization mapping unit, a knowledge graph-driven natural language interpretation unit, and an intelligent intervention recommendation unit. The multi-level visualization mapping unit uses the Grad-CAM++ algorithm to calculate the gradient of the diagnostic results with respect to the neural network feature layer, generates a semi-transparent heat map on the original tongue image, highlights the tongue surface area that contributes the most to the AI ​​decision, normalizes the physiological data matrix output by the biological tissue optical inversion module, and maps it into a pseudo-color image so that users can view the microscopic distribution of the indicators.

[0013] As a further description of the above technical solution: The knowledge graph-driven natural language interpretation unit traces back the complete reasoning path from the syndrome node determined by the TCM syndrome inference decision module back to the feature node in the knowledge graph. Through natural language generation technology, using a template-based generator, the traced reasoning path is transformed into a text description that conforms to TCM logic. The text description includes locating abnormal areas, describing the corresponding abnormal morphological and physiological parameters, explaining the corresponding pathogenesis, and giving the final diagnostic conclusion.

[0014] As a further description of the above technical solution: The intelligent intervention recommendation unit retrieves corresponding intervention plans from the prescription and treatment nodes in the knowledge graph based on the diagnosed syndrome type. It also retrieves corresponding prescription and treatment nodes in the traditional Chinese medicine knowledge graph based on the syndrome diagnosis results. Combining the basic information input by the user, it performs safety filtering on the retrieved plans and outputs a comprehensive intervention report. The intervention report includes recommended Chinese herbal prescriptions, acupoint guidance with a 3D location map and massage techniques, and dietary suggestions with a list of suitable or unsuitable foods.

[0015] The present invention has the following beneficial effects: 1. In this invention, firstly, by introducing multispectral imaging technology, rich spectral information far exceeding that of ordinary visible light images can be obtained from the tongue, providing a data foundation for the subsequent inversion of in-depth physiological and biochemical indicators. Through the adaptive spectral super-resolution reconstruction module, discrete narrow-band data is expanded into a continuous hyperspectral data cube, greatly improving spectral resolution and detail representation. The biological tissue optical inversion module, through the modified Kubelka-Munk model and differential evolution algorithm, achieves accurate quantification of key physiological and biochemical parameters such as total hemoglobin concentration, blood oxygen saturation, tissue water content, scattering parameters, and melanin and bilirubin concentrations inside the tongue, solving the problem that traditional tongue diagnosis cannot deeply explore the microscopic physiological state of the tongue.

[0016] 2. In this invention, the multidimensional feature fusion analysis module combines spatial texture flow and physiological feature flow, and deeply integrates macroscopic morphological features and microscopic physiological features through a multimodal cross-attention mechanism, achieving complementarity and enhancement of information dimensions, and improving the comprehensiveness and accuracy of diagnosis. The TCM syndrome inference decision module utilizes graph attention network to perform message passing and probability propagation in the knowledge graph, realizing intelligent inference and multi-label output of complex TCM syndromes, and improving the level of intelligent diagnosis. The result interpretation and visualization interaction module provides multi-level visualization mapping, knowledge graph-driven natural language interpretation and intelligent intervention recommendation, making the diagnostic results objective and easy to understand, and providing users with personalized intervention suggestions, significantly improving user experience and clinical practical value. The entire system integrates multispectral acquisition and intelligent diagnostic functions, with flexible deployment methods, and can achieve integrated operation, effectively overcoming the subjectivity and limitations of traditional TCM tongue diagnosis, and providing strong technical support for the modernization, objectification and standardization of TCM. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the system architecture of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Reference Figure 1 One embodiment of the present invention is a multispectral intelligent diagnostic system for tongue appearance in traditional Chinese medicine, comprising a data acquisition and preprocessing interface, an adaptive spectral super-resolution reconstruction module, a biological tissue optical inversion module, a multidimensional feature fusion analysis module, a traditional Chinese medicine syndrome deduction and decision-making module, and a result interpretation and visualization interaction module. The modules interact through a data bus. The system is deployed on a cloud server or a high-performance edge computing terminal and is connected to a multispectral imaging module via USB, GigE, or wireless means to access a multispectral acquisition device, thereby realizing the integrated operation of multispectral acquisition and system diagnosis.

[0020] The data acquisition and preprocessing interface connects to a multispectral acquisition device, reads raw RAW image data containing several discrete narrow bands, and uses a semantic segmentation network to classify the raw RAW image data pixel by pixel, identifying and segmenting the oral cavity background, lips, teeth, and tongue. A tongue region mask is generated, and non-tongue region pixels are removed before outputting the tongue region data. Standard whiteboard data and dark current data acquired simultaneously by the acquisition device during shooting are used to correct the raw RAW image data of the tongue region. By subtracting dark current noise from the RAW image data and comparing it with the whiteboard data, the pixel grayscale values ​​of the image are converted into spectral reflectance values, using the formula: , For spectral reflectance, This is the original data. This is dark current noise. To obtain whiteboard data, the effects of uneven light source illumination and sensor response characteristics are eliminated, and a reflectance image is output.

[0021] Example 1: The data acquisition and preprocessing interface is the primary component of the system. Its main function is to acquire raw multispectral image data from the multispectral acquisition device and perform a series of preprocessing steps on this data to ensure data quality and the availability of subsequent modules. This interface connects to the multispectral acquisition device and can read raw RAW image data containing several discrete narrow bands. After data acquisition, the interface first uses a semantic segmentation network to classify the raw RAW image data pixel by pixel. The semantic segmentation network is a deep learning model that classifies each pixel in the image to identify and segment four regions: oral cavity background, lips, teeth, and tongue. The purpose of this step is to accurately determine the boundaries of the tongue and generate a tongue region mask to limit the diagnostic area to the tongue itself and avoid interference from other oral tissues. After generating the tongue region mask, the system removes pixels from non-tongue regions, retaining only the data from the tongue region for subsequent processing. To eliminate environmental influences and equipment errors during the acquisition process, the interface uses standard whiteboard data and dark current data acquired synchronously by the acquisition device during shooting to correct the raw RAW image data of the tongue region. By subtracting dark current noise from RAW image data, the inherent noise of the sensor in the absence of light input can be eliminated. The RAW data, after dark current subtraction, is compared with whiteboard data to convert the image's pixel grayscale values ​​into spectral reflectance values. The whiteboard data represents the response under ideal reflectance conditions; comparison with the whiteboard eliminates the influence of uneven light source illumination and sensor response characteristics. The interface outputs a calibrated and normalized reflectance image, providing high-precision, highly consistent input for subsequent spectral reconstruction and biological tissue parameter inversion.

[0022] The adaptive spectral super-resolution reconstruction module receives pre-processed discrete narrow-band spectral reflectance images, loads a pre-trained spectral reconstruction Transformer network, and uses a multi-head self-attention mechanism to analyze the correlation between different discrete bands in the spectral reflectance image, as well as the texture context information in the image space. It performs feature extraction and nonlinear interpolation through residual dense blocks to extract spatial features and complete the reconstruction from discrete narrow bands to continuous spectrum. The image data is expanded into a continuous hyperspectral data cube covering the entire visible to near-infrared band. During the reconstruction process, the spectral absorption peaks and reflection valleys corresponding to the tongue tissue are restored, and the spectral gaps between discrete bands are filled, so that each pixel corresponds to a continuous high-resolution spectral curve.

[0023] Example 2: The adaptive spectral super-resolution reconstruction module receives preprocessed discrete narrow-band spectral reflectance images output from the data acquisition and preprocessing interface. The core task of this module is to reconstruct these discrete narrow-band data into continuous hyperspectral data covering a wider band range and possessing higher spectral resolution. To this end, the module first loads a pre-trained spectral reconstruction Transformer network. The Transformer network, utilizing its powerful multi-head self-attention mechanism, can deeply analyze the intrinsic correlations between different discrete bands in the spectral reflectance image, as well as the texture context information in the image space. Using a public hyperspectral tongue image database as the ground truth (GrundTruth), and based on the Spectral Response Function (SRF) of the acquisition device, the hyperspectral ground truth is downsampled into simulated multispectral inputs to construct training data pairs. A hyperspectral data cube is generated, covering the 400nm-1000nm range, with a spectral resolution better than 10nm. The multi-head self-attention mechanism allows the model to simultaneously focus on different information subspaces during the reconstruction process, thereby capturing spectral and spatial features more comprehensively. During data processing, this module performs feature extraction and nonlinear interpolation using residual dense blocks. Residual dense blocks effectively learn deep image features and alleviate the gradient vanishing problem in deep network training through residual connections. Nonlinear interpolation techniques are used to generate continuous spectral data across discrete bands. Through these techniques, the system can extract spatial features and reconstruct the spectrum from discrete narrow bands to a continuous spectrum, expanding the original image data into a continuous hyperspectral data cube covering the entire visible to near-infrared band. Each pixel in this data cube corresponds to a continuous high-resolution spectral curve. During reconstruction, the system particularly emphasizes restoring the spectral absorption peaks and reflection valleys corresponding to the tongue tissue. These details are crucial for accurately retrieving physiological and biochemical parameters, while also filling in spectral gaps between discrete bands, ensuring that each pixel corresponds to a physically complete and continuous high-resolution spectral curve. Hybrid loss function formula: , : Represents a hybrid loss function used to train a Transformer network for spectral reconstruction, making the reconstruction results from discrete narrow bands to continuous hyperspectral data approximate the true value. : Represents Mean Squared Error, ensuring numerical accuracy. : Represents the SpectralAngle Mapper loss, ensuring that the shape (peak and trough positions) of the reconstructed spectral curve is consistent with the true value, thus guaranteeing color fidelity. The weighting coefficients in the formula are adjustment parameters that balance the mean square error and the loss of spectral angle mapping. , Mean squared error (MSE) measures the pixel-by-pixel difference between the reconstructed hyperspectral data cube and the true value. The number of pixels (or pixel samples) involved in training. : No. The spectral vector of the reconstructed hyperspectral data cube is obtained from the spectral reconstruction Transfrmer network. : No. The spectral vector of the spectral image of the Grund Truth corresponding to each pixel. The square norm of the difference between two spectral vectors is used to measure the magnitude of the numerical error; this error is averaged across all pixels. , : Spectral angle plotting loss, which measures the degree of consistency in shape between the reconstructed spectrum and the true spectrum (the smaller the angle, the more consistent the curve shape). , , The meanings are the same as above, namely the number of pixels, the reconstructed spectral vector, and the true spectral vector, respectively. : No. The inner product of the reconstructed spectrum and the true spectrum at each pixel measures the consistency of the directions of the two spectral curves. , : These are the L2 norms of the reconstructed spectrum and the true spectrum, respectively, used to normalize the inner product to obtain the angle. Convert the normalized similarity into spectral angles.

[0024] The biological tissue optical inversion module is based on the multilayer modified Kubelka-Munk model. It constructs a photon transmission simulation environment for tongue tissue, defining the tongue as a layered structure containing a surface liquid film layer, an epithelial layer, and a lamina propria. Standard absorption and scattering characteristics of biochemical substances such as hemoglobin, melanin, and water are set, and inversion constraints consistent with the spectral response of tongue tissue are established. The differential evolution algorithm is used to solve the parameters of the hyperspectral data cube in reverse. Taking the pixels of the hyperspectral data cube as the optimization object, the physiological parameters in the multilayer modified Kubelka-Munk model are iteratively adjusted to generate theoretical spectra. The error between the theoretical spectrum and the reconstructed spectrum of the corresponding pixels is calculated, and the physiological parameters are determined when the error reaches the minimum. Physiological parameters include total hemoglobin concentration, blood oxygen saturation, tissue water content, scattering parameters, and melanin and bilirubin concentrations. These inverted physiological parameters are mapped back to a two-dimensional image space to generate physiological and biochemical feature maps. These maps include blood perfusion maps, body fluid distribution maps, tongue coating structure maps, and pigment distribution maps. The blood perfusion map records total hemoglobin concentration and blood oxygen saturation, reflecting the congestion or ischemia of the tongue. The body fluid distribution map records tissue water content, reflecting the dryness or moisture of the tongue. The tongue coating structure map records scattering parameters, reflecting the thickness and fineness of the tongue coating. The pigment distribution map records melanin and bilirubin concentrations, reflecting the distribution of ecchymosis or jaundice. Differential optimization algorithm: Find the optimal This minimizes the error between the theoretical spectrum calculated by the multilayer modified Kubelka-Munk model and the actual reconstructed spectrum output by the adaptive spectral super-resolution reconstruction module (i.e., the true value used as the comparison benchmark). : Refers to each pixel in an image, The referencing algorithm targets pixels. The goal is to find the optimal parameter vector (i.e., the true set of physiological parameters). The component in the parameter vector refers to the concentration of oxyhemoglobin. The component in the parameter vector refers to the concentration of deoxyhemoglobin. The components in the parameter vector represent the tissue water content. The components in the parameter vector represent melanin concentration. The components in the parameter vector refer to scattering parameters, which are related to the density of the scattering volume. The component in the parameter vector refers to the scattering parameter, which is related to the size of the scattering particles and corresponds to the thickness of the tongue coating.

[0025] Example 3: The biological tissue optical inversion module is one of the core functions of the system, aiming to convert hyperspectral data into physiological and biochemical parameters of the tongue. This module is based on a multilayer modified Kubelka-Munk model to construct a photon transport simulation environment for tongue tissue. The Kubelka-Munk model is a classic radiative transport theory, modified to better suit the photon transport characteristics of multilayered biological tissues. In this model, the tongue is defined as a layered structure containing a surface liquid film layer, an epithelial layer, and a lamina propria. This layered simulation more closely resembles the actual physiological structure of the tongue. The system pre-sets the standard absorption and scattering characteristics of biochemical substances such as hemoglobin, melanin, and water; these standard characteristics form the basis for constructing the inversion model. Based on these settings, the inversion module establishes inversion constraints consistent with the spectral response of the tongue tissue, ensuring that the theoretical spectrum inverted by the model is physically consistent with the actual observed spectrum. To solve for physiological parameters from the hyperspectral data cube, this module uses a differential evolution algorithm for inverse parameter solving. The differential evolution algorithm is a highly efficient global optimization algorithm that treats each pixel of the hyperspectral data cube as the optimization object. The algorithm iteratively adjusts physiological parameters in a multi-layered modified Kubelka-Munk model, such as total hemoglobin concentration, blood oxygen saturation, and tissue water content, to generate theoretical spectra. Subsequently, it calculates the error between the generated theoretical spectra and the reconstructed spectra of the corresponding pixels. When the error reaches its minimum, the algorithm determines that the physiological parameters at this point are the corresponding physiological and biochemical parameters for that pixel. Detailed data on these physiological parameters include total hemoglobin concentration, blood oxygen saturation, tissue water content, scattering parameters, and melanin and bilirubin concentrations. These parameters are key indicators reflecting the microscopic physiological state of the tongue. Mapping the inverted physiological parameters back to a two-dimensional image space generates various physiological and biochemical feature maps, thus visually displaying the physiological information inside the tongue. These maps specifically include blood perfusion maps, body fluid distribution maps, tongue coating structure maps, and pigment distribution maps. Blood perfusion maps record total hemoglobin concentration and blood oxygen saturation, directly reflecting the congestion or ischemia state of the tongue, which is crucial for diagnosing the circulation of qi and blood. The body fluid distribution map records tissue water content, clearly reflecting the tongue's moisture level and serving as a crucial basis for assessing abnormalities in body fluid distribution and metabolism. The tongue coating structure map records scattering parameters, reflecting the thickness and fineness of the tongue coating, which is significant for evaluating the true composition of the tongue coating and its pathological changes. The pigment distribution map records melanin and bilirubin concentrations, reflecting the distribution of ecchymosis or jaundice on the tongue surface, providing objective evidence for judging pathological states such as blood stasis and damp-heat. These physiological and biochemical maps overcome the shortcomings of traditional tongue diagnosis in its insufficient perception of deep tissue information, providing strong data support for subsequent TCM syndrome diagnosis.

[0026] The multidimensional feature fusion analysis module performs dual-stream feature extraction of spatial texture flow and physiological feature flow. Spatial texture flow consists of macroscopic morphological features, while physiological feature flow consists of microscopic physiological features. Through a spatial texture flow network based on the EfficientNet architecture, crack, tooth mark, dot, and old / young morphological texture features are extracted from the synthesized RGB image. Through a physiological feature flow network based on a three-dimensional convolutional neural network, the spatial distribution features of physiological parameters in different reflex zones of the tongue are extracted from the physiological and biochemical feature map. A multimodal cross-attention mechanism is used to align and fuse the spatial texture flow and physiological feature flow. Using spatial texture features as a query guide, the corresponding explanatory basis is searched in the physiological feature map, and the fused features are output and mapped into a high-dimensional latent space feature vector.

[0027] Example 4: The multidimensional feature fusion analysis module aims to deeply integrate macroscopic morphological features and microscopic physiological features to form a comprehensive diagnostic basis. It performs dual-stream feature extraction of spatial texture flow and physiological feature flow. Spatial texture flow primarily focuses on macroscopic morphological features, which can be obtained from visible light images or synthesized RGB images. Using a spatial texture flow network based on the EfficientNet architecture, morphological texture features such as cracks, tooth marks, pitting, and age are extracted from synthesized RGB images. EfficientNet is an efficient convolutional neural network architecture that optimizes model size while maintaining model accuracy, making it very suitable for image feature extraction. Physiological feature flow focuses on microscopic physiological features, which originate from the physiological and biochemical feature atlas generated by the biological tissue optical inversion module. Using a physiological feature flow network based on a three-dimensional convolutional neural network (3D CNN), this module can extract the spatial distribution features of physiological parameters in different reflective areas of the tongue surface from the physiological and biochemical feature atlas. 3D CNN can effectively process input data with spatial and multi-channel information (such as different physiological parameters), capturing deeper physiological patterns. To effectively integrate these two different dimensional features, the system employs a multimodal cross-attention mechanism to align and fuse the spatial texture flow and the physiological feature flow. This mechanism allows the model to learn how to correlate information from different modalities, using spatial texture features as query guides to find corresponding explanatory evidence in the physiological feature map, thereby achieving semantic alignment and enhancement between macroscopic morphology and microscopic physiology. Finally, the module outputs fused features and maps these features to high-dimensional latent space feature vectors, providing comprehensive and abstract input for subsequent TCM syndrome inference decisions.

[0028] The TCM syndrome inference and decision-making module loads a pre-built TCM knowledge graph and maps the high-dimensional feature vectors output by the multi-dimensional feature fusion analysis module to tongue feature nodes in the graph. Using a graph attention network, message passing and probability propagation are performed in the knowledge graph, so that the weights of the tongue feature nodes are passed along the graph edge relationships to the pathogenesis nodes and converge to the syndrome nodes, forming the syndrome activation probability. Based on a preset dynamic threshold, syndrome nodes whose syndrome activation probabilities reach the threshold are selected as syndrome diagnosis results, and multi-label output is supported.

[0029] Example 5: The TCM syndrome deduction and decision-making module is the core of intelligent diagnosis. This module loads a pre-built TCM knowledge graph and maps the high-dimensional feature vectors output by the multi-dimensional feature fusion analysis module to tongue feature nodes in the graph. The TCM knowledge graph is a structured knowledge base containing complex relationships between TCM theories, symptoms, pathogenesis, and prescriptions. Through a graph attention network (GAT), this module can perform message passing and probability propagation within the knowledge graph. The GAT learns the weight relationships between different nodes to determine the intensity of information transmission, allowing the weights of tongue feature nodes to be passed along the graph's edge relationships to pathogenesis nodes and ultimately converge to syndrome nodes. This mechanism simulates the logical reasoning process of TCM's four diagnostic methods (inspection, auscultation, inquiry, and palpation), linking objective tongue features with pathogenesis and syndromes in TCM theory. Through message passing and probability propagation, the system forms the activation probability of each syndrome. Based on a preset dynamic threshold, the system selects syndrome nodes whose activation probabilities reach the threshold as the syndrome diagnosis result. This module supports multi-label output, meaning that a single diagnosis result may contain multiple TCM syndromes, which better reflects the complexity of TCM clinical practice.

[0030] The results interpretation and visualization interaction module visualizes and textualizes the diagnostic process and results. It includes a multi-level visualization mapping unit, a knowledge graph-driven natural language interpretation unit, and an intelligent intervention recommendation unit. The multi-level visualization mapping unit uses the Grad-CAM++ algorithm to calculate the gradient of the diagnostic results with respect to the neural network feature layer, generating a semi-transparent heatmap on the original tongue image, highlighting the tongue surface areas that contribute most to the AI ​​decision-making. It normalizes the physiological data matrix output from the biological tissue optical inversion module and maps it to a pseudo-color image, allowing users to view the microscopic distribution of indicators. The knowledge graph-driven natural language interpretation unit traces the complete reasoning path from the syndrome node determined by the TCM syndrome deduction decision module back to the feature node in the knowledge graph. Through natural language generation technology, using a template-based generator, it transforms the traced reasoning path into a text description conforming to TCM logic. The text description includes locating abnormal areas, describing the corresponding morphological and physiological parameter abnormalities, explaining the corresponding pathogenesis, and providing the final diagnostic conclusion. The intelligent intervention recommendation unit retrieves corresponding intervention plans from the prescription and treatment nodes in the knowledge graph based on the diagnosed syndrome type. It also retrieves corresponding prescription and treatment nodes in the traditional Chinese medicine knowledge graph based on the syndrome diagnosis results. Combining the basic information input by the user, it performs safety filtering on the retrieved plans and outputs a comprehensive intervention report. The intervention report includes recommended Chinese herbal prescriptions, acupoint guidance with a 3D location map and massage techniques, and dietary suggestions with a list of suitable or unsuitable foods.

[0031] Example 6: The results interpretation and visualization interaction module is a crucial user-facing component, designed to present the complex diagnostic process and results clearly and intuitively to users, and to provide intelligent intervention suggestions. This module visualizes and textualizes the diagnostic process and results, primarily comprising a multi-level visualization mapping unit, a knowledge graph-driven natural language interpretation unit, and an intelligent intervention recommendation unit. The multi-level visualization mapping unit uses the Grad-CAM++ algorithm to calculate the gradient of the diagnostic results with respect to the neural network feature layers, generating a semi-transparent heatmap on the original tongue image. The heatmap highlights the tongue surface areas that contribute most to the AI's decision-making, allowing users to intuitively understand the basis for the AI's diagnosis. Simultaneously, this unit normalizes the physiological data matrix output from the biological tissue optical inversion module and maps it to a pseudo-color image, enabling users to clearly view the microscopic distribution of various physiological indicators, such as blood oxygen saturation and water content in different areas of the tongue. The knowledge graph-driven natural language interpretation unit further enhances the comprehensibility of the diagnostic results. This unit traces the complete reasoning path from the syndrome node determined by the TCM syndrome deduction decision module back to the feature node within the knowledge graph. This means it can trace back to which specific tongue features (including morphological and physiological / biochemical characteristics) led to a diagnosis of a particular syndrome. Through natural language generation technology, particularly template-based generators, this unit can transform the backtracking reasoning path into textual descriptions consistent with Traditional Chinese Medicine (TCM) logic. These textual descriptions typically include locating abnormal areas (such as the tip of the tongue or the middle of the tongue coating), describing the corresponding morphological and physiological parameter abnormalities (such as "red tongue tip, high blood perfusion"), explaining the corresponding pathogenesis (such as "excessive heart fire"), and providing a final diagnostic conclusion, thus offering an "explainable AI" diagnostic process. The intelligent intervention recommendation unit then provides personalized intervention suggestions to users based on the diagnostic results. According to the diagnosed syndrome type, this unit retrieves corresponding intervention plans from the formula and treatment nodes in the knowledge graph. In the TCM knowledge graph, it retrieves corresponding formula and treatment nodes based on the syndrome diagnosis results to ensure the matching of recommended plans with the diagnostic results. To ensure the safety and personalization of the recommendations, this unit combines basic information input by the user (such as age, gender, allergy history, etc.) to perform safety filtering on the retrieved plans to avoid potential side effects or discomfort. Finally, the system outputs a comprehensive intervention report. This report not only includes recommended traditional Chinese medicine prescriptions but also displays 3D location maps and acupoint guidance for massage techniques, facilitating self-care for users. The report also generates dietary recommendations, including lists of suitable and unsuitable foods, providing users with a comprehensive health management plan.

[0032] Example 7: Intelligent Diagnostic Process for a Single Tongue Image (Multispectral) In this embodiment, the system is a "Traditional Chinese Medicine Tongue Image Multispectral Intelligent Diagnosis System", which includes a data acquisition and preprocessing interface, an adaptive spectral super-resolution reconstruction module, a biological tissue optical inversion module, a multidimensional feature fusion analysis module, a traditional Chinese medicine syndrome deduction and decision-making module, and a result interpretation and visualization interaction module. Each module interacts through a data bus. The system is deployed on a cloud server and connects to a multispectral imaging module via USB to access the multispectral acquisition device.

[0033] S1. Data Acquisition and Preprocessing Subjects naturally extended their tongues in front of the multispectral acquisition device, which acquired raw RAW image data containing several discrete narrow bands (center wavelength covering 450nm to 950nm) and simultaneously acquired standard whiteboard data and dark current data.

[0034] Data acquisition and preprocessing interface execution: The semantic segmentation network (U-Net) classifies the original RAW image pixel by pixel, segmenting the oral cavity background, lips, teeth and tongue, generating a tongue region mask, removing non-tongue region pixels and outputting tongue region data; noise reduction of RAW image data is performed using dark current data, and radiometric calibration and reflectance correction are completed by comparing with standard whiteboard data, converting pixel grayscale values ​​into spectral reflectance values ​​and outputting a reflectance image.

[0035] S2. Adaptive Spectral Super-Resolution Reconstruction After receiving the reflectance image, the adaptive spectral super-resolution reconstruction module loads a pre-trained spectral reconstruction Transfrmer network (SpectralReconstructinTransfrmer, SRT); performs feature extraction and nonlinear interpolation using the multi-head self-attention mechanism of SpectralAttentin and the residual dense block of ResidualDenseBlck; and expands the discrete narrow-band reflectance image into a continuous hyperspectral data cube (400nm-1000nm, with a spectral resolution better than 10nm) covering the entire visible to near-infrared band.

[0036] S3. Optical Inversion of Biological Tissues The biological tissue optical inversion module performs pixel-by-pixel inversion of hyperspectral data cubes: it adopts a multi-layer modified Kubelka-Munk (KM) model, defining the tongue as a three-layer structure consisting of a surface fluid film layer, an epithelial layer, and a lamina propria, and sets the absorption and scattering characteristics of biochemical substances such as hemoglobin, melanin, and water; it uses a differential evolution (DE) algorithm to iteratively adjust the model's physiological parameters to minimize the error between the theoretical spectrum and the reconstructed spectrum; it obtains and outputs physiological parameters: total hemoglobin concentration, blood oxygen saturation, tissue water content, scattering parameters, and melanin and bilirubin concentrations; it maps the parameters back to two-dimensional space to generate blood perfusion maps, body fluid distribution maps, lichen structure maps, and pigment distribution maps.

[0037] In this data, the blood perfusion map showed that the total hemoglobin concentration in the tip of the tongue was high and the blood oxygen saturation was abnormal; the body fluid distribution map showed that the overall tissue water content was low.

[0038] S4. Multidimensional Feature Fusion Analysis The multidimensional feature fusion analysis module performs dual-stream feature extraction and fusion: Spatial Stream: taking a synthetic RGB image as input, it extracts morphological features such as pinpoints, age, cracks, and tooth marks based on the EfficientNet architecture; Physiological Stream: taking a physiological and biochemical feature map as input, it extracts the spatial distribution features of physiological parameters based on a three-dimensional convolutional neural network (3D-CNN); a multimodal cross-attention mechanism is used to guide the alignment and fusion of physiological features with spatial texture features as the query; the fused features are output and mapped to a high-dimensional latent space feature vector.

[0039] In this case, the spatial texture flow detected obvious dotted patterns at the tip of the tongue; the cross-attention mechanism further linked this to abnormal blood perfusion in the area.

[0040] S5. TCM Syndrome Deduction and Decision Making The TCM syndrome inference and decision-making module loads a pre-built TCM knowledge graph (TCM-KG); maps the high-dimensional latent space feature vectors to the initial state of tongue feature nodes in the graph; uses a graph attention network (GAT) for message passing and probability propagation, so that the weights of the tongue feature nodes are passed along the edge relationships to the pathogenesis nodes (heat toxicity, dampness turbidity, qi deficiency) and converge to the syndrome nodes (heart fire excess, liver fire excess, spleen deficiency with dampness / spleen deficiency with dampness excess, liver fire excess); and selects nodes whose syndrome activation probability exceeds the threshold as diagnostic results based on a dynamic threshold.

[0041] The syndrome diagnosis result output in this example is: excessive heart fire (the activation probability of the inferred pathway of increased blood perfusion exceeds the threshold due to tongue tip pricking).

[0042] S6. Results Interpretation and Visual Interaction The results interpretation and visualization interaction module outputs an interpretable report: Multi-level visualization mapping unit: The Grad-CAM++ algorithm is used to generate a significant heat map on the original tongue image, with the heat area concentrated on the tip of the tongue; after normalizing the physiological parameter matrices such as blood perfusion map and body fluid distribution map, a pseudo-color image is generated to intuitively present the abnormal blood flow at the tip of the tongue and the low water content of the whole tongue.

[0043] Knowledge graph-driven natural language interpretation unit: traces back the high-weighted reasoning path from tongue appearance feature nodes to pathogenesis nodes to syndrome nodes; generates text interpretation based on templates, describing the correspondence between abnormal areas on the tip of the tongue, puncture morphology, and increased blood perfusion with the syndrome of "toxic / heart fire excess".

[0044] The intelligent intervention recommendation unit retrieves relevant treatment methods and prescription nodes from the TCM knowledge graph based on the syndrome diagnosis results; performs security filtering based on the user's input information; and outputs a comprehensive intervention report, including: TCM prescription suggestions (prescription nodes related to the graph, such as the "Gentian Liver-Clearing Decoction" type of prescription suggestion in the example); acupoint guidance (e.g., displaying 3D location diagrams and massage techniques for Taichong and Xingjian acupoints); and dietary recommendations (generating a "Do Not Eat / Avoid" list).

[0045] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A multispectral intelligent diagnostic system for tongue appearance in Traditional Chinese Medicine, characterized in that: The system includes a data acquisition and preprocessing interface, an adaptive spectral super-resolution reconstruction module, a biological tissue optical inversion module, a multi-dimensional feature fusion analysis module, a traditional Chinese medicine syndrome deduction and decision-making module, and a result interpretation and visualization interaction module. The modules interact through a data bus. The system is deployed on a cloud server or a high-performance edge computing terminal and connects to the multispectral imaging module via USB, GigE, or wireless to access the multispectral acquisition device, realizing the integrated operation of multispectral acquisition and system diagnosis.

2. The TCM tongue image multispectral intelligent diagnostic system according to claim 1, characterized in that: The data acquisition and preprocessing interface connects to a multispectral acquisition device, reads raw RAW image data containing several discrete narrow bands, and uses a semantic segmentation network to classify the raw RAW image data pixel by pixel, identifying and segmenting the oral cavity background, lips, teeth, and tongue, generating a tongue region mask, and outputting tongue region data after removing pixels from non-tongue regions. Using standard whiteboard data and dark current data acquired simultaneously by the acquisition device during shooting, the raw RAW image data of the tongue region is corrected. By subtracting dark current noise from the RAW image data and comparing it with the whiteboard data, the pixel grayscale values ​​of the image are converted into spectral reflectance values, eliminating the influence of uneven light source illumination and sensor response characteristics, and outputting a reflectance image.

3. The TCM tongue image multispectral intelligent diagnostic system according to claim 1, characterized in that: The adaptive spectral super-resolution reconstruction module receives pre-processed discrete narrow-band spectral reflectance images, loads a pre-trained spectral reconstruction Transfrmer network, and uses a multi-head self-attention mechanism to analyze the correlation between different discrete bands in the spectral reflectance image, as well as the texture context information in the image space. It performs feature extraction and nonlinear interpolation through residual dense blocks to extract spatial features and complete the reconstruction from discrete narrow bands to continuous spectrum. The image data is expanded into a continuous hyperspectral data cube covering the entire visible to near-infrared band. During the reconstruction process, the spectral absorption peaks and reflection valleys corresponding to the tongue tissue are restored, and the spectral gaps between discrete bands are filled, so that each pixel corresponds to a continuous high-resolution spectral curve.

4. The TCM tongue image multispectral intelligent diagnostic system according to claim 1, characterized in that: The biological tissue optical inversion module is based on the multilayer modified Kubelka-Munk model. It constructs a photon transmission simulation environment for tongue tissue, defining the tongue as a layered structure containing a surface liquid film layer, an epithelial layer, and a lamina propria. Standard absorption and scattering characteristics of biochemical substances such as hemoglobin, melanin, and water are set, and inversion constraints consistent with the spectral response of tongue tissue are established. The differential evolution algorithm is used to solve the parameters of the hyperspectral data cube in reverse. Taking the pixels of the hyperspectral data cube as the optimization object, the physiological parameters in the multilayer modified Kubelka-Munk model are iteratively adjusted to generate theoretical spectra. The error between the theoretical spectrum and the reconstructed spectrum of the corresponding pixels is calculated, and the physiological parameters are determined when the error reaches the minimum.

5. The TCM tongue image multispectral intelligent diagnostic system according to claim 4, characterized in that: Physiological parameters include total hemoglobin concentration, blood oxygen saturation, tissue water content, scattering parameters, and melanin and bilirubin concentrations. The inverted physiological parameters are mapped back to a two-dimensional image space to generate physiological and biochemical feature maps. These maps include blood perfusion maps, body fluid distribution maps, tongue coating structure maps, and pigment distribution maps. The blood perfusion map records total hemoglobin concentration and blood oxygen saturation, reflecting the congestion or ischemia of the tongue. The body fluid distribution map records tissue water content, reflecting the moisture level of the tongue. The tongue coating structure map records scattering parameters, reflecting the thickness and fineness of the tongue coating. The pigment distribution map records melanin and bilirubin concentrations, reflecting the distribution of ecchymosis or jaundice.

6. The TCM tongue image multispectral intelligent diagnostic system according to claim 1, characterized in that: The multidimensional feature fusion analysis module performs dual-stream feature extraction of spatial texture flow and physiological feature flow. Spatial texture flow consists of macroscopic morphological features, while physiological feature flow consists of microscopic physiological features. Through a spatial texture flow network based on the EfficientNet architecture, crack, tooth mark, dot, and old / young morphological texture features are extracted from the synthesized RGB image. Through a physiological feature flow network based on a three-dimensional convolutional neural network, the spatial distribution features of physiological parameters in different reflex zones of the tongue are extracted from the physiological and biochemical feature map. A multimodal cross-attention mechanism is used to align and fuse the spatial texture flow and physiological feature flow. Using spatial texture features as a query guide, the corresponding explanatory basis is searched in the physiological feature map, and the fused features are output and mapped into a high-dimensional latent space feature vector.

7. The TCM tongue image multispectral intelligent diagnostic system according to claim 1, characterized in that: The TCM syndrome inference and decision-making module loads a pre-built TCM knowledge graph and maps the high-dimensional feature vectors output by the multi-dimensional feature fusion analysis module to tongue feature nodes in the graph. Using a graph attention network, message passing and probability propagation are performed in the knowledge graph, so that the weights of the tongue feature nodes are passed along the graph edge relationships to the pathogenesis nodes and converge to the syndrome nodes, forming the syndrome activation probability. Based on a preset dynamic threshold, syndrome nodes whose syndrome activation probabilities reach the threshold are selected as syndrome diagnosis results, and multi-label output is supported.

8. The TCM tongue image multispectral intelligent diagnostic system according to claim 1, characterized in that: The results interpretation and visualization interaction module visualizes and textualizes the diagnostic process and results. It includes a multi-level visualization mapping unit, a knowledge graph-driven natural language interpretation unit, and an intelligent intervention recommendation unit. The multi-level visualization mapping unit uses the Grad-CAM++ algorithm to calculate the gradient of the diagnostic results with respect to the neural network feature layer, generates a semi-transparent heat map on the original tongue image, highlights the tongue surface area that contributes the most to the AI ​​decision, normalizes the physiological data matrix output by the biological tissue optical inversion module, and maps it into a pseudo-color image so that users can view the microscopic distribution of the indicators.

9. The TCM tongue image multispectral intelligent diagnostic system according to claim 8, characterized in that: The knowledge graph-driven natural language interpretation unit traces back the complete reasoning path from the syndrome node determined by the TCM syndrome inference decision module back to the feature node in the knowledge graph. Through natural language generation technology, using a template-based generator, the traced reasoning path is transformed into a text description that conforms to TCM logic. The text description includes locating abnormal areas, describing the corresponding abnormal morphological and physiological parameters, explaining the corresponding pathogenesis, and giving the final diagnostic conclusion.

10. The TCM tongue image multispectral intelligent diagnostic system according to claim 8, characterized in that: The intelligent intervention recommendation unit retrieves corresponding intervention plans from the prescription and treatment nodes in the knowledge graph based on the diagnosed syndrome type. It also retrieves corresponding prescription and treatment nodes in the traditional Chinese medicine knowledge graph based on the syndrome diagnosis results. Combining the basic information input by the user, it performs safety filtering on the retrieved plans and outputs a comprehensive intervention report. The intervention report includes recommended Chinese herbal prescriptions, acupoint guidance with a 3D location map and massage techniques, and dietary suggestions with a list of suitable or unsuitable foods.