Traditional Chinese medicine tongue diagnosis auxiliary system based on artificial intelligence

The AI-based TCM tongue diagnosis assistance system enables efficient and accurate analysis of tongue appearance, solves the problems of deep pathological information capture and misjudgment in tongue diagnosis tools, and provides personalized tongue diagnosis assistance and disease monitoring.

CN121709201APending Publication Date: 2026-03-20QIDI HEALTH TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511804663.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing tongue diagnosis aids are unable to effectively capture the deep pathological information of the tongue appearance, and lack the ability to identify deep features such as tongue texture and tongue coating morphology, resulting in low diagnostic efficiency and easy misjudgment.

Method used

An AI-based TCM tongue diagnosis auxiliary system is adopted, including an image processing module, a chromatographic analysis module, a zoning module, and a color difference analysis module. Through video acquisition, texture stabilization processing, tongue zoning, and zoning color difference analysis, a zoning difference feature set is generated to realize the digitization and standardization of tongue surface texture and color features. Combined with a TCM pathology knowledge base, intelligent auxiliary analysis is performed.

Benefits of technology

It improves the accuracy and reliability of tongue diagnosis, reduces misjudgments, provides quantifiable tongue diagnosis indicators, facilitates disease tracking and personalized diagnosis, and enhances the scientific nature and verifiability of TCM tongue diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121709201A_ABST
    Figure CN121709201A_ABST
Patent Text Reader

Abstract

The invention relates to the field of tongue diagnosis image analysis, in particular to a traditional Chinese medicine tongue diagnosis auxiliary system based on artificial intelligence. The traditional Chinese medicine tongue diagnosis auxiliary system based on artificial intelligence comprises an image processing module, a chromatographic analysis module, a partition module, a chromatic aberration analysis module and a tongue diagnosis analysis module. The image processing module is used for collecting a user tongue surface recognition video; performing image processing on the user lingual surface recognition video, and constructing a texture-stable tongue image; the chromatographic analysis module is used for performing chromatographic analysis on the texture of the texture-stabilized tongue image to generate a texture chromatographic analysis result; the partitioning module is used for performing tongue body partitioning recognition based on the texture-stabilized tongue image and outputting a tongue surface partitioning map; and the color difference analysis module is used for performing partition color difference analysis on the lingual surface partition map based on a texture chromatography analysis result to generate a partition difference feature set. According to the invention, objective quantification and intelligent analysis of lingual surface features are realized, and a preliminary tongue diagnosis analysis result of a patient is rapidly and accurately obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tongue diagnosis image analysis, and in particular to an artificial intelligence-based auxiliary system for traditional Chinese medicine tongue diagnosis. Background Technology

[0002] During the acquisition and analysis of tongue images, factors such as lighting conditions, differences in imaging equipment, environmental interference, and individual differences often result in noise, color deviations, and subtle, difficult-to-quantify features in tongue images. Tongue features themselves are complex and diverse, including subtle differences in tongue color, continuous variations in tongue coating thickness, and inconsistent morphologies of cracks and petechiae, all of which place high demands on manual identification. Furthermore, traditional methods relying on experience-based judgment struggle to achieve efficient and consistent pathological analysis when faced with large-scale tongue image data. In addition, because tongue diagnosis involves numerous non-linear and multi-dimensional feature relationships, conventional image processing methods struggle to effectively capture deep pathological information, easily leading to diagnostic biases or even misjudgments. Existing tongue diagnosis aids are mostly based on simple image enhancement or color analysis methods, which, while improving tongue image observation to some extent, generally lack the ability to recognize deep features such as tongue texture, tongue coating morphology, and tongue edge characteristics, and also struggle to establish a reliable mapping relationship between tongue image and pathological state. Moreover, existing systems rely heavily on manual annotation and result judgment in data processing, making real-time analysis, intelligent inference, and automated pathological analysis difficult, resulting in low overall diagnostic efficiency. Summary of the Invention

[0003] To address the aforementioned technical problems, this invention proposes an artificial intelligence-based auxiliary system for TCM tongue diagnosis, thereby resolving at least one of the aforementioned technical issues.

[0004] To achieve the above objectives, the present invention provides an artificial intelligence-based TCM tongue diagnosis auxiliary system, which includes an image processing module, a chromatographic analysis module, a zoning module, a color difference analysis module, and a tongue diagnosis analysis module. The image processing module is used to acquire user tongue surface recognition video; perform image processing on the user tongue surface recognition video to construct a texture-stabilized tongue image; The chromatographic analysis module is used to perform chromatographic analysis on textured tongue images with stable textures and generate chromatographic analysis results. The partitioning module is used to perform tongue body partitioning recognition based on texture-stabilized tongue images and output a tongue surface partitioning map. The color difference analysis module is used to perform zonal color difference analysis on the tongue surface zonal map based on the texture chromatographic analysis results, and generate a zonal difference feature set; The tongue diagnosis analysis module is used to perform pathological analysis based on the partition difference feature set and output an intelligent tongue diagnosis analysis report.

[0005] The specific beneficial effects of this invention are as follows: It obtains continuous dynamic images of the tongue surface within a short period, avoiding lighting, angle, and motion errors that may exist in single-frame static images. Video acquisition can capture subtle texture changes and natural postures of the tongue; compared to static photographs, video data better reflects the true state of the tongue surface, improving the accuracy and reliability of diagnosis. Image processing can eliminate interference such as motion blur, uneven lighting, and noise, generating high-quality, texture-stable tongue images. Texture-stable tongue images can realistically and clearly reflect the texture structure of the tongue surface, providing a reliable basis for texture and color analysis. It improves the accuracy of tongue zoning and color difference analysis, reducing misjudgments. Through texture chromatographic analysis, the texture and color features of the tongue surface are digitized and standardized, providing a basis for quantitative analysis. It can automatically identify tongue surface texture types (such as tongue coating thickness, cracks, teeth marks, etc.) and color distribution (such as red, light, purple, etc.), enhancing objectivity. It provides quantifiable indicators for traditional Chinese medicine tongue diagnosis, facilitating long-term tracking and monitoring of disease changes. The tongue surface is divided into different functional regions (e.g., the heart, liver, spleen, lungs, and kidneys correspond to the tip, sides, and root of the tongue), forming a structured tongue surface image. Regional identification improves the spatial accuracy of the analysis, providing a foundation for personalized diagnosis and making TCM tongue diagnosis more scientific and repeatable. Regional color difference analysis can accurately quantify the color and texture differences in different areas of the tongue surface, such as a reddish tip and a purplish root. The generated "regional difference feature set" is the core data for pathological analysis and can be used to identify abnormal tongue patterns. This transforms subjective observation into an objective feature set, improving the scientific rigor and verifiability of TCM tongue diagnosis. Matching quantitative tongue surface features with a TCM pathology knowledge base enables intelligent auxiliary analysis for tongue diagnosis. Attached Figure Description

[0006] Figure 1 This is a schematic diagram of the structure of a TCM tongue diagnosis auxiliary system based on artificial intelligence according to the present invention. Detailed Implementation

[0007] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0008] This application provides an AI-based TCM tongue diagnosis assistance system. The executing entities of the AI-based TCM tongue diagnosis assistance system include, but are not limited to, the following: mechanical equipment, data processing platform, cloud server node, network upload device, etc., which can be considered general computing nodes of this application. The data processing platform includes, but is not limited to, at least one of: an audio-image management system, an information management system, and a cloud data management system.

[0009] Please see Figure 1 This invention provides an artificial intelligence-based TCM tongue diagnosis auxiliary system, which includes an image processing module, a chromatographic analysis module, a zoning module, a color difference analysis module, and a tongue diagnosis analysis module. The image processing module is used to acquire user tongue surface recognition video; perform image processing on the user tongue surface recognition video to construct a texture-stabilized tongue image; The chromatographic analysis module is used to perform chromatographic analysis on textured tongue images with stable textures and generate chromatographic analysis results. The partitioning module is used to perform tongue body partitioning recognition based on texture-stabilized tongue images and output a tongue surface partitioning map. The color difference analysis module is used to perform zonal color difference analysis on the tongue surface zonal map based on the texture chromatographic analysis results, and generate a zonal difference feature set; The tongue diagnosis analysis module is used to perform pathological analysis based on the partition difference feature set and output an intelligent tongue diagnosis analysis report.

[0010] In the embodiments of the present invention, see Figure 1 This paper provides an AI-based TCM tongue diagnosis auxiliary system. In this embodiment, the AI-based TCM tongue diagnosis auxiliary system includes an image processing module, a chromatographic analysis module, a zoning module, a color difference analysis module, and a tongue diagnosis analysis module. The image processing module is used to acquire user tongue surface recognition video; perform image processing on the user tongue surface recognition video to construct a texture-stabilized tongue image; In this embodiment, a suitable device for tongue image acquisition is selected, such as a high-resolution camera (1080p or higher resolution is recommended), equipped with an adjustable brightness ring light source to ensure uniform illumination of the tongue surface during the shooting process and reduce shadows and reflection interference. During the acquisition process, the subject is required to keep their tongue extended, with the tongue as flat and slightly relaxed as possible to avoid blurring of texture caused by tongue curling or trembling. The frame rate during video acquisition is usually set between 30 and 60 frames per second to capture subtle changes in the tongue surface. The acquisition process also needs to control the background and ambient light to avoid interference from other light sources on the tongue surface color; a dark or neutral-colored background can be used. The video duration is generally controlled between 2 and 5 seconds to ensure that enough frames (e.g., 60-300 frames) are acquired for image processing without data redundancy due to slight tongue movements. After acquisition, preliminary frame screening can be performed to remove abnormal frames caused by blinking, saliva reflection, or large tongue movements. Frame-by-frame preprocessing of the video frames is performed, including noise reduction, color correction, and alignment. Noise reduction typically employs multi-frame averaging or spatiotemporal filtering methods. For example, weighted averaging is performed on each pixel across 3–5 adjacent frames to reduce the impact of transient reflection noise and lighting flicker, while preserving subtle texture information on the tongue surface. For color correction, white balance correction and color space conversion (such as RGB to Lab or HSV) can be used to ensure color consistency across different frames and eliminate interference from lighting fluctuations in color analysis. Subsequently, inter-frame registration is required to correct for slight movements or jitters of the tongue in the video. Image registration methods based on feature point matching or optical flow tracing can be used. This involves calculating tongue contour feature points (such as key points on the tongue tip and sides) and performing affine or rigid transformations to ensure consistent tongue surface geometry across frames. After registration, brightness compensation is achieved through local brightness reflectance calculation. This involves standardizing each pixel based on its brightness values ​​across multiple frames to generate a brightness-stable frame sequence, thereby eliminating the influence of locally high-reflectance or low-light areas on the tongue surface. The brightness-stable frame sequence is further compensated for texture perturbation. By analyzing and filtering the texture gradient of multiple frames, short-term texture changes caused by fluctuations in tongue surface humidity or saliva reflection are removed, resulting in a uniform and stable textured tongue image.

[0011] The chromatographic analysis module is used to perform chromatographic analysis on textured tongue images with stable textures and generate chromatographic analysis results. In this embodiment, the tongue image is converted to a multi-channel feature space suitable for texture and color analysis, such as Lab or HSV space. The L channel reflects brightness, and the a and b channels reflect hue shift, or the H channel represents hue and the S channel represents saturation. Multi-channel processing allows for more accurate capture of tongue color differences and moss color characteristics. Subsequently, local texture calculations are performed on the tongue image, typically defining small unit windows of 8×8 or 16×16 pixels. Texture gradient analysis and color difference statistics are performed on the pixels within each unit window, such as calculating pixel grayscale or brightness gradients, hue change rates, and local standard deviations, to obtain information on the directionality, coarseness, and intensity of the local texture. Furthermore, these local texture indicators are combined with color deviations using multi-frame accumulation or weighted averaging methods to generate texture chromatographic feature values ​​for each region, including the principal color vector, local color difference value, principal texture direction, and local color stability. For example, in the tip of the tongue, a predominantly red color can be extracted, with a fine and uniform texture; in the root of the tongue, a predominantly yellow or grayish-white color can be identified, with a slightly higher texture roughness. Finally, the texture chromatographic features of all local units are summarized to form a texture chromatographic analysis result matrix, which includes not only the local texture information of the tongue body and tongue coating, but also color distribution patterns, local abnormal textures, and overall texture consistency indicators.

[0012] The partitioning module is used to perform tongue body partitioning recognition based on texture-stabilized tongue images and output a tongue surface partitioning map. In this embodiment, the tongue contour is extracted using edge detection algorithms (such as Sobel or Canny algorithms) to identify the geometric features of the tongue tip, middle, root, and left and right boundaries. Subsequently, tongue tip curvature analysis and root color variation are used to identify vertical key points, serving as anchor points for longitudinal tongue segmentation. For example, the peak value of the tongue tip curvature marks the starting point of the tongue tip, and the abrupt change in root color marks the starting position of the root. The tongue length is then divided longitudinally into three equal parts in a 1:1:1 ratio, resulting in the tongue tip region, tongue middle region, and tongue root region. Horizontal segmentation relies on the tongue midline, dividing the left and right sides of the tongue into left and right tongue edge regions. Tongue midline recognition employs a combination of pixel symmetry analysis, brightness gradient, and texture direction consistency. By calculating the point of minimum difference between the left and right pixel columns and supplementing with local texture direction vector field optimization, a smooth and continuous tongue midline curve is obtained. After completing the longitudinal and horizontal segmentation, the contours of each region are smoothed and closed using spline interpolation or Gaussian filtering to ensure natural and jagged partition boundaries. The information on the tongue tip, middle, root, and left and right sides is integrated into a label matrix to generate a tongue surface partition map. In this structure map, each pixel corresponds to a unique partition identifier, while preserving the tongue surface texture and color information.

[0013] The color difference analysis module is used to perform zonal color difference analysis on the tongue surface zonal map based on the texture chromatographic analysis results, and generate a zonal difference feature set; In this embodiment, the texture chromatogram analysis results are mapped onto the tongue surface partition map to achieve a precise correspondence between local color and texture features within each partition. This step typically employs pixel-level mapping, associating the texture chromatogram information of each 8×8 or 16×16 pixel unit with its corresponding partition, thereby achieving spatial aggregation of color and texture features within the partition. Subsequently, local color difference statistics are performed within each partition, calculating the color difference ΔE between each pixel or unit and the dominant color of the partition to obtain the degree of color deviation within the partition. The mean brightness, standard deviation of saturation, and texture direction deviation within the partition can be calculated to reflect local texture stability and color consistency. For example, the tip of the tongue may exhibit a reddish dominant color, with a mean color difference ΔE of approximately 6–10; the root of the tongue may exhibit a yellowish-white coating, with a color difference ΔE of approximately 8–12. After completing the statistics within a single partition, inter-regional comparisons are performed. By calculating the mean color difference, maximum difference, and texture gradient changes between partitions, inter-regional difference indicators are formed. For example, the difference in hue, saturation, and texture direction between the left and right sides of the tongue can be quantified to determine whether there is asymmetry or local anomalies. The color difference statistics, inter-regional difference indicators, local texture deviation, and abnormal color point information of each zone are integrated to form a complete set of zone difference features.

[0014] The tongue diagnosis analysis module is used to perform pathological analysis based on the partition difference feature set and output an intelligent tongue diagnosis analysis report.

[0015] In this embodiment, the set of regional difference features is input into a preset pathological feature model, which includes various tongue diagnosis pathological templates, such as Qi deficiency, blood stasis, damp heat, phlegm dampness, Yin deficiency, and Yang deficiency. Each template consists of typical regional difference indicators, including the regional main color deviation value ΔE, abnormal color point density, texture direction deviation, saturation, and brightness distribution characteristics. The system uses similarity calculation methods, such as Euclidean distance, cosine similarity, or Mahalanobis distance, to match each template with the current tongue surface regional difference features, generating similarity scores for various pathological templates. Subsequently, confidence is calculated based on the similarity scores, identifying high-similarity templates as primary pathological signs, while simultaneously assessing secondary pathological trends to quantify the predominant pathological direction and potential development trends. During the analysis process, the differences between regional sections are also considered, such as a stronger red tongue tip with increased blood stasis points, and a thick tongue root coating with a higher damp heat score, thereby providing a regional basis for pathological judgment. Finally, the system outputs the structured results, including pathological type, description of regional characteristics, abnormal point distribution map, and pathological confidence score, and generates suggested health tips or treatment plans. The report is presented in a combination of visual charts and text descriptions, such as tongue surface regional heatmaps, abnormal point distribution diagrams, confidence values ​​of each pathological template, and comprehensive assessment conclusions. This allows users or physicians to intuitively understand the relationship between tongue surface characteristics and potential pathological states, thereby realizing AI-based TCM tongue diagnosis-assisted diagnosis and personalized health management.

[0016] In this embodiment, the image processing module is used to acquire a user's tongue surface recognition video; and to perform image processing on the user's tongue surface recognition video to construct a texture-stable tongue image, specifically for: Acquire a video of user tongue surface recognition; decompose the user's tongue surface recognition video into image frames to generate a tongue surface image frame sequence; The local luminance reflectance of each frame is calculated by performing frame-by-frame local luminance reflectance calculation on the tongue surface image frame sequence to obtain the local luminance reflectance of each frame. Brightness compensation is performed based on the local brightness reflectivity to output a brightness-stable image sequence. Texture perturbation compensation is performed on brightness-stable image sequences to construct texture-stable tongue images.

[0017] In this embodiment, the acquisition of user tongue surface recognition video is the foundation of the entire tongue diagnosis-assisted pathological analysis process. Therefore, during the acquisition phase, strict control needs to be exercised over multiple aspects, including the light source, camera equipment, patient posture stability, and shooting environment, to ensure the acquisition of tongue surface images with high stability and high diagnostic value. The shooting equipment typically uses a CMOS camera with high dynamic range (HDR) capability and high sensitivity, with a recommended resolution of 1080p or higher and a frame rate controlled at 30–60 fps to capture subtle changes in tongue surface texture and luster. Regarding the light source, it is recommended to use a standard LED light source with a color temperature of 5000–5500K and a color rendering index (CRI ≥ 95) to provide natural illumination close to that of a clinical tongue diagnosis lightbox, making the tongue color more realistic. The light intensity can be set in the range of 300–500 lux, and the light should evenly illuminate the tongue surface to avoid creating localized highlights or dark corners. The shooting distance is generally maintained at 12–18 cm. Jaw support or lens fixation structures are used to reduce head movement, thereby minimizing tongue displacement in the video sequence. During shooting, the subject is guided to naturally extend their tongue and remain still for approximately 1.5–2 seconds to minimize slight shaking caused by breathing. The acquisition system automatically monitors image brightness and sharpness. If the average brightness is too high (>230) or edge sharpness is insufficient (e.g., low Laplacian variance), the system will prompt for repositioning or re-acquisition. After video acquisition, the video content needs to be decomposed frame by frame to obtain a stable sequence of tongue images suitable for brightness and texture processing. A decoder with the corresponding encoding format (such as an H.264 or H.265 decoding module) is used to convert the short video stream into single-frame static images, forming the initial frame set. Due to potential natural slight tongue movement, changes in light reflection within the oral cavity, and brightness instability caused by the automatic adjustment of the camera system, the initial frames need to be screened. The screening strategy is generally based on structural similarity (SSIM). When the SSIM between adjacent frames drops significantly, such as below 0.6, it can be considered that the tongue is moving rapidly, and the frame needs to be removed. Simultaneously, brightness variation detection is used; when the brightness variation between frames exceeds a set range (e.g., ±15%), the frame is identified as having unstable lighting and is excluded. To ensure that processing is focused on the tongue region, tongue region detection algorithms, such as lightweight convolutional neural networks (MobileNet, YOLO-Nano, etc.), are used for tongue surface localization. The tongue region is then cropped to a fixed size, such as 512×512 or 640×640, to ensure structural consistency between different frames. For positional offsets in the tongue edge region, active contour models can be used for fine-tuning to maintain geometrically stable continuity in the frame sequence.

[0018] The tongue surface region in the image is divided into several small patches according to rules, such as 16×16 or 32×32 sub-regions, so that the light reflection behavior of each region can be analyzed independently. Before calculation, the tongue surface frame is usually converted from RGB space to CIELab or HSV color space, and the luminance channel (such as L or V) is used as the core calculation index. Then, according to the light reflection model (such as Retinex theory, multi-scale Retinex, or luminance estimation method based on local contrast), the deviation ratio of the local luminance value of each region to the luminance of the surrounding neighborhood is calculated to obtain the regional reflectance. For example, when the luminance of a certain region is significantly higher than that of its neighborhood (such as reaching more than 1.3 times), it can be regarded as enhanced local reflection of the tongue surface, which is usually related to factors such as the moisture of the tongue surface, the luster of the tongue coating, and the microstructure of the tongue surface protrusions. To observe the stability of brightness over time, the brightness sequence of the same area in multiple consecutive frames (e.g., 30 frames) is analyzed. Standard deviation, trend, and frequency pattern are used to determine whether the changes are caused by micro-perturbations in illumination or by variations in the tongue's texture itself. The purpose of brightness compensation is to eliminate uneven brightness in the image caused by factors such as shooting angle, changes in tongue curvature, local highlight reflections, or changes in automatic exposure, ensuring a natural and stable overall and local brightness in each frame of the tongue image. Based on the local brightness reflectance map obtained in the previous step, a compensation coefficient can be calculated for each area: when reflectance is high, the compensation coefficient is set to less than 1 (e.g., 0.7–0.9) to reduce brightness; when reflectance is low, the compensation coefficient is set to greater than 1 (e.g., 1.1–1.3) to increase brightness. To avoid unnatural brightness discontinuities between areas, the compensation coefficient map needs to be smoothed, such as by using bilateral filtering or guided filtering, to create a smooth transition in brightness changes. After local compensation, sequence-level global brightness alignment is required to ensure consistent overall brightness across frames. This is often achieved through brightness mean normalization or histogram matching, adjusting the brightness distribution of the sequence to match that of a "reference frame," such as selecting the frame with the most balanced brightness as the baseline. After this series of compensations, the tongue images will exhibit greater consistency in brightness, and local details such as differences in tongue coating thickness, tongue surface gloss levels, and color differences between pale red and pale white tongues will be easier to identify. Even after sufficient brightness correction, the tongue image may still show slight temporal misalignment or jitter due to minor tongue movements, changes in oral cavity humidity, camera noise, and minute lens shake. Therefore, further texture correction is needed to ensure a continuous and stable texture structure. Texture perturbation compensation is typically performed from three perspectives: temporal smoothing, directional consistency correction, and local motion field restoration.For temporal smoothing, the brightness or texture response of each pixel in consecutive frames is statistically analyzed. Median filtering or bilateral temporal filtering over 7–11 frames is used to suppress random texture noise while maintaining the original fine texture structure. Regarding directional consistency, texture energy in different directions (e.g., 0°, 45°, 90°, 135°) is extracted using Gabor filters. If an abnormal jump occurs in the response value of an adjacent frame in a certain direction (e.g., a change exceeding 20%), it is determined to be a non-realistic texture change and requires repair using directional field interpolation between consecutive frames to ensure a natural transition in texture direction. In areas with drastic texture changes, such as the junction of the tongue edge and tongue coating, and the edges of cracks, dense optical flow is used to estimate local motion vectors. Optical flow correction is then used to prevent texture stretching, drift, or breakage.

[0019] In this embodiment, the specific steps for performing texture perturbation compensation on the brightness-stable image sequence to construct a texture-stable tongue image are as follows: Calculate the multi-frame texture gradient change value of each pixel based on a brightness-stable image sequence; Based on the multi-frame texture gradient change values, regions are divided to obtain stable and unstable texture regions. Texture components are identified and extracted from stable texture regions to obtain principal component features of the texture. Interference analysis is performed on the unstable region to extract texture interference information; the texture interference information includes transient reflection noise and humidity disturbance components; Perturbation compensation is performed based on the texture interference information to obtain low-stability texture features; We construct a textured tongue image by weighted reconstruction of the principal component features and low-stability texture features.

[0020] In this embodiment, each frame of the tongue surface image is converted into a space suitable for texture analysis, such as grayscale space or the L channel in CIELab space, to reduce the interference of color factors on texture calculation. For each frame, gradient operators such as Sobel, Scharr, or Laplacian are used to calculate the local texture gradients in the horizontal and vertical directions, forming a gradient magnitude map. To enhance sensitivity to subtle texture changes, multi-scale gradient analysis can be used, such as calculating gradient responses at different scales of 3×3, 5×5, and 7×7, and then normalizing and synthesizing the results. Next, statistical analysis is performed on the gradient magnitude sequence of the same pixel location in consecutive frames (e.g., 30–50 frames). By calculating indicators such as gradient change rate, local variance, and texture stability index in the temporal direction, the texture changes of that pixel in the temporal dimension are described. If a pixel exhibits small gradient fluctuations across consecutive frames and maintains a low standard deviation (e.g., below a set threshold of 3–5), it indicates stable texture in the region where that pixel is located. Conversely, frequent and significant gradient changes, or short-term peak jumps, often indicate that the texture is affected by humidity, tongue coating reflections, or micro-displacements of the tongue. Smooth clustering of pixel-level gradient changes can be performed using K-means, region growing, or local density-based clustering methods. The changes can be categorized into several classes based on stability, such as "high stability," "medium stability," and "low stability." In practical applications, "high stability" regions typically exhibit small texture variations and low gradient sequence standard deviations (e.g., below the threshold of 3), often appearing in dry areas of the central tongue or areas with uniform tongue surface. "Low stability" regions are characterized by dramatic gradient changes and large fluctuations, commonly found at strong reflective points on the tongue tip, moist areas of the tongue coating, or locations where the angle of light exposure on the tongue edge changes significantly. During the region segmentation process, to avoid judgment errors caused by single-point noise, a neighborhood consistency check is employed. For example, a weighted average of the gradient stability of 5×5 pixel blocks is used, and regions are classified on a block-by-block basis. Furthermore, superpixel-based region segmentation (such as SLIC superpixel segmentation) can be introduced to create more natural spatial boundaries between stable and unstable regions. After region segmentation, the tongue surface is divided into two main structures: one type, the "texture-stable region," which represents the true biological texture of the tongue; and the other type, the "texture-unstable region," which is mainly affected by light reflection, humidity disturbance, surface liquid flow, and local motion.

[0021] Within stable texture regions, texture signals accurately reflect the true surface structure of the tongue, thus requiring in-depth texture component identification and feature extraction. To enhance the comprehensiveness of texture representation, multidimensional texture analysis methods are typically employed, including statistically based Gray-Level Co-occurrence Matrix (GLCM), frequency-domain based Gabor filter banks, and local pattern-based LBP and its extended forms (such as CLBP and LBP-TOP). GLCM can describe parameters such as contrast, energy, entropy, and correlation of texture, demonstrating good discriminative power for features such as tongue coating thickness, granulation degree, and tongue surface roughness. Gabor filter banks capture the directionality and periodicity of tongue surface texture through convolution responses at different frequencies and directions (such as 0°, 45°, 90°, and 135°), enabling accurate identification of structures such as cracks, petechiae, and granular distribution. LBP-like methods can describe local microstructural features and are suitable for distinguishing between tongue coating granularity and subtle texture variations. To enhance the temporal consistency of texture features, a weighted fusion process is applied to stable regions in the temporal dimension. This involves taking a weighted average of the texture responses at the same location across consecutive frames, thereby reducing the impact of random noise on the features. Based on the gradient change characteristics of this region, interference is categorized into two types: one is transient reflection noise, commonly found on the tip of the tongue, wet and bright areas of the tongue surface, and near highlights. Its main characteristic is a sudden jump in gradient value within a short period, exhibiting a peak-like change over time. The other type is humidity disturbance, often caused by a thin layer of liquid on the tongue surface. Its texture typically exhibits slow fluctuations, gentle brightness changes, and a gradual blurring trend over time. Time series analysis methods can be used to extract interference components, such as peak detection of gradient values ​​across consecutive frames, calculation of local change rates, and analysis of the correlation between brightness and gradient. For transient reflection noise, typical characteristics include: a sudden increase in gradient amplitude exceeding 2–4 times the original stable value, and a short duration (e.g., 1–3 frames). For humidity disturbance components, low-frequency texture change models are used for identification, such as time low-pass filters, local energy analysis, or optical flow-based consistency detection, to identify humidity-induced texture diffusion, micro-blurring, and low-frequency fluctuation features.

[0022] Differential compensation is performed based on the interference type obtained in the previous step. For transient reflection noise, the main causes are excessive light reflection or specular reflection formed by liquid on the tongue surface. Therefore, the compensation strategy focuses on suppressing high-frequency fluctuations. This is typically achieved by using gradient amplitude suppression, local brightness regression, or reflection suppression filters to weaken the sudden jumps in high-frequency gradient responses, bringing them back to a level similar to the surrounding texture. Additionally, a guide map-based filtering method can be used to maintain structural consistency with the surrounding area during texture restoration. For humidity disturbance components, due to their low-frequency, slow-spreading texture changes, the compensation method focuses on restoring the main structural contours of the restored area. This can be achieved by inferring the original texture direction through optical flow, restoring blurred texture details using multi-frame texture interpolation, or enhancing texture clarity using local structure enhancement algorithms. The compensation process also incorporates neighborhood consistency principles, such as correcting the consistency of texture direction and amplitude in 5×5 or 7×7 neighborhoods to avoid texture misalignment or breakage. After extracting the principal components of the texture in stable regions and compensating for the low-stability textures in unstable regions, the two types of texture information need to be weighted and fused according to their stability to construct a texture-stabilized tongue image. Based on the texture stability index of each region (which can be composed of indicators such as texture gradient standard deviation and temporal consistency score), different weights are assigned to the principal component features and low-stability features of the texture. For example, stable regions can be assigned a weight of 0.7–0.9, while low-stability regions can be assigned a weight of 0.1–0.3, to ensure that the overall texture structure is mainly determined by real and stable texture information. A combination of spatial and temporal weighting is used in the fusion process: in the spatial dimension, the weights are smoothed through pixel neighborhood consistency checks; in the temporal dimension, multi-frame fusion strategies (such as weighted averaging, confidence fusion, and texture direction consistency correction) are used to ensure the texture continuity of the tongue image. For key areas such as tongue surface cracks, ecchymosis, and granularity, structure-preserving filters are used to ensure that their details are preserved at a high proportion during fusion.

[0023] In this embodiment, the specific steps for constructing a textured, stable tongue image by weighted reconstruction of the principal component features and low-stability texture features are as follows: Calculate the energy distribution of the principal component features of the texture; calculate the steady-state contribution of the principal components based on the energy distribution; The perturbation residual amount of the low-stability texture features is evaluated to obtain the confidence level of the compensation domain; The steady-state contribution of the principal components and the confidence level of the compensation domain are normalized to generate a set of weighted coefficients; Linear fusion is performed based on the weighted coefficient set to construct a texture-stable tongue image.

[0024] In this embodiment, after extracting the principal component features of the stable texture region, it is necessary to further evaluate the energy distribution of these features to determine their steady-state contribution to the entire tongue surface texture structure. "Energy distribution" mainly refers to the energy intensity of texture features in the spatial or frequency domain, used to measure the importance and stability of the texture structure. Common methods include Discrete Wavelet Transform (DWT), Gabor energy analysis, and local gradient energy integration. Principal component features can typically be analyzed hierarchically by direction, scale, and spatial location: for example, using a Gabor filter bank with 4 scales and 6 directions to convolve the principal component texture and calculating the energy integral of each response map; or using wavelet decomposition to decompose the texture principal components into four sub-bands: LL, LH, HL, and HH, and recording their energy ratio distribution. To make the energy distribution more stable, a weighted average of the principal component texture energy at the same location is calculated over multiple time frames, for example, using the stable region of the most recent 10–20 frames as a reference, and reducing random fluctuations through temporal smoothing. After obtaining the energy distribution, the steady-state contribution of the principal components is further calculated. Steady-state contribution reflects the importance of a texture component in the overall structure of the tongue surface and can be calculated using indicators such as energy percentage, directional consistency, and texture intensity persistence. For example, principal components with an energy percentage exceeding 15% can be assigned a higher base contribution; textures with high directional consistency (variation range less than 5°) can further increase their steady-state weight; and parts whose intensity remains stable across multiple frames can be superimposed with a temporal stabilization factor. Low-stability texture features originate from areas affected by humidity, reflection, highlights, or slight motion. Although these features have undergone prior compensation, they may still retain some residual perturbations, thus requiring further evaluation to calculate the "compensation domain confidence." The temporal recovery consistency of low-stability textures is evaluated, for example, by calculating the structural similarity (SSIM) change rate, local gradient consistency index, and brightness stability index for the compensation results across multiple consecutive frames. If the fluctuation range of these indicators is small (e.g., SSIM fluctuation below 0.05, gradient variance below 4), the compensation is relatively reliable; otherwise, the confidence needs to be reduced. Furthermore, residual interference in the spatial dimension is detected, such as identifying areas where local highlights have not been completely eliminated, edge areas where low-frequency blur caused by humidity has not been fully recovered, or texture misalignment areas caused by motion. Residual disturbances can be detected by the difference between low-frequency energy and high-frequency texture response. For example, the proportion of high-frequency texture energy can be calculated (e.g., using a high-pass filter). If the proportion is more than 40% lower than that of the normal tongue-like texture area, it is considered to have significant residual blur and should be given a lower confidence level. At the same time, local optical flow analysis can be used to determine whether the compensated texture still has unstructured dynamic fluctuations, such as frequent changes in optical flow direction (exceeding 35°) or abnormally increased optical flow amplitude (more than twice the stable value). By comprehensively calculating factors such as temporal consistency, spatial consistency, and frequency structure integrity, the confidence level of the compensation domain can be obtained.

[0025] Two weighting metrics from different sources—"principal component steady-state contribution" and "compensation domain confidence"—are normalized to generate a unified set of weighting coefficients. The core objective of normalization is to achieve uniform weight scale, controllable range, and smoothness between regions. When the principal component steady-state contribution and compensation domain confidence appear in different intervals (e.g., 0.2–0.9 or 0.1–0.8), direct use in weighted fusion can lead to structural bias. Therefore, linear normalization or a relative normalization method based on Softmax is required. Common processing methods include Min-Max normalization using the target interval 0–1 as the standard, and smoothing regularization of local regions, such as using 3×3 or 5×5 neighborhood mean filtering, to make the weights change continuously in space and avoid sudden jumps. For the steady-state contribution weights, nonlinear compression can be introduced, such as γ correction based on their energy magnitude (e.g., γ=0.8 or 1.2), so that high-energy textures receive slightly enhanced weights. For the confidence level of the compensation domain, a confidence suppression strategy can be adopted, that is, a lower limit on the weights of extremely low confidence regions (e.g., a minimum limit of 0.05) can be imposed to avoid the appearance of completely missing texture holes during the fusion process. The two normalized results are processed uniformly using a normalized sum method, such as w_main + w_low = 1, thus forming a paired set of weight coefficients, ensuring that both stable and compensated textures can obtain reasonable positions in the fusion. After obtaining the complete set of weighted coefficients, the principal component features of the texture and the low-stability texture features can be linearly fused to construct a textured, stable tongue image. The core idea of ​​linear fusion is to allocate weights at each pixel position, allowing the stable texture to dominate the main structure, while the compensated texture is used to fill in the details of the originally unstable regions. The fusion model typically adopts a pixel-wise linear combination form: T_final = w_main · T_main + w_low · T_low, where w_main and w_low come from the normalized set of weighted coefficients and satisfy w_main + w_low = 1. To maintain the natural continuity of the tongue's surface structure, neighborhood consistency constraints are incorporated during the fusion process. For example, a light structure-preserving filter is applied to the fused tongue image to smooth the edge transitions. Furthermore, in areas with prominent directional textures (such as cracks or particle orientation), a directional consistency correction method is used to further optimize the fusion result, ensuring that the directional information remains consistent with the actual tongue structure. For areas with significant differences in tongue coating thickness, local energy constraints are introduced to ensure that the texture density of thick coating areas remains realistic after fusion, while thin coating areas are not overly smoothed.

[0026] In this embodiment, the chromatographic analysis module is used to perform chromatographic analysis on the texture of a texture-stable tongue image, generating texture chromatographic analysis results, specifically for... Multi-channel color deconstruction is performed on textured tongue images to extract multi-color feature space; Adaptive threshold segmentation based on multi-color feature space is used to identify the tongue body and tongue coating. Calculate the mean hue and median saturation of the tongue body to generate the base color of the tongue body; The main color of the tongue coating is identified, and the main color feature of the tongue coating is generated; an 8×8 pixel unit scanning window is defined; Based on the 8×8 pixel unit scanning window, local texture color difference scanning calculation is performed on the multi-color feature space to generate local color deviation values; Multi-position texture chromatographic analysis was performed on the local color deviation values, the basic color of the tongue body, and the main color characteristics of the tongue coating to generate texture chromatographic analysis results.

[0027] In this embodiment, the tongue's color possesses complex physiological characteristics, including both the inherent color derived from the tongue's tissue and the surface color influenced by factors such as the degree of tongue coating, moisture content, and blood circulation. Therefore, a single color space cannot fully represent the tongue's color structure. Typically, textured tongue images are converted to multiple color spaces, including RGB, HSV, CIELab, and YCbCr, and combined to form a multi-color feature space. The HSV space is suitable for describing hue and saturation changes, the Lab space has good color difference consistency and can be used to capture subtle color changes, while the RGB space can serve as the original color reference. During the deconstruction process, the R, G, and B channels, H, S, and V channels, and L*, a*, and b* components are extracted from each pixel, forming a color description vector of at least 9 dimensions. To enhance the stability of color features, slight smoothing can be performed within a small neighborhood (e.g., 3×3 or 5×5) to reduce the impact of local noise on color analysis. Simultaneously, normalization operations can be performed on each color space to ensure consistency in the numerical scales of different channels. For example, by controlling the a* and b* values ​​in the Lab color space within the range of -128 to 127, and linearly stretching the S and V components in the HSV color space, subtle color differences in low-contrast regions can be expressed. After constructing the multi-color feature space, it is necessary to distinguish between the tongue body and tongue coating in the tongue surface region. The tongue coating and tongue body differ in hue, saturation, brightness, and Lab color difference, so these features can be used to implement adaptive thresholding. A common approach is to use the S and V components in the HSV color space to identify differences in tongue coating thickness. Typically, the tongue coating region has lower saturation and higher brightness; while in the Lab color space, the a* and b* component values ​​of the tongue coating are closer to the yellow or white regions, while the tongue body leans towards the red region. Otsu thresholding, adaptive local thresholding, or K-means-based color clustering methods can be used to combine multi-dimensional features for classification. For example, four-dimensional features (H, S, a*, b*) can be selected to construct a clustering vector, with K=2 used to separate the tongue body and tongue coating. To avoid color shifts caused by localized humidity, a local smoothing strategy can be employed, such as calculating the color mean of a 7×7 region and using it as the segmentation basis to reduce the impact of random noise. Furthermore, to prevent missegmentation, morphological processing is needed to correct the connectivity of the tongue coating region, such as using opening or closing operations to remove small-area false spots. Additionally, contour constraints can be used to correct the segmentation results; the tongue coating boundary must be continuous and consistent with the tongue body contour, and irregular protruding shapes are not allowed.

[0028] After segmenting the tongue region, it is necessary to calculate the base color of the tongue in that region. Tongue color often has strong physiological significance; for example, a pale color may indicate insufficient qi and blood, a reddish color may be related to heat, and a purplish color is often associated with blood stasis. Therefore, accurate base color of the tongue is crucial. H (hue) and S (saturation) in the HSV color space are usually chosen as the main color description indicators because they are more consistent with visual perception. The mean H value of all pixels in the tongue region is calculated to represent the overall hue direction of the tongue. Since there may be a few areas of abnormally high saturation due to local lighting effects, the median is usually used instead of the mean when calculating saturation to improve stability. To further enhance the representativeness of the color, H and S values ​​can be locally smoothed, for example, by weighting the average over a 5×5 neighborhood, to reduce color jumps caused by local bright reflections, fine cracks, or microvessels. Subsequently, the basic color of the tongue can be constructed into a composite feature vector, including H_mean, S_median, and, if necessary, the average L* value in Lab, to form a more complete tongue color index. For example, the basic color of the tongue can be represented as C_body = (H_mean, S_median, L_mean). This basic color will serve as a reference base color in local texture color difference scanning to determine whether there are abnormal fluctuations in the color of the tongue surface, such as whether reddish areas form dot-like clusters or whether purplish areas are distributed along cracks. Changes in tongue coating color are highly correlated with tongue diagnosis; for example, white, yellow, gray, and black coatings, as well as differences in thickness, are all important bases for judging cold, heat, deficiency, and excess. Therefore, after obtaining the tongue coating area, it is necessary to perform principal color identification on the tongue coating. This process usually uses Lab or HSV color space, especially Lab color difference space, which can more accurately describe the coating color bias. Color clustering (such as K-means or GMM) can be used to group the color distribution of the tongue coating area. For example, setting K=2 or K=3, the coating color can be classified into primary and secondary colors, and the color cluster with the largest area can be taken as the primary color of the tongue coating. Tongue coating thickness information can also be incorporated into primary color identification; for example, low-brightness and low-saturation areas tend to have thicker coatings, which can be corrected by combining with L* values. After obtaining the primary color features, the structure of the scanning window needs to be defined. An 8×8 pixel unit is usually used as the scanning window size because this scale can cover a local coating texture unit without losing too much detail; in tongue images with a resolution of approximately 720p or 1080p, an 8×8 unit can balance spatial resolution and local computational efficiency. When the scanning window slides across the tongue surface, its step size can be set to 4 pixels or 8 pixels, thus forming a color difference grid scan with moderate density. After defining the scanning window, a local color difference scan of the entire tongue surface needs to be performed to calculate the color deviation value of each window area.The core purpose of local color difference analysis is to identify the differences between the colors of various parts of the tongue surface and the base color of the tongue body and the main color of the tongue coating, in order to analyze features such as local pathological pigmentation, moisture accumulation, or petechiae. Specifically, each 8×8 unit is treated as a sampling block, and the Lab values ​​(L*, a*, b*) or HSV values ​​of its internal pixels are statistically analyzed. For example, the mean (average color), standard deviation (color dispersion), and color difference ΔE between the pixels within the window and the reference color (base color of the tongue body or main color of the tongue coating) are calculated. ΔE can be calculated using the CIEDE2000 model to enhance the ability to capture color differences sensitive to the human eye. During the scanning process, each window will generate a local color deviation value, including brightness deviation, hue deviation, and color saturation deviation. For example, if a window area is bright and highly saturated, it may represent a moist area of ​​the tongue surface; if it is purplish or dark, it may represent an area related to local blood stasis or cold symptoms. To avoid the influence of noise, the color difference between consecutive windows will be smoothed. For example, when the color difference between adjacent windows exceeds a certain threshold (such as ΔE>6), the boundary jump can be reduced by interpolation.

[0029] After obtaining the local color deviation values, they need to be integrated with the basic color of the tongue body and the main color characteristics of the tongue coating. A complete tongue surface chromatographic structure is generated through multi-position texture chromatographic analysis. The core idea of ​​this analysis process is to map the color changes of the tongue surface into a quantifiable chromatographic model, thereby revealing the spatial variation trend, regional characteristics, and potential pathological significance of the tongue surface color. By mapping the local color difference matrix to the corresponding regional labels (tongue body or tongue coating), two-dimensional chromatographic modeling is performed: the tongue body chromatogram mainly reflects the color distribution related to blood color and the circulation of qi and blood; the tongue coating chromatogram reflects the coverage state of pathological factors such as dampness, heat, cold, and dryness. The chromatographic trend can be described using color distribution histograms, two-dimensional color difference grids, and local color difference gradient fields. When integrating local color differences, areas that deviate far from the basic color of the tongue body (e.g., ΔE>10) are marked as "abnormal color gamuts," and further determined to be redder, purpler, or darker based on the hue direction. For the tongue coating area, it is classified into types such as "uniform coating color", "local dark coating", "local light coating", and "mixed coating color" according to the degree of deviation from the main color. In addition, by combining the position of the scanning window, a tongue surface topological chromatogram can be constructed to statistically analyze the color difference trends of the tongue tip, middle, root, and sides. For example, a reddish tongue tip may indicate heat, while a purplish tongue side may indicate liver stagnation and blood stasis.

[0030] In this embodiment, the partitioning module is used to perform tongue body partitioning recognition based on texture-stabilized tongue images and output a tongue surface partitioning map, specifically for: Calculate tongue tip curvature and identify tongue root color change points based on texture-stabilized tongue images; Based on the curvature of the tongue tip and the color change point of the tongue root, the length of the tongue body is divided into three equal parts longitudinally to determine the tongue tip region, the middle region, and the tongue root region. Based on the texture-stabilized tongue image, the tongue midline is identified and marked. The tongue is divided into horizontal sections based on the midline of the tongue, resulting in the left and right sides of the tongue. Perform contour line analysis on the left and right sides of the tongue, the tip of the tongue, the middle of the tongue, and the root of the tongue, and output a tongue surface partition map.

[0031] In this embodiment, edge extraction is performed on the overall contour of the tongue using Canny edge detection or Sobel gradient directional methods to obtain the outer contour of the tongue. The tip of the tongue is usually located at the front of the tongue and its geometric features are characterized by a high-curvature tip structure. Therefore, the curvature κ(s) of the anterior contour curve can be calculated, where s is the contour arc length. The curvature can be calculated using the three-point estimation method or the second derivative method. The peak curvature point usually corresponds to the tip of the tongue. In multi-frame stable images, this point has high consistency, and the curvature value is usually higher than other parts of the tongue (e.g., more than 2-3 times the average value). On the other hand, the identification of the tongue root region can be completed based on the color change points. The tongue root is close to the pharynx and often presents a region with rapid color transitions due to light occlusion and changes in tissue structure, especially in Lab space where the L* and a* values ​​will show obvious jumps. For example, L* may suddenly drop by 10-20 units at the tongue root, or the a* value may show a rapid transition from yellow to dark red. By analyzing the longitudinal color profile, the most obvious color difference abrupt change points ΔE peak areas (e.g., locations where ΔE>12) can be identified. The direction of the tongue's main axis is determined, typically by connecting the tip of the tongue to the color abrupt change point at the root of the tongue as the longitudinal reference axis. The effective length of the tongue is measured along this direction, using Euclidean distance or arc length measurement methods. Once the main axis is obtained, the tongue length can be divided into three equidistant segments in a 1:1:1 ratio. The first segment (the foremost segment) is defined as the tip region, corresponding to the "cardiopulmonary reflection area" in tongue diagnosis; the middle segment is the mid-tongue region, corresponding to the spleen and stomach function reflection area; and the last segment is the root region, often associated with factors such as the kidneys, bladder, and lower abdominal damp-heat. To make the region boundaries more closely resemble the actual tongue surface structure, local projection corrections are made to the straight-line partition boundaries based on the tongue's edge contour. For example, the straight-line anchor points are mapped along the nearest points on the left and right edges of the tongue, thus constructing natural partition boundaries. Simultaneously, the tongue surface texture gradient can be utilized to ensure that the boundaries do not cross areas with excessive structural abrupt changes (such as areas with dense cracks). The region division results will be recorded in the tongue structure matrix in the form of region labels.

[0032] Tongue midline recognition can be based on a comprehensive assessment of multi-dimensional features such as geometric symmetry, brightness symmetry, and texture direction consistency. The tongue region is projected horizontally to find the global axis of symmetry. This can be achieved using the minimum pixel difference method, which calculates the color and texture differences between each column of pixels and its mirror image; the column with the smallest difference is the approximate midline of the tongue. Further refinement is then achieved using texture direction fields (e.g., Gabor direction response or gradient field direction distribution), ensuring the directional difference between the two sides of the tongue midline is within a reasonable range (e.g., direction difference <10°). Additionally, brightness distribution characteristics can be used to assist in recognition; the area near the tongue midline is typically slightly brighter and has symmetrical shadow distribution due to the slight bulge of the tongue surface. Combining these methods, candidate regions for the tongue midline can be constructed, and then optimized using smoothing filters or Bayesian models to obtain a continuous and natural tongue midline. To avoid interference from cracks, crack masks can be used to protect areas with deep cracks, preventing the tongue midline from being misled by crack morphology. After identifying the tongue midline, it can be used as a reference for lateral partitioning, dividing the tongue surface into left and right tongue side regions. In Traditional Chinese Medicine (TCM) tongue diagnosis, the edges of the tongue are closely related to the liver and gallbladder system, Qi circulation, and the risk of blood stasis. Therefore, accurately distinguishing the tongue edge region is crucial. The basic method of horizontal partitioning is to use the tongue midline as the boundary in a tongue mask, defining all pixels to the left as the left tongue edge region and pixels to the right as the right tongue edge region. However, to make the partitioning more consistent with the actual shape of the tongue, it is necessary to optimize it by considering the left-right width variations of the outline. For example, the left-right width may differ in different parts of the tongue, with the tip being sharper and the root wider. Therefore, the width of each row of pixels can be measured vertically along the tongue midline, and the actual width of the tongue edge region can be determined based on a distance threshold (e.g., the tongue edge region can be defined as occupying 20-30% of the width of each row). This constructs left and right tongue edge regions that more closely resemble the actual structure of the tongue. Furthermore, to avoid unreasonable partitioning due to local offset of the tongue midline, the horizontal boundaries can be smoothed locally, such as averaging the boundaries of adjacent 5-10 rows to make the partition lines more natural. The partition contour analysis includes detecting, smoothing, closing, and labeling the boundaries of each region to present a clear and structured partition shape. Edge contours of each region are extracted using region boundary tracing algorithms, such as Moore-Neighbor tracing or SLIC superpixel-based region boundary extraction methods. When smoothing the contours, Gaussian smoothing or spline interpolation can be used to make the region boundaries more consistent with the natural shape of the tongue, avoiding jagged, discontinuous structures. Subsequently, unique color or marker codes are assigned to the tongue tip, middle, and root regions, as well as the left and right sides of the tongue (only for structural map representation, not for image coloring purposes), thus constructing a complete tongue surface structure map.To enhance the readability of the structural diagram in medical analysis, region labels can be added to the diagram, such as "tongue tip region", "tongue side region", "tongue middle region", "tongue root region", etc., so that the tongue diagnosis model can extract local features according to the region in pathological inference, such as whether the tongue tip is reddish, whether the tongue side is purplish, and whether the tongue root has a thick coating.

[0033] In this embodiment, the color difference analysis module is used to perform zonal color difference analysis on the tongue surface zonal map based on the texture chromatographic analysis results, and generate a zonal difference feature set, specifically for: Based on the texture chromatographic analysis results, the tongue surface partition map is localized using texture chromatographic analysis, thereby generating texture chromatographic information for multiple regions; The texture chromatographic information is partitioned to identify texture differences, and inter-region texture difference features are generated. Abnormal feature detection is performed on the tongue surface partition map, and abnormal color points in different tongue surface partitions are marked; the abnormal color points include petechiae, ecchymosis, and cracks; A partitioned color difference perception assessment is performed on the texture difference features between regions and the abnormal color points to generate a partitioned difference feature set.

[0034] In this embodiment, after completing the texture chromatogram analysis and obtaining the corresponding full-tongue texture color distribution, the texture chromatogram information needs to be spatially aligned with the tongue surface partition map to obtain the texture chromatogram information corresponding to different partitions. The tongue surface partition map has clearly located the tip, middle, root, left, and right sides of the tongue, while the texture chromatogram analysis map provides a composite result of local color difference features, main texture color, basic tongue color, and main tongue coating color for each location. Aligning the texture chromatogram with the tongue surface partition map at the pixel level allows for direct mapping based on the row and column consistency between the partition matrix and the texture matrix. Subsequently, the texture chromatogram values ​​are aggregated and analyzed within each partition. For example, key indicators such as the partition mean color difference, texture direction consistency, and local color shift can be statistically analyzed to ensure that each region has a unique chromatogram structure description. To make the positioning more stable, a 3×3 or 5×5 pixel local smoothing window can be used to slightly smooth the texture chromatogram values ​​within the partition, reducing the impact of noise on the partition edges. The distribution of characteristic color bands is calculated for different regions. For example, the tip of the tongue may show an increase in red components, while the coating at the root of the tongue may show an increase in yellow or white components. By statistically analyzing the probability of each color band appearing within a region, the textural chromatographic information of different regions is quantified. After obtaining the textural chromatographic information of the tip, middle, root, and left and right sides of the tongue, further textural differences between regions can be identified to analyze structural differences in color, texture coarseness, moisture levels, color difference changes, and directional consistency. By comparing the textural chromatographic statistical indicators of each region, including the regional average ΔE color difference, regional hue shift, saturation change rate, and differences in the main direction of texture distribution, for example, by comparing the textural chromatographic information of the tip and root of the tongue, it is possible to identify whether the tip of the tongue has a reddish tendency; by comparing the left and right sides of the tongue, it is possible to identify whether there is a purple or bluish tendency, thus assisting in the judgment of pathological manifestations such as qi stagnation and blood stasis. Secondly, the statistical distance between regions can be further calculated, such as using Euclidean distance, Mahalanobis distance, or Bhattacharyya distance to determine the intensity of differences in textural chromatographic distribution, making the difference features not only directional but also quantitative. For data that may exhibit characteristics such as thickened tongue coating, rough texture, and fluctuating humidity, multidimensional differences can be measured by comprehensively considering texture roughness, brightness fluctuation range, and density of texture details on the tongue surface. For example, if the degree of disorder in the direction of the tongue coating increases by more than 20% between the middle and root areas, it can be recorded as a texture difference.

[0035] To further enhance the pathological specificity of tongue diagnosis analysis, it is necessary to detect abnormal features in the tongue surface zoning map, focusing on marking key abnormal color points such as petechiae, ecchymoses, and fissures. A detailed scan of color brightness, saturation, and hue is performed in the stable tongue image, using abnormal color point detection methods in HSV or Lab space. For example, petechiae often appear as small, dot-like structures with reduced local saturation and a hue bias towards purple or dark red, which can be detected by statistically analyzing the color difference peaks within a 3×3 or 5×5 pixel window. Ecchymoses are usually larger and can be identified through connected component analysis of local low-brightness, dark red areas exceeding a certain threshold (e.g., more than 20 pixels). Fissures appear as thin, high-contrast linear structures, which can be extracted using gradient direction consistency analysis or linear structure enhancement methods based on the Hessian matrix. After identification, all abnormal color points are mapped back to the tongue surface zoning map, marking their location as the tip, side, middle, or root of the tongue. To avoid mislabeling, confidence scores can be assigned to outliers, such as assigning a confidence score of 0–1 to each outlier based on factors like color difference amplitude, connectivity, and shape characteristics. After obtaining the texture difference features between regions and the abnormal color points, a unified color difference perception assessment needs to be performed on these features to generate a regional difference feature set. The core objective of this assessment process is to integrate the texture differences between regions with the degree of color shift of local outliers, thereby making a more accurate regional judgment on the overall pathological state of the tongue surface. The texture difference matrix between regions is fused with the outlier information. For example, if the average ΔE texture color difference of a certain region is higher than that of other regions (e.g., exceeding 8–10 units), and the region has dense petechiae, purplish patches, or deep cracks, then the pathological perception level of that region can be assessed as high. Secondly, a regional score is constructed based on the distribution density of abnormal color points and the amplitude of regional texture differences. For example, the outlier contribution of a region can be calculated based on information such as the number of outliers, connectivity, average color difference, and saturation shift. Furthermore, by combining the normal texture deviation characteristics of the region itself, such as the tip of the tongue being slightly reddish which is normal, and the root of the tongue being slightly less bright which is normal, corrections are made through rules or models to make the assessment results more consistent with the rules of tongue diagnosis.

[0036] In this embodiment, the tongue diagnosis analysis module is used to perform pathological analysis based on the partition difference feature set and output an intelligent tongue diagnosis analysis report, specifically for: The feature vectors of the partition difference feature set are standardized to generate a standardized vector set. Based on a pre-defined theoretical tongue pathology model, pathological similarity is calculated on a standardized vector set to generate similarity scores for various pathological templates. Confidence is calculated based on the similarity score to generate the pathological symptom with the highest confidence. Based on the pathological signs with the highest confidence level, pathological analysis is performed, and an intelligent tongue diagnosis analysis report is output.

[0037] In this embodiment, after obtaining the regional difference feature set, it is necessary to perform feature vector standardization processing to enable the pathological model discrimination process to be analyzed on a uniform scale. The regional difference feature set usually contains information in multiple dimensions, including regional average color difference ΔE, abnormal color point density, crack density, local texture direction deviation, saturation shift, brightness stability, and other numerical features. The units, orders of magnitude, and distribution ranges of each feature differ, therefore a unified normalization process is necessary. Standardization can employ z-score standardization, min-max linear normalization, or interval compression methods based on robust statistics; the specific method can be selected according to the distribution of the features. For example, for tongue brightness difference features exhibiting an approximately normal distribution, z-score form can be used to transform it into a standardized variable with a mean of 0 and a standard deviation of 1; while for features with strong skewed distributions, such as abnormal point density, logarithmic transformation or quantile normalization can be used to make its distribution more stable. To ensure the comparability of feature vectors from different regions in spatial coordinates, a unified feature arrangement needs to be constructed according to the region order, such as arranging the tip, middle, root, left side, and right side of the tongue in sequence, giving the overall vector structure a fixed dimension. After obtaining the standardized vector set, it needs to be input into a pre-defined theoretical tongue pathology model for similarity calculation to evaluate its matching degree with various TCM tongue diagnosis pathological conditions. The theoretical tongue pathology model is usually constructed through expert annotation, tongue image databases, tongue quantitative indicators, and typical case statistics, including various common tongue diagnosis pathological patterns such as Qi deficiency, blood stasis, damp heat, Yin deficiency, phlegm dampness, blood stasis, and Yang deficiency. Each pathological pattern has a corresponding feature vector template. For example, Qi deficiency may manifest as pale tongue edges, high brightness, and fewer abnormal points; blood stasis is usually accompanied by a high purplish-dark color deviation and increased density of petechiae. Similarity calculation can use Euclidean distance, cosine similarity, Mahalanobis distance, or nonlinear similarity measures based on kernel functions. For example, if cosine similarity is used, a value closer to 1 indicates greater similarity; if Euclidean distance is used, a smaller value indicates closer proximity to the pathological template. To improve the stability of the discrimination, a weight matrix can be established so that different features contribute differently to certain pathological patterns. For example, in the blood stasis model, the feature weight for purplish color can be set to 0.4, while the feature weight for cracks can be 0.1, to accommodate the characteristics of different pathologies.

[0038] After obtaining similarity scores from multiple pathological templates, further confidence calculations are needed to determine the most likely pathological signs. Confidence calculations consider not only the numerical value of the similarity score itself but also a comprehensive analysis of factors such as the interrelationships between different pathological patterns, the dispersion of the similarity distribution, and the contribution of each feature dimension. For example, if the similarity of a certain pathological template is higher than that of other templates, and its peak similarity exceeds a set threshold (e.g., 0.75), this pattern can be directly marked as a high-confidence pathological trend. If the similarity of multiple pathological patterns is close, a confidence normalization method is needed for synthesis, such as using the softmax function to convert the similarity into a probability distribution, so that pathological patterns with higher similarity receive higher confidence. Furthermore, an uncertainty assessment module can be established, such as calculating the variance of the similarity scores. When the variance is small and the scores of different pathological patterns are not significantly different, it can indicate that the current tongue image has atypical manifestations, requiring more cautious conclusions. After determining the highest-confidence pathological sign, a systematic pathological analysis of the entire tongue image is performed based on this, generating an intelligent tongue diagnosis analysis report. The analysis process includes interpreting the correspondence between pathological signs and tongue features. For example, when the system identifies "blood stasis," it will point out typical manifestations such as an increased tendency towards purplish-darkness on the tongue surface, higher density of petechiae, and concentrated color differences in the tongue edges. When identified as "damp-heat," it will show a yellowish tongue coating, increased roughness of the tongue coating in the central area, and decreased brightness stability. The report not only provides the pathological type but also needs to include a tongue surface structure diagram, chromatographic analysis of the texture of each region, a schematic diagram of the distribution of abnormal color points, and even list the quantitative values ​​of key indicators, such as ΔE=14.2 in the tongue root region, an 18% increase in petechiae density on the tongue edges, and higher saturation of the tongue coating color in the central area, making the report interpretable. In addition, the pathological signs can be described in relation to corresponding theories in traditional Chinese medicine, such as blood stasis being related to qi stagnation, cold coagulation, and pain risk, and damp-heat being related to excessive dampness in the body and excessive gallbladder heat.

[0039] In this embodiment, the specific steps for performing pathological analysis based on the highest confidence pathological signs and outputting an intelligent tongue diagnosis analysis report are as follows: The highest confidence level pathological signs are analyzed for syndrome attributes to generate pathological syndrome characteristics. Based on the characteristics of pathological syndromes, potential predominance is predicted, and potential pathological data is generated; the potential pathological data includes the direction of predominance, the quantitative score of the degree of predominance, and the expected evolution time window. Personalized treatment suggestions are made based on potential pathological data, and a treatment plan is generated; based on the treatment plan and potential pathological data, a comprehensive statistical analysis is performed to output an intelligent tongue diagnosis report.

[0040] In this embodiment, after obtaining the pathological signs with the highest confidence level, it is necessary to perform in-depth syndrome attribute analysis to generate quantifiable pathological syndrome characteristics. Syndrome attribute analysis involves correlating information such as the texture, color, and distribution of abnormal points in different areas of the tongue surface with traditional Chinese medicine pathology theories, further refining a single pathological sign into multiple measurable feature dimensions. For example, for blood stasis signs, the analysis can be performed on attributes such as the degree of purplish-darkness of the tongue, the density of petechiae, the length of cracks, the distribution area, and the difference between the tongue edges and the center of the tongue. Each attribute is expressed numerically, such as a purplish-dark color deviating from the basic tongue color by ΔE=12.5, a petechiae density of 15 per square centimeter, and a total crack length of 18 mm. For damp-heat signs, the analysis can be performed on attributes such as the yellowness, thickness, brightness variation, and regional distribution range of the tongue coating, such as a tongue root coating color that is more than 20% yellowish, a decrease in brightness in the center of the tongue by ΔL=8.2, and a coating roughness index of 0.45. By summarizing these attributes into a unified feature vector, a pathological syndrome feature set can be formed, including both local indicators and global trends. Each attribute can be weighted according to its pathological nature; for example, color shift has a higher weight for identifying blood stasis, and tongue coating thickness has a higher weight for identifying damp-heat. After obtaining the pathological syndrome features, potential predominance prediction can be further conducted to infer the direction and trend of pathological changes reflected in the tongue appearance. Potential predominance prediction comprehensively considers the time-series characteristics of each region's features, inter-regional differences, and known TCM pathological evolution patterns. For example, indicators such as historical tongue appearance change trends, regional color change rates, and texture roughness changes can be used, combined with syndrome attribute features, to determine the direction of predominance, i.e., whether it is predominance due to heat, cold, blood stasis, dampness, or phlegm-dampness. The degree of predominance is scored using feature quantification methods, usually expressed in a range of 0–1 or 0–100, such as 0.68 for blood stasis and 0.55 for damp-heat, reflecting the current pathological load. Furthermore, by combining dynamic models of tongue surface evolution or statistical regression models, the development speed of potential pathologies can be predicted within a time window. For example, it can be predicted that the tongue appearance may change in the next 7–10 days, such as darkening of color, increase of petechiae, or thickening of the tongue coating.

[0041] After generating potential pathological data, personalized treatment plans can be developed based on the direction, degree, and time window of the predominance. These plans combine traditional Chinese medicine theory with modern health management principles, transforming pathological predictions and tongue characteristics into actionable treatment recommendations. Specific actions include recommending corresponding dietary adjustments, lifestyle modifications, and traditional Chinese medicine or dietary interventions based on the predominance type. For example, for users with a high score indicating predominance of blood stasis, the treatment plan might include increasing blood-activating and stasis-removing dietary components, moderate exercise, and maintaining a healthy sleep schedule; for those with predominance of damp-heat, the emphasis is on clearing heat and dampness, controlling high-sugar and oily foods, and maintaining a balanced tongue surface moisture. The development of treatment plans also needs to consider the time window of potential pathological evolution; for example, areas that may worsen within the next 7 days can be prioritized for targeted measures. The quantitative score of the potential pathological data is converted into treatment intensity and frequency; for example, a predominance degree within the range of 0.6–0.8 can be defined as moderate intervention, while a degree exceeding 0.8 indicates intensive intervention, ensuring that the plan is both scientific and personalized. After obtaining a personalized treatment plan, it needs to be combined with potential pathological data for comprehensive statistical analysis to generate a complete intelligent tongue diagnosis report. This comprehensive statistical analysis includes summarizing the potential pathological scores, evolution trends, and treatment measures for each area, integrating tongue surface area differences, abnormal color points, pathological syndrome characteristics, potential predominance predictions, and treatment plans into a structured output. The report typically includes an overall pathological assessment, a summary of abnormal characteristics in each area, the direction and degree of potential predominance, the expected evolution time window, treatment suggestions, and priority ranking. For example, if the tip of the tongue is reddish with a damp-heat score of 0.62 and the edge of the tongue has a blood stasis score of 0.71, a comprehensive analysis text such as "Damp-heat predominates at the tip of the tongue; attention should be paid to clearing heat and promoting diuresis; the tongue tip coating may thicken in the next week; blood stasis occurs at the edge of the tongue; treatment to promote blood circulation and remove blood stasis should be increased" can be generated. The report can provide quantitative charts, such as tongue surface area heatmaps, abnormal point density distribution maps, and pathological score trend curves, enabling users or physicians to intuitively understand the relationship between the tongue appearance and potential pathological states.

[0042] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.

[0043] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein are implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. An artificial intelligence-based auxiliary system for tongue diagnosis in Traditional Chinese Medicine, characterized in that, The AI-based TCM tongue diagnosis auxiliary system includes an image processing module, a chromatographic analysis module, a zoning module, a color difference analysis module, and a tongue diagnosis analysis module. The image processing module is used to acquire user tongue surface recognition video; perform image processing on the user tongue surface recognition video to construct a texture-stabilized tongue image; The chromatographic analysis module is used to perform chromatographic analysis on textured tongue images with stable textures and generate chromatographic analysis results. The partitioning module is used to perform tongue body partitioning recognition based on texture-stabilized tongue images and output a tongue surface partitioning map. The color difference analysis module is used to perform zonal color difference analysis on the tongue surface zonal map based on the texture chromatographic analysis results, and generate a zonal difference feature set; The tongue diagnosis analysis module is used to perform pathological analysis based on the partition difference feature set and output an intelligent tongue diagnosis analysis report.

2. The artificial intelligence-based TCM tongue diagnosis auxiliary system according to claim 1, characterized in that, The image processing module is used to acquire user tongue surface recognition video; and to perform image processing on the user tongue surface recognition video to construct a texture-stabilized tongue image, specifically for: Acquire a video of user tongue surface recognition; decompose the user's tongue surface recognition video into image frames to generate a tongue surface image frame sequence; The local luminance reflectance of each frame is calculated by performing frame-by-frame local luminance reflectance calculation on the tongue surface image frame sequence to obtain the local luminance reflectance of each frame. Brightness compensation is performed based on the local brightness reflectivity to output a brightness-stable image sequence. Texture perturbation compensation is performed on brightness-stable image sequences to construct texture-stable tongue images.

3. The artificial intelligence-based TCM tongue diagnosis auxiliary system according to claim 2, characterized in that, The specific steps for performing texture perturbation compensation on the brightness-stable image sequence to construct a texture-stable tongue image are as follows: Calculate the multi-frame texture gradient change value of each pixel based on a brightness-stable image sequence; Based on the multi-frame texture gradient change values, regions are divided to obtain stable and unstable texture regions. Texture components are identified and extracted from stable texture regions to obtain principal component features of the texture. Interference analysis is performed on the unstable region to extract texture interference information; the texture interference information includes transient reflection noise and humidity disturbance components; Perturbation compensation is performed based on the texture interference information to obtain low-stability texture features; We construct a textured tongue image by weighted reconstruction of the principal component features and low-stability texture features.

4. The artificial intelligence-based TCM tongue diagnosis auxiliary system according to claim 2, characterized in that, The specific steps for constructing a textured, stable tongue image by weighted reconstruction of principal component features and low-stability texture features are as follows: Calculate the energy distribution of the principal component features of the texture; calculate the steady-state contribution of the principal components based on the energy distribution; The perturbation residual amount of the low-stability texture features is evaluated to obtain the confidence level of the compensation domain; The steady-state contribution of the principal components and the confidence level of the compensation domain are normalized to generate a set of weighted coefficients; Linear fusion is performed based on the weighted coefficient set to construct a texture-stable tongue image.

5. The artificial intelligence-based TCM tongue diagnosis auxiliary system according to claim 1, characterized in that, The chromatographic analysis module is used for chromatographic analysis of textured, stable tongue images, and generating texture chromatographic analysis results. Specifically, it is used for... Multi-channel color deconstruction is performed on textured tongue images to extract multi-color feature space; Adaptive threshold segmentation based on multi-color feature space is used to identify the tongue body and tongue coating. Calculate the mean hue and median saturation of the tongue body to generate the base color of the tongue body; The main color of the tongue coating is identified, and the main color feature of the tongue coating is generated; Define an 8×8 pixel unit scanning window; Based on the 8×8 pixel unit scanning window, local texture color difference scanning calculation is performed on the multi-color feature space to generate local color deviation values; Multi-position texture chromatographic analysis was performed on the local color deviation values, the basic color of the tongue body, and the main color characteristics of the tongue coating to generate texture chromatographic analysis results.

6. The artificial intelligence-based TCM tongue diagnosis auxiliary system according to claim 1, characterized in that, The partitioning module is used for tongue body partition recognition based on texture-stabilized tongue images and outputs a tongue surface partition map, specifically for: Calculate tongue tip curvature and identify tongue root color change points based on texture-stabilized tongue images; Based on the curvature of the tongue tip and the color change point of the tongue root, the length of the tongue body is divided into three equal parts longitudinally to determine the tongue tip region, the middle region, and the tongue root region. Based on the texture-stabilized tongue image, the tongue midline is identified and marked. The tongue is divided into horizontal sections based on the midline of the tongue, resulting in the left and right sides of the tongue. Perform contour line analysis on the left and right sides of the tongue, the tip of the tongue, the middle of the tongue, and the root of the tongue, and output a tongue surface partition map.

7. The artificial intelligence-based TCM tongue diagnosis auxiliary system according to claim 1, characterized in that, The color difference analysis module is used to perform zonal color difference analysis on the tongue surface zonal map based on the texture chromatographic analysis results, and generate a zonal difference feature set, specifically for: Based on the texture chromatographic analysis results, the tongue surface partition map is localized using texture chromatographic analysis, thereby generating texture chromatographic information for multiple regions; The texture chromatographic information is partitioned to identify texture differences, and inter-region texture difference features are generated. Abnormal feature detection is performed on the tongue surface partition map, and abnormal color points in different tongue surface partitions are marked; the abnormal color points include petechiae, ecchymosis, and cracks; A partitioned color difference perception assessment is performed on the texture difference features between regions and the abnormal color points to generate a partitioned difference feature set.

8. The artificial intelligence-based TCM tongue diagnosis auxiliary system according to claim 1, characterized in that, The tongue diagnosis analysis module is used to perform pathological analysis based on the partition difference feature set and output an intelligent tongue diagnosis analysis report, specifically for: The feature vectors of the partition difference feature set are standardized to generate a standardized vector set. Based on a pre-defined theoretical tongue pathology model, pathological similarity is calculated on a standardized vector set to generate similarity scores for various pathological templates. Confidence is calculated based on the similarity score to generate the pathological symptom with the highest confidence. Based on the pathological signs with the highest confidence level, pathological analysis is performed, and an intelligent tongue diagnosis analysis report is output.

9. The method according to claim 8, characterized in that, The pre-defined theoretical tongue pathology model specifically includes: multiple standard tongue image templates, multiple versions of tongue image reference sets, syndrome logic rules, and basic classification standards for tongue body and tongue coating.

10. The artificial intelligence-based TCM tongue diagnosis auxiliary system according to claim 8, characterized in that, The specific steps for performing pathological analysis based on the highest confidence level pathological signs and outputting an intelligent tongue diagnosis analysis report are as follows: The highest confidence level pathological signs are analyzed for syndrome attributes to generate pathological syndrome characteristics. Based on the characteristics of pathological syndromes, potential predominance is predicted, and potential pathological data is generated. The potential pathological data includes the direction of predominance, the quantitative score of the degree of predominance, and the expected evolution time window; Personalized treatment suggestions are made based on potential pathological data, and treatment plans are generated. Based on the treatment plan and potential pathological data, a comprehensive statistical analysis is performed to output an intelligent tongue diagnosis report.