Multi-spectrum-based traditional Chinese medicine tongue diagnosis image enhancement system
By using multispectral image acquisition and color calibration elements, combined with color complexity modeling and structural texture analysis, and utilizing semantic segmentation networks for precise segmentation and intelligent annotation of the tongue region, this approach solves the problems of insufficient segmentation accuracy and difficulty in identifying structural abnormalities in traditional Chinese medicine tongue diagnosis. It achieves high-precision tongue image feature extraction and intelligent annotation, thereby improving the objectivity and intelligence of diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-10
AI Technical Summary
Traditional Chinese medicine tongue diagnosis is easily affected by ambient lighting, equipment performance, and individual differences. It lacks precision in tongue region segmentation, has limited ability to represent color information, low accuracy in identifying subtle structural abnormalities, and lacks multimodal information fusion, making it difficult to achieve high-precision tongue image feature extraction and intelligent annotation.
By employing multispectral image acquisition combined with color calibration elements, color complexity modeling and structural texture analysis are used to accurately segment and intelligently label the tongue region using a semantic segmentation network. This is combined with temporal evolution trend analysis and non-image diagnostic information for comprehensive judgment.
It improves the accuracy of tongue region segmentation, enhances color representation capabilities, improves the recognition accuracy of structural abnormalities such as cracks and teeth marks, realizes intelligent labeling of pixel-level abnormal regions, and improves the objectivity and intelligence of TCM tongue diagnosis.
Smart Images

Figure CN121838152A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent traditional Chinese medicine, and more specifically to a multispectral image enhancement system for tongue diagnosis in traditional Chinese medicine. Background Technology
[0002] Tongue diagnosis, a crucial diagnostic tool in traditional Chinese medicine, assesses a person's health by observing changes in the tongue's color, texture, and shape. It boasts a long history and rich clinical experience. However, traditional tongue diagnosis methods heavily rely on the doctor's subjective experience and are easily affected by factors such as ambient lighting, equipment performance, and individual differences, leading to inconsistent and inaccurate diagnostic results. With the development of computer vision and artificial intelligence technologies, image processing-based auxiliary systems for TCM tongue diagnosis have become a research hotspot. Existing technologies, while some systems attempt to extract tongue features through image acquisition and simple analysis, still have many shortcomings. For example, precise segmentation of the tongue region is often affected by complex background interference; the ability to represent color information is limited due to a lack of multispectral support; and the accuracy in identifying subtle structural abnormalities such as cracks and teeth marks is low. Furthermore, existing systems lack multi-source information fusion mechanisms for pixel-level segmentation and intelligent annotation of abnormal areas, resulting in insufficient confidence and clinical reference value of the annotation results. Simultaneously, due to the lack of temporal evolution trend analysis and multimodal feature fusion capabilities, these systems struggle to comprehensively integrate non-image diagnostic information from patients for integrated judgment, limiting their effectiveness in actual clinical practice. Therefore, developing a multispectral TCM tongue diagnosis image enhancement system that can overcome the above-mentioned defects, so as to achieve high-precision tongue image feature extraction, intelligent labeling of abnormal areas and multimodal information fusion, is of great significance for improving the objectivity, standardization and intelligence of TCM tongue diagnosis. Summary of the Invention
[0003] To address the shortcomings of existing technologies, the present invention aims to provide a multispectral TCM tongue diagnosis image enhancement system, which has the advantages of improving the accuracy of tongue region segmentation, enhancing color representation capabilities, improving the accuracy of structural anomaly identification such as cracks and tooth marks, and achieving pixel-level anomaly region segmentation and intelligent annotation.
[0004] To achieve the above objectives, the present invention provides the following technical solution: A multispectral-based image enhancement system for traditional Chinese medicine tongue diagnosis, comprising: The image acquisition module is used to acquire tongue image information, which includes the tongue body area and color calibration elements. The tongue extraction and standardization module is used to extract the tongue region from the tongue image information and perform geometric standardization based on the tongue contour. The color complexity modeling module is used to perform color space transformation on the tongue region and divide it into multiple sub-regions. Based on the color distribution dispersion and brightness abrupt change of each sub-region, the color complexity parameters are calculated. The structural texture analysis module is used to extract the texture features of each block in the tongue region, including gray-level co-occurrence matrix features, directional gradient information and high-frequency response indicators, to identify structural anomalies such as cracks and tooth marks. The abnormal region segmentation module is used to fuse color complexity parameters with structural texture features and input them into the semantic segmentation network to generate pixel-level segmentation masks for suspected abnormal regions in tongue image information. The intelligent annotation module for abnormal regions is used to classify the attributes of each abnormal region in the segmentation mask and overlay the annotations onto the original image.
[0005] Furthermore, this application also proposes a color space conversion unit for converting tongue image information into HSV color space and performing color calibration based on color calibration elements; Tongue segmentation unit, used to divide the tongue region into multiple equal-scale segmented regions; The color distribution feature extraction unit is used to statistically analyze pixel hue differences, brightness gradient distribution, and saturation range in each block region. The color complexity parameter generation unit is used to fuse color distribution features to generate a representation of the degree of color dispersion in the region.
[0006] Furthermore, this application also proposes a gray-scale statistics submodule for constructing the gray-scale co-occurrence matrix of each block region and extracting energy, contrast and entropy values; The multi-scale response submodule is used to filter the image based on multiple orientation and scale parameters to obtain a local orientation texture enhancement map. The crack morphology extraction submodule is used to refine and analyze the direction of continuous linear low-brightness structures and extract target areas that conform to crack morphology characteristics. The structural feature fusion unit is used to fuse grayscale statistical features and texture response features into a unified structural texture description vector.
[0007] Furthermore, this application also proposes a spectral saliency assessment unit, which is used to perform saliency analysis on the response characteristics of candidate structural regions under multiple spectral channels, calculate the structural saliency score for each channel, and assign differentiated confidence levels to the structural features output by different channels based on the saliency scores.
[0008] Furthermore, this application proposes that the anomaly region segmentation module adopts a semantic segmentation network structure with an attention-guided mechanism, and the network structure receives the following multi-source inputs: Grayscale channel map of the original image; Color complexity parameter diagram; Structural texture feature response map; Tongue morphology mask, output by the tongue extraction and standardization module, is used to limit the segmentation range to within the tongue region; Furthermore, a multi-scale spatial attention module is embedded in the decoding stage to enhance sensitivity to irregular regions.
[0009] Furthermore, this application also proposes an attribute feature analysis unit for extracting the main color, average brightness deviation, texture directionality index, and edge complexity index of each candidate anomaly region. The type determination unit is used to determine the label of abnormal areas based on preset rules or multi-class classification models and comprehensive attribute features. The labels include, but are not limited to: tooth marks, cracks, ecchymosis, moss peeling, water-wet deposits, and dark purple areas. The annotation layer generation unit is used to draw the label and segmentation boundary corresponding to each abnormal area as a visual layer, and display them by category legend; The reference image comparison module is used to filter out reference images that are close to the current image in terms of color consistency, brightness distribution similarity, and texture balance, and outputs the relative deviation index of abnormal areas as an auxiliary factor for label judgment.
[0010] Furthermore, this application also proposes a manual annotation correction interface, which is used by doctors to perform editing operations such as deleting, merging, and reclassifying the system's annotation areas; The metadata feedback mechanism is used to record the image feature context and original annotation results corresponding to each correction operation, and to generate incremental samples; The online fine-tuning and update module is used to perform low-frequency online learning on the semantic segmentation network, gradually improving the system's generalization ability in specific tongue image patterns; The temporal evolution trend analysis module is used to perform time series alignment and image registration on tongue images uploaded multiple times by the same user, and to track the evolution trajectory of the size, shape, color complexity and texture indicators of specific abnormal regions in historical images.
[0011] Furthermore, this application also proposes receiving non-image-based diagnostic information input synchronously by the user, including but not limited to the tongue diagnosis doctor's verbal description, the patient's self-reported symptoms, and basic physical signs data; Semantic keywords are extracted based on natural language processing models and matched with the types of anomaly regions already labeled in the image; A semantic consistency discrimination mechanism is used to enhance the recognition or correct the labels of some confidence boundary regions; The multimodal auxiliary results are fed back to the abnormal region intelligent annotation module to optimize the final output annotation results and region classification labels.
[0012] Furthermore, this application also proposes to generate an initial confidence score for each anomalous region by fusing the region prediction probability output by the semantic segmentation model, the reference image bias index, the historical doctor correction rate, and the anomalous attribute category stability index. Based on the significance scores of structural features in different spectral channels, weighting coefficients for multi-channel prediction probabilities are set, and the segmentation results are weighted and fused with confidence to enhance the interpretation and reliability output of structural response differences.
[0013] Furthermore, this application proposes that the confidence score be overlaid on the original image annotation layer in a pseudo-color transparency manner, and the scoring results are used to assist doctors in quickly judging high-confidence abnormal areas and triggering a review mechanism for low-confidence areas.
[0014] As can be seen from the above, the TCM tongue diagnosis image enhancement system and method based on multispectral imaging provided in this application achieves accurate segmentation of the tongue region and intelligent annotation of abnormal regions by combining multispectral image acquisition, color complexity modeling and structural texture analysis. It has the advantages of improving the segmentation accuracy of the tongue region, enhancing color representation ability, improving the recognition accuracy of structural anomalies such as cracks and tooth marks, and achieving pixel-level abnormal region segmentation and intelligent annotation. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the overall system architecture of the present invention; Figure 2 This is a schematic diagram of the color complexity modeling process of the present invention; Figure 3 This is a flowchart illustrating the structural texture analysis and segmentation network architecture in this invention; Figure 4 This is a schematic diagram of the intelligent annotation module in this invention; Figure 5 This is a schematic diagram of the multimodal feature fusion mechanism in this invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] It should be noted that when a component is described as "fixed to" another component, it can be directly on the other component or may have a component in between. When a component is considered "connected to" another component, it can be directly connected to the other component or may have a component in between. When a component is considered "set on" another component, it can be directly set on the other component or may have a component in between. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0019] In existing technologies, traditional Chinese medicine tongue diagnosis has long relied on the doctor's subjective experience, which is easily affected by ambient lighting and equipment performance, leading to inconsistent diagnostic results. Traditional image-assisted systems suffer from problems such as background interference in tongue segmentation, limited color representation capabilities, and low accuracy in recognizing fine structures. For example, in clinical settings, due to changes in lighting conditions and complex background interference, tongue region segmentation often results in blurred edges or misjudgments; color information is distorted due to equipment differences; and structural abnormalities such as cracks or teeth marks are difficult to accurately locate due to the lack of multi-source feature fusion.
[0020] To address the aforementioned issues, it is necessary to overcome technical bottlenecks such as insufficient tongue segmentation accuracy, color information distortion, and difficulty in identifying structural anomalies. Analysis reveals that existing methods neglect multispectral information during color space conversion, leading to color calibration deviations; single feature analysis struggles to capture complex texture variations; and the segmentation network lacks a multi-source input fusion mechanism, affecting the accuracy of anomaly region localization. Based on this, a proposed approach introduces multispectral image acquisition and color calibration components, combined with geometric normalization to eliminate deformation interference. Through block-based color distribution analysis and multi-scale texture feature extraction, a segmentation network fusing color complexity and structural response is constructed to achieve accurate segmentation and intelligent annotation of anomaly regions.
[0021] Therefore, this application proposes a multispectral-based image enhancement system for traditional Chinese medicine tongue diagnosis, combined with the attached... Figure 1 To be continued Figure 5The following is a description of the system: An image acquisition module, used to acquire tongue image information, including the tongue body region and color calibration elements; a tongue extraction and standardization module, used to extract the tongue body region from the tongue image information and perform geometric standardization based on the tongue contour; a color complexity modeling module, used to perform color space transformation on the tongue body region and divide it into multiple sub-regions, calculating color complexity parameters based on the color distribution dispersion and brightness abrupt change degree of each sub-region; a structure and texture analysis module, used to extract the texture features of each sub-region of the tongue body region, including gray-level co-occurrence matrix features, directional gradient information, and high-frequency response indicators, identifying structural anomalies such as cracks and tooth marks; an anomaly region segmentation module, used to fuse the color complexity parameters and structure and texture features, inputting them into a semantic segmentation network to generate pixel-level segmentation masks for suspected anomaly regions in the tongue image information; and an anomaly region intelligent annotation module, used to classify the attributes of each anomaly region in the segmentation mask and overlay annotations onto the original image.
[0022] The color calibration element refers to a standard color card or object with known reflectivity embedded in the image acquisition environment. Specifically, it can be implemented using checkerboard color blocks or a spectral reflectivity calibration board. This provides a stable color reference during image acquisition, eliminating the influence of equipment differences and ambient lighting on color information. Tongue contour geometric normalization refers to shape correction of the segmented tongue region using affine transformation or thin-plate spline interpolation algorithms. This can be achieved using keypoint registration methods to eliminate geometric deformation caused by shooting angle or tongue posture, ensuring the stability of subsequent feature analysis. Color distribution dispersion refers to the statistical variance of pixel hue, saturation, and brightness values within a segmented region. This can be achieved by calculating the dispersion of each channel in the HSV color space, quantifying the uniformity of color distribution in the region and assisting in identifying local color anomalies. High-frequency response indicators refer to image detail features extracted through a high-pass filter. This can be implemented using the Gaussian-Laplacian operator or wavelet transform, enhancing the salience of high-frequency texture structures such as cracks and dents. Semantic segmentation networks are deep learning models capable of fusing multi-source inputs. Specifically, they can employ an encoder-decoder structure with embedded attention mechanisms to integrate color complexity parameters and structural texture features, generating pixel-level segmentation results for abnormal regions. Attribute classification, on the other hand, determines the type of anomaly based on features such as dominant region color and texture directionality. This can be implemented using support vector machines or convolutional neural network models to assign clinically interpretable label information to the segmentation results.
[0023] Specifically, the image acquisition module uses a multispectral camera to acquire tongue images containing color calibration elements, ensuring the accuracy of color information acquisition. The tongue extraction and standardization module uses an edge detection algorithm to segment the tongue region and eliminates deformation interference through geometric transformations. The color complexity modeling module converts the tongue region into the HSV color space, divides it into equal-scale blocks, and calculates the color distribution dispersion of each block to generate a color complexity parameter map. The structure and texture analysis module extracts texture statistical features through the gray-level co-occurrence matrix and combines multi-scale filtering to enhance structural responses such as cracks, generating a structural feature map. The abnormal region segmentation module inputs the color parameter map, structural feature map, and tongue mask into a semantic segmentation network, and generates an abnormal region mask through multi-source feature fusion. The intelligent annotation module extracts attributes such as main color and texture directionality from the mask region, combines them with a classification model to generate label information, and overlays it onto the original image to assist doctors in quickly locating abnormal regions.
[0024] Compared to existing technologies, current systems often employ single color space conversion and lack calibration components, leading to color calibration biases; texture analysis relies solely on grayscale statistical features, making it difficult to distinguish similar structures; and segmentation networks fail to integrate color complexity and structural response, resulting in missegmentation of abnormal regions. This solution ensures color information accuracy through multispectral acquisition and color calibration, enhances feature representation capabilities by combining block-based color dispersion analysis and multi-scale texture response extraction, improves the accuracy of abnormal region localization by employing a multi-source input fusion segmentation network, and achieves clinically interpretable intelligent assisted diagnosis through attribute classification and annotation overlay.
[0025] Through the above technical solutions, this application solves the problems of tongue segmentation being affected by background interference, color information distortion, and difficulty in recognizing fine structures. It achieves high-precision tongue region segmentation, quantitative analysis of complex color measurements, detection of structural anomalies, and intelligent annotation, thereby improving the objectivity and diagnostic efficiency of TCM tongue diagnosis.
[0026] This application further proposes a color complexity modeling module including a color space conversion unit, a tongue segmentation unit, a color distribution feature extraction unit, and a color complexity parameter generation unit. The color space conversion unit is used to convert tongue image information into HSV color space and perform color calibration based on color calibration elements; the tongue segmentation unit is used to divide the tongue region into multiple equal-scale segmented regions; the color distribution feature extraction unit is used to statistically analyze pixel hue differences, brightness gradient distribution, and saturation range in each segmented region; and the color complexity parameter generation unit is used to fuse color distribution features to generate a representation of the color dispersion degree of the region.
[0027] HSV color space conversion refers to converting an image from the RGB color model into three independent components: hue, saturation, and lightness. This can be achieved using color space transformation algorithms. This conversion method better aligns with human visual perception of color, facilitating the separation of color information from brightness interference. Color calibration elements refer to standard color cards embedded in the image acquisition environment. This can be implemented using physical calibration objects with known chromaticity values. A color calibration matrix is established by comparing the actual color of the calibration object with the image color. Uniform-scale segmented regions refer to dividing the tongue-like region into rectangular grids of fixed sizes. This can be implemented using a sliding window segmentation algorithm. Local region analysis avoids the loss of detail caused by global statistics. Hue difference refers to the dispersion of pixel hue values within the segmented regions. This can be achieved using hue histogram variance calculation to quantify the uniformity of color distribution. Brightness gradient distribution refers to the rate of change of pixel brightness values in space. This can be achieved using the Sobel operator to calculate the gradient magnitude histogram, used to capture local abrupt changes in brightness. Saturation range refers to the difference between the maximum and minimum saturation values within a segmented region. It can be calculated by traversing the extreme values of the pixel saturation channels and is used to characterize the drastic change in color purity. Color dispersion is a quantitative indicator generated by integrating multi-dimensional features of hue, brightness, and saturation. It can be implemented using principal component analysis or weighted fusion methods and is used to reflect the complexity of color changes in a region.
[0028] Specifically, in the HSV color space conversion process, the original tongue image is first transformed into three independent channels: hue, saturation, and lightness, converting the RGB value of each pixel. Then, using known color patches in the image and calculating the deviation between the actual measured chromaticity value and the imaged chromaticity value, a color calibration matrix is established to correct the color of the entire image. After calibration, the tongue region is divided into multiple rectangular blocks of equal size, and color features are extracted independently from each block. Within each block, the histogram variance of the hue channel is calculated to characterize the dispersion of the color distribution, the gradient amplitude distribution of the lightness channel is calculated to reflect the intensity of light and dark changes, and the difference between the maximum and minimum values of the saturation channel is recorded. Finally, these three dimensions of features are used to generate a color complexity parameter through linear weighting or nonlinear fusion. This parameter comprehensively reflects the complexity of the block region in terms of hue diversity, lightness abrupt changes, and saturation fluctuations.
[0029] Compared to existing technologies, traditional methods often use the RGB color space for simple histogram statistics, which cannot effectively separate color and brightness information and are prone to color distortion under complex lighting conditions. This solution significantly improves the accuracy of color representation through HSV color space conversion and physical calibration; it achieves refined capture of local color variations through block region division and multi-dimensional feature fusion; and by introducing brightness gradient distribution analysis, it can effectively distinguish between natural color gradations and pathological mutation regions.
[0030] Through the above technical solutions, this application can accurately quantify the color complexity of each region of the tongue, providing a reliable basis for subsequent abnormal region segmentation. The color calibration mechanism effectively eliminates color deviation caused by differences in device imaging, the block statistical method avoids the masking of local abnormalities by global averaging, and multidimensional feature fusion enhances the robustness of color complexity assessment, ultimately improving the accuracy and consistency of tongue pathological feature detection.
[0031] This application further proposes a structural texture analysis module, including a grayscale statistics submodule, a multi-scale response submodule, a crack morphology extraction submodule, and a structural feature fusion unit. The grayscale statistics submodule is used to construct the grayscale co-occurrence matrix of each block region and extract energy, contrast, and entropy values; the multi-scale response submodule is used to filter the image based on multiple orientation and scale parameters to obtain a local orientation texture enhancement map; the crack morphology extraction submodule is used to refine and analyze the orientation of continuous linear low-brightness structures to extract target regions that conform to crack morphology characteristics; the structural feature fusion unit is used to fuse grayscale statistical features and texture response features into a unified structural texture description vector.
[0032] The gray-level co-occurrence matrix (GLCM) refers to the probability distribution matrix of gray values appearing at specific distances and angles in an image, which can be implemented using the Haralick feature extraction algorithm. It is used to quantify the texture uniformity, contrast, and information entropy of segmented regions. The multi-scale response submodule involves convolutional operations on the image using filter banks of different directions and scales, specifically implemented using Gabor filters or histogram-based directional gradient filters, to enhance texture features at different directions and frequencies. The crack morphology extraction submodule identifies linear low-brightness structures through morphological refinement and orientation field analysis, specifically implemented using skeletonization algorithms and histogram-based statistical methods, to distinguish between natural textures and pathological cracks. The structural feature fusion unit normalizes and weights multiple heterogeneous feature vectors, specifically implemented using principal component analysis or feature concatenation, to eliminate feature redundancy and enhance classification discriminative power.
[0033] Specifically, the grayscale statistics submodule constructs a grayscale co-occurrence matrix for segmented regions, calculating energy, contrast, and entropy values to characterize texture roughness and regularity. The multi-scale response submodule filters the image using multiple sets of direction and scale parameters, generating gradient response maps in multiple directions to capture texture details at different frequencies. The crack morphology extraction submodule performs morphological refinement on low-brightness areas, extracting linear structures of single-pixel width, and uses histogram analysis to filter out regions that match crack orientation characteristics. The structural feature fusion unit normalizes grayscale statistical features and multi-scale response features, generating a unified structural texture description vector through feature concatenation or dimensionality reduction fusion, which is then input into the subsequent segmentation network for abnormal region identification.
[0034] Compared to existing technologies, traditional methods typically employ only single texture features or fixed-scale filtering, making it difficult to effectively distinguish between natural tongue coating textures and pathological structural abnormalities. This proposed solution, through the synergistic analysis of gray-level co-occurrence matrices and multi-scale directional filtering, can simultaneously capture macroscopic texture distribution patterns and microscopic directional details. Combined with the geometric constraints of crack morphology, it significantly reduces the misclassification rate of tooth marks and normal textures.
[0035] Through the above technical solution, this application can accurately identify structural abnormalities such as cracks and teeth marks on the surface of the tongue. By fusion analysis of multi-dimensional texture features, it enhances the sensitivity to subtle pathological changes and provides high-discrimination structural feature input for subsequent semantic segmentation networks, thereby improving the boundary accuracy and classification reliability of abnormal region segmentation masks.
[0036] This application further proposes that the structural texture analysis module also includes a spectral saliency evaluation unit. The spectral saliency evaluation unit is used to perform saliency analysis on the response characteristics of candidate structural regions under multiple spectral channels, calculate the structural saliency score for each channel, and assign differentiated confidence levels to the structural features output by different channels based on the saliency scores.
[0037] The spectral saliency assessment unit is a processing module that calculates the saliency weights of structural regions using multispectral channel data. Specifically, it can be implemented using a multi-channel response value weighted fusion algorithm to quantify the contribution of different spectral dimensions to structural features. Multiple spectral channels refer to the imaging data of tongue images at different wavelengths, such as visible light and near-infrared. This can be achieved by simultaneously acquiring images of different bands using a multispectral sensor to capture the differences in the tongue's reflectivity under different spectra. Saliency analysis involves statistically modeling the texture response intensity of candidate regions in different channels. This can be achieved using inter-channel response variance calculation or regional contrast analysis methods to identify preferred channels sensitive to specific structures. The structural saliency score is a quantitative indicator characterizing the reliability of a spectral channel in detecting target structural features. It can be generated through a comprehensive calculation of response intensity and regional consistency indicators to screen high-confidence feature channels. Differential confidence refers to the dynamic weight allocation of multi-channel features based on the saliency score. This can be achieved by using normalized scores as weighting coefficients to suppress noise channel interference and enhance the representational ability of effective features.
[0038] Specifically, the spectral saliency evaluation unit first receives image data of the tongue region from various channels acquired by a multispectral imaging device. Candidate structural regions, initially identified by the crack morphology extraction submodule or grayscale statistics submodule, are input into this unit for multispectral analysis. For each candidate region, its texture response features, such as high-frequency component energy or directional gradient distribution, are extracted in the red, green, blue, and near-infrared spectral channels. The structural saliency score for each channel is generated by calculating the dispersion of the response values and the regional consistency index within each channel. When a channel exhibits a high response intensity and uniform spatial distribution in the target region, its saliency score is correspondingly improved. Finally, by weighted and fused with the structural feature vectors of each channel and their saliency scores, a comprehensive feature description with channel-optimized characteristics is formed, providing more reliable feature input for subsequent segmentation networks.
[0039] Compared to existing technologies, traditional methods typically use only a single visible light channel for texture analysis, failing to effectively utilize multispectral information to improve the robustness of feature representation. Our proposed solution, however, introduces a spectral saliency evaluation mechanism to automatically select advantageous channels sensitive to specific structural anomalies. For example, the near-infrared channel enhances the visualization of subcutaneous blood vessel morphology, or the blue channel exhibits high contrast response to fine cracks on the tongue surface. This multispectral data-driven channel optimization strategy overcomes the low information utilization problem caused by fixed channel selection, significantly improving the accuracy of structural feature detection, especially under complex lighting conditions or in scenarios with varying tongue color.
[0040] Through the above technical solution, this application effectively solves the problems of channel feature redundancy and noise interference in multispectral tongue image data, and optimizes the expression quality of structural texture features through a dynamic weight allocation mechanism. This solution can improve the segmentation accuracy of structural anomalies such as cracks and tooth marks, reduce misjudgments caused by missing or contaminated information in a single channel, and enhance the system's efficiency in utilizing information acquired by multispectral imaging equipment.
[0041] This application further proposes an anomaly region segmentation module employing a semantic segmentation network structure with an attention-guided mechanism. The network structure receives the following multi-source inputs: grayscale channel images of the original image; color complexity parameter maps; structural texture feature response maps; and a tongue morphology mask map, output by the tongue extraction and normalization module, used to limit the segmentation range to within the tongue region. A multi-scale spatial attention module is embedded in the decoding stage to enhance sensitivity to irregularly shaped regions. The volume extraction and normalization module is responsible for converting the original tongue image information acquired by the image acquisition module into a normalized tongue image suitable for subsequent depth analysis. This module first uses an edge detection algorithm to automatically and accurately segment the tongue contour region from the original image, which includes color calibration elements and complex backgrounds, effectively eliminating interference from non-target regions. Next, based on the extracted tongue contour, the module uses geometric correction algorithms such as affine transformation or thin-plate spline interpolation to perform shape correction and spatial normalization processing on the tongue. This step aims to eliminate geometric deformations introduced by differences in shooting angles or changes in the tongue's own posture, thereby ensuring that tongue images acquired at different times and under different conditions have a comparable spatial benchmark. The standardized tongue image processed by this module not only provides a clean and stable input for subsequent color complexity modeling and structural texture analysis, but its output tongue morphology mask will also serve as a key spatial constraint, inputting into the abnormal region segmentation network to ensure that all advanced analysis and annotation are strictly limited to physiologically relevant tongue regions, fundamentally laying the data foundation for the system to achieve high-precision and objective tongue diagnosis analysis.
[0042] The attention guidance mechanism refers to dynamically adjusting the network's attention weights for different regional features. This can be achieved by combining channel attention and spatial attention, focusing on the salient features of anomalous regions. Multi-source input involves using grayscale information, color complexity parameters, structural texture features, and a tongue morphology mask as parallel inputs. This can be implemented using a multi-branch feature fusion architecture to comprehensively represent the color, texture, and morphological information of the tongue region. The multi-scale spatial attention module assigns spatial weights to feature maps at different scales during decoding. This can be achieved using a combination of dilated convolution and adaptive pooling, enhancing the network's ability to identify small anomalous regions and irregular boundaries.
[0043] Specifically, the semantic segmentation network processes the grayscale channel image, color complexity parameter image, and structural texture feature response image separately during the encoding stage, using the tongue morphology mask as a spatial constraint through cross-layer connections. During the decoding stage, a multi-scale spatial attention module fuses features from different resolution layers, dynamically enhancing the feature response of abnormal regions by calculating spatial location relevance weights. The tongue morphology mask limits the segmentation range through pixel-level multiplication operations, avoiding interference from non-tongue regions.
[0044] Compared to existing technologies, traditional tongue segmentation methods typically rely on a single image channel or simple feature overlay, making it difficult to effectively integrate multi-dimensional information such as color complexity and structural texture. Existing semantic segmentation networks lack spatial constraints on tongue morphology masks and are susceptible to background noise interference. This solution achieves multi-scale feature enhancement of irregularly shaped regions through multi-source input fusion and attention mechanisms, while precisely defining the segmentation range using morphology masks.
[0045] Through the above technical solution, this application solves the problem of insufficient segmentation accuracy of abnormal regions in existing tongue diagnosis systems, effectively improving the segmentation accuracy of fine structural abnormalities such as cracks and teeth marks. By combining multi-source information fusion with spatial attention mechanisms, the probability of missegmentation of non-tongue regions is reduced, and the ability to recognize irregular boundaries and low-contrast regions is enhanced.
[0046] This application further proposes an intelligent annotation module for abnormal regions, comprising an attribute feature analysis unit, a type determination unit, an annotation layer generation unit, and a reference image comparison module. The attribute feature analysis unit extracts the dominant color, average brightness deviation, texture directionality index, and edge complexity index for each candidate abnormal region. The type determination unit labels the abnormal regions based on preset rules or a multi-class classification model that integrates attribute features; labels include serrations, cracks, ecchymosis, moss removal, waterlogged deposits, and dark purple areas. The annotation layer generation unit draws the label and segmentation boundary corresponding to each abnormal region as a visual layer and displays them separately according to category legends. The reference image comparison module selects reference images that are close to the current image in terms of color consistency, brightness distribution similarity, and texture uniformity, and outputs the relative deviation index of the abnormal region as an auxiliary factor for label determination.
[0047] The dominant color of the region refers to the color value with the highest frequency determined by statistically analyzing the histogram of pixel hue distribution within the abnormal region. This can be achieved using the mode statistics method of the hue channel in the HSV color space, representing the main color characteristics of the abnormal region. The average brightness deviation refers to the difference between the average brightness of pixels within the abnormal region and the average brightness of the entire tongue. This can be achieved by comparing the average values of the L channel in the Lab color space, reflecting the difference in brightness between the local area and the overall tongue image. The texture directionality index is a quantitative indicator derived from analyzing the texture direction distribution pattern using the gray-level co-occurrence matrix. This can be achieved by calculating the variance of the contrast values in four directions: 0°, 45°, 90°, and 135°, representing the consistency of the texture arrangement direction. The edge complexity index is the ratio of the perimeter to the area after approximating the boundary of the abnormal region as a polygon. This can be achieved by using the Douglas-Peucker algorithm to simplify the boundary and calculate the shape complexity parameter, quantifying the complexity of the abnormal region's outline. Reference image selection refers to constructing a comprehensive similarity score based on color histogram similarity, brightness distribution KL divergence, and texture feature Euclidean distance. Specifically, it can be implemented using a multi-feature weighted fusion retrieval algorithm to select the comparison sample that is closest to the current tongue image from historical data.
[0048] Specifically, after preliminary identification of candidate anomaly regions by a semantic segmentation network, the attribute feature analysis unit extracts multi-dimensional features from each segmented region. The main color analysis uses HSV hue channel histogram statistics to locate the main distribution range of the anomaly region in the hue dimension. The average brightness deviation is calculated by comparing the L channel mean in Lab color space to quantify the brightness difference between local areas and the overall tongue. The texture directionality index is calculated using the variance of the contrast values in four directions of the gray-level co-occurrence matrix to reflect the directional consistency of texture arrangement. The edge complexity index is calculated using the perimeter-to-area ratio after boundary simplification to assess the complexity of the region's outline. The type determination unit inputs the above features into a preset rule engine or a multi-class classification model, such as a classifier based on a random forest algorithm, and outputs the type label for the anomaly region. The annotation layer generation unit draws bounding boxes and category legends of different colors based on the classification results, and overlays them onto the original image to form a visual annotation. The reference image comparison module retrieves samples from the historical database that are closest to the current tongue image in terms of color, brightness, and texture features, and calculates the relative deviation index between the current anomaly region and the corresponding region of the reference sample, for example, using a weighted summation formula of feature differences, to provide auxiliary judgment criteria for type determination.
[0049] Compared to existing technologies, current tongue diagnosis systems rely solely on single image features and lack historical reference comparisons when annotating abnormalities, making the annotation results susceptible to individual differences and shooting conditions. This application introduces a multi-dimensional attribute feature analysis and reference image comparison mechanism, integrating the comparison information of current image features with historically similar samples during the annotation process, effectively reducing the probability of misjudgment caused by equipment differences or fluctuations in individual physiological characteristics. Existing technologies typically use fixed color annotations for the annotation layers; this application uses category legends to differentiate and display these layers, enabling doctors to intuitively identify the spatial distribution characteristics of different types of abnormal areas.
[0050] Through the above technical solutions, this application can improve the accuracy of abnormal region classification and reduce the risk of misjudgment based on a single feature by fusing multi-dimensional features. The reference image comparison mechanism provides objective auxiliary judgment criteria for type determination, enhancing the interpretability of the annotation results. The visual annotation layer displays different types of abnormal regions through differentiated legends, helping doctors quickly locate key lesion areas and analyze their distribution patterns. The introduction of the relative deviation index enables the system to indicate the degree of difference between the current abnormal region and typical cases, providing a quantitative reference for clinical diagnosis.
[0051] This application further proposes a doctor-interactive correction module, including a manual annotation correction interface for doctors to perform editing operations such as deletion, merging, and reclassification of system-annotated areas; a metadata feedback mechanism for recording the image feature context and original annotation results corresponding to each correction operation and generating incremental samples; an online fine-tuning update module for performing low-frequency online learning on the semantic segmentation network to gradually improve the system's generalization ability in specific tongue image patterns; and a temporal evolution trend analysis module for performing time-series alignment and image registration on tongue image images uploaded multiple times by the same user, tracking the evolution trajectory of the size, shape, color complexity, and texture indicators of specific abnormal regions in historical images.
[0052] The manual annotation correction interface is an interactive interface that allows doctors to manually intervene in the annotation results automatically generated by the system. This can be implemented using a graphical editing tool combined with a region selection algorithm. This interface can correct annotation errors caused by model misjudgment. The metadata feedback mechanism is a data processing flow that associates and stores the doctor's correction operations with the original image features. This can be implemented using a key-value database combined with feature hashing encoding technology. This mechanism preserves the relationship between the correction operation and the image context, providing a data foundation for subsequent model optimization. The online fine-tuning update module is an algorithm module that adjusts model parameters based on incremental data. This can be implemented using a transfer learning framework combined with a mini-batch gradient descent algorithm. This module can adapt to changes in tongue features under specific scenarios without compromising the original model's generalization ability. The temporal evolution trend analysis module is a component that performs temporal correlation analysis on tongue image data collected multiple times from the same patient. This can be implemented using a dynamic time warping algorithm combined with multimodal registration technology. This module can quantify the dynamic changes in abnormal areas, assisting doctors in judging the trend of disease progression.
[0053] Specifically, the doctor-interactive correction module receives adjustment instructions from doctors regarding the annotation results through a manual annotation correction interface, such as deleting misjudged areas or merging adjacent abnormal areas of the same type. The metadata feedback mechanism compares the image region features involved in the correction operation with the original model output, generating an incremental sample set containing the differences before and after correction. The online fine-tuning and update module periodically extracts data from the incremental sample set and updates the parameters of the semantic segmentation network by limiting the learning rate and update frequency, allowing the model to gradually adapt to newly discovered tongue image patterns. The temporal evolution trend analysis module performs spatial registration on the tongue image data from each patient's examination, establishing correspondences for the same anatomical locations. By calculating the area change rate, shape similarity index, and color parameter fluctuation values of abnormal areas, it generates a visualized evolution trend map.
[0054] Compared to existing technologies, traditional tongue diagnosis systems lack a closed-loop interaction between doctors and algorithmic models, failing to feed clinical experience into the system optimization process and neglecting the dynamic changes in the patient's tongue appearance as the disease progresses. Existing technologies typically employ a static analysis model, with each diagnosis being independent and unable to establish temporal correlations. This solution introduces a closed-loop feedback mechanism for doctors to correct data, achieving collaborative optimization between the artificial intelligence model and clinical experience. Furthermore, through time-series analysis, it overcomes the limitations of single-diagnosis approaches, providing data support for the diagnosis and treatment of chronic or cyclical diseases.
[0055] Through the above technical solutions, this application effectively solves the problem of lacking a correction channel when there is a discrepancy between the annotation results of existing systems and clinical judgment, thus improving the clinical applicability of abnormal region annotation. The incremental learning mechanism enables performance optimization of the model during continuous use, avoiding the performance degradation of traditional models caused by changes in data distribution. The time series analysis function provides doctors with quantitative assessment of the development trend of abnormal regions, enhancing the system's application value in disease monitoring.
[0056] This application further proposes a multimodal feature fusion module, which receives non-image-related diagnostic information input synchronously by the user, including the language description of the tongue diagnosis doctor, the patient's self-reported symptoms, and basic physical signs data. Based on a natural language processing model, semantic keywords are extracted and matched with the types of abnormal regions already labeled in the image. A semantic consistency discrimination mechanism is used to enhance the recognition or correct the labels of some confidence boundary regions. The multimodal auxiliary results are fed back to the abnormal region intelligent labeling module to optimize the final output labeling results and region classification labels.
[0057] The multimodal feature fusion module refers to a collaborative analysis architecture that integrates visual and textual information. Specifically, it can be implemented using a cross-modal attention mechanism, achieving information complementarity by establishing an association mapping matrix between image and text features. The natural language processing model refers to a deep learning model used to parse unstructured text. Specifically, it can use the pre-trained language model BERT to extract semantic keywords and extract core medical terms from symptom descriptions using entity recognition technology. The semantic consistency discrimination mechanism is a cross-modal feature alignment verification method. Specifically, it can use cosine similarity to calculate the matching degree between image and text feature vectors, filtering out low-relevance information by setting a dynamic threshold.
[0058] Specifically, when a patient reports symptoms of "dry mouth and tongue," the natural language processing model extracts the keyword "insufficient body fluids" and semantically associates it with the "peeled tongue coating" area detected in the image. When the system detects a low-confidence superficial crack area on the edge of the tongue, it combines the patient's complaint of "night sweats" and uses a semantic consistency discrimination mechanism to correct the area to the "yin deficiency crack" category. When a dark spot with abnormal color complexity is detected in the middle of the tongue, combined with the doctor's input of the diagnosis of "blood stasis constitution," the system upgrades the initial label of the dark spot area from ordinary pigmentation to "ecchymosis."
[0059] Compared to existing technologies, traditional tongue diagnosis systems rely solely on single-modal image data for analysis, failing to integrate key symptom descriptions from patient records. Existing methods suffer from high mislabeling rates due to a lack of textual information to aid judgment when dealing with ambiguous abnormal areas. This solution constructs a cross-modal feature interaction mechanism, enabling the tongue image analysis system to combine the "four diagnostic methods" principle from traditional Chinese medicine to achieve collaborative verification of image features and clinical symptoms.
[0060] Through the above technical solutions, this application effectively solves the labeling bias problem caused by isolated information in existing tongue diagnosis systems, and improves the recognition accuracy of complex pathological features. By establishing semantic relationships between multimodal data, the system's ability to analyze TCM syndrome elements is enhanced, making the labeling results of abnormal areas more consistent with clinical diagnostic logic. For controversial boundary areas, this technical solution can utilize non-image information to achieve dynamic correction, significantly reducing the frequency of manual review for repeated labeling.
[0061] This application further proposes that the system also includes a confidence output module, which is used to generate an initial confidence score for each abnormal region by fusing the region prediction probability output by the semantic segmentation model, the reference image deviation index, the historical doctor correction rate, and the abnormal attribute category stability index; and based on the significance scores of structural features in different spectral channels, setting weighting coefficients for multi-channel prediction probabilities, and performing confidence-weighted fusion of the segmentation results to enhance the interpretation and reliability output of structural response differences.
[0062] Among these, the region prediction probability refers to the probability prediction value of a pixel belonging to an anomaly category by the semantic segmentation network. Specifically, it can be obtained by normalizing the network output using the Softmax function, used to quantify the model's certainty regarding the segmentation result. The reference image bias index refers to the degree of attribute difference between the current anomaly region and the corresponding region of a similar tongue image in the reference image library. Specifically, it can be achieved by calculating the color histogram similarity and texture feature distance, used to assess the anomaly salience of the current region. The historical doctor correction rate refers to the proportion of this type of anomaly region that has been manually modified in historical annotations. Specifically, it can be obtained by statistically analyzing the ratio of correction operations to initial annotations in the annotation system log, used to reflect the reliability of the model's judgment of this type of anomaly. The anomaly attribute category stability index refers to the probability that the same region is classified into the same category under different spectral channels. Specifically, it can be calculated through the consistency test of multispectral feature classification results, used to characterize the cross-modal stability of anomaly attributes. The multi-channel prediction probability weighting coefficient refers to the weight parameters dynamically adjusted according to the salience scores of structural features under different spectral channels. Specifically, it can be generated using linear weighting or adaptive attention mechanisms, used to strengthen the influence of high-salience channels on the final confidence score.
[0063] Specifically, the confidence output module first receives the initial segmentation results from the semantic segmentation network and extracts the predicted probability distribution data for each abnormal region. Simultaneously, it obtains the deviation index between the current region and historical cases from the reference image comparison module, and combines this with correction rate data from the doctor's correction record database to construct a multi-dimensional feature vector. A weighted fusion algorithm integrates these indicators into an initial confidence score, where the weight coefficients of each indicator can be dynamically adjusted based on clinical validation data. Furthermore, the multispectral channel saliency score output from the structural texture analysis module is used as an auxiliary weight to perform secondary optimization of the initial score. For example, when a crack region has a higher saliency response in the near-infrared channel, the predicted probability corresponding to that channel will be assigned a higher weighting coefficient, thereby improving the segmentation confidence for subtle structural anomalies.
[0064] In some specific implementations, the confidence score can be converted into a visual heatmap using a color transparency mapping algorithm and overlaid on the original tongue image. For example, red tones can represent high-confidence abnormal areas, and blue tones can represent low-confidence areas, with transparency positively correlated with the score. Doctors can adjust the weighting coefficient combination strategy through an interactive interface, for example, prioritizing the salience of texture features in the edge areas of the tongue and emphasizing color complexity parameters in the central area of the tongue.
[0065] Compared to existing technologies, current tongue diagnosis systems typically rely solely on the predicted probability of a single model as the basis for confidence levels, without considering the impact of physician correction feedback and differences in multispectral feature responses on the results. For example, when identifying tongue fissures, traditional methods may result in artificially high prediction confidence levels for a single visible light channel due to ambient light interference, while ignoring the more significant structural response features in the near-infrared channel.
[0066] Through the above technical solution, this application effectively solves the problem of the single dimension of confidence assessment in existing tongue diagnosis systems. By integrating model prediction, doctor correction records, cross-modal stability, and multispectral response characteristics, a multi-dimensional quantitative assessment of the reliability of abnormal areas is achieved. Doctors can quickly identify high-confidence abnormal areas for focused diagnosis, while initiating a review process for low-confidence areas to avoid the risk of missed diagnoses due to model misjudgment. Furthermore, through a multispectral channel weighting mechanism, the system's ability to resolve complex structural abnormalities is enhanced. For example, in areas where tongue coating peeling and water deposits are difficult to distinguish in the visible light channel, the classification confidence can be improved through significant response differences in the ultraviolet spectral channel.
[0067] This application further proposes to overlay the confidence score onto the original image annotation layer using a pseudo-color transparency method. The score results are used to assist doctors in quickly identifying high-confidence abnormal areas and triggering a review mechanism for low-confidence areas.
[0068] The pseudo-color transparency overlay display maps different confidence scores to visual overlay layers with different colors and transparency levels. This can be achieved by combining the HSV color space mapping algorithm with a transparency gradient adjustment algorithm, visually distinguishing labeled areas of different confidence levels through color differences and transparency changes. The rapid identification of high-confidence abnormal areas prioritizes the display of areas with confidence scores higher than a preset threshold based on color mapping rules. This can be achieved by setting color saturation and brightness thresholds, allowing doctors to quickly focus on more reliable test results. The verification mechanism trigger automatically generates a verification prompt signal when a low-confidence area is detected. This can be achieved by combining confidence score distribution histogram analysis with edge region continuity detection algorithms, ensuring secondary verification of controversial labeled areas.
[0069] Specifically, the confidence score is generated through multi-source feature fusion calculation, and then color-coded to map the values to a preset pseudo-color range. Simultaneously, the layer transparency is dynamically adjusted based on the score. In the display interface, high-confidence areas are covered with a semi-transparent, highly saturated color, while low-confidence areas are covered with a highly transparent, low-saturation color. When observing the annotation results, doctors can quickly identify abnormal areas with higher confidence levels through color comparison. Simultaneously, the system automatically triggers a review process for areas with confidence levels below a threshold, prompting doctors to conduct manual review or utilize multimodal data for auxiliary verification.
[0070] Compared to existing technologies, current tongue diagnosis systems typically only output binary segmentation results and lack confidence visualization, making it difficult for doctors to distinguish the reliability of the test results. This solution, however, uses a pseudo-color transparency overlay mechanism to transform the algorithm's internal confidence information into intuitive visual cues. Simultaneously, it establishes a review trigger rule, effectively addressing the risk of misjudgment caused by ignoring prediction uncertainty in traditional methods.
[0071] Through the above technical solution, this application realizes the visualization of the confidence level of the detection results of abnormal areas of tongue image, enabling doctors to prioritize the treatment of high-confidence abnormal areas and specifically review low-confidence areas, reducing the problem of missed or false detections caused by algorithm misjudgment, and improving the efficiency and reliability of the diagnostic process through an automated review triggering mechanism.
[0072] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principle of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A multispectral-based image enhancement system for traditional Chinese medicine tongue diagnosis, characterized in that, include: The image acquisition module is used to acquire tongue image information, which includes the tongue body area and color calibration elements; The tongue extraction and standardization module is used to extract the tongue region from the tongue image information and perform geometric standardization based on the tongue contour. The color complexity modeling module is used to perform color space conversion on the tongue region and divide it into multiple sub-regions. Based on the color distribution dispersion and brightness abrupt change of each sub-region, the color complexity parameters are calculated. The structural texture analysis module is used to extract the texture features of each block in the tongue region, including gray-level co-occurrence matrix features, directional gradient information and high-frequency response indicators, to identify cracks, tooth marks and structural anomalies. The abnormal region segmentation module is used to fuse color complexity parameters with structural texture features and input them into the semantic segmentation network to generate pixel-level segmentation masks for suspected abnormal regions in tongue image information. The intelligent annotation module for abnormal regions is used to classify the attributes of each abnormal region in the segmentation mask and overlay the annotations onto the original image.
2. The multispectral-based TCM tongue diagnosis image enhancement system according to claim 1, characterized in that, The color complexity modeling module further includes: The color space conversion unit is used to convert tongue image information into the HSV color space and perform color calibration based on the color calibration element. Tongue segmentation unit, used to divide the tongue region into multiple equal-scale segmented regions; The color distribution feature extraction unit is used to statistically analyze pixel hue differences, brightness gradient distribution, and saturation range in each block region. The color complexity parameter generation unit is used to fuse the color distribution features to generate a characterization of the color dispersion degree of the region.
3. The multispectral-based TCM tongue diagnosis image enhancement system according to claim 1, characterized in that, The structural texture analysis module further includes: The grayscale statistics submodule is used to construct the grayscale co-occurrence matrix of each block region and extract energy, contrast and entropy values; The multi-scale response submodule is used to filter the image based on multiple orientation and scale parameters to obtain a local orientation texture enhancement map. The crack morphology extraction submodule is used to refine and analyze the direction of continuous linear low-brightness structures and extract target areas that conform to crack morphology characteristics. The structural feature fusion unit is used to fuse grayscale statistical features and texture response features into a unified structural texture description vector.
4. The multispectral-based TCM tongue diagnosis image enhancement system according to claim 3, characterized in that, The structural texture analysis module also includes a spectral saliency evaluation unit, which is used to perform saliency analysis on the response characteristics of candidate structural regions under multiple spectral channels, calculate the structural saliency score for each channel, and assign differentiated confidence levels to the structural features output by different channels based on the saliency scores.
5. The multispectral-based TCM tongue diagnosis image enhancement system according to claim 3, characterized in that, The abnormal region segmentation module adopts a semantic segmentation network structure with an attention-guided mechanism, which receives the following multi-source inputs: Grayscale channel map of the original image; Color complexity parameter diagram; Structural texture feature response map; Tongue morphology mask, which is output by the tongue extraction and standardization module, is used to limit the segmentation range to within the tongue region; Furthermore, a multi-scale spatial attention module is embedded in the decoding stage to enhance sensitivity to irregular regions.
6. The multispectral-based TCM tongue diagnosis image enhancement system according to claim 4, characterized in that, The intelligent annotation module for abnormal regions includes: The attribute feature analysis unit is used to extract the main color, average brightness deviation, texture directionality index and edge complexity index of each candidate anomaly region. The type determination unit is used to determine the label of abnormal areas based on preset rules or multi-class classification models and comprehensive attribute features. The labels include, but are not limited to: tooth marks, cracks, ecchymosis, moss peeling, water-wet deposits, and dark purple areas. The annotation layer generation unit is used to draw the label and segmentation boundary corresponding to each abnormal area as a visual layer, and display them by category legend; The reference image comparison module is used to filter out reference images that are close to the current image in terms of color consistency, brightness distribution similarity, and texture balance, and outputs the relative deviation index of abnormal areas as an auxiliary factor for label judgment.
7. The multispectral-based TCM tongue diagnosis image enhancement system according to claim 6, characterized in that, It also includes a doctor interaction correction module, which includes: The manual annotation correction interface is used by doctors to delete, merge, and reclassify the system's annotation areas. The metadata feedback mechanism is used to record the image feature context and original annotation results corresponding to each correction operation, and to generate incremental samples; The online fine-tuning and update module is used to perform low-frequency online learning on the semantic segmentation network, gradually improving the system's generalization ability in specific tongue image patterns; The time evolution trend analysis module is used to perform time series alignment and image registration on tongue images uploaded multiple times by the same user, and to track the evolution trajectory of the size, shape, color complexity and texture indicators of specific abnormal regions in historical images.
8. The multispectral-based TCM tongue diagnosis image enhancement system according to claim 6, characterized in that, It also includes a multimodal feature fusion module, which is used for: Receive non-image-based diagnostic information input by users synchronously, including but not limited to the verbal descriptions of tongue diagnosis doctors, patient self-reported symptoms, and basic physical signs data; Semantic keywords are extracted based on natural language processing models and matched with the types of anomaly regions already labeled in the image; A semantic consistency discrimination mechanism is used to enhance the recognition or correct the labels of some confidence boundary regions; The multimodal auxiliary results are fed back to the abnormal region intelligent annotation module to optimize the final output annotation results and region classification labels.
9. The multispectral-based TCM tongue diagnosis image enhancement system according to claim 8, characterized in that, The system further includes a confidence output module, which is used for: For each anomalous region, an initial confidence score is generated by fusing the region prediction probability output by the semantic segmentation model, the reference image bias index, the historical doctor correction rate, and the anomalous attribute category stability index. Based on the significance scores of the structural features in different spectral channels, weighting coefficients for multi-channel prediction probabilities are set, and the segmentation results are weighted and fused with confidence to enhance the interpretation and reliability output of structural response differences.
10. The multispectral TCM tongue diagnosis image enhancement system according to claim 9, characterized in that, The confidence score is overlaid on the original image annotation layer in a pseudo-color transparency manner. The score result is used to assist doctors in quickly judging high-confidence abnormal areas and triggering a review mechanism for low-confidence areas.