Tongue picture analysis method and system based on RGB image and spectral data fusion
By selecting dominant modality information for asymmetric fusion in tongue image analysis, the problem of information dilution caused by global fusion strategies is solved, and the comprehensiveness and accuracy of tongue image analysis are improved. In particular, in the diagnosis of specific areas of the tongue surface, high-fidelity expression of key features is ensured.
Patent Information
- Application Number
- CN202511394221.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-28
AI Technical Summary
In existing technologies, global fusion strategies for RGB and hyperspectral images cannot fully account for the heterogeneity of information contributions from different regions of the tongue surface. This results in inaccurate or insufficient representation of local lesions or key physiological features in the fused image, posing a challenge, especially when it is necessary to perform refined analysis and differential diagnosis of subtle physiological or pathological features of specific functional zones or specific anatomical structures of the tongue surface.
By identifying the target local region and tongue features in RGB and hyperspectral images, dominant modality information is selected based on a pre-set medical knowledge rule base, and an asymmetric fusion algorithm is used to enhance the expression of the target tongue features, thereby achieving local adaptive fusion.
It improves the comprehensiveness and accuracy of tongue image analysis, ensures high-fidelity representation of the most valuable local subtle features in the fused image for diagnosis, and enhances the accuracy and reliability of subsequent feature extraction, quantitative analysis and intelligent recognition.
Smart Images

Figure CN120876484A_ABST
Abstract
Description
Technical Field
[0001] This application relates to tongue image processing and analysis technology, and in particular to a tongue image analysis method and system based on the fusion of RGB images and spectral data. Background Technology
[0002] Tongue observation is commonly used in Traditional Chinese Medicine to assess health status. By observing the shape of the tongue, the color of the tongue coating, and its moisture and dryness, it reflects the function of the internal organs and the state of qi, blood, and body fluids. Therefore, it is an important auxiliary means for TCM diagnosis and monitoring of sub-health and chronic diseases.
[0003] Because doctors have varying levels of experience, standardization is impossible. Therefore, the relevant technology employs intelligent tongue image detection. First, it captures RGB and hyperspectral images of the subject's tongue. RGB images primarily record the color and texture information of the tongue surface, offering high spatial resolution and visually showcasing the macroscopic features of the tongue. Hyperspectral images, on the other hand, record spectral reflectance information from multiple continuous narrow bands at different locations on the tongue surface. Their core advantage lies in providing continuous spectral curves for each pixel, revealing subtle differences in tissue biochemical components that are difficult to discern with the naked eye. After receiving these two types of image data, a fusion algorithm combines the spatial details of the RGB image with the spectral features of the hyperspectral image to generate a fused tongue image. Based on this fused image, the system further analyzes tongue color, tongue coating, tongue morphology, and other features, outputting a detection report.
[0004] However, the tongue itself is not a homogeneous object of observation; its different physiological regions exhibit significant differences in tissue structure, vascular distribution, and pathological manifestations. This leads to uneven information contribution and effectiveness of RGB and hyperspectral images in representing these different regions. When a global, unchanging fusion strategy is used in the image fusion process, it cannot fully account for the heterogeneity of information contribution from different regions of the tongue, nor can it adapt to the different degrees of dependence of different types of local features on the two modalities. As a result, for local lesions or key physiological features that are very significant in one modality but relatively weak in another, their information expression in the fused image may be diluted, weakened, or even submerged, leading to inaccurate and insufficient representation of local information in the fused image. The challenges are particularly prominent when this tongue image analysis method based on the fusion of RGB and hyperspectral images is applied to scenarios requiring refined analysis and differential diagnosis of subtle physiological or pathological features of specific functional areas or anatomical structures of the tongue. Summary of the Invention
[0005] This application provides a tongue image analysis method and system based on the fusion of RGB images and spectral data, which has the advantages of improving the comprehensiveness and accuracy of tongue image analysis.
[0006] On the one hand, this application provides a tongue image analysis method based on the fusion of RGB images and spectral data, including: RGB and hyperspectral images of the tongue surface of the same subject were captured separately, and RGB and hyperspectral image information was obtained through image analysis. Determine at least one target local region and at least one target tongue feature within the target local region in the RGB image and hyperspectral image; Based on a preset medical knowledge rule base, for the target local region and the target tongue image features, the dominant modality information used to characterize the target tongue image features is selected from the RGB image information and the hyperspectral image information. Guidance data is generated based on the dominant modality information, and target data related to the target tongue image features is extracted from the RGB image information and the hyperspectral image information. The target data is processed by an asymmetric fusion algorithm using the guiding data to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to enhance the expression of the target tongue image features in the fusion result of the target local region; Based on the local fusion results of each target local region, a fused tongue image is generated.
[0007] The above scheme identifies at least one target local region and at least one target tongue feature within the target local region in both RGB and hyperspectral images. This is a crucial step in achieving local adaptive fusion. By focusing on specific regions and features, information fusion can be more targeted, avoiding information dilution caused by global fusion. Then, for the target local region and target tongue features, based on a pre-defined medical knowledge rule base, the dominant modality information used to characterize the target tongue features is selectively determined from both RGB and hyperspectral image information. This step is the core of achieving differentiated fusion. The medical knowledge rule base determines the importance of different modal information in characterizing specific regions and features, thus providing a basis for the selection of the dominant modality and making the fusion process more in line with medical understanding. Next, guiding data is generated based on the dominant modality information, and target data related to the target tongue features is extracted from the RGB and hyperspectral image information. Generating guiding data based on the dominant modality information effectively guides the fusion direction of the target data, highlighting the features of the dominant modality, while extracting target data related to the target tongue features, ensuring that the fused information is effective information relevant to diagnosis. A differentiated fusion algorithm is used to process the feature based on the specific contribution of the two modal data, thereby preserving and enhancing the most valuable local subtle features for diagnosis in the fused image with high fidelity, so as to improve the accuracy and reliability of subsequent feature extraction, quantitative analysis and intelligent recognition.
[0008] Optionally, after acquiring RGB and hyperspectral images of the tongue surface captured separately for the same subject, the method further includes: The RGB image and hyperspectral image are positionally calibrated to ensure that the coordinates of each point on the tongue surface in the spatial coordinate system remain consistent in the RGB image and hyperspectral image. The first preprocessing operation is performed on the RGB image and hyperspectral image after position calibration. The first preprocessing operation includes at least one of noise filtering, image enhancement, geometric correction and illumination correction.
[0009] The above approach can improve the accuracy and quality of image registration, providing a more accurate data foundation for subsequent fusion.
[0010] Optionally, the step of determining at least one target local region and at least one target tongue feature within the target local region in the RGB image and hyperspectral image specifically includes: An edge-detection-based tongue segmentation algorithm is used to segment the tongue region and obtain multiple target local regions. The feature detection algorithm is applied to each local region of the target to obtain the type information of the target tongue image features. The feature detection algorithm includes at least one of spot detection, texture analysis, and edge detection.
[0011] The above method can accurately segment the tongue region and detect the target tongue image features, providing more accurate regional and feature information for subsequent fusion.
[0012] Optionally, the step of determining the dominant modality information for characterizing the target tongue image features, based on a preset medical knowledge rule base, selectively choosing between the RGB image information and the hyperspectral image information, specifically includes: For the target local region and the target tongue image features, calculate the RGB image structural feature evaluation value obtained based on the RGB image information, and the hyperspectral image spectral feature evaluation value obtained based on the hyperspectral image information; The structural feature evaluation value of the RGB image and the spectral feature evaluation value of the hyperspectral image are compared, and the one with the larger evaluation value is determined as the dominant mode of the image data. The dominant knowledge base modality corresponding to the target local region and the target tongue image features is obtained from the preset medical knowledge rule base. Determine whether the dominant mode of the knowledge base is consistent with the dominant mode of the image data. If they are inconsistent, obtain the knowledge base weight of the target local region from the preset medical knowledge rule base, determine the data weight based on the preprocessing quality of the RGB image information and the hyperspectral image information, perform weighted fusion of the dominant mode of the knowledge base based on the knowledge base weight, and perform weighted fusion of the dominant mode of the image data based on the data weight to obtain the target dominant mode information. If they match, the dominant modality of the knowledge base is used as the dominant modality information in the RGB image information and the hyperspectral image information to characterize the features of the target tongue image.
[0013] The above approach can comprehensively consider image data and medical knowledge to more accurately determine the dominant modality information and improve the accuracy of fusion.
[0014] Optionally, the steps of obtaining the knowledge base weight of the target local region from the preset medical knowledge rule base, and determining the data weight based on the preprocessing quality of the RGB image information and the hyperspectral image information, specifically include: Obtain tongue image feature type information related to the target local region; Based on the acquired tongue image feature type information, an evaluation criterion is selected to quantify the reliability of the target local region; Based on the selected evaluation criteria, the reliability of the target local region is quantified to obtain a value for the reliability of the target local region. The reliability value of the target local region is converted into the knowledge base weight using a preset conversion method; and, Calculate the quality index values of the first preprocessing operation of the RGB image information and the hyperspectral image information in the target local area; Based on the quality index value, the credibility of the image data is quantified, and the credibility is converted into the data weight through a second preset conversion method.
[0015] The above approach allows for a more reasonable determination of knowledge base weights and data weights, thereby improving the reliability of the fusion process.
[0016] Optionally, the step of generating guidance data based on the dominant modality information and extracting target data related to the target tongue features from the RGB image information and the hyperspectral image information specifically includes: When RGB image information is used as the dominant modal information, the RGB image is filtered using a guided filter to generate a structured guided map that preserves edge details and smooths unstructured areas; and color space components are extracted as a color guided map; the structured guided map or the color guided map is used as guided data. Extract the spectral feature vector map from the hyperspectral image that corresponds to the local region of the target in the RGB image and is related to the tongue features of the target, as the target data.
[0017] The above scheme can effectively utilize the structural and color information of RGB images to guide the fusion of hyperspectral images and enhance the details of the fused image.
[0018] Optionally, the target data is processed using the guiding data using an asymmetric fusion algorithm to obtain a local fusion result. The step of using the asymmetric fusion algorithm to enhance the expression of the target tongue image features in the fusion result of the target local region specifically includes: Using the spectral feature vector map as input to the guided filter and the structural guided map as the guided image, an asymmetric fusion algorithm is applied to align the output spectral features spatially with the RGB structure and enhance details; or, The spectral feature vector map is used as the input image for guided filtering, and the color guide map is used as the guide image. The brightness, saturation, or intensity of a specific color channel of the corresponding pixel in the color guide map is adjusted using hyperspectral feature values.
[0019] The above methods can achieve alignment and detail enhancement of spectral features with RGB structures, or adjust the color guide map using hyperspectral features, thereby enhancing the expression of the target tongue image features.
[0020] Optionally, the step of generating guidance data based on the dominant modality information and extracting target data related to the target tongue features from the RGB image information and the hyperspectral image information specifically includes: When hyperspectral image information is used as the dominant modal information, the spectral index map obtained by band operation on the hyperspectral image, or the principal component score map obtained by principal component analysis on the hyperspectral image, is used as the guiding data. The fine texture and color information corresponding to the target local region in the hyperspectral image and related to the target tongue features are extracted from the RGB image as target data.
[0021] The above scheme can effectively utilize the spectral information of hyperspectral images to guide the fusion of RGB images, thereby enhancing the texture and color information of the fused image.
[0022] Optionally, the target data is processed using the guiding data using an asymmetric fusion algorithm to obtain a local fusion result. The step of using the asymmetric fusion algorithm to enhance the expression of the target tongue image features in the fusion result of the target local region specifically includes: Using the fine texture and color information as input to the guided filter, and the spectral index map as the guiding image, an asymmetric fusion algorithm is applied to visually enhance abnormal regions indicated by the spectral index in the fused tongue image; or... The fine texture and color information are used as input to the guided filter, and the principal component score map is used as the guided image. The information of each local region after guided fusion processing is weighted and combined or selectively stitched together.
[0023] The above scheme can achieve visual enhancement of spectrally indicated abnormal regions in fused tongue images, or weighted combination or selective splicing of information from local regions, thereby strengthening the expression of target tongue image features.
[0024] Secondly, a tongue image analysis system based on the fusion of RGB images and spectral data, the system comprising: The acquisition module is used to acquire RGB and hyperspectral images of the tongue surface of the same subject, and obtain RGB and hyperspectral image information through image analysis. A region feature determination module is used to determine at least one target local region and at least one target tongue image feature within the target local region in the RGB image and hyperspectral image; The dominant mode determination module is used to determine the dominant mode information for characterizing the target tongue image features based on a preset medical knowledge rule base, choosing one from the RGB image information and the hyperspectral image information; The data preparation module is used to generate guidance data based on the dominant modality information and extract target data related to the target tongue image features from the RGB image information and the hyperspectral image information; The local fusion module is used to process the target data using the guiding data with an asymmetric fusion algorithm to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to enhance the expression of the target tongue image features in the fusion result of the target local region; An image generation module is used to generate a fused tongue image based on the local fusion results of each of the target local regions.
[0025] The above scheme enables the implementation of each step in tongue image analysis, improving the efficiency and accuracy of tongue image analysis.
[0026] As can be seen from the above, the tongue image analysis method and system based on the fusion of RGB images and spectral data provided in this application adaptively selects the dominant modal information of RGB images and hyperspectral images according to the characteristics of different regions of the tongue surface, and performs asymmetric fusion. This solves the problem of dilution of local key diagnostic information caused by the global fusion strategy in the prior art, and has the advantage of improving the comprehensiveness and accuracy of tongue image analysis. Attached Figure Description
[0027] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 The diagram above exemplifies a flowchart of a tongue image analysis method based on the fusion of RGB images and spectral data according to an embodiment. Figure 2 The diagram illustrates, by way of example, a block diagram of a tongue image analysis system 100 based on the fusion of RGB image and spectral data according to an embodiment. Detailed Implementation
[0029] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of this application.
[0030] It should be understood that the terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate, for example, to allow implementation in orders other than those given in the embodiments illustrated or described in this application.
[0031] Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclusively include, for example, a product or device that includes a series of components is not necessarily limited to those that are explicitly listed, but may include other components that are not explicitly listed or that are inherent to such product or device.
[0032] Traditional tongue image analysis systems utilize RGB and hyperspectral images for fusion analysis. RGB images provide spatial color texture information, while hyperspectral images provide spectral biochemical composition information. However, the physiological and pathological characteristics of different regions of the tongue vary, leading to an imbalance in the information contribution of RGB and hyperspectral images in characterizing these regions. Related fusion strategies typically employ a globally uniform approach, failing to adequately consider the heterogeneity of information contributions from different regions of the tongue and the varying degrees of dependence of different types of local features on the two modalities. This globally uniform fusion strategy results in inaccurate or insufficient representation of key local diagnostic information, especially features significant in one modality but weak in another, in the fused image.
[0033] For example, suppose a TCM-assisted diagnostic system requires a detailed analysis of a user's tongue image. The system acquires RGB and hyperspectral images of the user's tongue. When analyzing the tip of the tongue, it needs to identify tiny red dots; when analyzing the sides of the tongue, it needs to identify the details of the teeth marks; when analyzing the root of the tongue, it needs to assess the thickness and color of the tongue coating; and when analyzing the sublingual region, it needs to observe the morphology and color of the veins. These local features differ significantly between the RGB and hyperspectral images. If a globally uniform fusion algorithm is used, it may not be able to effectively highlight these local features that are significant in specific modalities. Tiny red dots at the tip of the tongue may become blurred due to interference from hyperspectral information, and spectral anomalies in the tongue coating at the root of the tongue may be weakened by the averaging of RGB texture information. This inaccuracy in information representation makes it difficult for the system to accurately identify and quantify these key local diagnostic features. If the above problems are not addressed, i.e., if a globally uniform fusion strategy is continued, the fused tongue image will fail to retain and enhance the most valuable local subtle features for diagnosis with high fidelity. This will directly affect subsequent tongue feature extraction algorithms, potentially leading to inaccurate feature extraction or missed detections.
[0034] To address the aforementioned challenges, this application explores ways to overcome the limitations of global fusion strategies and achieve differentiated processing of different regions and features of the tongue. The application proposes combining medical knowledge and the inherent characteristics of the image data to determine which modality—RGB or hyperspectral—is dominant in a specific local region and for a particular tongue feature. Then, based on the determined dominant modality, an asymmetric fusion algorithm is designed to guide the fusion process and enhance the expression of the target tongue features in the fusion result. This asymmetric fusion strategy selectively utilizes information from different modalities according to the characteristics of local regions and features, avoiding information dilution and thus highlighting key diagnostic information in the fused image.
[0035] like Figure 1As shown, an exemplary flowchart of a tongue image analysis method based on the fusion of RGB images and spectral data is illustrated. This application proposes a tongue image analysis method based on the fusion of RGB images and spectral data, comprising: S10: Acquire RGB and hyperspectral images of the tongue surface of the same subject, respectively, and obtain RGB image information and hyperspectral image information through image analysis.
[0036] First, RGB and hyperspectral images are acquired, and corresponding image information is obtained through image analysis. This is the basis for subsequent fusion, providing a data source for subsequent analysis and fusion.
[0037] Furthermore, after acquiring RGB and hyperspectral images of the tongue surface of the same subject, the process also includes: The RGB image and hyperspectral image are positionally calibrated to ensure that the coordinates of each point on the tongue surface in the spatial coordinate system remain consistent in the RGB image and hyperspectral image. The first preprocessing operation is performed on the position-calibrated RGB image and hyperspectral image, the first preprocessing operation including at least one of noise filtering, image enhancement, geometric correction and illumination correction.
[0038] Position calibration refers to calculating and applying geometric transformations to ensure that pixels corresponding to the same physical point in two images have the same spatial coordinates. This can be achieved using feature point matching methods, calculating the transformation matrix between images, and then applying this transformation to resample one or both images. Alternatively, registration methods based on regional cross-correlation or mutual information can be used.
[0039] The first preprocessing operation refers to a set of image processing steps performed on the original image before image fusion, aimed at improving image quality, eliminating interference, or correcting distortion. Specifically, noise filtering refers to removing randomly distributed abnormal pixel values from the image using filtering algorithms. Image enhancement refers to improving the visual effect of an image or highlighting specific features by adjusting its contrast, brightness, sharpness, or color distribution. Geometric correction refers to correcting geometric distortions caused by camera lens or imaging angle by applying pre-calibrated or calculated geometric transformation parameters. Illumination correction refers to compensating for brightness or color variations in the image caused by uneven light sources or shadows, resulting in a more uniform illumination distribution.
[0040] In some embodiments, this application adds position calibration and a first preprocessing step after acquiring RGB and hyperspectral images to provide high-quality, registered input data for subsequent image analysis and fusion. Position calibration ensures that the spatial structural information of the RGB image is accurately aligned with the spectral information of the hyperspectral image. This is crucial for subsequent feature extraction in specific local regions, determination of dominant modes, and asymmetric fusion. If the images are misaligned, even the most sophisticated subsequent fusion algorithm cannot combine the correct spatial information with the correct spectral information, resulting in distorted fusion results that fail to accurately reflect the local features of the tongue image. The first preprocessing operation improves the quality of the image itself.
[0041] Noise filtering reduces random interference, making the texture, color, and spectral curves of the tongue image clearer and more discernible, thus preventing noise from interfering with subsequent feature extraction and analysis. Image enhancement can highlight subtle features of the tongue image, such as color changes or fine cracks in early lesions, making them easier to detect and identify in subsequent processing. Geometric correction eliminates image distortion, ensuring the accuracy of tongue morphology and local region segmentation. Illumination correction compensates for uneven illumination, making the color information of the tongue and tongue coating more realistic and reliable, avoiding misjudgments caused by differences in illumination. These preprocessing steps work together with a local asymmetric fusion strategy.
[0042] In a specific embodiment, after acquiring the RGB and hyperspectral images, position calibration can be performed first. Position calibration can employ a SIFT-based method to extract feature points from the RGB and hyperspectral images, remove mismatched points using the RANSAC algorithm, and calculate the affine transformation matrix between the two images. The hyperspectral image is then resampled according to the calculated affine transformation matrix to spatially align it with the RGB image. Next, a first preprocessing operation is performed. Noise filtering can be performed by applying a mean-mode filter to both the aligned RGB and hyperspectral images. Image enhancement can involve contrast-limited adaptive histogram equalization of the RGB image and linear stretching of each band of the hyperspectral image. Geometric correction can be performed to correct distortion in the original image using camera calibration parameters. Illumination correction can be achieved by applying the Retinex algorithm to the RGB image and the gray-world algorithm to the hyperspectral image. After these steps, the RGB and hyperspectral images, after position calibration and the first preprocessing operation, are obtained for subsequent image analysis and fusion.
[0043] S20, determine at least one target local region and at least one target tongue feature within the target local region in the RGB image and hyperspectral image.
[0044] The identification of at least one target local region and at least one target tongue feature within that region refers to recognizing tongue areas of specific diagnostic significance, such as the tip, middle, root, and sides of the tongue, and the specific tongue features exhibited within these regions, such as spots, cracks, tongue coating, and veins, in the acquired RGB and hyperspectral images. This can be achieved using image segmentation techniques, such as edge-, color-, or texture-based segmentation algorithms, to identify the tongue and divide it into regions. Alternatively, feature detection algorithms, such as spot detection, texture analysis, edge detection, or morphological analysis, can be used to identify and locate specific features. The main purpose is to focus subsequent information processing on these local regions and features that are crucial for diagnosis.
[0045] In some embodiments, determining the target local region and target tongue features can specifically involve analyzing image content to identify specific regions on the tongue and visible features within those regions. This provides a focus for subsequent image fusion.
[0046] In some embodiments, the step of determining at least one target local region and at least one target tongue feature within the target local region in the RGB image and hyperspectral image specifically includes: S201, using an edge detection-based tongue segmentation algorithm, the tongue region is segmented to obtain multiple target local regions; S202, Perform feature detection algorithm processing on each target local region to obtain the type information of the target tongue image features, wherein the feature detection algorithm includes at least one of spot detection, texture analysis, and edge detection.
[0047] In some embodiments, an edge-detection-based tongue segmentation algorithm is used to segment the tongue region. The boundaries of the object are determined by identifying areas in the image where pixel values change drastically. In tongue image analysis, this algorithm can effectively separate the tongue region from the background because the tongue and its surroundings typically exhibit significant color or texture edge differences. This segmentation can yield multiple target local regions.
[0048] These local regions can be predefined anatomical divisions of the tongue or regions with similar attributes automatically segmented based on image features. Feature detection algorithms are applied to each target local region. These algorithms are used to identify and extract specific visual patterns in the image, corresponding to various clinical features of the tongue. The feature detection algorithms used here include at least one of speckle detection, texture analysis, and edge detection. Specifically, speckle detection algorithms are used to identify small regions in the image, suitable for detecting punctate lesions on the tongue. Texture analysis algorithms are used to quantify the texture attributes of image regions, suitable for assessing the thickness and texture of the tongue coating, as well as cracks on the tongue surface. Edge detection algorithms, in addition to segmentation, can also be used to identify morphological features of the tongue edges, such as teeth marks and abnormalities in the tongue contour. By combining these algorithms, appropriate algorithms or combinations can be selected based on the characteristics of different local regions and the types of features to be detected, thereby improving the targeting and accuracy of feature extraction.
[0049] In some embodiments, based on the acquired, calibrated, and preprocessed RGB and hyperspectral images, the process of determining the target local region and target tongue features is further refined. First, an edge-detection-based tongue segmentation algorithm is used to process the image. This algorithm analyzes the edge information of the image to accurately delineate the outline of the tongue, distinguishing it from non-tongue regions. This segmentation effectively narrows the scope of subsequent analysis, eliminating interference from background noise and irrelevant information on feature detection, allowing the analysis to focus more on the tongue itself. After segmentation, the tongue region is divided into multiple target local regions. This division considers the physiological and pathological differences of different parts of the tongue, laying the foundation for subsequent refined analysis.
[0050] Next, for these segmented target local regions, this application flexibly employs at least one of several feature detection algorithms, such as spot detection, texture analysis, and edge detection, based on the characteristics of the region and the potential types of tongue image features. For example, spot detection is prioritized in areas where punctate lesions may appear, texture analysis is emphasized in areas where tongue coating or cracks need to be evaluated, and edge detection is applied in areas where tongue shape or teeth marks are of interest. This targeted, multi-algorithm collaborative feature detection approach can capture various subtle features of the tongue image more comprehensively and accurately, including information on color, shape, and texture. Compared to performing global or single-algorithm feature analysis directly on the entire tongue image, the target local regions and target tongue image feature information determined in this way significantly improve the accuracy and reliability of feature extraction. This precise and comprehensive feature information provides high-quality input for subsequent generation of guiding data based on dominant modality information, extraction of target data, and asymmetric fusion processing, enabling the final fused tongue image to more accurately express key local tongue image features, thereby improving the effectiveness of the overall tongue image analysis method.
[0051] As a specific implementation method, Canny edge detection can be performed on the acquired RGB and hyperspectral images to extract edge information. Then, based on the extracted edge information, active contour models or graph cut algorithms are used to segment the tongue region, separating it from the background to obtain a tongue mask. According to tongue anatomy, the segmented tongue region can be further divided into multiple target local regions, such as the tip, middle, root, left, and right sides. For the tip region, a speckle detection algorithm based on the LoG operator can be applied to identify small red dots that may appear in this area. For the middle region, texture analysis algorithms such as gray-level co-occurrence matrix or local binary pattern can be used to quantify the texture features of the tongue coating and assess its thickness and texture. In the side regions, edge detection algorithms combined with morphological operations can be used again to detect and analyze the serrated structure of the tongue edges to identify teeth marks. In this way, appropriate feature detection algorithms are selected and applied for different local regions and the types of features of interest, thereby obtaining type information about the target tongue image features within these regions.
[0052] By utilizing an edge-detection-based tongue segmentation algorithm, this application effectively separates the tongue from the background, reducing interference from irrelevant regions and thus improving the accuracy of subsequent feature detection. Furthermore, for each target local region, at least one of several feature detection algorithms, including speckle detection, texture analysis, and edge detection, is employed for processing. This allows for the selection of appropriate algorithms based on different regions and feature types, overcoming the limitations of relying on a single algorithm. Therefore, this application can extract various subtle feature information of the tongue image more comprehensively and accurately, providing a more reliable and precise data foundation for subsequent image fusion and contributing to the generation of a fused tongue image that more accurately represents key local features.
[0053] S30, targeting the local area of the target and the characteristics of the target tongue image, based on a preset medical knowledge rule base, selects one of the RGB image information and hyperspectral image information to determine the dominant modality information used to characterize the characteristics of the target tongue image.
[0054] For the target local region and target tongue features, based on a preset medical knowledge rule base, the dominant modality information used to characterize the target tongue features is selected from RGB image information and hyperspectral image information. This step is the core of achieving differentiated fusion. By using the medical knowledge rule base to determine the importance of different modal information in characterizing specific regions and features, the selection of the dominant modality is provided, making the fusion process more in line with medical cognition.
[0055] In some embodiments, generating guiding data based on dominant modality information and extracting target data related to the target tongue image features from RGB image information and hyperspectral image information refers to generating auxiliary data, such as structural guidance maps, color guidance maps, or spectral index maps, from the dominant modality image information to guide the fusion process, based on the determined dominant modality. These data contain key information about the dominant modality. Simultaneously, data that spatially corresponds to the target tongue image features and needs to be fused or enhanced, such as spectral feature vector maps or fine texture and color information, is extracted from non-dominant modality image information. This is primarily to provide input for the asymmetric fusion algorithm, clarifying the direction and content of the fusion.
[0056] In some implementations, for the target local region and the target tongue image features, the dominant modality information used to characterize the target tongue image features is selected from RGB image information and hyperspectral image information. The dominant modality information can be determined by comparing the overall information content of the RGB image and the hyperspectral image, for example, by calculating the Shannon entropy of the image, and selecting the modality corresponding to the image with a larger entropy value as the dominant modality. In this way, it can be preliminarily determined which modality is more suitable for characterizing the features based on the quality of the image itself.
[0057] Based on this, step S30 can be specifically expanded as follows: S301, For the target local area and target tongue image features, calculate the RGB image structural feature evaluation value obtained based on RGB image information, and the hyperspectral image spectral feature evaluation value obtained based on hyperspectral image information; S302, compares the evaluation values of the structural features of the RGB image and the evaluation values of the spectral features of the hyperspectral image, and determines the dominant mode of the image data by the evaluation value with the larger value. S303, Obtain the dominant modality of the knowledge base corresponding to the target local area and the target tongue features from the preset medical knowledge rule base; S304, determine whether the dominant modality of the knowledge base is consistent with the dominant modality of the image data. If they are inconsistent, obtain the knowledge base weight of the target local region from the preset medical knowledge rule base, and determine the data weight based on the preprocessing quality of the RGB image information and the hyperspectral image information. Then, perform weighted fusion of the dominant modality of the knowledge base based on the knowledge base weight, and perform weighted fusion of the dominant modality of the image data based on the data weight to obtain the target dominant modality information. S304, if consistent, then the dominant modality of the knowledge base shall be used as the dominant modality information in the RGB image information and hyperspectral image information to characterize the features of the target tongue image.
[0058] Among them, the RGB image structural feature evaluation value refers to the numerical value that evaluates the ability or quality of an RGB image to express structural features in a local target region, and can be calculated using indicators such as structural similarity index, edge intensity, or texture contrast. The hyperspectral image spectral feature evaluation value refers to the numerical value that evaluates the ability or quality of a hyperspectral image to express spectral features in a local target region, and can be calculated using indicators such as spectral angle mapping, spectral information divergence, or signal-to-noise ratio.
[0059] Image data-dominant modality refers to the image modality that is more suitable for characterizing the features of the target tongue image, determined based on the evaluation results of the image data itself.
[0060] A pre-defined medical knowledge rule base refers to a collection of medical knowledge rules related to tongue features and local regions, which can be stored in the form of lookup tables or expert system rule sets. The dominant modality of the knowledge base refers to the image modality recommended by medical knowledge based on the pre-defined medical knowledge rule base, targeting specific local regions and tongue features.
[0061] Knowledge base weight refers to a numerical value used to measure the credibility or importance of a pre-defined medical knowledge rule base in a specific target local region and target tongue image feature. It is determined based on factors such as the depth of research or clinical validation of the region or feature in medical knowledge, and can also be obtained through expert scoring or statistical analysis based on historical data. Data weight refers to a numerical value used to measure the preprocessing quality or credibility of RGB image information and hyperspectral image information in a specific target local region. It is determined based on preprocessing quality indicators such as image signal-to-noise ratio, sharpness, or registration accuracy, and can also be obtained based on the direct output of image quality assessment algorithms.
[0062] When the dominant modality of the knowledge base and the dominant modality of the image data are inconsistent, the tendency of the two modalities is weighted by comprehensively considering the knowledge base weight and the data weight, thus obtaining the final dominant modality information. This can be achieved using algorithms such as weighted averaging or weighted voting. The dominant modality information is the one that is finally determined after comprehensive judgment and used to characterize the features of the target tongue image.
[0063] For example, when determining the dominant mode of red dot features in the tongue tip region, the structural similarity index of the RGB image of that region can be calculated as the RGB image structural feature evaluation value, and the spectral reflectance contrast of the hyperspectral image of that region in a specific band can be calculated as the hyperspectral image spectral feature evaluation value. By comparing these two evaluation values, the mode with the larger evaluation value is determined as the dominant mode of the image data.
[0064] Simultaneously, a pre-defined medical knowledge rule base is queried. Based on the "tip of the tongue" region and the "red dot" feature, the dominant modality recommended by the knowledge base is obtained. For example, the rule base might recommend "RGB" as the dominant modality. It is then determined whether the dominant modality of the knowledge base ("RGB") is consistent with the dominant modality of the image data. If they are consistent, the final dominant modality is determined to be "RGB". If the dominant modality of the image data is "hyperspectral" (e.g., the quality of an RGB image in this region is very poor), then the two are inconsistent. In this case, the knowledge base weight for the tongue tip region is obtained from the rule base, for example, 0.8, indicating that the medical knowledge's judgment on the tongue tip region has high credibility.
[0065] Simultaneously, data weights are determined based on the sharpness score of the RGB image and the signal-to-noise ratio score of the hyperspectral image in the region; for example, RGB data has a weight of 0.3, and hyperspectral data has a weight of 0.9. Weighted fusion can be performed using a weighted voting method; for example, the knowledge base favors RGB with a weight of 0.8, and the image data favors hyperspectral data with a weight of 0.9. Because the data weights are higher than the knowledge base weights, the final dominant modality may favor hyperspectral data, or a more complex weighted calculation may be used to obtain the final result. In this way, even if the image data evaluation results do not match medical knowledge, a comprehensive judgment can be made based on their respective credibility to obtain more reliable dominant modality information.
[0066] The above method improves the accuracy of dominant mode information determination. First, the structural feature evaluation value of the RGB image and the spectral feature evaluation value of the hyperspectral image are calculated separately, and their values are compared to determine the dominant mode of the image data. This step considers the characteristics of the two image modes and uses the information in the image data itself to preliminarily determine which mode has a higher information contribution.
[0067] Then, the dominant modality corresponding to the target local region and target tongue features is obtained from a pre-defined medical knowledge rule base. Introducing a medical knowledge rule base can reduce the risk of inaccurate judgments that may result from relying solely on image data and improve the reliability of the dominant modality information determination.
[0068] Next, it is determined whether the dominant modality of the knowledge base is consistent with the dominant modality of the image data. If they are consistent, the dominant modality of the knowledge base is directly used as the final dominant modality information. This indicates that the judgment results of the image data and medical knowledge are consistent in this case and can be directly used as the final result.
[0069] If the dominant modality of the knowledge base is inconsistent with the dominant modality of the image data, weighted fusion processing is required. Knowledge base weights for the target local region are obtained from a pre-defined medical knowledge rule base, and data weights are determined based on the preprocessing quality of the RGB and hyperspectral image information. Knowledge base weights represent the credibility of the medical knowledge, while data weights represent the quality level of the image data. Through weighted fusion, the credibility of both medical knowledge and image data can be comprehensively evaluated, resulting in more accurate dominant modality information.
[0070] Specifically, the dominant modality of the knowledge base is weighted according to its weight, and the dominant modality of the image data is weighted according to its weight, ultimately yielding the target dominant modality information. This weighted fusion approach reconciles the differences between medical knowledge and image data, making the final dominant modality information more reasonable. This comprehensive method of determining the dominant modality provides a more reliable input for subsequent generation of guiding data based on dominant modality information, extraction of target data, and asymmetric fusion. This allows the final local fusion result to more accurately enhance the expression of the target tongue image features, thereby generating a higher-quality fused tongue image. This overcomes the problem in related technologies where global fusion strategies fail to consider the heterogeneity of local information contributions, leading to the dilution, weakening, or submersion of key local diagnostic information.
[0071] In some embodiments, by determining knowledge base weights and data weights, and using these weights to perform weighted fusion of the knowledge base dominant modality and the image data dominant modality, a more accurate target dominant modality information determination can be obtained. Specifically, the steps of obtaining the knowledge base weights of the target local region from the preset medical knowledge rule base, and determining the data weights based on the preprocessing quality of the RGB image information and the hyperspectral image information, specifically include: Obtain tongue image feature type information related to the target local region; Based on the acquired tongue image feature type information, an evaluation criterion is selected to quantify the reliability of the target local region; Based on the selected evaluation criteria, the reliability of the target local region is quantified to obtain a value for the reliability of the target local region. The reliability value of the target local region is converted into the knowledge base weight using a preset conversion method; and, Calculate the quality index values of the first preprocessing operation of the RGB image information and the hyperspectral image information in the target local area; Based on the quality index value, the credibility of the image data is quantified, and the credibility is converted into the data weight through a second preset conversion method.
[0072] The first preset conversion method refers to a rule or function that maps one numerical range to another, which can be implemented using lookup tables, linear functions, nonlinear functions, or piecewise functions. The quality index value refers to a quantified value used to measure the quality of image data within a specific region, which can be implemented using signal-to-noise ratio, sharpness, contrast, artifact severity, or the detection confidence of specific features. Quantifying the credibility of image data refers to converting the image data's quality index value into a value representing the degree of trustworthiness of the data in the current application scenario, which can be achieved by normalizing the quality index value to a specific range (e.g., 0 to 1). The second preset conversion method refers to a rule or function that converts the image data's credibility value into data weights, which can be implemented using lookup tables, linear functions, nonlinear functions, or piecewise functions.
[0073] In some embodiments, information on tongue feature types related to the target local region is first acquired to provide a basis for selecting appropriate evaluation criteria. Different tongue feature types have different manifestations and diagnostic significance in different regions, thus requiring targeted evaluation of the reliability of that region. Then, based on the acquired tongue feature type information, an evaluation criterion for quantifying the reliability of the target local region is selected. This selection process ensures the relevance and effectiveness of the evaluation. Next, based on the selected evaluation criterion, the reliability of the target local region is quantified to obtain an objective reliability value. This value reflects the reliability of the knowledge base guidance for that local region under the current feature type. Subsequently, this reliability value is converted into a knowledge base weight through a preset transformation method. The higher the reliability value, the larger the converted knowledge base weight, indicating a stronger guiding role of the knowledge base for that region and feature. Simultaneously, the quality index values of the first preprocessing operation of RGB image information and hyperspectral image information within the target local region are calculated. These quality indices reflect the quality status of the image data itself, such as the presence of noise, blur, or distortion. Based on these quality index values, the credibility of the image data is quantified. The higher the credibility, the more reliable the information in that region. Finally, the credibility is converted into data weights using a second preset transformation method. Higher credibility results in greater data weights, indicating a greater contribution from the image data in that region. Through these steps, this application provides a weight determination method that comprehensively considers tongue image feature type, target local region reliability, and image data quality. This method is combined with a step of determining whether the dominant modality of the knowledge base and the dominant modality of the image data are consistent. When they are inconsistent, these adaptively determined knowledge base weights and data weights are used to perform a weighted fusion of the knowledge base dominant modality and the image data dominant modality, thereby obtaining more accurate target dominant modality information. This combination makes the dominant modality determination process more refined and adaptive, better balancing the prior knowledge of the knowledge base and the actual information of the image data. Especially when there is a conflict between the two, the influence of each can be adjusted according to the specific situation, improving the accuracy of dominant modality determination.
[0074] For example, in one specific implementation, suppose the target local region is the tip of the tongue, and the target tongue image feature is "red dot on the tip of the tongue." First, the tongue image feature type information is obtained as "red dot on the tip of the tongue." Based on this information, an evaluation criterion for quantifying the reliability of the tongue tip region is selected. For example, the reliability of the knowledge base in this region can be evaluated based on physiological indicators related to the red dot on the tip of the tongue, such as the blood vessel distribution density and tissue thickness. Suppose the reliability value of the tongue tip region after evaluation is 0.7. Through a preset conversion method, such as a lookup table or function, 0.7 is converted into a knowledge base weight, for example, a weight value of 0.6. Simultaneously, the quality index values of the RGB image and hyperspectral image within the tongue tip region after a first preprocessing (such as denoising and enhancement) are calculated. For example, the sharpness index of the RGB image is 0.85, and the signal-to-noise ratio index of the hyperspectral image is 0.8. Based on these quality index values, the credibility of the image data is quantified, for example, taking the average value of 0.825 as the credibility value. Through a second preset conversion method, such as another lookup table or function, 0.825 is converted into a data weight, for example, a weight value of 0.4. In this way, when the dominant modality of the knowledge base and the dominant modality of the image data are inconsistent on the feature of the red dot on the tip of the tongue, the calculated knowledge base weight of 0.6 and the data weight of 0.4 can be used for weighted fusion to determine the final target dominant modality information.
[0075] Through the aforementioned technical means, this application provides a method for adaptively determining knowledge base weights and data weights based on tongue image feature type, target local region reliability, and image data quality when the dominant modality of the knowledge base and the dominant modality of the image data are inconsistent. This method enables the subsequent weighted fusion process to more accurately balance knowledge base information and image data information, thereby obtaining more accurate target dominant modality information and improving the accuracy and reliability of tongue image feature representation.
[0076] S40 generates guidance data based on dominant modality information and extracts target data related to the target tongue image features from RGB image information and hyperspectral image information.
[0077] Guiding data is generated based on dominant modality information, and target data related to the target tongue image features are extracted from RGB image information and hyperspectral image information. The generation of guiding data based on dominant modality information can effectively guide the fusion direction of target data, highlight the features of the dominant modality, and at the same time extract target data related to the target tongue image features, ensuring that the fused information is effective information related to diagnosis.
[0078] S50, the target data is processed by an asymmetric fusion algorithm using the guiding data to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to enhance the expression of the target tongue image features in the fusion result of the target local region.
[0079] The target data is processed by an asymmetric fusion algorithm using the guiding data to obtain local fusion results. The asymmetric fusion algorithm is used to enhance the expression of the target tongue image features in the fusion results of the target local region. The asymmetric fusion algorithm can selectively enhance or suppress information of different modalities according to the guiding data, thereby enhancing the expression of the target tongue image features in the fusion results and avoiding information overload.
[0080] S60, based on the local fusion results of each target local region, generates a fused tongue image.
[0081] In some embodiments, by combining the determination of local areas and target tongue features on the tongue surface, the determination of the dominant mode based on medical knowledge and image analysis, and the asymmetric fusion based on the dominant mode, adaptive and differentiated fusion for different regions and features of the tongue surface is achieved, thereby achieving the effect of high-fidelity preservation and enhancement of key local diagnostic information.
[0082] Specifically, the method of this application first acquires RGB and hyperspectral images of the same subject and performs image analysis. Next, it identifies at least one target local region requiring focused attention and at least one target tongue feature within that region in these images. Then, for these specific regions and features, based on a pre-defined medical knowledge rule base, it determines whether RGB or hyperspectral information is more effective in representing the feature, thereby determining the dominant modality information. Based on the determined dominant modality information, guiding data for fusion is generated from the dominant modality image information, and target data related to the target tongue feature is extracted from the other modality image information. Subsequently, an asymmetric fusion algorithm, guided by the guiding data, is used to process the target data to obtain a local fusion result. This algorithm is specifically designed to enhance the expression of the target tongue feature in the fusion result. Finally, the local fusion results of each target local region are integrated to generate the final fused tongue image. The entire process is a flow from acquiring raw data, to local identification and analysis, to local adaptive fusion based on the analysis results, and finally to the integration into a holistic image. The core lies in the differentiated and asymmetric processing of local regions and features.
[0083] In some embodiments, the steps of generating guidance data based on dominant modality information and extracting target data related to the target tongue features from the RGB image information and the hyperspectral image information specifically include: When RGB image information is used as the dominant modal information, the RGB image is filtered using a guided filter to generate a structured guided map that preserves edge details and smooths unstructured areas; and color space components are extracted as a color guided map; the structured guided map or the color guided map is used as guided data. Extract the spectral feature vector map from the hyperspectral image that corresponds to the local region of the target in the RGB image and is related to the tongue features of the target, as the target data.
[0084] The guiding filter employs a local linear model-based implementation. The structure guiding map refers to the image obtained after processing by the guiding filter, which preserves the edge information of the original image while smoothing unstructured regions. Color space components refer to the components obtained after converting an RGB image to other color spaces; this can be achieved by converting the RGB image to the HSV color space and extracting one or more of the H, S, and V components. The color guiding map refers to the extracted color space components. The spectral feature vector map is a vector set composed of the reflectance or absorptivity of each pixel in a hyperspectral image across different wavelengths; it can be achieved by extracting the spectral curves of pixels corresponding to the RGB image region from a hyperspectral cube.
[0085] In some embodiments, when RGB image information is the dominant mode, a guiding filter is used to filter the RGB image to generate a structure guiding map. This structure guiding map preserves edge details and smooths unstructured areas, thereby highlighting the structural information of the image. Secondly, color space components are extracted as color guiding maps. Color space components can provide color information of the image, thereby providing color guidance for subsequent fusion.
[0086] The reason for processing the RGB image to generate the guide map is that directly using the original RGB image may introduce noise or irrelevant details, affecting the fusion effect. By using guided filtering or extracting color space components, more useful information about the target tongue image features can be extracted, thereby improving the accuracy of the fusion.
[0087] Secondly, spectral feature vector maps corresponding to the target local region in the RGB image and related to the target tongue image features are extracted from the hyperspectral image as target data. The reason for extracting spectral feature vector maps from hyperspectral images is that they provide rich spectral information that reflects the biochemical composition and tissue structure of the tongue. By extracting spectral feature vector maps corresponding to the target local region, the spectral features of that region can be obtained, thus providing spectral information for subsequent fusion. Simultaneously, emphasis is placed on correlation with the target tongue image features to ensure that the extracted spectral features are relevant to the features of interest, avoiding the introduction of irrelevant information.
[0088] This application, by combining steps of determining the target region, target features, and dominant mode, selectively extracts data most beneficial to subsequent fusion from both modes based on the determined dominant mode (RGB). By extracting structural or color guidance information from RGB, the spatial resolution advantage of RGB can be utilized; by extracting spectral feature vector maps from hyperspectral data, the spectral resolution advantage of hyperspectral data can be utilized. This targeted data preparation provides high-quality input for subsequent asymmetric fusion, enabling the fusion result to better enhance the expression of the target tongue image features and solving the problem of diluted local key information caused by the globally uniform fusion strategy in existing technologies.
[0089] Therefore, by extracting structural or color information from RGB images as a guide, and extracting spectral information related to the target features from hyperspectral images as the target, targeted and high-quality input data is provided for subsequent asymmetric fusion. This data preparation method can fully utilize the spatial details of RGB images and the spectral features of hyperspectral images, enabling the subsequent fusion process to more effectively highlight and enhance the expression of the target tongue image features, thereby improving the accuracy of tongue image analysis.
[0090] In some embodiments, the step of using the guiding data to perform an asymmetric fusion algorithm on the target data to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to enhance the expression of the target tongue image features in the fusion result of the target local region, specifically includes: Using the spectral feature vector map as input to the guided filter and the structural guided map as the guided image, an asymmetric fusion algorithm is applied to align the output spectral features spatially with the RGB structure and enhance details; or, The spectral feature vector map is used as the input image for guided filtering, and the color guide map is used as the guide image. The brightness, saturation, or intensity of a specific color channel of the corresponding pixel in the color guide map is adjusted using hyperspectral feature values.
[0091] In some embodiments, after determining that RGB image information is the dominant mode and generating a structure guide map or color guide map as guide data, and a spectral feature vector map as target data, the asymmetric fusion stage is entered.
[0092] If a structural guide map is chosen as the guiding data, and the spectral feature vector map is used as the input image, with the structural guide map serving as the guiding image, guided filtering is performed. Guided filtering utilizes the spatial structure information of the RGB image provided by the structural guide map to guide the filtering process of the spectral feature vector map, ensuring that the output spectral features spatially align with the structure of the RGB image, preserving details in structural regions and smoothing in smooth regions. This approach combines hyperspectral spectral information with the spatial structure of RGB, presenting an image in the fusion result that possesses both spectral information and clear structure, which is beneficial for identifying and analyzing tongue image features related to structure.
[0093] If a color guide map is chosen as the guiding data, and the spectral feature vector map is used as the input image, then the color guide map serves as the guiding image. Hyperspectral feature values are used to adjust the color attributes of corresponding pixels in the color guide map, such as brightness, saturation, or the intensity of a specific color channel. This means that the spectral information of the hyperspectral image is used to modulate the color representation of the RGB image. For example, if the hyperspectral feature value of a pixel indicates an abnormality in a certain biochemical component, this hyperspectral feature value can be used to increase the intensity or saturation of the red channel of the corresponding pixel in the color guide map, highlighting the color change of that abnormal area in the fusion result.
[0094] Both approaches embody the idea of asymmetric fusion, utilizing guiding data generated by the dominant modality to direct the fusion of target data from non-dominant modalities, thereby enhancing the expression of target tongue features in the fusion result. This fusion method avoids information dilution that may result from simple superposition or averaging, and can more effectively combine the advantages of the two modalities. The preliminary step determines the dominant modality and generates appropriate guiding and target data based on it. The asymmetric fusion algorithm in this step is based on this prepared data. The determination of the dominant modality ensures that the fusion process is guided by a more reliable and representative modality, while the extraction of guiding and target data provides the necessary and relevant inputs for asymmetric fusion. This combination allows the entire method to select a suitable modality as the dominant one based on the characteristics of the tongue features and the region, and to adopt corresponding fusion strategies. This more effectively highlights and enhances the expression of target tongue features in the fusion result of local regions, solving the problem that a globally uniform fusion strategy cannot adapt to the heterogeneity of information in different regions of the tongue.
[0095] For example, in one implementation scenario, for a target local area on the tongue surface, such as the tongue coating area, and a target tongue image feature within that area, such as abnormal tongue coating color, the RGB image information is determined as the dominant modal information after preliminary steps, and a color guide map and a spectral feature vector map are generated. At this point, a second fusion method can be used. The spectral feature vector map of the tongue coating area extracted from the hyperspectral image is used as the input image for the guide filter. The color guide map of the tongue coating area extracted from the RGB image, such as the red channel intensity map of the RGB image, is used as the guide image. Hyperspectral feature values from the spectral feature vector map, such as the intensity of specific bands or spectral indices related to tongue coating components, are used to adjust the intensity of corresponding pixels in the color guide map. For example, if the hyperspectral feature value indicates the presence of yellow pigment deposition in the pixel, the intensity of the yellow-related channel of the corresponding pixel in the color guide map can be increased, thereby highlighting the abnormal yellow area of the tongue coating in the fusion result.
[0096] By using spectral feature vector maps as input to guided filtering and employing different fusion strategies based on different guided data (structural guided maps or color guided maps), this application effectively utilizes the spatial structure or color information of RGB images to guide the fusion of spectral features from hyperspectral images. When using a structural guided map, the spectral features in the fusion result are spatially aligned with the structure of the RGB image, and details are enhanced. This helps to clearly present structural tongue features with spectral characteristics, improving the recognition accuracy of these features. When using a color guided map and adjusting color attributes in conjunction with hyperspectral feature values, the fusion result can highlight color changes related to spectral characteristics, making these features more visually identifiable and reflecting differences in their biochemical composition. This asymmetric fusion method can selectively fuse hyperspectral information based on the information advantage of the RGB dominant mode, avoiding information dilution or distortion that may result from simple fusion. This more effectively strengthens the expression of target tongue features in the fusion result of local areas, improving the fusion image's ability to represent subtle local tongue features and providing a more accurate and information-rich image foundation for subsequent tongue analysis and diagnosis.
[0097] In some embodiments, the step of generating guidance data based on the dominant modality information and extracting target data related to the target tongue features from the RGB image information and the hyperspectral image information specifically includes: When hyperspectral image information is used as the dominant modal information, the spectral index map obtained by band operation on the hyperspectral image, or the principal component score map obtained by principal component analysis on the hyperspectral image, is used as the guiding data. The fine texture and color information corresponding to the target local region in the hyperspectral image and related to the target tongue features are extracted from the RGB image as target data.
[0098] Based on the aforementioned method for determining dominant modality information, a specific data preparation strategy is adopted when hyperspectral image information is determined to be the dominant modality of the current target local region and target tongue image features. Specifically, in order to fully leverage the advantages of hyperspectral information in spectral feature recognition when it is dominant, and to utilize the spatial details of RGB images, this application extracts guiding data that can represent key spectral information from hyperspectral images, and extracts corresponding spatial details from RGB images as target data. The spectral index map or principal component score map of hyperspectral images can effectively condense diagnostic spectral information in hyperspectral data, such as reflecting tongue coating thickness and sublingual vein fullness. Using these as guiding data, the spectral information can be used to guide the fusion of spatial details in RGB images during subsequent fusion processes.
[0099] Simultaneously, fine texture and color information corresponding to the target region in the hyperspectral image are extracted from the RGB image as target data. This information provides high spatial resolution details of the local area of the tongue surface. In this way, when hyperspectral information is the dominant modality, the fusion process no longer relies solely on the structural guidance of the RGB image. Instead, it utilizes the spectral advantages of the hyperspectral image itself to guide the expression of details in the RGB image, ensuring that the fusion result highlights the key spectral features that are significant in the hyperspectral image while preserving the spatial details of the RGB image. This allows for a more accurate reflection of the true physiological and pathological state of the local area of the tongue surface. This data preparation method, combined with the aforementioned method for determining the dominant modality, enables adaptive data selection and preparation based on the information advantages of different regions and features. This overcomes the limitations of a globally unified fusion strategy and improves the accuracy of expressing key local information.
[0100] Further, the target data is processed using the guiding data using an asymmetric fusion algorithm to obtain a local fusion result. The step of using the asymmetric fusion algorithm to enhance the expression of the target tongue image features in the fusion result of the target local region specifically includes: Using the fine texture and color information as input to the guided filter, and the spectral index map as the guiding image, an asymmetric fusion algorithm is applied to visually enhance abnormal regions indicated by the spectral index in the fused tongue image; or... The fine texture and color information are used as input to the guided filter, and the principal component score map is used as the guided image. The information of each local region after guided fusion processing is weighted and combined or selectively stitched together.
[0101] By utilizing hyperspectral image information as a guide, asymmetric fusion processing is performed on the texture and color information of RGB images to enhance the expression of target tongue features in the fusion result. Specifically, when hyperspectral image information is determined to be the dominant mode, a spectral index map or principal component score map is generated as guiding data, while fine texture and color information is extracted from the RGB image as target data. Subsequently, an asymmetric fusion algorithm is used to process these data.
[0102] One approach involves using fine texture and color information as input for guided filtering, and a spectral index map as the guide image. The spectral index map reflects spectral anomalies and their spatial distribution related to pathological changes. Using it as the guide image, the guided filtering process adjusts the filtering intensity of texture and color information based on the abnormal regions indicated by the spectral index map. This ensures that the texture and color information corresponding to these spectral anomaly regions are preserved or enhanced in the fusion result, thus visually highlighting these abnormal regions.
[0103] Another processing method involves using fine texture and color information as input for guided filtering, and the principal component score (PCS) map as the guide image. The PCS map captures the main variation patterns of the hyperspectral data and provides spectral structure information as guidance. Using the PCS map as guidance allows the fusion result to incorporate the spectral features of the hyperspectral data while preserving the texture and color of the RGB image. Based on this, the information from each local region after guided fusion is weighted and combined or selectively stitched together. This allows for processing based on the characteristics of different regions or the fusion effect, further optimizing the overall quality of the fused image and avoiding the potential impact of globally uniform processing on information representation. Through these two methods, the spectral information of the hyperspectral image guides the visual information representation of the RGB image, enabling the fused tongue image to accurately present the target tongue features, especially those features that are abnormal in the spectrum but difficult to detect in the RGB image. This improves the recognition of the target tongue features, allowing the fused image to accurately reflect the physiological and pathological state of the tongue surface, providing an image foundation for subsequent tongue image analysis and diagnosis.
[0104] On the other hand, such as Figure 2 As shown, this application provides a tongue image analysis system 100 based on the fusion of RGB images and spectral data. The system includes: The acquisition module 11 is used to acquire RGB images and hyperspectral images captured on the tongue surface of the same subject, and to obtain RGB image information and hyperspectral image information through image analysis. The region feature determination module 12 is used to determine at least one target local region and at least one target tongue image feature within the target local region in the RGB image and the hyperspectral image; The dominant mode determination module 13 is used to determine the dominant mode information for characterizing the target tongue image features based on a preset medical knowledge rule base, choosing between RGB image information and hyperspectral image information, for the target local area and target tongue image features. Data preparation module 14 is used to generate guidance data based on dominant modality information and extract target data related to the target tongue image features from RGB image information and hyperspectral image information; The local fusion module 15 is used to process the target data with an asymmetric fusion algorithm using the guiding data to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to enhance the expression of the target tongue image features in the fusion result of the target local region; Image generation module 16 is used to generate a fused tongue image based on the local fusion results of each target local region.
[0105] In some embodiments, this system can be implemented as an integrated system. The acquisition module may include an RGB camera and a hyperspectral camera, configured to synchronously or near-synchronously capture tongue images of the same subject and transmit the captured raw image data to a processing unit. A region feature determination module, as part of a software program running on the processing unit, is designed to execute image segmentation algorithms to identify tongue regions and further execute feature detection algorithms, such as speckle detection or texture analysis, within the tongue regions to determine target local regions and target tongue features. A dominant modality determination module, also as part of a software program running on the processing unit, is configured to access a preset medical knowledge rule base stored in a storage medium and, based on the logic in the rule base and combining region and feature information, determine whether RGB information or hyperspectral information is more suitable to characterize the feature, thereby determining the dominant modality information. The data preparation module, also part of the software program, is configured to generate guiding data by calling image processing functions based on the dominant modality information. For example, when RGB information is the dominant modality, guiding filtering or color space conversion can be performed on the RGB image to generate a guiding map. Simultaneously, it extracts target data related to the target features from the original image data, such as extracting the spectral vector of the corresponding region in the hyperspectral image. The local fusion module can implement asymmetric fusion algorithms. For example, when the guiding data is an RGB structural guiding map and the target data is a hyperspectral spectral vector map, this module can execute a guiding filtering algorithm, using the spectral vector map as input and the structural guiding map as the guiding image for processing, to achieve spatial alignment and enhancement of spectral features. The image generation module can be configured to receive the local fusion results from the local fusion module and integrate these results, for example, through stitching or weighted overlay, to ultimately generate a complete fused tongue image.
[0106] The entire processing flow can be completed on one or more processing units, which may include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), or a digital signal processor (DSP). Similar or identical parts between the various embodiments in this specification can be referred to mutually. In particular, the embodiments are basically similar to the method embodiments, so the descriptions are relatively simple, and relevant details can be found in the descriptions within the method embodiments.
[0107] The embodiments of the present invention described above do not constitute a limitation on the scope of protection of the present invention.
Claims
1. A tongue image analysis method based on the fusion of RGB images and spectral data, characterized in that, include: RGB and hyperspectral images of the tongue surface of the same subject were captured separately, and RGB and hyperspectral image information was obtained through image analysis. Determine at least one target local region and at least one target tongue feature within the target local region in the RGB image and hyperspectral image; Based on a preset medical knowledge rule base, for the target local region and the target tongue image features, the dominant modality information used to characterize the target tongue image features is selected from the RGB image information and the hyperspectral image information. Guidance data is generated based on the dominant modality information, and target data related to the target tongue image features is extracted from the RGB image information and the hyperspectral image information. The target data is processed by an asymmetric fusion algorithm using the guiding data to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to enhance the expression of the target tongue image features in the fusion result of the target local region; Based on the local fusion results of each target local region, a fused tongue image is generated.
2. The tongue image analysis method according to claim 1, characterized in that, After acquiring RGB and hyperspectral images of the tongue surface of the same subject, the process also includes: The RGB image and hyperspectral image are positionally calibrated to ensure that the coordinates of each point on the tongue surface in the spatial coordinate system remain consistent in the RGB image and hyperspectral image. The first preprocessing operation is performed on the position-calibrated RGB image and hyperspectral image, the first preprocessing operation including at least one of noise filtering, image enhancement, geometric correction and illumination correction.
3. The tongue image analysis method according to claim 2, characterized in that, The step of determining at least one target local region in the RGB image and hyperspectral image, and at least one target tongue feature within the target local region, specifically includes: An edge-detection-based tongue segmentation algorithm is used to segment the tongue region and obtain multiple target local regions. The feature detection algorithm is applied to each local region of the target to obtain the type information of the target tongue image features, wherein the feature detection algorithm includes at least one of spot detection, texture analysis and edge detection.
4. The tongue image analysis method according to claim 1, characterized in that, The step of determining the dominant modality information for characterizing the target tongue features, based on a preset medical knowledge rule base, from the RGB image information and the hyperspectral image information, specifically includes: For the target local region and the target tongue image features, calculate the RGB image structural feature evaluation value obtained based on the RGB image information, and the hyperspectral image spectral feature evaluation value obtained based on the hyperspectral image information; The structural feature evaluation value of the RGB image and the spectral feature evaluation value of the hyperspectral image are compared, and the one with the larger evaluation value is determined as the dominant mode of the image data. The dominant knowledge base modality corresponding to the target local region and the target tongue image features is obtained from the preset medical knowledge rule base. Determine whether the dominant mode of the knowledge base is consistent with the dominant mode of the image data. If they are inconsistent, obtain the knowledge base weight of the target local region from the preset medical knowledge rule base, determine the data weight based on the preprocessing quality of the RGB image information and the hyperspectral image information, perform weighted fusion of the dominant mode of the knowledge base based on the knowledge base weight, and perform weighted fusion of the dominant mode of the image data based on the data weight to obtain the target dominant mode information. If they match, the dominant modality of the knowledge base is used as the dominant modality information in the RGB image information and the hyperspectral image information to characterize the features of the target tongue image.
5. The tongue image analysis method according to claim 4, characterized in that, The steps of obtaining the knowledge base weight of the target local region from the preset medical knowledge rule base, and determining the data weight based on the preprocessing quality of the RGB image information and the hyperspectral image information, specifically include: Obtain tongue image feature type information related to the target local region; Based on the acquired tongue image feature type information, an evaluation criterion is selected to quantify the reliability of the target local region; Based on the selected evaluation criteria, the reliability of the target local region is quantified to obtain a value for the reliability of the target local region. The reliability value of the target local area is converted into the knowledge base weight through a preset conversion method; as well as, Calculate the quality index values of the first preprocessing operation of the RGB image information and the hyperspectral image information in the target local area; Based on the quality index value, the credibility of the image data is quantified, and the credibility is converted into the data weight through a second preset conversion method.
6. The tongue image analysis method according to claim 1, characterized in that, The step of generating guidance data based on the dominant modality information and extracting target data related to the target tongue features from the RGB image information and the hyperspectral image information specifically includes: When RGB image information is used as the dominant modal information, the RGB image is filtered using a guided filter to generate a structured guided map that preserves edge details and smooths unstructured areas; and color space components are extracted as a color guided map; the structured guided map or the color guided map is used as guided data. Extract the spectral feature vector map from the hyperspectral image that corresponds to the local region of the target in the RGB image and is related to the tongue features of the target, as the target data.
7. The tongue image analysis method according to claim 6, characterized in that, The step of using the guiding data to process the target data using an asymmetric fusion algorithm to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to enhance the expression of the target tongue image features in the fusion result of the target local region, specifically includes: Using the spectral feature vector map as input to the guided filter and the structural guided map as the guided image, an asymmetric fusion algorithm is applied to align the output spectral features spatially with the RGB structure and enhance details; or, The spectral feature vector map is used as the input image for guided filtering, and the color guide map is used as the guide image. The brightness, saturation, or intensity of a specific color channel of the corresponding pixel in the color guide map is adjusted using hyperspectral feature values.
8. The tongue image analysis method according to claim 1, characterized in that, The step of generating guidance data based on the dominant modality information and extracting target data related to the target tongue features from the RGB image information and the hyperspectral image information specifically includes: When hyperspectral image information is used as the dominant modal information, the spectral index map obtained by band operation on the hyperspectral image, or the principal component score map obtained by principal component analysis on the hyperspectral image, is used as the guiding data. The fine texture and color information corresponding to the target local region in the hyperspectral image and related to the target tongue features are extracted from the RGB image as target data.
9. The tongue image analysis method according to claim 8, characterized in that, The step of using the guiding data to process the target data using an asymmetric fusion algorithm to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to enhance the expression of the target tongue image features in the fusion result of the target local region, specifically includes: Using the fine texture and color information as input to the guided filter, and the spectral index map as the guiding image, an asymmetric fusion algorithm is applied to visually enhance abnormal regions indicated by the spectral index in the fused tongue image; or... The fine texture and color information are used as input to the guided filter, and the principal component score map is used as the guided image. The information of each local region after guided fusion processing is weighted and combined or selectively stitched together.
10. A tongue image analysis system based on the fusion of RGB images and spectral data, characterized in that, The system includes: The acquisition module is used to acquire RGB and hyperspectral images of the tongue surface of the same subject, and obtain RGB and hyperspectral image information through image analysis. A region feature determination module is used to determine at least one target local region and at least one target tongue image feature within the target local region in the RGB image and hyperspectral image; The dominant mode determination module is used to determine the dominant mode information for characterizing the target tongue image features by selecting one from the RGB image information and the hyperspectral image information based on a preset medical knowledge rule base, for the target local region and the target tongue image features. The data preparation module is used to generate guidance data based on the dominant modality information and extract target data related to the target tongue image features from the RGB image information and the hyperspectral image information; The local fusion module is used to process the target data using the guiding data with an asymmetric fusion algorithm to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to enhance the expression of the target tongue image features in the fusion result of the target local region; An image generation module is used to generate a fused tongue image based on the local fusion results of each of the target local regions.
Citation Information
Patent Citations
Hyperspectral image super-resolution reconstruction method and device and electronic equipment
CN113139902A
Attention mechanism guided matching association hyperspectral and RGB video fusion tracking method
CN117689692A
Optical filter sampling and fusion-based video hyperspectral imaging method and system, and medium
WO2024178879A1
Cited By
Medical hyperspectral image enhancement method and system based on multi-domain fusion
CN121582077A