A tongue image analysis method and system based on RGB image and spectral data fusion

By selecting dominant modality information for asymmetric fusion in tongue image analysis, the problem of information dilution caused by the global unified fusion strategy is solved, and the comprehensiveness and accuracy of tongue image analysis are improved, especially in the feature expression of different regions of the tongue surface.

CN120876484BActive Publication Date: 2025-12-12HEFEI YUNZHEN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511394221.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-12-12
Estimated Expiration
2045-09-28

AI Technical Summary

Technical Problem

In existing technologies, the global unified fusion strategy for RGB images and hyperspectral images in tongue image analysis results in an uneven information contribution from different regions of the tongue surface, which fails to fully express local lesions or key physiological characteristics and affects diagnostic accuracy.

Method used

By identifying the target local region and features in RGB and hyperspectral images, selecting the dominant modality information based on a preset medical knowledge rule base, and using an asymmetric fusion algorithm to enhance the expression of the target tongue image features, a local fusion result is generated.

Benefits of technology

It improves the comprehensiveness and accuracy of tongue image analysis, better preserves and enhances local diagnostic information, and improves the accuracy of feature extraction and recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876484B_ABST
    Figure CN120876484B_ABST
Patent Text Reader

Abstract

The application discloses a tongue image analysis method and system based on RGB image and spectral data fusion. The method comprises the following steps: acquiring RGB image information and hyperspectral image information; determining at least one target local area in the RGB image and the hyperspectral image, and a target tongue image feature in the target local area; based on a preset medical knowledge rule base, determining dominant modal information for representing the target tongue image feature in the RGB image information and the hyperspectral image information; generating guide data based on the dominant modal information; extracting target data related to the target tongue image feature from the RGB image information and the hyperspectral image information; adopting the guide data to perform an asymmetric fusion algorithm processing on the target data to obtain a local fusion result; and finally generating a fusion tongue image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to tongue image processing and analysis technology, and in particular to a tongue image analysis method and system based on RGB image and spectral data fusion. BACKGROUND

[0002] Tongue observation can be used in traditional Chinese medicine to assess health status by observing tongue shape, tongue fur color and moisture, and other characteristics, reflecting the function of viscera and the state of blood and body fluid, and thus is an important auxiliary means for traditional Chinese medicine syndrome differentiation and sub-health and chronic disease monitoring.

[0003] Due to the different experiences of doctors, it cannot be standardized, and in the related art, a tongue image intelligent detection method is adopted, which first captures the RGB image and hyperspectral image of the tongue of the examinee, wherein the RGB image mainly records the color and texture information of the tongue surface, has high spatial resolution, and can intuitively display the macroscopic characteristics of the tongue surface; and the hyperspectral image records the spectral reflection information of different positions of the tongue surface in multiple continuous narrow wave bands, and its core advantage lies in that it can provide continuous spectral curve of each pixel point, and reveal the subtle differences of biochemical components that cannot be distinguished by the naked eye. After receiving the two types of image data, a fusion algorithm is used to combine the spatial details of the RGB image and the spectral characteristics of the hyperspectral image to generate a fused tongue image. Based on the fused image, the system further analyzes the characteristics such as tongue color, tongue fur and tongue shape, and outputs a detection report.

[0004] However, the tongue itself is not a homogeneous observation object, and the tissue structure, blood vessel distribution and manifestation in pathological state of different physiological regions are significantly different. This leads to the fact that the information contribution and effectiveness of the RGB image and the hyperspectral image in representing the characteristics of these different regions are each focused and often uneven. When a global and unchangeable fusion strategy is adopted in the image fusion link, the heterogeneity of the information contribution of the different regions of the tongue surface cannot be fully considered, and different types of local characteristics cannot be adapted to different degrees of dependence on the two modal data. As a result, for those local lesions or key physiological characteristics whose feature expression is very significant in one modality but relatively weak in another modality, the information expression in the fused image may be diluted, weakened or even submerged, resulting in the problem that the local information is not accurately and sufficiently expressed in the fused image. When this tongue image analysis method based on RGB image and hyperspectral image fusion is applied to a scene that requires fine analysis and differential diagnosis of the subtle physiological or pathological characteristics of specific functional partitions or specific anatomical structures of the tongue surface, the challenge is particularly prominent. SUMMARY

[0005] The present application provides a tongue image analysis method and system based on RGB image and spectral data fusion, which has the advantages of improving the comprehensiveness and accuracy of tongue image analysis.

[0006] In one aspect, the application provides a tongue image analysis method based on RGB image and spectral data fusion, comprising:

[0007] Obtaining an RGB image and a hyperspectral image captured respectively on the tongue surface of the same subject, and obtaining RGB image information and hyperspectral image information through image analysis;

[0008] Determining at least one target local area in the RGB image and the hyperspectral image, and at least one target tongue image feature in the target local area;

[0009] For the target local area and the target tongue image feature, based on a preset medical knowledge rule base, determining dominant modality information for representing the target tongue image feature in the RGB image information and the hyperspectral image information;

[0010] Generating guide data based on the dominant modality information, and extracting target data related to the target tongue image feature from the RGB image information and the hyperspectral image information;

[0011] Using the guide data to perform asymmetric fusion algorithm processing on the target data to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to strengthen the expression of the target tongue image feature in the fusion result of the target local area;

[0012] Generating a fused tongue image based on the local fusion result of each target local area.

[0013] By the above scheme, at least one target local area in the RGB image and the hyperspectral image is determined, and at least one target tongue image feature in the target local area is determined, which is a key step for realizing local adaptive fusion. By focusing on a specific area and feature, information fusion can be more targeted, avoiding information dilution caused by global fusion. Then, based on a preset medical knowledge rule base, dominant modality information for representing the target tongue image feature is determined from the RGB image information and the hyperspectral image information. This step is the core of realizing differentiated fusion. The importance of different modalities in representing a specific area and feature is determined by the medical knowledge rule base, thereby providing a basis for selecting the dominant modality, so that the fusion process is more in line with medical cognition. Then, guided data is generated based on the dominant modality information, and target data related to the target tongue image feature is extracted from the RGB image information and the hyperspectral image information. The guided data generated based on the dominant modality information can effectively guide the fusion direction of the target data, highlighting the features of the dominant modality, while the target data related to the target tongue image feature is extracted, ensuring that the information for fusion is effective information related to diagnosis. The specificity contribution of the two modalities to the feature is processed by a differentiated fusion algorithm, thereby preserving and enhancing the local subtle features most valuable for diagnosis in the fused image with high fidelity, to improve the accuracy and reliability of subsequent feature extraction, quantitative analysis and intelligent recognition.

[0014] Optionally, after acquiring the RGB image and the hyperspectral image respectively captured on the tongue surface of the same subject, the method further comprises:

[0015] Positionally calibrating the RGB image and the hyperspectral image so that the coordinates of each point on the tongue surface in a spatial coordinate system are consistent in the RGB image and the hyperspectral image;

[0016] and performing a first preprocessing operation on the RGB image and the hyperspectral image after position calibration, the first preprocessing operation comprising at least one of noise filtering, image enhancement, geometric correction and illumination correction.

[0017] By the above scheme, the accuracy of image registration and the image quality can be improved, providing a more accurate data basis for subsequent fusion.

[0018] Optionally, the step of determining at least one target local area in the RGB image and the hyperspectral image, and at least one target tongue image feature in the target local area, specifically comprises:

[0019] using an edge detection-based tongue segmentation algorithm to segment the tongue area to obtain a plurality of target local areas;

[0020] performing feature detection algorithm processing on each target local region to obtain type information of target tongue image features, wherein the feature detection algorithm includes at least one of spot detection, texture analysis, and edge detection.

[0021] Through the above scheme, the tongue region can be accurately segmented and the target tongue image features can be detected, and more accurate region and feature information is provided for subsequent fusion.

[0022] Optionally, based on a preset medical knowledge rule base, dominant modality information for representing the target tongue image features is determined in the RGB image information and the hyperspectral image information for the target local region and the target tongue image features, and the step specifically includes:

[0023] For the target local region and the target tongue image features, an RGB image structure feature evaluation value obtained based on the RGB image information and a hyperspectral image spectral feature evaluation value obtained based on the hyperspectral image information are calculated;

[0024] The RGB image structure feature evaluation value and the hyperspectral image spectral feature evaluation value are compared, and the evaluation value with a larger value is determined as the dominant modality of image data;

[0025] The knowledge base dominant modality corresponding to the target local region and the target tongue image features is obtained from the preset medical knowledge rule base;

[0026] It is determined whether the knowledge base dominant modality is consistent with the image data dominant modality, if not, the knowledge base weight of the target local region is obtained from the preset medical knowledge rule base, the data weight is determined according to the preprocessing quality of the RGB image information and the hyperspectral image information, the knowledge base dominant modality is weighted and fused according to the knowledge base weight, and the image data dominant modality is weighted and fused according to the data weight, to obtain target dominant modality information;

[0027] If consistent, the knowledge base dominant modality is taken as the dominant modality information for representing the target tongue image features in the RGB image information and the hyperspectral image information.

[0028] Through the above scheme, the image data and the medical knowledge can be comprehensively considered, the dominant modality information can be more accurately determined, and the accuracy of fusion is improved.

[0029] Optionally, the step of obtaining the knowledge base weight of the target local region from the preset medical knowledge rule base and determining the data weight according to the preprocessing quality of the RGB image information and the hyperspectral image information specifically includes:

[0030] acquire tongue feature type information related to the target local region;

[0031] select an evaluation criterion for quantifying reliability of the target local region according to the acquired tongue feature type information;

[0032] quantify reliability of the target local region based on the selected evaluation criterion to obtain a target local region reliability value;

[0033] convert the target local region reliability value into the knowledge base weight through a preset conversion mode; and

[0034] calculate a quality index value of a first preprocessing operation of the RGB image information and the hyperspectral image information in the target local region;

[0035] quantify a credibility of the image data according to the quality index value, and convert the credibility into the data weight through a second preset conversion mode.

[0036] Through the above scheme, the knowledge base weight and the data weight can be more reasonably determined, and the reliability of fusion can be improved.

[0037] Optionally, the step of generating guide data based on the dominant modality information and extracting target data related to the target tongue feature from the RGB image information and the hyperspectral image information specifically includes:

[0038] in the case of using the RGB image information as the dominant modality information, using a guide filter to filter the RGB image to generate a structure guide map that retains edge details and smooths non-structural areas, and extracting a color space component as a color guide map; and using the structure guide map or the color guide map as the guide data;

[0039] extracting a spectral feature vector map corresponding to the target local region in the RGB image and related to the target tongue feature from the hyperspectral image as the target data.

[0040] Through the above scheme, the structure and color information of the RGB image can be effectively used to guide the fusion of the hyperspectral image, and the details of the fused image can be enhanced.

[0041] Optionally, the step of using the guide data to perform an asymmetric fusion algorithm on the target data to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to strengthen expression of the target tongue feature in the fusion result of the target local region specifically includes:

[0042] The spectral feature vector image is taken as an input of guided filtering, and the structure guide image is taken as a guide image, and an asymmetric fusion algorithm is processed, so that the output spectral feature is aligned with the RGB structure in spatial distribution and is enhanced in detail; or

[0043] The spectral feature vector image is taken as an input image of guided filtering, and the color guide image is taken as a guide image, and the hyperspectral feature value is used to adjust the brightness, saturation or intensity of a specific color channel of the corresponding pixel in the color guide image.

[0044] Through the above scheme, the alignment of the spectral feature and the RGB structure and the detail enhancement can be realized, or the color guide image is adjusted by using the hyperspectral feature, so that the expression of the target tongue image feature is strengthened.

[0045] Optionally, the step of generating guide data based on the dominant modal information and extracting target data related to the target tongue image feature from the RGB image information and the hyperspectral image information specifically includes:

[0046] In the case that the hyperspectral image information is the dominant modal information, a spectral index image obtained by performing band operation on the hyperspectral image, or a main component score image obtained by performing principal component analysis on the hyperspectral image, is taken as the guide data.

[0047] The fine texture and color information corresponding to the target local region in the hyperspectral image and related to the target tongue image feature are extracted from the RGB image as the target data.

[0048] Through the above scheme, the spectral information of the hyperspectral image can be effectively used to guide the fusion of the RGB image, and the texture and color information of the fused image can be enhanced.

[0049] Optionally, the target data is processed by using the guide data through an asymmetric fusion algorithm, and a local fusion result is obtained, and the asymmetric fusion algorithm is used to strengthen the expression of the target tongue image feature in the fusion result of the target local region, and the step specifically includes:

[0050] The fine texture and color information is taken as an input of guided filtering, and the spectral index image is taken as a guide image, and an asymmetric fusion algorithm is processed, so that the abnormal region indicated by the spectral index is visually enhanced in the fused tongue image; or

[0051] The fine texture and color information is taken as an input of guided filtering, and the main component score image is taken as a guide image, and the information of each local region after guided fusion processing is weighted combined or selectively spliced.

[0052] Through the above scheme, visual enhancement of the abnormal area indicated by the spectrum in the fused tongue image can be realized, or the information of the local area is weighted combined or selectively spliced, so as to strengthen the expression of the target tongue feature.

[0053] In a second aspect, a tongue image analysis system based on RGB image and spectral data fusion is provided, and the system comprises:

[0054] An acquisition module is configured to acquire an RGB image and a hyperspectral image captured respectively for a tongue surface of a same subject, and obtain RGB image information and hyperspectral image information through image analysis;

[0055] A region feature determination module is configured to determine at least one target local area in the RGB image and the hyperspectral image, and at least one target tongue feature in the target local area;

[0056] A dominant modality determination module is configured to determine dominant modality information for representing the target tongue feature in the RGB image information and the hyperspectral image information based on a preset medical knowledge rule base for the target local area and the target tongue feature;

[0057] A data preparation module is configured to generate guide data based on the dominant modality information, and extract target data related to the target tongue feature from the RGB image information and the hyperspectral image information;

[0058] A local fusion module is configured to perform non-symmetrical fusion algorithm processing on the target data by using the guide data, to obtain a local fusion result, wherein the non-symmetrical fusion algorithm is used to strengthen the expression of the target tongue feature in the fusion result of the target local area;

[0059] An image generation module is configured to generate a fused tongue image based on the local fusion result of each target local area.

[0060] Through the above scheme, each step of the tongue image analysis method can be realized, and the efficiency and accuracy of tongue image analysis can be improved.

[0061] As can be seen from the above, the tongue image analysis method and system based on RGB image and spectral data fusion provided by the present application adaptively selects the dominant modality information of the RGB image and the hyperspectral image according to the features of different areas of the tongue surface, and performs non-symmetrical fusion, thereby solving the problem that the local key diagnostic information is diluted due to the global fusion strategy in the prior art, and having the advantages of improving the comprehensiveness and accuracy of tongue image analysis. BRIEF DESCRIPTION OF DRAWINGS

[0062] In order to more clearly illustrate the technical solutions of the present application, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without any creative effort.

[0063] Figure 1 Fig. 1 exemplarily shows a flow diagram of a tongue image analysis method based on RGB image and spectral data fusion according to an embodiment;

[0064] Figure 2 Fig. 2 exemplarily shows a module configuration block diagram of a tongue image analysis system 100 based on RGB image and spectral data fusion according to an embodiment. DETAILED DESCRIPTION

[0065] In order to make the technical personnel in the art better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort should be within the scope of protection of the present application.

[0066] It should be understood that the terms "first", "second", "third" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, for example, those given in the embodiment illustrations or descriptions of the present application can be implemented in an order other than that given.

[0067] In addition, the terms "include" and "have" and any variations thereof are intended to cover but not exclusively include, for example, a product or device containing a series of components without being limited to those components clearly listed, but can include other components not clearly listed or inherent to such products or devices.

[0068] In the traditional existing tongue analysis system, RGB images and hyperspectral images are used for fusion analysis. RGB images provide spatial color texture information, and hyperspectral images provide spectral biochemical component information. However, the physiological and pathological characteristics of different regions of the tongue are different, resulting in uneven information contribution of RGB images and hyperspectral images in representing the characteristics of these regions. The related fusion strategy usually adopts a global unified manner, which fails to fully consider the heterogeneity of information contribution of different regions of the tongue surface and the different dependence of different types of local features on two modal data. This global unified fusion strategy leads to inaccurate or insufficient expression of local key diagnostic information, especially features that are significant in one modality and weak in another modality, in the fused image.

[0069] For example, assuming that in a traditional Chinese medicine auxiliary diagnosis system, it is necessary to perform fine analysis on the tongue image of a user. The system collects the RGB image and the hyperspectral image of the tongue of the user. When analyzing the tip of the tongue, it is necessary to identify tiny red dots; when analyzing the edge of the tongue, it is necessary to identify the details of the tooth marks; when analyzing the root of the tongue, it is necessary to evaluate the thickness and color of the tongue coating; and when analyzing the sublingual region, it is necessary to observe the morphology and color of the collaterals. These local features differ greatly in their expression in RGB images and hyperspectral images. If a global unified fusion algorithm is used, it may not be able to effectively highlight these local features that are significant in a specific modality. Tiny red dots on the tip of the tongue may become blurred due to the interference of hyperspectral information, and spectral abnormalities of the tongue coating in the root of the tongue may be weakened due to the averaging of RGB texture information. This inaccuracy in information expression makes it difficult for the system to accurately identify and quantify these key local diagnostic features. If the above problem is not solved, i.e., if the global unified fusion strategy is continued to be used, the fused tongue image will not be able to accurately preserve and enhance the local subtle features that are most valuable for diagnosis. This will directly affect the subsequent tongue feature extraction algorithm, which may result in inaccurate feature extraction or missed detection.

[0070] In the face of the above problems, the present application considers how to overcome the limitations of the global fusion strategy and achieve differential processing of different regions and features of the tongue surface. The present application proposes that, in combination with medical knowledge and the characteristics of the image data itself, it can be determined which of the RGB information and the hyperspectral information is the dominant modality in a specific local region and for a specific tongue feature. Then, based on the determined dominant modality, a non-symmetrical fusion algorithm is designed, so that the information of the dominant modality can guide the fusion process and strengthen the expression of the target tongue feature in the fusion result. This non-symmetrical fusion strategy can selectively use the information of different modalities according to the characteristics of the local region and feature, avoid information dilution, and thus highlight the key diagnostic information in the fused image.

[0071] For example, as shown in FIG. 1, the tongue image is divided into different regions, and the RGB image and the hyperspectral image are fused in different regions. In the tip region of the tongue, the RGB image is the dominant modality, and the hyperspectral image is the auxiliary modality. In the edge region of the tongue, the hyperspectral image is the dominant modality, and the RGB image is the auxiliary modality. In the root region of the tongue, the RGB image is the dominant modality, and the hyperspectral image is the auxiliary modality. In the sublingual region, the RGB image is the dominant modality, and the hyperspectral image is the auxiliary modality. Figure 1As shown, an exemplary flowchart of a tongue image analysis method based on RGB image and spectral data fusion is shown. The present application proposes a tongue image analysis method based on RGB image and spectral data fusion, comprising:

[0072] S10, obtaining an RGB image and a hyperspectral image respectively captured on the tongue surface of the same subject, and obtaining RGB image information and hyperspectral image information through image analysis.

[0073] First, the RGB image and the hyperspectral image are obtained, and the corresponding image information is obtained through image analysis, which is the basis for subsequent fusion and provides data sources for subsequent analysis and fusion.

[0074] Further, after obtaining the RGB image and the hyperspectral image respectively captured on the tongue surface of the same subject, further comprising:

[0075] Position calibration is performed on the RGB image and the hyperspectral image to make the coordinates of each point on the tongue surface in the spatial coordinate system consistent in the RGB image and the hyperspectral image;

[0076] and performing a first preprocessing operation on the position-calibrated RGB image and hyperspectral image, the first preprocessing operation comprising at least one of noise filtering, image enhancement, geometric correction, and illumination correction.

[0077] Wherein, the position calibration refers to calculating and applying geometric transformation to make the pixels corresponding to the same physical point in the two images have the same spatial coordinates. A feature point matching-based method can be used to calculate the transformation matrix between the images, and then the transformation is applied to resample one or both images. A registration method based on regional cross-correlation or mutual information can also be used.

[0078] The first preprocessing operation refers to a set of image processing steps performed on the original image before image fusion, aiming to improve image quality, eliminate interference, or correct distortion. Specifically, noise filtering refers to removing randomly distributed pixel value abnormalities in the image through a filtering algorithm. Image enhancement refers to improving the visual effect of the image or highlighting specific features by adjusting the contrast, brightness, sharpness, or color distribution of the image. Geometric correction refers to correcting the geometric distortion of the image caused by the camera lens or imaging angle by applying pre-calibrated or calculated geometric transformation parameters. Illumination correction refers to compensating for brightness or color changes in the image caused by uneven light sources or shadows through an algorithm, making the light distribution of the image more uniform.

[0079] In some embodiments, the present application aims to provide high-quality, registered input data for subsequent image analysis and fusion by adding a position calibration and a first pre-processing operation step after acquiring the RGB image and the hyperspectral image. Position calibration ensures that the spatial structure information of the RGB image and the spectral information of the hyperspectral image are accurately aligned in space. This is crucial for subsequent feature extraction, dominant modality judgment, and asymmetric fusion in a specific local area. If the images are not aligned, even the most sophisticated subsequent fusion algorithm cannot combine the correct spatial information with the correct spectral information, resulting in distorted fusion results that cannot accurately reflect the local features of the tongue image. The first pre-processing operation improves the quality of the image itself.

[0080] Noise filtering reduces random interference, making the texture, color, and spectral curve features of the tongue image more clear and distinguishable, avoiding noise interference in subsequent feature extraction and analysis. Image enhancement can highlight weak features of the tongue image, such as color changes or fine cracks in early lesions, making them easier to detect and identify in subsequent processing. Geometric correction eliminates image distortion, ensuring the accuracy of tongue morphology and local area division. Illumination correction compensates for uneven lighting, making the color information of the tongue and tongue coating more accurate and reliable, avoiding misjudgment due to lighting differences. These pre-processing steps work together with the local asymmetric fusion strategy.

[0081] In a specific embodiment, after acquiring the RGB image and the hyperspectral image, position calibration can be performed first. Position calibration can use a method based on the SIFT algorithm to extract feature points in the RGB image and the hyperspectral image, eliminate mis-matching points through the RANSAC algorithm, and calculate the affine transformation matrix between the two images. The hyperspectral image is resampled according to the calculated affine transformation matrix to align it with the RGB image in space. Then, the first pre-processing operation is performed. Noise filtering can apply a median filter to the aligned RGB image and hyperspectral image respectively. Image enhancement can perform limited contrast adaptive histogram equalization on the RGB image and linear stretching on each band of the hyperspectral image. Geometric correction can use camera calibration parameters to correct the distortion of the original image. Illumination correction can apply the Retinex algorithm to the RGB image and the gray world algorithm to the hyperspectral image. After these steps, the RGB image and the hyperspectral image that have undergone position calibration and the first pre-processing operation are obtained, which are used for subsequent image analysis and fusion.

[0082] S20, determining at least one target local area in the RGB image and the hyperspectral image, and at least one target tongue image feature in the target local area.

[0083] The determining of the at least one target local region and the at least one target tongue feature in the target local region refers to identifying tongue surface regions with specific diagnostic significance, such as the tongue tip, the middle of the tongue, the tongue root, the tongue edge, and specific tongue features in these regions, such as spots, cracks, tongue fur, and blood vessels, in the acquired RGB image and hyperspectral image. This can be achieved by using image segmentation techniques, such as edge-based, color-based, or texture-based segmentation algorithms to identify the tongue body and divide the regions, or by using feature detection algorithms, such as spot detection, texture analysis, edge detection, or morphological analysis to identify and locate specific features. The main purpose is to focus the subsequent information processing on these local regions and features that are crucial for diagnosis.

[0084] In some embodiments, the determining of the target local region and the target tongue feature can be specifically achieved by analyzing the image content to identify specific regions on the tongue body and visible features in these regions. This can provide focus points for subsequent image fusion.

[0085] In some embodiments, the determining of the at least one target local region and the at least one target tongue feature in the target local region in the RGB image and the hyperspectral image specifically includes:

[0086] S201, performing tongue body region segmentation using an edge detection-based tongue body segmentation algorithm to obtain a plurality of target local regions;

[0087] S202, performing feature detection algorithm processing on each target local region to obtain type information of the target tongue feature, wherein the feature detection algorithm includes at least one of spot detection, texture analysis, and edge detection.

[0088] In some embodiments, the tongue body region segmentation is performed using an edge detection-based tongue body segmentation algorithm. The object boundary is determined by identifying regions with sharp changes in pixel values in the image. In tongue analysis, this algorithm can effectively separate the tongue body region from the background, as there is usually a significant color or texture edge difference between the tongue body and the surrounding environment. Through this segmentation, a plurality of target local regions can be obtained.

[0089] These local regions can be predefined anatomical partitions of the tongue or regions with similar attributes automatically partitioned according to image features. For each target local region, a feature detection algorithm is applied. Feature detection algorithms are used to identify and extract specific visual patterns in the image, which correspond to various clinical features of the tongue image. The feature detection algorithms used here include at least one of spot detection, texture analysis, and edge detection. Among them, the spot detection algorithm is used to identify small areas in the image, which is suitable for detecting point lesions on the tongue. The texture analysis algorithm is used to quantify the texture properties of the image region, which is suitable for evaluating the thickness, texture, and cracks of the tongue fur, etc. The edge detection algorithm can be used to identify morphological features of the tongue edge, such as tooth marks, abnormalities in the tongue contour, etc. in addition to segmentation. By combining the use of these algorithms, the appropriate algorithm or combination of algorithms can be selected according to the characteristics of different local regions and the types of features to be detected, thereby improving the relevance and accuracy of feature extraction.

[0090] In some embodiments, based on the acquired and preprocessed RGB images and hyperspectral images, the determination process of target local regions and target tongue image features is further refined. First, an edge detection-based tongue segmentation algorithm is used to process the images. This algorithm analyzes the edge information of the image to accurately outline the contour of the tongue, distinguishing it from non-tongue regions. This segmentation operation effectively reduces the scope of subsequent analysis, eliminating background noise and irrelevant information that interferes with feature detection, making the analysis more focused on the tongue itself. After segmentation, the tongue region is divided into multiple target local regions. This division takes into account the differences in physiology and pathology of different parts of the tongue, laying the foundation for subsequent detailed analysis.

[0091] Next, for these divided target local regions, at least one of the spot detection, texture analysis, edge detection, and other feature detection algorithms is flexibly used according to the characteristics of the region and the potential types of tongue image features. For example, spot detection is preferred in areas where point lesions may occur, texture analysis is emphasized in areas where tongue fur or cracks need to be evaluated, and edge detection is applied in areas where tongue morphology or tooth marks are concerned. This targeted, multi-algorithm collaborative feature detection method can more comprehensively and accurately capture various subtle features of the tongue image, including color, morphology, texture, and other information. The target local regions and target tongue image feature information determined in this way significantly improve the accuracy and reliability of feature extraction compared to global or single-algorithm feature analysis directly on the entire tongue image. This accurate and comprehensive feature information provides high-quality input for subsequent generation of guide data based on dominant modal information, extraction of target data, and asymmetric fusion processing, enabling the final fusion tongue image to more accurately express key local tongue image features, thereby improving the effectiveness of the overall tongue image analysis method.

[0092] As a specific implementation, Canny edge detection can be first performed on the acquired RGB image and hyperspectral image to extract edge information of the images. Then, based on the extracted edge information, tongue body region segmentation is performed using an active contour model or a graph cut algorithm to separate the tongue body region from the background and obtain a tongue body mask. According to tongue anatomy knowledge, the segmented tongue body region can be further divided into multiple target local regions such as a tongue tip region, a tongue middle region, a tongue root region, a tongue left edge region, and a tongue right edge region. For the tongue tip region, a LoG operator-based spot detection algorithm can be applied to identify possible small red spots in the region. For the tongue middle region, a gray level co-occurrence matrix or a local binary pattern texture analysis algorithm can be used to quantify texture features of tongue fur and assess its thickness and texture. In the tongue edge region, an edge detection algorithm combined with morphological operations can be used again to detect and analyze the sawtooth structure of the tongue edge to identify tooth marks. In this way, for different local regions and feature types of interest, corresponding feature detection algorithms are selected and applied to obtain type information about target tongue image features in these regions.

[0093] By using the tongue body segmentation algorithm based on edge detection to segment the tongue body region, the tongue body can be effectively separated from the background, reducing the interference of irrelevant regions and improving the accuracy of subsequent feature detection. Further, for each target local region, at least one of a plurality of feature detection algorithms including spot detection, texture analysis, and edge detection is used for processing, so that appropriate algorithms can be selected according to different regions and feature types, overcoming the limitations of a single algorithm. Thus, the present application can more comprehensively and accurately extract various subtle feature information of tongue images, providing a more reliable and accurate data basis for subsequent image fusion, which helps to generate a fused tongue image that more accurately expresses local key features.

[0094] S30, for the target local region and the target tongue image feature, based on the preset medical knowledge rule base, the dominant modal information used to represent the target tongue image feature is determined from the RGB image information and the hyperspectral image information.

[0095] For the target local region and the target tongue image feature, based on the preset medical knowledge rule base, the dominant modal information used to represent the target tongue image feature is determined from the RGB image information and the hyperspectral image information. This step is the core of differential fusion, which determines the importance of different modal information in representing specific regions and features through the medical knowledge rule base, thereby providing a basis for selecting the dominant modal, making the fusion process more consistent with medical cognition.

[0096] In some embodiments, the guiding data is generated based on the dominant modal information, and the target data related to the target tongue image features is extracted from the RGB image information and the hyperspectral image information, which means that auxiliary data such as structure guide map, color guide map or spectral index map is generated from the dominant modal image information to guide the fusion process according to the determined dominant modal, and the auxiliary data contains the key information of the dominant modal. At the same time, data such as spectral feature vector map or fine texture and color information corresponding to the target tongue image features in space and needing to be fused or enhanced is extracted from the non-dominant modal image information. The main purpose is to provide input for the asymmetric fusion algorithm, and to clearly indicate the direction and content of fusion.

[0097] In some embodiments, for the target local area and the target tongue image features, the dominant modal information used to represent the target tongue image features is determined in the RGB image information and the hyperspectral image information. The dominant modal information can be determined by comparing the overall information amount of the RGB image and the hyperspectral image, such as calculating the Shannon entropy of the image, and selecting the modal corresponding to the image with the larger entropy value as the dominant modal, so that it can be preliminarily judged which modal is more suitable for representing the features according to the quality of the image itself.

[0098] Based on this, step S30 can be specifically expanded as:

[0099] S301, for the target local area and the target tongue image features, calculating the RGB image structure feature evaluation value obtained based on the RGB image information, and the hyperspectral image spectral feature evaluation value obtained based on the hyperspectral image information;

[0100] S302, comparing the RGB image structure feature evaluation value and the hyperspectral image spectral feature evaluation value, and determining the image data dominant modal with the larger evaluation value;

[0101] S303, obtaining the knowledge base dominant modal corresponding to the target local area and the target tongue image features from the preset medical knowledge rule base;

[0102] S304, judging whether the knowledge base dominant modal is consistent with the image data dominant modal, if not, obtaining the knowledge base weight of the target local area from the preset medical knowledge rule base, and determining the data weight according to the preprocessing quality of the RGB image information and the hyperspectral image information, and performing weighted fusion on the knowledge base dominant modal according to the knowledge base weight, and performing weighted fusion on the image data dominant modal according to the data weight, to obtain the target dominant modal information;

[0103] S304, if consistent, the knowledge base dominant modal is used as the dominant modal information used to represent the target tongue image features in the RGB image information and the hyperspectral image information.

[0104] The RGB image structure feature evaluation value refers to a value for evaluating the expression ability or quality of the structure feature of the RGB image in the target local region, which can be calculated by using a structure similarity index, edge intensity, or texture contrast, etc. The hyperspectral image spectral feature evaluation value refers to a value for evaluating the expression ability or quality of the spectral feature of the hyperspectral image in the target local region, which can be calculated by using a spectral angle mapping, spectral information divergence, or signal-to-noise ratio, etc.

[0105] The image data dominant modality refers to an image modality determined according to the evaluation result of the image data itself, which is more suitable for representing the target tongue image feature.

[0106] The preset medical knowledge rule base refers to a set of medical knowledge rules related to the tongue image feature and the local region, which can be stored in the form of a lookup table or an expert system rule set, etc. The knowledge base dominant modality refers to an image modality recommended by the medical knowledge according to the preset medical knowledge rule base, the specific target local region, and the target tongue image feature.

[0107] The knowledge base weight refers to a value for measuring the credibility or importance of the preset medical knowledge rule base in the specific target local region and the target tongue image feature, which is determined according to the research depth or clinical verification situation of the region or feature in the medical knowledge, or obtained by expert scoring or statistical analysis based on historical data. The data weight refers to a value for measuring the preprocessing quality or credibility of the RGB image information and the hyperspectral image information in the specific target local region, which is determined according to the preprocessing quality indicators such as the signal-to-noise ratio, definition, or registration accuracy of the image, or obtained based on the direct output of the image quality evaluation algorithm.

[0108] When the knowledge base dominant modality is inconsistent with the image data dominant modality, the knowledge base weight and the data weight are comprehensively considered, and the tendency of the two modalities is weighted and calculated, so as to obtain the final dominant modality information, which can be realized by using a weighted average or weighted voting algorithm. The target tongue image feature finally determined after comprehensive judgment is used to represent the dominant modality information.

[0109] For example, in determining the dominant modality of the red dot feature of the tongue tip region, the structure similarity index of the RGB image of the region can be calculated as the RGB image structure feature evaluation value, and the spectral reflectance contrast of the hyperspectral image of the region in a specific waveband can be calculated as the hyperspectral image spectral feature evaluation value. By comparing the two evaluation values, the modality with a larger evaluation value is determined as the image data dominant modality.

[0110] Meanwhile, a preset medical knowledge rule base is queried, and a knowledge base recommended dominant mode is obtained according to the “tongue tip” area and “red dot” features. For example, the rule base may recommend “RGB” as the dominant mode. It is determined whether the knowledge base dominant mode (“RGB”) is consistent with the image data dominant mode. If consistent, the final dominant mode is determined as “RGB”. If the image data dominant mode is “hyperspectral” (for example, the quality of the RGB image in this area is very poor), the two are inconsistent. At this time, the knowledge base weight of the tongue tip area is obtained from the rule base, for example, 0.8, indicating that the medical knowledge has a high credibility for the judgment of the tongue tip area.

[0111] Meanwhile, according to the definition score of the RGB image in this area and the signal-to-noise ratio score of the hyperspectral image in this area, the data weight is determined, for example, the RGB data weight is 0.3 and the hyperspectral data weight is 0.9. Weighted fusion can be performed in a weighted voting manner, for example, the knowledge base tends to RGB, with a weight of 0.8; the image data tends to hyperspectral, with a weight of 0.9. Since the data weight is higher than the knowledge base weight, the final dominant mode may tend to hyperspectral, or a more complex weighted calculation is adopted to obtain the final result. In this way, even if the image data evaluation result does not match the medical knowledge, a comprehensive judgment can be made according to their respective credibility to obtain more reliable dominant mode information.

[0112] The above-mentioned manner improves the accuracy of the determination of the dominant mode information. First, the structural feature evaluation value of the RGB image and the spectral feature evaluation value of the hyperspectral image are calculated respectively, and the numerical values of the two are compared to determine the image data dominant mode. This step considers the characteristics of the two image modalities, and uses the information of the image data itself to preliminarily determine which modality has a higher information contribution.

[0113] Then, the knowledge base dominant mode corresponding to the target local area and the target tongue image features is obtained from the preset medical knowledge rule base. The introduction of the medical knowledge rule base can reduce the risk of inaccurate judgment caused by simply relying on image data and improve the credibility of the determination of the dominant mode information.

[0114] Next, it is determined whether the knowledge base dominant mode is consistent with the image data dominant mode. If consistent, the knowledge base dominant mode is directly taken as the final dominant mode information, which indicates that the image data and the medical knowledge are consistent in this case, and can be directly used as the final result.

[0115] If the knowledge base dominant mode is inconsistent with the image data dominant mode, a weighted fusion process is needed. The knowledge base weight of the target local area is obtained from the preset medical knowledge rule base, and the data weight is determined according to the pretreatment quality of the RGB image information and the hyperspectral image information. The knowledge base weight represents the credibility of the medical knowledge, and the data weight represents the quality level of the image data. Through weighted fusion, the credibility of the medical knowledge and the image data can be comprehensively evaluated, and the dominant mode information with higher accuracy can be obtained.

[0116] Specifically, the knowledge base dominant mode is weighted according to the knowledge base weight, and the image data dominant mode is weighted according to the data weight, and finally the target dominant mode information is obtained. This weighted fusion processing method can coordinate the differences between medical knowledge and image data, so that the final dominant mode information has higher rationality. Through this method of comprehensively judging the dominant mode, more reliable input is provided for subsequent generation of guide data based on the dominant mode information, extraction of target data and asymmetric fusion, so that the final local fusion result can more accurately strengthen the expression of target tongue image features, and further generate a higher quality fusion tongue image, overcoming the problem that the global fusion strategy in the related art cannot take into account the heterogeneity of local information contribution, resulting in dilution, weakening or submergence of local key diagnostic information.

[0117] In some embodiments, by determining the knowledge base weight and the data weight, and using these weights to weight the knowledge base dominant mode and the image data dominant mode, more accurate target dominant mode information is determined. Specifically, the steps of obtaining the knowledge base weight of the target local area from the preset medical knowledge rule base, and determining the data weight according to the pretreatment quality of the RGB image information and the hyperspectral image information, specifically include:

[0118] Obtain tongue feature type information related to the target local area;

[0119] According to the obtained tongue feature type information, select an evaluation standard for quantifying the reliability of the target local area;

[0120] Based on the selected evaluation standard, the reliability of the target local area is quantified to obtain a target local area reliability value;

[0121] The target local area reliability value is converted into the knowledge base weight through a preset conversion method; and

[0122] Calculate the quality index value of the first pretreatment operation of the RGB image information and the hyperspectral image information in the target local area;

[0123] According to the quality index value, quantize the reliability of the image data, and convert the reliability into the data weight through a second preset conversion mode.

[0124] The preset conversion mode refers to a rule or function that maps one value range to another value range, which can be implemented by a lookup table, a linear function, a nonlinear function, or a piecewise function. The quality index value refers to a quantized value for measuring the quality of image data in a specific region, which can be implemented by signal-to-noise ratio, sharpness, contrast, artifact degree, or detection confidence of a specific feature. Quantizing the reliability of the image data refers to converting the quality index value of the image data into a value that represents the degree of reliability of the data in the current application scenario, which can be implemented by normalizing the quality index value to a specific range (e.g., 0 to 1). The second preset conversion mode refers to a rule or function that converts the reliability value of the image data into a data weight, which can be implemented by a lookup table, a linear function, a nonlinear function, or a piecewise function.

[0125] In some embodiments, first, the tongue feature type information related to the target local region is obtained, providing a basis for subsequent selection of appropriate evaluation criteria. Different tongue feature types have different manifestations and diagnostic significance in different regions, so it is necessary to evaluate the reliability of the region in a targeted manner. Then, according to the obtained tongue feature type information, the evaluation criteria for quantifying the reliability of the target local region are selected. This selection process ensures the pertinence and effectiveness of the evaluation. Next, based on the selected evaluation criteria, the reliability of the target local region is quantified, obtaining an objective reliability value. This value reflects the reliability of the local region under the current feature type in the knowledge base guidance. Subsequently, the reliability value is converted into a knowledge base weight through a preset conversion method. The higher the reliability value, the greater the knowledge base weight obtained by conversion, indicating that the guidance of the knowledge base is stronger in this region and under this feature. At the same time, the quality index value of the first preprocessing operation of the RGB image information and the hyperspectral image information in the target local region is calculated. These quality indicators reflect the quality of the image data itself, such as whether there is noise, blur or distortion. According to these quality index values, the credibility of the image data is quantified. The higher the credibility, the more reliable the image data information in the region. Finally, the credibility is converted into a data weight through a second preset conversion method. The higher the credibility, the greater the data weight obtained by conversion, indicating that the contribution of the image data in this region should be greater. Through the above steps, the present application provides a weight determination method that comprehensively considers tongue feature types, target local region reliability, and image data quality. This method is combined with the step of determining whether the knowledge base dominant mode and the image data dominant mode are consistent, and when they are inconsistent, the knowledge base dominant mode and the image data dominant mode are weighted and fused using the adaptively determined knowledge base weight and data weight, thereby obtaining more accurate target dominant mode information. This combination makes the determination process of the dominant mode more refined and adaptive, and can better balance the prior knowledge of the knowledge base and the actual information of the image data, especially when there is a conflict between the two, the influence of each can be adjusted according to the specific circumstances, improving the accuracy of the dominant mode determination.

[0126] For example, in a specific embodiment, assuming that the target local area is the tip of the tongue and the target tongue image feature is "tongue tip red dot". First, the tongue image feature type information is obtained as "tongue tip red dot". According to this information, the evaluation criteria for quantifying the reliability of the tongue tip area are selected, for example, the reliability of the knowledge base in guiding the area can be evaluated based on the physiological indicators related to the tongue tip red dot, such as the blood vessel distribution density, tissue thickness, etc. of the area. Assuming that the tongue tip area reliability value obtained after evaluation is 0.7. Through a preset conversion method, such as a lookup table or a function, 0.7 is converted into a knowledge base weight, for example, a weight value of 0.6 is obtained. At the same time, the quality index value of the RGB image and the hyperspectral image in the tongue tip area after the first preprocessing (such as denoising, enhancement) is calculated, for example, the clarity index of the RGB image is 0.85, and the signal-to-noise ratio index of the hyperspectral image is 0.8. According to these quality index values, the credibility of the image data is quantified, for example, the average value 0.825 is taken as the credibility value. Through a second preset conversion method, such as another lookup table or function, 0.825 is converted into a data weight, for example, a weight value of 0.4 is obtained. In this way, when the knowledge base dominant modality and the image data dominant modality are inconsistent in the tongue tip red dot feature, the calculated knowledge base weight 0.6 and data weight 0.4 can be used for weighted fusion to determine the final target dominant modality information.

[0127] Through the above technical means, the application provides a method for adaptively determining the knowledge base weight and data weight according to the tongue image feature type, target local area reliability and image data quality when the knowledge base dominant modality and the image data dominant modality are inconsistent. This method enables the subsequent weighted fusion process to more accurately balance the knowledge base information and image data information, thereby obtaining more accurate target dominant modality information and improving the accuracy and reliability of tongue image feature representation.

[0128] S40, generating guide data based on the dominant modality information, and extracting target data related to the target tongue image feature from the RGB image information and the hyperspectral image information.

[0129] Generating guide data based on the dominant modality information and extracting target data related to the target tongue image feature from the RGB image information and the hyperspectral image information can effectively guide the fusion direction of the target data, highlight the features of the dominant modality, and ensure that the fused information is effective information related to diagnosis.

[0130] S50, using the guide data to perform an asymmetric fusion algorithm on the target data to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to strengthen the expression of the target tongue image feature in the fusion result of the target local area.

[0131] The target data is processed by using the guide data to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to strengthen the expression of the target tongue feature in the fusion result of the target local region, and the asymmetric fusion algorithm can selectively enhance or suppress information of different modalities, thereby strengthening the expression of the target tongue feature in the fusion result and avoiding information flooding.

[0132] S60, generating a fused tongue image based on the local fusion results of the target local regions.

[0133] In some embodiments, by combining the determination of the tongue surface local region and the target tongue feature, the determination of the dominant modality based on medical knowledge and image analysis, and the asymmetric fusion based on the dominant modality, adaptive and differentiated fusion is realized for different regions and features of the tongue surface, and the effect of high-fidelity preservation and enhancement of local key diagnostic information is achieved.

[0134] Specifically, the method of the present application first acquires the RGB image and the hyperspectral image of the same subject and performs image analysis. Then, at least one target local region and at least one target tongue feature in the region that need to be focused on are identified in the images. Then, for these specific regions and features, it is determined whether the RGB information or the hyperspectral information can more effectively represent the feature by combining a pre-set medical knowledge rule base, so as to determine the dominant modality information. Based on the determined dominant modality information, guide data for guiding fusion is generated from the dominant modality image information, and target data related to the target tongue feature is extracted from the image information of the other modality. Subsequently, the target data is processed by using an asymmetric fusion algorithm guided by the guide data to obtain a local fusion result, and the algorithm is specifically used to strengthen the expression of the target tongue feature in the fusion result. Finally, the local fusion results of each target local region are integrated to generate a final fused tongue image. The whole process is a process from acquiring original data, to local identification and analysis, to local adaptive fusion based on the analysis results, and finally to integration to form a whole image, and the core is to perform differentiated and asymmetric processing for local regions and features.

[0135] In some embodiments, the step of generating guide data based on the dominant modality information and extracting target data related to the target tongue feature from the RGB image information and the hyperspectral image information specifically comprises:

[0136] In the case where the RGB image information is the dominant modality information, a structure guide image that retains edge details and smooths non-structural regions is generated by using a guide filter to filter the RGB image; and a color space component is extracted as a color guide image; the structure guide image or the color guide image is used as the guide data;

[0137] A spectral feature vector map corresponding to the target local region in the RGB image and related to the target tongue image feature is extracted from the hyperspectral image as target data.

[0138] The guided filter adopts a local linear model-based implementation. The structure-guided map refers to an image obtained after processing by the guided filter, which retains the edge information of the original image while smoothing the non-structure region. The color space component refers to each component obtained after converting the RGB image to other color spaces, which can be implemented by converting the RGB image to the HSV color space and extracting one or more of the H, S, and V components. The color-guided map refers to the extracted color space component. The spectral feature vector map refers to a set of vectors composed of the reflectivity or absorbance of each pixel point in the hyperspectral image at different wavebands, which can be implemented by extracting the spectral curve of the pixel points corresponding to the region of the RGB image from the hyperspectral cube.

[0139] In some embodiments, when the RGB image information is the dominant modality, the RGB image is filtered using a guided filter to generate a structure-guided map that retains edge details and smooths non-structure regions, thereby highlighting the structural information of the image; second, the color space component is extracted as a color-guided map, which can provide color information of the image, thereby providing color guidance for subsequent fusion.

[0140] The reason for processing the RGB image to generate a guided map is that directly using the original RGB image may introduce noise or irrelevant details, affecting the fusion effect. By guided filtering or extracting color space components, more useful information for the target tongue image feature can be extracted, thereby improving the accuracy of fusion.

[0141] Secondly, a spectral feature vector map corresponding to the target local region in the RGB image and related to the target tongue image feature is extracted from the hyperspectral image as target data. The reason for extracting the spectral feature vector map from the hyperspectral image is that the hyperspectral image can provide rich spectral information, which can reflect the biochemical composition and tissue structure of the tongue image. By extracting the spectral feature vector map corresponding to the target local region, the spectral features of the region can be obtained, thereby providing spectral information for subsequent fusion. At the same time, it is emphasized that the target tongue image feature is related, which ensures that the extracted spectral features are related to the current concerned features, avoiding the introduction of irrelevant information.

[0142] The present application, in combination with the steps of determining the target region, target feature and dominant modality, extracts data that is most conducive to subsequent fusion from the two modalities according to the determined dominant modality (RGB). By extracting structural or color guide information from RGB, the spatial resolution advantage of RGB can be utilized; by extracting a spectral feature vector map from hyperspectral, the spectral resolution advantage of hyperspectral can be utilized. This targeted data preparation provides high-quality input for subsequent asymmetric fusion, enabling the fusion result to better enhance the expression of target tongue image features, and solving the problem of dilution of local key information caused by the global unified fusion strategy in the prior art.

[0143] Therefore, by extracting structural or color information from the RGB image as a guide and extracting spectral information related to the target feature from the hyperspectral image as a target, targeted and high-quality input data is provided for subsequent asymmetric fusion. This data preparation method can fully utilize the spatial details of the RGB image and the spectral features of the hyperspectral image, enabling the subsequent fusion process to more effectively highlight and enhance the expression of target tongue image features, thereby improving the accuracy of tongue image analysis.

[0144] In some embodiments, the target data is processed by the asymmetric fusion algorithm using the guide data to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to enhance the expression of the target tongue image feature in the fusion result of the target local region, specifically including:

[0145] The spectral feature vector map is used as the input of the guide filtering, and the structural guide map is used as the guide image, and the asymmetric fusion algorithm is processed, so that the output spectral feature is aligned with the RGB structure in spatial distribution and is enhanced in detail; or,

[0146] The spectral feature vector map is used as the input image of the guide filtering, and the color guide map is used as the guide image, and the hyperspectral feature value is used to adjust the brightness, saturation or intensity of a specific color channel of the corresponding pixel in the color guide map.

[0147] In some embodiments, after determining that the RGB image information is the dominant modality and generating the structural guide map or the color guide map as the guide data and the spectral feature vector map as the target data, the asymmetric fusion stage is entered.

[0148] If the structure guide map is selected as the guide data, the spectral feature vector map is taken as the input image, and the structure guide map is taken as the guide image to perform the guided filtering processing. The guided filtering utilizes the spatial structure information of the RGB image provided by the structure guide map to guide the filtering process of the spectral feature vector map, so that the output spectral features are aligned with the structure of the RGB image in space, while details are preserved in the structure region and smoothing is performed in the smooth region. In this way, the spectral information of the hyperspectral image is combined with the spatial structure of the RGB image, and an image with both spectral information and clear structure is presented in the fusion result, which is beneficial to identifying and analyzing the tongue features related to the structure.

[0149] If the color guide map is selected as the guide data, the spectral feature vector map is taken as the input image, and the color guide map is taken as the guide image. The hyperspectral feature value is used to adjust the color attributes of the corresponding pixel in the color guide map, such as brightness, saturation, or intensity of a specific color channel. This means that the spectral information of the hyperspectral image is used to modulate the color representation of the RGB image. For example, the hyperspectral feature value of a certain pixel indicates that there is an abnormal biochemical component in the region, and this hyperspectral feature value can be used to increase the red channel intensity or saturation of the corresponding pixel in the color guide map, highlighting the color change of the abnormal region in the fusion result.

[0150] Both of the above-mentioned ways embody the idea of asymmetric fusion, which uses the guide data generated by the dominant modality to guide the fusion of the target data of the non-dominant modality, and strengthens the expression of the target tongue features in the fusion result. This fusion method avoids the information dilution caused by simple superposition or averaging, and can more effectively combine the advantages of the two modalities. The pre-step determines the dominant modality and generates appropriate guide data and target data based on the dominant modality. The asymmetric fusion algorithm of this step is based on these prepared data. The determination of the dominant modality ensures that the fusion process is guided based on the modality with more reliable and more representative information, and the extraction of the guide data and the target data provides necessary and relevant inputs for the asymmetric fusion. This combination enables the entire method to select the appropriate modality as the dominant modality according to the characteristics of the tongue features and the region, and to adopt the corresponding fusion strategy to more effectively highlight and strengthen the expression of the target tongue features in the fusion result of the local region, solving the problem that the global unified fusion strategy cannot adapt to the information heterogeneity of different regions on the tongue surface.

[0151] For example, in one implementation scenario, for a target local area on the tongue surface, such as a tongue coating area, and a target tongue image feature in the area, such as a tongue coating color abnormality, the RGB image information is determined as the dominant modality information through a pre-step, and a color guide map and a spectral feature vector map are generated. At this time, the second fusion mode can be used. The spectral feature vector map of the tongue coating area extracted from the hyperspectral image is taken as the input image of the guided filtering. The color guide map of the tongue coating area extracted from the RGB image, such as the red channel intensity map of the RGB image, is taken as the guide image. The hyperspectral feature values in the spectral feature vector map, such as the specific band intensity or spectral index related to the tongue coating composition, are used to adjust the intensity of the corresponding pixels in the color guide map. For example, if the hyperspectral feature value indicates that the pixel has yellow pigment deposition, the intensity of the yellow related channel of the corresponding pixel in the color guide map can be increased, so as to highlight the yellow abnormal area of the tongue coating in the fusion result.

[0152] By taking the spectral feature vector map as the input of the guided filtering and using different fusion strategies according to different guide data (structure guide map or color guide map), the present application can effectively use the spatial structure or color information of the RGB image to guide the fusion of the spectral features of the hyperspectral image. When the structure guide map is used, the spectral features in the fusion result are aligned with the structure of the RGB image in the spatial distribution, and the details are enhanced, which helps to clearly present the structural tongue image features with spectral properties and improve the recognition accuracy of these features. When the color guide map is used and the color properties are adjusted according to the hyperspectral feature values, the fusion result can highlight the color changes related to the spectral properties, so that these features are more easily identified visually, reflecting the differences in their biochemical composition. This asymmetric fusion mode can selectively fuse the hyperspectral information according to the information advantage of the RGB dominant modality, avoiding the information dilution or distortion that may be caused by simple fusion, so as to more effectively strengthen the expression of the target tongue image features in the fusion result of the local area, and improve the representation ability of the fusion image for the local subtle tongue image features, providing a more accurate and information-rich image basis for subsequent tongue image analysis and diagnosis.

[0153] In some embodiments, the steps of generating guide data based on the dominant modality information and extracting target data related to the target tongue image feature from the RGB image information and the hyperspectral image information specifically include:

[0154] In the case where the hyperspectral image information is the dominant modality information, a spectral index map obtained by band operation on the hyperspectral image, or a main component score map obtained by principal component analysis on the hyperspectral image, is taken as the guide data;

[0155] Extracting fine texture and color information corresponding to the target local region in the hyperspectral image and related to the target tongue image feature from the RGB image as target data.

[0156] Based on the foregoing method of determining the dominant modal information, when it is judged that the hyperspectral image information is the dominant modal of the current target local region and the target tongue image feature, a specific data preparation strategy is adopted. Specifically, in order to give full play to the advantage of hyperspectral information in spectral feature recognition and make use of the spatial details of the RGB image, the present application extracts guide data representing key spectral information from the hyperspectral image, and extracts corresponding spatial details from the RGB image as target data. The spectral index map or the main component score map of the hyperspectral image can effectively condense the diagnostic spectral information in the hyperspectral data, such as reflecting the thickness of tongue fur and the filling degree of sublingual collateral, etc. Taking it as guide data, these spectral information can be used to guide the fusion of RGB image spatial details in the subsequent fusion process.

[0157] At the same time, fine texture and color information corresponding to the target region of the hyperspectral image are extracted from the RGB image as target data, which provides high spatial resolution details of the local region of the tongue surface. In this way, when the hyperspectral information is the dominant modal, the fusion process is no longer simply guided by the structure of the RGB image, but uses the spectral advantage of the hyperspectral image itself to guide the expression of the details of the RGB image, ensuring that the fusion result can highlight the key spectral features that are significant in the hyperspectral image, while retaining the spatial details of the RGB image, so as to more accurately reflect the real physiological and pathological state of the local region of the tongue surface. This data preparation method, combined with the foregoing method of determining the dominant modal, realizes adaptive data selection and preparation according to the information advantage of different regions and features, overcomes the limitations of the global unified fusion strategy, and improves the expression accuracy of local key information.

[0158] Further, the target data is processed by an asymmetric fusion algorithm using the guide data to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to strengthen the expression of the target tongue image feature in the fusion result of the target local region, specifically including:

[0159] The fine texture and color information are taken as input of guided filtering, and the spectral index map is taken as a guide image, and the asymmetric fusion algorithm is processed, so that the abnormal region of the spectral index is visually enhanced in the fusion tongue image; or,

[0160] The fine texture and color information are taken as input of guided filtering, and the main component score map is taken as a guide image, and the information of each local region after guided fusion processing is weighted combined or selectively spliced.

[0161] By utilizing hyperspectral image information as guidance, the texture and color information of the RGB image is asymmetrically fused, thereby strengthening the expression of target tongue image features in the fusion result. Specifically, when the hyperspectral image information is determined as the dominant modality, a spectral index map or a principal component score map is generated as the guidance data, and fine texture and color information is extracted from the RGB image as the target data. Subsequently, an asymmetric fusion algorithm is used to process these data.

[0162] One of the processing methods is to use the fine texture and color information as the guided filtering input and the spectral index map as the guide image. The spectral index map can reflect the spectral abnormalities related to pathological changes and their spatial distribution. By using it as a guide image, the guided filtering process will adjust the filtering strength of the texture and color information according to the abnormal areas indicated by the spectral index map, so that the texture and color information corresponding to these spectral abnormal areas are retained or enhanced in the fusion result, thereby visually highlighting these abnormal areas.

[0163] Another processing method is to use the fine texture and color information as the guided filtering input and the principal component score map as the guide image. The principal component score map captures the main variation mode of the hyperspectral data and can provide spectral structure information as guidance. By using the principal component score map as guidance, the fusion result can retain the texture and color of the RGB image while incorporating the spectral features of the hyperspectral data. On this basis, the information of each local region after guided fusion is weighted combined or selectively spliced, which can be processed according to the characteristics or fusion effect of different regions, further optimizing the quality of the overall fusion image and avoiding the impact on information expression caused by global uniform processing. Through the above two methods, the spectral information of the hyperspectral image guides the visual information expression of the RGB image, so that the fused tongue image can accurately present the target tongue image features, especially those features that are abnormal in the spectrum but not easily detected in the RGB image, improving the recognition of the target tongue image features and enabling the fusion image to accurately reflect the physiological and pathological state of the tongue surface, providing an image basis for subsequent tongue image analysis and diagnosis.

[0164] On the other hand, as shown in Figure 2 The present application provides a tongue image analysis system 100 based on RGB image and spectral data fusion, which comprises:

[0165] The acquisition module 11 is used to acquire the RGB image and the hyperspectral image captured respectively for the tongue surface of the same subject, and obtain the RGB image information and the hyperspectral image information through image analysis;

[0166] The regional feature determination module 12 is configured to determine at least one target local region in the RGB image and the hyperspectral image, and at least one target tongue image feature in the target local region.

[0167] The dominant modal determination module 13 is configured to determine, based on a preset medical knowledge rule base, dominant modal information for representing the target tongue image feature in the RGB image information and the hyperspectral image information for the target local region and the target tongue image feature.

[0168] The data preparation module 14 is configured to generate guide data based on the dominant modal information, and extract target data related to the target tongue image feature from the RGB image information and the hyperspectral image information.

[0169] The local fusion module 15 is configured to perform asymmetric fusion algorithm processing on the target data by using the guide data, to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to strengthen the expression of the target tongue image feature in the fusion result of the target local region.

[0170] The image generation module 16 is configured to generate a fused tongue image based on the local fusion result of each target local region.

[0171] In some embodiments, the system can be implemented as an integrated system. In this system, the acquisition module can include an RGB camera and a hyperspectral camera configured to capture images of the tongue surface of the same subject synchronously or near-synchronously, and transmit the captured raw image data to a processing unit. The region feature determination module can be part of a software program running on the processing unit, which is designed to execute an image segmentation algorithm to identify the tongue region, and further execute a feature detection algorithm, such as spot detection or texture analysis, within the tongue region to determine the target local region and the target tongue image feature. The dominant modality determination module can also be part of the software program running on the processing unit, which is configured to access a pre-set medical knowledge rule base stored in a storage medium, and determine which of the RGB information or the hyperspectral information is more suitable for representing the feature according to the logic in the rule base, in combination with the region and feature information, to determine the dominant modality information. The data preparation module can also be part of the software program, which is configured to generate guided data according to the dominant modality information, such as performing guided filtering or color space conversion on the RGB image to generate a guided image when the RGB information is the dominant modality; at the same time, it extracts target data related to the target feature from the raw image data, such as extracting the spectral vector of the corresponding region in the hyperspectral image. The local fusion module can implement an asymmetric fusion algorithm, such as performing guided filtering algorithm when the guided data is the structural guided image of RGB and the target data is the spectral vector image of hyperspectral, taking the spectral vector image as input and the structural guided image as guided image to achieve spatial alignment and enhancement of spectral features. The image generation module can be configured to receive the local fusion results from the local fusion module, and integrate these results, such as through splicing or weighted superposition, to finally generate a complete fusion tongue image.

[0172] The entire processing flow can be completed on one or more processing units, which can include but are not limited to central processing units (CPUs), graphics processing units (GPUs) or digital signal processors (DSPs). The same or similar parts among various embodiments in this specification can be referred to each other. In particular, for the embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiments.

[0173] The above-described embodiments of the present application do not constitute a limitation on the protection scope of the present application.

Claims

1. A tongue image analysis method based on the fusion of RGB images and spectral data, characterized in that, include: RGB and hyperspectral images of the tongue surface of the same subject were captured separately, and RGB and hyperspectral image information was obtained through image analysis. Determine at least one target local region and at least one target tongue feature within the target local region in the RGB image and hyperspectral image; Based on a preset medical knowledge rule base, for the target local region and the target tongue image features, the dominant modality information used to characterize the target tongue image features is selected from the RGB image information and the hyperspectral image information. Guidance data is generated based on the dominant modality information, and target data related to the target tongue image features is extracted from the RGB image information and the hyperspectral image information. The target data is processed by an asymmetric fusion algorithm using the guiding data to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to enhance the expression of the target tongue image features in the fusion result of the target local region; Based on the local fusion results of each target local region, a fused tongue image is generated.

2. The tongue image analysis method according to claim 1, characterized in that, After acquiring RGB and hyperspectral images of the tongue surface of the same subject, the process also includes: The RGB image and hyperspectral image are positionally calibrated to ensure that the coordinates of each point on the tongue surface in the spatial coordinate system remain consistent in the RGB image and hyperspectral image. The first preprocessing operation is performed on the position-calibrated RGB image and hyperspectral image, the first preprocessing operation including at least one of noise filtering, image enhancement, geometric correction and illumination correction.

3. The tongue image analysis method according to claim 2, characterized in that, The step of determining at least one target local region in the RGB image and hyperspectral image, and at least one target tongue feature within the target local region, specifically includes: An edge-detection-based tongue segmentation algorithm is used to segment the tongue region and obtain multiple target local regions. The feature detection algorithm is applied to each local region of the target to obtain the type information of the target tongue image features, wherein the feature detection algorithm includes at least one of spot detection, texture analysis and edge detection.

4. The tongue image analysis method according to claim 1, characterized in that, The step of determining the dominant modality information for characterizing the target tongue features, based on a preset medical knowledge rule base, from the RGB image information and the hyperspectral image information, specifically includes: For the target local region and the target tongue image features, calculate the RGB image structural feature evaluation value obtained based on the RGB image information, and the hyperspectral image spectral feature evaluation value obtained based on the hyperspectral image information; The structural feature evaluation value of the RGB image and the spectral feature evaluation value of the hyperspectral image are compared, and the one with the larger evaluation value is determined as the dominant mode of the image data. The dominant knowledge base modality corresponding to the target local region and the target tongue image features is obtained from the preset medical knowledge rule base. Determine whether the dominant mode of the knowledge base is consistent with the dominant mode of the image data. If they are inconsistent, obtain the knowledge base weight of the target local region from the preset medical knowledge rule base, determine the data weight based on the preprocessing quality of the RGB image information and the hyperspectral image information, perform weighted fusion of the dominant mode of the knowledge base based on the knowledge base weight, and perform weighted fusion of the dominant mode of the image data based on the data weight to obtain the target dominant mode information. If they match, the dominant modality of the knowledge base is used as the dominant modality information in the RGB image information and the hyperspectral image information to characterize the features of the target tongue image.

5. The tongue image analysis method according to claim 4, characterized in that, The steps of obtaining the knowledge base weight of the target local region from the preset medical knowledge rule base, and determining the data weight based on the preprocessing quality of the RGB image information and the hyperspectral image information, specifically include: Obtain tongue image feature type information related to the target local region; Based on the acquired tongue image feature type information, an evaluation criterion is selected to quantify the reliability of the target local region; Based on the selected evaluation criteria, the reliability of the target local region is quantified to obtain a value for the reliability of the target local region. The reliability value of the target local area is converted into the knowledge base weight through a preset conversion method; as well as, Calculate the quality index values ​​of the first preprocessing operation of the RGB image information and the hyperspectral image information in the target local area; Based on the quality index value, the credibility of the image data is quantified, and the credibility is converted into the data weight through a second preset conversion method.

6. The tongue image analysis method according to claim 1, characterized in that, The step of generating guidance data based on the dominant modality information and extracting target data related to the target tongue features from the RGB image information and the hyperspectral image information specifically includes: When RGB image information is used as the dominant modal information, the RGB image is filtered using a guided filter to generate a structured guided map that preserves edge details and smooths unstructured areas; and color space components are extracted as a color guided map; the structured guided map or the color guided map is used as guided data. Extract the spectral feature vector map from the hyperspectral image that corresponds to the local region of the target in the RGB image and is related to the tongue features of the target, as the target data.

7. The tongue image analysis method according to claim 6, characterized in that, The step of using the guiding data to process the target data using an asymmetric fusion algorithm to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to enhance the expression of the target tongue image features in the fusion result of the target local region, specifically includes: Using the spectral feature vector map as input to the guided filter and the structural guided map as the guided image, an asymmetric fusion algorithm is applied to align the output spectral features spatially with the RGB structure and enhance details; or, The spectral feature vector map is used as the input image for guided filtering, and the color guide map is used as the guide image. The brightness, saturation, or intensity of a specific color channel of the corresponding pixel in the color guide map is adjusted using hyperspectral feature values.

8. The tongue image analysis method according to claim 1, characterized in that, The step of generating guidance data based on the dominant modality information and extracting target data related to the target tongue features from the RGB image information and the hyperspectral image information specifically includes: When hyperspectral image information is used as the dominant modal information, the spectral index map obtained by band operation on the hyperspectral image, or the principal component score map obtained by principal component analysis on the hyperspectral image, is used as the guiding data. The fine texture and color information corresponding to the target local region in the hyperspectral image and related to the target tongue features are extracted from the RGB image as target data.

9. The tongue image analysis method according to claim 8, characterized in that, The step of using the guiding data to process the target data using an asymmetric fusion algorithm to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to enhance the expression of the target tongue image features in the fusion result of the target local region, specifically includes: Using the fine texture and color information as input to the guided filter, and the spectral index map as the guiding image, an asymmetric fusion algorithm is applied to visually enhance abnormal regions indicated by the spectral index in the fused tongue image; or... The fine texture and color information are used as input to the guided filter, and the principal component score map is used as the guided image. The information of each local region after guided fusion processing is weighted and combined or selectively stitched together.

10. A tongue image analysis system based on the fusion of RGB images and spectral data, characterized in that, The system includes: The acquisition module is used to acquire RGB and hyperspectral images captured on the tongue surface of the same subject, and obtain RGB and hyperspectral image information through image analysis. A region feature determination module is used to determine at least one target local region and at least one target tongue image feature within the target local region in the RGB image and hyperspectral image; The dominant mode determination module is used to determine the dominant mode information for characterizing the target tongue image features based on a preset medical knowledge rule base, choosing one from the RGB image information and the hyperspectral image information; The data preparation module is used to generate guidance data based on the dominant modality information and extract target data related to the target tongue image features from the RGB image information and the hyperspectral image information; The local fusion module is used to process the target data using the guiding data with an asymmetric fusion algorithm to obtain a local fusion result, wherein the asymmetric fusion algorithm is used to enhance the expression of the target tongue image features in the fusion result of the target local region; An image generation module is used to generate a fused tongue image based on the local fusion results of each of the target local regions.

Citation Information

Patent Citations

  • Hyperspectral image super-resolution reconstruction method and device and electronic equipment

    CN113139902A

  • Attention mechanism guided matching association hyperspectral and RGB video fusion tracking method

    CN117689692A