Multispectral tongue picture feature extraction and classification method based on multidimensional feature fusion
Through multi-spectral tongue image feature extraction and classification method of multi-dimensional feature fusion, the spatial and spectral features of tongue image were extracted in combination with CNN and MLP, and multi-scale feature fusion was used to perform multi-scale feature fusion, which solved the problem of insufficient fusion of image data and spectral data in traditional methods, and improved the classification accuracy and reliability of tongue image analysis.
Patent Information
- Application Number
- CN202510316848.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional tongue image analysis methods rely on a single data source, resulting in insufficient fusion of image data and spectral data, affecting the comprehensiveness of feature extraction and classification accuracy.
The multi-spectral tongue image feature extraction and classification method using multi-dimensional feature fusion is used to extract the tongue area through the K-means segmentation algorithm, and spatial and spectral features are extracted in combination with 2D convolutional neural network (CNN) and multi-layer perceptron (MLP), and multi-scale feature fusion is used to fusion with the improved HiFuse network. Finally, the fusion features are input to the Transformer module for global modeling.
It improves the multimodal fusion accuracy and classification accuracy of tongue image data, reduces the influence of lighting and subjective judgment, and enhances the accuracy and reliability of analysis.
Smart Images

Figure CN120107698A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing, and in particular to a multi-spectral tongue image feature extraction and classification method with multi-dimensional feature fusion; Background Art Traditional Chinese medicine tongue image analysis relies on doctors to observe tongue color, tongue coating thickness, tongue shape and other information through naked eye for diagnosis; however, due to the lack of unified standards and standardized processes, the analysis results are often highly subjective and easily affected by factors such as lighting, angle and doctor experience, leading to judgment errors; in addition, the existing tongue image analysis methods still have the problem of insufficient fusion of image data and spectral data, which affects the comprehensiveness of feature extraction and classification accuracy; Existing methods often rely on only a single data source, such as RGB images or spectral data. Although RGB images can provide basic color information of the tongue, they are weak in capturing the details of the tongue coating and are easily affected by lighting conditions and equipment errors. Spectral data can reveal subtle changes in the tongue coating, but due to the complex coupling relationship between its spectral characteristics and spatial features, traditional methods fail to effectively integrate these two types of data, resulting in the inability to fully extract multi-dimensional information of the tongue image, which ultimately affects the accuracy and reliability of the analysis. Summary of the invention In order to solve the problems of the existing technology, this study proposed a multi-spectral tongue image feature extraction and classification method with multi-dimensional feature fusion, aiming to solve the problems of insufficient fusion of image data and spectral data and poor feature extraction effect in current tongue image analysis; specifically, this method adopts the following innovative steps: first, the K-means segmentation algorithm is used to accurately extract the tongue body area to ensure the quality of image and spectral data, and the elastic registration technology is used to solve the spatial position offset problem between images of different bands; secondly, the 2D convolutional neural network (CNN) is combined to extract the spatial features of the tongue image and the multi-layer perceptron (MLP) is used to extract the spectral features, and the improved HiFuse network is used for multi-scale feature fusion, which optimizes the complementarity of image data and spectral data and improves the accuracy of data fusion; finally, the fused features are input into the Transformer module for global modeling, thereby further improving the classification accuracy of the tongue image; In view of the above-mentioned defects of the prior art, the present invention aims to solve the problem that although the traditional RGB visible light imaging method can capture the basic color information of the tongue, when analyzing the subtle changes in tongue coating and tongue quality, the amount of information is insufficient, and it is easily affected by lighting conditions and subjective judgment, and the results caused by single modality data are inaccurate; the present invention provides a multi-spectral tongue image feature extraction and classification method with multi-dimensional feature fusion, and the specific steps are as follows: First, image and spectral data were acquired and preprocessed (including spatial alignment, spectral calibration, denoising, normalization, etc.); the tongue area was extracted using the K-means segmentation algorithm to ensure data quality; Use 2D CNN to extract the depth and spatial information of tongue images, such as tongue shape and texture distribution. In the spectral data part, use multi-layer perceptron (MLP) to extract spectral features. It is used to process the global nonlinear relationship in spectral data and the coupling and feature correlation between different bands. Use the improved HiFuse network's multi-scale feature fusion architecture to fuse the multi-layer spatial features extracted by 2D CNN and the global spectral features extracted by MLP. Broadcast the spectral features to the image space for feature alignment. Use the CBAM module to weight the attention of image features and spectral features at different levels, dynamically adjust the weights of channels and spaces, and ensure the complementarity of multimodal data. Output the fused multimodal features. Inputting the fused features into the Trans-CNN classifier, globally modeling the fused features, generating a one-dimensional feature sequence, classifying the feature sequence, and outputting a final diagnosis or symptom prediction result; Since the tongue is photographed in different bands, there is a spatial position offset due to the influence of equipment or external factors. The base band (usually the G channel in the RGB band) is used as the reference image; the images of other bands are rigidly registered (translated, rotated, scaled) or elastically registered. Suppose the reference image is The image to be registered is , the registration process can be expressed as: Where is the transformation function (including rotation matrix and translation vector); For each pixel, is the transformation function (including rotation matrix and translation vector); For each pixel (For example, ), calculate its correlation with each cluster center The Euclidean distance of: in, is the dimension of the feature, assigning pixels to the nearest cluster center: Calculate the cluster center of each cluster as the mean of all points in the cluster, and repeat the assignment and update steps until the cluster center no longer changes or the maximum number of iterations is reached; ensure that the pixel positions of the image data and the spectral data are consistent; In the process of spectral feature extraction, the ID CNN model is used to perform multi-layer convolution and maximum pooling operations to obtain a 24-dimensional tongue coating spectral feature vector; for the area of interest in the tongue coating, PCA is used for preprocessing to obtain an input image composed of the first three principal components; a two-dimensional CNN model with four layers of convolution and maximum pooling is proposed; finally, a 48-dimensional spatial feature vector is obtained; after multi-level feature extraction of image data and spectral data in their respective network modules, feature information of different coarseness and fineness is extracted through multi-scale convolution layers, and then these features are weighted fused to obtain multimodal data fusion features of different levels and scales; The hierarchical multi-scale feature fusion module of the HiFuse network is used to combine the features of different levels in the image data and spectral data. In the feature extraction process of each level, the features of different scales will be effectively fused to maximize the complementary advantages of image data and spectral data, thereby improving the model's ability to process complex data. After the multi-scale features are fused, the fused features are further modeled and processed through the Transformer module; the fused multi-scale features are sent to the Transformer module for global modeling; the feature encoding layer converts the fused features of the image data and the spectral data into one-dimensional sequence features, which are sent to the classifier for final prediction; this step ensures that the cross-modal dependencies in the image and spectral data can be effectively captured by the Transformer network, thereby improving the classification accuracy and model performance; the performance of the proposed method is evaluated using statistical indicators, including overall accuracy, precision, recall and Fl-scorel45; these four indicators can be defined as follows: Among them, true positive (TP) represents the number of correctly classified tongue coating grades, true negative (TN) represents the number of correctly classified tongue coating grades; false positive (TN) represents the number of correctly classified tongue coating grades. (FP) indicates the number of misclassifications as surface grades, and false negatives (FN) indicate the number of misclassifications as other surface grades; BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 1 is a schematic diagram of a process for multi-spectral tongue image feature extraction and classification according to an embodiment of the present invention; Figure 2 is a system architecture diagram of multi-spectral tongue image feature extraction and classification according to an embodiment of the present invention; Figure 3 is a flow chart of tongue image multimodal analysis for multispectral tongue image feature extraction and classification according to an embodiment of the present invention; Figure 4 It is a fusion network structure diagram of multi-spectral tongue image feature extraction and classification in an embodiment of the present invention; DETAILED DESCRIPTION The present invention is described in detail below in conjunction with the accompanying drawings; In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention; based on the embodiments of the present invention, all examples obtained by ordinary technicians in the field without making creative work belong to the scope of protection of the present invention; Embodiment 1: like Figure 1 As shown, a multi-spectral tongue image feature extraction and classification method with multi-dimensional feature fusion is provided. In view of the above-mentioned defects of the prior art, the present invention aims to solve the problem that although the traditional RGB visible light imaging method can capture the basic color information of the tongue, when analyzing the subtle changes of tongue coating and tongue quality, the amount of information is insufficient, and it is easily affected by lighting conditions and subjective judgment, and the results caused by single modality data are inaccurate; the present invention provides a multi-spectral tongue image feature extraction and classification method with multi-dimensional feature fusion, and the specific steps are as follows: First, image and spectral data were acquired and preprocessed (including spatial alignment, spectral calibration, denoising, normalization, etc.); the tongue area was extracted using the K-means segmentation algorithm to ensure data quality; Use 2D CNN to extract the depth and spatial information of tongue images, such as tongue shape and texture distribution. In the spectral data part, use multi-layer perceptron (MLP) to extract spectral features. It is used to process the global nonlinear relationship in spectral data and the coupling and feature correlation between different bands. Use the improved HiFuse network's multi-scale feature fusion architecture to fuse the multi-layer spatial features extracted by 2D CNN and the global spectral features extracted by MLP. Broadcast the spectral features to the image space for feature alignment. Use the CBAM module to weight the attention of image features and spectral features at different levels, dynamically adjust the weights of channels and spaces, and ensure the complementarity of multimodal data. Output the fused multimodal features. Inputting the fused features into the Trans-CNN classifier, globally modeling the fused features, generating a one-dimensional feature sequence, classifying the feature sequence, and outputting a final diagnosis or symptom prediction result; Since the tongue is photographed in different bands, there is a spatial position offset due to the influence of equipment or external factors. The base band (usually the G channel in the RGB band) is used as the reference image; the images of other bands are rigidly registered (translated, rotated, scaled) or elastically registered. Suppose the reference image is The image to be registered is , the registration process can be expressed as: Where is the transformation function (including rotation matrix and translation vector); For each pixel, is the transformation function (including rotation matrix and translation vector); For each pixel (For example, ), calculate its correlation with each cluster center The Euclidean distance of: in, is the dimension of the feature, assigning pixels to the nearest cluster center: Calculate the cluster center of each cluster as the mean of all points in the cluster, and repeat the assignment and update steps until the cluster center no longer changes or the maximum number of iterations is reached; ensure that the pixel positions of the image data and the spectral data are consistent; In the process of spectral feature extraction, the ID CNN model is used to perform multi-layer convolution and maximum pooling operations to obtain a 24-dimensional tongue coating spectral feature vector; for the area of interest in the tongue coating, PCA is used for preprocessing to obtain an input image composed of the first three principal components; a two-dimensional CNN model with four layers of convolution and maximum pooling is proposed; finally, a 48-dimensional spatial feature vector is obtained; after multi-level feature extraction of image data and spectral data in their respective network modules, feature information of different coarseness and fineness is extracted through multi-scale convolution layers, and then these features are weighted fused to obtain multimodal data fusion features of different levels and scales; In the input stage, the RGB image of the tongue (size is H×W×3) is input to extract the spatial and depth features of the tongue, such as morphology and texture distribution. The multispectral data of the tongue, which contains spectral information of different bands, is input to capture the spectral characteristics of the tongue surface reflection. A 2D CNN is constructed using multiple convolutional layers and pooling layers to capture local features (such as texture distribution and edge information) and global features (such as morphology and depth changes) of the tongue image. Multi-scale spatial features are gradually extracted from the shallow layer (texture details) to the deep layer (semantic morphology). A set of multi-level tongue image spatial feature maps with dimensions of C×H×W is obtained, where C is the number of channels. When using a multi-layer perceptron (MLP) to extract spectral features, the spectral data is converted into a vector representation, with each spectral point as input; the MLP consists of multiple fully connected layers and activation functions (such as ReLU) to extract nonlinear features in the spectral data; capture the coupling relationship between different bands, as well as the global characteristics and correlations in the spectral data; generate a global spectral feature vector with a dimension of 1×D, where D is the spectral feature dimension; The improved HiFuse feature fusion module is used to combine different levels of features in image data and spectral data; in the feature extraction process of each level, features of different scales will be effectively fused to maximize the complementary advantages of image data and spectral data, thereby improving the model's ability to process complex data; input multi-layer tongue image spatial features from 2D CNN and global spectral features extracted from MLP; repeat and expand the spectral feature vector to the dimension C×H×W that matches the image spatial features; achieve alignment of spectral features in image space through interpolation or other methods; use the multi-scale feature fusion mechanism of the HiFuse network to gradually interact low-level and high-level spatial features with spectral features to retain the complementary information of the two; use feature weighting and attention mechanism to adjust the fusion weight; Calculate weights for the channel dimension (C) to give more important features higher weights; calculate weights for the spatial dimension (H×W) to highlight the features of key areas in the image space; combine channel and spatial weights to dynamically adjust image features and spectral features to ensure the complementarity of multimodal features; After the multi-scale features are fused, the fused features are further modeled and processed through the Transformer module; the fused multi-scale features are sent to the Transformer module for global modeling; the feature encoding layer converts the fused features of the image data and the spectral data into one-dimensional sequence features, which are sent to the classifier for final prediction; this step ensures that the cross-modal dependencies in the image and spectral data can be effectively captured by the Transformer network, thereby improving the classification accuracy and model performance; the performance of the proposed method is evaluated using statistical indicators, including overall accuracy, precision, recall and Fl-scorel45; these four indicators can be defined as follows: Among them, true positive (TP) represents the number of correctly classified tongue coating grades, true negative (TN) represents the number of correctly classified tongue coating grades; false positive (TN) represents the number of correctly classified tongue coating grades. (FP) indicates the number of misclassifications as surface grades, and false negatives (FN) indicate the number of misclassifications as other surface grades; Finally, in the output stage, a multimodal feature map with dimensions of C×H×W is output, which integrates the spatial information of the tongue image and the global spectral information; this feature includes the color characteristics of the tongue image, the category of the tongue coating, and the nonlinear correlation of the spectral data.
Claims
1. A multi-spectral tongue image feature extraction and classification method with multi-dimensional feature fusion, characterized in that: The following steps are involved: S1: Obtain image and spectral data and perform preprocessing; extract the tongue area using the K-means segmentation algorithm; S2: 2DCNN is used to extract the depth and spatial information of tongue images. In the spectral data part, a multi-layer perceptron is used to extract spectral features to process the global nonlinear relationship in the spectral data and the coupling and feature correlation between different bands; The improved HiFuse network’s multi-scale feature fusion architecture is used to fuse the multi-layer spatial features extracted by 2DCNN and the global spectral features extracted by MLP. The spectral features are broadcasted to the image space for feature alignment. Use the CBAM module to weight the image features and spectral features at different levels, dynamically adjust the weights of channels and spaces, and ensure the complementarity of multimodal data; output the fused multimodal features; S3: Input the fused features into the Trans-CNN classifier, perform global modeling on the fused features, generate a one-dimensional feature sequence, classify the feature sequence, and output the final diagnosis or symptom prediction result.
2. The multispectral tongue image feature extraction and classification method based on multidimensional feature fusion according to claim 1 is characterized in that: The preprocessing includes spatial alignment, spectral calibration, denoising and normalization.
3. The multi-spectral tongue image feature extraction and classification method based on multi-dimensional feature fusion according to claim 2 is characterized in that: The depth and spatial information of the tongue image includes the tongue shape and texture distribution.
4. The multi-spectral tongue image feature extraction and classification method based on multi-dimensional feature fusion according to claim 3 is characterized in that: In the step S1, the reference band is used as the reference image; rigid registration or elastic registration is performed on the images of other bands. Assume that the reference image is , the image to be registered is , the registration process can be expressed as: in is the transformation function; for each pixel , calculate the center of each cluster The Euclidean distance of: in, is the dimension of the feature, assigning pixels to the nearest cluster center: The cluster center of each cluster is calculated as the mean of all points in the cluster, and the allocation and update steps are repeated until the cluster center no longer changes or the maximum number of iterations is reached; ensure that the pixel positions of the image data and the spectral data are consistent.
5. The multi-spectral tongue image feature extraction and classification method based on multi-dimensional feature fusion according to claim 4 is characterized in that: In the step S2, during the spectral feature extraction process, the IDCNN model is used to perform multi-layer convolution and maximum pooling operations to obtain a 24-dimensional tongue coating spectral feature vector.
6. The multi-spectral tongue image feature extraction and classification method based on multi-dimensional feature fusion according to claim 5 is characterized in that: For the area of interest in the tongue coating, PCA is used for preprocessing to obtain an input image composed of the first three principal components; a two-dimensional CNN model with four layers of convolution and maximum pooling is proposed; finally, a 48-dimensional spatial feature vector is obtained; after multi-level feature extraction of image data and spectral data in their respective network modules, feature information of different coarseness and fineness is extracted through multi-scale convolution layers, and these features are weightedly fused to obtain multimodal data fusion features of different levels and scales.
7. The multi-spectral tongue image feature extraction and classification method based on multi-dimensional feature fusion according to claim 6 is characterized in that: The S2 uses the hierarchical multi-scale feature fusion module of the HiFuse network to combine different levels of features in image data and spectral data; in the feature extraction process of each level, features of different scales will be effectively fused to maximize the complementary advantages of image data and spectral data, thereby improving the model's ability to process complex data.
8. The multispectral tongue image feature extraction and classification method based on multidimensional feature fusion according to any one of claims 1 to 7, characterized in that: In the S3 step, after the multi-scale features are fused, the fused features are further modeled and processed through the Transformer module; the fused multi-scale features are sent to the Transformer module for global modeling; the feature encoding layer converts the fused features of the image data and the spectral data into one-dimensional sequence features, which are sent to the classifier for final prediction.