Dynamic closed-loop tongue moss separation method based on improved UNet

By improving the UNet network combined with multiple loss functions and feature visualization technology, the problems of low precision and poor feature interpretability in traditional tongue coating separation are solved, high-precision and robust tongue image quantitative analysis is achieved, and the clarity of tongue segmentation and the credibility of diagnostic results are improved.

CN120635135APending Publication Date: 2025-09-12SHUIFA SMART IND GRP CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510667864.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional tongue coating separation technology has problems such as low accuracy, poor feature interpretability and sensitivity to edge noise. Especially when dealing with areas with blurred tongue edges and complex textures, the segmentation effect is not ideal and it is difficult to adapt to tongue image inputs with different lighting and postures.

Method used

An improved UNet network was adopted, combined with multi-classification cross entropy loss and Dice loss function, and Grad-CAM feature visualization and Sobel operator edge enhancement were introduced to form a dynamic closed-loop tongue coating separation method. The model convergence was optimized by a hybrid loss function, and the key feature areas were located using heat maps and the weights were dynamically adjusted. The dataset was optimized by combining grayscale dimensionality reduction and edge enhancement techniques.

Benefits of technology

It achieves high-precision and high-robustness quantitative analysis of tongue images, improves the clarity of tongue and background segmentation and adaptability to lighting changes, solves the problems of fuzzy boundaries and category imbalance in tongue image segmentation, and enhances feature interpretability and the credibility of diagnostic results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635135A_ABST
    Figure CN120635135A_ABST
Patent Text Reader

Abstract

The invention relates to a dynamic closed-loop tongue moss separation method based on an improved UNet, and belongs to the technical field of medical information and health data processing. The method comprises the following steps: training a tongue body segmentation UNet model by adopting a mixed loss function, segmenting to obtain an image only containing a tongue body, further carrying out coating separation on the image only containing the tongue body by adopting a UNet network, introducing gradient weighting class activation mapping to generate a high-resolution thermodynamic diagram, carrying out gray dimension reduction, and carrying out Sobel operator edge enhancement to generate a gradient amplitude image; and merging the gradient magnitude image and the original tongue-only-containing image, inputting the merged image into the tongue coating separation UNet network to obtain a tongue coating segmentation image, dynamically feeding back and adjusting the weight of each layer of the tongue coating separation UNet network according to the importance of the quantitative characteristics of the thermodynamic diagram, and finally realizing accurate separation of the tongue coating. According to the method, high-precision and high-robustness tongue picture quantitative analysis is realized through a collaborative optimization mechanism of mixed loss function design, dynamic feature visualization and edge enhancement feedback.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a dynamic closed-loop tongue coating separation method based on an improved UNet, and belongs to the technical field of medical information and health data processing. Background Art

[0002] Tongue diagnosis, a distinctive diagnostic method in Traditional Chinese Medicine (TCM), has a comprehensive theoretical basis. It is a clear, accessible, and effective diagnostic method for syndrome diagnosis, and plays a vital role in today's clinical practice.

[0003] Traditional tongue diagnosis's qualitative and quantitative criteria are influenced by the physician's academic level and clinical experience, resulting in significant uncertainty. With the advancement of computer technology, researchers have begun utilizing methods such as deep learning and machine vision, combined with the extensive clinical experience of TCM experts, to conduct qualitative, quantitative, and localized research on tongue images in Traditional Chinese Medicine. However, due to the varying shapes of tongues, and the close proximity of tongue color to lips and face color when using machine vision for tongue diagnosis, tongue segmentation is less than ideal. Tongue coating separation is also more susceptible to features such as tongue surface color, resulting in less than ideal results. Patent CN 116071373A discloses a method for automatic tongue segmentation based on a U-net model fused with PCA. The method comprises: acquiring tongue image data; preprocessing the acquired tongue image data; constructing an improved U-net segmentation model based on PCA for the preprocessed tongue image data; calculating the weights of the various parameters of the U-net segmentation model using principal component analysis; and pruning the U-net segmentation model. The improved U-net segmentation model is trained, tested, and verified; the tongue image data is segmented and processed using the trained improved U-net segmentation model to obtain the segmentation results. This patent directly applies PCA to the calculation of UNet parameter weights, which may lead to feature information loss, especially when processing areas with blurred tongue edges and complex textures, which may weaken the model's ability to capture details. The convolution kernel parameters of UNet have local spatial correlation, and the global dimensionality reduction of PCA may destroy this local correlation, resulting in a decrease in segmentation accuracy. Moreover, the pruning strategy is difficult to adapt to tongue image inputs with different lighting and postures, which may lead to unstable segmentation results. Summary of the Invention

[0004] The purpose of the present invention is to overcome the shortcomings of existing tongue coating separation technology, and to provide a dynamic closed-loop tongue coating separation method based on improved UNet to achieve high-precision and high-robustness tongue image quantitative analysis, targeting the problems of low coating separation accuracy, poor feature interpretability and edge noise sensitivity in traditional tongue image analysis.

[0005] The technical solution adopted by the present invention is:

[0006] The dynamic closed-loop tongue coating separation method based on the improved UNet includes the following steps:

[0007] S1. Collect and preprocess tongue images, input the labeled tongue image data into the UNet network to train the tongue segmentation model, and use a hybrid loss function to improve the convergence model;

[0008] S2. After the tongue image data is processed through the tongue segmentation UNet model, an image dataset containing only the tongue is obtained;

[0009] S3. For images containing only the tongue, a UNet network is further used to perform preliminary moss separation. Gradient-weighted class activation mapping (Grad-CAM) is introduced in the final layer of the UNet network to generate a high-resolution heat map (i.e., class activation map) to visualize important regions in the image, quantify feature contributions, and locate key feature regions.

[0010] S4. Convert the heat map-guided feature map into a grayscale image for grayscale dimensionality reduction, and then use the Sobel operator to enhance the edge of the grayscale image to generate a gradient magnitude image;

[0011] S5. The gradient magnitude image generated after edge enhancement is combined with the original tongue image as the input of a new round of tongue coating separation UNet network to obtain a tongue coating segmentation image;

[0012] S6. Quantify the feature importance (i.e., the activation intensity of the heat map) based on the high-contribution feature areas located by the heat map, dynamically feedback-adjust the weights of each layer of the tongue coating separation UNet network, iteratively strengthen the activation response of tongue coating-related features, and convert the iterative feature map into a grayscale image and a gradient amplitude image. The gradient amplitude image is merged with the original tongue body image as the input of the tongue coating separation UNet network to finally obtain a high-precision tongue coating segmentation image.

[0013] In the above method, the hybrid loss function described in step S1 is a weighted sum of the multi-classification cross entropy loss function and the Dice coefficient loss function:

[0014] L=ωL CCE +(1-ω)L DL ,

[0015] In the formula, L CCE is the multi-classification cross entropy loss function; L DL is the Dice loss function, ω and (1-ω) are L CCE 、L DL The weight of

[0016] The formula for the above multi-classification cross entropy loss function is:

[0017]

[0018] In the formula, m is the number of samples, k is the number of categories, and y ij is the one-hot encoding of sample i in category j, Represents the predicted probability of sample i in category j;

[0019] The Dice loss function formula is:

[0020]

[0021] In the formula, K is the set of all categories, y ij is the one-hot encoding of sample i in category j, represents the predicted probability of sample i in category j, and ε is a small constant.

[0022] In step S3 above, weights are first calculated based on the gradients of each feature channel of the feature map generated by the penultimate layer of the UNet network. The feature map is then multiplied by the calculated weights, and then nonlinear processing is performed through the ReLU function to finally obtain a high-resolution heat map, which is the Grad-CAM visualization result. The specific process is as follows:

[0023] (1) Assume that the penultimate layer of the UNet network generates K feature maps A with width μ and height ν k (k∈K), these feature maps are globally pooled and then linearly combined to produce the confidence score y of the tongue coating category c c ;

[0024] (2) Tongue coating category c is back-propagated on feature layer A to obtain gradient information Perform global average pooling on the calculated gradient in the dimensions of width i and height j to obtain the weight

[0025]

[0026] Where Z is the product of the width i and height j of the feature layer; c is the tongue coating category; y c is the confidence score corresponding to the category obtained by forward propagation; k is the selected channel; A is the feature layer, which is the output of the penultimate convolutional layer here; is the data of feature layer A with coordinates (i, j) in channel k; The gradient information of tongue coating category c is obtained by backpropagation on feature layer A;

[0027] (3) Multiply the feature map with the calculated weights, and then perform nonlinear processing through the ReLU function. The final heat map is the Grad-CAM visualization result. Partial linearization of the downstream deep network representing feature map A and obtaining the importance of feature map k of tongue coating category c; Grad-CAM then performs a weighted combination of the forward activation maps and obtains the discriminant localization map of tongue coating category c through ReLU

[0028]

[0029] Last passed A heatmap of the same size as the convolutional feature map is generated, where darker colors indicate greater importance for moss separation.

[0030] The grayscale dimensionality reduction formula in the above step S4 is:

[0031]

[0032] In the formula, G represents the grayscale image obtained, k represents the channel, β is the heat map weight adjustment factor, α k is the gradient weight of the kth channel, and m is the channel index variable.

[0033] The generation process of the gradient magnitude image is:

[0034]

[0035] G smooth =G*K Gaussian ,

[0036]

[0037] K y =K x T ,

[0038] G x =G smooth *K x ,

[0039] G y =G smooth *K y ,

[0040]

[0041] In the formula, * is the convolution operation, K Gaussian is a Gaussian smoothing kernel, with a higher weight for the central pixel and gradually decaying towards the surrounding pixels, and the sum of the weights is 1, ensuring that the overall brightness of the image remains unchanged; G represents the grayscale image obtained, G smooth It is the image obtained by convolution operation between the original grayscale image and the Gaussian kernel; K x , K yis the Sobel operator convolution kernel, which detects horizontal and vertical edges respectively; E is the gradient amplitude image, and the high-value area corresponds to the tongue contour mutation point.

[0042] In step S5 of the above method, the channel weights of the feature map are dynamically adjusted according to the activation intensity of the heat map, the weights of high-contribution channels are increased, and the weights of low-contribution channels are reduced.

[0043] Convolution weight adjustment formula:

[0044]

[0045] is the weight matrix of layer t; η is the learning rate; Indicates that when the convolution kernel weight W conv The rate of change of total loss L when a small change occurs.

[0046] The beneficial effects of the present invention are:

[0047] The present invention adopts an improved UNet network, combined with multi-class cross entropy loss (categorical crossentropy) and Dice loss, to achieve accurate segmentation of the tongue and background. The Dice loss function improves the boundary clarity of the tongue-background segmentation by strengthening the pixel matching of the tongue contour area, thus solving the problem of blurred boundaries in tongue segmentation. The multi-class cross entropy loss function further improves the segmentation accuracy through dynamic weighting, and has strong robustness to interference such as uneven lighting and tongue deformation, solving the problem of imbalanced tongue segmentation categories. This tongue segmentation technology solves the two major technical bottlenecks of the traditional single loss function in tongue segmentation, as well as the problem of difficult segmentation of lips and face due to the similarity of color and texture to the tongue.

[0048] The present invention introduces Grad-CAM feature re-optimization technology in tongue coating separation, establishes a synergistic mechanism of feature visualization and edge enhancement, and forms a closed loop of "segmentation-visualization-edge enhancement-parameter tuning-segmentation", which solves the problems of poor feature interpretability and inaccurate tongue coating separation. Through gradient backpropagation, a class activation heat map is generated to quantitatively display deep features such as tongue surface texture and color gradient areas that tongue coating discrimination depends on, making the model decision basis transparent. Doctors can intuitively verify the consistency of features with tongue diagnosis theory, thereby improving the credibility of the diagnosis results; the key discriminant feature areas are located through heat maps, and combined with grayscale conversion to eliminate color interference. The Sobel operator is combined with Gaussian smoothing preprocessing technology to extract high-contrast edge features while suppressing noise, directly improving the model's sensitivity to the subtle structure of the tongue, and avoiding missegmentation caused by noise interference in traditional methods. After locating the high-contribution feature areas based on the heat map, the activation response of the relevant convolution kernel is strengthened through weight tuning, so that the model adaptively focuses on anatomical features that are strongly related to tongue coating discrimination, avoiding overfitting redundant information. The present invention integrates grayscale dimensionality reduction and edge enhancement preprocessing to achieve high information density and low noise input optimization, and optimizes the data set to guide the model to further improve the accuracy of moss separation.

[0049] The present invention constructs a solution that emphasizes both enhanced interpretability and lightweight deployment, meeting the dual needs of clinical diagnosis and industrialization. This technology breaks through the limitations of traditional models that rely on static parameter updates and realizes the co-evolution of feature selection and model training. The present invention achieves high-precision and high-robustness tongue image quantitative analysis through a collaborative optimization mechanism of hybrid loss function design, dynamic feature visualization and edge enhancement feedback. The model supports tongue image data input and embedded device deployment, ensuring the interpretability of clinical diagnosis while meeting real-time processing requirements, providing reliable technical support for the objectivity and intelligence of traditional Chinese medicine tongue diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 Schematic diagram of the process of the present invention;

[0051] Figure 2 This is a schematic diagram of the UNet network structure of the present invention. DETAILED DESCRIPTION

[0052] The present invention is further described below with reference to specific embodiments.

[0053] Example 1 A dynamic closed-loop tongue coating separation method based on an improved UNet comprises the following steps (eg Figure 1 (shown) as follows:

[0054] S1. Collect and preprocess tongue images, input the labeled tongue image data into the UNet network to train the tongue segmentation model, and use the hybrid loss function to improve the convergence model:

[0055] After determining the equipment and lighting environment, we collected a number of tongue image data, screened the collected images, removed the blurred and incomplete images, and then annotated the tongue to obtain a self-made tongue image dataset.

[0056] The labeled tongue image data is input into the UNet network to train the tongue segmentation model. The UNet network is a U-shaped network, which is divided into two parts (such as Figure 2 ).

[0057] The left half of UNet is the downsampling module, which extracts high-dimensional features by gradually reducing the spatial dimensions of the input data. Its core consists of five groups of convolutional layers and a maximum pooling layer (MaxPooling). The first and second groups use two 3*3 convolution operations with 64 and 128 convolution kernels, respectively. The third, fourth, and fifth groups use three 3*3 convolution operations with 256, 512, and 512 convolution kernels, respectively. A Batch Normalization (BN) layer is added after each convolution operation to normalize the features of each layer of the network, making the feature distribution of each layer more uniform. This improves the model's convergence speed and fault tolerance.

[0058] The right half of UNet is centrally symmetrical with the left half. It consists of a series of upsampling layers. Its core is the five groups of upsampling modules and convolutional layers corresponding to the downsampling. In addition to the deep abstract features obtained by the upsampling module in the previous layer, the input of each group of convolutional layers also includes the shallow local features output by the corresponding downsampling layer. The deep features and shallow features are fused by splicing, thereby restoring the details of the feature map and ensuring that the corresponding spatial information dimension remains unchanged.

[0059] Image segmentation is to classify pixels one by one. Since the linear model has insufficient expressive power, the ReLU activation function is introduced in the hidden layer in the middle of the network to add nonlinear factors to improve the model's expressive power and solve the classification problem:

[0060] f(x)=max(0,x),

[0061] Where x represents the linear weighted sum of the previous layer of the network, and f(x) is the output value after ReLU activation, and it is a non-negative value.

[0062] In the last layer of the network, the softmax activation function is used to calculate the probability that the pixel of the input tongue image belongs to the tongue class, which is defined as follows:

[0063]

[0064] Where r represents the pixel; o represents the category; W represents the weight; ρ(ο|r,W) represents the probability that the pixel point r belongs to category o under the weight parameter; f o (r, W) represents the input pixel r and the weight W that together determine the calculated features of category o. m is the index that traverses all categories, M is the total number of categories, and exp performs an exponential operation on the feature value, amplifying the feature difference and making the probability distribution more significant.

[0065] In order to improve the performance of the model, a hybrid loss function is used to ensure the correctness of the learning direction from both global and local aspects, and accelerate the convergence of the model.

[0066] The hybrid loss function is a weighted sum of the multi-classification cross entropy loss function and the Dice coefficient loss function:

[0067] L=ωL CCE +(1-ω)L DL ,

[0068] In the formula, L CCE is the multi-classification cross entropy loss function; L DL is the Dice loss function, ω and (1-ω) are L CCE 、L DL The weight of .

[0069] The formula for the multi-classification cross entropy loss function is:

[0070]

[0071] In the formula, m is the number of samples, k is the number of categories, and y ij is the one-hot encoding of sample i in category j, Represents the predicted probability of sample i in category j;

[0072] The Dice loss function formula is:

[0073]

[0074] In the formula, K is the set of all categories, y ij is the one-hot encoding of sample i in category j, represents the predicted probability of sample i in category j, and ε is a small constant.

[0075] S2. After the tongue image data is processed by the tongue segmentation UNet model, an image dataset containing only the tongue is obtained.

[0076] S3. A UNet network is further used to perform preliminary moss separation on images containing only the tongue. Gradient-weighted class activation mapping (Grad-CAM) is introduced in the last layer of the UNet network to generate a high-resolution heat map (i.e., class activation map) to visualize important areas in the image, quantify feature contributions, and locate key feature areas.

[0077] A UNet network is used to further classify tongue coating and tongue texture in images containing only the tongue body. Grad-CAM analyzes the gradients in the last convolutional layer of the UNet to decode the importance of each feature map to the tongue coating category. Grad-CAM is used in the last layer of the UNet to generate a class activation map to visualize important areas in the image. By analyzing the focus of different tongue body feature extraction methods, the model is guided and the dataset is optimized to further improve model performance and enhance the accuracy of tongue body segmentation.

[0078] Assume that the penultimate layer of the UNet network generates K feature maps A with width μ and height ν k (k∈K), these feature maps are globally pooled and then linearly combined to produce the confidence score y of the tongue coating category c c .

[0079] Tongue coating category c is back-propagated on feature layer A to obtain gradient information Perform global average pooling on the calculated gradient in the dimensions of width i and height j to obtain the weight

[0080]

[0081] Where Z is the product of the width i and height j of the feature layer; c is the tongue coating category; y c is the confidence score corresponding to the category obtained by forward propagation; k is the selected channel; A is the feature layer, which is the output of the penultimate convolutional layer here; is the data of feature layer A with coordinates (i, j) in channel k; The gradient information of tongue coating category c is obtained by backpropagation on feature layer A.

[0082] The feature map is multiplied by the calculated weights, and then nonlinear processing is performed through the ReLU function to highlight the areas that have the greatest impact on the classification results. The final heat map is the Grad-CAM visualization result. The darker the color, the greater the impact on the tongue body and tongue coating classification results. Partial linearization of the downstream deep network representing feature map A and obtaining the importance of feature map k of tongue coating category c; Grad-CAM then performs a weighted combination of the forward activation maps and obtains the discriminant localization map of tongue coating category c through ReLU

[0083]

[0084] Last passed The heat map generated from the convolution feature map has the same size. The darker the color, the more important it is for the moss separation operation. The higher the value, the greater the contribution of the feature map to the classification, and the corresponding convolution kernel weight needs to be increased.

[0085] In the formula, the role of ReLU is to make the final output greater than zero and suppress the weight parts that are not of interest; c is the selected tongue coating category; k is the selected channel; A is the feature layer; is the weight of tongue coating category c on the kth channel of feature layer A; k is the weight matrix of feature layer A in channel k.

[0086] S4. Convert the heat map-guided feature map into a grayscale image for grayscale dimensionality reduction, and then use the Sobel operator to enhance the edge of the grayscale image to generate a gradient magnitude image:

[0087] Heat map-guided feature grayscale dimensionality reduction reduces computational energy consumption and removes color interference:

[0088]

[0089] In the formula, G represents the grayscale image obtained, k represents the channel, β is the heat map weight adjustment factor, α k is the gradient weight of the kth channel, m is the channel index variable, such as when k = 1, the denominator traverses all channels m = 1, 2, L.

[0090] The grayscale image is then used to generate a gradient magnitude image using the Sobel operator, thereby enhancing the target contour:

[0091]

[0092] G smooth =G*K Gaussian ,

[0093]

[0094] K y =K x T ,

[0095] G x =G smooth *K x ,

[0096] G y =G smooth *Ky ,

[0097]

[0098] In the formula, * is the convolution operation, K Gaussian is a Gaussian smoothing kernel, with a higher weight for the central pixel and gradually decaying towards the surrounding pixels, and the sum of the weights is 1, ensuring that the overall brightness of the image remains unchanged; G represents the grayscale image obtained, G smooth It is the image obtained by convolution operation between the original grayscale image and the Gaussian kernel; K x ,K y is the Sobel operator convolution kernel, detecting horizontal and vertical edges respectively; E is the gradient magnitude image, with high-value areas corresponding to tongue contour mutation points. The high-contrast feature of E directly increases the network's sensitivity to tongue coating contours, ultimately resulting in a more accurate tongue coating segmentation image.

[0099] S5. The gradient amplitude image generated by edge enhancement is merged with the original tongue image containing only the tongue body as the input of a new round of tongue coating separation UNet network to obtain the tongue coating segmentation image.

[0100] S6. Quantify the feature importance (i.e., the activation intensity of the heat map) based on the high-contribution feature regions located by the heat map. Dynamic feedback is used to adjust the weights of each layer of the tongue coating separation UNet network. The activation response of tongue coating-related features is iteratively strengthened. The feature map after iteration is converted into a grayscale image and a gradient amplitude image. The gradient amplitude image is merged with the original tongue body-only image as the input of the tongue coating separation UNet network to finally obtain a high-precision tongue coating segmentation image:

[0101] According to the activation intensity of the heat map, the channel weights of the feature map are dynamically adjusted, the weights of high-contribution channels are increased, and the weights of low-contribution channels are reduced.

[0102] Convolution weight adjustment formula:

[0103]

[0104] is the weight matrix of layer t; η is the learning rate; Indicates that when the convolution kernel weight W conv The rate of change of total loss L when a small change occurs.

[0105] The above is a further description of the present invention in combination with the embodiments, and the protection scope of the present invention is not limited thereto.

Claims

1. A dynamic closed-loop tongue coating separation method based on improved UNet is characterized by: The steps are as follows: S1. Collect and preprocess tongue images, input the labeled tongue image data into the UNet network to train the tongue segmentation model, and use a hybrid loss function to improve the convergence model; S2. After the tongue image data is processed through the tongue segmentation UNet model, an image dataset containing only the tongue body is obtained; S3. For images containing only the tongue, a UNet network is further used to perform preliminary moss separation. Gradient-weighted class activation mapping is introduced in the final layer of the UNet network to generate a high-resolution heat map to visualize important regions in the image, quantify feature contributions, and locate key feature regions. S4. Convert the heat map-guided feature map into a grayscale image for grayscale dimensionality reduction, and then use the Sobel operator to enhance the edge of the grayscale image to generate a gradient magnitude image; S5. The gradient magnitude image generated after edge enhancement is combined with the original tongue image as the input of a new round of tongue coating separation UNet network to obtain a tongue coating segmentation image; S6. Quantify the feature importance based on the high-contribution feature areas located by the heat map, dynamically feedback and adjust the weights of each layer of the tongue coating separation UNet network, iteratively strengthen the activation response of tongue coating-related features, and convert the iterative feature map into a grayscale image and a gradient amplitude image. The gradient amplitude image is merged with the original tongue body image as the input of the tongue coating separation UNet network to finally obtain a high-precision tongue coating segmentation image.

2. The dynamic closed-loop tongue coating separation method based on improved UNet according to claim 1 is characterized in that the steps The hybrid loss function described in S1 is a weighted sum of the multi-classification cross entropy loss function and the Dice coefficient loss function: L=ωL CCE +(1-ω)L DL , In the formula, L CCE is the multi-classification cross entropy loss function; L DL is the Dice loss function, ω and (1-ω) are L CCE , L DL The weight of .

3. The dynamic closed-loop tongue coating separation method based on improved UNet according to claim 2 is characterized in that: The formula of the multi-classification cross entropy loss function is: In the formula, m is the number of samples, k is the number of categories, and y ij is the one-hot encoding of sample i in category j, Represents the predicted probability of sample i in category j; The Dice loss function formula is: In the formula, K is the set of all categories, y ij is the one-hot encoding of sample i in category j, represents the predicted probability of sample i in category j, and ε is a small constant.

4. The dynamic closed-loop tongue coating separation method based on improved UNet according to claim 1 is characterized in that: In step S3, the weights are first calculated based on the gradients of each feature channel of the feature map generated by the penultimate layer of the UNet network, and then the feature map is multiplied by the calculated weights. The nonlinear processing is then performed through the ReLU function, and finally a high-resolution heat map is obtained, which is the Grad-CAM visualization result.

5. The dynamic closed-loop tongue coating separation method based on improved UNet according to claim 4 is characterized in that: The specific process in step S3 is as follows: (1) Assume that the penultimate layer of the UNet network generates K feature maps A with width μ and height ν k , k∈K, these feature maps are globally pooled and then linearly combined to produce the confidence score y of the tongue coating category c c ; (2) Tongue coating category c is back-propagated on feature layer A to obtain gradient information Perform global average pooling on the calculated gradient in the dimensions of width i and height j to obtain the weight Where Z is the product of the width i and height j of the feature layer; c is the tongue coating category; y c is the confidence score corresponding to the category obtained by forward propagation; k is the selected channel; A is the feature layer, which is the output of the penultimate convolutional layer here; is the data of feature layer A with coordinates (i, j) in channel k; The gradient information of tongue coating category c is obtained by backpropagation on feature layer A; (3) Multiply the feature map with the calculated weights, and then perform nonlinear processing through the ReLU function. The final heat map is the Grad-CAM visualization result. Partial linearization of the downstream deep network representing feature map A and obtaining the importance of feature map k of tongue coating category c; Grad-CAM then performs a weighted combination of the forward activation maps and obtains the discriminant localization map of tongue coating category c through ReLU Last passed A heatmap of the same size as the convolutional feature map is generated, where darker colors indicate greater importance for moss separation.

6. The dynamic closed-loop tongue coating separation method based on improved UNet according to claim 1 is characterized in that: The grayscale dimensionality reduction formula in step S4 is: In the formula, G represents the grayscale image obtained, k represents the channel, β is the heat map weight adjustment factor, α k is the gradient weight of the kth channel, m is the channel index, G represents the grayscale image obtained, k represents the channel, β is the heat map weight adjustment factor, α k is the gradient weight of the kth channel, and m is the channel index variable.

7. The dynamic closed-loop tongue coating separation method based on improved UNet according to claim 1 is characterized in that: The generation process of the gradient magnitude image in step S4 is: G smooth =G*K Gaussian , K γ =K x T , G x =G smooth *K x , G y =G smooth *K y , In the formula, * is the convolution operation, K Gaussian is a Gaussian smoothing kernel, with a higher weight for the central pixel and gradually decaying towards the surrounding pixels, and the sum of the weights is 1, ensuring that the overall brightness of the image remains unchanged; G represents the grayscale image obtained, G smooth It is the image obtained by convolution operation between the original grayscale image and the Gaussian kernel; K x , K y is the Sobel operator convolution kernel, which detects horizontal and vertical edges respectively; E is the gradient amplitude image, and the high-value area corresponds to the tongue contour mutation point.

8. The dynamic closed-loop tongue coating separation method based on improved UNet according to claim 1 is characterized in that: In step S5, the channel weights of the feature map are dynamically adjusted according to the activation intensity of the heat map, the weights of high-contribution channels are increased, and the weights of low-contribution channels are reduced. Convolution weight adjustment formula: is the weight matrix of layer t; η is the learning rate; Indicates that when the convolution kernel weight W conv The rate of change of total loss L when a small change occurs.

Citation Information

Patent Citations

  • U-net model tongue body automatic segmentation method based on fusion PCA

    CN116071373A