A Method for Identifying the Tenderness of Traditional Chinese Medicine Tongue Coats Based on Image Inpainting and Convolutional Neural Network
Through image repair and convolutional neural network technology, the tongue image is separated and repaired, and the problems of tongue coating interference and inaccurate recognition are solved, achieving more efficient recognition of tongue quality.
Patent Information
- Application Number
- CN202111572065.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-21
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-12-21
AI Technical Summary
The prior art is easily disturbed by tongue coating in the recognition of old and tender tongue quality, and lacks quantitative and objective judgment standards, resulting in poor recognition effect.
Using a method based on image repair and convolutional neural network, tongue coating and tongue mass are separated by tongue semantic segmentation and Gaussian hybrid model, the generated image repair network is used to repair tongue image, and features are extracted and classified through improved residual networks.
It effectively avoids the impact of tongue coating on the recognition of old and tender tongue quality, extracts richer tongue color and texture features, and improves the accuracy of old and tender tongue quality recognition.
Smart Images

Figure CN114372926B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tongue image analysis and diagnosis, and specifically relates to a method for identifying the maturity of traditional Chinese medicine tongue quality based on image restoration and convolutional neural network. Background Technique
[0002] Tongue diagnosis is one of the important diagnostic methods in traditional Chinese medicine and has high clinical value
[35] . Among them, the maturity of the tongue quality is an important judgment index in tongue diagnosis. The texture of the old tongue is rough, firm and old, indicating excess syndrome; the texture of the tender tongue is delicate, floating and delicate, indicating deficiency syndrome. However, in clinical diagnosis, the judgment of the maturity of the tongue quality mainly relies on the doctor's naked eye observation and subjective judgment, lacking quantitative and objective judgment criteria. Therefore, it is very necessary to apply modern computer technology to study the identification method of the maturity characteristics of the tongue quality and realize the objectification of the identification of the maturity characteristics of the tongue quality.
[0003] At present, some scholars have carried out research on the identification of the maturity characteristics of the tongue quality. The identification methods mainly include: using the method of gray difference statistics to describe the texture characteristics of the tongue image, and identifying the maturity characteristics of the tongue quality according to the different change trends of the texture characteristic description parameters in three tongue images of old tongue, medium tongue quality and tender tongue; extracting the color and texture fusion characteristics of the tongue image, and establishing a tongue quality maturity identification model based on the AdaBoost algorithm of the k-nearest neighbor classifier (k-Nearest Neighbor, KNN) for classification
[29] ; constructing a multi-task convolutional neural network to identify multiple characteristics including tongue color, coating color and the maturity of the tongue quality.
[0004] The first method mentioned above only extracts the eigenvalue describing the texture characteristics as the classification basis at first, while ignoring that the maturity characteristics of the tongue image actually include two parts: color characteristics and texture characteristics, resulting in an incomplete description of the maturity characteristics of the tongue image and affecting the classification effect; secondly, the method of delimiting the discrimination threshold for the maturity classification of the tongue quality according to the change trend of the texture characteristics is easily affected by subjective factors; extracting the overall texture characteristics of the tongue image as the discrimination basis for the maturity of the tongue quality is easily interfered by the tongue coating, and only realizes the identification of old tongue and tender tongue, ignoring the situation of medium tongue quality. The second method combines the color characteristics and texture characteristics of the tongue image, and then establishes a tongue quality maturity classification model based on the AdaBoost algorithm of KNN to obtain a better classification effect. However, the samples of old tongue and tender tongue are very few in the experimental data, and extracting the overall color and texture characteristics of the tongue image as the classification basis for the maturity of the tongue quality is easily interfered by the tongue coating.
[0005] With the continuous development of convolutional neural networks, deep learning has been increasingly widely applied in the field of image classification. By simulating the structure of the human brain's nervous system, convolutional neural networks transmit information layer by layer and automatically extract corresponding features. Compared with specific color and texture feature description methods, convolutional neural networks can extract color and texture information in images more fully and completely, achieving better classification results. Although the third method uses a convolutional neural network for tongue image feature recognition, it uses the entire tongue image to identify the two categories of old and tender tongue substances, and is also easily interfered by the tongue coating. Therefore, we provide a method for identifying the old and tender tongue substances in traditional Chinese medicine based on image restoration and convolutional neural networks. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] A method for identifying the old and tender tongue substances in traditional Chinese medicine based on image restoration and convolutional neural networks of the present invention includes the following steps:
[0008] Step 1: Obtain the original tongue image, and use the tongue body semantic segmentation model to segment the tongue image to obtain the tongue body segmentation image;
[0009] Step 2: Use the Gaussian mixture model to separate the tongue coating and the tongue substance of the tongue body segmentation image to obtain the tongue substance image;
[0010] Step 3: Establish a tongue substance image restoration model based on the generative image restoration network, and use the tongue substance image restoration model to restore the tongue substance image to obtain a tongue substance restoration image with continuous texture features and color changes;
[0011] Step 4: Use the improved residual network to extract features and classify the dataset of the restored tongue substance restoration image, and establish a model for identifying the old and tender tongue substances; use the model for identifying the old and tender tongue substances to identify the old and tender tongue substances.
[0012] As a preferred technical solution of the present invention, the method for separating the tongue coating and the tongue substance of the tongue body segmentation image by the Gaussian mixture model is that, assuming there is a d-dimensional random variable x = (x 1 , x 2 ,......, x w ), then the Gaussian mixture model containing K components can be expressed as a formula: T , then the Gaussian mixture model containing K components can be expressed as the formula
[0013]
[0014] where N(x|μ k , ∑ k ) is the Gaussian probability density function, ω k , μ k , ∑ kThey are respectively the weight, mean, and covariance matrix of the k-th component in the Gaussian mixture model.
[0015] As a preferred technical solution of the present invention, the method for the training data required to establish the tongue texture image restoration model is to traverse the tongue texture image in the form of N×N small patches, extract the small patches with the tongue texture proportion exceeding 80%, and use them as the training data required to establish the tongue texture image restoration model; a total of M tongue texture small patch images are obtained, and the data set is expanded to H images through methods such as mirroring, where J images are used as the training set and I images are used as the validation set.
[0016] As a preferred technical solution of the present invention, a new content-aware layer CAL is introduced after the generation of the image restoration network. The new content-aware layer CAL is used to learn to borrow feature information from a certain place in the known area of the image to generate the missing patches. The new content-aware layer CAL calculates the matching score between the foreground patch and the background patch using convolution, then applies the softmax function for comparison and obtains the attention score of each background pixel. Finally, the foreground patch with the background patch is reconstructed by deconvolving the attention score of each background pixel.
[0017] As a preferred technical solution of the present invention, the method for extracting features and classifying the data set of the tongue texture restoration image obtained after restoration using an improved residual network to establish a tongue texture oldness and tenderness recognition model is as follows: First, zero-padding is performed on the input tongue texture restoration image, and convolution is performed on the zero-padded tongue texture restoration image.
[0018] layer, then batch normalization is performed. Subsequently, the regularized tongue texture restoration image is activated using an activation function, and max pooling is performed on the maximum value. Then, after feature operations, the obtained feature vectors are subjected to average pooling, and the multi-dimensional feature vectors are one-dimensionalized to obtain one-dimensional feature vectors. Then, they are input into the fully connected layer to obtain an output. Finally, the probabilities of each category are calculated by the softmax classifier to obtain the final tongue texture oldness and tenderness classification result, and a tongue texture oldness and tenderness recognition model is established based on the tongue texture oldness and tenderness classification result.
[0019] The beneficial effects of the present invention are:
[0020] The method for identifying the oldness and tenderness of traditional Chinese medicine tongue based on image inpainting and convolutional neural network obtains the original tongue image, uses the tongue body semantic segmentation model to segment the tongue image to obtain the tongue body segmentation image, and uses the Gaussian mixture model to separate the tongue coating and tongue substance from the tongue body segmentation image; obtains the tongue substance image, establishes a tongue substance image inpainting model based on the generative image inpainting network, uses the tongue substance image inpainting model to inpaint the tongue substance image, obtains the tongue substance inpainting image with continuous texture features and color changes, uses the improved residual network to extract features and classify the dataset of the tongue substance inpainting image obtained after inpainting, and establishes a model for identifying the oldness and tenderness of the tongue substance; uses the model for identifying the oldness and tenderness of the tongue substance to identify the oldness and tenderness of the tongue substance. The method proposed by the present invention can avoid the influence of the tongue coating on the identification of the oldness and tenderness of the tongue substance compared with the previous methods for identifying the oldness and tenderness of the tongue substance. The features obtained through self-learning can reflect richer tongue substance color and texture features, and achieve a good effect in identifying the oldness and tenderness of the tongue substance. Description of the Drawings
[0021] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0022] Figure 1 is a flowchart of a method for identifying the oldness and tenderness of traditional Chinese medicine tongue based on image inpainting and convolutional neural network of the present invention;
[0023] Figure 2 is a schematic diagram of tongue body separation;
[0024] Figure 3 is a schematic diagram of the Gaussian mixture model;
[0025] Figure 4 is an effect diagram of tongue coating separation;
[0026] Figure 5 is an effect diagram of tongue substance separation;
[0027] Figure 6 is a schematic diagram of the coarse-to-fine network architecture;
[0028] Figure 7 is a schematic diagram of the content-aware layer;
[0029] Figure 8 is a schematic diagram of parallel encoder fusion based on the rough result;
[0030] Figure 9 is a schematic diagram of the tongue substance image structure;
[0031] Figure 10 is a schematic diagram of the tongue substance inpainting image structure;
[0032] Figure 11It is a schematic diagram of the incomplete block structure in the improved residual network;
[0033] Figure 12 It is a schematic diagram of the performance comparison between ResNet50 and ResNet101;
[0034] Figure 13 It is a schematic diagram of the structure of ResNet101;
[0035] Figure 14 It is a schematic diagram of the recognition results of different networks;
[0036] Figure 15 It is a schematic diagram of the recognition results of different networks;
[0037] Figure 16 It is a schematic diagram of the classification results of different preprocessing methods;
[0038] Figure 17 It is a schematic diagram of the classification results of different recognition methods. Detailed implementation manners
[0039] The following is a description of the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0040] Embodiment: As Figure 1 shown, a method for identifying the oldness and tenderness of traditional Chinese medicine tongue quality based on image inpainting and convolutional neural network of the present invention includes the following steps.
[0041] Step 1: Obtain the original tongue image, and use the tongue body semantic segmentation model to segment the tongue image to obtain the tongue body segmentation image; Since the collected tongue image contains the background such as the lips and the surrounding skin in addition to the tongue body, which will interfere with the identification of the oldness and tenderness of the tongue quality, it is necessary to segment the tongue body from the tongue image. In this paper, a tongue body semantic segmentation model is trained based on the DeepLab v3+ semantic segmentation network to realize the automatic segmentation of the tongue body. The tongue body outline segmented by the DeepLab v3+ tongue body semantic segmentation model is clear and accurate, and can better remove the lips, skin and other backgrounds, which is beneficial to the subsequent analysis and recognition of the tongue image features. Among them Figure 2 is the schematic diagram of tongue body separation.
[0042] Step 2: Use the Gaussian mixture model to separate the tongue coating and tongue quality of the tongue body segmentation image; obtain the tongue quality image;
[0043] Step 3: Establish a tongue quality image inpainting model based on the generative image inpainting network, and use the tongue quality image inpainting model to inpaint the tongue quality image to obtain a tongue quality inpainted image with continuous texture features and color changes;
[0044] Step 4: Use an improved residual network to extract features and classify the dataset of the repaired tongue body repair images, and establish a recognition model for the maturity of the tongue body; use the recognition model for the maturity of the tongue body to recognize the maturity of the tongue body.
[0045] The method for separating the tongue coating and the tongue body of the tongue body segmentation image by the Gaussian mixture model is that it is assumed that there is a d-dimensional random variable x = (x 1 , x 2 ,......, x w ) T , then the Gaussian mixture model containing K components can be expressed as a formula, Figure 3 is a schematic diagram of the Gaussian mixture model;
[0046]
[0047] In the formula, N(x|μ k , ∑ k ) is the Gaussian probability density function, ω k , μ k , ∑ k are respectively the weight, mean and covariance matrix of the k-th component in the Gaussian mixture model.
[0048] Among the features of the maturity of the tongue body, an old tongue refers to rough texture of the tongue body, firm and old in shape and color, mostly belonging to excess syndromes; a tender tongue refers to delicate texture of the tongue body, floating, plump and delicate in shape and color, mostly belonging to deficiency syndromes. Thus, it can be seen that the maturity of the tongue body is closely related to the texture and shape / color of the tongue body, and the presence of the tongue coating will interfere with the recognition of the maturity of the tongue body. Therefore, this paper uses the GMM algorithm to separate the tongue coating and the tongue body, and obtains the tongue body and tongue coating images respectively, preparing for the subsequent repair of the tongue body image. GMM refers to the linear combination of multiple Gaussian distribution functions (normal distribution), and it is also the probability model with the fastest learning speed. GMM tries to find the mixture of multi-dimensional Gaussian distributions that can best simulate the input dataset.
[0049] Practice has proved that many datasets conform to the Gaussian distribution. Even if the original dataset does not conform to the Gaussian distribution, but with the increase of samples, according to the central limit theorem, it will approach the Gaussian distribution. And theoretically, by increasing the number of models, GMM can approach any continuous probability density distribution, which makes it applicable to more flexible cluster shapes. During the classification process of GMM clustering, the output after training is not a specific value like K-Means clustering, but a set of probability values (probabilities belonging to different classes). Judging which class the data specifically belongs to by the size of the probability values belonging to different classes is more accurate than directly assigning the data to a certain class. Especially when there is an overlap between different classes of clustering, it is extremely easy to produce confusion when directly determining the class to which the data belongs. Therefore, the accuracy of GMM clustering is higher than that of K-Means clustering. Figure 4 is the effect diagram of tongue coating separation;Figure 5 This is the effect diagram of tongue separation;
[0050] In order to eliminate the influence of tongue coating on the identification of tongue texture and tenderness, and maintain the continuity of tongue texture and color changes, this paper uses the GIICA network to repair tongue images to avoid the interference of discontinuous texture and color changes on the identification of tongue texture and tenderness. Since tongue texture and tenderness are closely related to tongue color and texture features, traditional image repair algorithms have poor repair effects on large-scale damaged areas and cannot restore detailed textures well. At present, deep learning-based image repair methods have achieved good results in challenging tasks such as repairing a large number of missing areas in images. Therefore, this paper uses deep learning methods to repair tongue images. Although general deep learning-based image repair methods can generate visually reasonable image structures and textures, since convolutional neural networks cannot clearly copy and borrow textures at a certain position in distant space, the repaired area is prone to produce blurred textures or distorted structures that are inconsistent with the surrounding areas. To address this problem, Yu et al.
[47] proposed the GIICA network. While generating new image structures, the GIICA network can also refer to the features of the background image during network training to obtain better prediction results. Therefore, the GIICA network has a good effect on the repair of large-scale defective areas of images, and the repair model established by the network can be used to repair any shape missing area of images of any resolution. Therefore, this paper uses the GIICA network to establish a tongue quality image repair model, repair the tongue quality image, and obtain a tongue quality image with the same shape and size as the tongue body image. Since the GIICA network training requires images of all tongue quality as a data set to establish a tongue quality image repair model, and the tongue body image basically contains two parts: tongue coating and tongue quality, it cannot be used directly as training data. Therefore, this paper separates the coating from the old and tender tongue images, traverses the tongue quality images in 96×96 small blocks, extracts small blocks with a tongue quality ratio of more than 80%, and uses them as the training data required to establish the tongue quality image repair model. A total of 222 tongue quality small block images are obtained. The data set is expanded to 444 by mirroring and other methods, of which 396 are used as training sets and 48 are used as validation sets.
[0051] The method of establishing a tongue quality image restoration model based on a generative image restoration network and using the tongue quality image restoration model to restore the tongue quality image is to first establish the training data required for the tongue quality image restoration model, and then build the basic generative image restoration network by copying and improving the image restoration algorithm based on global and local content consistency, and then introduce a coarse-to-fine network architecture in which the first network makes a rough prediction of the missing area, and the second network uses the rough prediction result as input and makes a fine prediction, and finally completes the restoration of the missing area of the image. This structure is also called a coarse-to-fine network architecture, as shown in the following figure. Figure 6 shown.
[0052] In addition, since convolutional neural networks process image features with local convolutional kernels layer by layer, it is difficult to obtain features from distant spatial positions. Therefore, in the next part of the GIICA network, to overcome this limitation, a perception mechanism is considered and a new Contextual Attention Layer (CAL) is introduced into the deep generative network. CAL is used to learn to borrow feature information from somewhere in the known region of the image to generate missing patches. Since CAL is differentiable, it can be trained in a fully convolutional network and a deep model and can be tested on images of any resolution. The working process of CAL Figure 7 is shown as follows.
[0053] To integrate the improved generative image inpainting network with the content-aware module, the GIICA network introduces two parallel encoders based on the rough prediction results of the first encoder-decoder. Figure 6 The bottom encoder in is specifically used for the imagined content of layer-by-layer (diffusion) convolution, while the top encoder attempts to participate in the content-aware function. Finally, the output features from the two encoders are aggregated and input into a single decoder to obtain the final inpainting result, as Figure 8 shown.
[0054] Figure 8 visualizes the Attention Map in . The colors in the figure represent the relative positions of the most interesting background color patches corresponding to each pixel in the foreground. The positions of the colors in the Attention Map Color Coding represent the relative positions of the most interesting background color patches corresponding to the pixels of the colors in the Attention Map. For example, white indicates that the most interesting area is on itself, pink indicates that the most interesting background color patch for the corresponding foreground area is in the lower left, and green indicates that after the tongue image is inpainted, the original tongue coating area is basically repaired to the tongue, and the detailed texture of the inpainted area is better restored, which is beneficial to excluding the interference of the tongue coating area on the recognition of the oldness and tenderness of the tongue and the subsequent recognition of the oldness and tenderness of the tongue. The inpainting effect diagrams are as Figure 9 and Figure 10 shown.
[0055] Establishment of the oldness and tenderness recognition model.
[0056] Among them, for the recognition of the oldness and tenderness of the tongue body based on the residual network, an improved residual network is used to extract features and classify the dataset of the repaired tongue body repair images. The method for establishing the recognition model of the oldness and tenderness of the tongue body is as follows: First, zero-padding is performed on the input tongue body repair images. Then, a convolutional layer is applied to the zero-padded tongue body repair images, followed by batch normalization. Subsequently, an activation function is used to activate the normalized tongue body repair images, and max pooling is performed on the maximum values. After feature operations, the obtained feature vectors are subjected to average pooling, and then the multi-dimensional feature vectors are one-dimensionalized to obtain one-dimensional feature vectors. Then, these are input into the fully connected layer to obtain an output. Finally, the softmax classifier calculates the probabilities of each category to obtain the final classification result of the oldness and tenderness of the tongue body, and the recognition model of the oldness and tenderness of the tongue body is established based on the classification result of the oldness and tenderness of the tongue body.
[0057] Convolutional Neural Networks (CNNs)
[49] are a type of deep neural network that includes convolutional calculations and are widely used in image recognition. Convolutional neural networks usually include convolutional layers, pooling layers, and fully connected layers.
[0058] Convolutional layer
[0059] The convolutional layer is mainly composed of convolution kernels with learnable parameters. The convolutional layer performs convolution on the input data through local connection and weight sharing between convolution kernels to extract data features. The first convolutional layer will extract some primary features of the input data. For example, when the input data is an image, the first convolutional layer will first extract its primary information such as edges, contours, and lines. These primary information will form multiple feature maps, which will then be used as the input for the next convolutional layer to extract deeper feature information. The calculation formula of the convolutional layer is
[0060]
[0061] In the formula: "*" represents the convolution operation, represents the operation result of the j-th convolution kernel in the k-th layer and is also the input of the (k + 1)-th layer; represents the j-th convolution kernel in the k-th layer; represents the bias value; f represents the activation function.
[0062] Pooling layer
[0063] After the convolution operation, the dimension of the feature vector increases. If directly trained, the required network computation amount and complexity are too large. Therefore, it is necessary to reduce the dimension of the feature maps obtained by the convolution operation. The pooling layer can not only reduce the feature dimension through local pooling of the feature maps, but also expand the receptive field, achieve non-linearity and invariance (translation invariance, rotation invariance, and scale invariance), which is helpful for reducing the overfitting problem of the network and increasing the robustness of the network.
[0064] Fully Connected Layers
[0065] Fully Connected Layers (FC) are generally in the last few layers of a deep learning network and act as a "classifier" in the entire convolutional neural network. Each node in the fully connected layer is connected to all nodes in the previous layer, connecting and learning all the features extracted by the previous network layers, and completing the learning of the specified classification target by perceiving global information, and finally obtaining the classification result.
[0066] The Residual Network (ResNet) not only draws on the advantages of traditional deep learning networks but also introduces the residual learning method, solving problems such as loss and dissipation of information during transmission; enabling the entire network to only learn the differences between the input and output, simplifying the learning objectives and difficulties of the network; effectively solving the problems of gradient dissipation and gradient explosion in deep networks, enabling the network to be deepened as much as possible.
[0067] The improved ResNet of the present invention uses the ResNet residual network, in which ResNet proposes the identity mapping and the residual mapping on the basis of the traditional convolutional neural network. Figure 11 is a typical Residual Block. The identity mapping is the process shown by the curve (Short Connection, SC) in the figure, which is a mapping that can directly send X to the ReLU layer after skipping 2 weight layers (the number of layers is uncertain and can be 3 layers or 4 layers). Since X directly skips the weight layer without any operation, it is called the identity mapping, that is, G(X)=X. The straight-line process represents the residual mapping, which is the difference between the input and the output, that is, F(X)=H(X)-X.
[0068] Through the SC process, the input and output of the residual block can perform element-wise operations. This simple operation does not increase the parameters and computational complexity of the network, but can greatly improve the training speed of the model and enhance the training effect. And due to the identity map, the gradient can directly return to the shallower layer during backpropagation, effectively solving the problem of model degradation when the network is deepened. Applying the residual block multiple times on the basis of the traditional convolutional neural network constitutes ResNet, making it possible to train extremely deep deep learning neural networks. The most common deep residual networks are mainly ResNet50 and ResNet101. The performance comparison of ResNet50 and ResNet101 on the ImageNet validation dataset is as Figure 12As shown in the figure, the ImageNet Image Classification Competition uses the Top-1 error rate or the Top-5 error rate as the evaluation criteria for model performance. Top-1 = (the number of samples with different correct labels and the best label output by the model) / total number of samples, and Top-5 = (the number of samples with correct labels not among the top 5 best labels output by the model) / total number of samples. Figure 12 As shown in the figure, it shows that on the ImageNet validation dataset, the Top-1 error rate and Top-5 error rate of ResNet101 are respectively 0.87% and 0.65% lower than those of ResNet50. Thus, it can be seen that ResNet101 has a better image classification effect on the ImageNet validation dataset compared to ResNet50. Therefore, in this chapter, ResNet101 is selected to establish the recognition model for the texture of the tongue being old or tender.
[0069] Recognition model for the texture of the tongue being old or tender.
[0070] ResNet101 is a residual network with a 101-layer network structure, including 100 convolutional layers and 1 fully connected layer. Its specific structure is as Figure 13 shown
[0071] ZERO PAD means padding zeros to the input repaired image of the tongue texture. CONV in stage1 represents a convolutional layer, Batch Norm represents batch normalization processing, ReLU represents an activation function, and MAX POOL represents max pooling; in stage2 - 5, CONV BLOCK represents a residual block that changes the scale of the feature vector, ID BLOCK×2 represents 2 residual blocks that do not change the feature scale, and ID BLOCK×3 represents 3 residual blocks that do not change the feature scale (in ID BLOCK×N, ID BLOCK represents a residual block that does not change the feature scale, and N represents the number of such residual blocks); and each residual block includes 3 convolutional layers. As Figure 13 can be seen, there are a total of 1 + 3×(3 + 4 + 23 + 3) = 100 convolutional layers. After the feature operations in stage1 - 5, the obtained feature vectors are subjected to AVG POOL (average pooling), and then the multi-dimensional feature vectors are flattened into one-dimensional vectors to obtain one-dimensional feature vectors. Then, these vectors are input into the fully connected layer (FC) to obtain the output, and finally, the softmax classifier calculates the probabilities of each category to obtain the final classification result for the texture of the tongue being old or tender.
[0072] The method of using an improved residual network to extract features and classify the dataset of the repaired tongue texture image to establish a recognition model for the oldness and tenderness of the tongue texture is as follows: First, zero-padding is performed on the input tongue texture repair image, and then a convolutional layer is applied to the zero-padded tongue texture repair image, followed by batch normalization. Subsequently, an activation function is used to activate the normalized tongue texture repair image, and max pooling is performed on the maximum value. After feature calculation, the obtained feature vectors are subjected to average pooling, and then the multi-dimensional feature vectors are one-dimensionalized to obtain one-dimensional feature vectors, which are then input into the fully connected layer to obtain an output. Finally, the softmax classifier calculates the probabilities of each category to obtain the final classification result of the oldness and tenderness of the tongue texture, and a recognition model for the oldness and tenderness of the tongue texture is established based on the classification result of the oldness and tenderness of the tongue texture.
[0073] In this paper, the classification accuracy (Accuracy, Ac) of the test set for identifying the oldness and tenderness features of the tongue texture is used as the experimental evaluation index, and its mathematical expression is:
[0074] R represents the number of correctly classified samples in the test set, and S represents the total number of samples in the test set.
[0075] In this paper, the ResNet101 residual network is implemented through the keras framework, and the parameters are set as follows: the number of training epochs is 1000, the initial learning rate (learning rate, lr) is 0.001, the batch size is 10, the training step for reducing the learning rate is 5 epochs, that is, when the val_loss does not decrease for 5 consecutive epochs, the learning rate is reduced, and the learning rate decay factor is 0.1, that is, the learning rate decays at a rate of 10 times. The step for early stopping of training is 15 epochs, that is, when the val_loss does not decrease for 15 consecutive epochs of training, the model training is terminated early. The software configuration of this experiment is Anaconda3.5 and Python3.6, and a keras deep learning environment is built on this basis; in the hardware configuration, the CPU is Intel Core i5-10300H, the memory is 16GB, and the GPU graphics card is NVIDIA GTX 1660Ti. To verify the effectiveness of the ResNet101 residual network selected in this paper in identifying the oldness and tenderness of the tongue texture, ResNet50 and ResNet101 recognition models for the oldness and tenderness of the tongue texture are established using the training set of the tongue texture repair image, and the recognition effect of the model is tested using the test set. The test results are as Figure 14 shown:
[0076] From Figure 14It can be seen that compared with the model established using ResNet50, the recognition accuracy of old tongue, moderate tongue quality, and tender tongue by the tongue texture oldness and tenderness recognition model established using ResNet101 has increased by 1.0%, 2.0%, and 2.0% respectively, and the overall recognition accuracy has increased by 1.7%. Thus, it can be known that in terms of tongue texture oldness and tenderness recognition, the model established using ResNet101 has a better classification effect. One old tongue image, one moderate tongue quality image, and one tender tongue image randomly selected were used as the inputs of the ResNet101 tongue texture oldness and tenderness recognition model to obtain the heat maps of some intermediate layers to analyze the effect of the convolutional layer's self-learning of the features of the three types of tongue texture images. The restored images of the three types of tongue textures, namely old tongue, moderate tongue quality, and tender tongue, basically conform to the traditional Chinese medicine's description of the color and texture of tongue texture oldness and tenderness. The texture of the restored old tongue image is relatively rough and the color is relatively dull. The texture of the restored tender tongue image is relatively delicate and the shape is slightly plump. Most of the heat maps of the old tongue images are orange-red, most of the heat maps of the moderate tongue quality images are light blue, and most of the heat maps of the tender tongue images are yellow-green. Through the heat maps of the three, it can be reflected that there are obvious differences in the characteristics of old tongue, moderate tongue quality, and tender tongue, indicating that ResNet101 can better learn the characteristics of the three types of tongue images and establish a tongue texture oldness and tenderness recognition model with a better classification effect. Using the GIICA tongue image restoration model to obtain the restored tongue texture images with continuous texture and color changes is beneficial to avoiding the interference of tongue coating on the recognition of tongue texture oldness and tenderness. The tongue body image and the restored tongue texture image were respectively input into the ResNe101 network to establish a tongue texture oldness and tenderness recognition model. The test results are shown in Figure 15 For the ResNet101 tongue texture oldness and tenderness recognition model established with the dataset preprocessed by algorithms such as GIICA restoration, the recognition accuracies of old tongue, moderate tongue quality, and tender tongue are 83.0%, 98.0%, and 92.0% respectively, which are 5.0%, 6.0%, and 4.0% higher than the recognition accuracies of the model established using the original tongue body image training set. The overall recognition accuracy is 91.0%, an increase of 5.0%. The experiment shows that repairing the tongue texture image based on the GIICA network to obtain the restored tongue texture image with continuous texture and color changes can effectively eliminate the interference of tongue coating on the recognition of tongue texture oldness and tenderness. Combining with the ResNet101 network to establish a tongue texture oldness and tenderness recognition model has achieved a better recognition effect of tongue texture oldness and tenderness compared with the model established using the original tongue image as the training set. To further verify the effect of the proposed tongue texture oldness and tenderness recognition method in this paper, the recognition effects of this method are compared with those of the gray difference statistics method and the AdaBoost algorithm based on KNN. The classification results are shown in Figure 16 as follows.
[0077] Comparing the recognition effect of the AdaBoost algorithm based on KNN, the classification results are shown in Figure 17Compared with the gray difference statistics method, the classification method of this paper has improved the recognition accuracy of old tongue, moderate tongue quality, and tender tongue by 23.0%, 26.0%, and 33.0% respectively, and the overall recognition accuracy has increased by 27.3%. Compared with KNN+AdaBoost, the recognition accuracy of old tongue, moderate tongue quality, and tender tongue has increased by 14.0%, 22.0%, and 25.0% respectively, and the overall recognition accuracy has increased by 20.0%. It can be seen that this paper uses the GIICA tongue quality image restoration model to obtain a tongue quality image with continuous color and texture changes, and then uses the ResNet101 network to self-learn the features of the restored tongue quality image. The obtained features more comprehensively reflect the color and texture features of different old and tender tongue qualities compared with the manually extracted old and tender tongue quality features. Finally, the established old and tender tongue quality recognition model has achieved a satisfactory recognition effect for old and tender tongue quality.
[0078] In summary, based on the GMM clustering algorithm, this paper separates the coating quality from the old and tender tongue quality dataset, then repairs the tongue quality image through the GIICA network, and then uses the repaired tongue quality image to establish a ResNet101 old and tender tongue quality recognition model, which has a better classification effect than the model established using the tongue body image. Compared with existing methods such as the gray difference statistics method and KNN+AdaBoost, the GIICA+ResNet101 method has achieved a higher classification accuracy, indicating that the method proposed in this paper has achieved a satisfactory recognition effect in the recognition of old and tender tongue quality.
[0079] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for identifying the oldness and tenderness of traditional Chinese medicine tongue quality based on image inpainting and convolutional neural network, characterized in that: It includes the following steps, Step 1: Obtain the original tongue image, and use the tongue body semantic segmentation model to segment the tongue image to obtain the tongue body segmentation image; Step 2: Use the Gaussian mixture model to separate the tongue coating and tongue quality from the tongue body segmentation image; obtain the tongue quality image; Step 3: Establish a tongue quality image inpainting model based on the generative image inpainting network, and use the tongue quality image inpainting model to inpaint the tongue quality image to obtain a tongue quality inpainted image with continuous texture features and color changes; Step 4: Use the improved residual network to extract features and classify the dataset of the tongue quality inpainted image obtained after inpainting, and establish a tongue quality oldness and tenderness recognition model; use the tongue quality oldness and tenderness recognition model to identify the oldness and tenderness of the tongue quality; The method of establishing a tongue quality image inpainting model based on the generative image inpainting network and using the tongue quality image inpainting model to inpaint the tongue quality image is as follows: First, establish the training data required for the tongue quality image inpainting model, and construct its basic generative image inpainting network by copying and improving the image inpainting algorithm based on global and local content consistency. Then, introduce a network architecture from rough to fine, where the first network makes a rough prediction of the missing area, and the second network takes the rough prediction result as input and makes a fine prediction, and finally completes the repair of the image missing area.
2. A method for identifying the oldness and tenderness of traditional Chinese medicine tongue quality based on image inpainting and convolutional neural network according to claim 1, characterized in that, The method for separating the tongue coating and the tongue body in the tongue body segmentation image by the Gaussian mixture model is as follows: Assume that there is a d-dimensional random variable x = (x 1 , x 2 , ……, x w ). T , then the Gaussian mixture model containing K components can be expressed as a formula. where \(N(x|\mu\ k ,\Sigma k )\) is the Gaussian probability density function, \(\omega k \), \(\mu k \), \(\Sigma k are the weight, mean, and covariance matrix of the \(k\)-th component in the Gaussian mixture model, respectively.
3. A method for identifying the oldness and tenderness of traditional Chinese medicine tongue quality based on image inpainting and convolutional neural network according to claim 2, characterized in that, The method of establishing the training data required for the tongue quality image inpainting model is to traverse the tongue quality image in the form of N×N small blocks, extract the small blocks with the tongue quality proportion exceeding 80%, and use them as the training data required for establishing the tongue quality image inpainting model; A total of M tongue quality small block images are obtained, and the dataset is amplified to H images by means of mirroring, etc., where J images are used as the training set and I images are used as the validation set.
4. A method for identifying the oldness and tenderness of traditional Chinese medicine tongue quality based on image inpainting and convolutional neural network according to claim 2, characterized in that, A new content-aware layer CAL is introduced after the generative image inpainting network, where the new content-aware layer CAL is used to learn to borrow feature information from a certain place in the known area of the image to generate the missing patches, and the new content-aware layer CAL uses convolutional calculation to obtain the matching score between the foreground patch and the background patch, then applies the softmax function for comparison and obtains the attention score of each background pixel, and finally deconvolves the attention score of each background pixel to reconstruct the foreground patch with the background patch.
5. A method for identifying the oldness and tenderness of traditional Chinese medicine tongue quality based on image inpainting and convolutional neural network according to claim 1, characterized in that, The method of using an improved residual network to extract features and classify the dataset of the repaired tongue texture repair image to establish a tongue texture oldness and tenderness recognition model is as follows: First, zero-padding is performed on the input tongue texture repair image, and then a convolutional layer is applied to the zero-padded tongue texture repair image, followed by batch normalization processing. Subsequently, an activation function is used to activate the normalized tongue texture repair image, and max pooling is performed on the maximum value. After feature operations, the obtained feature vectors are subjected to average pooling, and then the multi-dimensional feature vectors are one-dimensionalized to obtain one-dimensional feature vectors, which are then input into a fully connected layer to obtain an output. Finally, the softmax classifier calculates the probabilities of each category to obtain the final tongue texture oldness and tenderness classification result, and a tongue texture oldness and tenderness recognition model is established based on the tongue texture oldness and tenderness classification result.
Citation Information
Patent Citations
Method for automatic tongue coating segmentation based on deep learning
CN107610087A
Tongue image recognition method and device, computer equipment and computer readable storage medium
CN110363072A