A method for body constitution recognition based on tongue image features based on MLP-like network

By combining high-wide channel interaction MLP network and graph convolutional label relationship learning, the problem of insufficient correlation between tongue image features and physique labels in traditional Chinese medicine physique recognition is solved, and more efficient physique recognition of tongue image features is achieved, improving the accuracy and interpretability of the model.

CN117237784BActive Publication Date: 2025-08-15SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311192091.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-15
Publication Date
2025-08-15
Estimated Expiration
2043-09-15

AI Technical Summary

Technical Problem

In the prior art, traditional Chinese medicine physique recognition methods lack consideration of the correlation between tongue image features and physique labels, resulting in insufficient generalization ability of the model, and the application of MLP-like models in tongue image features physique recognition has not been fully explored, especially multi-label physique recognition with the theory of physique is difficult to effectively carry out.

Method used

High-width channel interaction MLP network is used to combine graph convolution label relationship learning, and the correlation between physique labels is learned through the bilayer graph convolution network, the correlation between tongue image features and physique labels is fused, and the model is optimized using binary cross entropy loss function to improve the accuracy of physique recognition.

Benefits of technology

It improves the accuracy of physical constitution recognition of tongue image features and the generalization ability of model, can better reflect the correlation between tongue images and physical constitution labels, and enhances the interpretability of physical constitution recognition and the effect of multi-label recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117237784B_ABST
    Figure CN117237784B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying constitutions based on tongue image features based on an MLP-like network, comprising the following steps: S1, collecting tongue images, and marking constitution labels corresponding to the tongue images to produce a training data set and a test data set. S2, constructing a high-width channel interaction-like MLP network model, which uses binary cross entropy as a loss function and uses constitution labels to perform supervised learning on tongue images. S3, using graph convolutional network branches to learn the relationship between constitution labels corresponding to tongue image features, and obtaining the correlation between combined constitution labels. The benefit of the present invention is that it uses a graph convolutional network to learn the correlation between combined constitutions corresponding to tongue image features, and constructs a high-width channel interaction-like MLP model to extract the constitution feature information of the tongue image, and the correlation performance of the extracted tongue image features and constitution labels is integrated to effectively improve the accuracy of combined constitution identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for identifying physical constitution based on tongue image characteristics based on an MLP-like network, and in particular to a method for identifying physical constitution based on tongue image characteristics based on graph convolution label relationship learning and a high-width channel interaction MLP-like network. Background Art

[0002] In the long tradition of Traditional Chinese Medicine (TCM), the Yellow Emperor's Classic of Internal Medicine is the earliest major work in medical history to discuss the phenomenon of constitution, and it has had a profound influence. In modern "Basic Theory of Traditional Chinese Medicine," constitution refers to the functions of the human body, manifested by both innate endowments and acquired qualities. It highlights the inherent characteristics of relatively stable psychological functions and morphological structures. Currently, the most widely used and recognized constitution classification is the nine-category classification proposed by Academician Wang Qi: balanced constitution, qi deficiency constitution, yang deficiency constitution, yin deficiency constitution, phlegm-damp constitution, damp-heat constitution, blood stasis constitution, qi stagnation constitution, and special constitution. In TCM, constitution identification plays a crucial role in disease prevention and conditioning. Furthermore, the concept of "pre-illness" is also a unique concept in TCM, and constitution identification also plays a key role. Furthermore, according to TCM's theory of combined constitutions, each person has at least one, and sometimes multiple, constitutions.

[0003] Traditional Chinese Medicine (TCM) constitution identification requires doctors to examine the tongue, the tongue, and a combination of both. Each person's constitution is unique, but not necessarily unique; a person may have multiple constitutions. Therefore, compared to multi-classification constitution identification, multi-label constitution identification is more rational and comprehensive, facilitating a comprehensive understanding of a patient's constitution and enabling targeted treatment or treatment. Artificial intelligence has developed rapidly in recent years, particularly with deep neural network machine learning technologies based on big data, which have seen significant progress and widespread application. Deep neural networks can be used to extract deep features from highly complex images. However, current machine learning methods, particularly deep supervised learning methods, require large amounts of labeled data for training. Otherwise, model training can easily suffer from overfitting, resulting in insufficient generalization and an inability to effectively represent the discriminative features of image data. Furthermore, collecting labeled data is difficult and expensive, limiting the number of training sets for each task and exposing them to the typical long-tail distribution problem. Furthermore, some labels are inherently small in number, making it unreasonable to improve model generalization solely through data augmentation.

[0004] In particular, the amount of labeled data required by machine learning methods is positively correlated with the complexity of the learning algorithm. That is, the more complex the algorithm, the more training data required, which dramatically increases the difficulty of implementation. Currently, the most representative deep neural network methods, such as those based on convolutional neural networks (CNNs) and transformers, are complex and require large amounts of training data, making them difficult to meet for our task. Fortunately, a simpler class of multi-layer perceptron (MLP)-like models and their variants, collectively referred to as MLP-like models, have been developed. These models offer similar performance to the first class, but with simpler methods and stronger generalization capabilities. These methods are supervised learning methods and require an objective function (loss function) to optimize the trained model. Furthermore, the mapping of sample labels across different spaces significantly influences recognition results, and their consistency and complementarity are not considered. Furthermore, the correlation between labels corresponding to input images is not considered, which prevents the trained network from learning the information characteristics of human cognition and thus hinders model interpretability.

[0005] Current body type recognition methods fail to integrate the correlation between tongue images and corresponding body types, and no MLP-like models have been applied to the task of tongue image body type recognition, particularly multi-label body type recognition based on the theory of combined body types. Furthermore, according to the theory of combined body types, body types can coexist with both mutual exclusion and high correlation. Therefore, a tongue image feature-based body type recognition method using graph convolutional label relationship learning and a high-width channel interactive MLP-like network is urgently needed by those skilled in the art. Summary of the Invention

[0006] The purpose of this invention is to provide a method for identifying body constitutions based on tongue image features using an MLP-like network. This method fully considers the correlation between body constitution labels corresponding to tongue image features, integrates graph convolutional label relationship features with a high-width channel interactive MLP-like network to extract tongue image features, and improves the accuracy of body constitution identification.

[0007] The technical solution of the present invention is a method for identifying body constitution based on tongue image features based on a MLP-like network, comprising the following steps:

[0008] Step S1: Data collection and processing: Collect tongue images, label the corresponding tongue images with physical labels, and create training and test data sets;

[0009] Step S2 constructs a high-width channel interaction MLP model: a high-width channel interaction MLP model is proposed that integrates graph convolution label relationship learning and uses body shape labels to supervise the learning of tongue images;

[0010] Step S3: Graph convolutional network model with embedded label relationships: Based on the TCM combined constitution theory, the graph convolutional network is used to learn the correlation encoding between constitution labels to avoid the mutual exclusivity of constitution labels, and then learn the potential connection between tongue image features and the correlation between constitution labels, and integrate the encoding of tongue image feature extraction of the high-width channel interactive MLP model.

[0011] In the aforementioned method for identifying constitution based on tongue image features of a MLP-like network, the constitution labels include balanced constitution, qi deficiency constitution, yang deficiency constitution, yin deficiency constitution, phlegm-damp constitution, damp-heat constitution, blood stasis constitution, qi stagnation constitution, and special constitution.

[0012] In the aforementioned method for identifying constitutions based on tongue image features using an MLP-like network, the tongue image is first scaled to a uniform standard size before being input into the model. A smaller image is then randomly cropped from the scaled image, and the cropped image is randomly horizontally flipped for data augmentation. The enhanced tongue image is then input into a high-width channel interactive MLP model, and the correlation between the tongue image and the corresponding constitution labels learned by graph convolution is integrated to automatically learn tongue image features to identify constitution categories.

[0013] In the aforementioned method for identifying body constitution based on tongue image features based on a quasi-MLP network, the step S2 of constructing a high-width channel interaction quasi-MLP model includes the following steps:

[0014] S21: The normalized batch tongue images are used as input images for the high-width channel interaction MLP model;

[0015] S22: Using a high-width channel interaction MLP model to learn and extract features from the input image, the high-width channel interaction MLP model includes several convolutional layers, normalization layers, downsampling layers, learnable high-width feature interaction layers, and fully connected layers;

[0016] S23: Using a two-layer graph convolutional network to learn the correlation between the physical labels corresponding to the input tongue image, the mutual exclusivity and high-intensity parallelism matrix of the physical labels are integrated into the extracted tongue image features;

[0017] S24: Binary cross entropy is used as the loss function as the objective function for tongue constitution label recognition of the high-width channel interaction MLP model. The tongue image extracted by the high-width channel interaction MLP model and the corresponding label relationship features are input into the Sigmoid classifier to obtain the constitution category:

[0018] q * =Sigmoid(FC(Z))

[0019] Among them, q *Represents the probability vector of the output physical multi-label category; Sigmoid represents the probability value between the output label category (0,1); FC represents the fully connected layer; Z represents the output feature related to the physical label category corresponding to the input tongue image.

[0020] In the aforementioned tongue image feature constitution recognition method based on a quasi-MLP network, the feature learning and extraction process of the high-width channel interactive quasi-MLP model in S22 can be expressed as:

[0021] Assume that the input is X=[x1,x2,…,x n ], the corresponding Token-FC feature can be expressed as:

[0022]

[0023] Where, Represents the spatial features of each feature interaction transformation layer, where Will Added to the weight to get the output interaction feature φ j , then the aggregated features H and W interact to output φ j for:

[0024]

[0025] Where W t Represents the learnable weight; k represents the k-th input Token; j represents the j-th output Token; s k Represents the feature vector after Token interaction transformation;

[0026] In summary, the mathematical expression for feature extraction based on the MLP-like model of H and W channel information interaction is:

[0027] Y=HWmixer FC (BN(X+X

[0028] Z=CH FC (BN(Y+Y

[0029] Where X represents the input features; BN represents batch normalization, in which Gaussian error GELU is used as the activation function; Y represents the intermediate features learned by the H and W channel interaction MLP module; Z represents the output features related to the physical label category corresponding to the tongue image.

[0030] In the aforementioned method for identifying constitution based on tongue image features based on a MLP-like network, the step S3 embeds a graph convolutional network model with label relationships, comprising the following steps:

[0031] S31: Input the body constitution labels corresponding to the tongue image into the graph convolutional network branch. The correlation between body constitution labels is learned through a two-layer graph convolutional network. The matrix is converted into a body constitution label correlation matrix through Sigmoid and then integrated into the high-width channel interaction MLP model to extract tongue image features.

[0032] S32: The training uses the loss function as the objective function to guide the multiple iterative training of the high-wide channel interaction MLP model that integrates graph convolution label relationship learning, and saves and outputs the optimal model parameters for physical identification testing.

[0033] In the aforementioned method for identifying body constitution based on tongue image features based on a MLP-like network, in step S32, the loss function is constructed as follows: the binary cross entropy classification loss of the multi-label category of tongue image features body constitution is calculated, assuming that the body constitution label Y = {y1, y2, ..., y C},y i ∈{0,1}, then the binary cross entropy loss function of the physical label is expressed as:

[0034]

[0035] Where C represents the number of body type labels corresponding to each tongue image, i.e., 9 types of body type labels; i Indicates the real physical label; q i Represents the predicted constitution label.

[0036] Beneficial effects of the present invention: Compared with the prior art, the method of the present invention has the following advantages:

[0037] (1) This invention innovatively uses the HWmixer-MLP model (a high-width channel interaction MLP model) to extract deep feature information from tongue images based on multi-label physical constitution recognition based on tongue images. The network learns deep features of tongue images through H and W interactive FC operations, which is conducive to improving the performance of physical constitution recognition. In addition, a two-layer graph convolutional network model branch embedded in label relationship learning learns the correlation encoding between labels through the graph convolutional network to avoid the mutual exclusivity of labels and the potential connection between tongue image features and label correlations. It also incorporates the tongue image features extracted by the HWmixer-MLP model to enhance the learning ability of the physical constitution recognition network.

[0038] (2) The present invention uses the binary cross entropy loss function as the objective function of the HWmixer-MLP model, utilizes a two-layer graph convolutional network to learn the correlation between the physical labels corresponding to the input tongue image, and integrates the mutual exclusivity and high-intensity parallelism matrix of the physical labels into the extracted tongue image features. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1This is a training flowchart for the tongue image feature constitution recognition method based on graph convolutional label relationship learning and high-width channel interaction MLP network;

[0040] Figure 2 This is a schematic diagram of the high-width channel interaction MLP network structure;

[0041] Figure 3 Schematic diagram of the high-width channel interaction module. DETAILED DESCRIPTION

[0042] The present invention will be further described below with reference to the accompanying drawings and embodiments, but they are not intended to limit the present invention.

[0043] Embodiment of the present invention: A tongue image feature constitution recognition method based on a MLP-like network, such as Figure 1 As shown, the following steps are included:

[0044] Step S1: Data acquisition and processing: Collect tongue images and label the corresponding constitution labels. Constitution labels include balanced constitution, qi deficiency constitution, yang deficiency constitution, yin deficiency constitution, phlegm-damp constitution, damp-heat constitution, blood stasis constitution, qi stagnation constitution, and special constitution. Create training and test datasets.

[0045] Step S2 constructs a high-width channel interaction MLP model: a high-width channel interaction MLP model is proposed that integrates graph convolution label relationship learning and uses body shape labels to supervise the learning of tongue images;

[0046] Step S3 embeds a two-layer graph convolutional network model for label relationship learning: According to the theory of combined constitution in Traditional Chinese Medicine, constitution labels are mutually exclusive. The neutral constitution label generally does not appear at the same time as other constitution labels. Therefore, the correlation encoding between labels is learned through the graph convolutional network to avoid the mutual exclusivity of labels, and then the potential connection between tongue image features and label correlation is learned, and the tongue image feature extraction encoding of the high-width channel interaction MLP model is integrated.

[0047] In step S2, before the tongue image is input into the model, it is first scaled to the standard size of 256×256. Then, a smaller image of 224×224 is randomly cropped from the scaled image. The cropped image is then randomly horizontally flipped for data augmentation.

[0048] Step S3 specifically adopts the following scheme:

[0049] Assume that there are N samples in the dataset D The physical label category is C con Class. We denote X as the input sample set and Y as the label set. is the category vector of each constitution. Divide the dataset D into D train and D test, respectively used for training and testing of the high-width channel interaction MLP model. For tongue images, assume that the body type set T = {t1, t2, ..., t n}, n represents the type of constitution; then, after visual modeling of the tongue image, the feature code f is obtained, and after inputting into the classifier, the output constitution decision vector category Q is obtained con =[q1,q2,…,q n ], whose dimension n is equal to the size of the set T. The elements of the decision vector are q i ∈{0, 1}, when q i = 1, the current input tongue image corresponds to the body type t i .

[0050] Use the high-width channel interaction MLP model to model the input tongue image X, model HWmixerF mlp (·, W k )The visual features C(X) extracted from the input X can be expressed as:

[0051] C(CX=HWmixerF mlp (X,,W k )

[0052] Where W k Represents the weight of the high-width channel interaction MLP model.

[0053] Then, a binary classifier is used to input the feature C(X) into the fully connected layer FC, which is implemented by the sigmoid activation function σ(·). The physical category is as follows:

[0054]

[0055] Where, Represents the parameters of the output FC layer; represents the entire parameter set; Q con (X, β con ) represents the output vector of physical identification, which is composed of the physical label category probability q(t i |X,β com )composition.

[0056] In step S3, the input tongue image is divided into blocks using patch embedding, and a convolutional layer with a 7×7 convolution kernel, a stride of 4, and a padding of 2 is used to extract preliminary features as the input features of the H and W interaction modules. To enhance the spatial feature interaction of the input image, we use the HWmixer-MLP model to extract deep features of the tongue image. The H and G interaction module consists of multiple blocks, each of which is composed of a 1×1 convolutional layer, normalization, ReLU activation function, and alternating 1×7 and 7×1 convolutional layers. The Token feature aggregation block is composed of alternating normalization, Channel-MLP, and GELU activation functions, and is downsampled by a 3×3 convolutional layer (stride = 2) to reduce computational complexity, thereby obtaining the input features of the next layer.

[0057] In step S3, the input image is subjected to feature learning and extraction using a high-width channel interaction MLP model. The high-width channel interaction MLP model includes several convolutional layers, normalization layers, downsampling layers, learnable high-width feature interaction layers, and fully connected layers. The feature learning and extraction process of the HWmixer-MLP model can be expressed as follows:

[0058] Assume that the input is X=[x1,x2,…,x n ], the corresponding Token-FC feature can be expressed as:

[0059]

[0060] Where, Represents the spatial features of each feature interaction transformation layer, where Will Add the weight to get the output interaction feature φ j , then the aggregated features H and W interact to output φ j for:

[0061]

[0062] Where W t Represents the learnable weight; k represents the k-th input Token; j represents the j-th output Token; s k Represents the feature vector after Token interaction transformation.

[0063] In summary, the mathematical expression for feature extraction based on the MLP-like model of H and W channel information interaction is:

[0064] Y=HWmixer FC (BN(X))+X

[0065] Z=CH FC(BN(Y))+Y

[0066] Where X represents the input features; BN represents batch normalization, in which Gaussian error GELU is used as the activation function; Y represents the intermediate features learned by the H and W channel interaction MLP module; Z represents the output features related to the physical label category corresponding to the tongue image.

[0067] The features extracted by the deep network are input into the Sigmoid classifier to obtain the physical category:

[0068] q * =Sigmoid(FC(Z))

[0069] Among them, q * Represents the probability vector of the output physical multi-label category; Sigmoid represents the probability value between the output label category (0,1); FC represents the fully connected layer; Z represents the output feature related to the physical label category corresponding to the input tongue image.

[0070] In step S3, a two-layer graph convolutional network model that embeds label relationships is trained using the training dataset, including the following steps:

[0071] S31: Input the body constitution labels corresponding to the tongue image into the graph convolutional network branch. The correlation between the body constitution labels is learned through a two-layer graph convolutional network. The matrix is converted into the body constitution label correlation matrix through Sigmoid and then integrated into the high-width channel interaction MLP network to extract the features of the tongue image.

[0072] S32: The training uses the loss function as the objective function to guide the multiple iterative training of the high-wide channel interaction MLP model that integrates graph convolution label relationship learning, and saves and outputs the optimal model parameters for physical identification testing.

[0073] Construction of loss function in step S32: Calculate the binary cross entropy classification loss of tongue image feature constitution multi-label category. Assume constitution label Y = {y1, y2, ..., y C},y i ∈{0, 1}, then the binary cross entropy loss function of the physical label is expressed as:

[0074]

[0075] Where C represents the number of body type labels corresponding to each tongue image, i.e., 9 types of body type labels; i Indicates the real physical label; q i Represents the predicted constitution label.

[0076] This implementation utilizes the deep learning framework PyTorch and the model library timm. All experiments were run on a server equipped with a GeForce GTX 1080 GPU, 12GB of video memory, a 1.2GHz CPU, and Ubuntu 14.04. The proposed method uses stochastic gradient descent (SGD) for training with the following parameters: weight decay of 5e-4, momentum of 0.9, and a batch size of 64. Furthermore, the model was trained for 200 epochs, with an initial learning rate of 0.01. This learning rate was decayed using cosine annealing from the 60th epoch to a minimum learning rate of 2e-4. The input tongue images were resized to 224×224, and both training and test images were normalized.

[0077] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

[0078] Table 1

[0079]

Claims

1. A method for identifying body constitution based on tongue image features based on a MLP-like network, characterized by: The following steps are involved: Step S1: Data collection and processing: Collect tongue images, label the corresponding tongue images with physical labels, and create training and test data sets; Step S2 constructs a high-width channel interaction MLP model: a high-width channel interaction MLP model is proposed that integrates graph convolution label relationship learning and uses body shape labels to supervise the learning of tongue images; Step S3: Graph Convolutional Network Model Embedding Label Relationships: Based on the TCM theory of combined constitutions, a graph convolutional network is used to learn the correlation encoding between constitution labels to avoid mutual exclusivity of constitution labels. Furthermore, the potential relationship between tongue image features and constitution label correlations is learned, and the encoding of tongue image feature extraction is incorporated into the high-width channel interactive MLP model. The step S2 of constructing the high-width channel interaction MLP model includes the following steps: S21: The normalized batch tongue images are used as input images for the high-width channel interaction MLP model; S22: Using a high-width channel interaction MLP model to learn and extract features from the input image, the high-width channel interaction MLP model includes several convolutional layers, normalization layers, downsampling layers, learnable high-width feature interaction layers, and fully connected layers; S23: Using a two-layer graph convolutional network to learn the correlation between the physical labels corresponding to the input tongue image, the mutual exclusivity and high-intensity parallelism matrix of the physical labels are integrated into the extracted tongue image features; S24: Binary cross entropy is used as the loss function as the objective function for tongue constitution label recognition of the high-width channel interaction MLP model. The tongue image extracted by the high-width channel interaction MLP model and the corresponding label relationship features are input into the Sigmoid classifier to obtain the constitution category: , in, Represents the probability vector of the output physical multi-label category; Sigmoid represents the probability value between the output label category (0,1); FC represents the fully connected layer; Represents the output features related to the physical label category corresponding to the input tongue image; In S22, the feature learning and extraction process of the high-width channel interaction MLP model can be expressed as: Assume the input is , the corresponding Token-FC feature can be expressed as: , Where, Indicates the characteristics after token aggregation; Represents the Token aggregation block function; Represents the network weight parameters; Represents a multiplication operation; k represents the kth input Token; j represents the jth output Token; The number of tokens representing the features; Represents the spatial features of each feature interaction transformation layer, where ,Will Add the weights to get the interactive features of the output , then the aggregated features H and W interact to output for: ; Where, Represents the learnable weight; k represents the kth input Token; j represents the jth output Token; Represents interactive features; Represents the feature vector after Token interaction transformation; In summary, the mathematical expression for feature extraction based on the MLP-like model of H and W channel information interaction is: , , Where, Represents the channel information interaction feature extraction block function; Represents the channel feature learning block function; Represents the characteristics of the input; represents batch normalization, where Gaussian error GELU is used as the activation function; Represents the intermediate features learned by the H and W channel interaction MLP modules; Represents the output features related to the physical label category corresponding to the tongue image.

2. The method for identifying body constitution based on tongue image features based on an MLP-like network according to claim 1, characterized in that: The constitution labels include balanced constitution, qi deficiency constitution, yang deficiency constitution, yin deficiency constitution, phlegm-damp constitution, damp-heat constitution, blood stasis constitution, qi stagnation constitution, and special constitution.

3. The method for identifying body constitution based on tongue image features based on an MLP-like network according to claim 1, characterized in that: Before the tongue image is input into the model, it is first scaled to a uniform standard size. Then, a smaller image is randomly cropped from the scaled image. The cropped image is randomly horizontally flipped for data augmentation. The enhanced tongue image is then input into a high-width channel interaction MLP model. The correlation between the tongue image and the corresponding physical label learned by graph convolution is integrated to automatically learn tongue image features to identify physical categories.

4. The method for identifying body constitution based on tongue image features based on an MLP-like network according to claim 1, characterized in that: The step S3 embeds a graph convolutional network model of label relationships, comprising the following steps: S31: Input the body constitution labels corresponding to the tongue image into the graph convolutional network branch. The correlation between body constitution labels is learned through a two-layer graph convolutional network. The matrix is converted into a body constitution label correlation matrix through Sigmoid and then integrated into the high-width channel interaction MLP model to extract tongue image features. S32: The training uses the loss function as the objective function to guide the multiple iterative training of the high-wide channel interaction MLP model that integrates graph convolution label relationship learning, and saves and outputs the optimal model parameters for physical identification testing.

5. The method for identifying body constitution based on tongue image features based on an MLP-like network according to claim 4, characterized in that: In step S32, the loss function is constructed by calculating the binary cross entropy classification loss of the tongue image feature constitution multi-label category, assuming that the constitution label , then the binary cross entropy loss function expression of the physical label is: , Where, Indicates the number of body type labels corresponding to each tongue image, i.e., 9 types of body type labels; Indicates a true physical label; represents the predicted constitution label, ( ) represents logarithmic operation.

Citation Information

Patent Citations

  • Tongue constitution recognition method based on wavelet attention and remodeling fusion

    CN115661047A

  • Visceral organ attribute coding method and system fusing multi-modal features

    CN116467675A