A Tongue Constitution Recognition Method Based on Wavelet Attention and Reshaping Fusion

Through the wavelet attention and remodeling fusion method, the features are extracted and enhanced from the tongue image, which solves the problem of insufficient feature extraction ability in the existing tongue image recognition methods, and achieves high-accuracy tongue physique recognition and interpretability prediction.

CN115661047BActive Publication Date: 2025-07-18SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211224472.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-08
Publication Date
2025-07-18
Estimated Expiration
2042-10-08

AI Technical Summary

Technical Problem

The existing tongue image recognition methods have weak feature extraction capabilities, making it difficult to extract accurate and discriminant features, and fail to efficiently utilize multi-level features in deep neural networks, lacking the diversity of features and adaptive fusion of them, resulting in low classification accuracy.

Method used

The wavelet attention module is used to extract and enhance features from the tongue image, feature fusion is performed by reshaping the fusion module, and the features are mapped to the attribute space of the tongue image using relative transformation, and the similarity between the predicted tongue image attribute vector and the real tongue image attribute vector is calculated to determine the physique category.

Benefits of technology

Effectively extract the global and local features of the tongue image, focus on important features and ignore noise, improve classification accuracy, and the predicted tongue image attribute vector is interpretable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115661047B_ABST
    Figure CN115661047B_ABST
Patent Text Reader

Abstract

The present invention discloses a tongue constitution recognition method based on wavelet attention and reshaping fusion, including: collecting tongue images to be recognized and generating image features; extracting tongue image features from the image features through a wavelet attention module, enhancing the tongue image features, and generating new features; performing feature fusion on the new features through a reshaping fusion module; enhancing the fused features by relative transformation; mapping the enhanced new features to the attribute space of the tongue image to obtain a predicted tongue image attribute vector; calculating the similarity between the predicted tongue image attribute vector and the true tongue image attribute vector in the hyperbolic space; and the constitution type pointed to by the tongue image attribute vector corresponding to the maximum similarity is the constitution category of the tongue image to be recognized. This method can effectively extract the global and local features of the tongue image, focus on important features while ignoring noise, has a high classification accuracy, and the predicted tongue image attribute vector is interpretable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tongue image classification, and particularly to a tongue constitution recognition method based on wavelet attention and reshaping fusion. Background Art

[0002] Tongue recognition is to segment and classify the acquired tongue image. The purpose of tongue segmentation is to effectively filter out the interference of background information and improve the classification performance of subsequent classifiers. The latest research applies deep learning to tongue image segmentation, including proposing a real-time automatic tongue image segmentation method that uses a lightweight architecture based on an encoder-decoder structure. And a network called Tongue U-Net is proposed, which combines the classic U-Net structure with squeeze-and-excitation blocks (SE), dense atrous convolution blocks (DAC), and residual multi-kernel pooling blocks (RMP). And a dilated encoding network (De-Net) is proposed for automatically segmenting tongue images acquired from mobile devices in an open environment. The recognized tongue image can reflect the constitution category. Existing methods use classical convolutional neural networks, gray-level co-occurrence matrices, minimum bounding rectangles, and edge curves to extract tongue image features, and then use different classifiers to classify the constitution.

[0003] Although these studies have achieved certain effects, the feature extraction ability of existing tongue image recognition methods is weak, and it is difficult to extract more accurate and discriminative features. And existing tongue image recognition methods do not efficiently utilize the multi-level features in deep neural networks and lack feature diversity. Existing tongue image recognition methods have not achieved adaptive fusion of different features, resulting in insufficient comprehensive expression of features.

[0004] Therefore, based on the existing tongue image classification technology, how to effectively extract the global and local features of tongue images, better focus on important features while ignoring noise, and improve the classification accuracy has become an urgent problem for those skilled in the art to solve. Summary of the Invention

[0005] In view of the above problems, the present invention proposes a tongue constitution recognition method based on wavelet attention and reshaping fusion that solves at least the above partial technical problems. The tongue attribute vector predicted by this method is interpretable, and can effectively extract the global and local features of tongue images, better focus on important features while ignoring noise, and improve the classification accuracy.

[0006] An embodiment of the present invention provides a tongue constitution recognition method based on wavelet attention and reshaping fusion, including the following steps:

[0007] S1. Collect the tongue image to be recognized and generate image features;

[0008] S2. Extract the tongue image features from the image features through a wavelet attention module, enhance the tongue image features, and generate new features; perform feature fusion on the new features through a reshaping and fusion module;

[0009] S3. Further enhance the fused features, map the enhanced new features to the attribute space of the tongue image, and obtain the predicted tongue image attribute vector;

[0010] S4. Calculate the similarity between the predicted tongue image attribute vector and the true tongue image attribute vector; the constitution type pointed to by the tongue image attribute vector corresponding to the maximum similarity is the constitution category of the tongue image to be recognized.

[0011] Further, in the step S2, extracting the tongue image features from the tongue image through a wavelet attention module and enhancing the tongue image features to generate new features includes:

[0012] Perform discrete wavelet transform on the image features to obtain feature components;

[0013] Generate a spatial attention mask corresponding to each feature in the feature components through spatial attention;

[0014] Perform spatial normalization on the spatial attention mask through position normalization to generate a new attention mask;

[0015] Adjust the weights of the feature components according to the new attention mask; perform feature aggregation according to the weights to generate new features.

[0016] Further, the generating a spatial attention mask corresponding to each feature in the feature components through spatial attention includes:

[0017] Generate a feature description operator by mean pooling the feature components along a preset channel dimension;

[0018] Generate an attention map using a standard convolutional layer with a convolution kernel size of 3×3 according to the feature description operator;

[0019] Generate a spatial attention mask corresponding to each feature in the feature components according to the attention map.

[0020] Further, in the step S2, performing feature fusion on the new features through a reshaping and fusion module includes:

[0021] Concatenate the new features through a concatenation operation to form a pooled feature;

[0022] Reshape the pooled feature into a three-dimensional intermediate feature through a reshaping operation;

[0023] Extract the weights between different features in the three-dimensional intermediate features through the relationship interaction module to generate a three-dimensional relationship mask;

[0024] Use the inverse reshaping operation to adjust the three-dimensional relationship mask to a one-dimensional relationship mask;

[0025] Scale each element in the pooled features by multiplying it with the attention weight corresponding to the one-dimensional relationship mask; the attention weight is calculated through the sigmoid function.

[0026] Further, the relationship interaction module extracts the weights between different features in the three-dimensional intermediate features through convolution operations.

[0027] Further, in step S3, the fused features are further enhanced through relative transformation, and the enhanced new features are converted into a predicted tongue image attribute vector using a multi-layer perceptron.

[0028] Further, in step S4, the similarity is calculated by computing the distance between the predicted tongue image attribute vector and the true attribute vectors corresponding to each preset constitution in the hyperbolic space.

[0029] Further, the distance between the predicted tongue image attribute vector and the true tongue image attribute vector in the hyperbolic space is shortened by calculating the attribute embedding loss.

[0030] The beneficial effects of the above technical solutions provided by the embodiments of the present invention at least include:

[0031] A tongue constitution recognition method based on wavelet attention and reshaping fusion provided by an embodiment of the present invention includes: collecting a tongue image to be recognized to generate image features; extracting tongue image features from the image features through a wavelet attention module and enhancing the tongue image features to generate new features; performing feature fusion on the new features through a reshaping fusion module; further enhancing the fused features, mapping the enhanced new features to the attribute space of the tongue image to obtain a predicted tongue image attribute vector; calculating the similarity between the predicted tongue image attribute vector and the true tongue image attribute vector; the constitution type pointed to by the tongue image attribute vector corresponding to the maximum similarity is the constitution category of the tongue image to be recognized. This method can effectively extract the global and local features of the tongue image, focus on important features while ignoring noise, has a high classification accuracy, and the predicted tongue image attribute vector is interpretable.

[0032] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained by the structures specifically pointed out in the written specification, claims, and drawings.

[0033] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0034] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the accompanying drawings:

[0035] Figure 1 It is a flowchart of the tongue constitution recognition method based on wavelet attention and reshaping fusion provided by the embodiment of the present invention;

[0036] Figure 2 It is an overall framework diagram provided by the embodiment of the present invention;

[0037] Figure 3 It is a structural diagram of the wavelet attention module provided by the embodiment of the present invention;

[0038] Figure 4 It is a structural diagram of the reshaping fusion module provided by the embodiment of the present invention;

[0039] Figure 5 It is a structural diagram of the reshaping operation and inverse reshaping operation provided by the embodiment of the present invention;

[0040] Figure 6 It is a schematic diagram of the relationship interaction operation for multi-level features provided by the embodiment of the present invention. Detailed Embodiments

[0041] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully communicated to those skilled in the art.

[0042] The embodiment of the present invention provides a tongue constitution recognition method based on wavelet attention and reshaping fusion. Referring to Figure 1 as shown, the method includes the following steps:

[0043] S1. Collect the tongue image to be recognized and generate image features;

[0044] S2. Extract the tongue image features from the image features through the wavelet attention module, enhance the tongue image features, and generate new features; perform feature fusion on the new features through the reshaping fusion module;

[0045] S3. Re-enhance the fused features, map the enhanced new features to the attribute space of the tongue image, and obtain the predicted tongue image attribute vector;

[0046] S4. Calculate the similarity between the predicted tongue image attribute vector and the true tongue image attribute vector; the physique type pointed to by the tongue image attribute vector corresponding to the maximum similarity is the physique category of the tongue image to be recognized.

[0047] The tongue physique recognition method based on wavelet attention and reshaping fusion provided in this embodiment can effectively extract the global and local features of the tongue image, focus on important features while ignoring noise, has a high classification accuracy, and the predicted tongue image attribute vector is interpretable.

[0048] The following elaborates in detail on the tongue physique recognition method based on wavelet attention and reshaping fusion provided in this embodiment:

[0049] S1: Input the tongue image.

[0050] S2: Automatically obtain the tongue image features corresponding to the input tongue image by using multiple-stage convolutional layers for feature extraction, wavelet attention modules for feature enhancement, and reshaping fusion modules for feature fusion, etc.

[0051] Refer to Figure 2 As shown, the backbone network uses ResNet18, which has 5 stages. Except for the first stage, the other 4 stages are composed of residual blocks with different output feature scales. The residual blocks in the Nth stage are denoted as Stage N. Given the input tongue image X ∈ R H×W×C , after the feature extraction by multiple residual blocks, features at different stages will be obtained, such as {X1, X2, X3, X4, X5}. These features are first decomposed, weighted, and aggregated by the wavelet attention module, and then cross-stage fused with the features of the next stage to obtain multiple features {F1, F2, F3, F4} with relatively rough cross-level information. These features will pass through global average pooling and the reshaping fusion module in sequence to obtain the final tongue image feature F′ ∈ R C .

[0052] Among them, the wavelet attention module in step S2 includes the following steps:

[0053] S21: Perform two-dimensional discrete wavelet transformation on the input features to obtain feature components;

[0054] S22: Pass these feature components through spatial attention and position normalization operations in sequence to obtain the corresponding attention masks;

[0055] S23: Weight and concatenate according to the attention masks, and aggregate to obtain the final output features.

[0056] On the one hand, tongue image data may contain meaningless redundant backgrounds, such as lips, teeth, nose, and facial regions. On the other hand, the image recognition model requires a strong ability to extract richer features. For this reason, a novel wavelet attention module is proposed, and its structure is shown in the appendix Figure 3 as follows. Specifically, given an input feature X of dimension H×W×C, after two-dimensional discrete wavelet (2D DWT) transform, 4 feature components of dimension are obtained. These feature components are successively passed through spatial attention and position normalization operations to obtain the corresponding attention masks Finally, the weighted concatenation is performed according to the attention masks to aggregate and obtain the final output feature The specific structure of the spatial attention operation is shown in the Figure 3 rounded rectangle box in the upper right corner, and the operation of position normalization is shown in the Figure 3 rounded rectangle box in the lower right corner.

[0057] S21: Generate feature components based on two-dimensional discrete wavelet transform.

[0058] The two-dimensional discrete wavelet transform decomposes the input data into various components of different frequencies, and the spatial dimension is exactly half of the original. Therefore, the two-dimensional discrete wavelet transform can replace the downsampling operation (max pooling, average pooling) in the convolutional neural network, and only need to set the level of the discrete wavelet transform according to the required downsampling factor.

[0059] The two-dimensional discrete wavelet transform in the proposed wavelet attention can be described as follows:

[0060]

[0061] where X ∈ R H×W×C is the input feature (input image feature) of the wavelet attention module, 2D DWT represents a two-dimensional discrete wavelet transform, and {X LL , X LH , X HL , X HH} are 4 different frequency feature components. The discrete wavelet transform decomposes the input feature into feature components of different frequencies, where the low-frequency component retains the main information of the original feature, and the high-frequency component often contains noise or texture information. More geometric and texture information is needed in tongue image recognition, and the high-frequency component can provide more detailed information for the neural network. In order to maximize the utilization of various features, these different frequency features are adaptively aggregated through the discrete wavelet transform to provide as comprehensive features as possible for the physical constitution classification task.

[0062] S22: These feature components are successively passed through spatial attention and position normalization operations to obtain corresponding attention masks.

[0063] To better aggregate the four features after discrete wavelet transform decomposition, spatial attention and position normalization operations are proposed. Spatial attention can focus on important positions in the features, thereby paying attention to important features and suppressing unnecessary features. Also, since the physical classification task of tongue images is closely related to the features in the image space, the wavelet attention adopts a spatial attention mechanism. The proposed spatial attention operation calculates the attention mask of the features through an additional module, providing a basis for the subsequent adaptive weighting process of the features. Four different attention masks will be obtained in the wavelet attention module. To consider the relationship between different frequencies during the feature aggregation process, position normalization is proposed to normalize the weights of each spatial position in the attention mask, so as to adjust their different contribution degrees.

[0064] Given four different intermediate features as input, the spatial attention of the wavelet attention module successively infers the spatial attention mask corresponding to each feature Normalize the attention mask in terms of spatial position, finely adjust the importance of each feature, and the aggregated features. As Figure 4 shown, this process can be simply summarized as:

[0065]

[0066]

[0067]

[0068]

[0069]

[0070]

[0071]

[0072]

[0073]

[0074]

[0075] Among them is the attention mask after position normalization, represents element-wise multiplication, is the fine-tuned feature, represents the concatenation operation, and X′ is the aggregated output feature.

[0076] As Figure 3 shown in the upper right corner, assume that the given feature F ∈ R H×W×C is used as the input. The spatial attention generates a spatial attention mask by extracting its internal relationships, where AvgPool is the mean pooling operation and 3×3Conv is the convolution operation with a kernel size of 3×3. First, the proposed spatial attention module applies mean pooling along the channel dimension to generate the feature descriptor F AvgPool ∈ R H×W×1 , which plays the role of feature aggregation, reducing the feature dimension, and reducing the subsequent computational consumption. For the feature descriptor F AvgPool , the spatial attention module applies a standard convolutional layer with a kernel size of 3×3 to generate the attention map M S ∈ R H×W . In short, the spatial attention module can be described by the following formula:

[0077] F AvgPool = f AvgPool (F) (12)

[0078] M S = f 3×3 (F AvgPool ) (13)

[0079] where f AvgPool (·) represents the mean pooling operation, and f 3×3 (·) represents the convolution operation with a kernel size of 3×3.

[0080] In the general attention mechanism, its ultimate goal is to calculate an attention mask and then use this mask to adjust the weights of the original features. Although different models obtain the attention mask in different ways, they all need to adjust the input features in the way of element-wise multiplication or element-wise addition. The wavelet attention module proposed in this embodiment is naturally no exception. The difference is that in order to better aggregate these 4 different frequency feature components, the wavelet attention module also introduces position normalization on this basis to normalize the 4 spatial attention masks in terms of spatial position.

[0081] The information contained in the 4 different frequency feature components obtained after two-dimensional discrete wavelet transform has different focuses, and the contributions of different feature components to different inputs are also different. The attention masks corresponding to the 4 feature components Only the attention weights of each feature component at each position are encoded, but no connection is established between them, that is, the relationship between different attention masks is ignored. Therefore, a position normalization operation is proposed to learn the relationship between these four different feature components, aiming to dynamically adjust the weights of each feature component at each spatial position for different input features, and finally learn the complementarity of different feature components to increase the feature richness.

[0082] The structure of position normalization is as Figure 3 shown in the lower right corner. Assuming that four input attention masks {M 1 , M 2 , M 3 , M 4} ∈ R H×W are given, the mathematical formula is described as follows:

[0083]

[0084] where represents the weight of the i-th input attention mask at the spatial coordinates (h, w), represents the weight of the i-th output attention mask at the spatial coordinates (h, w), represents the transformation with exponent , represents the transformation with exponent , H represents the height of the image, and W represents the width of the image. That is, the position normalization operation performs the softmax function on the four weights belonging to the same spatial position so that their weight sum is 1, and then different features have the property of mutual complementarity.

[0085] S23: Weighted concatenation is performed according to the attention mask, and the final output feature is aggregated.

[0086] After completing the two steps of generating the attention mask and position normalization of the attention mask, the last step of wavelet attention is feature aggregation. The wavelet attention module uses the concatenation method. Specifically, the new attention mask obtained by spatial attention and position normalization is used to adjust the weights of the four feature components {X LL , X LH , X HL , X HH}, and then feature aggregation is performed to obtain a new feature

[0087] The reshaping and fusion module in step S3 includes the following steps:

[0088] S31: Obtain a one-dimensional aggregated feature by concatenating multiple features;

[0089] S32: Apply a reshaping operation to rearrange the multi-level features in the three-dimensional spatial dimension;

[0090] S33: Apply a relational interaction operation to the reshaped features to learn the relationships between the features;

[0091] S34: Use an inverse reshaping operation to adjust the obtained three-dimensional relational mask into a one-dimensional relational mask that can be weighted with the one-dimensional pooled features;

[0092] S35: Finally, adjust the contribution degrees of different features according to the one-dimensional relational mask encoding the relationships between the encoded features;

[0093] S36: Re-enhance the features fused in step S35 through relative transformation, and then map the enhanced new features to the attribute space of the tongue image to obtain the predicted tongue image attribute vector.

[0094] The Reshape Fusion module is a highly efficient and powerful feature fusion method that can be used to dynamically integrate multi-level features from different network layers, enhance the importance of key features, and reduce the noise brought by irrelevant features. The specific framework of the Reshape Fusion module is as Figure 4 shown. It dynamically fuses features by means of a concatenation operation, a reshaping operation, a relational interaction operation, an inverse reshaping operation, and an element-wise multiplication operation in sequence, and finally obtains a more discriminative and rich-semantic representation.

[0095] S31: Given multiple features from different levels, first use a concatenation operation to splice them into the pooled features Assume the number of input features is 4. The concatenation operation is described by the following mathematical formula:

[0096]

[0097] C = C1 + C2 + C3 + C4 (16)

[0098] where C1, C2, C3, and C4 are the 4 one-dimensional input features respectively, and C1, C2, C3, and C4 are the dimensions of the 4 features respectively, represents the concatenation operation, and F C represents the simple pooling of the 4 features. Next, the reshaping operation rearranges the one-dimensional pooled feature F C along the spatial dimension to obtain the reshaped intermediate feature F R ∈R H×W×K .

[0099] S32: Apply a reshaping operation to rearrange the multi-level features in the three-dimensional spatial dimension

[0100] The Reshape operation is a very simple and commonly used dimensionality adjustment method. It is a function that transforms a given matrix into a matrix with the target dimensions, while the number of elements in the input matrix and the output matrix is the same during the operation. Through the Reshape operation, the matrix can be transformed from one space to another without introducing additional parameters.

[0101] The input features at multiple different levels are all one-dimensional, and they aggregate different context information. To make full use of this context information, reduce the complexity of the fully connected operation, and make full use of multiple features, a very simple feature preprocessing method - the Reshape operation is introduced. The specific process is as Figure 5 shown. It can re-adjust the input one-dimensional features into three-dimensional features. Given the input one-dimensional aggregated feature F C ∈R C , the Reshape operation reshapes it into a three-dimensional intermediate feature F R ∈R H ×W×K , and the dimensions of the two have the following relationship:

[0102] C = H × W × K (17)

[0103] where H is the height of the intermediate feature F R , W is the width of the intermediate feature F R , and K is the number of channels of the intermediate feature F R . That is, the Reshape operation reshapes the one-dimensional feature into K feature maps with dimensions of H × W. The whole process is described formulaically as follows:

[0104]

[0105] where φ R (·) represents the Reshape function, [u 1, u2,…,u C represents each component of the one-dimensional feature F C , and [f1,f2,…,f K represents the K feature maps of the intermediate feature F R .

[0106] S33: Apply the relational interaction operation to the reshaped features to learn the relationships between features

[0107] The contributions of different features are different. As Figure 6 shown, in this embodiment, a relational interaction module is proposed to learn the weights between different features.

[0108] The relational interaction module is used to model the three-dimensional relational mask M of the reshaped feature F R ∈R H×W×K 3D ​∈R H×W×K , The specific process is shown in the following formula:

[0109]

[0110]

[0111] Where represents the convolution operation, and θ 3×3 represents a convolution kernel with a scale of 3×3. M 3D is a three-dimensional relationship mask obtained through learning, and its dimension is consistent with that of the feature F R . Each element v i corresponds to the adaptive weight of the element u R in F i .

[0112] When extracting the relationship between different features, the operation used by the relationship interaction module is the convolution operation instead of the fully connected operation, which can reduce the number of parameters introduced by this operation. For a feature with an input dimension of C = H×W×K, the following formulas are used to describe the number of parameters required to model the feature relationship using the fully connected (Fully Connected) operation and the convolution (Convolution) operation respectively:

[0113] Fully connected operation: Params = (HWK)×(HWK) = H 2 W 2 K 2 (20)

[0114] Convolution operation: Params = K×3×3×K = 9K 2 (21)

[0115] When H 2 W 2 > 9, the number of parameters of the fully connected operation will be more than that of the convolution operation, and in reality, H 2 W 2 is much larger than 9. The above comparison proves that the proposed relationship interaction module has the advantage of having fewer parameters. As shown in the areas outlined by multiple squares in Figure 6 , the relationship interaction module can extract relationships from multiple different spatial neighborhoods through the convolution operation. Since the convolution operation has the property of sharing, the learned weights consider not only the neighborhood of the feature itself but also the neighborhoods of other features, that is, the internal relationships of multiple different neighborhoods are modeled simultaneously, so as to be able to model the relationship between features more precisely.

[0116] S34: Use the inverse reshaping operation to adjust the obtained three-dimensional relationship mask into a one-dimensional relationship mask that can be weighted with the one-dimensional aggregated feature.

[0117] Because the input aggregated feature F C ∈R C belongs to a one-dimensional space, and the three-dimensional relation mask cannot directly perform element-wise multiplication with it to adjust the contribution intensity of each feature. Therefore, it is necessary to reshape the three-dimensional relation mask M 3D ∈R H×W×K into a one-dimensional relation mask M 1D ∈R C . An Inverse Reshape (IR) operation is proposed to complete this process. As Figure 5 shown, the inverse reshape operation is the inverse process of the reshape operation, and its specific description is as follows:

[0118]

[0119] where φ IR (·) represents the inverse reshape operation. After the inverse reshape operation, the dimension of the three-dimensional relation mask is reshaped into one dimension, and finally it is consistent with the dimension of the aggregated feature F C , facilitating the adjustment of the weight of each feature in the next step.

[0120] S35: Finally, adjust the contribution degree of different features according to the one-dimensional relation mask encoding the relationship between the encoded features.

[0121] Each element in the aggregated feature F C is scaled by multiplying the corresponding attention weight of the one-dimensional relation mask M 1D . This operation is completed through element-wise multiplication, and the form is as follows:

[0122]

[0123] where represents the element-wise multiplication operation, and σ(·) is defined as the sigmoid function. By adaptively adjusting the features at each level, the reshape fusion module can enhance the diversity of the output features. At the same time, the complementary properties between features are considered in the relationship interaction process, and a more comprehensive representation is further fused.

[0124] S36: Re-enhance the features fused in step S35 through relative transformation, and then map the enhanced new features to the attribute space of the tongue image to obtain the predicted tongue image attribute vector.

[0125] The method of re-enhancement is to recalculate the center vector {c1,..., c i ,..., c m} of each category for each sample feature X obtained in step S2, whose category is c. Assume there are m categories. Then calculate the distance from the category center vector c to all category center vectors in the hyperbolic space:

[0126]

[0127] Then these distances are used as the new feature C:

[0128] C = [d(c, c1)... d(c, c i )... d(c, c m )]

[0129] Finally, the new feature C and X are concatenated to form a new feature vector Z = [X C].

[0130] Use a multi-layer perceptron to convert the enhanced feature vector Z into a predicted tongue image attribute vector.

[0131] Next, map the tongue image features to the attribute space of the tongue image to obtain a predicted tongue image attribute vector.

[0132] The real tongue image attribute vector needs to be constructed in advance, which is the vector of the attribute space W Attribute of the tongue image. The construction of W Attribute is pre-constructed by traditional Chinese medicine experts according to domain knowledge, associating each constitution type with the attributes of the tongue image.

[0133] Classification and Determination of Traditional Chinese Medicine Constitutions gives the connotations and classification criteria of 9 constitutions. Each constitution describes its corresponding tongue image features. For example, "the tongue coating of the balanced constitution is thin and white, and the tongue is light red; the tongue coating of the yang-deficiency constitution is moist, the tongue is plump and tender, with tooth marks; the tongue coating of the yin-deficiency constitution is less, dry, and the tongue is red; the tongue coating of the damp-heat constitution is red and yellow and greasy; the tongue coating of the phlegm-dampness constitution is white and greasy, and the tongue is large; the tongue coating of the blood stasis constitution has varicose veins under the tongue and purple-black stasis spots". The tongue image attributes constructed in this embodiment altogether include 15 attributes, which are respectively light red, red, purple-black, large, tender, cracked, tooth marks, varicose veins, thin white, thick white and greasy, yellow and greasy, less coating, moist, dry and others.

[0134] After step S2, the input tongue image will obtain features at different stages after being extracted by multiple residual blocks, such as {X1, X2, X3, X4, X5}. These features are first decomposed, weighted and aggregated by the wavelet attention module, and then cross-stage fused with the features of the next stage to obtain multiple features {F1, F2, F3, F4} with relatively rough cross-level information. The above process enhances the expression of different-scale information in the intermediate features. Next, these intermediate features will successively pass through the global average pooling and reshaping fusion module to obtain the final tongue image feature F' ∈ R C . The above process can be described by the formula:

[0135] X1 = φ Stage1 (X) (24)

[0136] Xi = φ Stagei (X i-1 ), i ∈ {2, 3, 4, 5} (25)

[0137] X′ i = f 1×1 (φ WA (X i ))), i ∈ {2, 3, 4} (26)

[0138] F1 = f GAP (X2) (27)

[0139] F i = f GAP (X′ i + X i+1 ), i ∈ {2, 3, 4} (28)

[0140] F′ = φ RF (F1, F2, F3, F4) (19)

[0141] where φ Stagei represents the feature extraction process performed by the network block in the i-th stage, f 1×1 (·) represents a 1×1 convolution operation, φ WA (·) represents the wavelet attention module, f GAP (·) represents global average pooling, and φ RF (·) represents the reshaping and fusion module.

[0142] After obtaining the tongue image feature F′, applying a multi-layer perceptron with a hidden layer of 1 to it can convert it into a predicted tongue image attribute vector where the calculation method is:

[0143]

[0144] where MLP(·) represents the multi-layer perceptron, and W A ∈ R C×15 are the learnable parameters of the multi-layer perceptron.

[0145] S4. Calculate the similarity between the predicted tongue image attribute vector and the true tongue image attribute vector; the constitution type pointed to by the tongue image attribute vector corresponding to the maximum similarity is the constitution category of the tongue image to be recognized.

[0146] S41: Calculate the similarity between the predicted tongue image attribute vector and the true tongue image attribute vector.

[0147] W AttributeIt is known. Therefore, calculate the similarity between the predicted attribute vector u and the labeled attribute vector v corresponding to different physical constitutions, and then output the probabilities of each predicted physical constitution according to the similarity. To speed up the running speed and simplify the calculation process, the calculation process is vectorized and simplified. Therefore, the entire calculation process is as follows:

[0148]

[0149]

[0150]

[0151] where is the similarity score without normalization operation, which is calculated by the inner product operation of the predicted attribute vector and the true attribute vector corresponding to the i-th physical constitution in the attribute matrix:

[0152] s(u, v) = exp(-d(u, v));

[0153]

[0154] S42: The physical constitution type pointed to by the tongue image attribute vector with the maximum similarity is the physical constitution category of the output tongue image.

[0155] After calculating the similarity of the input tongue image belonging to each physical constitution category, take the physical constitution category corresponding to the maximum similarity as the physical constitution category of the input tongue image.

[0156] The physical constitution category of the input image obtained based on the above steps can be used to optimize the entire network. Physical constitution recognition belongs to the image classification task. Therefore, it is necessary to use the backpropagation optimization algorithm to optimize the physical constitution classification loss to make it as small as possible. The physical constitution classification loss L CLS is implemented by the cross-entropy loss, which is defined as:

[0157]

[0158] where Y ∈ R 1×9 represents the true physical constitution label of the input data, and Y i represents the probability of the i-th physical constitution, so its value is either 0 or 1. If the i-th physical constitution is the true physical constitution of the tongue image, then Y i = 1, otherwise 0.

[0159] To better constrain the predicted tongue image attribute v = F Attribute ∈ R 1×15 , making its distance from the true tongue image attribute v = W Attribute closer, introduce the Attribute Embedding (AE) loss LAE to narrow the distance between the predicted attributes and the true attributes. Use the distance between two attribute vectors in the hyperbolic space to calculate L AE :

[0160] L AE = d(u, v) (35)

[0161] The overall loss combines the physique classification loss L CLS and the attribute embedding loss L AE , as follows:

[0162] L = L CLS + L AE (36)

[0163] where the weight coefficients of the two losses are both 1.

[0164] This embodiment is implemented using the deep learning framework PyTorch and the model library timm. All experiments are run on a server equipped with 2 NVIDIA RTX 3090 GPUs. In addition, its CPU is an Intel i9-10850K, the memory is 64G, and the operating system is Ubuntu 18.04. And the same hyperparameter settings as ResNet18 are adopted. During training, use the stochastic gradient descent algorithm with a weight decay of 1e-4, a momentum of 0.9, and a batch size of 64 to learn the parameters. The total number of training epochs of the model is 300. The learning rate starts from 0.1 and is decayed to the original When training the model, only the general data augmentation method is used, that is, the size of the input image is scaled to 224×224, then 4 pixels are filled symmetrically around the image with the edge as the axis, then a copy of size 224×224 is randomly cropped, and finally the image copy is horizontally flipped with a probability of 0.5. In the test phase, the image is scaled to 224×224. In addition, this embodiment also normalizes the training and test images, that is, subtracts the mean and divides by the standard deviation. Experiments show that this embodiment outperforms all the compared physique recognition methods, not only with a small number of parameters, but also with better performance, which also shows the efficiency of the proposed method.

[0165] The tongue constitution recognition method based on wavelet attention and reshaping fusion provided in this embodiment first separates multi-scale features through two-dimensional discrete wavelet transform and uses the attention mechanism to perform weighted fusion on features of multiple scales, improving the feature extraction ability of the neural network. Secondly, the reshaping operation is used to explore the associations of various hierarchical features, and then the features are efficiently fused. Finally, wavelet attention and reshaping fusion are integrated into the convolutional neural network to generate more accurate tongue image attributes, and the constitution recognition task is completed based on the tongue image attributes. The advantages are as follows: improving the accuracy of recognizing the constitution type through tongue images and providing the interpretability of the constitution recognition results.

[0166] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and modifications.

Claims

1. A tongue constitution recognition method based on wavelet attention and reshaping fusion, characterized in that It includes the following steps: S1. Collect the tongue image to be recognized and generate image features; S2. Extract tongue image features from the image features through a wavelet attention module, enhance the tongue image features, and generate new features; perform feature fusion on the new features through a reshaping fusion module; In step S2, performing feature fusion on the new features through the reshaping fusion module includes: Concatenate the new features into aggregated features through a concatenation operation; Reshape the aggregated features into three-dimensional intermediate features through a reshaping operation; Extract the weights between different features in the three-dimensional intermediate features through a relationship interaction module to generate a three-dimensional relationship mask; Adjust the three-dimensional relationship mask to a one-dimensional relationship mask using an inverse reshaping operation; Scale each element in the aggregated features by multiplying it with the attention weight corresponding to the one-dimensional relationship mask; the attention weight is calculated through a sigmoid function; S3. Further enhance the fused features, map the enhanced new features to the attribute space of the tongue image, and obtain a predicted tongue image attribute vector; S4. Calculate the similarity between the predicted tongue image attribute vector and the true tongue image attribute vector; the constitution type pointed to by the tongue image attribute vector corresponding to the maximum similarity is the constitution category of the tongue image to be recognized.

2. The tongue constitution recognition method based on wavelet attention and reshaping fusion according to claim 1, characterized in that In step S2, extracting tongue image features from the tongue image through a wavelet attention module and enhancing the tongue image features to generate new features includes: Perform discrete wavelet transform on the image features to obtain feature components; Generate a spatial attention mask corresponding to each feature in the feature components through spatial attention; Perform spatial normalization on the spatial attention mask through position normalization to generate a new attention mask; Adjust the weights of the feature components according to the new attention mask; perform feature aggregation according to the weights to generate new features.

3. The tongue constitution recognition method based on wavelet attention and reshaping fusion according to claim 2, wherein The generating a spatial attention mask corresponding to each feature in the feature components through spatial attention includes: Generate a feature description operator by mean pooling the feature components along a preset channel dimension; Generate an attention map using a standard convolutional layer with a convolutional kernel size of 3×3 according to the feature description operator; Generate a spatial attention mask corresponding to each feature in the feature components according to the attention map.

4. The tongue constitution recognition method based on wavelet attention and reshaping fusion according to claim 1, characterized in that The relationship interaction module extracts the weights between different features in the three-dimensional intermediate features through a convolutional operation.

5. The tongue constitution recognition method based on wavelet attention and reshaping fusion according to claim 1, wherein In step S3, further enhance the fused features through relative transformation, and use a multi-layer perceptron to convert the enhanced new features into a predicted tongue image attribute vector.

6. The tongue constitution recognition method based on wavelet attention and reshaping fusion according to claim 1, characterized in that In step S4, the similarity is calculated by calculating the distance between the predicted tongue image attribute vector and the true attribute vectors corresponding to each preset constitution in the hyperbolic space.

7. The tongue constitution recognition method based on wavelet attention and reshaping fusion according to claim 1, characterized in that Pull closer the distance between the predicted tongue image attribute vector and the true tongue image attribute vector in the hyperbolic space by calculating an attribute embedding loss.

Citation Information

Patent Citations

  • Coated tongue physique recognition method based on zero sample learning

    CN110664373A

  • Equipment control method based on motor imagery EEG (Electroencephalograph), and terminal

    CN113143295A