Tongue image multi-attribute classification method based on CNN-Transform network

The CNN-Transformer network enhances tongue image classification by modeling global dependencies and applying a group masking strategy, improving classification accuracy and addressing attribute imbalance and semantic extraction challenges.

CN120318584APending Publication Date: 2025-07-15SHANGHAI UNIVERSITY OF INTERNATIONAL BUSINESS AND ECONOMICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510470730.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art has problems such as attribute imbalance, difficulty in extracting pixel-level semantic information and difficulty in modeling attribute correlation in the multi-attribute classification of tongue images, which affects the repeatability and accuracy of the diagnosis.

Method used

The multi-attribute classification method of tongue image based on CNN-Transformer network is adopted, and the tongue image features and attribute embedding are extracted through the embedding extraction module, and the global dependency of image features and attribute embedding is modeled using the double-layer attention Transformer module, and the joint loss function is constructed through the grouping masking strategy to improve the model's learning ability of low-frequency attributes.

Benefits of technology

The accuracy of multi-attribute classification of tongue images is improved, the problems of attribute imbalance and difficulty in extracting semantic information are solved, and the reliability and accuracy of diagnosis are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318584A_ABST
    Figure CN120318584A_ABST
Patent Text Reader

Abstract

The invention discloses a tongue image multi-attribute classification method based on a CNN-Transform network, and the method comprises the steps: carrying out the preprocessing of a tongue image, obtaining an image data set, dividing the image data set into a training set and a test set, inputting the tongue image in the training set and a corresponding attribute tag into a tongue image multi-attribute classification-double-layer attention network model; respectively extracting image features and attribute embedding of the tongue image and the attribute tag through an embedding extraction module; modeling is carried out on the global dependency relationship between image features and attribute embedding through a double-layer attention Transform module, and attribute embedding is updated; and the updated attributes are embedded and input into an attribute classifier module to carry out attribute probability output, a grouping mask strategy is added to construct a joint loss function, and a multi-attribute classification model is trained. According to the method, the advantages of CNN and Transform can be combined, and the accuracy of tongue image multi-attribute classification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically to a multi-attribute classification method for tongue images based on a CNN-Transformer network. Background Technique

[0002] The four diagnostic methods in traditional Chinese medicine, "inspection, auscultation and olfaction, interrogation, and palpation", are the essence of syndrome differentiation and treatment in traditional Chinese medicine, and inspection ranks first among the "four diagnostic methods". Tongue diagnosis is an important part of inspection in traditional Chinese medicine and has been emphasized by medical experts of all dynasties. Traditional Chinese medicine observes the changes in the color, texture, maturity, and tooth marks of the tongue to assist in diagnosing and differentiating the physiological functions and pathological changes of the body. For example, the tongue color can reflect the functions of the zang-fu organs, the circulation of qi and blood, and the pathological state. However, traditional tongue diagnosis is based on basic theories and is understood and judged by the human eye with the help of personal experience, so there will be certain differences, affecting the repeatability of diagnosis. The computer-aided tongue image analysis and diagnosis technology emerged as the times require. In intelligent tongue diagnosis in traditional Chinese medicine, the classification of tongue attributes has become a key research field.

[0003] However, due to the complexity and diversity of tongue features, classical image processing methods have problems such as time-consuming and space-consuming algorithms, difficulty in automated high-throughput processing, and weak transfer ability in correlation research, and cannot comprehensively analyze tongue images.

[0004] In recent years, methods mainly based on deep learning have been introduced into the research of multi-attribute classification of the tongue. Compared with traditional methods, they can learn favorable features from raw data without manually constructing features, thus obtaining more effective feature representations. However, the problems of attribute imbalance in tongue images, difficulty in extracting pixel-level semantic information in images, and difficulty in modeling attribute correlations are major challenges in the current multi-attribute classification task of tongue images. Summary of the Invention

[0005] The purpose of the present invention is to provide a multi-attribute classification method for tongue images based on a CNN-Transformer network, which is used to improve the classification accuracy of multi-attributes of tongue images, thereby solving the problems in the prior art.

[0006] To achieve the above purpose, the present invention provides the following technical solutions:

[0007] A multi-attribute classification method for tongue images based on a CNN-Transformer network, which is used to perform multi-attribute classification on tongue images to obtain the corresponding attribute categories of tongue color, tongue shape, coating color, coating texture, and glossiness at one time, including:

[0008] Step S1, preprocess the tongue images to obtain an image data set, which is divided into a training set and a test set, and input the tongue images and corresponding attribute labels in the training set into the multi-attribute classification - double attention network model for tongue images;

[0009] Step S2, use the embedding extraction module to extract the image features and attribute embeddings of the tongue image and the attribute label respectively;

[0010] Step S3, use the double-layer attention Transformer module to model the global dependence between the image features and the attribute embeddings, and update the attribute embeddings;

[0011] Step S4, input the updated attribute embeddings into the attribute classifier module to output the attribute probabilities, add the grouped mask strategy to construct the joint loss function, and train the multi-attribute classification model.

[0012] Furthermore, the tongue image multi-attribute classification - double-layer attention network model includes:

[0013] In the embedding extraction module, use the pixel-level image feature extraction backbone network TResNet to extract the tongue image features, and use the torch.nn library in Pytorch to learn the attribute label embedding vectors from scratch;

[0014] In the double-layer attention Transformer module, model the global dependence and interaction between the image features and the attribute embeddings through the self-attention and cross-attention mechanisms, and use the attribute embeddings as the query vectors to check the existence of each attribute by adaptively performing multi-head cross-attention on the pooled target attribute features;

[0015] In the attribute classifier module, input the updated attribute embedding vectors for binary classification and predict the probabilities of the corresponding attributes;

[0016] Add the tongue attribute grouped mask training strategy to group and mask the tongue attributes, improve the learning ability for low-probability attributes, predict the masked attribute probabilities, and construct the joint loss function to train the model.

[0017] Furthermore, the method for the embedding extraction module to extract image features includes:

[0018] Improve the TResNet network by replacing the second convolutional layer in the basic block and the bottleneck block with the pyramid convolutional layer. For the input feature map FM i , apply different convolutional kernels to each layer {1, 2, 3,..., n} of the pyramid convolution, and each layer has different spatial dimensions and different convolutional kernel depths and output different numbers of output feature maps

[0019] Furthermore, the method for the double-layer attention Transformer module to update the attribute embedding vectors includes:

[0020] Step S301, embed and splice the image features and attributes \(H = \{h_1, h_2, \ldots, h\) n \} into the Transformer encoder, which consists of \(L1\) identical encoder layers, each containing a self-attention layer and a feed-forward network;

[0021] Step S302, the updated embedding set \(H'=\{h_1', h_2', \ldots, h\) n '\} is input into the Transformer decoder, which is composed of a total of \(L2\) identical decoder layers, each containing a self-attention layer, a cross-attention layer and a feed-forward network.

[0022] Furthermore, the step S301 includes:

[0023] Calculate the normalized scalar attention coefficient \(\alpha\) i between \(h\) j and \(h\) ij . After calculating \(\alpha\) i for all \(h\) j and \(h\) ij , use the weighted sum of all embeddings, and then use the non-linear ReLU function to update each embedding \(h\) i to \(h\) i ':

[0024]

[0025]

[0026] where \(W\) q is the query weight matrix, \(W\) k is the key weight matrix, \(d\) is the feature dimension, \(W\) v is the value weight matrix, \(W\) r and \(W\) o are transformation matrices, and \(b1\) and \(b2\) are bias vectors.

[0027] Furthermore, the step S302 includes:

[0028] The attribute embedding part \(L'=\{l_1', l_2', \ldots, l\) l '\} in \(H'\) enters the self-attention layer to update its own characteristics to \(L''=\{l_1'', l_2'', \ldots, l\) l ''\}. The cross-attention layer uses the updated attribute embedding \(L''\) as the query vector to extract and aggregate attribute-specific information from the image features \(Z'=\{z_1', z_2', \ldots, z\) h×w '\} in \(H\). The spatial features are calculated as key-value pairs:

[0029]

[0030] L″′ = {l1″′, l2″′,..., l l ″′} is the finally updated attribute embedding vector.

[0031] Furthermore, the attribute classifier module includes:

[0032] Using an independent feed-forward network FFN i to predict the probability that each final attribute embedding l i ″′ is positive, where FFN i contains a single linear layer, the weight matrix w of attribute i i is a d×1 vector, b i is the bias, and σ is the sigmoid function:

[0033]

[0034] Using an improved asymmetric loss function to calculate the loss function L raw :

[0035]

[0036] L raw where y i is the binary label indicating whether the input tongue image has attribute i, is the predicted probability of each attribute, and γ+ = 0, γ- = 1 are set.

[0037] Furthermore, the method for the grouped mask strategy to solve attribute imbalance includes:

[0038] Randomly masking a certain number of attribute embeddings using the grouped mask ratio ɑ = (ɑ1, ɑ2, α3, α4, α5). The mask ratio α1 determines the proportion of attribute embeddings to be masked in the tongue color group, α2 represents the tongue shape group, a3 represents the coating color group, a4 represents the coating texture group, and a5 represents the glossiness group. The masked attribute embeddings are represented as [mask], and a new masked attribute embedding M = (M(l1), M(l2),..., M(l l )) is input into the double-layer attention Transformer module and the attribute classifier module again, and the loss function for masked attribute prediction is additionally calculated:

[0039]

[0040] The final loss function is calculated by combining L raw and L mask with λ used to balance these two losses:

[0041] L total= L raw + λL mask 。

[0042] Furthermore, the method for the grouping mask strategy to solve the attribute imbalance further includes:

[0043] Obtain the tongue image and its corresponding attribute label dataset, train the multi-attribute classification - double attention network model for the tongue image, and train it on an NVIDIA A100 GPU through the Pytorch framework. The number of Transformer encoder - decoder layers is set to 2 and 3 respectively, the number of attention heads is set to 4, the grouping mask ratio is set to α = (α1 = 0.25, α2 = 0.3, α3 = 0.5, α4 = 0.2, α5 = 0.5), in the joint loss L total , λ is set to 1. If the output probability is greater than 0.5, the predicted attribute is positive. The optimization method uses Adam for 30 epochs of training, the batch size is 4, True - weight - decay is 1e - 2, (β1, β2) = (0.9, 0.9999), the learning rate is 1×10 -5 , and dropout with p = 0.1 is used for regularization.

[0044] Furthermore, the tongue image pre - processing method includes the following steps:

[0045] Augment the number of tongue images by flipping, translating, randomly splitting, rotating by a predetermined angle, scaling by a predetermined ratio, and adding Gaussian noise to the tongue image.

[0046] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0047] The present invention constructs a multi - attribute classification model for the tongue based on the CNN - Transformer framework. The embedding extraction module takes the tongue image as input, converts it into image - space features, and generates initial attribute embedding vectors for all input attributes during training. The double - attention Transformer module models the global dependencies and interactions between the image features and the attribute embeddings, and updates the attribute embeddings. The attribute classifier module predicts the multi - attributes of the tongue image according to the updated attribute embeddings. By adding the grouping mask strategy, the model's learning ability for low - frequency attributes is improved, and the classification accuracy is improved. The present invention can effectively solve the problems of attribute imbalance in tongue images, difficulty in extracting pixel - level semantic information in images, and difficulty in modeling attribute correlations, integrates the advantages of CNN and Transformer, and is beneficial to the doctor's diagnosis work. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1Flow chart of the tongue image multi-attribute classification method based on the CNN-Transformer network according to the embodiment of the present invention;

[0049] Figure 2 Partial example diagram of the tongue image and its attribute label data set according to the embodiment of the present invention;

[0050] Figure 3 Schematic flow diagram of the working process of the tongue image multi-attribute classification - double attention network according to the embodiment of the present invention;

[0051] Figure 4 Schematic diagram of the improved TResNet according to the embodiment of the present invention.

[0052] Figure 5 Schematic diagram of the double attention Transformer module according to the embodiment of the present invention.

[0053] Figure 6 Flow chart of the grouped mask strategy according to the embodiment of the present invention. Detailed implementation manners

[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] Figure 1 It is a flow chart of the tongue image multi-attribute classification method based on the CNN-Transformer (CNN: Convolutional Neural Network, which is a convolutional neural network; Transformer: a deep learning model based on the self-attention mechanism) network provided by the embodiment of the present invention, specifically including:

[0056] Step S1, preprocess the tongue image to obtain an image data set, and divide it into a training set and a test set, and input the tongue images and corresponding attribute labels in the training set into the tongue image multi-attribute classification - double attention network model;

[0057] Step S2, respectively extract the image features and attribute embeddings of the tongue image and the attribute label through the embedding extraction module;

[0058] Step S3, model the global dependence relationship between the image features and the attribute embeddings through the double attention Transformer module, and update the attribute embeddings;

[0059] Step S4: Embed the updated attributes into the input attribute classifier module for attribute probability output, add a grouping mask strategy to construct a joint loss function, and train the multi-attribute classification model.

[0060] As Figure 2 shown, it is a local example diagram of the tongue image and its attribute label dataset in the embodiment of the present invention.

[0061] In this embodiment, some data is as Figure 2 shown. The tongue image dataset includes images collected from the tongues of different people, totaling 500 images. Different colors are used to distinguish the primary attributes of the tongue. Orange represents the tongue color group, blue represents the tongue shape group, green represents the coating color group, purple represents the coating texture group, and gray represents the glossiness group. Further, the tongue color includes pale white, light red, red, and purple. The tongue shape includes tooth marks, petechiae, and cracks. The coating color includes white coating and yellow coating. The coating texture includes thin coating, thick coating, moist coating, dry coating, greasy coating, peeled coating, and the glossiness includes dull and bright, totaling 17 secondary attributes. All images are resized to 448×448.

[0062] As Figure 3 shown, it is a schematic flowchart of the working process of the multi-attribute classification - double attention network for the overall tongue image in the embodiment of the present invention; the modules of the network include: the embedding extraction module takes the tongue image as input, converts it into image space features, and generates initial attribute embeddings for all input attribute labels during training; the double attention Transformer module models the global dependencies and interactions between the image features and the attribute embeddings, and updates the attribute embeddings; the attribute classifier module predicts all attribute categories of the tongue image based on the updated attribute embeddings. By adding a grouping mask strategy, the model's ability to learn low-frequency attributes is improved, thereby improving the classification accuracy.

[0063] Figure 4 It is a schematic diagram of the improved TResNet in the embodiment of the present invention.

[0064] As Figure 4As shown, the improved TResNet is mainly used for pixel-level feature extraction of tongue images. TResNet retains the traditional basic blocks and bottleneck blocks, and makes improvements in five aspects: SpaceToDepth stem (SpaceToDepth stem is an architecture improvement method for convolutional neural networks (CNNs), mainly used to improve the accuracy and efficiency of the network. SpaceToDepth stem is part of the ResNet architecture), Anti-Alias downsampling (referred to as AA for anti-aliasing downsampling), InPlaceActivated BatchNorm (referred to as Inplace ABN, a technique for optimizing the memory occupancy of deep neural network (DNN) training), block selection, and SE module (fully named Squeeze-and-Excitation Module). The main advantages are greater GPU throughput and higher accuracy and efficiency. On this basis, the embodiments of the present invention further improve the TResNet network by replacing the second convolutional layer in the basic block and bottleneck block with pyramid convolution (Py_Conv). For the input feature map FM i , each layer {1, 2, 3,..., n} of the pyramid convolution applies different convolutional kernels, and each layer has different spatial dimensions and different convolutional kernel depths and outputs different numbers of output feature maps It has better perception ability for the pixel-level information of tongue images and can provide more accurate feature representations for subsequent tongue image analysis and recognition tasks. Py_Conv uses convolutional kernels with filters of different sizes and depths to capture features at multiple scales. Compared with standard convolution, it does not increase the computational cost and the number of parameters.

[0065] As Figure 5 shown, it is a schematic diagram of the double-layer attention Transformer module of the embodiments of the present invention.

[0066] The encoder component of the Transformer consists of the same encoder layers. Each encoder layer contains a self-attention layer and a feed-forward network. Concatenate the image features Z = {z1, z2,..., z h×w} and the attribute embeddings L = {l1, l2,..., l l} to get H = {z1,..., z h×w , l1,..., l l} = {h1, h2,..., h n} as the input to the Transformer encoder. In the Transformer, the embedding h j ∈H is relative to hh The importance or weight of ∈H is learned through self-attention. The embedding h h and h j The attention score between is calculated as follows. First, the present invention calculates the normalized scalar attention coefficient α between the embedding h h and h j . After calculating α for all h h and h j , the embodiments of the present invention use the weighted sum of all embeddings, and then use the non-linear ReLU function to update each embedding h ij to h′ h : i :

[0067]

[0068]

[0069]

[0070] where is W q The query weight matrix, W k is the key weight matrix, d is the feature dimension, W v is the value weight matrix, W r and W o are transformation matrices, and b1 and b2 are bias vectors. This update process can be repeated for this layer, where the updated embedding h′ i enters the Transformer encoder layer as input. The learned weight matrices W k , W q , W v , W r , are not shared between layers. After multi-layer updates, the final output of the Transformer encoder is represented as H′ = {z1′,..., z h×w ′, l1′,..., l l ′}.

[0071] The decoder component of the Transformer consists of the same decoder layers. Each decoder layer contains a self-attention layer, a cross-attention layer, and a feed-forward network. The Transformer decoder takes the output H′ of the Transformer encoder as input, uses the learnable label embedding as the query vector, and probes and pools the attribute-related features through the cross-attention mechanism in the Transformer decoder. The pooled features have strong adaptability and discriminability, and can well improve the performance of multi-attribute classification. The attribute embedding part L′ = {l1′, l2′,..., ll ′} enters the self-attention layer. Similar to the above self-attention layer method, its own characteristics are updated to L″ = {l1″, l2″,..., l l ″}. Then, the cross-attention layer uses the updated attribute embedding L″ as the query vector to extract and aggregate attribute-specific information from the image features Z′ = {z1′, z2′,..., z h×w ′} in H′. These spatial features are calculated as key-value pairs:

[0072]

[0073] L″′ = {l1″′, l2″′,..., l l ″′} is the finally updated attribute embedding vector.

[0074] The attribute classifier module described in the present invention uses an independent feed-forward network FFN i to predict the probability that each final attribute embedding l i ″′ is positive. Among them, FFN i contains a single linear layer. The weight matrix w of attribute i is a d×1 vector, and b i is the bias, and σ is the sigmoid function:

[0075]

[0076] An improved asymmetric loss function is used to calculate the loss function L raw :

[0077]

[0078] L raw where y i in is the binary label indicating whether the input tongue image has attribute i, is the predicted probability of each attribute. Let

[0079] γ+ = 0 and γ- = 1 be set.

[0080] As Figure 6 shown, it is the flowchart of the grouped mask strategy of the embodiment of the present invention.

[0081] Using the same probability mask for all attributes will result in insufficient learning of low-frequency attributes, thereby reducing the overall prediction accuracy. The embodiment of the present invention maps different attributes to different tokens and constructs a set of masked attribute embeddings M. The advantage of this strategy can help the model learn various attributes with different frequencies more fully and evenly.

[0082] In one embodiment of the present invention, a grouping mask ratio ɑ = (ɑ1, ɑ2, ɑ3, α4, α5) is used to randomly mask a certain number of attribute embeddings. The mask ratio α1 determines the proportion of attribute embeddings to be masked in the tongue color group, α2 represents the tongue shape group, α3 represents the tongue coating color group, α4 represents the tongue coating texture group, and α5 represents the glossiness group. The masked attribute embeddings are represented as [mask], and a new masked attribute embedding M = (M(l1), M(l2),..., M(l l )) Input the double-layer attention Transformer module and the attribute classifier module again, and additionally calculate the loss function of the masked attribute prediction:

[0083]

[0084] The final loss function is composed of L raw and L mask Combined calculation, λ is used to balance these two losses:

[0085] L total = L raw + λL mask

[0086] Furthermore, it also includes:

[0087] Obtain the tongue image and its corresponding attribute label dataset, train the multi-attribute classification - double-layer attention network model for the tongue image, and perform training on an NVIDIA A100 GPU through the Pytorch framework. The number of Transformer encoder and decoder layers are set to 2 and 3 respectively, the number of attention heads is set to 4, the grouping mask ratio is set to ɑ = (ɑ1 = 0.25, α2 = 0.3, ɑ3 = 0.5, ɑ4 = 0.2, ɑ5 = 0.5), and in the joint loss L total λ is set to 1. If the output probability is greater than 0.5, the predicted attribute is positive. The optimization method uses Adam for 30 epochs of training, the batch size is 4, True-weight-decay is 1e-2, (β1, β2) = (0.9, 0.9999), the learning rate is 1×10 -5 , and dropout with p = 0.1 is used for regularization.

[0088] Furthermore, the tongue image preprocessing method includes the following steps:

[0089] Augment the number of tongue images by flipping (in both horizontal and vertical directions), translation, random segmentation, a predetermined rotation angle, a predetermined scaling ratio, and adding Gaussian noise to the tongue image.

[0090] In some embodiments, the Average Precision (AP for short, average precision rate), Average Recall (AR for short, average recall rate), Average F1-score (AF1 for short, average F1 score), Overall Precision (OP for short, overall precision), Overall Recall (OR for short, overall recall rate), and Overall F1-score (OF1 for short, overall F1 score) are used as model evaluation metrics.

[0091] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A multi-attribute classification method for tongue images based on a CNN-Transformer network, which is used to perform multi-attribute classification on tongue images to obtain the corresponding attribute categories of tongue color, tongue shape, coating color, coating texture, and glossiness at one time. It is characterized in that, Including: Step S1: Preprocess the tongue images to obtain an image dataset, which is divided into a training set and a test set. Input the tongue images and corresponding attribute labels in the training set into the tongue image multi-attribute classification - double attention network model; Step S2: Respectively extract the image features of the tongue images and the attribute embeddings of the attribute labels through the embedding extraction module; Step S3: Model the global dependence relationship between the image features and the attribute embeddings through the double attention Transformer module, and update the attribute embeddings; Step S4: Input the updated attribute embeddings into the attribute classifier module for attribute probability output, add a grouped masking strategy to construct a joint loss function, and train the multi-attribute classification model.

2. The multi-attribute classification method of tongue images based on the CNN-Transformer network according to claim 1, characterized in that The tongue image multi-attribute classification - double attention network model includes: In the embedding extraction module, use the pixel-level image feature extraction backbone network TResNet to extract tongue image features, and use the torch.nn library provided by Pytorch to learn the attribute label embedding vectors from scratch; In the double attention Transformer module, model the global dependence relationship and interaction between the image features and the attribute embeddings through the self-attention and cross-attention mechanisms, and use the attribute embeddings as query vectors to check the existence of each attribute by adaptively performing multi-head cross-attention on the pooled target attribute features; In the attribute classifier module, input the updated attribute embedding vectors for binary classification and predict the probabilities of the corresponding attributes; Add the tongue attribute grouped masking training strategy to perform grouped masking on the tongue attributes, improve the learning ability for low-probability attributes, predict the masked attribute probabilities, and construct a joint loss function to train the model.

3. The multi-attribute classification method for tongue images based on the CNN-Transformer network according to claim 2, characterized in that, The method for the embedding extraction module to extract image features includes: Improve the TResNet network by replacing the second convolutional layer in the basic block and bottleneck block with a pyramid convolutional layer for the input feature map FM i , apply different convolutional kernels to each layer {1, 2, 3,..., n} of the pyramid convolution, and each layer has a different spatial size and different depths of convolutional kernels and output different numbers of output feature maps 4. The multi-attribute classification method of tongue images based on the CNN-Transformer network according to claim 1, characterized in that, The method for the double attention Transformer module to update the attribute embedding vectors includes: Step S301, input the concatenated image features and attributes \(H = \{h_1, h_2, \ldots, h\) n \} into the Transformer encoder, which consists of \(L_1\) identical encoder layers, each layer containing a self-attention layer and a feed-forward network; Step S302, the updated embedding set H′={h1′, h2′,..., h n ′} is input into the Transformer decoder, and the Transformer decoder is composed of a total of L2 identical decoder layers, each layer containing a self-attention layer, a cross-attention layer, and a feed-forward network.

5. The multi-attribute classification method of tongue images based on the CNN-Transformer network according to claim 4, wherein The said step S301 includes: Compute the embedded h i and h j between the normalized scalar attention coefficient α ij , when calculating all h i and h j of α ij after that, use the weighted sum of all embeddings, and then use the non-linear ReLU function to update each embedding h i to h i ': Among them is W q Query weight matrix, W k is the key weight matrix, d is the feature dimension, W v is the value weight matrix, W r and W o is the transformation matrix, b1 and b2 are bias vectors.

6. The multi-attribute classification method of tongue images based on the CNN-Transformer network according to claim 4, wherein, The said step S302 includes: The attribute embedding part \(L'=\{l_1', l_2', \ldots, l\) in \(H'\) enters the self-attention layer and updates its own characteristics to \(L'' = \{l_1'', l_2'', \ldots, l\) l ''\}. The cross-attention layer uses the updated attribute embedding \(L''\) as the query vector to extract and aggregate attribute-specific information from the image features \(Z'=\{z_1', z_2', \ldots, z\) l '\} in \(H'\). The spatial features are calculated as key-value pairs: h×w '\} L″′ = {l1″′, l2″′,..., l l ″′} is the finally updated attribute embedding vector.

7. The multi-attribute classification method of tongue images based on the CNN-Transformer network according to claim 1, characterized in that, The attribute classifier module includes: Use an independent feed-forward network FFN i to predict the probability that each final attribute embedding l i ″′ is positive, where FFN i contains a single linear layer, the weight matrix w of attribute i i is a d×1 vector, b i is the bias, and σ is the sigmoid function: Use an improved asymmetric loss function to calculate the loss function L raw : L raw y in i is a binary label indicating whether the input tongue image has the attribute i, is the predicted probability for each attribute, with γ+ = 0 and γ- = 1.

8. The multi-attribute classification method for tongue images based on the CNN-Transformer network according to claim 1, characterized in that The method for the grouped masking strategy to solve attribute imbalance includes: Randomly mask a certain number of attribute embeddings using the grouping mask ratio α=(α1, α2, α3, α4, α5). The mask ratio α1 determines the proportion of attribute embeddings to be masked in the tongue color group, α2 represents the tongue shape group, α3 represents the coating color group, α4 represents the coating texture group, and α5 represents the glossiness group. The masked attribute embeddings are represented as [mask], and a new masked attribute embedding M=(M(l1), M(l2),..., M(l l )) Input the double attention Transformer module and the attribute classifier module again, and additionally calculate the loss function of the masked attribute prediction: The final loss function is composed of L raw and L mask which are combined and calculated, and λ is used to balance these two losses: L total = L raw + λL mask .

9. The multi-attribute classification method of tongue images based on the CNN-Transformer network according to claim 8, characterized in that, The method for the grouped masking strategy to solve attribute imbalance also includes: Obtain a dataset of tongue images and their corresponding attribute labels, train the multi-attribute classification - double attention network model for the tongue images, and conduct the training on an NVIDIA A100 GPU through the Pytorch framework. The number of Transformer encoder and decoder layers is set to 2 and 3 respectively, the number of attention heads is set to 4, the grouped mask ratio is set to α = (α1 = 0.25, α2 = 0.3, α3 = 0.5, α4 = 0.2, α5 = 0.5), and in the joint loss L total , λ is set to 1. If the output probability is greater than 0.5, the predicted attribute is positive. The optimization method uses Adam for 30 epochs of training, the batch size is 4, True-weight-decay is 1e-2, (β1, β2) = (0.9, 0.9999), and the learning rate is 1×10 -5 , and dropout with p = 0.1 is used for regularization.

10. The multi-attribute classification method of tongue images based on the CNN-Transformer network according to claim 8, characterized in that, The tongue image preprocessing method includes the following steps: Augment the number of tongue images by flipping, translating, randomly splitting, rotating by a predetermined angle, scaling by a predetermined ratio, and adding Gaussian noise to the said tongue images.