Tongue picture analysis method based on hybrid model architecture
Through the tongue image analysis method based on the hybrid model architecture, the improved backbone network and multi-head self-attention mechanism are used, combined with the convolutional layer and the Sobel operator, the features of the tongue image image are extracted and decoded, which solves the subjectivity and limitations of tongue image analysis in the traditional method, and achieves higher analysis comprehensiveness and diagnostic accuracy.
Patent Information
- Application Number
- CN202510132536.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-06
AI Technical Summary
Traditional tongue image analysis methods rely on doctors’ experience and manually extracted image features, and have subjectivity and limitations. It is difficult to fully explore the complex relationship between tongue image features and healthy state. In the face of complex and heterogeneous tongue image data, accuracy and stability are limited.
The tongue image analysis method based on the hybrid model architecture is adopted, and the improved backbone network and multi-head self-attention mechanism are constructed, combined with the convolutional layer and the Sobel operator, the features of the tongue image are extracted, and the transformer architecture decoder and pixel decoder are decoded, and the metabolic correlation fatty liver disease corresponding to the tongue image is finally obtained.
It significantly improves the comprehensiveness and accuracy of tongue image analysis, can effectively deal with complex textures and heterogeneity characteristics, making disease diagnosis more reliable and accurate, and improving the adaptability and generalization ability of the model.
Smart Images

Figure CN120070968A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image analysis, and particularly to a tongue image analysis method based on a hybrid model architecture. Background Art
[0002] In the field of healthcare, tongue image analysis, as an important clinical diagnostic tool, has been widely used especially in traditional Chinese medicine diagnosis. By observing features such as tongue coating and tongue body, tongue image analysis can help doctors evaluate and judge the health status of patients, and has important clinical value. However, traditional tongue image analysis methods mainly rely on doctors' experience and manually extracted image features, and this dependence often has limitations when facing complex tongue image and patients' personalized health conditions. Specifically, the manually extracted image features are often too subjective and one-sided, and it is difficult to comprehensively and deeply explore the complex relationship between tongue image features and patients' health status. In addition, in the face of heterogeneous and diverse tongue image data, traditional methods are also significantly limited in terms of accuracy, stability and processing ability, and cannot meet the requirements of efficient and accurate diagnosis in modern medicine.
[0003] Therefore, how to use advanced technologies to conduct more comprehensive tongue image analysis to improve the accuracy and objectivity of diagnosis has become an important research direction in the medical field. This provides room for the development of tongue image analysis systems based on deep learning and intelligent algorithms, which can break through the limitations of traditional methods, fully explore the deep connection between tongue image features and health status, and thus provide more effective support for personalized diagnosis and treatment and health assessment.
[0004] With the rapid development of artificial intelligence technology, deep learning has shown great potential in the field of medical image analysis. In recent years, tongue image analysis methods based on deep learning have gradually attracted wide attention. In particular, convolutional neural networks (CNNs) have achieved remarkable breakthroughs in medical image analysis. Many researchers have improved the analysis ability of medical images through innovative network structures. For example, in 2019, Zhuoling Li et al. published a paper titled "CLU-CNNs: Object Detection for Medical Images" in the journal Neurocomputing. In the paper, the CLU-CNN network was constructed to detect lesions in medical images, significantly improving the accuracy of image detection; in 2020, Kang Zhou et al. published a paper titled "Automated Prostate Cancer Diagnosis Based on Gleason Grading Using Convolutional Neural Network" in the journal Computer Vision and Pattern Recognition. In the paper, the PBIR Network was constructed, and by reconstructing abnormal conditions in medical images, the accuracy of target detection was significantly improved; in 2024, Xu Qiao et al. published a paper titled "Intelligent tongue diagnosis model for gastrointestinal diseases based on tongue images" in the journal Biomedical Signal Processing and Control. In the paper, a new information fusion detection method was introduced, integrating manually crafted and auto-encoded features, as well as disease detection using squeeze-and-excitation (SE) and slot attention mechanisms.
[0005] Although existing methods have achieved remarkable results in specific tasks, most models relying on the CNN architecture still have certain limitations. These models are insufficient in capturing global information in tongue image, and traditional neural network structures are limited in the face of complex and heterogeneous tongue image features. To solve these problems, more innovative hybrid model architectures and deep learning methods are needed to comprehensively capture the detailed features and global patterns of tongue images, so as to improve the application level of tongue image analysis in medical images and provide more powerful support for personalized medicine and accurate diagnosis. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides a tongue image analysis method based on a hybrid model architecture, including the following steps:
[0007] S1. Construct a hybrid model architecture, obtain the tongue image of the patient, and form a tongue image dataset; input the tongue images of the patients in the tongue image dataset into the improved backbone network for feature extraction;
[0008] S2. Input the extracted features into the encoder for encoding to obtain the encoded feature vector;
[0009] S3. Extract features from the pre-annotated mask through the convolutional layer, and combine the extracted mask image features with the encoded feature vector obtained in the previous step to obtain the combined feature vector;
[0010] S4. Input the combined feature vector into the decoder for decoding to obtain the finally segmented tongue image. The decoder includes a Transformer architecture decoder and a pixel decoder;
[0011] S5. Scale the segmented tongue image to 224*224 pixels and keep the original 3 color channels;
[0012] S6. Input the scaled tongue image into the block encoding layer for encoding to obtain the encoded feature vector;
[0013] S7. Connect the pre-annotated category with the encoded feature vector obtained in the previous step to obtain the feature vector with category information;
[0014] S8. Encode the position information, add it to the feature vector with category information, and then complete regularization through the DropPath layer to keep the vector dimension unchanged;
[0015] S9. Input the feature vector obtained in the previous step into the encoder for encoding to obtain the encoded feature vector;
[0016] S10. Input the encoded feature vector into the decoder for decoding to obtain the decoded feature vector. The decoder is the same as the Transformer architecture decoder in step S4;
[0017] S11. Input the decoded feature vector into the normalization layer and then extract the classification information to obtain the decoded classification vector;
[0018] S12. Input the decoded classification vector into the detection layer to obtain the rating of the metabolic associated fatty liver disease corresponding to the tongue image. The rating includes mild fatty liver, moderate fatty liver, and severe fatty liver.
[0019] The further limited technical solution of the present invention is:
[0020] Further, in the improved backbone network of step S1, the tongue image of the patient passes through convolutional kernels with kernel sizes of 1 / 32*scales, 1 / 16*scales, 1 / 8*scales, and 1 / 4*scales, and strides of 2, 1, 1 / 2, and 1 / 4 respectively, from top to bottom to obtain an image feature matrix with a four-layer pyramid structure, where scales represents the image pixel size.
[0021] As described above, a tongue image analysis method based on a hybrid model architecture, step S2 specifically includes the following sub-steps:
[0022] S2.1. Divide the image feature matrix into blocks and embed position encoding, expressed as follows:
[0023] Emb i = Linear(Patch i )
[0024] z l-1 = Emb + PE
[0025] Where Emb i represents the calculation result of the linear layer for each small block of the image, Patch i represents the i-th non-overlapping matrix block, z l-1 represents the calculation result of the linear layer with added position encoding, and PE represents position encoding;
[0026] S2.2. Pass the non-overlapping feature matrix with position encoding through the multi-head self-attention mechanism, expressed as follows:
[0027] Q = z l-1 W Q
[0028] K = z l-1 W K
[0029] V = z l-1 W V
[0030]
[0031] MHA(z l-1 ) = concat(head 1 , head 2 ,..., head n )W O
[0032] z' l = LayerNorm(z l-1 + MHA(z l-1 ))
[0033] Among them, Q, K, and V respectively represent the query matrix Query, the key matrix Key, and the value matrix Value, and W Q , W K , W V respectively represent the weight matrices of the query matrix Query, the key matrix Key, and the value matrix Value. l represents the current self-attention layer number, and d k represents the dimension of the K vector, and W O represents the weight matrix of the linear transformation of the multi-head output. z′ l represents the normalized output of the l-th layer;
[0034] S2.3. Pass the calculation result obtained in the previous step through the feed-forward neural network layer to obtain the finally encoded feature vector, which is expressed as follows:
[0035] FFN(z′ l ) = ReLU(z′ l W 1 + b 1 )W 2 + b 2
[0036] z l = LayerNorm(z′ l + FFN(z′ l ))
[0037] Among them, FFN(z′ l ) represents the output of the feed-forward neural network of z′ l . W 1 represents the weight matrix of the first-layer linear layer, and b 1 represents the bias term of the first-layer linear layer. W 2 represents the weight matrix of the second-layer linear layer, and b 2 represents the bias term of the second-layer linear layer. z l represents the output result of the l-th layer of residual connection and normalization.
[0038] As described above, in a tongue image analysis method based on a hybrid model architecture, in step S3, the Sobel operator and the convolutional layer are combined to extract features for the mask, which specifically includes the following sub-steps:
[0039] S3.1. Obtain the mask image features through the convolutional layer, which is expressed as follows:
[0040] Z l = Conv(P, W 11 ) + b 11
[0041] I(x, y) = ReLU(BatchNorm(Zl ))
[0042] Among them, P represents the pre-annotated mask image, W 11 represents the convolutional kernel weight, b 11 represents the bias term, Z l represents the initial feature map;
[0043] S3.2. Obtain the mask edge features through the Sobel operator:
[0044]
[0045] Among them, I(x, y) represents the input mask; the edge information in the horizontal and vertical directions is obtained through calculation;
[0046] S3.3. Calculate and merge G x with G y into the edge intensity G, which is used to express the overall intensity of the mask at the edge:
[0047]
[0048] S3.4. Concatenate the mask image features and the mask edge features, expressed as follows:
[0049]
[0050] Among them, I(x, y) represents the input mask image features, G(x, y) represents the overall intensity of the mask at the edge, and X represents the mask features;
[0051] S3.5. Concatenate the mask feature X with the picture encoding feature z l to obtain a new feature vector, expressed as follows:
[0052] F low-level = concat(X, z l )
[0053] Among them, F low-level represents the new comprehensive feature.
[0054] As described above, a tongue image analysis method based on a hybrid model architecture, step S4 specifically includes the following sub-steps:
[0055] S4.1. Decode the combined feature vector through the decoder of the transformer architecture, and obtain the mask classification result through the mask loss function. Among them, the self-attention mechanism has the same structure as the multi-head self-attention mechanism in step S2.2 except for the output layer. In this step, the output layer of the self-attention mechanism uses the softmax function to obtain the probability distribution, expressed as follows:
[0056] y t_mask = softmax(z L W out + b out )
[0057] where z L represents the output of the last layer of the decoder, W out represents the weights of the output layer, and b out represents the bias term of the output layer;
[0058] S4.2. Decode the combined feature vector through a pixel decoder, expressed as follows:
[0059]
[0060] F fused = F low-level + F upsampled
[0061] F conv = ReLU(Conv(F fused , W) + b)
[0062] y t_pixel = softmax(Conv 1*1 (F conv , W out ) + b out )
[0063] where F low-level (x i , y i ) represents the values of neighboring pixels in the comprehensive feature, w i,j represents the interpolation weight, F low-level represents the combined feature vector output in step S3, W represents the convolution kernel, b represents the bias term, and Conv 1*1 represents the convolution calculation of a 1*1 convolution kernel, W out represents the output layer convolution kernel, and b out represents the output layer bias term;
[0064] S4.3. Multiply the mask prediction result by the pixel prediction result to obtain the segmentation result, expressed as follows:
[0065] y seg = y t_mask ⊙ y t_pixel
[0066] where y t_mask represents the mask prediction result, and y t_pixel represents the pixel prediction result.
[0067] As described above, in a tongue image analysis method based on a hybrid model architecture, in step S6, the scaled tongue image is input into the block encoding layer for encoding to obtain an encoded feature vector, which is expressed as follows:
[0068] Z l =Conv(X, W 1 ) + b 1
[0069] I(x, y) = ReLU(BatchNorm(Z l ))
[0070] where W 1 represents the convolutional kernel weight, b 1 represents the bias term, and Z l represents the initial image.
[0071] As described above, in a tongue image analysis method based on a hybrid model architecture, in steps S7 to S8, the class feature and the position feature are fused into the feature vector to obtain a feature vector with complete information, which is expressed as follows:
[0072] z concat =Concat(z encoded , y label )
[0073] where z encoded represents the encoded feature vector, and y label represents the corresponding label;
[0074] The position information is encoded and added to the feature vector with class information, and then regularization is completed through the DropPath layer to keep the vector dimension unchanged, which is expressed as follows:
[0075] z combined =z concat +z pos
[0076] z dropath =Dropath(z combined )
[0077] where z pos represents the position vector, z combined represents the fusion of the encoded feature vector and the position vector, and z dropath represents the output of the DropPath layer.
[0078] As described above, in a tongue image analysis method based on a hybrid model architecture, step S9 specifically includes the following sub-steps:
[0079] S9.1. The image feature matrix is divided into blocks and embedded with position encoding, which is expressed as follows:
[0080] Emb i = Linear(Patch i )
[0081] z l-1 = Emb + PE
[0082] Among them, Emb i is the calculation result of the linear layer for each small block of the image, Patch i represents the i-th non-overlapping matrix block, and PE represents the position encoding;
[0083] S9.2. Express the non-overlapping feature matrix with position encoding through the sparse multi-head self-attention mechanism as follows:
[0084] Q = z l-1 W Q
[0085] K = z l-1 W K
[0086] V = z l-1 W V
[0087]
[0088] MHA(z l-1 ) = concat(head 1 , head 2 ,..., head n )W O
[0089] z' l = LayerNorm(z l-1 + MHA(z l-1 ))
[0090] Among them, Q, K, and V respectively represent the query matrix Query, the key matrix Key, and the value matrix Value, W Q , W K , W V respectively represent the weight matrices of the query matrix Query, the key matrix Key, and the value matrix Value, l represents the current self-attention layer number, d k represents the dimension of the K vector, W O represents the linear transformation weight matrix of the multi-head output, z l-1 represents the calculation result of the linear layer with position encoding added, and z' l represents the output of the normalization layer;
[0091] S9.3. Take z'l An input decoder decodes the feature vector through a sparse attention layer and a feed-forward layer, expressed as follows:
[0092] Q′ = z′ l W Q′
[0093] K′ = z′ l W K′
[0094] V′ = z′ l W V′
[0095]
[0096] Among them, Q′, K′, and V′ respectively represent the query matrix Query′, the key matrix Key′, and the value matrix Value′, and W Q′ , W K′ , W V′ respectively represent the weight matrices of the query matrix Query′, the key matrix Key′, and the value matrix Value′, l represents the current self-attention layer number, d k′ represents the dimension of the K′ vector, and M is the sparse mask matrix;
[0097] S9.4. Pass the calculation result of step S9.2 through the feed-forward neural network layer to obtain the finally encoded feature vector, expressed as follows:
[0098] FFN(z′ l ) = ReLU(z′ l W 1 + b 1 )W 2 + b 2
[0099] z l = LayerNorm(z′ l + FFN(z′ l ))
[0100] Among them, z′ l represents the output of the normalization layer in step S9.2, W 1 represents the weight matrix of the first linear layer, b 1 represents the bias term of the first linear layer, W 2 represents the weight matrix of the second linear layer, b 2 represents the bias term of the second linear layer, and z l represents the normalization output.
[0101] For a tongue image analysis method based on a hybrid model architecture as described above, in step S11, the decoded feature vector zl After the input normalization layer, classification information is extracted to obtain the decoded classification vector, which specifically includes the following steps:
[0102] S11.1. Input the decoded feature vector into the normalization layer, expressed as follows:
[0103]
[0104] where μ represents the mean of the feature vector z l and σ represents the standard deviation of the feature vector z l ∈ represents a small value to prevent division by zero, with a value of 10 -6 ;
[0105] S11.2. Input the normalization result into the activation function layer, expressed as follows:
[0106]
[0107] where z represents the input vector, and e z represents the exponential function of z, and e -z represents the exponential function of -z;
[0108] S11.3. Extract classification information from the decoded classification vector through a linear layer, expressed as follows:
[0109] z class = z norm W class + b class
[0110] where W class represents the weight matrix of the linear layer, and b class represents the bias vector.
[0111] In a tongue image analysis method based on a hybrid model architecture as described above, in step S12, the decoded classification vector is input into the detection layer to obtain the rating of metabolic-related fatty liver disease corresponding to the tongue image, expressed as follows:
[0112]
[0113] where z i represents the i-th element in the input vector z.
[0114] The beneficial effects of the present invention are:
[0115] (1) In the present invention, the global information of the tongue image can be captured more meticulously, and the grasp of the macroscopic and microscopic features of the tongue image can be improved through precise analysis, thereby enhancing the comprehensiveness and accuracy of tongue image analysis;
[0116] (2) In the present invention, when facing complex image features, it has stronger performance capabilities, can effectively process complex textures, morphological changes, and heterogeneity features, making its application in disease diagnosis more reliable and accurate;
[0117] (3) In the present invention, for large-scale data sets, it demonstrates excellent representation and learning capabilities. Through the optimization of the deep learning architecture, it realizes the comprehensive learning and processing of complex tongue image information in a big data environment, and improves the adaptability and generalization ability of the model. Overall, this analysis method has significant advantages in the ability to capture details and global information, can better serve tongue image analysis and disease screening, and provide more scientific and accurate support for clinical diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0118] Figure 1 is the overall flow schematic diagram of the present invention;
[0119] Figure 2 is the flow schematic diagram of the improved backbone network in the embodiment of the present invention;
[0120] Figure 3 is the flow schematic diagram of mask feature extraction in the embodiment of the present invention;
[0121] Figure 4 is the flow schematic diagram of the decoder in the tongue image segmentation stage in the embodiment of the present invention;
[0122] Figure 5 is the flow schematic diagram of feature vector encoding in the embodiment of the present invention;
[0123] Figure 6 is the flow schematic diagram of the judgment layer in the tongue image judgment stage in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0124] A tongue image analysis method based on a hybrid model architecture provided in this embodiment is used to analyze the tongue image of a patient to determine whether the patient has metabolic associated fatty liver disease. The method in this embodiment is divided into three main stages, namely the tongue image segmentation stage, the tongue image detection stage, and the disease judgment stage, as Figure 1 shown, and specifically includes the following steps:
[0125] (1) Tongue image segmentation stage:
[0126] S1. Construct a hybrid model architecture, obtain the tongue image of the patient, form a tongue image data set, and divide the training set and the test set in a ratio of 8:2; input the tongue image of the patient in the tongue image data set into the improved backbone network for feature extraction.
[0127] As Figure 2As shown in the figure, in the improved backbone network, the tongue image of the patient passes through convolutional kernels with kernel sizes of 1 / 32*scales, 1 / 16*scales, 1 / 8*scales, and 1 / 4*scales, and strides of 2, 1, 1 / 2, and 1 / 4 respectively from top to bottom to obtain an image feature matrix with a four-layer pyramid structure, where scales represents the pixel size of the image.
[0128] S2. Input the extracted features into the encoder for encoding to obtain the encoded feature vector, which specifically includes the following sub-steps:
[0129] S2.1. Divide the image feature matrix into blocks and embed the position encoding, which is expressed as follows:
[0130] Emb i =Linear(Patch i )
[0131] z l-1 =Emb+PE
[0132] Among them, Emb i represents the calculation result of the linear layer for each small block of the image, Patch i represents the i-th non-overlapping matrix block, z l-1 represents the calculation result of the linear layer with the position encoding added, and PE represents the position encoding.
[0133] S2.2. Pass the non-overlapping feature matrix with the position encoding through the multi-head self-attention mechanism, which is expressed as follows:
[0134] Q=z l-1 W Q
[0135] K=z l-1 W K
[0136] V=z l-1 W V
[0137]
[0138] MHA(z l-1 )=concat(head 1 ,head 2 ,...,head n )W O
[0139] z′ l =LayerNorm(z l-1 +MHA(z l-1 ))
[0140] Among them, Q, K, and V respectively represent the query matrix Query, the key matrix Key, and the value matrix Value, and W Q , W K , W V respectively represent the weight matrices of the query matrix Query, the key matrix Key, and the value matrix Value. l represents the current self-attention layer number, and d k represents the dimension of the K vector, and W O represents the linear transformation weight matrix for the multi-head output. z′ l represents the normalized output of the l-th layer.
[0141] S2.3. Pass the calculation result obtained in the previous step through the feed-forward neural network layer to obtain the finally encoded feature vector, which is expressed as follows:
[0142] FFN(z′ l ) = ReLU(z′ l W 1 + b 1 )W 2 + b 2
[0143] z l = LayerNorm(z′ l + FFN(z′ l ))
[0144] Among them, FFN(z′ l ) represents the output of the feed-forward neural network for z′ l . W 1 represents the weight matrix of the first linear layer, b 1 represents the bias term of the first linear layer, W 2 represents the weight matrix of the second linear layer, b 2 represents the bias term of the second linear layer, and z l represents the output result of the l-th layer's residual connection and normalization.
[0145] S3. As Figure 3 shown, extract features from the pre-annotated mask through the convolutional layer, and combine the extracted mask image features with the encoded feature vector obtained in the previous step to obtain the combined feature vector; combine the Sobel operator and the convolutional layer to extract features for the mask, which specifically includes the following sub-steps:
[0146] S3.1. Obtain the mask image features through the convolutional layer, which is expressed as follows:
[0147] Z l = Conv(P, W 11 ) + b 11
[0148] I(x, y) = ReLU(BatchNorm(Z l ))
[0149] Where P represents the pre-annotated masked image, W 11 represents the convolutional kernel weight, b 11 represents the bias term, and Z l represents the initial feature map.
[0150] S3.2. Obtain the masked edge features through the Sobel operator:
[0151]
[0152] Where I(x, y) represents the input mask; the edge information in the horizontal and vertical directions is obtained through calculation.
[0153] S3.3. Calculate to merge G x with G y into the edge intensity G to represent the overall intensity of the mask at the edge:
[0154]
[0155] S3.4. Concatenate the masked image features and the masked edge features, expressed as follows:
[0156]
[0157] Where I(x, y) represents the input masked image features, G(x, y) represents the overall intensity of the mask at the edge, and X represents the masked features.
[0158] S3.5. Concatenate the masked feature X with the picture encoding feature z l to obtain a new feature vector, expressed as follows:
[0159] F low-level = concat(X, z l )
[0160] Where F low-level represents the new combined feature.
[0161] S4. Input the combined feature vector into the decoder for decoding to obtain the finally segmented tongue image. The decoder includes a Transformer architecture decoder and a pixel decoder.
[0162] The merged features are used by the decoder of the Transformer architecture to judge the categories of the segmented pixel groups. At the same time, the decoding result of the Transformer architecture decoder is multiplied by the pixel decoder, and then the mask probability is judged to finally obtain the segmentation result, as Figure 4 shown, specifically including the following steps:
[0163] S4.1. Decode the combined feature vector through the decoder of the Transformer architecture, and obtain the mask classification result through the mask loss function. The self-attention mechanism has the same structure as the multi-head self-attention mechanism in step S2.2 except for the output layer. In this step, the output layer of the self-attention mechanism uses the softmax function to obtain the probability distribution, expressed as follows:
[0164] y t_mask = softmax(z L W out + b out )
[0165] where z L represents the output of the last layer of the decoder, W out represents the weight of the output layer, and b out represents the bias term of the output layer.
[0166] S4.2. Decode the combined feature vector through the pixel decoder, expressed as follows:
[0167]
[0168] F fused = F low-level + F upsampled
[0169] F conv = ReLU(Conv(F fused , W)+ b)
[0170] y t_pixel = softmax(Conv 1*1 (F conv , W out )+ b out )
[0171] where F low-level (x i , y i ) represents the value of adjacent pixels in the comprehensive feature, w i,j represents the interpolation weight, F low-level represents the combined feature vector output in step S3, W represents the convolution kernel, b represents the bias term, and Conv 1*1Denote the convolution calculation of the 1*1 convolution kernel, W out Denote the convolution kernel of the output layer, b out Denote the bias term of the output layer.
[0172] S4.3. Multiply the mask prediction result by the pixel prediction result to obtain the segmentation result, expressed as follows:
[0173] y seg = y t_mask ⊙ y t_pixel
[0174] where y t_mask denotes the mask prediction result, and y t_pixel denotes the pixel prediction result.
[0175] (2) Tongue image detection stage, as Figure 5 shown:
[0176] S5. Scale the segmented tongue image to 224*224 pixels and keep the original 3 color channels.
[0177] S6. Input the scaled tongue image into the block encoding layer for encoding to obtain the encoded feature vector, expressed as follows:
[0178] Z l = Conv(X, W 1 ) + b 1
[0179] I(x, y) = ReLU(BatchNorm(Z l ))
[0180] where W 1 denotes the convolution kernel weight, b 1 denotes the bias term, and Z l denotes the initial image.
[0181] S7. Concatenate the pre-annotated category with the encoded feature vector obtained in the previous step to obtain the feature vector with category information, expressed as follows:
[0182] z concat = Concat(z encoded , y label )
[0183] where z encoded denotes the encoded feature vector, and y label denotes the corresponding label.
[0184] S8. Encode the location information and add it to the feature vector with category information. Then, complete regularization through the DropPath layer while keeping the vector dimension unchanged, expressed as follows:
[0185] z combined =z concat +z pos
[0186] z dropath =Dropath(z combined )
[0187] where z pos represents the location vector, z combined represents the fusion of the encoded feature vector and the location vector, and z dropath represents the output of the DropPath layer.
[0188] S9. Input the feature vector obtained in the previous step into the encoder for encoding to obtain the encoded feature vector, which specifically includes the following sub-steps:
[0189] S9.1. Divide the image feature matrix into blocks and embed the position encoding, expressed as follows:
[0190] Emb i =Linear(Patch i )
[0191] z l-1 =Emb+PE
[0192] where Emb i is the calculation result of the linear layer for each small block of the image, Patch i represents the i-th non-overlapping matrix block, and PE represents the position encoding.
[0193] S9.2. Pass the non-overlapping feature matrix with position encoding through the sparse multi-head self-attention mechanism, expressed as follows:
[0194] Q=z l-1 W Q
[0195] K=z l-1 W K
[0196] V=z l-1 W V
[0197]
[0198] MHA(z l-1 )=concat(head 1 , head2 ,..., head n )W O
[0199] z' l = LayerNorm(z l-1 + MHA(z l-1 ))
[0200] Among them, Q, K, and V respectively represent the query matrix Query, the key matrix Key, and the value matrix Value, W Q , W K , W V respectively represent the weight matrices of the query matrix Query, the key matrix Key, and the value matrix Value, l represents the current self-attention layer number, d k represents the dimension of the K vector, W O represents the linear transformation weight matrix of the multi-head output, z l-1 represents the calculation result of the linear layer with positional encoding added, z' l represents the output of the normalization layer.
[0201] S9.3. Input z' l into the decoder, and decode the feature vector through the sparse attention layer and the feed-forward layer, expressed as follows:
[0202] Q' = z' l W Q′
[0203] K' = z' l W K′
[0204] V' = z' l W V′
[0205]
[0206] Among them, Q', K', and V' respectively represent the query matrix Query', the key matrix Key', and the value matrix Value', W Q′ , W K′ , W V′ respectively represent the weight matrices of the query matrix Query', the key matrix Key', and the value matrix Value', l represents the current self-attention layer number, d k′ represents the dimension of the K' vector, and M is the sparse mask matrix.
[0207] The sparse mask matrix M is composed of a diagonal matrix and is expressed as follows:
[0208]
[0209] S9.4. Obtain the finally encoded feature vector by passing the calculation result of step S9.2 through a feed-forward neural network layer, expressed as follows:
[0210] FFN(z′ l ) = ReLU(z′ l W 1 +b 1 )W 2 +b 2
[0211] z l = LayerNorm(z′ l +FFN(z′ l ))
[0212] where z′ l represents the output of the normalization layer in step S9.2, W 1 represents the weight matrix of the first linear layer, b 1 represents the bias term of the first linear layer, W 2 represents the weight matrix of the second linear layer, b 2 represents the bias term of the second linear layer, and z l represents the normalized output.
[0213] S10. Input the encoded feature vector into the decoder for decoding to obtain the decoded feature vector. The decoder is the same as the decoder of the transformer architecture in step S4.
[0214] (3) Tongue image judgment stage:
[0215] S11. Input the decoded feature vector z l into the normalization layer to extract classification information and obtain the decoded classification vector, as Figure 6 shown, specifically including the following sub-steps:
[0216] S11.1. Input the decoded feature vector into the normalization layer, expressed as follows:
[0217]
[0218] where μ represents the mean of the feature vector z l , σ represents the standard deviation of the feature vector z l , and ∈ represents a small value to prevent division by zero, with a value of 10 -6 .
[0219] S11.2. Input the normalization result into the activation function layer, expressed as follows:
[0220]
[0221] Among them, z represents the input vector, and e z represents the exponential function of z, and e -z represents the exponential function of -z.
[0222] S11.3. Extract classification information from the decoded classification vector through a linear layer, expressed as follows:
[0223] z class = z norm W class + b class
[0224] Among them, W class represents the weight matrix of the linear layer, and b class represents the bias vector.
[0225] S12. Input the decoded classification vector into the detection layer to obtain the rating of the metabolic-related fatty liver disease corresponding to the tongue image, expressed as follows:
[0226]
[0227] Among them, z i represents the i-th element in the input vector z, and the ratings include mild fatty liver, moderate fatty liver, and severe fatty liver.
[0228] In this embodiment, taking the data of the physical examination center of Changhai Hospital as the dataset, comparing this embodiment with U-Net, DeepLabV3+, and ResNet-50, the results are shown in Table 1 below. It can be seen from Table 1 that this embodiment performs excellently in the tongue image segmentation task, and key indicators such as the Dice coefficient and IoU are better than other comparison models, fully demonstrating its strong ability in processing fine-grained medical image tasks; this indicates that this embodiment can capture and segment complex details in the tongue image more accurately, improving the overall segmentation effect.
[0229] At the same time, DeepLabV3+ also performs well in the segmentation performance, and its results are close to this embodiment, showing its significant advantages in capturing the tongue image edges and local details. However, the performance of ResNet-50 in the segmentation task is relatively poor, probably because it is mainly optimized for classification tasks and has certain limitations in fine-grained image segmentation; therefore, this embodiment, by fusing multi-scale feature extraction and deep learning techniques, is particularly suitable for high-precision medical image tasks such as tongue image segmentation, showing stronger segmentation effects and detail capture capabilities, providing a solid foundation for subsequent disease analysis and diagnosis.
[0230] Table 1
[0231]
[0232] As shown in Table 2 below, this embodiment performs excellently in the disease judgment task, and its key indicators such as accuracy, F1 score, and AUC significantly exceed those of other comparison models, fully demonstrating its outstanding effect in combining tongue image features with disease judgment; in contrast, ResNet-50 performs relatively well in the classification task, but its overall performance is still slightly inferior to this embodiment, indicating that although ResNet-50 has good classification ability, it still has deficiencies in the specific analysis of the association between tongue images and diseases.
[0233] U-Net and DeepLabV3+ perform excellently in the segmentation task, especially having advantages in capturing image boundaries and details. However, since they are not optimized for disease classification, their performance in the classification task is relatively poor; these results indicate that by integrating multi-scale feature extraction and deep learning techniques, this embodiment is more suitable for the disease judgment task, can more comprehensively utilize tongue image information for accurate classification, and provides stronger support and guarantee for clinical diagnosis.
[0234] Table 2
[0235]
[0236] To address the complexity challenges in the field of tongue image analysis, this embodiment proposes a tongue image analysis method based on a hybrid model architecture, which combines the advantages of convolutional neural networks (CNNs) in local feature extraction with the sensitivity of vision transformers (ViTs) in capturing global information to form a more efficient tongue image processing architecture.
[0237] Specifically, CNNs deeply excavate the detailed features in tongue images, such as color, texture, cracks, etc., through layer-by-layer convolution operations, enabling the model to identify subtle changes and local differences; at the same time, the global attention mechanism of ViTs can effectively capture the macroscopic structure and overall pattern in tongue images, making up for the deficiencies of traditional models in processing global information; through this deep learning strategy of combining local and global information, this hybrid model can automatically learn complex disease association features and improve the ability to identify and analyze the personalized features of different patients; the method of this embodiment not only significantly improves the accuracy and reliability of tongue image analysis, but also provides more accurate support for clinical diagnosis, and is expected to play an important role in applications such as medical image analysis, health assessment, and early disease screening.
[0238] In addition to the above embodiments, the present invention may also have other implementation manners. All technical solutions formed by equivalent replacement or equivalent transformation fall within the protection scope required by the present invention.
Claims
1. A tongue image analysis method based on a hybrid model architecture, characterized in that: The following steps are involved: S1. Build a hybrid model architecture and obtain patient tongue images to form a tongue image dataset; input the patient tongue image images in the tongue image dataset into the improved backbone network for feature extraction; S2, input the extracted features into the encoder for encoding to obtain the encoded feature vector; S3, extracting features from the pre-annotated mask through a convolutional layer, and combining the extracted mask image features with the encoded feature vector obtained in the previous step to obtain a combined feature vector; S4, inputting the combined feature vector into a decoder for decoding to obtain the final segmented tongue image, the decoder including a transformer architecture decoder and a pixel decoder; S5, scaling the segmented tongue image to 224*224 pixels, and keeping the original three color channels; S6, inputting the scaled tongue image into the block coding layer for encoding to obtain an encoded feature vector; S7, connecting the pre-marked category with the encoded feature vector obtained in the previous step to obtain a feature vector with category information; S8, encode the position information and add it to the feature vector with category information, and then complete the regularization through the DropPath layer to keep the vector dimension unchanged; S9, inputting the feature vector obtained in the previous step into the encoder for encoding to obtain an encoded feature vector; S10, input the encoded feature vector into the decoder for decoding to obtain a decoded feature vector, and the decoder is the same as the transformer architecture decoder in step S4; S11, input the decoded feature vector into the normalization layer and extract the classification information to obtain a decoded classification vector; S12. Input the decoded classification vector into the detection layer to obtain the rating of metabolism-related fatty liver disease corresponding to the tongue image, which includes mild fatty liver, moderate fatty liver and severe fatty liver.
2. The tongue image analysis method based on a hybrid model architecture according to claim 1, characterized in that: In the improved backbone network of step S1, the patient's tongue image is respectively convolved through convolution kernels with kernel sizes of 1 / 32*scales, 1 / 16*scales, 1 / 8*scales, and 1 / 4*scales, and step sizes of 2, 1, 1 / 2, and 1 / 4, to obtain an image feature matrix with a four-layer pyramid structure from top to bottom, where scales represents the image pixel size.
3. The tongue image analysis method based on a hybrid model architecture according to claim 2, characterized in that: The step S2 specifically includes the following sub-steps: S2.
1. Divide the image feature matrix into blocks and embed the position code, which is expressed as follows: Emb i =Linear(Patch i ) z l-1 =Pack+PE Among them, Emb i Patch represents the linear layer calculation result of each small block of the image. i represents the i-th non-overlapping matrix block, z l-1 It represents the calculation result of the linear layer with position encoding added, and PE represents position encoding; S2.2, the non-overlapping feature matrix with position encoding is expressed as follows through the multi-head self-attention mechanism: Q=z l-1 W Q K=z l-1 IN K V=z l-1 W V MHA(z l-1 )=concat(head1,head2,...,head n )W O With' l =LayerNorm(from l-1 +MHA(from l-1 )) Among them, Q, K, and V represent the query matrix Query, the key matrix Key, and the value matrix Value, respectively. Q , W K , W V Represents the weight matrices of the query matrix Query, the key matrix Key, and the value matrix Value, l represents the current number of self-attention layers, d k represents the dimension of K vector, W O Represents the linear transformation weight matrix of multi-head output, z′ l represents the normalized output of the lth layer; S2.
3. Pass the calculation result obtained in the previous step through the feedforward neural network layer to obtain the final encoded feature vector, which is expressed as follows: <h2 style=";text-align:left;direction:ltr">FFN(z′<h2 style=";text-align:left;direction:ltr"> l <h2 style=";text-align:left;direction:ltr"> )=ReLU(z′<h2 style=";text-align:left;direction:ltr"> l <h2 style=";text-align:left;direction:ltr"> W1+b1)W2+b2 z l =LayerNorm(z′ l +FFN(z′ l )) Among them, FFN(z′ l ) represents z′ l The output of the feedforward neural network, W1 represents the weight matrix of the first linear layer, b1 represents the bias term of the first linear layer, W2 represents the weight matrix of the second linear layer, b2 represents the bias term of the second linear layer, z l Represents the output result of the lth layer residual connection and normalization.
4. The tongue image analysis method based on a hybrid model architecture according to claim 3, characterized in that: In step S3, the Sobel operator and the convolution layer are combined to extract features from the mask, which specifically includes the following sub-steps: S3.
1. The mask image features are obtained through the convolution layer, which is expressed as follows: Z l =Conv(P,W 11 )+b 11 I(x,y)=ReLU(BatchNorm(Z l )) Among them, P represents the pre-labeled mask image, W 11 represents the convolution kernel weight, b 11 represents the bias term, Z l represents the initial feature map; S3.
2. Obtain mask edge features through Sobel operator: Where I(x, y) represents the input mask; edge information in the horizontal and vertical directions is obtained by calculation; S3.3, calculate G by the following formula x With G y Merged into edge strength G, which is used to express the overall strength of the mask at the edge: S3.
4. Concatenate the mask image features with the mask edge features, as shown below: Among them, I(x, y) represents the input mask image feature, G(x, y) represents the overall strength of the mask at the edge, and X represents the mask feature; S3.
5. Combine the mask feature X with the image encoding feature z l After concatenation, a new feature vector is obtained, which is expressed as follows: F low-level =concat(X,z l ) Among them, F low-level Indicates new comprehensive features.
5. The tongue image analysis method based on a hybrid model architecture according to claim 4, characterized in that: The step S4 specifically includes the following sub-steps: S4.
1. Decode the combined feature vector through the transformer architecture decoder, and obtain the mask classification result through the mask loss function. The self-attention mechanism is the same as the multi-head self-attention mechanism in step S2.2 except for the output layer. In this step, the output layer of the self-attention mechanism uses the softmax function to obtain the probability distribution, which is expressed as follows: y t_mask =softmax(z L W out +b out ) Among them, z L represents the output of the last layer decoder, W out represents the weight of the output layer, b out Represents the bias term of the output layer; S4.2, the combined feature vector is decoded by a pixel decoder, and is expressed as follows: F fused =F low-level +F upsampled F conv =ReLU(Conv(F fused ,W)+b) y t_pixel =softmax(Conv 1*1 (F conv ,W out )+b out ) Among them, F low-level (x i ,y i ) represents the value of the neighboring pixels in the comprehensive feature, w i,j represents the interpolation weight, F low-level represents the combined feature vector output in step S3, W represents the convolution kernel, b represents the bias term, Conv 1*1 Represents the convolution calculation of 1*1 convolution kernel, W out represents the output layer convolution kernel, b out Represents the output layer bias term; S4.3, cross-multiply the mask prediction result and the pixel prediction result to obtain the segmentation result, which is expressed as follows: and seg =and t_mask ⊙and t_pixel Among them, y t_mask Represents the mask prediction result, y t_pixel Represents the pixel prediction result.
6. The tongue image analysis method based on a hybrid model architecture according to claim 5, characterized in that: In step S6, the scaled tongue image is input into the block coding layer for coding to obtain the encoded feature vector, which is expressed as follows: Z l =Conv(X,W1)+b1 I(x,y)=ReLU(BatchNorm(Z l )) Among them, W1 represents the convolution kernel weight, b1 represents the bias term, and Z l Indicates the initial image.
7. The tongue image analysis method based on a hybrid model architecture according to claim 6, characterized in that: In step S7 to step S8, the category feature and the position feature are fused into the feature vector to obtain a feature vector with complete information, which is expressed as follows: from concat =Concat(from encoded ,y label ) Among them, z encoded represents the encoded feature vector, y label Indicates the corresponding label; The position information is encoded and added to the feature vector with category information, and then regularized through the DropPath layer to keep the vector dimension unchanged, as shown below: With combined =Z concat +Z pos With dropath =Dropath(from combined ) Among them, z pos represents the position vector, z combined Indicates the fusion of the encoded feature vector and the position vector, z dropath Represents the output of the Dropat layer.
8. The tongue image analysis method based on a hybrid model architecture according to claim 7, characterized in that: The step S9 specifically includes the following sub-steps: S9.
1. Divide the image feature matrix into blocks and embed the position code, which is expressed as follows: Emb i =Linear(Patch i ) z l-1 =Pack+PE Among them, Emb i The linear layer calculation results for each small patch of the image, Patch i represents the i-th non-overlapping matrix block, PE represents the position encoding; S9.
2. The non-overlapping feature matrix with position encoding is expressed as follows through the sparse multi-head self-attention mechanism: Q=z l-1 W Q K=z l-1 IN K V=z l-1 W V MHA(z l-1 )=concat(head1,head2,...,head n )W O With' l =LayerNorm(from l-1 +MHA(from l-1 )) Among them, Q, K, and V represent the query matrix Query, the key matrix Key, and the value matrix Value respectively. Q , W K , W V Represents the weight matrices of query matrix Query, key matrix Key, and value matrix Value respectively, l represents the current number of self-attention layers, d k represents the dimension of K vector, W O Represents the linear transformation weight matrix of multi-head output, z l-1 represents the calculation result of the linear layer with position encoding added, z′ l Represents the normalization layer output; S9.
3. z′ l Input decoder, decode the feature vector through sparse attention layer and feedforward layer, expressed as follows: Q′=z′ l W Q′ K′=z′ l IN K′ V′=z′ l W V′ Among them, Q′, K′, and V′ represent the query matrix Query′, the key matrix Key′, and the value matrix Value′, respectively. Q′ , W K′ , W V′ Respectively represent the weight matrices of query matrix Query′, key matrix Key′, and value matrix Value′, l represents the current number of self-attention layers, d k′ represents the dimension of K′ vector, M is the sparse mask matrix; S9.4, pass the calculation result of step S9.2 through the feedforward neural network layer to obtain the final encoded feature vector, which is expressed as follows: <h2 style=";text-align:left;direction:ltr">FFN(z′<h2 style=";text-align:left;direction:ltr"> l <h2 style=";text-align:left;direction:ltr"> )=ReLU(z′<h2 style=";text-align:left;direction:ltr"> l <h2 style=";text-align:left;direction:ltr"> W1+b1)W2+b2 z l =LayerNorm(z′ l +FFN(z′ l )) Among them, z′ l represents the normalized layer output in step S9.2, W1 represents the weight matrix of the first linear layer, b1 represents the bias term of the first linear layer, W2 represents the weight matrix of the second linear layer, b2 represents the bias term of the second linear layer, z l Represents normalized output.
9. The tongue image analysis method based on a hybrid model architecture according to claim 8, characterized in that: In step S11, the decoded feature vector z l After inputting the normalization layer, the classification information is extracted to obtain the decoded classification vector, which specifically includes the following sub-steps: S11.
1. Input the decoded feature vector into the normalization layer, expressed as follows: Where μ represents the eigenvector z l The mean of σ represents the eigenvector z l The standard deviation of ∈ indicates a small value to prevent division by zero, and its value is 10 -6 ; S11.
2. Input the normalized result into the activation function layer, expressed as follows: Where z represents the input vector, e z represents the exponential function of z, e -z represents the exponential function of -z; S11.
3. Extract classification information from the decoded classification vector through a linear layer, as shown below: With class =from nnorm IN class +b class Among them, W class represents the weight matrix of the linear layer, b class Represents the bias vector.
10. The tongue image analysis method based on a hybrid model architecture according to claim 9, characterized in that: In step S12, the decoded classification vector is input into the detection layer to obtain the rating of metabolic-related fatty liver disease corresponding to the tongue image, which is expressed as follows: Among them, z i Represents the i-th element in the input vector z.
Citation Information
Patent Citations
Efficient modeling three-dimensional medical image segmentation method based on mask supervision strategy
CN117333497A
Lumbar vertebra CT image segmentation and identification method based on multi-channel attention
CN118781336A
Sparse code multiple access encoding and decoding system based on generative adversarial network
WO2024016424A1