Method for guiding double-flow graph neural network to identify multi-label tongue picture based on hierarchical knowledge

Through the dual-flow graph neural network guided by hierarchical knowledge, the visual representation of tongue images is optimized, which solves the problem of neglecting label relationships in multi-label recognition of tongue images and improves the accuracy of tongue images classification.

CN120298743APending Publication Date: 2025-07-11HENAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510178535.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify multi-label attributes in tongue images, ignoring the relationship between labels, resulting in insufficient recognition accuracy.

Method used

A dual-stream graph neural network based on hierarchical knowledge is used to guide visual representations through hierarchical semantics, enhance the interaction between categories, optimize the expression of visual representations, and extract tongue image features using the ResNet-101 model, and generate category representations by iteratively propagating node information.

Benefits of technology

It improves the accuracy of multi-label recognition of tongue icons and enhances the discrimination ability of tongue icon classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298743A_ABST
    Figure CN120298743A_ABST
Patent Text Reader

Abstract

A method for guiding a double-flow graph neural network to identify a multi-label tongue picture based on hierarchical knowledge comprises the following steps: acquiring semantic embedding vectors of all tongue picture categories from a tongue picture corpus, and optimizing the semantic embedding vectors to obtain hierarchical semantics; constructing a graph # imgabs0 # according to an association relationship between hierarchical semantics and tongue picture categories, and obtaining category representation; obtaining a tongue picture image, and performing feature extraction on the tongue picture image by using the trained ResNet-101 model to obtain a visual feature map; aligning the visual feature map and the hierarchical semantics in a feature space by using an alignment module to obtain visual representation guided by the hierarchical semantics; constructing a graph # imgabs1 # based on visual representation, and obtaining an output vector of each category; the category representation and the output vector of each category are combined, and a tongue picture classification result is obtained.Visual representation is obtained through hierarchical semantic guidance, interaction between the categories is enhanced, and expression of the visual representation is optimized; and the accuracy of identification results is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-label tongue image recognition, and specifically to a method for recognizing multi-label tongue images based on a hierarchical knowledge-guided two-stream graph neural network. Background Art

[0002] Tongue diagnosis is one of the most important diagnostic methods in traditional Chinese medicine, which can provide an effective and non-invasive way to assist in evaluating the physical condition of patients. Traditional tongue diagnosis largely relies on the observation skills and experience of Chinese medicine practitioners, and environmental factors and factors of the examinees lead to unstable and inaccurate tongue diagnosis results. Therefore, it is crucial to establish an objective and accurate tongue image classification model.

[0003] In the prior art, certain achievements have been made in recognizing tongue images using deep learning. For example, the recognition of teeth-marked tongue images is modeled as multi-instance learning, the concave information in the tongue image is used to locate the suspected teeth-marked area, the pre-trained VGG-16 convolutional neural network is used to extract feature vectors from these suspected areas, and a multi-instance support vector machine is used for classification to obtain the classification effect of teeth-marked tongue images; a CNNs model is used to extract tongue color features, and a center loss function is added during the training stage to enhance the discriminative ability of the four tongue color features; input integration is used to connect the original image and the feature map as the input of the network, and then feature integration is used to splice the features obtained from the segmentation branch to the classification branch, so that the segmented features assist the classification features to classify the size and shape of the tongue; a single-shot multi-box detector is used to locate the crack area, an improved network is used to extract global and local features, and these features are fused to complete the classification of cracked tongue using a gradient boosting decision tree. However, the above methods are all for single-label binary classification or single-label multi-classification of tongue images.

[0004] In actual work, doctors often need to combine multiple tongue image features to judge the condition of patients, and tongue images naturally have multi-label attributes. Tongue images usually contain multiple features, such as teeth marks, cracks, peeled tongue coating, prickles, etc. Therefore, in the research of tongue images, multi-label recognition of tongue images is a more fundamental and practical task. At present, a simple multi-label classification of tongue images has been completed using the method of a multi-CNN fusion model. However, the method it uses is to fit a single-label classification model to a multi-label task, which ignores the relationship between labels, and the number of predicted labels will increase exponentially with the increase in the number of categories. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for recognizing multi-label tongue images based on a hierarchical knowledge-guided two-stream graph neural network, to obtain visual representations through hierarchical semantic guidance, enhance the interaction between categories, and optimize the expression of visual representations; and to improve the accuracy of recognition results.

[0006] To achieve the above object, the specific solution adopted by the present invention is as follows: A method for identifying multi-label tongue images based on a hierarchical knowledge-guided dual-stream graph neural network, comprising the following steps:

[0007] Obtain the semantic embedding vectors of all tongue image categories from the tongue image corpus, and optimize the semantic embedding vectors by hierarchical operations to obtain hierarchical semantics;

[0008] Construct a graph based on the association relationship between the hierarchical semantics and the tongue image categories And finally obtain the category representation by iteratively propagating node information;

[0009] Obtain a tongue image, use the trained ResNet-101 model to extract features from the tongue image, and obtain a visual feature map;

[0010] Use a hierarchical knowledge-guided alignment module to align the visual feature map and the hierarchical semantics in the feature space to obtain a hierarchical semantics-guided visual representation;

[0011] Construct a graph based on the visual representation And use a gated recurrent update mechanism to propagate messages to obtain an output vector for each category;

[0012] Combine the category representation and the output vector for each category to obtain the tongue image classification result.

[0013] As an optimized scheme of the above method for identifying multi-label tongue images based on a hierarchical knowledge-guided dual-stream graph neural network, the method for obtaining the semantic embedding vectors of all tongue image categories from the tongue image corpus and optimizing the semantic embedding vectors by hierarchical operations includes:

[0014] Use the Glove algorithm to obtain the semantic embedding vectors of all categories;

[0015] Construct a tree-shaped class hierarchy, including multiple leaf nodes, and map the semantic embedding vectors to the leaf node semantic vectors in the class hierarchy;

[0016] Cluster the leaf node semantic vectors upward to obtain superclass layer nodes, and obtain the superclass layer node semantic vectors;

[0017] Decouple the leaf node semantic vectors downward, and layer-by-layer splice the superclass layer node semantic vectors and the leaf node semantic vectors to obtain updated leaf node semantic vectors;

[0018] Map the updated leaf node semantic vectors back to the semantic embedding vectors to obtain hierarchical semantics.

[0019] As another optimization scheme for the above method of identifying multi-label tongue images based on hierarchical knowledge-guided dual-stream graph neural network, the semantic vector of the superclass layer node is:

[0020]

[0021] where l ∈ [1, L - 1] represents the level, and C(T l ) represents the set of child nodes of the parent node T l .

[0022] As another optimization scheme for the above method of identifying multi-label tongue images based on hierarchical knowledge-guided dual-stream graph neural network, the updated semantic vector of the leaf node is:

[0023]

[0024] where represents the weight matrix implemented by the fully connected layer.

[0025] As another optimization scheme for the above method of identifying multi-label tongue images based on hierarchical knowledge-guided dual-stream graph neural network: The method for obtaining the class representation by iteratively propagating node information is:

[0026] Construct a graph using hierarchical semantics and class dependencies

[0027] Initialize the hidden state of each node using hierarchical semantics;

[0028] During the iteration process, each node aggregates the messages of its neighbor nodes, and then updates the hidden state of each node, finally generating the output hidden state;

[0029] Concatenate the output hidden state with the hierarchical semantics to obtain the class representation.

[0030] As another optimization scheme for the above method of identifying multi-label tongue images based on hierarchical knowledge-guided dual-stream graph neural network, the method for obtaining the visual representation includes:

[0031] Use the trained ResNet-101 model to extract features from the tongue image to obtain the visual feature map

[0032] Align the visual feature map with the class representation in the latent space, and use low-rank bilinear pooling to fuse the visual features and class representations based on positions to obtain multiple position features;

[0033] Calculate the importance score of each position feature;

[0034] Normalize the importance scores of all positions;

[0035] Combining visual representations based on all positions, position features, and importance scores.

[0036] As another optimization scheme of the above method for identifying multi-label tongue images based on a hierarchical knowledge-guided two-stream graph neural network, the position features are:

[0037]

[0038] where tanh(·) is the hyperbolic tangent function, denotes learnable parameters, ⊙ denotes element-wise multiplication, and d1 and d2 are the dimensions of the joint embedding and output features respectively.

[0039] As another optimization scheme of the above method for identifying multi-label tongue images based on a hierarchical knowledge-guided two-stream graph neural network, the importance scores are:

[0040]

[0041] where f a (·) is a fully connected function.

[0042] As another optimization scheme of the above method for identifying multi-label tongue images based on a hierarchical knowledge-guided two-stream graph neural network: the output vector for each category is:

[0043]

[0044] where f o () represents the output function, denotes the category representation, and f c denotes the category-related visual feature vector.

[0045] As another optimization scheme of the above method for identifying multi-label tongue images based on a hierarchical knowledge-guided two-stream graph neural network: The ResNet-101 model includes multiple average pooling layers, and the last average pooling layer is 2×2 with a stride of 2.

[0046] Compared with the prior art, the present invention has the following beneficial effects:

[0047] The present invention provides a method for identifying multi-label tongue images based on a hierarchical knowledge-guided dual-stream graph neural network. Semantic embedding vectors of all tongue image categories are obtained from a tongue image corpus, and hierarchical semantics are generated based on the semantic embedding vectors, and category representations are obtained through the hierarchical semantics. A tongue image is acquired and feature extraction is performed. An alignment module is used to align the visual feature map and the hierarchical semantics in the feature space to obtain a hierarchical semantics-guided visual representation, and an output vector for each category is obtained based on the visual representation. The category representation and the output vector for each category are combined to obtain the classification result of the tongue image. That is, the present invention optimizes the visual representation by guiding the category-related visual representation through hierarchical semantics, and improves the accuracy of the recognition method. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is the class hierarchy constructed by the present invention;

[0049] Figure 2 is the framework diagram of the present invention;

[0050] Figure 3 is the upward clustering in the hierarchical operation;

[0051] Figure 4 is the downward decoupling in the hierarchical operation. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The technical solutions of the present invention will be further elaborated in detail below in conjunction with specific embodiments. For parts that are not detailedly recorded and disclosed in the following embodiments of the present invention, they should all be understood as the prior art known or should be known to those skilled in the art.

[0053] Embodiment

[0054] A method for identifying multi-label tongue images based on a hierarchical knowledge-guided dual-stream graph neural network. First, introduce the figure where V represents graph nodes and E represents graph edges. Specifically, assume that the data set covers C categories, and corresponding labels are assigned to each category. V is represented as {v0, v1, v2,..., v C-1}, where the element v c is the feature representation of category c, and E is represented as {e 00 , e 01 ,..., e 0(C-1) ,..., e (C-1)(C-1)}, where e cc′Denote the probability that an object belongs to class C' when there is an object belonging to class c. Calculate the probabilities between all pairs of classes using the sample annotations in the training set, and call this label dependence, that is, the association relationship between classes. There is an obvious internal structure in the classes of tongue images, and there is an obvious co-occurrence relationship between classes. For example, a swollen tongue and teeth marks often appear simultaneously, and this co-occurrence relationship has a certain impact on the recognition result of tongue images.

[0055] The methods for recognizing multi-label tongue images include:

[0056] Obtain the semantic embedding vectors of all tongue image classes from the tongue image corpus, and optimize the semantic embedding vectors using hierarchical operations to obtain hierarchical semantics.

[0057] Specifically include:

[0058] Use the semantic extraction module in the Glove algorithm to obtain the semantic embedding vectors of all classes

[0059] T c = f g (w c ) ;

[0060] where C represents the number of classes, d s represents the semantic embedding dimension; w c represents the semantic word of class c, and f g (·) represents the Glove algorithm.

[0061] Construct a tree-shaped class hierarchy to encode the semantic relationships between different classes. As Figure 1 shown, the class hierarchy includes a leaf layer, a superclass layer, and a top layer. The leaf layer has multiple leaf nodes, and the superclass layer has multiple superclass layer nodes. This class hierarchy can cover all classes. Although different classes do not overlap at the leaf layer, they share nodes at the top layer. Map the semantic embedding vectors to the leaf node semantic vectors in the class hierarchy:

[0062] T L = T c ;

[0063] where c ∈ {1, 2,..., C}, and L represents the level number of the leaf layer in the class hierarchy.

[0064] The semantic vectors of leaf nodes are clustered upward to obtain superclass layer nodes. Specifically, when two categories belong to the same superclass layer, there must be a certain similarity in the feature information of the two categories; when two categories belong to different superclass layers, there must be a certain difference in the feature information of the two categories. These similarity and difference information will be aggregated in the superclass layer nodes and the top layer nodes, and the semantic vectors of the superclass layer nodes will be obtained. The semantic vectors of the superclass layer nodes are the average values of the semantic vectors of their subclasses:

[0065]

[0066] where l ∈ [1, L - 1] represents the level in the class hierarchy, and C(T l ) represents the set of child nodes of the parent node T l .

[0067] The semantic vectors of leaf nodes are decoupled downward, and the semantic vectors of leaf nodes are reconstructed. As Figure 4 shown, each leaf node aggregates information along the hierarchical path, so as to learn the correlation between its own features and the global semantics, that is, the semantic vectors of the superclass layer nodes and the semantic vectors of the leaf nodes are concatenated layer by layer to obtain the updated semantic vectors of the leaf nodes:

[0068]

[0069] where represents the weight matrix implemented by the fully connected layer.

[0070] The update of the semantic vectors of leaf nodes is realized by a two-way strategy of first clustering upward and then decoupling downward. The updated semantic vectors of leaf nodes are mapped back to the semantic embedding vectors to obtain the hierarchical semantics:

[0071]

[0072] Through the above method, the interaction between each category point and other category points is realized.

[0073] Construct a graph based on the hierarchical semantics and the association relationship of categories and propagate the node information through graph iteration. Finally, the category representation is obtained; thus, a more discriminative category representation is generated, providing more powerful guidance for the subsequent recognition of tongue images. Specifically:

[0074] Use the hierarchical semantics to initialize the hidden state of each node; use the hierarchical semantics and label dependencies after hierarchical operations to construct a graph and propagate node messages through graph iteration. During the iteration process, each node has a hidden state and use the hierarchical semantics to initialize the hidden state:

[0075]

[0076] During the iterative process, each node aggregates the messages of its neighbor nodes, and then updates the hidden state of each node, and finally generates an output hidden state; specifically: during the iterative process, each node aggregates the messages from its neighbor nodes:

[0077]

[0078] Promote the message propagation of highly relevant categories in the semantic space, otherwise, suppress the propagation. In this way, the hidden state of each node is updated, and each node aggregates the hierarchical semantics from other categories to learn richer hierarchical semantics and generate an output hidden state

[0079] Concatenate the output hidden state and the hierarchical semantics to obtain a class representation:

[0080]

[0081] where w o is the output function, which and are concatenated and output to w c .

[0082] Obtain a tongue image, use the trained ResNet-101 model to extract features from the tongue image, and obtain a visual feature map;

[0083] Use the hierarchical knowledge-guided alignment module to align the visual feature map and the hierarchical semantics in the feature space to obtain a hierarchical semantics-guided visual representation;

[0084] Specifically:

[0085] Use the trained ResNet-101 model to extract features from the tongue image to obtain a visual feature map:

[0086] f I = f cnn (I);

[0087] where d f , W, and H are the number of channels, width, and height of the visual feature map respectively, and f cnn (·) is the visual encoder implemented by the ResNet-101 model.

[0088] Align the visual feature map and the class representation in the latent space, and use low-rank bilinear pooling to fuse the visual features and class representations based on positions to obtain multiple position features. For each position (w, h), use low-rank bilinear pooling to fuse and Obtain the position feature:

[0089]

[0090] where tanh(·) is the hyperbolic tangent function, and represent learnable parameters, ⊙ represents element-wise multiplication, and d1 and d2 are the dimensions of the joint embedding and the output position feature respectively.

[0091] Calculate the importance score of each position feature:

[0092]

[0093] where f a (·) is the fully connected function.

[0094] To compare the importance scores of different position features, use the softmax function to normalize the importance scores of all position features:

[0095]

[0096] Based on all positions, position features, and importance scores, combine to obtain the visual representation. Specifically, perform a dot product on the importance scores and position features of all positions to obtain the feature vector f c :

[0097]

[0098] Finally, obtain the hierarchical semantic-guided visual representation {f0, f1,..., f C-1} which pays more attention to the semantic regions of each category.

[0099] Construct a graph based on the visual representation and adopt a gated recurrent update mechanism to propagate messages to obtain the output vector of each category;

[0100] Specifically, construct a graph from the visual representation in a manner based on statistical label co-occurrence Each node v c ∈V has a hidden state at the time step Meanwhile, since each leaf node corresponds to the hierarchical semantics of a specific category, and to explore the interaction between the visual representation and the hierarchical semantics, therefore, it is initialized using the feature vector corresponding to its corresponding category:

[0101]

[0102] At the time step t, each node v c will aggregate the messages from its neighbor nodes:

[0103]

[0104] When node v c ′ is highly correlated with node v c the gated recurrent update mechanism encourages message propagation, otherwise it inhibits message propagation between nodes, based on the aggregated feature vector and its hidden state at the previous time step update the hidden state:

[0105]

[0106]

[0107]

[0108]

[0109] where {W z , U z , W r , U r , W h , U h} are learnable parameters, σ(·) is the logistic sigmoid function, tanh(·) is the hyperbolic tangent function, and ⊙ represents element-wise multiplication.

[0110] Each node aggregates messages from other nodes and simultaneously propagates information through the graph to achieve interaction between all feature vectors corresponding to all classes. This process is repeated T f times to generate the final hidden state Each hidden state not only encodes information from class c but also carries context information from other classes. Concatenate the output hidden state with the feature vector f c to obtain the output vector for each class:

[0111]

[0112] where f o () represents the output function used to concatenate and output to O c and f c .

[0113] Combine the class representation and the output vector for each class to obtain the tongue image classification result:

[0114] S = {s0, s1,..., s C-1};

[0115] S = (o c ⊙ w c ) + b f ;

[0116] Wherein, is a learnable parameter.

[0117] In this embodiment, the ResNet-101 model is used as the visual encoder f cnn (·), including multiple average pooling layers, and the last average pooling layer is 2×2 with a stride of 2, and the other average pooling layers remain unchanged. For low-rank bilinear pooling, N, d s , d1 and d2 are set to 2048, 300, 1024, and 1024 respectively. Therefore, f a is implemented by a fully connected layer from 1024 to 1, which maps 1024 feature vectors to a single attention coefficient.

[0118] In this embodiment, the hidden state dimensions of the class representation and the output vector of each class are both 2048, and the number of iterations T f and T s are both 3, and the output vector o c of each class is 2048. Therefore, the output functions f o (·) and w o (·) are implemented by a fully connected layer from 4096 to 2048 followed by the hyperbolic tangent function.

[0119] Optimize the ResNet-101 model. Given a dataset containing M training samples , where I i is the i-th tongue image, and y i = {y i0 , y i1 ,..., y i(c-1)} is the class (label) corresponding to the tongue image I i . If the sample contains the class c, then y ic is assigned 1, otherwise it is assigned 0. Given a tongue image I i , a tongue image classification result s i = {s i0 , s i1 ,..., s i(C-1)} can be obtained, and the cross-entropy is used as the objective loss function:

[0120]

[0121] Wherein, y ic represents the true label of each sample, and s icRepresents the predicted label for each sample.

[0122] Using the above objective loss function for training, first, initialize the parameters of the corresponding layers in f with the parameters of the ResNet 101 model pre-trained on the ImageNet dataset, and then initialize the parameters of other layers in the model. Since the parameters of the lower layers pre-trained on the ImageNet dataset have good generalization on different datasets, the parameters of the first 92 convolutional layers are frozen. This model uses the ADAM algorithm as the optimizer, with a Batch size of 16, momentum of 0.999 and 0.9, and the learning rate is initialized to 10 cnn , when the error reaches a steady state, the learning rate will be reduced by 10 times; during the training process, the size of the tongue image is adjusted to 640×640, and a random number is randomly selected from {640, 576, 512, 384, 320} as the image height and width for random cropping, and the cropped patch is adjusted to 576×576. In the test stage, only the tongue image needs to be adjusted to 640×640, and a central crop of 576×576 is performed for evaluation. -5 For the above description of the disclosed embodiments, those skilled in the art can implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

[0123] For the above description of the disclosed embodiments, those skilled in the art can implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for identifying multi-label tongue images based on a hierarchical knowledge-guided dual-stream graph neural network, characterized in that: It includes the following steps: Obtain the semantic embedding vectors of all tongue image categories from the tongue image corpus, and optimize the semantic embedding vectors through hierarchical operations to obtain hierarchical semantics; Construct a graph based on the association relationship between hierarchical semantics and tongue image categories And finally obtain the category representation by iteratively propagating node information; Obtain a tongue image, and use the trained ResNet-101 model to extract features from the tongue image to obtain a visual feature map; Use the hierarchical knowledge-guided alignment module to align the visual feature map and the hierarchical semantics in the feature space to obtain a hierarchical semantics-guided visual representation; Construct a graph based on visual representations And use a gated recurrent update mechanism to propagate messages to obtain output vectors for each category; Combine the class representation and the output vector of each class to obtain the tongue image classification result.

2. The method for identifying multi-label tongue images based on a hierarchical knowledge-guided two-stream graph neural network according to claim 1, wherein: The method for obtaining the semantic embedding vectors of all tongue image categories from the tongue image corpus and optimizing the semantic embedding vectors through hierarchical operations includes: Use the Glove algorithm to obtain the semantic embedding vectors of all categories; Construct a tree-shaped class hierarchy including multiple leaf nodes, and map the semantic embedding vectors to the leaf node semantic vectors in the class hierarchy; Cluster the leaf node semantic vectors upward to obtain superclass layer nodes and obtain the superclass layer node semantic vectors; Decouple the leaf node semantic vectors downward, and layer by layer splice the superclass layer node semantic vectors and the leaf node semantic vectors to obtain the updated leaf node semantic vectors; Map the updated leaf node semantic vectors back to the semantic embedding vectors to obtain hierarchical semantics.

3. The method for identifying multi-label tongue images based on a hierarchical knowledge-guided dual-stream graph neural network according to claim 2, wherein: The superclass layer node semantic vectors are: Among them, l ∈ [1, L - 1] represents the level, and C(Tl) represents the set of child nodes of the parent node T l ​ 4. The method for identifying multi-label tongue images based on a hierarchical knowledge-guided two-stream graph neural network according to claim 2, wherein: The updated leaf node semantic vectors are: Among them, represents the weight matrix implemented by full connection.

5. The method for identifying multi-label tongue images based on a hierarchical knowledge-guided dual-stream graph neural network according to claim 1, characterized in that: The method for obtaining the class representation by iteratively propagating node information is: Construct a graph using hierarchical semantics and category dependencies Use hierarchical semantics to initialize the hidden state of each node; During the iterative process, each node aggregates the messages of its neighbor nodes, and then updates the hidden state of each node, and finally generates an output hidden state; Splice the output hidden state and the hierarchical semantics to obtain the class representation.

6. The method for identifying multi-label tongue images based on a hierarchical knowledge-guided two-stream graph neural network according to claim 1, characterized in that: The method for obtaining the visual representation includes: Extract features of tongue images using the trained ResNet-101 model to obtain visual feature maps Align the visual feature maps with the class representations in the latent space, and use low-rank bilinear pooling to fuse the visual features and class representations based on positions to obtain multiple position features; Calculate the importance score of each position feature; Normalize the importance scores of all positions; Combine the visual representation based on all positions, position features, and importance scores.

7. The method for identifying multi-label tongue images based on a hierarchical knowledge-guided two-stream graph neural network according to claim 1, characterized in that The position features are: where tanh(·) is the hyperbolic tangent function, denotes learnable parameters, ⊙ denotes element-wise multiplication, and d1 and d2 are the dimensions of the joint embedding and output features, respectively.

8. A method for identifying multi-label tongue images based on a hierarchical knowledge-guided dual-stream graph neural network as described in claim 1, characterized in that: The importance scores are: where f a (·) is a fully connected function.

9. The method for identifying multi-label tongue images based on a hierarchical knowledge-guided dual-stream graph neural network according to claim 1, wherein: The output vector of each class is: Among them, f o () represents the output function, represents the class representation, and f c represents the visual feature vector related to the class.

10. The method for identifying multi-label tongue images based on a hierarchical knowledge-guided dual-stream graph neural network according to claim 1, characterized in that: The ResNet-101 model includes multiple average pooling layers, and the last average pooling layer is 2×2 with a stride of 2.