Tongue surface image multi-label classification method, storage medium, electronic equipment and product
By constructing a multi-label classification model for tongue surfaces and combining multiple modules to process tongue surface images, the problem of unstable diagnosis results in tongue diagnosis is solved, the correlation relationship between tongue image features is realized, and the accuracy and efficiency of tongue diagnosis in traditional Chinese medicine is improved.
Patent Information
- Application Number
- CN202510520378.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-05
AI Technical Summary
In the prior art, tongue diagnosis mainly relies on the clinical experience of doctors. The diagnosis results are greatly affected by environmental and individual differences. The existing research has failed to effectively explore the correlation between multiple labels of tongue images, resulting in insufficient reliability of tongue image diagnosis in traditional Chinese medicine.
By constructing a multi-label classification model for tongue surfaces, combining basic feature extraction, cross-scale feature processing, graph convolution network and fully connected classifier, multi-label classification of tongue surface images is realized, probability values of different tongue surface classification labels are output, and correlation between tongue image features is captured.
It improves the accuracy and efficiency of tongue image diagnosis, provides more reliable Chinese medicine diagnosis suggestions, reduces dependence on artificial experience, and realizes the intelligence and standardization of tongue image diagnosis.
Smart Images

Figure CN120431382A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, storage medium, electronic device, and product for multi-label classification of tongue surface images. Background Art
[0002] Tongue diagnosis, as an important part of visual diagnosis in traditional Chinese medicine, reflects the physiological state and pathological changes of the human body by observing multi-dimensional characteristics such as the tongue shape, tongue texture, color, and thickness of tongue coating.
[0003] Currently, tongue diagnosis primarily relies on the physician's clinical experience and expertise, and its results are easily affected by a variety of factors, including environmental conditions and individual differences. With the development of artificial intelligence technology, existing TCM tongue image research has primarily focused on classifying single features of tongue images, failing to explore the correlations between other tongue features and, consequently, failing to provide reliable tongue diagnosis recommendations for TCM practitioners.
[0004] Therefore, how to provide a technical solution for an accurate multi-label classification method for tongue surface images has become a technical problem that needs to be solved urgently. Summary of the Invention
[0005] The purpose of some embodiments of the present application is to provide a method, storage medium, electronic device and product for multi-label classification of tongue images. Through the technical solutions of the embodiments of the present application, accurate multi-label classification of tongue images can be achieved, and the correlation between labels can be displayed, thereby providing Chinese medicine practitioners with reliable suggestions for tongue diagnosis and improving the efficiency of tongue diagnosis.
[0006] In a first aspect, some embodiments of the present application provide a method for multi-label classification of tongue images, comprising: obtaining a tongue image of a target object; inputting the tongue image into a pre-trained tongue multi-label classification model to obtain a probability value of at least one category of tongue classification labels to which the tongue image belongs; wherein the tongue classification labels include: tongue quality labels, tongue shape labels, tongue coating labels, tooth mark labels, and crack labels; the tongue multi-label classification model is obtained by training a tongue feature model, and the tongue multi-label classification model includes: a basic feature extraction module, a cross-scale feature processing module, a graph convolutional network module, and a fully connected classifier.
[0007] Some embodiments of the present application classify tongue images through a tongue multi-label classification model obtained by training a tongue feature model composed of multiple modules, and obtain the probability value of at least one category of tongue classification label to which the tongue image belongs. This can determine the probability that the tongue image contains different tongue classification labels, achieve accurate multi-label classification of tongue images, and show the existence relationship between labels, thereby providing Chinese medicine with reliability recommendations for tongue diagnosis and improving the efficiency of tongue diagnosis.
[0008] In some embodiments, the tongue image is input into a pre-trained tongue multi-label classification model to obtain the probability value of at least one category of tongue classification labels to which the tongue image belongs, including: extracting the tongue image through the basic feature extraction module to obtain initial tongue features; processing the initial tongue features through the cross-scale feature processing module to obtain multi-scale tongue features; calculating and globally averaging the multi-scale tongue features through the graph convolutional network module to obtain a global feature vector; analyzing the global feature vector through the fully connected classifier to output the probability value of each category of tongue classification labels in the at least one category of tongue classification labels.
[0009] Some embodiments of the present application process tongue images in sequence through multiple cascade modules in a tongue multi-label classification model, and finally output the probability value of each type of tongue classification label, thereby realizing multi-label classification of tongue images and achieving comprehensive classification.
[0010] In some embodiments, the cross-scale feature processing module includes a first encoder and a second encoder; wherein, the processing of the initial tongue surface features by the cross-scale feature processing module to obtain multi-scale tongue surface features includes: processing the initial tongue surface features by the first encoder to obtain local feature information; and downsampling the initial tongue surface features by the second encoder to obtain context image information; and using a multi-head cross-attention mechanism to fuse the local feature information and the context image information to obtain the multi-scale tongue surface features.
[0011] Some embodiments of the present application use a cross-scale feature processing module to perform dual-channel feature extraction on the initial tongue surface features and then fuse them to obtain multi-scale tongue surface features, providing data support for subsequent accurate classification.
[0012] In some embodiments, the multi-scale tongue surface features are calculated and globally averaged pooled by the graph convolutional network module to obtain a global feature vector, including: constructing an adjacency matrix based on the similarity between feature nodes in the multi-scale tongue surface features; and performing a global average pooling operation on the adjacency matrix to obtain the global feature vector.
[0013] Some embodiments of the present application process multi-scale tongue surface features through a graph convolutional network module, output a global feature vector, and provide data support for subsequent accurate classification.
[0014] In some embodiments, before inputting the tongue surface image into a pre-trained tongue surface multi-label classification model, the method further includes: constructing a tongue surface sample dataset, wherein the tongue surface sample dataset includes: multiple tongue surface sample images and multiple labels for each of the multiple tongue surface sample images; using the tongue surface sample dataset to train the tongue surface feature model to obtain the tongue surface multi-label classification model.
[0015] Some embodiments of the present application train the tongue feature model by constructing a rich tongue sample dataset to obtain a tongue multi-label classification model, thereby improving the generalization ability and robustness of the model.
[0016] In some embodiments, the multiple tongue surface sample images are obtained by the following method: collecting multiple initial tongue surface images; annotating the multiple initial tongue surface images to obtain tongue surface annotated images; segmenting the tongue surface annotated images using a segmentation model to obtain tongue surface segmented images; normalizing and data enhancing the tongue surface segmented images to obtain the multiple tongue surface sample images.
[0017] Some embodiments of the present application process multiple initial tongue surface images to obtain multiple tongue surface sample images, which can improve the diversity of sample data and the accuracy of model training.
[0018] In some embodiments, the tongue surface feature model is trained using the tongue surface sample dataset to obtain the tongue surface multi-label classification model, including: training the tongue surface feature model using the tongue surface sample dataset to obtain the model to be verified; confirming that the value of the evaluation index of the model to be verified meets the preset conditions, then using the model to be verified as the tongue surface multi-label classification model; wherein the evaluation index includes: label-level index and model-wide index; the label-level index includes: label accuracy, classification accuracy, recall rate or harmony index F1 value; the model-wide index includes: overall accuracy mean, overall accuracy value, overall recall rate or overall F1 value.
[0019] Some embodiments of the present application verify the model to be verified through a variety of different evaluation indicators to obtain a tongue surface multi-label classification model, which can obtain a tongue surface multi-label classification model with higher accuracy and ensure the accuracy of tongue surface label classification.
[0020] In a second aspect, some embodiments of the present application provide a device for multi-label classification of tongue images, comprising: an image acquisition module for acquiring a tongue image of a target object; a label classification module for inputting the tongue image into a pre-trained tongue multi-label classification model to obtain the probability value of at least one category of tongue classification labels to which the tongue image belongs; wherein the tongue classification labels include: tongue quality labels, tongue shape labels, tongue coating labels, tooth mark labels, and crack labels; the tongue multi-label classification model is obtained by training a tongue feature model, and the tongue multi-label classification model includes: a basic feature extraction module, a cross-scale feature processing module, a graph convolutional network module, and a fully connected classifier.
[0021] In a third aspect, some embodiments of the present application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the method described in any embodiment of the first aspect.
[0022] In a fourth aspect, some embodiments of the present application provide an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor can implement a method as described in any embodiment of the first aspect when executing the program.
[0023] In a fifth aspect, some embodiments of the present application provide a computer program product, comprising a computer program, wherein the computer program, when executed by a processor, can implement the method described in any embodiment of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of some embodiments of the present application, the following is a brief introduction to the drawings required for use in some embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0025] Figure 1 System diagram for multi-label classification of tongue images provided for some embodiments of the present application;
[0026] Figure 2 A flowchart of a method for obtaining a tongue surface multi-label classification model provided in some embodiments of the present application;
[0027] Figure 3 A structural diagram of a tongue surface feature model provided for some embodiments of the present application;
[0028] Figure 4 A flowchart of a method for multi-label classification of tongue images provided in some embodiments of the present application;
[0029] Figure 5 A block diagram of the apparatus for multi-label classification of tongue images provided in some embodiments of the present application;
[0030] Figure 6 A schematic diagram of an electronic device is provided for some embodiments of the present application. DETAILED DESCRIPTION
[0031] The technical solutions in some embodiments of the present application will be described below in conjunction with the drawings in some embodiments of the present application.
[0032] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0033] Among related technologies, research on tongue images in traditional Chinese medicine mainly focuses on the classification task of a single feature, while there are relatively few studies on multi-label classification of tongue images. Specifically, most studies on multi-label classification of tongue images improve classic convolutional neural networks such as VGGNet and ResNet, combining attention mechanisms and multi-scale feature extraction technology. Although some progress has been made, it has failed to deeply explore the correlation between labels. Existing research also faces the following shortcomings: insufficient modeling of the relationship between labels, that is, existing research mostly focuses on a single feature, there are few studies on multi-label classification, and multi-label models are difficult to fully capture the interactions between these features; the sample size of the tongue image dataset is small, which limits the generalization ability of the model and makes it difficult to cope with complex clinical application scenarios; in the model training stage, there is a lack of unified model performance evaluation standards, and there is a lack of comparability between different research results.
[0034] In view of this, some embodiments of the present application provide a method for multi-label classification of tongue images. The method constructs a new tongue feature model structure, and obtains a tongue multi-label classification model by training the tongue feature model. The tongue multi-label classification model integrates multiple modules, so as to realize the multi-label classification task of tongue images. At the same time, the tongue multi-label classification model can output the probabilities of different tongue classification labels, capture the correlation between tongue image features, and has good model generalization ability. Moreover, the tongue multi-label classification model can realize the intelligent, objective and standardized classification of traditional Chinese medicine tongue images, improve the accuracy and efficiency of tongue image auxiliary diagnosis, and provide more reliable technical support for traditional Chinese medicine clinical diagnosis.
[0035] The following is combined with Figure 1The overall structure of the system for multi-label classification of tongue images provided by some embodiments of the present application is exemplified.
[0036] like Figure 1 As shown, some embodiments of the present application provide a system diagram for multi-label classification of tongue images. The system may include: a terminal 100 and a server 200. Terminal 100 may send a captured or received tongue image of a target subject to server 200. After acquiring the tongue image, server 200 inputs it into a tongue multi-label classification model and outputs a probability value for each of at least one tongue classification label to which the tongue image belongs. The results are then fed back to terminal 100 for review by the target subject or a doctor.
[0037] In some embodiments of the present application, the tongue and surface multi-label classification model is pre-trained and deployed to the server 200. The training phase of the tongue and surface multi-label classification model can also be performed by the server 200.
[0038] In some embodiments of the present application, the terminal 100 may be a mobile terminal or a non-portable computer terminal, which is not specifically limited in the embodiments of the present application.
[0039] In order to achieve accurate label classification of tongue surface images, we first need to obtain a tongue surface multi-label classification model. Figure 2 The implementation process of obtaining a tongue and surface multi-label classification model performed by the server 200 provided in some embodiments of the present application is exemplified.
[0040] Please see the attached Figure 2 , Figure 2 A flowchart of a method for obtaining a tongue surface multi-label classification model is provided for some embodiments of the present application. The method for obtaining a tongue surface multi-label classification model may include:
[0041] S210 , constructing a tongue surface sample dataset, wherein the tongue surface sample dataset includes: a plurality of tongue surface sample images and a plurality of labels for each of the plurality of tongue surface sample images.
[0042] For example, in some embodiments of the present application, each tongue sample image is labeled to obtain its corresponding sample label. Among them, the sample labels may include: tongue quality label, tongue shape label, tongue coating label, tooth mark label or crack label. Tongue quality labels include: light red tongue, crimson tongue, etc.; tongue shape labels include: fat tongue, normal tongue, thin tongue, etc.; tongue coating labels include: white coating, yellow coating, etc.; tooth mark labels include: tooth mark tongue, normal tongue, etc.; crack labels include: cracked tongue, normal tongue, etc. In addition, the tongue surface sample data set may also include the proportion of each sample label (as a specific example of the sample label probability), which may be given by an expert or obtained through feature calculation, and the embodiments of the present application are not specifically limited here. It is understandable that the sample labels and the subcategories contained in the sample labels can be flexibly expanded, and the embodiments of the present application are not limited to this.
[0043] In some embodiments of the present application, S210 may include: collecting multiple initial tongue surface images; annotating the multiple initial tongue surface images to obtain tongue surface annotated images; segmenting the tongue surface annotated images using a segmentation model to obtain tongue surface segmented images; normalizing and data enhancing the tongue surface segmented images to obtain the multiple tongue surface sample images.
[0044] For example, in some embodiments of the present application, a professional four-diagnostic instrument is used to collect frontal images of the patient's tongue under a constant built-in light source, and a total of thousands of initial tongue images in PNG format are obtained. The initial tongue image samples are annotated by Chinese medicine clinical experts or related algorithms to obtain tongue annotated images, ensuring that the tongue classification labels are accurate and have clinical reference value. The trained Segment Anything Model (SAM) model is used to automatically segment the tongue image area of the tongue annotated image, and the segmentation results are optimized by manual screening to obtain tongue segmentation images. The tongue segmentation image is converted into a tensor format and normalized, and data enhancement strategies such as random horizontal flipping and random rotation are used to obtain multiple tongue sample images. Through a rich sample data set, the robustness and generalization performance of the subsequent training model are improved.
[0045] S220: Train the tongue surface feature model using the tongue surface sample dataset to obtain the tongue surface multi-label classification model. For example, model training is performed in Python 3.9 using the PyTorch framework. The AdamW optimizer and the BCE loss function are used for 50 rounds of training, with a batch size of 16.
[0046] For example, in some embodiments of the present application, the tongue surface feature model is trained using the prepared tongue surface sample dataset to obtain a tongue surface multi-label classification model.
[0047] For ease of understanding, the present application provides the following examples: Figure 3 The structural diagram of the tongue surface feature model includes: Input module, EfficientNetB0 module, cross-scale Transformer module, Multi-Head Attention cross-scale multi-head attention mechanism, graph convolutional network GCN module and Output module.
[0048] Specifically, the pre-trained EfficientNetB0 module is used as the basic feature extraction network. EfficientNetB0 uses the moving inverted residual convolution (MBConv) module and the Squeeze-and-Excitation (SE) attention mechanism, which can effectively extract key information such as tongue quality, tongue coating, and cracks while maintaining a low computational cost. After each tongue sample image in the tongue sample data is input through the Input module, the basic features are first extracted by a 3×3 standard convolution (Conv3×3), and then deep features are extracted by multiple MBConv modules (i.e., multiple MBConv1, k3×3, MBConv6, k3×3, MBConv6, k5×5). Finally, preliminary feature extraction is completed through 1×1 convolution optimization, global average pooling layer dimensionality reduction, and fully connected layer.
[0049] In response to the problem that the receptive field of convolutional neural networks is limited and it is difficult to capture long-distance feature interactions, the embodiment of the present application uses a dual-branch cross-scale Transformer module to further perform global information modeling and multi-scale feature fusion on the original features initially extracted by the EfficientNetB0 module. The cross-scale Transformer module includes two TransformerEncoders. Among them, the two encoders in the dual-branch cross-scale Transformer module constitute two branches, and these two branches include Add&Norm layers, i.e. residual connections and layer normalization, and Feed Forward forward propagation layers. One branch is responsible for processing the original features of the original resolution, retaining key local detail information, and the other branch extracts a wider range of contextual information through downsampling operations, and then restores it to the original resolution through upsampling. With the help of the cross-scale multi-head attention mechanism Multi-Head Attention, the complementary fusion of the above-mentioned two-branch features is achieved, the model's ability to recognize subtle features is enhanced, and multi-scale feature information is effectively integrated. Among them, the original feature representation is regarded as the query Q, the downsampled features are regarded as the key K and the value V, so as to achieve complementary fusion of two scale information. The multi-scale information obtained by splicing the fused features and adjusting the number of channels through 1×1 convolution is input into the GCN module for label relationship modeling.
[0050] Considering that there is a certain correlation between the tongue surface classification labels in tongue image classification, such as cracked tongue is often accompanied by a specific tongue coating color, a graph convolutional neural network architecture based on the fused features (i.e. Figure 3 The GCN module in the network). The adjacency matrix is constructed by calculating the cosine similarity between the feature points in the fused multi-scale information. For example, a sparse adjacency matrix is constructed using the K-nearest neighbor strategy to reduce the computational complexity. The features in the adjacency matrix are iteratively updated using a multi-layer graph convolutional network Conv, the dimension is reduced through a global average pooling operation, and then the classification prediction of the tongue image in the tongue sample image is realized through a fully connected layer, so as to deeply learn the intrinsic relationship between different labels. Specifically, the GCN module flattens the feature map corresponding to the fused multi-scale information, regards each spatial position as a node in the graph, and constructs a sparse adjacency matrix by calculating the cosine similarity between the nodes to reflect the connection relationship between the nodes in the graph structure. After the global average pooling operation, the global feature vector is input into the fully connected classifier, and the Sigmoid activation function is used to output the tongue classification labels (such as tongue quality, tongue shape, tongue coating, tooth marks, and cracks) and the predicted probability of each tongue classification label. The predicted probability of each tongue classification label can intuitively show the correlation between the tongue classification labels. For example, the predicted probability of fissured tongue is 80%, and the predicted probability of the proportion of a specific tongue coating color is 90%, indicating that fissured tongue is often accompanied by a specific tongue coating color.
[0051] In some embodiments of the present application, S220 may include: using the tongue sample data set to train the tongue feature model to obtain the model to be verified; confirming that the value of the evaluation index of the model to be verified meets the preset conditions, and then using the model to be verified as the tongue multi-label classification model; wherein the evaluation index includes: label-level index and model overall-level index; the label-level index includes: label accuracy, classification accuracy, recall rate or harmonization index F1 value; the model overall-level index includes: overall accuracy mean, overall accuracy value, overall recall rate or overall F1 value.
[0052] For example, in some embodiments of the present application, during the training process, it is necessary to verify the model to be verified obtained from each training in order to output a tongue multi-label classification model that meets the requirements. A variety of evaluation indicators are introduced in the verification stage. Among them, the harmonic index F1 value is the harmonic mean of the precision (i.e., classification accuracy) and the recall rate (Recall), which aims to balance the performance of these two indicators. The overall F1 value represents the harmonic mean of the overall precision value and the overall recall rate. The label accuracy rate represents the accuracy of the labels assigned to the tongue sample images, the classification accuracy standard represents the accuracy of the label categories assigned to the tongue sample images, and the recall rate represents the correct ratio of the labels assigned to multiple tongue sample images and the annotated results. The overall model level label can be calculated for the entire tongue sample data set. In addition, the preset condition in the above can be a set threshold, that is, when the value of the evaluation indicator exceeds the set threshold, it is considered that the preset condition is met, and the tongue multi-label classification model can be output at this time. In practical applications, one or more indicators can be selected from the label level indicators and the overall model level indicators as evaluation indicators for the verification model according to actual conditions. The embodiments of the present application are not specifically limited here.
[0053] The following is combined with Figure 4 The implementation process of obtaining multi-label classification of tongue surface images performed by the server 200 provided in some embodiments of the present application is exemplified.
[0054] Please see the attached Figure 4 , Figure 4 A flowchart of a method for multi-label classification of tongue images provided in some embodiments of the present application may include:
[0055] S410: Acquire a tongue surface image of the target object.
[0056] For example, in some embodiments of the present application, the tongue image of the target object is collected by a collection device connected to the terminal 100 or the collection function of the terminal 100 itself to obtain a tongue surface image.
[0057] S420, input the tongue surface image into a pre-trained tongue surface multi-label classification model to obtain the probability value of at least one category of tongue surface classification labels to which the tongue surface image belongs; wherein the tongue surface classification labels include: tongue quality label, tongue shape label, tongue coating label, tooth mark label, and crack label; the tongue surface multi-label classification model is obtained by training the tongue surface feature model, and the tongue surface multi-label classification model includes: a basic feature extraction module, a cross-scale feature processing module, a graph convolutional network module, and a fully connected classifier.
[0058] For example, in some embodiments of the present application, the tongue surface image is input to the Figure 2 and Figure 3In the tongue surface multi-label classification model trained by the method embodiment, the tongue surface classification labels contained in the tongue surface image and the predicted probability of each type of tongue surface classification label (as a specific example of the probability value) are obtained. Among them, the basic feature extraction module in the tongue surface multi-label classification model can be Figure 3 The EfficientNetB0 module in
[15] is used; the cross-scale feature processing module is a cross-scale Transformer module, which also includes Multi-Head Attention; and the graph convolutional network module is a GCN module. Tongue surface classification labels remain the same as those used during training, and can include five major categories (e.g., tongue texture, tongue shape, tongue coating, tooth marks, and cracks), each of which contains different subcategories.
[0059] The above process is described below as an example.
[0060] In some embodiments of the present application, S420 may include:
[0061] S421: Extract the tongue surface image using the basic feature extraction module to obtain initial tongue surface features.
[0062] For example, in some embodiments of the present application, the EfficientNetB0 module can perform feature extraction on the tongue surface image to obtain initial tongue surface features.
[0063] S422: Process the initial tongue surface features through the cross-scale feature processing module to obtain multi-scale tongue surface features.
[0064] For example, in some embodiments of the present application, the initial tongue surface features are input into a cross-scale Transformer module, and multi-scale tongue surface features can be output.
[0065] In some embodiments of the present application, the cross-scale feature processing module includes a first encoder and a second encoder; S422 may include: processing the initial tongue surface features through the first encoder to obtain local feature information; and performing a downsampling operation on the initial tongue surface features through the second encoder to obtain context image information; and using a multi-head cross-attention mechanism to fuse the local feature information and the context image information to obtain the multi-scale tongue surface features.
[0066] For example, in some embodiments of the present application, the feature map corresponding to the initial tongue surface features is flattened and then sent to two parallel Transformer encoders (as a specific example of the first encoder and the second encoder) at the same time to process the initial tongue surface features and the downsampled features at the original resolution, respectively; then, Multi-Head Attention (i.e., multi-head cross attention mechanism) is used to fuse the features output by the two encoders, and the number of channels is adjusted through 1×1 convolution to obtain multi-scale tongue surface features.
[0067] S423: Calculate and perform global average pooling operations on the multi-scale tongue surface features through the graph convolutional network module to obtain a global feature vector.
[0068] For example, in some embodiments of the present application, the GCN module can calculate and process the feature maps corresponding to the multi-scale tongue surface features and output a global feature vector.
[0069] In some embodiments of the present application, S423 may include: constructing an adjacency matrix based on the similarity between feature nodes in the multi-scale tongue surface features; performing a global average pooling operation on the adjacency matrix to obtain the global feature vector.
[0070] For example, in some embodiments of this application, the GCN module flattens the feature map corresponding to the fused multi-scale tongue surface features, treats each spatial location as a node in the map, and constructs a sparse adjacency matrix by calculating the cosine similarity between the nodes. After the adjacency matrix undergoes a global average pooling operation, a global feature vector is obtained.
[0071] S424 , analyzing the global feature vector through the fully connected classifier, and outputting a probability value of each tongue surface classification label in the at least one category of tongue surface classification labels.
[0072] For example, in some embodiments of the present application, the global feature vector is input into a fully connected classifier to output each tongue surface classification label (at least one of the five labels of tongue quality, tongue shape, tongue coating, tooth marks, and cracks) and the predicted probability of each tongue surface classification label.
[0073] Through some of the above embodiments of the present application, it can be seen that the present application combines EfficientNetB0, cross-scale Transformer and graph convolutional network, uses the cross-scale Transformer module to realize the fusion of local and global features, introduces the graph convolutional network to model the dependency relationship between labels, and realizes the multi-label classification of traditional Chinese medicine tongue images. Compared with the basic model in the prior art, the tongue surface multi-label classification model proposed in this application has significantly improved the performance in the tongue image multi-label classification task, with an average precision of 79.91% and a label-level accuracy of 78.61%. It can more accurately identify tongue image features such as tongue quality, tongue shape, tongue coating, tooth marks and cracks, providing a more reliable basis for traditional Chinese medicine diagnosis. Using large-scale data sets for training and adopting data enhancement strategies effectively reduces the risk of overfitting, improves the stability and generalization performance of the tongue surface multi-label classification model under different data distributions, and makes it more suitable for actual clinical applications. With the help of deep learning technology, the dependence of tongue diagnosis results on manual experience is reduced, and the standardization and standardized application of tongue diagnosis are promoted.
[0074] Please refer to Figure 5 , Figure 5 The following is a block diagram illustrating the components of an apparatus for multi-label classification of tongue images provided in some embodiments of the present application. It should be understood that this apparatus for multi-label classification of tongue images corresponds to the aforementioned method embodiments and is capable of performing each of the steps involved in the aforementioned method embodiments. The specific functions of this apparatus for multi-label classification of tongue images can be found in the description above, and a detailed description is omitted here to avoid repetition.
[0075] Figure 5 The device for multi-label classification of tongue images includes at least one software functional module that can be stored in a memory in the form of software or firmware or solidified in the device for multi-label classification of tongue images. The device for multi-label classification of tongue images includes: an image acquisition module 510, used to acquire the tongue image of the target object; a label classification module 520, used to input the tongue image into a pre-trained tongue multi-label classification model to obtain the probability value of at least one category of tongue classification labels to which the tongue image belongs; wherein the tongue classification labels include: tongue quality labels, tongue shape labels, tongue coating labels, tooth mark labels, and crack labels; the tongue multi-label classification model is obtained by training a tongue feature model, and the tongue multi-label classification model includes: a basic feature extraction module, a cross-scale feature processing module, a graph convolutional network module, and a fully connected classifier.
[0076] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method, and will not be described in detail here.
[0077] Some embodiments of the present application further provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the operations corresponding to any of the above methods provided in the above embodiments.
[0078] Some embodiments of the present application further provide a computer program product, which includes a computer program, wherein when the computer program is executed by a processor, it can implement the operations corresponding to any of the above methods provided in the above embodiments.
[0079] like Figure 6 As shown, some embodiments of the present application provide an electronic device 600, which includes: a memory 610, a processor 620, and a computer program stored in the memory 610 and executable on the processor 620, wherein the processor 620 can implement a method as described in any of the above embodiments when reading the program from the memory 610 through the bus 630 and executing the program.
[0080] Processor 620 can process digital signals and can include various computing architectures, such as a complex instruction set computer architecture, a reduced instruction set computer architecture, or an architecture that implements a combination of multiple instruction sets. In some examples, processor 620 can be a microprocessor.
[0081] The memory 610 can be used to store instructions executed by the processor 620 or data related to the execution of instructions. These instructions and / or data may include code for implementing some or all functions of one or more modules described in the embodiments of this application. The processor 620 of the embodiment of the present disclosure can be used to execute the instructions in the memory 610 to implement the method shown above. The memory 610 includes dynamic random access memory, static random access memory, flash memory, optical storage, or other memory known to those skilled in the art.
[0082] The foregoing is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.
[0083] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0084] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
Claims
1. A method for multi-label classification of tongue surface images, characterized in that: include: Acquire a tongue surface image of a target object; The tongue surface image is input into a pre-trained tongue surface multi-label classification model to obtain the probability value of at least one category of tongue surface classification labels to which the tongue surface image belongs; wherein the tongue surface classification labels include: tongue quality label, tongue shape label, tongue coating label, tooth mark label, and crack label; the tongue surface multi-label classification model is obtained by training the tongue surface feature model, and the tongue surface multi-label classification model includes: a basic feature extraction module, a cross-scale feature processing module, a graph convolutional network module, and a fully connected classifier.
2. The method according to claim 1, wherein Inputting the tongue surface image into a pre-trained tongue surface multi-label classification model to obtain a probability value of at least one tongue surface classification label to which the tongue surface image belongs includes: Extracting the tongue surface image by the basic feature extraction module to obtain initial tongue surface features; Processing the initial tongue surface features by the cross-scale feature processing module to obtain multi-scale tongue surface features; Calculating and performing a global average pooling operation on the multi-scale tongue surface features through the graph convolutional network module to obtain a global feature vector; The global feature vector is analyzed by the fully connected classifier, and a probability value of each tongue surface classification label in the at least one category of tongue surface classification labels is output.
3. The method according to claim 2, wherein The cross-scale feature processing module includes a first encoder and a second encoder; wherein, processing the initial tongue surface features by the cross-scale feature processing module to obtain multi-scale tongue surface features includes: Processing the initial tongue surface features by the first encoder to obtain local feature information; and performing a downsampling operation on the initial tongue surface features by the second encoder to obtain context image information; The local feature information and the context image information are fused using a multi-head cross-attention mechanism to obtain the multi-scale tongue surface features.
4. The method according to claim 2 or 3, wherein: The multi-scale tongue surface features are calculated and globally averaged pooled by the graph convolutional network module to obtain a global feature vector, including: constructing an adjacency matrix based on the similarity between feature nodes in the multi-scale tongue surface features; A global average pooling operation is performed on the adjacency matrix to obtain the global eigenvector.
5. The method according to any one of claims 1 to 3, wherein Before inputting the tongue surface image into a pre-trained tongue surface multi-label classification model, the method further includes: Constructing a tongue surface sample dataset, wherein the tongue surface sample dataset includes: a plurality of tongue surface sample images and a plurality of labels for each of the plurality of tongue surface sample images; The tongue surface feature model is trained using the tongue surface sample dataset to obtain the tongue surface multi-label classification model.
6. The method according to claim 5, wherein The multiple tongue surface sample images are obtained by the following method: Collect multiple initial tongue surface images; annotating the multiple initial tongue surface images to obtain an annotated tongue surface image; Segmenting the tongue surface annotated image using a segmentation model to obtain a tongue surface segmentation image; Normalization processing and data enhancement processing are performed on the tongue surface segmented image to obtain the multiple tongue surface sample images.
7. The method according to claim 5, wherein The method of training the tongue surface feature model using the tongue surface sample dataset to obtain the tongue surface multi-label classification model includes: Using the tongue surface sample data set to train the tongue surface feature model to obtain a model to be verified; If it is confirmed that the value of the evaluation index of the model to be verified meets the preset conditions, the model to be verified will be used as the tongue multi-label classification model; wherein the evaluation index includes: label-level index and model-wide index; the label-level index includes: label accuracy, classification accuracy, recall rate or harmonic index F1 value; the model-wide index includes: overall accuracy mean, overall accuracy value, overall recall rate or overall F1 value.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program is executed by a processor to perform the method according to any one of claims 1 to 7.
9. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the computer program executes the method according to any one of claims 1 to 7 when run by the processor.
10. A computer program product, characterized in that The computer program product comprises a computer program, wherein the computer program is executed by a processor to perform the method according to any one of claims 1 to 7.