Traditional Chinese medicine tongue diagnosis feature classification method and system based on million-level private data set

By building a million-level private tongue diagnosis data set and deep learning model, combined with privacy protection technology, the problem of insufficient data in the traditional Chinese medicine tongue diagnosis system is solved, high-precision and rapid classification and diagnosis of tongue diagnosis characteristics is achieved, and the objective analysis ability of traditional Chinese medicine tongue diagnosis is improved.

CN120339216APending Publication Date: 2025-07-18BEIJING ZANGBAOTANG TRADITIONAL CHINESE MEDICINE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510409717.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, due to the small scale of the data set, the traditional Chinese medicine tongue diagnosis system is difficult to fully reflect the diversity and complexity of tongue diagnosis characteristics, making it difficult for the model to learn subtle differences and potential correlations, and insufficient classification accuracy, which affects the comprehensiveness and reliability of the diagnosis.

Method used

A million-level private tongue diagnosis image dataset is constructed, images are collected and preprocessed through distributed devices, multi-dimensional features are extracted and classified using deep learning models, and combined with privacy protection technology to generate diagnostic reports and personalized suggestions.

Benefits of technology

It significantly improves the accuracy of feature classification, realizes comprehensive characterization of tongue diagnosis characteristics, improves the reliability and efficiency of classification results, ensures data security, and controls the processing time of a single image within 1-2 seconds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005342222360000011
    Figure HDA0005342222360000011
  • Figure HDA0005342222360000012
    Figure HDA0005342222360000012
Patent Text Reader

Abstract

The invention relates to the field of traditional Chinese medicine diagnostics, and discloses a traditional Chinese medicine tongue diagnosis feature classification method and system based on a million-level private data set, and the method comprises the following steps: S1, collecting tongue diagnosis images through distributed equipment, and constructing and dynamically updating a million-level private tongue diagnosis image data set; s2, the tongue diagnosis image is preprocessed, a deep learning model is utilized to extract and classify multi-dimensional tongue diagnosis features, tongue diagnosis feature categories and association between the tongue diagnosis features and health states are output, and the multi-dimensional features comprise color features, shape features, coating features and texture features; and S3, a privacy protection technology is integrated in data processing and model training and is responsible for data security, and a diagnosis report and personalized health suggestions are generated in real time according to a classification result. According to the method, by constructing and utilizing the million-level private tongue diagnosis image data set, the diversity of tongue diagnosis features can be more comprehensively captured, and the overfitting problem caused by insufficient samples can be effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of traditional Chinese medicine diagnostics, and particularly to a method and system for classifying tongue diagnosis features in traditional Chinese medicine based on a private dataset of millions of images. Background Art

[0002] Tongue diagnosis in traditional Chinese medicine is an important diagnostic method. By observing features such as the color, shape, coating, and texture of the tongue, the health status of patients can be judged. With the development of artificial intelligence technology, in recent years, there have been various automated tongue diagnosis methods based on machine learning and deep learning. These methods usually use pre-trained models to segment, extract features, and classify tongue diagnosis images, aiming to reduce the subjectivity of manual diagnosis and improve the analysis efficiency. In the prior art, the implementation of tongue diagnosis systems mostly relies on image processing technology and neural network models, and through the digital representation of tongue image features, a preliminary assessment of the health status is provided.

[0003] However, the prior art has significant drawbacks, mainly reflected in the relatively small scale of the dataset, usually only containing hundreds to thousands of images, which is difficult to fully reflect the diversity and complexity of tongue diagnosis features. Due to insufficient data volume, rare or subtle tongue image features (such as specific color changes or complex texture patterns) appear with low frequency in training samples, resulting in the model being difficult to learn the laws of these features and prone to overfitting. This data limitation restricts the classification accuracy and is difficult to reach the high-precision level required for clinical diagnosis. The existing methods are insufficient in capturing the subtle differences and potential correlations of tongue diagnosis features, thus affecting the comprehensiveness and reliability of tongue diagnosis analysis. Summary of the Invention

[0004] To make up for the above deficiencies, the present invention provides a method and system for classifying tongue diagnosis features in traditional Chinese medicine based on a private dataset of millions of images, aiming to improve the problem that due to insufficient data volume, the model is difficult to learn the laws of these features and is prone to overfitting.

[0005] In the first aspect, the present invention provides the following technical solution. A method for classifying tongue diagnosis features in traditional Chinese medicine based on a private dataset of millions of images includes the following steps:

[0006] S1. Collect tongue diagnosis images through distributed devices, and construct and dynamically update a private tongue diagnosis image dataset of millions of images;

[0007] S2. Preprocess the tongue diagnosis images, extract multi-dimensional tongue diagnosis features using a deep learning model and classify them, and output the tongue diagnosis feature categories and their associations with the health status. The multi-dimensional features include color features, shape features, coating features, and texture features;

[0008] S3. Integrate privacy protection technologies in data processing and model training, be responsible for data security, and generate diagnostic reports and personalized health advice in real time according to the classification results.

[0009] Preferably, the step S1 includes:

[0010] Collect tongue diagnosis images through a mobile device or medical institution device authorized by the user, along with metadata, where the metadata includes the shooting time, region, patient age, and gender;

[0011] Use a pre-trained segmentation model to automatically segment the tongue body area and combine it with manual verification to generate labeled data. Design a distributed database to support real-time data upload and update, and perform clustering analysis regularly to identify new tongue diagnosis patterns.

[0012] Preferably, the extraction of multi-dimensional tongue diagnosis features in the step S2 includes:

[0013] Perform standardized preprocessing on the tongue diagnosis images, adjust the lighting and remove background noise;

[0014] Adopt an improved deep convolutional neural network based on ResNet-50 combined with an attention mechanism to extract the color features, shape features, coating features, and texture features.

[0015] Preferably, the classification in the step S2 includes:

[0016] Construct a multi-task classifier based on the Transformer architecture, and output the tongue diagnosis feature categories and their associations with traditional Chinese medicine health status;

[0017] Use a private dataset of millions for supervised learning, and combine transfer learning and data augmentation techniques to optimize the model performance.

[0018] Preferably, the integration of privacy protection technologies in the step S3 includes:

[0019] Adopt federated learning to train the model locally on the user device and only upload the model parameters to the central server;

[0020] Add differential privacy noise in data aggregation and model update so that individual data cannot be reverse-derived.

[0021] Preferably, the real-time generation of diagnostic reports and personalized health advice in the step S3 includes:

[0022] Complete the feature extraction and classification of a single tongue diagnosis image within 1 - 2 seconds;

[0023] Generate a personalized tongue diagnosis model through clustering analysis according to the user metadata, provide targeted health advice, and visually display the tongue diagnosis features in the form of a heat map.

[0024] In a second aspect, the present invention provides the following technical solution. A traditional Chinese medicine tongue diagnosis feature classification system based on a million-level private dataset includes:

[0025] A data acquisition and management module for acquiring tongue diagnosis images through distributed devices, constructing, and dynamically updating a million-level private dataset;

[0026] A feature extraction and classification module for extracting multi-dimensional features of tongue diagnosis images based on a deep learning model and classifying them, outputting the tongue diagnosis feature categories and their associations with health status;

[0027] A privacy protection and result output module for integrating privacy protection technology during data processing and generating a diagnostic report and personalized health advice in real time according to the classification results.

[0028] Preferably, the data acquisition and management module includes a data acquisition unit, a data annotation unit, and a data update unit for acquiring tongue diagnosis images and metadata, generating annotation data, and performing real-time data updates;

[0029] The feature extraction and classification module includes a preprocessing unit, a feature extraction unit, and a classification unit for image preprocessing, feature extraction, and classification output;

[0030] The privacy protection and result output module includes a real-time processing unit and a personalized generation unit for completing processing and generating personalized suggestions within 1-2 seconds.

[0031] In a third aspect, the invention provides the following technical solution. A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned traditional Chinese medicine tongue diagnosis feature classification method based on a million-level private dataset.

[0032] In a fourth aspect, the present invention provides the following technical solution. A readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned traditional Chinese medicine tongue diagnosis feature classification method based on a million-level private dataset.

[0033] The present invention has the following beneficial effects:

[0034] 1. In the present invention, by constructing and utilizing a million-level private tongue diagnosis image dataset, the accuracy of feature classification is significantly improved, and the diversity of tongue diagnosis features can be captured more comprehensively. The million-level data provides rich training samples for deep learning models (such as improved ResNet-50 and Transformer), enabling the classification accuracy of the model on the test set to reach over 95%, effectively reducing the overfitting problem caused by insufficient samples, and providing technical support for the objective analysis of traditional Chinese medicine tongue diagnosis.

[0035] 2. In the present invention, by extracting and classifying multi-dimensional tongue diagnosis features (including color, shape, coating, and texture), a comprehensive characterization of tongue images is achieved. Using a deep convolutional neural network combined with an attention mechanism to simultaneously process the interrelationships among multiple features, a multi-dimensional feature vector (with a dimension of 1024) is classified through a Transformer model. The output results not only include feature categories but also cover their relevance to the health status, improving the integrity of diagnostic information. The F1 score of the classification results is increased by approximately 10%, providing a more reliable basis for clinical diagnosis.

[0036] 3. In the present invention, federated learning and differential privacy technologies are integrated in the processing of a million-scale private dataset to ensure data security. By locally training the model on the client side (using MobileNetV3 and only uploading parameters instead of the original images) and adding Gaussian noise (σ = 0.1, privacy budget ε = 1.0) on the server side, the risk of data leakage is effectively prevented.

[0037] 4. In the present invention, through distributed computing and optimized model design, the real-time processing ability of tongue diagnosis images is achieved. The total time for preprocessing, feature extraction, and classification of a single image is controlled within 1 - 2 seconds. Relying on a cloud GPU cluster (such as 4 NVIDIA V100s), 500 images can be processed per second. Compared with traditional manual tongue diagnosis (taking several minutes) or existing automated systems (with a processing time of about 5 - 10 seconds), the efficiency of this method is significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a method flow chart of the traditional Chinese medicine tongue diagnosis feature classification method based on a million-scale private dataset proposed by the present invention;

[0039] Figure 2 is a system framework diagram of the traditional Chinese medicine tongue diagnosis feature classification system based on a million-scale private dataset proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0041] Embodiment 1

[0042] Referring to Figure 1 , in the first embodiment of the present invention, the present invention provides a traditional Chinese medicine tongue diagnosis feature classification method based on a million-scale private dataset, including the following steps:

[0043] S1. Collect tongue diagnosis images through distributed devices, construct and dynamically update a private tongue diagnosis image dataset of millions of levels.

[0044] S2. Preprocess the tongue diagnosis images, extract multi-dimensional tongue diagnosis features using a deep learning model and classify them, output the tongue diagnosis feature categories and their associations with the health status. The multi-dimensional features include color features, shape features, coating features, and texture features.

[0045] S3. Integrate privacy protection technologies in data processing and model training, be responsible for data security, and generate a diagnostic report and personalized health advice in real time according to the classification results.

[0046] Specifically, in step S1, the construction of the private dataset of millions of levels relies on distributed devices (such as smartphones and tongue diagnosis instruments). The camera API (resolution ≥ 1920×1080, frame rate 30fps) is called through the client software to collect tongue diagnosis images and store them in the JPEG format. The size of a single image is about 2MB. The data is uploaded to the cloud and encrypted for transmission through the HTTPS protocol, with a bandwidth requirement of 50Mbps. Step S2 uses a deep learning model. The specific implementation includes image preprocessing (performing grayscale conversion and histogram equalization using the OpenCV library), feature extraction (based on the improved ResNet-50 network, with a convolution kernel size of 3×3 and a stride of 2), and classification (Transformer model, 12-layer encoder, 8 attention heads). The entire process runs in parallel on a GPU server (NVIDIA A100, video memory 40GB), and the processing time for a single image is 1.5 seconds. Step S3 protects the data through federated learning (aggregating parameters using the FedAvg algorithm, with a weight update period of 24 hours) and differential privacy (noise standard deviation σ = 0.1, privacy budget ε = 1.0). The diagnostic report is generated in the JSON format and pushed to the user side in real time through WebSocket. The personalized advice calculates the user feature similarity based on K-means clustering (number of clusters K = 10).

[0047] Step S1 includes:

[0048] Collect tongue diagnosis images through mobile devices or medical institution devices authorized by the user, along with metadata, where the metadata includes the shooting time, region, patient age, and gender.

[0049] Use a pre-trained segmentation model to automatically segment the tongue body area and combine it with manual verification to generate labeled data. Design a distributed database to support real-time data upload and update, and perform clustering analysis regularly to identify new tongue diagnosis patterns.

[0050] Specifically, the technical implementation of step S1 starts from data collection. The mobile device captures images through the Android Camera2 API or iOS AVCaptureSession, enabling autofocus and exposure compensation during shooting to ensure that the tongue body occupies more than 70% of the picture. The metadata is recorded in a JSON structure, such as {"time":"2025-03-24 10:00","location":"latitude 23.1, longitude 113.2","age":45,"gender":"male"}. The pre-trained segmentation model is based on U-Net. The input image is resized to 512×512, with 32 convolutional layers, trained using the ReLU activation function and Dice loss function. The training dataset consists of 50,000 tongue diagnosis images, iterated 100 times, and the segmentation accuracy reaches 98%. Manual verification is achieved through an online platform (such as a Web application developed with Django). Experts use the mouse to select the tongue body area, and the annotation time is about 20 seconds per image. The distributed database adopts the Hadoop HDFS architecture, with a single node storing 10TB, a data shard size of 128MB, and the upload is processed through Spark Streaming at a speed of 100Mbps. The clustering analysis uses the K-means algorithm, with the input features being the RGB values of the tongue color and texture entropy (calculated based on the gray-level co-occurrence matrix), iterated 20 times, to identify new tongue diagnosis patterns such as "pale tongue with cracks".

[0051] The extraction of multi-dimensional tongue diagnosis features in step S2 includes:

[0052] Perform standardized preprocessing on the tongue diagnosis images to adjust the lighting and remove background noise;

[0053] Adopt an improved deep convolutional neural network based on ResNet-50 combined with an attention mechanism to extract color features, shape features, coating features, and texture features.

[0054] Specifically, the preprocessing in step S2 is implemented using the OpenCV library. The specific process includes: reading the RGB image, converting it to a grayscale image (formula Gray = 0.299R + 0.587G + 0.114B), applying Gaussian blur (kernel size 5×5, σ = 1.5) for denoising, and then enhancing the contrast through histogram equalization to ensure the visibility of tongue coating details. Feature extraction is based on an improved ResNet-50 model. The specific implementation is as follows: adding an SE attention module (compression ratio 16) to the third residual block, increasing the number of convolutional layer parameters to 25M, resizing the input image to 224×224, and outputting a 1024-dimensional feature vector; the color features are calculated through the RGB mean and standard deviation (range 0-255), the shape features are calculated by calculating the area / perimeter ratio after extracting the contour using Canny edge detection (low threshold 50, high threshold 150), the coating features are calculated based on the gray-level co-occurrence matrix (distance 1, angle 0° / 45° / 90°) for contrast and uniformity, and the texture features are extracted through Gabor filters (frequency 0.1-0.5, direction 0°-135°) for crack direction. Model training uses the PyTorch framework, optimizer Adam (learning rate 0.001, β1 = 0.9), batch size 64, and the training time is about 72 hours.

[0055] The classification in step S2 includes:

[0056] Constructing a multi-task classifier based on the Transformer architecture to output the tongue diagnosis feature categories and their associations with traditional Chinese medicine health status;

[0057] Using a private dataset of millions for supervised learning, and optimizing the model performance by combining transfer learning and data augmentation techniques.

[0058] Specifically, the classification technology in step S2 is implemented based on the Transformer model. The specific process is as follows: Input a 1024-dimensional feature vector with an embedding layer dimension of 512. Process it through 12 encoder layers (each layer has 8 attention heads and a feed-forward network dimension of 2048). Use LayerNorm normalization and Dropout (rate 0.1) to prevent overfitting. Output 10 types of tongue diagnosis features (such as "damp-heat tongue") and probabilities (such as 0.85). The model is trained on a dataset of millions. The dataset is divided into 800,000 training images, 100,000 validation images, and 100,000 test images. The cross-entropy loss function is adopted, and the weight initialization is based on ImageNet pre-training; for transfer learning, the first 8 encoder layers are frozen, and only the last 4 layers are fine-tuned for 50 epochs. Data augmentation includes random cropping (ratio 0.9 - 1.0), brightness adjustment (±20%), and horizontal flipping, which is implemented using the Albumentations library. After augmentation, the sample size increases to 1.2 million. During implementation, the model is deployed on TensorFlow Serving, with a single inference time of 0.7 seconds, and the accuracy of the test set reaches 95%. The recognition rate for complex tongue images (such as "purple tongue with white greasy coating") is increased by 20%.

[0059] The integrated privacy protection technology in step S3 includes:

[0060] Adopt federated learning to train the model locally on user devices and only upload the model parameters to the central server;

[0061] Add differential privacy noise during data aggregation and model update so that individual data cannot be reverse-derived.

[0062] Specifically, in the implementation of the privacy protection technology in step S3, the FedAvg algorithm is adopted for federated learning. The specific process is as follows: Client devices (CPU ≥ 2GHz, RAM ≥ 4GB) load the MobileNetV3 model (parameter quantity 4M), input local tongue diagnosis images (batch size 32), and use the SGD optimizer (learning rate 0.01) to train for 5 epochs to generate a parameter update vector (size approximately 10MB); the client uploads the parameters to the server through the gRPC protocol, and the server uses weighted average aggregation (weight = local data volume / total data volume) with an update period of 24 hours. Differential privacy is implemented on the server side. Gaussian noise (mean 0, σ = 0.1) is added during parameter aggregation, and the privacy budget ε = 1.0. The noise scale is calculated through the TensorFlowPrivacy library to ensure compliance with the ε-differential privacy definition. During implementation, the system supports concurrent training of 10 million clients. The server is equipped with a 128-core CPU and 512GB of RAM, and the aggregation takes 30 seconds. Privacy tests show that the data leakage risk is lower than 0.01%, meeting the requirements of GDPR and the Personal Information Protection Law of the People's Republic of China.

[0063] The real-time generation of diagnostic reports and personalized health advice in step S3 includes:

[0064] Complete the feature extraction and classification of a single tongue diagnosis image within 1 - 2 seconds;

[0065] Generate a personalized tongue diagnosis model through cluster analysis based on user metadata, provide targeted health advice, and visually display tongue diagnosis features in the form of a heat map.

[0066] Specifically, the real-time diagnosis technology in step S3 is implemented through a cloud GPU cluster. The specific process is as follows: The user uploads the tongue diagnosis image to the server (bandwidth 100Mbps), and preprocessing (OpenCV, taking 0.3 seconds), feature extraction (ResNet-50, taking 0.5 seconds), and classification (Transformer, taking 0.7 seconds) run in parallel. The total time consumption is 1.5 seconds. The cluster is equipped with 4 NVIDIA V100s (video memory 32GB), supporting the processing of 500 images per second. The personalized model is based on K-means clustering. Input user metadata (such as {"age": 50, "region": "south"}) and the tongue diagnosis feature vector, set K = 10, and iterate 20 times to generate a subset model (such as "southern middle-aged group"), which is trained using PyTorch, and the output advice is such as "Damp-heat constitution, it is recommended to have a heat-clearing diet". The visualization heat map is generated through the Matplotlib library. The tongue color is mapped to an RGB gradient (red 0 - 255), and the coating thickness is represented by transparency (0 - 1). The result is returned to the user App in PNG format. In the embodiment, tests on 1000 users show that the average diagnosis time is 1.8 seconds, and the accuracy of personalized advice reaches 90%.

[0067] Embodiment 2:

[0068] Refer to Figure 2 In the second embodiment of the present invention, the present invention provides a traditional Chinese medicine tongue diagnosis feature classification system based on a million-level private dataset, including:

[0069] A data collection and management module for collecting tongue diagnosis images through distributed devices, constructing, and dynamically updating a million-level private dataset;

[0070] A feature extraction and classification module for extracting multi-dimensional features of tongue diagnosis images based on a deep learning model and classifying them, outputting the tongue diagnosis feature categories and their associations with health status;

[0071] A privacy protection and result output module for integrating privacy protection technology during data processing and generating diagnostic reports and personalized health advice in real time according to the classification results.

[0072] Specifically, the data collection and management module calls the camera API through edge devices (smartphones, Android 10+ or iOS 14+) to collect images, and uploads them to the server (HTTPS, certificate SHA-256 encryption) using the OkHttp library. The server-side receives data based on the Spring Boot framework and stores it in a MySQL database (single-table capacity of 10 million records). The feature extraction and classification module is deployed in a Docker container, running the improved ResNet-50 and Transformer models, and uses Kubernetes to manage the container cluster (10 nodes, 16-core CPU and 64GB RAM per node), supporting the processing of 500,000 images per day. The privacy protection and result output module achieves load balancing through Nginx reverse proxy, enables TLS 1.3 encryption for data transmission, generates diagnostic reports in JSON format (example: {"tongue_type": "damp-heat tongue", "confidence": 0.85, "advice": "light diet"}), and pushes them to the user side through WebSocket. The system supports multiple languages (Chinese, English), can be docked with the hospital HIS system through the HL7 protocol to improve clinical efficiency.

[0073] The data collection and management module includes a data collection unit, a data annotation unit, and a data update unit, which are used to collect tongue diagnosis images and metadata, generate annotated data, and perform real-time data updates.

[0074] The feature extraction and classification module includes a preprocessing unit, a feature extraction unit, and a classification unit, which are used for image preprocessing, feature extraction, and classification output.

[0075] The privacy protection and result output module includes a real-time processing unit and a personalized generation unit, which are used to complete processing and generate personalized suggestions within 1-2 seconds.

[0076] Specifically, in the implementation of the data collection and management module, the data collection unit detects the image quality through the CameraX library (clarity threshold 0.8), and automatically retakes unqualified images; the data annotation unit develops a web platform based on Flask, and experts use the Canvas API to select the tongue body and save the annotated data as JSON (such as {"color": "red", "coating": "thick yellow"}), which is synchronized to HDFS (sharding 128MB); the data update unit processes uploads through the Kafka queue (number of partitions 10), with a throughput of 10,000 messages per second. In the feature extraction and classification module, the preprocessing unit uses OpenCV for batch processing (batch size 1000), the feature extraction unit dynamically loads the ResNet model (depth 18 - 101 optional), and the classification unit calls the Transformer service through gRPC to output results with a confidence level ≥ 0.9. In the privacy protection and result output module, the real-time processing unit caches intermediate results based on Redis (TTL 60 seconds), the personalized generation unit uses Spark MLlib to calculate K-means (K = 10), and the generation of suggestions takes 0.5 seconds. During implementation, the system is piloted in the hospital, processing 5,000 images per day, with a 35% improvement in diagnostic consistency.

[0077] Embodiment 3

[0078] In the third embodiment of the present invention, based on the same inventive concept, a computer-readable storage medium is proposed. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the above-mentioned method for classifying traditional Chinese medicine tongue diagnosis features based on a million-level private dataset.

[0079] Embodiment 4

[0080] In the fourth embodiment of the present invention, based on the same inventive concept, a computer device is proposed. The terminal includes: a processor and a memory; the processor and the memory communicate with each other; the memory is used to store instructions; the processor is used to execute the instructions in the memory to implement the above-mentioned method for classifying traditional Chinese medicine tongue diagnosis features based on a million-level private dataset.

[0081] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0082] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for classifying the characteristics of traditional Chinese medicine tongue diagnosis based on a private dataset of millions of levels, characterized in that, It includes the following steps: S1. Collect tongue diagnosis images through distributed devices, construct and dynamically update a private tongue diagnosis image dataset of millions of levels; S2. Preprocess the tongue diagnosis images, extract multi-dimensional tongue diagnosis features using a deep learning model and classify them, output the tongue diagnosis feature categories and their associations with the health status, where the multi-dimensional features include color features, shape features, coating features, and texture features; S3. Integrate privacy protection technologies in data processing and model training, be responsible for data security, and generate a diagnosis report and personalized health advice in real time according to the classification results.

2. The method for classifying traditional Chinese medicine tongue diagnosis features based on a million-level private data set according to claim 1, wherein The step S1 includes: Collect tongue diagnosis images through mobile devices or medical institution devices authorized by users, with attached metadata, where the metadata includes shooting time, region, patient age, and gender; Automatically segment the tongue body area using a pre-trained segmentation model and generate annotation data in combination with manual verification, design a distributed database to support real-time data upload and update, and perform clustering analysis regularly to identify new tongue diagnosis patterns.

3. The method for classifying traditional Chinese medicine tongue diagnosis features based on a million-level private dataset according to claim 1, wherein In the step S2, the extraction of multi-dimensional tongue diagnosis features includes: Perform standardized preprocessing on the tongue diagnosis images, adjust the illumination and remove background noise; Adopt an improved deep convolutional neural network based on ResNet-50 combined with an attention mechanism to extract the color features, shape features, coating features, and texture features.

4. The traditional Chinese medicine tongue diagnosis feature classification method based on a private data set of millions, according to claim 1, is characterized in that In the step S2, the classification includes: Construct a multi-task classifier based on the Transformer architecture, and output the tongue diagnosis feature categories and their associations with the traditional Chinese medicine health status; Use the private dataset of millions of levels for supervised learning, and optimize the model performance by combining transfer learning and data augmentation techniques.

5. The method for classifying traditional Chinese medicine tongue diagnosis features based on a million-level private dataset according to claim 1, wherein In the step S3, the integration of privacy protection technologies includes: Adopt federated learning to train the model locally on user devices and only upload the model parameters to the central server; Add differential privacy noise in data aggregation and model update, so that individual data cannot be deduced reversely.

6. The method for classifying traditional Chinese medicine tongue diagnosis features based on a million-level private data set according to claim 1, wherein In the step S3, the real-time generation of a diagnosis report and personalized health advice includes: Complete the feature extraction and classification of a single tongue diagnosis image within 1-2 seconds; Generate a personalized tongue diagnosis model through clustering analysis according to user metadata, provide targeted health advice, and visually display the tongue diagnosis features in the form of a heat map.

7. A traditional Chinese medicine tongue diagnosis feature classification system based on a private dataset of millions, characterized in that For the traditional Chinese medicine tongue diagnosis feature classification method based on the private dataset of millions of levels according to any one of claims 1-6, it includes: A data collection and management module for collecting tongue diagnosis images through distributed devices, constructing and dynamically updating a private dataset of millions of levels; A feature extraction and classification module for extracting multi-dimensional features of tongue diagnosis images based on a deep learning model and classifying them, outputting the tongue diagnosis feature categories and their associations with the health status; A privacy protection and result output module for integrating privacy protection technologies in data processing and generating a diagnosis report and personalized health advice in real time according to the classification results.

8. The traditional Chinese medicine tongue diagnosis feature classification system based on a million-level private data set according to claim 7, wherein The data collection and management module includes a data collection unit, a data annotation unit, and a data update unit, which are used to collect tongue diagnosis images and metadata, generate annotation data, and perform real-time data update; The feature extraction and classification module includes a preprocessing unit, a feature extraction unit, and a classification unit, which are used for image preprocessing, feature extraction, and classification output; The privacy protection and result output module includes a real-time processing unit and a personalized generation unit, which are used to complete the processing and generate personalized suggestions within 1-2 seconds.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the traditional Chinese medicine tongue diagnosis feature classification method based on a million-level private dataset according to any one of claims 1 to 6.

10. A readable storage medium, characterized in that, A computer program is stored on the readable storage medium, and when the computer program is executed by the processor, it implements the traditional Chinese medicine tongue diagnosis feature classification method based on a million-level private dataset according to any one of claims 1 to 6.