A method for classifying traditional Chinese medicine syndromes and a database system thereof

CN117807534BActive Publication Date: 2026-09-11SHANDONG UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410003204.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2026-09-11
Estimated Expiration
2044-01-02

AI Technical Summary

Technical Problem

但尚缺乏将AI应用于中医证候分类判断方面的研究

Benefits of technology

[0033] In existing technologies, AI diagnostic and treatment models often overlook the relationships between symptoms/signs, leading to inaccurate data sources and limiting the depth of machine learning. Unsupervised clustering algorithms are a classic AI method that does not rely on predefined labels and has the potential to classify symptoms/signs based on inherent human principles. By eliminating human bias in the classification process, these algorithms can more accurately establish objective subtypes of diseases. In the MeSH library and the "Chinese Thesaurus of Traditional Chinese Medicine," each subject term can be described by one or more tree structure numbers, indicating its position in the hierarchical tree structure and its relationships with other subject terms. Combined with this invention, the tree structure numbers can be used to reveal potential deep relationships between symptoms/signs, meeting the requirements of applying TCM syndrome differentiation concepts to different types of diseases in modern medicine. This invention establishes a special AI computational model to mine all symptoms/signs of patients in clinical literature research, objectively classify patients, and uncover currently unknown syndrome types and medication patterns in TCM treatment, providing new pathways and methods for research and datafication in the field of TCM. Furthermore, this invention can also provide assistance in establishing new digital TCM diagnosis and treatment models, and promote the inheritance and development of TCM.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117807534B_ABST
    Figure CN117807534B_ABST
Patent Text Reader

Abstract

The application discloses a TCM syndrome attribution classification method and a database system thereof. The classification method adopts an unsupervised K-Means calculation model, collects and screens a symptom / physical sign data set of a certain disease and a corresponding treatment traditional Chinese medicine data set of a MySQL database, represents symptom / physical sign information in a vector form, reclassifies diseases into different subtypes based on association implied by a hierarchical rule of MeSH, and classifies diseases into different subtypes through an unsupervised clustering algorithm. Furthermore, the database system is established based on the above method. The application can be used for mining all symptoms and physical signs of patients in clinical literature research, objectively classifying patients, scientifically analyzing syndrome classification of the disease, mining unknown syndrome types and medication rules of traditional Chinese medicine treatment, and providing a new way and method for research and dataization in the field of traditional Chinese medicine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to digital medical technology, and specifically relates to a traditional Chinese medicine syndrome attribution classification method and a database system thereof. Background Art

[0002] Personalized diagnosis and treatment have attracted increasing attention in the medical field. The syndrome-based syndrome differentiation and treatment of traditional Chinese medicine (TCM) represents the oldest personalized medical system, which focuses on diagnosis and treatment based on individual factors. The "syndrome" in TCM refers to a summary understanding of pathophysiological differences at different stages of a disease, taking into account factors such as the nature and location of the disease. Specific traditional Chinese medicine prescriptions are issued according to the "syndrome" manifested by each patient. For example, according to TCM theory, depression can be classified into multiple subtypes ("syndromes") based on patients' symptoms / signs, including liver qi stagnation syndrome, liver stagnation and spleen deficiency syndrome, qi deficiency and blood stasis syndrome, etc. Different treatment strategies are then provided according to the patient's specific "syndrome". For example, prescriptions for regulating qi are used for the liver qi stagnation subtype, and prescriptions for tonifying qi and activating blood are used for the qi deficiency and blood stasis subtype. A number of modern studies have proven that TCM syndrome differentiation and treatment has superiority in terms of efficacy. However, the determination of "syndrome" largely relies on the experience of TCM physicians in extracting and analyzing the implicit information in symptoms / signs. This subjective method may lack sufficient objectivity and accuracy, and ultimately the limited experience and depth of understanding of humans in correct diagnosis and treatment limit the development of personalized treatment based on TCM theory. Therefore, it is necessary to find a method and paradigm that can promote the objectification and accuracy of TCM "syndromes". The rapid development of artificial intelligence (AI) has enabled the training and analysis of large data sets, thereby promoting the progress of personalized medicine. At present, AI has been applied to explore the specific relationship between the properties of Chinese herbal medicines and their effects. However, there is still a lack of research on applying AI to TCM syndrome classification and judgment. This method and paradigm obtained by using AI to calculate a large amount of experience and data accumulated in personalized diagnosis and treatment of clinical diseases in TCM can diagnose different syndrome types of diseases and determine the corresponding optimal treatment plan. The application of this method will promote the inheritance and development of clinical application of TCM. Summary of the Invention

[0003] To overcome the above technical drawbacks, the present invention provides a traditional Chinese medicine syndrome attribution classification method. The main concept of the method is as follows: collecting and screening the symptom / sign data set of a certain disease and the corresponding therapeutic traditional Chinese medicine data set, establishing a MySQL database of symptoms / signs and used prescriptions for depression patients, representing symptom / sign information in a vectorized form, and performing reclassification of different subtypes of the disease through an unsupervised clustering algorithm based on the implicit associations in the hierarchical rules of Medical Subject Headings (MeSH) and Chinese Traditional Medicine and Materia Medica Subject Headings (CTMMMSH), so as to scientifically analyze the syndrome classification of the disease and discover currently unrecognized syndrome types.

[0004] The specific technical solution includes the following steps:

[0005] 1. Literature collection and screening: Collect and screen clinical literature on traditional Chinese medicine treatment that is related to the disease, has a clear therapeutic effect, and records the symptoms and signs of cases;

[0006] 2. Extraction and standardization of literature information: Extract and standardize the symptoms, signs and corresponding medication information of each case recorded in the above literature; perform frequency analysis on the above information, create disease symptom and sign datasets and medication information datasets, and import them into a MySQL database for storage;

[0007] 3. Disease patient clustering: Based on the above disease symptom and sign dataset, an unsupervised K-Means calculation model is used to divide disease patients into different groups;

[0008] 4. Analyze the medication patterns of patient groups and deduce the syndrome categories of the corresponding groups: Based on the above medication information dataset, deduce the core prescriptions and efficacy of the Chinese herbal medicines used in each group, and thus deduce the syndrome categories of the corresponding groups.

[0009] 5. Verification of syndrome categories: Compare the similarity of symptoms / signs between different groups and the same syndrome in traditional Chinese medicine to verify the syndrome categories derived above.

[0010] The sources of the literature include CNKI, VIP, China Medical Information Network, Wanfang, PubMed, and Web of Science databases.

[0011] The preferred method for selecting the literature is as follows: the collected literature is processed using Knime 5.4, and the data tables are connected with uniform columns; redundant literature is eliminated by comparing the title and abstract columns according to the literature inclusion criteria; the improved literature list is imported into Zotero 6.0, and the full text of the articles is downloaded to create a local literature database for data extraction.

[0012] The preferred standardization method for the literature information is as follows: For the standardization of symptoms and signs, the classification standards provided by ICD-11 International Classification of Diseases, 11th Revision, China Traditional Chinese Medicine Subject Headings (CTMMMSH), and SymMap database are used, and then the tree structure number of each symptom / sign is obtained according to MeSH and CTMMMMSH; For the standardization of medication information, the names of Chinese medicines in the medication information are standardized based on the Pharmacopoeia of the People's Republic of China and the Chinese Materia Medica.

[0013] The preferred method for establishing the unsupervised K-Means computation model includes the following steps:

[0014] 1. Establishing the hierarchical relationship matrix of symptoms or signs: Convert the hierarchical relationships of the tree structure to Boolean hierarchies. For each relationship, calculate according to the following formula and represent it as "0" or "1".

[0015]

[0016] Where x and y are the tree structure numbers of disease symptoms or signs; their Boolean values ​​form a two-dimensional matrix representing the hierarchical relationship of the tree structure numbers of disease symptoms or signs; through this matrix, it can be determined whether a given x belongs to y in the hierarchical structure;

[0017] 2. Construction of patient fingerprint vectors: The symptoms / signs of each disease patient are transformed into phenotypic vectors, where the values ​​in the vectors are represented by 0 or 1. A value of 1 indicates that the patient exhibits a specific symptom / sign, while a value of 0 indicates that the patient does not have that symptom / sign. By performing a dot product between the phenotypic vector and the symptom / sign grading matrix, all grading information of a single disease patient is captured. The resulting dot product matrix is ​​flattened, and the dimensions are pruned with values ​​of 0 to create fingerprint vectors.

[0018] 3. Unsupervised Machine Learning: The K-Means clustering algorithm was implemented using the `sklearn.cluster.KMeans` package in Python 3.9. Fingerprint vectors were used as input to the clustering algorithm. During K-Means clustering, cosine distance was calculated to measure the similarity between fingerprint vectors. The results of unsupervised clustering were visualized as a 3D scatter plot. The silhouette coefficient was chosen as the evaluation metric for the clustering algorithm and calculated using the following formula:

[0019]

[0020] In the formula, s is the average silhouette coefficient of all samples; a is the average intra-cluster distance of each sample; and b is the distance between a sample and the nearest cluster that does not belong to that sample.

[0021] The preferred method for analyzing the medication patterns of patient groups is as follows: retrieve prescriptions corresponding to patients from the medication information dataset, use Knime 5.4 to analyze the category and contribution of each Chinese herbal medicine in each patient group; select the top 20 most frequently used Chinese herbal medicines in each patient group, and refer to "Advanced Series of Traditional Chinese Medicine: Formulae" to determine the efficacy categories of core formulas and Chinese herbal medicines, thereby deriving the syndrome categories of the corresponding groups;

[0022] The preferred method for verifying the syndrome categories is as follows: using Knime 5.4, calculate the contribution of symptoms / signs in the corresponding group syndrome categories derived in step 1.4, where the contribution of a symptom / sign is defined as the frequency of occurrence of the symptom / sign divided by the total number of symptoms / signs observed in a certain group; and calculate the squared deviation (VSD) of symptom / sign values ​​between the same syndrome in traditional Chinese medicine and each group, where the squared deviation (VSD) is calculated according to the following equation:

[0023]

[0024] Among them, A i Indicates: The contribution of traditional Chinese medicine to the same syndrome, B i This indicates the contribution of symptoms / signs to the syndrome categories of the corresponding groups derived from the derivation. Finally, the correctness of the syndrome categories of the corresponding groups derived in step 1.4 is verified by the squared deviation VSD.

[0025] Based on the above classification method, the present invention further discloses its database system.

[0026] The database system includes the following modules:

[0027] 1. Literature Collection and Screening Module: This module is linked to a literature database or local database to collect...

[0028] We collected and screened clinical literature on traditional Chinese medicine treatments that were related to the disease, had clear therapeutic effects, and recorded case symptoms / signs.

[0029] 2. Document Information Extraction and Standardization Module: This module extracts the symptoms / signs and corresponding medication information for each case recorded in the above-mentioned documents and assigns a tree structure number to each symptom / sign for standardization in accordance with the current standard classification guidelines; then performs frequency analysis on the above information, and stores the analyzed data in the form of independent datasets in a MySQL database to create disease symptom / sign datasets and medication information datasets;

[0030] 3. Disease Patient Clustering Analysis Processing Module: Based on the above disease symptom / sign dataset, this processing module vectorizes the symptoms / signs by assigning tree structure numbers, constructs disease patient fingerprint vectors, uses the fingerprint vectors as input information for the clustering algorithm, and employs the unsupervised K-Means calculation model described in the above classification method to calculate cosine distance to measure the similarity between fingerprint vectors, and divides disease patients into different groups.

[0031] 4. Patient Group Medication Pattern Analysis Module: This module retrieves prescriptions corresponding to patients with diseases from the medication information dataset, uses Knime 5.4 to analyze the category and contribution of each Chinese medicine in each patient group, determines the Chinese medicine with high usage frequency in each group based on the contribution size, derives the core prescriptions and efficacy of each group, and thus derives the syndrome category of the corresponding group.

[0032] 5. Syndrome Category Verification Module: This module collects traditional Chinese medicine literature on syndromes that are consistent with the syndrome categories of the corresponding groups derived above, further extracts and standardizes the symptom / sign names in the literature, performs frequency analysis on the symptoms / signs, establishes a symptom / sign dataset of the traditional Chinese medicine syndrome, verifies the similarity between the symptom / signs of the syndrome categories of the corresponding groups derived above, and outputs the results.

[0033] In existing technologies, AI diagnostic and treatment models often overlook the relationships between symptoms / signs, leading to inaccurate data sources and limiting the depth of machine learning. Unsupervised clustering algorithms are a classic AI method that does not rely on predefined labels and has the potential to classify symptoms / signs based on inherent human principles. By eliminating human bias in the classification process, these algorithms can more accurately establish objective subtypes of diseases. In the MeSH library and the "Chinese Thesaurus of Traditional Chinese Medicine," each subject term can be described by one or more tree structure numbers, indicating its position in the hierarchical tree structure and its relationships with other subject terms. Combined with this invention, the tree structure numbers can be used to reveal potential deep relationships between symptoms / signs, meeting the requirements of applying TCM syndrome differentiation concepts to different types of diseases in modern medicine. This invention establishes a special AI computational model to mine all symptoms / signs of patients in clinical literature research, objectively classify patients, and uncover currently unknown syndrome types and medication patterns in TCM treatment, providing new pathways and methods for research and datafication in the field of TCM. Furthermore, this invention can also provide assistance in establishing new digital TCM diagnosis and treatment models, and promote the inheritance and development of TCM. Attached Figure Description

[0034] Figure 1 This is a schematic diagram of the technical process for classifying depression syndromes according to an embodiment of the present invention;

[0035] Figure 2 This is a schematic diagram of the literature screening process and analysis results for depression according to an embodiment of the present invention, wherein A is a flowchart of the literature screening process for depression, B is a graph of the frequency analysis results of the depression symptom / sign dataset, and C is a graph of the frequency analysis results of depression medication.

[0036] Figure 3This is a diagram showing the symptom classification results of patients with depression according to an embodiment of the present invention. A is a diagram showing the relationship between K-Means silhouette coefficient and number of clusters, B is a diagram showing the spatial distribution results of patients with depression after unsupervised clustering, and C is a diagram showing the symptom characteristics of patients with various types of depression.

[0037] Figure 4 This is a diagram illustrating the analysis results of medication patterns in various types of patients with depression according to embodiments of the present invention, where A is a cluster-herbal medicine network diagram and B is a schematic diagram of syndrome classification for various types of patients with depression.

[0038] Figure 5 This is a flowchart of the screening process for clinical literature on patients with Qi deficiency according to an embodiment of the present invention, and a graph showing the frequency analysis results of the dataset of symptoms / signs of patients with Qi deficiency.

[0039] Figure 6 This is a comparative illustration of symptoms / signs of nine subtypes of depression and patients with qi deficiency in this invention.

[0040] Figure 7 This is a flowchart of the Knime workflow according to an embodiment of the present invention;

[0041] Figure 8 This is a diagram illustrating the category and contribution of each herb in the treatment of patients in each group according to embodiments of the present invention;

[0042] Figure 9 This is a diagram illustrating the contribution of the present invention to the calculation of symptoms / signs in an embodiment of the present invention; Detailed Implementation

[0043] The following example, using the study of syndrome classification of depression and verification of Qi deficiency type depression, further illustrates the present invention. This embodiment does not constitute a limitation on the scope of protection of the present invention.

[0044] To delve deeper into disease classification in clinical diagnosis, this invention develops a comprehensive workflow database system utilizing unsupervised clustering algorithms. This system consists of the following five functional modules.

[0045] 1. Literature Collection and Screening Module: This module can collect and screen clinical literature related to depression from various databases according to predetermined criteria.

[0046] 2. Symptom and Sign Extraction and Standardization Module: This module is used to extract and standardize symptoms and signs recorded in selected literature. Referring to the MeSH database and the *Chinese Traditional Medicine Subject Headings*, an appropriate tree structure number is assigned to each symptom / sign.

[0047] 3. Depression Patient Clustering Module: This module uses an unsupervised K-Means computation model to divide patients into different groups.

[0048] 4. Traditional Chinese Medicine Extraction and Standardization Module: this module extracts and standardizes the herbal medicines for treating depression recorded in literatures, and compiles and analyzes the herbal prescriptions and medication rules corresponding to each group of patients.

[0049] 5. Symptom / sign comparison module for each group of patients and qi-deficiency patients: this module is used to compare the similarity between the selected clusters and traditional Chinese medicine diagnostic typing, verify the scientificity and rationality of the method, and evaluate the efficacy of the system.

[0050] The system uses a specific unsupervised clustering K-Means algorithm model to learn and process literature information. The modeling method is as follows:

[0051] 1. Data collection and screening steps

[0052] 1.1. Literature collection and data cleaning

[0053] In databases including CNKI, VIP, Chinese Medical Information Network, Wanfang Data, PubMed, Web of Science, etc.

[0054] Comprehensive retrieval is conducted on clinical literatures of traditional Chinese medicine for depression. The Chinese search term is "depression", and the English search terms are: "depression", "depressive disorders", "single episode depressive disorder", "current depressive disorder", "dysthymic disorder", "mixed depressive and anxiety disorder" and "premenstrual dysphoric disorder". Articles are included according to the following criteria:

[0055] (1) Study type: clinical studies focusing on the treatment of depression with traditional Chinese medicine;

[0056] (2) Participants: depression patients of any age or gender;

[0057] (3) Intervention method: studies using traditional Chinese medicine alone or in combination with western medicine;

[0058] (4) Outcome: with clear reports on therapeutic effect.

[0059] Databases including CNKI, VIP, Sino Med, Wangfang, PubMed and Web of Science are used to retrieve qi-deficiency related literatures published before April 2023. The inclusion criteria of literatures are as follows:

[0060] (1) Research type: Clinical literature related to Qi deficiency;

[0061] (2) Participants: Patients of any age or gender with "Qi deficiency";

[0062] (3) Intervention: Unrestricted;

[0063] (4) Result: No limit.

[0064] The obtained literature was processed using Knime 5.4, and the data tables were connected using uniform columns. To ensure data accuracy, redundant literature was removed by comparing the title and abstract columns according to the inclusion and exclusion criteria mentioned above. The Knime workflow used in this process is shown in the diagram below. Figure 7 As shown. Import the list of references into Zotero 6.0 to download the full text of the articles, creating a local reference database for data extraction.

[0065] 1.2. Data Extraction and Standardization

[0066] Cases conforming to TCM syndrome differentiation and treatment were considered as independent samples, and their symptoms / signs and prescriptions were extracted and standardized. For symptom / sign standardization, guidelines provided by ICD-11, the Chinese Thesaurus of Traditional Chinese Medicine (CTMMMSH), and the SymMap database were referenced. Tree structure numbers for each symptom or sign were collected from the CTMMMMSH and MeSH databases. If a symptom or sign did not have a specified tree structure number, a final decision was made according to the following rules: (1) if a synonym or similar symptom / sign exists, the tree structure number of the synonym or similar symptom / sign is used; or (2) if no alternative term is found, the parent tree structure number in the tree structure is used. Herbal names in prescriptions were standardized according to the Pharmacopoeia of the People's Republic of China and the Chinese Materia Medica. The data extraction and cleaning processes mentioned above were completed independently by two researchers, and any disagreements were resolved through consultation with a third researcher. Frequency analysis of symptoms / signs and herbs in the depression and Qi deficiency samples was performed using Knime 5.4. The analyzed data were stored in a MySQL database as independent datasets, creating datasets for depressive symptoms / signs, depressive herbal medicines, and qi deficiency symptoms / signs.

[0067] 1.3. Establishment of the Symptom / Sign Hierarchical Relationship Matrix

[0068] Use a Python 3.9 script to convert the hierarchical relationships of the tree structure numbers to Boolean hierarchies, and calculate them according to the following formula, representing them as "0" or "1":

[0069]

[0070] Where x and y are the tree structure numbers of the symptoms / signs of depression.

[0071] The above Boolean values ​​form a two-dimensional matrix, representing the hierarchical relationship of the tree structure of depressive symptoms / signs. This matrix can be used to determine whether a given x belongs to y in the hierarchical structure.

[0072] 1.4. Construction of Patient Fingerprint Vectors

[0073] The symptoms / signs of each patient with depression (sample) are transformed into a phenotypic vector, where the values ​​in the vector are represented as 0 or 1. A value of 1 indicates that the patient exhibits a specific symptom / sign, while a value of 0 indicates that the symptom / sign is not present. All hierarchical information of a single patient with depression is captured by performing a dot product between the phenotypic vector and the symptom / sign hierarchy matrix. The resulting dot product matrix is ​​further flattened and dimensionality reduced to create a fingerprint vector for the patient with depression.

[0074] 1.5. Unsupervised Machine Learning

[0075] To classify patients with depression based on their symptoms / signs and other information, the K-Means clustering algorithm was implemented using the sklearn.cluster.KMeans package in Python 3.9. Fingerprint vectors, obtained by flattening the dot product matrix of all samples and pruning the dimensions with values ​​of 0, were used as input to the clustering algorithm. During K-Means clustering, cosine distance was calculated to measure the similarity between fingerprint vectors. The results of unsupervised clustering were visualized as 3D scatter plots, allowing for a visual representation of the formed clusters. Furthermore, the proportion of samples in each cluster was analyzed using Origin 2021 software. The silhouette coefficient was chosen as the evaluation metric for the clustering algorithm. It was calculated using the following formula:

[0076]

[0077] In the formula, s:

[0078] The average silhouette coefficient of all samples; a: the average intra-cluster distance of each sample; b: the distance between a sample and the nearest cluster that does not belong to that sample.

[0079] 1.6 Analysis of Medication Use Patterns in Clusters of Depression Patients

[0080] Prescriptions corresponding to different groups of patients with depression were retrieved from the herbal medicine dataset for depression. Knime 5.4 was used to analyze the category and contribution of each herb in the treatment of patients in each group. To determine the core prescriptions, the top 20 most frequently used herbs in each cluster were selected. These core prescriptions represent commonly used prescriptions for treating patients with similar clinical manifestations in specific clusters according to traditional Chinese medicine theory. The candidate list of core prescriptions was obtained from "Advanced Series of Traditional Chinese Medicine: Prescriptions". The selection criteria for core prescriptions are as follows: (1) If a candidate prescription consists of 4 or fewer herbs, all herbs must match the composition of the core prescription; (2) If a candidate prescription consists of more than 4 herbs, at least 80% of the herbs must match the composition of the core prescription. The contributions of herbs and prescriptions were calculated using Knime 5.4 (results are shown in the figure). Figure 8 (As stated above). The contribution of a herb is defined as the frequency of that herb divided by the total number of herbs used to treat a particular group. The contribution of a prescription is defined as the sum of the contributions of each herb in the prescription. The relationship between patient groups and herbs / core prescriptions in the treatment of each group was visualized using Cytoscape 3.8. In addition, the efficacy categories of these herbs were analyzed according to the definitions in the "Advanced Series of Traditional Chinese Medicine: Materia Medica".

[0081] 1.7 Comparison of the similarity of symptoms / signs between Qi deficiency and depression

[0082] The contribution of symptoms / signs was calculated using Knime 5.4 (results are shown below). Figure 9 The contribution of a symptom / sign is defined as the frequency of its occurrence divided by the total number of symptoms / signs observed in a given cluster. To calculate the value-squared deviation (VSD) between the Qi deficiency symptom / sign dataset and each depression cluster, the contribution values ​​of all symptoms / signs were used. The VSD is calculated according to the following equation:

[0083]

[0084] A i B: Contribution to Qi deficiency syndrome; i The contribution of symptoms / signs of patients in a particular group of depression.

[0085] 1.8 Medical Case Verification

[0086] Case studies related to patients with depression are downloaded from the medical record cloud platform, standardized, and then input into the model for training.

[0087] The specific classification method process is as follows: Figure 1 As shown, the steps are as follows:

[0088] 1. Collect and screen clinical or experimental literature related to depression.

[0089] A total of 685,244 articles related to depression published before April 2023 were retrieved through searches of databases including CNKI, VIP, China Medical Information Network, Wanfang Data, PubMed, and Web of Science. The screening process was as follows: Figure 2 -A. Specific keywords were selected based on the inclusion criteria, and the titles and abstracts of the articles were screened based on these keywords, resulting in a total of 26,684 full-text articles. After comprehensive evaluation, 3,522 high-quality clinical articles on traditional Chinese medicine (TCM) treatment of depression were included for further analysis. To analyze the data, the names of symptoms / signs and TCM herbs mentioned in the articles were extracted and standardized. Subsequently, frequency analysis was performed on the symptoms / signs and herbs, thus creating a dataset of depression symptoms / signs and a dataset of TCM herbs used by patients with depression. Among the frequently observed symptoms / signs of depression, the top 30 were identified, including insomnia, depression, boredom, anorexia, tight pulse, anger, irritability, anxiety, depression, chest pain, etc. Figure 2 -B (as shown). The 30 most commonly used Chinese herbs for treating depression were determined based on their frequency of occurrence. These herbs include Bupleurum, Licorice, Peony, Ligusticum chuanxiong, Curcuma longa, Angelica sinensis, Poria cocos, Pinellia ternata, etc. (e.g., Figure 2 -C is shown).

[0090] 2. Unsupervised clustering of patients with depression based on symptoms and signs

[0091] All symptom / sign information was extracted from the literature and arranged into a tree structure to form a table. Each row in the table was converted into Boolean data to indicate the presence or absence of a symptom / sign. The correlation information between symptoms / signs was then integrated by calculating the dot product of the hierarchical relationship matrix of symptoms / signs. Subsequently, the dot product matrix was flattened and converted into a fingerprint vector for each patient with depression, and then cluster analysis was performed using the K-Means algorithm. The results show that the relationship between the K-Means silhouette coefficient and the number of clusters is as follows: Figure 3 As shown in Figure A, the silhouette coefficient reaches a local maximum at nine clusters, indicating that the optimal subtype classification model for patients with depression is to divide them into nine subtypes. The three-dimensional spatial distribution of these nine subtypes of depression is as follows: Figure 3 -B is shown. Figure 3 -C shows the top 10 symptoms / signs observed in each of the 9 clusters. Each cluster exhibits significant differences from the other clusters. For example, in cluster 6, the most common symptoms / signs are depression, boredom, anxiety, sleep-wake disorder, sleep initiation and maintenance disorder, anorexia, nausea, irritability, pessimism, and depression. Conversely, cluster 7's symptoms / signs include sweating, insomnia, arrhythmia, vivid dreams, dyspnea, weakness, weak pulse, white tongue coating, and abdominal and chest pain.

[0092] 3. The medication patterns of patients with nine subtypes of depression were analyzed to deduce the syndrome types to which the nine subtypes belong.

[0093] To further examine the similarities and differences in prescriptions for each type of depression patient, the next step involved analyzing the core prescriptions. Prescriptions for depression patients were obtained from a dataset of herbs used to treat depression. Referring to the results of unsupervised clustering, the 20 most frequently used herbs in each cluster were selected, and the core formula for each cluster was derived. The clustering-formula / herb network is shown below. Figure 4 -A is shown. The analysis results show that clusters 1, 2, 4, 5, 6, 8, and 9 primarily use Dang Gui Shao Yao San (a formula for promoting blood circulation and removing blood stasis), Liu Wei Di Huang Wan (a formula for nourishing Yin), Si Ni San (a formula for regulating Qi), and Liu Jun Zi Tang (a formula for tonifying Qi). Specific additions and subtractions of herbs vary among different patients. For example, Si Ni San was identified as the core formula for cluster 3, while Liu Jun Zi Tang was identified as the core formula for cluster 7. Further analysis of the efficacy of each type of herb revealed that clusters 1, 2, and 6 mainly use herbs for promoting blood circulation and removing blood stasis. In contrast, clusters 3, 4, 5, 8, and 9 mainly use herbs with Qi-regulating effects, while cluster 7 primarily focuses on tonifying Qi (such as...). Figure 4 -B is shown).

[0094] 4. Discovery of a new subtype of depression characterized primarily by Qi deficiency

[0095] Unsupervised clustering results showed that patients in cluster 7 most frequently used Qi-tonifying formulas and traditional Chinese medicines, suggesting a potential relationship between this subtype and patients with Qi deficiency in Traditional Chinese Medicine (TCM). To further explore the medical significance of cluster 7, symptoms / signs of patients diagnosed with Qi deficiency in TCM were collected and compared with the sample from cluster 7. 98,819 articles related to Qi deficiency were retrieved from databases including CNKI, VIP, China Medical Information Network, Wanfang, PubMed, and Web of Science, and were then filtered. The screening process is as follows: Figure 5shown in Part A. Screening keywords are selected according to the inclusion criteria, and the titles and abstracts of the literature are screened using the specified keywords. A total of 5847 articles meeting the full-text screening criteria are obtained. After comprehensive full-text screening, 2608 high-quality clinical studies related to qi deficiency are identified. Further extraction and standardization of symptom / sign names related to the literature are conducted, frequency analysis is performed thereon, and a symptom / sign dataset of patients with qi deficiency is established. The 30 most common symptoms / signs in patients with qi deficiency are selected from the dataset, including weakness, dyspnea, thready pulse, pale tongue, WTC, anorexia, arrhythmia, sweating, fatigue, etc. (as Figure 5 shown in Part B).

[0096] All symptoms and signs of patients with 9 depression subtypes are compared with the top 30 symptoms / signs of patients with qi deficiency, based on their contribution to clustering. The heatmap of symptoms and signs of 9 clusters and qi deficiency patients shows that the distribution of symptoms and signs in cluster 7 is similar to that of the qi deficiency group (as Figure 6 shown in Part A). In addition, statistical analysis of VSD values shows that the difference between cluster 7 and qi deficiency patients is smaller than that of the other 8 categories (0.008 for cluster 7, 0.015, 0.018, 0.017, 0.037, 0.013, 0.019, 0.020 and 0.054 for categories 1-6, 8 and 9). These results indicate that cluster 7 may represent a depression subtype with qi deficiency as the main manifestation (as Figure 6 shown in Part B).

[0097] 5. Verification of personalized diagnosis and treatment cases

[0098] A total of 36 cases of depression patient information were collected from the medical case cloud platform for verification, and the diagnosis and treatment information is shown in the table below. The traditional Chinese medicine syndromes diagnosed in the cases include deficiency of both heart and spleen, the symptoms include insomnia and dreaminess, and the Western medicine diagnosis includes depression. Preprocessing of personalized traditional Chinese medicine diagnosis and treatment information for depression mainly includes unifying medical term names and removing unnecessary symbols, for example, "gloomy and depressed" and "low mood" are unified into "low mood". As a result, after extracting and training the traditional Chinese medicine diagnosis and treatment information of 5 depression patients in the qi deficiency group, the diagnosis and treatment information of 3 depression patients fell into the 7th category, with a true positive rate of 60%. Among 22 patients in the non-qi deficiency group, 1 patient fell into the 7th category, with a false positive rate of 4.5%.

[0099] Personalized traditional Chinese medicine diagnosis and treatment information table for depression

[0100]

Claims

1. A method for classifying syndromes in Traditional Chinese Medicine, characterized by comprising the following steps: 1.1 Literature Collection and Screening: Literature related to the disease and possessing clear therapeutic effects was collected and screened. Clinical or experimental literature on traditional Chinese medicine treatment that records information on symptoms and signs of patients; 1.2 Extraction and Standardization of Literature Information: The symptoms, signs, and corresponding medication information for each case recorded in the aforementioned literature were extracted and standardized; frequency analysis was then performed on the above information to create... Disease symptom and sign datasets and medication information datasets; 1.3 Disease Patient Clustering: Based on the above disease symptom and sign dataset, unsupervised K-means clustering was used. The Means calculation model categorizes disease patients into different groups; The method for establishing the unsupervised K-Means computation model includes the following steps: A. Establishing the symptom or sign hierarchy matrix: Convert the MeSH tree numbering of the membership relationships into Boolean membership relationships. For each relationship, calculate according to the following formula and represent it as "0" or "1". Where x and y are the MeSH tree numbers of disease symptoms or signs; their Boolean values ​​form a two-dimensional matrix representing the hierarchical relationship of the MeSH tree numbers of disease symptoms or signs; this matrix can be used to determine whether a given x belongs to y in the hierarchical structure; B. Construction of patient fingerprint vectors: Converting the symptoms / signs of each disease patient into phenotypic vectors. The values ​​of the vector are represented as 0 or 1, where a value of 1 indicates that the patient exhibits specific symptoms / signs, and a value of 0 indicates that the patient exhibits specific symptoms / signs. This indicates that the symptom / sign is not present; By performing a cross-reference between phenotypic vectors and symptom / sign grading matrix The dot product captures all levels of information from a single disease patient, and the resulting dot product matrix is ​​further flattened. Dimensionality reduction was used to create fingerprint vectors for patients with depression. C. Unsupervised Machine Learning: Using sklearn.cluster.KMeans in Python 3.9 scripts This package implements the K-Means clustering algorithm, using fingerprint vectors as input. During K-Means clustering, cosine distance is calculated to measure the similarity between fingerprint vectors. The results of unsupervised clustering are visualized as 3D scatter plots, allowing for a visual representation of the formed clusters. Furthermore, the proportion of samples in each cluster is analyzed using Origin 2021 software, and the silhouette coefficient is selected as the evaluation metric for the clustering algorithm, calculated using the following formula: In the formula, s: the average silhouette coefficient of all samples; a: the average intra-cluster distance of each sample; b: the sample and... The distance between the nearest clusters that do not belong to this sample; 1.4 Analysis of medication patterns in patient groups and deduction of syndrome categories for corresponding groups: Based on the above medication information dataset, the most frequently used Chinese medicines in each group are identified, and the core prescriptions and their additions and subtractions for each group are deduced, thereby deriving the syndrome categories for the corresponding groups. 1.5 Verification of syndrome categories: Compare the similarity of symptoms and signs between different groups and the same syndrome in traditional Chinese medicine to verify the syndrome categories derived above.

2. The method as described in claim 1, characterized in that... Step 1.1 The literature sources include CNKI, VIP, Sino Med, Wanfang, PubMed, or Web of Science databases.

3. The method as described in claim 1, characterized in that, the literature screening method in step 1.1 is as follows: Knime 5.4 is used to process the collected literature, and the data table is connected using a unified column; redundant literature is eliminated by comparing the title and abstract columns according to the literature inclusion criteria; the improved literature list is imported into Zotero 6.0, and the full text of the articles is downloaded to create a local literature database for data extraction; wherein, the literature inclusion criteria are: 3.1 Research type: Clinical research on the treatment of depression with traditional Chinese medicine; 3.2 Participants: Patients of any age or gender; 3.3 Intervention: Studies on the use of traditional Chinese medicine alone or in combination with Western medicine; 3.4 Results: Research literature with clear reports of treatment efficacy.

4. The method as described in claim 1, characterized in that... Step 1.2 The standardization method for the literature information is as follows: For the standardization of medication information, refer to ICD-11 and the "Chinese Traditional Medicine Subject Headings". Using guidelines provided by the SymMap database, the tree structure number of each symptom or sign is collected from the "Chinese Traditional Medicine Subject Headings" and the Medical Subject Headings Database, and the names of Chinese medicines or herbs used in medication information are standardized based on the "Pharmacopoeia of the People's Republic of China" and "Chinese Materia Medica".

5. The method as described in claim 1, characterized in that, The patient group described in step 1.4 is used The specific method for drug pattern analysis is as follows: retrieve prescriptions corresponding to patients with diseases from the drug use information dataset. The method was analyzed using Knime 5.4 to determine the category and contribution of each traditional Chinese medicine in each patient group; Select each The top 20 most frequently used traditional Chinese medicines in each patient group were selected, with reference to "Advanced Series of Traditional Chinese Medicine: Formulas". The study determines the core formula and efficacy categories of Chinese medicine, thereby deriving the corresponding syndrome categories; The screening criteria for the core formula are as follows: (1) If the candidate formula consists of 4 or fewer ingredients... (1) If the composition of Chinese medicine is as follows, then all Chinese medicines must match the components of the core prescription; (2) If the candidate prescription is composed of If the formula consists of four or more Chinese herbs, then at least 80% of the herbs must match the components of the core formula. use Knime 5.4 calculates the contribution of Chinese herbal medicines and prescriptions. The contribution of a Chinese herbal medicine is defined as the frequency of that medicine divided by the total number of Chinese herbal medicines used to treat a specific cluster. The contribution of a prescription is defined as the sum of the contributions of each Chinese herbal medicine in the prescription. Cytoscape 3.8 is used to visualize the relationship between patient clusters and the Chinese herbal medicines and core prescriptions in the treatment of each patient group.

6. The method as described in claim 1, characterized in that, The syndrome categories mentioned in step 1.5 The verification method is as follows: Use Knime 5.4 to calculate the syndrome class of the corresponding group derived in step 1.

4. The contribution of symptoms / signs to the overall health status, where the contribution of a symptom / sign is defined as the occurrence of that symptom / sign. The frequency is divided by the total number of symptoms / signs observed in a given group; and the same syndrome in traditional Chinese medicine is compared with... The squared deviation (VSD) of symptom / sign values ​​between each group is calculated according to the following equation: calculate: Where Ai represents the contribution of the same syndrome in traditional Chinese medicine, Bi represents the contribution of symptoms / signs in the syndrome category of the corresponding group derived by the deduced method, and finally the correctness of the syndrome category of the corresponding group derived in step 1.4 is verified by the squared deviation VSD.

7. A TCM syndrome classification database system, characterized in that... The database system includes: 7.1 Literature Collection and Screening Module: This module is linked to a literature database or local database to collect and screen clinical literature on traditional Chinese medicine treatment that is related to the disease, has a clear therapeutic effect, and records case symptoms / signs. 7.2 Document Information Extraction and Standardization Module: This module extracts the symptoms / signs and corresponding medication information for each case recorded in the above-mentioned documents and assigns MeSH tree numbers to each symptom / sign for standardization in accordance with the current standard classification guidelines; then performs frequency analysis on the above information, and stores the analyzed data in the form of independent datasets in a MySQL database to create disease symptom / sign datasets and medication information datasets; 7.3 Disease Patient Cluster Analysis Processing Module: This processing module clusters patients based on the above-mentioned disease symptoms / signs. Based on the data, an unsupervised K-Means computational model was used to divide disease patients into different groups; The construction method of the unsupervised K-Means computation model includes the following steps: A. Establishing the symptom or sign hierarchy matrix: Convert the MeSH tree numbering of the membership relationships into Boolean membership relationships. For each relationship, calculate according to the following formula and represent it as "0" or "1". Where x and y are the MeSH tree numbers of disease symptoms or signs; their Boolean values ​​form a two-dimensional matrix representing the hierarchical relationship of the MeSH tree numbers of disease symptoms or signs; this matrix can be used to determine whether a given x belongs to y in the hierarchical structure; B. Construction of patient fingerprint vectors: The symptoms / signs of each disease patient are transformed into phenotypic vectors, where the values ​​of the vectors are represented as 0 or 1, where a value of 1 indicates that the patient exhibits a specific symptom / sign, and a value of 0 indicates that the patient does not have that symptom / sign; by performing a dot product between the phenotypic vector and the symptom / sign hierarchical relationship matrix, all hierarchical information of a single disease patient is captured, and the resulting dot product matrix is ​​further flattened and dimensionality reduced to create fingerprint vectors for depressed patients; C. Unsupervised machine learning: using sklearn.cluster in Python 3.9 scripts. The KMeans package implements the K-Means clustering algorithm, using fingerprint vectors as input information for the clustering algorithm. In the K-Means clustering process, cosine distance is calculated to measure the similarity between fingerprint vectors; the results of unsupervised clustering are visualized as 3D scatter plots, allowing for a visual representation of the formed clusters; furthermore, the proportion of samples in each cluster is analyzed using Origin 2021 software, and the silhouette coefficient is selected as the evaluation metric for the clustering algorithm, calculated using the following formula: In the formula, s: the average silhouette coefficient of all samples; a: the average intra-cluster distance of each sample; b: the sample and... The distance between the nearest clusters that do not belong to this sample; 7.4 Patient Group Medication Pattern Analysis Module: This module retrieves prescriptions corresponding to patients with diseases from the medication information dataset, uses Knime 5.4 to analyze the category and contribution of each Chinese medicine in each patient group, determines the Chinese medicine with high usage frequency in each group based on the contribution size, derives the core prescriptions and efficacy of each group, and thus derives the syndrome category of the corresponding group. 7.5 Syndrome Category Verification Module: This module collects traditional Chinese medicine literature on syndromes that are consistent with the syndrome categories of the corresponding groups derived above, further extracts and standardizes the symptom / sign names in the literature, performs frequency analysis on the symptoms / signs, establishes a symptom / sign dataset of the traditional Chinese medicine syndrome, verifies the similarity of the symptoms / signs of the syndrome categories of the corresponding groups derived above, and outputs the results.