Method and system for constructing tongue coating-syndrome type associated knowledge graph

By constructing a knowledge graph linking tongue coating and syndrome types using topological manifold learning technology, the problem of heterogeneous data association between tongue images and syndrome descriptions is solved, enabling personalized diagnostic reasoning and improved interpretability, thereby enhancing the accuracy and scalability of TCM diagnosis.

CN120895267AInactive Publication Date: 2025-11-04SHENZHEN TRADITIONAL CHINESE MEDICINE HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511409893.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-11-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In existing technologies, it is difficult to effectively link heterogeneous data such as tongue images and syndrome descriptions, and the knowledge base lacks interpretability and reasoning ability, making it impossible to achieve personalized diagnosis and treatment.

Method used

Using topological manifold learning technology, multi-level topological and semantic topological features of tongue images and TCM syndrome descriptions are extracted. These features are then mapped to a common embedding space through a topological manifold alignment model to construct a knowledge graph of tongue coating syndrome associations. The relationships between nodes are determined based on topological similarity, and knowledge from the TCM domain is introduced to optimize the graph structure, enabling personalized diagnostic reasoning.

Benefits of technology

It effectively captures the inherent structural characteristics of tongue images and syndrome descriptions, achieves precise alignment and fusion of tongue image and syndrome data, improves the structural characteristics and semantic associations of knowledge graphs, enhances the interpretability and accuracy of syndrome differentiation and treatment, and has good scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895267A_ABST
    Figure CN120895267A_ABST
Patent Text Reader

Abstract

The invention discloses a tongue coating-syndrome type associated knowledge graph construction method and system, and belongs to the field of traditional Chinese medicine informatization, and the method comprises the steps: obtaining tongue picture image data and traditional Chinese medicine syndrome description data; extracting multi-level tongue picture topological features from the tongue picture image by using a feature extraction network, and extracting multi-level semantic topological features from the syndrome description by using a semantic analysis network; constructing a topological manifold alignment model, performing parameterized representation on the manifolds through a local linear embedding technology, designing adaptive distance measurement to calculate a geodesic distance, determining manifold anchor points based on tongue picture-syndrome pairs labeled by experts, and realizing alignment of the two manifolds; taking the feature representation as a knowledge graph node, determining an association relationship based on topological similarity, and constructing a tongue coating syndrome type association knowledge graph; tongue image topological features of a new patient are extracted and mapped to a knowledge graph, similar syndrome types are identified through topological path reasoning, a syndrome differentiation result is generated in combination with individual conditions of the patient, the syndrome differentiation accuracy is improved, and personalized syndrome differentiation treatment is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of traditional Chinese medicine informatization, in particular to a tongue fur-syndrome type association knowledge graph construction method and system, and specifically relates to a method and system for associating tongue image with traditional Chinese medicine syndrome description by using topological manifold learning technology, constructing a knowledge graph, and realizing individualized syndrome differentiation and treatment. BACKGROUND

[0002] Tongue diagnosis is an important part of traditional Chinese medicine diagnosis, which can determine the health status and disease type of the human body by observing the characteristics of tongue body and tongue fur. Traditional tongue diagnosis mainly relies on the experience of doctors, which has strong subjectivity, non-uniform standards, and knowledge inheritance difficulties. With the development of artificial intelligence technology, combining tongue diagnosis with modern technology to construct a knowledge base of tongue fur and syndrome type association has become an important direction of traditional Chinese medicine modernization research.

[0003] In the prior art, simple image processing and machine learning methods are usually used to extract tongue image features, text analysis methods are used to process syndrome description, and then statistical correlation or simple neural networks are used to establish the association between tongue image and syndrome type. These methods have the following problems: first, they cannot effectively capture the internal structural characteristics of tongue image and syndrome description; second, it is difficult to fuse and align the two heterogeneous data of tongue image and syndrome; third, the constructed knowledge base lacks explainability and reasoning ability; fourth, it cannot realize individualized syndrome differentiation and treatment.

[0004] At present, knowledge graph technology has been widely applied in the field of medical health, but in the field of traditional Chinese medicine tongue diagnosis, there is still a lack of effective knowledge graph construction methods, especially a lack of technical solutions that can process two kinds of heterogeneous data of image and text and establish their deep semantic association. SUMMARY

[0005] The purpose of the present application is to provide a tongue fur-syndrome type association knowledge graph construction method and system, which solves the technical problems of the prior art that tongue image and syndrome description heterogeneous data are difficult to effectively associate, and the knowledge base lacks explainability and reasoning ability.

[0006] The present application provides a tongue fur-syndrome type association knowledge graph construction method, which comprises:

[0007] Obtaining tongue image data and traditional Chinese medicine syndrome description data;

[0008] Extracting topological features of the tongue image data and semantic topological features of the traditional Chinese medicine syndrome description data; wherein the extraction comprises:

[0009] Extracting multi-level tongue image topological features from the tongue image data by using a feature extraction network, the multi-level tongue image topological features including tongue body shape features, tongue fur distribution features, and tongue body texture features;

[0010] extracting, by using a semantic analysis network, multi-level semantic topology features from the TCM syndrome description data, the multi-level semantic topology features including symptom level features, syndrome level features, and syndrome type level features;

[0011] constructing a topology manifold alignment model to map the tongue appearance topology features and the semantic topology features to a common embedding space; wherein the constructing includes:

[0012] constructing manifold structures for the tongue appearance topology features and the semantic topology features, respectively;

[0013] realizing alignment between the two manifolds based on topology preservation constraints;

[0014] optimizing feature representations in the common embedding space through topology-aware contrastive learning;

[0015] constructing a tongue fur syndrome type correlation knowledge graph based on the feature representations in the common embedding space; wherein the constructing includes:

[0016] taking the feature representations in the common embedding space as node features of the knowledge graph;

[0017] determining correlation relationships between nodes based on topology similarity;

[0018] introducing TCM domain knowledge to optimize the structure of the graph;

[0019] realizing individualized syndrome differentiation reasoning based on the tongue fur syndrome type correlation knowledge graph; wherein the realizing includes:

[0020] mapping tongue appearance images of new patients to the tongue fur syndrome type correlation knowledge graph;

[0021] recognizing similar syndrome types based on topology path reasoning;

[0022] generating individualized syndrome differentiation results in combination with individual conditions of patients.

[0023] As a preferred, the extracting the topology features of the tongue appearance image data specifically includes:

[0024] extracting basic visual features of the tongue appearance image data by using an efficient convolutional neural network;

[0025] constructing a topology-aware layer to calculate topology invariants of the tongue appearance image data at different scales, the topology invariants including connected components, voids, and high-dimensional cavities;

[0026] designing multi-scale filters to capture topology features of the tongue appearance image data at multiple scale parameters, and generating a scale-topology feature matrix;

[0027] The attention mechanism integrates the topological features of each level to generate a unified topological feature representation of tongue appearance.

[0028] As preferred, the extracting semantic topological features of the TCM syndrome description data specifically includes:

[0029] The TCM syndrome description data is converted into a semantic vector using a bidirectional long short-term memory network.

[0030] A neighborhood relation graph of the TCM syndrome description data is constructed based on semantic similarity.

[0031] A semantic topological space of the TCM syndrome description data is constructed through the neighborhood relation graph, which includes semantic connected components and semantic voids.

[0032] The structure information of the semantic topological space is enriched in combination with TCM syndrome correlation knowledge.

[0033] Through topological persistence analysis, the semantic topological features of different levels are integrated into a unified semantic topological representation.

[0034] As preferred, the constructing topological manifold alignment model specifically includes:

[0035] The tongue appearance topological features are regarded as data points on a first high-dimensional manifold, and the semantic topological features are regarded as data points on a second high-dimensional manifold.

[0036] Through local linear embedding technology, the two high-dimensional manifolds are represented in low-dimensional parameters.

[0037] An adaptive distance metric is designed to calculate the geodesic distance between data points on the two high-dimensional manifolds.

[0038] Based on expert-labeled tongue appearance-syndrome pairs, corresponding anchor points on the two high-dimensional manifolds are determined.

[0039] A local structure preserving constraint is designed to ensure that the local neighborhood relationship remains unchanged during the alignment process.

[0040] Through a global optimization algorithm, the overall structural difference between the two high-dimensional manifolds is minimized.

[0041] As preferred, the optimizing feature representation in the common embedding space through topologically aware contrastive learning specifically includes:

[0042] The dimensional homology groups of the tongue appearance topological feature space and the semantic topological feature space are calculated.

[0043] A mapping function that preserves the homology group structure is designed to ensure that the topological structure remains unchanged during the mapping process.

[0044] Select tongue and syndrome pairs with similar topological structures as positive samples, and select samples with similar topological structures but different semantics as difficult negative samples;

[0045] Design a contrast loss function that integrates topological consistency constraints;

[0046] Based on the topological complexity, dynamically adjust the gradient scale to avoid overfitting simple samples;

[0047] Construct a bidirectional topological preserving mapping network to ensure the cyclic consistency of the mapping.

[0048] As preferred, the constructing a tongue fur syndrome type association knowledge graph based on the feature representation in the common embedding space specifically includes:

[0049] Define tongue features, syndrome descriptions, and syndrome types as knowledge graph nodes;

[0050] Embed the topological features in the common embedding space as the initial feature representation of the nodes;

[0051] Define a corresponding attribute set for each type of node, and initialize the node attribute values based on domain knowledge and data analysis;

[0052] Define relationship types such as represents, belongs to, and is associated with as knowledge graph edges;

[0053] Based on topological similarity, construct the association strength between entities to determine the weight of the edges;

[0054] Verify the rationality of the automatically constructed relationships through expert knowledge;

[0055] Based on topological redundancy analysis, prune redundant relationships in the graph;

[0056] Strengthen the key semantic paths in the graph to improve reasoning efficiency.

[0057] As preferred, the individualized syndrome differentiation reasoning based on the tongue fur syndrome type association knowledge graph specifically includes:

[0058] Extract the topological feature representation of the new patient's tongue;

[0059] Map the topological feature representation of the new patient's tongue to the common embedding space;

[0060] Based on topological similarity, retrieve similar cases in the tongue fur syndrome type association knowledge graph;

[0061] Generate a set of possible reasoning paths based on the knowledge graph structure;

[0062] Design a path scoring function that integrates topological structure information;

[0063] Select the reasoning path with the highest score for syndrome identification;

[0064] Analyze the distribution of similar cases to generate a personalized syndrome result;

[0065] Combine individual patient conditions to generate personalized treatment recommendations.

[0066] As a preferred embodiment, it also includes the step of dynamically updating the tongue syndrome type association knowledge graph:

[0067] Collect new tongue image data and syndrome description data;

[0068] Extract the topological features of the new data;

[0069] Map the topological features of the new data to the public embedding space;

[0070] According to expert feedback, verify and adjust the association of the new data in the knowledge graph;

[0071] Based on the time decay mechanism, balance the weight of new and old knowledge;

[0072] Periodically re-optimize the overall structure of the tongue syndrome type association knowledge graph.

[0073] As a preferred embodiment, before obtaining tongue image data and TCM syndrome description data, it also includes a data preprocessing step:

[0074] Standardize the tongue image data and adjust the image size to a uniform specification;

[0075] Extract the tongue area through tongue positioning algorithm;

[0076] Quality assessment of the tongue image data, filtering low-quality images;

[0077] Text cleaning of the TCM syndrome description data to remove noise and redundant information;

[0078] Standardize the TCM syndrome description data according to TCM terminology standards;

[0079] Data augmentation of the tongue image data and the TCM syndrome description data to expand the training samples.

[0080] Tongue syndrome type association knowledge graph construction system, comprising:

[0081] Data acquisition and preprocessing module for obtaining tongue image data and TCM syndrome description data, and standardizing the data;

[0082] a topology feature extraction module, configured to extract topology features of the tongue image data and semantic topology features of the TCM syndrome description data; wherein the topology feature extraction module comprises:

[0083] a tongue feature extraction submodule, configured to extract multi-level tongue topology features from the tongue image data by using a feature extraction network;

[0084] a syndrome feature extraction submodule, configured to extract multi-level semantic topology features from the TCM syndrome description data by using a semantic analysis network;

[0085] a multi-modal feature alignment module, configured to construct a topology manifold alignment model and map the tongue topology features and the semantic topology features to a common embedding space; wherein the multi-modal feature alignment module comprises:

[0086] a manifold construction submodule, configured to construct manifold structures for the tongue topology features and the semantic topology features, respectively;

[0087] a manifold alignment submodule, configured to realize alignment between the two manifolds based on topology preservation constraints;

[0088] a contrast learning submodule, configured to optimize feature representations in the common embedding space through topology-aware contrast learning;

[0089] a knowledge graph construction module, configured to construct a tongue coating syndrome type correlation knowledge graph based on the feature representations in the common embedding space; wherein the knowledge graph construction module comprises:

[0090] a graph construction submodule, configured to take the feature representations in the common embedding space as node features of the knowledge graph and determine correlation relationships between nodes based on topology similarity;

[0091] a graph optimization submodule, configured to introduce TCM domain knowledge to optimize graph structure;

[0092] an individualized syndrome reasoning module, configured to realize individualized syndrome reasoning based on the tongue coating syndrome type correlation knowledge graph; wherein the individualized syndrome reasoning module comprises:

[0093] a feature mapping submodule, configured to map tongue image of a new patient to the tongue coating syndrome type correlation knowledge graph;

[0094] a path reasoning submodule, configured to identify similar syndromes based on topology path reasoning;

[0095] a result generation submodule, configured to generate an individualized syndrome result in combination with individual conditions of a patient.

[0096] The beneficial effects of the present application include:

[0097] 1. By introducing topological manifold learning theory, the internal topological structure characteristics of tongue image and syndrome description can be effectively captured, and the robustness and expression ability of feature representation are improved.

[0098] 2. The topological preserving manifold alignment technology is adopted to realize accurate alignment and fusion of tongue and syndrome, which solves the problem of heterogeneous data fusion.

[0099] 3. The knowledge graph constructed based on topological features has better structure characteristics and semantic association, and can support complex reasoning and query.

[0100] 4. The topological path reasoning mechanism makes the personalized syndrome differentiation have better interpretability and accuracy, and improves the effect of syndrome differentiation and treatment.

[0101] 5. The whole system has good scalability, can continuously absorb new tongue-syndrome samples, and continuously optimize the knowledge graph structure. BRIEF DESCRIPTION OF DRAWINGS

[0102] Figure 1 the flow chart of the tongue-syndrome related knowledge graph construction method of the present application;

[0103] Figure 2 the structure schematic diagram of the tongue image topological feature extraction module of the present application;

[0104] Figure 3 the structure schematic diagram of the syndrome semantic topological feature extraction module of the present application;

[0105] Figure 4 the personalized syndrome differentiation reasoning flow chart of the present application;

[0106] Figure 5 the framework diagram of the tongue-syndrome related knowledge graph construction system of the present application. DETAILED DESCRIPTION

[0107] Please refer to Figure 1 - Figure 5 , the present application will be described in detail below with reference to the accompanying drawings.

[0108] As Figure 1 shown, the present application provides a tongue-syndrome related knowledge graph construction method, mainly including the following steps: obtaining tongue image data and traditional Chinese medicine syndrome description data; extracting topological features of tongue image data and semantic topological features of traditional Chinese medicine syndrome description data; constructing a topological manifold alignment model, mapping the tongue topological features and semantic topological features to a common embedding space; based on the feature representation in the common embedding space, constructing a tongue-syndrome related knowledge graph; based on the tongue-syndrome related knowledge graph, realizing personalized syndrome differentiation reasoning.

[0109] In one embodiment of the present application, tongue image data and TCM syndrome description data are first acquired. Tongue image data is mainly collected through standardized tongue diagnosis equipment, including peeled fur tongue fur images and normal tongue fur images. TCM syndrome description data is mainly extracted from TCM electronic medical record system, including text description of patient's syndrome.

[0110] Preferably, before acquiring data, a data preprocessing step is performed, including: standardizing tongue image data, adjusting image size to a uniform specification of 512x512 pixels; extracting tongue body area through tongue body positioning algorithm; quality assessment of tongue image data, filtering low-quality images; text cleaning of TCM syndrome description data, removing noise and redundant information; standardizing TCM syndrome description data according to TCM terminology specifications; data augmentation of tongue image data and TCM syndrome description data, expanding training samples.

[0111] Data augmentation techniques include rotating tongue image data (±10°), translating (±5%), adjusting brightness (±10%), adjusting contrast (±10%), and other operations, and processing syndrome description text with synonym replacement, sentence transformation, etc. These preprocessing steps can improve the effectiveness of subsequent feature extraction and model training.

[0112] As shown in Figure 2 The present application extracts topological features of tongue image data, specifically including: using an efficient convolutional neural network to extract basic visual features of tongue image data; constructing a topological perception layer to calculate topological invariants of tongue image data at different scales; designing a multi-scale filter to capture topological features of tongue image data at multiple scale parameters; integrating topological features at each level through an attention mechanism to generate a unified representation of tongue topological features.

[0113] In one embodiment of the present application, EfficientNet-B3 network is first used to extract basic visual features of tongue image. EfficientNet-B3 is an efficient convolutional neural network with about 12 million parameters, which can effectively extract visual features of images. For an input tongue image of 512x512 pixels, after processing by EfficientNet-B3, a 512-dimensional visual feature vector can be obtained.

[0114] Next, a topological perception layer is constructed to calculate the topological invariants of tongue image at different scales. Topological invariants are important features that describe the topological structure of data, including connected components (0-dimensional homology group), voids (1-dimensional homology group), and high-dimensional cavities (high-dimensional homology group), etc.

[0115] In the topology-aware layer, we first treat the tongue image as a high-dimensional point cloud, and then analyze its topology by constructing a simplicial complex. Specifically, for a pixel value matrix , we first convert it to a grayscale image , and then construct binary images under different thresholds :

[0116] ,

[0117] where is the binary image under threshold , and denotes the pixel value of the binary image at position ; is the grayscale image, and denotes the pixel value of the grayscale image at position ; is the threshold parameter controlling the binarization process; and are the row and column indices of the pixel, respectively, taking values in the ranges .

[0118] Based on the binary images , we can construct a simplicial complex and compute its Betti numbers:

[0119] ,

[0120] ,

[0121] ,

[0122] where denotes the 0-dimensional Betti number under threshold , i.e., the number of connected components; denotes the 1-dimensional Betti number under threshold , i.e., the number of holes; denotes the 2-dimensional Betti number under threshold , i.e., the number of 2-dimensional cavities. Betti numbers are important invariants that describe the topological structure of a space. Different dimensions of Betti numbers correspond to different types of topological features.

[0123] Then, we design a multi-scale filter to capture the topological features of tongue images at multiple scale parameters. Specifically, we choose a series of thresholds For example, 20 thresholds are selected in the range of [0.1, 2.0] with a step of 0.1, and the Betti numbers under each threshold are calculated to obtain the scale-topological feature matrix:

[0124] ,

[0125] wherein: is the scale-topological feature matrix, with a size of , wherein is the number of thresholds (e.g., 20); each row in the matrix represents a topological feature vector under a scale parameter ; each column represents the 0-dimensional, 1-dimensional, and 2-dimensional Betti numbers under different thresholds, respectively.

[0126] Finally, the basic visual features and multi-scale topological features are integrated through an attention mechanism. Specifically, we first map the scale-topological feature matrix to a topological feature vector through a fully connected layer, and then calculate the attention weights of the visual feature vector and the topological feature vector using a self-attention mechanism:

[0127] ,

[0128] wherein: is the attention weight, which is a scalar value representing the importance of the topological features; and are the attention weight matrices, used to map the feature vectors to the query space and the key space, wherein is the dimension of the attention mechanism, usually set to 128; is the visual feature vector, with a dimension of 512; is the topological feature vector, with a dimension of 256; represents the dot product operation; represents function, used to normalize the dot product result to a probability distribution.

[0129] Based on the attention weight , we weight and integrate the visual features and topological features:

[0130] ,

[0131] wherein: is the integrated image topological feature, with a dimension of 512; is the visual feature vector; is the attention weight; This is a value mapping matrix used to map topological feature vectors to the same dimensional space as visual features; These are topological feature vectors; + indicates matrix multiplication; + indicates vector addition.

[0132] In practical applications, we usually further... To enhance its expressive power, a fully connected layer is used to map the feature vector to a higher dimension, such as 768 dimensions.

[0133] like Figure 3 As shown, this invention extracts semantic topological features from TCM syndrome description data, specifically including: converting TCM syndrome description data into semantic vectors using a bidirectional long short-term memory network; constructing a neighborhood relationship graph of TCM syndrome description data based on semantic similarity; constructing a semantic topological space of TCM syndrome description data through the neighborhood relationship graph; enriching the structural information of the semantic topological space by combining TCM syndrome association knowledge; and integrating semantic topological features at different levels into a unified semantic topological representation through topological persistence analysis.

[0134] In one embodiment of the present invention, a Bidirectional Long Short-Term Memory (BiLSTM) network is first used to transform the TCM syndrome description text into semantic vectors. The BiLSTM network is configured with a 3-layer structure, a hidden layer dimension of 256, and a dropout rate of 0.3. The input syndrome description text, after word segmentation and word embedding processing, is fed into the BiLSTM network to obtain the semantic vectors. .

[0135] Next, a neighborhood relationship graph for syndrome descriptions is constructed based on semantic similarity. Specifically, for the set of semantic vectors... Calculate the cosine similarity between any two vectors:

[0136] .

[0137] in: Semantic vectors and The cosine similarity between them ranges from [-1, 1]. and They represent the first The and the first Each syndrome description has a semantic vector with a dimension of 512; Represents the dot product operation of vectors; and Representing vectors respectively and Euclidean norm ( Norm); and are the indices of syndrome descriptions, whose value ranges are and , respectively. is the total number of syndrome descriptions.

[0138] Based on the similarity matrix, we construct a K-Nearest Neighbor graph , where is the node set, representing all the syndrome descriptions, is the edge set, representing the relationship between syndrome descriptions. If is one of the K-Nearest Neighbors of , there exists an edge . In practical applications, the value of K is usually set to 15.

[0139] Then, we construct the semantic topological space of syndrome descriptions through the neighborhood relationship graph. Specifically, based on the neighborhood relationship graph , we can construct the Vietoris-Rips complex and calculate its Persistent Homology. Persistent Homology is a method to describe the persistence of topological features at different scales, which can capture the multi-scale topological structure of data.

[0140] For different distance thresholds , we construct the Vietoris-Rips complex and calculate its Betti numbers , , etc. By recording the changes of Betti numbers with the threshold , we can get the Persistence Barcode or Persistence Diagram.

[0141] The Persistence Barcode can be represented as a series of intervals , where represents the birth time of the topological feature, represents the death time. The duration represents the importance or significance of the topological feature.

[0142] Combining the knowledge of TCM syndromes, we can further enrich the structural information of the semantic topological space. Specifically, we introduce TCM expert knowledge to establish additional connections between semantically related syndrome descriptions in the topological space, thereby enhancing the expressive ability of the semantic topological space.

[0143] Finally, through topological persistence analysis, we integrate the semantic topological features of the symptom level, syndrome level, and syndrome type level into a unified semantic topological representation. Specifically, we calculate the Persistent Homology for the features of different levels, and then integrate these features through weighted fusion:

[0144] ,

[0145] wherein: is the fused syndrome semantic topological feature, with a dimension of 512; is the feature vector of the layer, respectively representing the symptom layer, the syndrome layer and the syndrome type layer; is the corresponding weight, satisfying , for controlling the importance of different levels of features; represents the multiplication operation of a scalar and a vector; represents the vector addition operation, and the three weighted feature vectors are summed. In actual application, we can determine the optimal weight through learning, or set it according to experience, such as setting the symptom layer weight to 0.3, the syndrome layer weight to 0.4, and the syndrome type layer weight to 0.3.

[0146] Similarly, we usually map to a higher-dimensional feature vector, such as 768 dimensions, to keep the same dimension as the tongue topological feature.

[0147] 4. Topological manifold alignment model construction

[0148] The present application constructs a topological manifold alignment model, specifically including: regarding the tongue topological feature as a data point on a first high-dimensional manifold, and regarding the semantic topological feature as a data point on a second high-dimensional manifold; performing low-dimensional parameterization of the two high-dimensional manifolds through a locally linear embedding technique; designing an adaptive distance metric to calculate the geodesic distance between data points on the two high-dimensional manifolds; determining corresponding anchor points on the two high-dimensional manifolds based on expert-labeled tongue-syndrome pairs; designing a local structure preserving constraint to ensure that the local neighborhood relationship remains unchanged during the alignment process; and minimizing the overall structural difference between the two high-dimensional manifolds through a global optimization algorithm.

[0149] In an embodiment of the present application, the tongue topological feature is first regarded as a data point on a first high-dimensional manifold , and the syndrome semantic topological feature is regarded as a data point on a second high-dimensional manifold .

[0150] Then, the two high-dimensional manifolds are parameterized in low dimension through a locally linear embedding (LLE) technique. LLE is a nonlinear dimensionality reduction method that can preserve the local structure of data. Specifically, for data points on a manifold ,We first find the K-nearest neighbors of each point, and then compute the reconstruction weights So that each point can be reconstructed by a linear combination of its neighbors:

[0151] ,

[0152] Where: is the reconstruction weight matrix; denotes the point ; is the reconstruction weight of point ; denotes the topological feature vector of the i-th tongue image, with dimension 768; denotes the set of neighbors of point , containing the indices of the K closest points to point ; denotes the summation over all neighbors of point ; denotes the squared Euclidean distance; denotes the optimization of the weight matrix to minimize the objective function. The reconstruction weights need to satisfy the constraint that the sum of the weights of all neighbors of each point is 1.

[0153] Based on the reconstruction weights , we can solve the low-dimensional embedding such that the points in the low-dimensional space can also be reconstructed by the same weights:

[0154] ,

[0155] Where: is the low-dimensional embedding matrix, denotes the coordinates of the i-th point in the low-dimensional space, typically with dimension 128; is the reconstruction weight computed from the high-dimensional space; other symbols are the same as in the previous formula. This optimization problem can be solved by solving an eigenvalue problem, resulting in the low-dimensional embedding .

[0156] Through LLE, we can obtain the low-dimensional parameterized representations and of the manifold and , typically with dimension 128 covering 99% of the variance.

[0157] Next, we design an adaptive distance metric to compute the geodesic distance between data points on the manifold. On a manifold, the shortest path between two points is not a straight line, but a geodesic. We can approximate the geodesic distance by constructing a nearest-neighbor graph and then using Dijkstra's algorithm to compute the shortest path on the graph:

[0158] ,

[0159] where: denotes the geodesic distance between points and ; denotes the selection of the path with the minimum distance among all paths from point to point ; denotes the summation over all adjacent pairs of points in the path; denotes the Euclidean distance between points and , calculated by ; , , and are feature vectors with the same dimension.

[0160] Based on the expert-labeled tongue sound-syndrome pairs , we can determine the corresponding anchor points on the two manifolds. These anchor points will guide the alignment process of the two manifolds.

[0161] We design a local structure preserving constraint to ensure that the local neighborhood relationships remain unchanged during the alignment process. Specifically, we define the local structure preserving loss as:

[0162] ,

[0163] where: is the local structure preserving loss; denotes the summation over all sample points; denotes the summation over all neighbors of point i; and denote the distance metrics on manifolds and , which can be Euclidean distance or geodesic distance; and denote the topological feature vectors of the i-th and j-th tongue images, respectively; and denote the semantic topological feature vectors of the i-th and j-th syndrome descriptions, respectively. This represents the square of the Euclidean distance. This loss function encourages maintaining consistent distance relationships between corresponding pairs of points on two manifolds.

[0164] Finally, a global optimization algorithm is used to minimize the overall structural difference between the two manifolds. Specifically, we define the global alignment loss:

[0165] ,

[0166] in: This is the global alignment loss; This indicates that the summation is performed on all labeled pairs; Represents the topological feature vector of the i-th tongue image; This represents the semantic topological feature vector of the syndrome description corresponding to the i-th tongue image; This represents the square of the Euclidean distance. This loss function encourages the corresponding tongue-symptom pairs to have the smallest possible distance in the feature space.

[0167] Combining the local structure preservation loss and the global alignment loss, we can obtain the total alignment loss:

[0168] ,

[0169] in: This represents the total alignment loss; This is the global alignment loss; To preserve the loss in local structures; These are balancing parameters used to control the relative importance of preserving local structures and global alignment. In practice, It is usually set to 0.5.

[0170] By minimizing alignment loss We can obtain the optimal alignment of the two manifolds. The optimization process uses gradient descent, with an initial learning rate of 0.001, a cosine annealing scheduling strategy, a maximum of 500 iterations, and an early stopping condition of no performance improvement on the validation set for 5 consecutive iterations.

[0171] In one embodiment of the present invention, feature representations in a common embedding space are optimized through topology-aware contrastive learning. Specifically, this includes: calculating homology groups for each dimension of the tongue-image topological feature space and the semantic topological feature space; designing a mapping function that preserves the homology group structure to ensure that the topological structure remains unchanged during the mapping process; selecting tongue-image-symptom pairs with similar topological structures as positive samples and selecting samples with similar topological structures but different semantics as difficult negative samples; designing a contrastive loss function that incorporates topological consistency constraints; dynamically adjusting the gradient scale based on topological complexity to avoid overfitting simple samples; and constructing a bidirectional topology-preserving mapping network to ensure the cyclic consistency of the mapping.

[0172] First, we compute the homology groups of each dimension of the tongue feature space and the syndrome feature space. Homology groups are algebraic tools to describe the topological structure of a space, and homology groups of different dimensions correspond to connected components, loops, cavities, and other topological features. We use persistent homology to compute the Betti number sequences of the two feature spaces as an approximation of the homology groups.

[0173] Next, we design a mapping function that preserves the homology group structure, ensuring that the topological structure remains unchanged during the mapping process. We construct a mapping function from the tongue feature space X to the syndrome feature space Y, and require that the mapping remains unchanged in the homological sense, i.e., for any k, there exists where denotes the k-dimensional homology group.

[0174] In practice, we implement the mapping function f through a neural network and constrain it to preserve the topological structure through a regularization term:

[0175] ,

[0176] where: is the topological preservation loss; denotes the sum over the homology groups of dimensions 0 to d; denotes the k-dimensional Betti number of the feature space X; denotes the k-dimensional Betti number of the space obtained by applying the inverse mapping of the mapping f to the feature space Y; denotes the square of the Euclidean distance; d is the highest dimension considered, usually set to 2. This loss function encourages the topological structure to remain consistent before and after the mapping.

[0177] We select tongue-syndrome pairs with similar topological structures as positive samples and samples with similar topological structures but different semantics as difficult negative samples. Specifically, we define the topological similarity:

[0178] ,

[0179] where: denotes the topological similarity between the tongue feature and the syndrome feature , with a value range of (0, 1]; denotes the exponential function; denotes the sum over the Betti numbers of dimensions 0 to d; denotes the k-dimensional Betti number of the tongue feature ; denotes the k-dimensional Betti number of the syndrome feature ; denotes the square of the Euclidean distance; d is the highest dimension considered, usually set to 2.

[0180] Based on the topological similarity, we can select the tongue image- syndrome pairs with high similarity as positive samples, and select the samples with similar topological structure but not corresponding relationship as difficult negative samples.

[0181] A contrast loss function is designed to fuse the topological consistency constraint. Traditional contrast loss functions mainly focus on the relative distance between samples, ignoring the geometric structure of the feature space. We introduce topological consistency constraints into the contrast loss:

[0182] ,

[0183] where: is the contrast loss; denotes the natural logarithm function; denotes the exponential function; denotes the image feature and its corresponding positive sample feature between them, usually using cosine similarity; is the temperature parameter, which controls the smoothness of the distribution; denotes the summation over all N samples; denotes the jth feature; is the topological regularization coefficient, which controls the weight of the topological preservation loss; is the topological preservation loss. The first term of this loss function is the standard InfoNCE loss, which encourages the similarity of positive sample pairs to be higher than that of negative sample pairs; the second term is the topological preservation loss, which ensures the consistency of the topological structure before and after mapping.

[0184] Based on the topological complexity, the gradient scale is dynamically adjusted to avoid overfitting simple samples. Specifically, we adjust the temperature parameter :

[0185] ,

[0186] where: is the dynamically adjusted temperature parameter; is the base temperature parameter, usually set to 0.1; is the adjustment coefficient, which controls the influence of complexity on temperature; is the topological complexity measure, which can be defined based on the Bayesian variance or entropy: where is the average Bayesian. Higher topological complexity will result in a larger temperature parameter, thereby reducing the gradient scale and avoiding overfitting.

[0187] We construct a bidirectional topologically preserving mapping network to ensure the cyclic consistency of the mappings. Specifically, we learn the mapping function from tongue features to syndrome features and the mapping function from syndrome features to tongue features at the same time, and require them to satisfy the cyclic consistency:

[0188] ,

[0189] where: is the cyclic consistency loss; denotes the feature obtained by mapping the tongue feature through the mapping and then through the mapping ; is the original tongue feature; denotes the feature obtained by mapping the syndrome feature through the mapping and then through the mapping ; is the original syndrome feature;

[0190] denotes the square of the Euclidean distance. This loss function encourages the bidirectional mappings to be able to reverse each other, ensuring that no information is lost. By minimizing the cyclic consistency loss

[0191] we can ensure the consistency of the bidirectional mappings, thus improving the quality of feature alignment.

[0192] ,

[0193] where: is the total loss; is the contrastive loss, which already contains the topological preserving loss; is the weight of the cyclic consistency loss, usually set to 0.5; is the cyclic consistency loss.

[0194] By minimizing the total loss we can optimize the feature representation in the common embedding space, making it have good semantic alignment and topological preservation at the same time.

[0195] The application is based on feature representation in a public embedding space, and builds a tongue coating syndrome type association knowledge graph, which specifically includes: defining tongue image features, syndrome description, syndrome type, and other entity types as nodes of the knowledge graph; embedding topological features in the public embedding space as the initial feature representation of the nodes; defining a corresponding attribute set for each type of node, initializing the node attribute values based on domain knowledge and data analysis; defining relationship types such as performance, belonging, and association as edges of the knowledge graph; constructing the association strength between entities based on topological similarity to determine the weight of the edges; verifying the rationality of the automatically constructed relationship through expert knowledge; based on topological redundancy analysis, pruning redundant relationships in the graph; strengthening the key semantic path in the graph to improve reasoning efficiency.

[0196] In an embodiment of the application, the node types of the knowledge graph are first defined, including tongue image feature nodes, syndrome description nodes, and syndrome type nodes. The tongue image feature nodes represent the features of the tongue image, such as tongue shape, tongue coating distribution, and tongue texture, etc.; the syndrome description nodes represent the description of traditional Chinese medicine syndromes, such as stomach heat, liver depression, etc.; and the syndrome type nodes represent traditional Chinese medicine syndromes, such as liver depression and spleen deficiency syndrome, stomach heat syndrome, etc.

[0197] The topological features in the public embedding space are embedded as the initial feature representation of the nodes. Specifically, for the tongue image feature nodes, we use tongue image topological features as their feature representation; for the syndrome description nodes, we use syndrome semantic topological features as their feature representation; and for the syndrome type nodes, we can construct their feature representation by aggregating the features of related syndrome description nodes.

[0198] A corresponding attribute set is defined for each type of node, and the node attribute values are initialized based on domain knowledge and data analysis. For example, the attributes of the tongue image feature nodes can include tongue color, tongue coating thickness, tongue shape state, etc.; the attributes of the syndrome description nodes can include symptom severity, occurrence site, etc.; and the attributes of the syndrome type nodes can include syndrome definition, common performance, etc.

[0199] The edge types of the knowledge graph are defined, including performance, belonging, and association relationships. The performance relationship represents the specific performance of a certain syndrome type, such as stomach heat syndrome showing red tongue and yellow fur; the belonging relationship represents that a certain syndrome belongs to a certain syndrome type, such as red tongue and yellow fur belonging to stomach heat syndrome; and the association relationship represents the correlation between different entities, such as stomach heat syndrome being associated with liver depression syndrome.

[0200] The association strength between entities is constructed based on topological similarity to determine the weight of the edges. Specifically, we calculate the topological similarity between different entity features:

[0201] ,

[0202] wherein: denotes the topological similarity between features and , with a value range of (0, 1]; denotes the exponential function; denotes the summation of the Betti numbers from 0 to d dimensions; denotes the k-dimensional Betti number of feature ; denotes the k-dimensional Betti number of feature ; denotes the square of the Euclidean distance; is the highest dimension considered, usually set to 2.

[0203] If the similarity is greater than a threshold (such as 0.75), a strong association is established; if the similarity is between and (such as 0.5 to 0.75), a weak association is established. The weight of the edge can directly use the similarity value, or can be adjusted through a function mapping.

[0204] Verify the rationality of the automatically constructed relationships through expert knowledge. Since machine learning methods may produce some relationships that do not conform to TCM theory, we need to invite TCM experts to verify the automatically constructed relationships, retain reasonable relationships, and modify or delete unreasonable relationships.

[0205] Based on topological redundancy analysis, prune redundant relationships in the graph. In a knowledge graph, if there are multiple paths between two nodes and these paths are semantically equivalent, we can retain the most important one and prune the other redundant paths. Specifically, we can calculate the transitive closure of the path, and if the similarity in the transitive closure is less than a threshold (such as 0.3), the relationship is considered redundant and can be pruned.

[0206] Strengthen the key semantic paths in the graph to improve reasoning efficiency. In a knowledge graph, some paths are particularly important for reasoning tasks, and we can strengthen these paths by increasing their weights or adding additional direct connections. This can improve the efficiency and accuracy of reasoning.

[0207] Through the above steps, we can construct a tongue coating syndrome association knowledge graph that is structurally reasonable and semantically rich, providing a foundation for subsequent personalized syndrome reasoning.

[0208] For example, Figure 4As shown, the present application is based on tongue coating syndrome type correlation knowledge graph, realizes individualized syndrome differentiation reasoning, specifically including: extracting the topological feature representation of the tongue image of the new patient; mapping the topological feature representation of the tongue image of the new patient to a public embedding space; searching for similar cases in the tongue coating syndrome type correlation knowledge graph based on topological similarity; generating a possible reasoning path set based on the knowledge graph structure; designing a path scoring function that fuses topological structure information; selecting the reasoning path with the highest score for syndrome type identification; comprehensively analyzing the syndrome type distribution of similar cases to generate an individualized syndrome differentiation result; combining the individual situation of the patient to generate an individualized treatment plan suggestion.

[0209] In an embodiment of the present application, the topological feature representation of the tongue image of the new patient is first extracted. Specifically, we use the same feature extraction network as in the training stage to extract the topological features from the tongue image of the new patient .

[0210] The topological feature representation of the tongue image of the new patient is mapped to a public embedding space. Specifically, we use the trained mapping function to map the tongue feature of the new patient to the public embedding space:

[0211] ,

[0212] where: is the mapped feature, with a dimension consistent with the public embedding space; is the mapping function from the tongue feature space to the public embedding space; is the topological feature of the tongue image of the new patient.

[0213] Based on topological similarity, similar cases in the tongue coating syndrome type correlation knowledge graph are searched. Specifically, we calculate the topological similarity between the new patient's feature and all tongue feature nodes in the knowledge graph:

[0214] ,

[0215] where: denotes the mapped feature of the new patient , denotes the topological similarity between the th tongue feature node in the knowledge graph and the new patient's feature; denotes the exponential function; denotes the summation of the 0 to dimensional Bernoulli numbers; denotes the dimensional Bernoulli number of feature ; denotes the dimensional Bernoulli number of feature ; denotes the square of the Euclidean distance; The highest dimension considered is usually set to 2.

[0216] Then select the Top-K cases with the highest similarity as the candidate set. In practice, K is usually set to 20, i.e., retrieve the 20 most similar historical cases.

[0217] Generate a set of possible reasoning paths based on the knowledge graph structure. Starting from the tongue feature node of the new patient, we can generate paths to various syndrome nodes through graph traversal algorithms such as depth-first search or breadth-first search. To control the reasoning complexity, we usually limit the path length to no more than 3 hops.

[0218] Design a path scoring function that integrates topological structure information. For each reasoning path , we define its scoring function as follows:

[0219] ,

[0220] Where: represents the score of path , with a value range of ; represents the product of the scores of all adjacent node pairs in the path; represents the weight of the edge between nodes and , with a value range of ; represents the topological similarity between nodes and , with a value range of ; represents the th node in the path, ; represents the path length.

[0221] Select the reasoning path with the highest score for syndrome identification. We calculate the scores of all possible paths and select the path with the highest score as the optimal reasoning path. The endpoint syndrome node of this path is the identification result.

[0222] Integrate the syndrome distribution of similar cases to generate personalized syndrome results. Specifically, we count the frequency of each syndrome in the Top-K similar cases to generate the syndrome distribution:

[0223] ,

[0224] Where: represents the probability of syndrome , with a value range of ; The number of occurrences of each syndrome type in the Top-K cases The number of occurrences of each syndrome type in the Top-K cases The total number of similar cases, for example, 20.

[0225] We only output the syndrome types with probabilities greater than the threshold (like 0.7) as the diagnosis results, while returning the top 3 most likely syndrome types and their probability distributions for the doctor's reference.

[0226] Based on the patient's individual situation, we generate personalized treatment plan recommendations. Based on the diagnosis results and the patient's individual characteristics (such as age, gender, medical history, etc.), we can retrieve the corresponding treatment plans from the knowledge graph and make personalized adjustments to generate the final treatment recommendations.

[0227] Through the above steps, we realize the personalized syndrome reasoning based on the tongue coating syndrome type association knowledge graph, providing intelligent assistance for TCM clinical diagnosis.

[0228] In one embodiment of the present invention, in order to maintain the timeliness and accuracy of the knowledge graph, we also design a dynamic updating mechanism, which includes: collecting new tongue image data and syndrome description data; extracting the topological features of the new data; mapping the topological features of the new data to a public embedding space; verifying and adjusting the association of the new data in the knowledge graph based on expert feedback; balancing the weights of new and old knowledge based on a time decay mechanism; periodically re-optimizing the overall structure of the tongue coating syndrome type association knowledge graph.

[0229] Collect new tongue image data and syndrome description data. As clinical practice progresses, we can continuously collect new tongue image and corresponding syndrome description, enriching the database.

[0230] Extract the topological features of the new data. Use the trained feature extraction network to extract the topological features from the new data.

[0231] Map the topological features of the new data to a public embedding space. Use the trained mapping function to map the features of the new data to a public embedding space.

[0232] Verify and adjust the association of the new data in the knowledge graph based on expert feedback. Add the new data to the knowledge graph and invite experts to verify the rationality of its association, and adjust if necessary.

[0233] Balance the weights of new and old knowledge based on a time decay mechanism. In order to avoid the old knowledge being completely replaced by the new knowledge, we introduce a time decay mechanism to adjust the weight of the knowledge:

[0234] ,

[0235] wherein represents the weight of the knowledge of time, taking a value range of represents an exponential decay function; represents the current time; represents the time when the knowledge is added to the graph; is a decay coefficient, controlling the speed of decay, usually set to 0.1 / month;- represents subtraction operation;· represents multiplication operation. With the passage of time, the weight of old knowledge will gradually decrease, but will not be 0, ensuring that important historical knowledge can still play a role. Periodically re-optimize the overall structure of the tongue coating syndrome-related knowledge graph. As new knowledge is continuously added, the structure of the knowledge graph may become less optimized. Therefore, we need to periodically (such as monthly) re-optimize the overall structure of the knowledge graph, including relationship pruning, path strengthening, etc.

[0236] Through the dynamic updating mechanism, we can make the tongue coating syndrome-related knowledge graph evolve continuously, adapt to new clinical findings and theoretical progress.

[0237] As shown in

[0238] , the present application also provides a tongue-coating syndrome-related knowledge graph construction system, comprising a data acquisition and preprocessing module 1, a topological feature extraction module 2, a multi-modal feature alignment module 3, a knowledge graph construction module 4, and a personalized syndrome reasoning module 5. Figure 5 The data acquisition and preprocessing module 1 is used to acquire tongue image data and traditional Chinese medicine syndrome description data, and to standardize the data. This module includes a tongue image acquisition sub-module, a tongue image preprocessing sub-module, a syndrome description acquisition sub-module, and a data quality control sub-module.

[0239] The topological feature extraction module 2 is used to extract the topological features of the tongue image data and the semantic topological features of the traditional Chinese medicine syndrome description data. This module includes a tongue feature extraction sub-module and a syndrome feature extraction sub-module. The tongue feature extraction sub-module uses a feature extraction network to extract multi-level tongue topological features from the tongue image data. The syndrome feature extraction sub-module uses a semantic analysis network to extract multi-level semantic topological features from the traditional Chinese medicine syndrome description data.

[0240] The multi-modal feature alignment module 3 is used to construct a topological manifold alignment model, which maps the tongue topological features and the semantic topological features to a common embedding space. This module includes a manifold construction sub-module, a manifold alignment sub-module, and a contrastive learning sub-module. The manifold construction sub-module constructs manifold structures for tongue topological features and semantic topological features respectively. The manifold alignment sub-module realizes the alignment between the two manifolds based on topological preservation constraints. The contrastive learning sub-module optimizes the feature representation in the common embedding space through topologically aware contrastive learning.

[0241] Through the dynamic updating mechanism, we can make the tongue coating syndrome-related knowledge graph evolve continuously, adapt to new clinical findings and theoretical progress.

[0242] The knowledge graph construction module 4 is used to construct a tongue coating syndrome type correlation knowledge graph based on the feature representations in the common embedding space. The module includes a graph construction submodule and a graph optimization submodule. The graph construction submodule takes the feature representations in the common embedding space as the node features of the knowledge graph, and determines the correlation between nodes based on topological similarity. The graph optimization submodule introduces TCM domain knowledge to optimize the graph structure.

[0243] The individualized syndrome reasoning module 5 is used to realize individualized syndrome reasoning based on the tongue coating syndrome type correlation knowledge graph. The module includes a feature mapping submodule, a path reasoning submodule, and a result generation submodule. The feature mapping submodule maps the tongue image of a new patient to the tongue coating syndrome type correlation knowledge graph. The path reasoning submodule identifies similar syndromes based on topological path reasoning. The result generation submodule generates an individualized syndrome result in combination with the individual situation of the patient.

[0244] The system also includes an optional knowledge graph updating module for dynamically updating the knowledge graph to maintain its timeliness and accuracy.

[0245] The above description is only preferred embodiments of the present application and is not intended to limit the present application. The present application can have various modifications and changes for those skilled in the art. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for constructing a knowledge graph linking tongue coating and syndrome types, characterized in that, include: Acquire tongue image data and TCM syndrome description data; Extracting the topological features of the tongue image data and the semantic topological features of the TCM syndrome description data; wherein, the extraction includes: A feature extraction network is used to extract multi-level tongue topological features from the tongue image data. The multi-level tongue topological features include tongue morphology features, tongue coating distribution features, and tongue texture features. A semantic analysis network is used to extract multi-level semantic topological features from the TCM syndrome description data. The multi-level semantic topological features include symptom-level features, syndrome-level features, and syndrome type-level features. A topological manifold alignment model is constructed to map the tongue image topological features and the semantic topological features to a common embedding space; wherein, the construction includes: Construct manifold structures for the tongue image topological features and the semantic topological features, respectively; Alignment between two manifolds is achieved based on topology-preserving constraints; The feature representation in the common embedding space is optimized through topology-aware contrastive learning. Based on the feature representations in the aforementioned public embedding space, a knowledge graph relating tongue coating syndrome types is constructed; wherein, the construction includes: The feature representations in the public embedding space are used as node features of the knowledge graph; Determine the relationships between nodes based on topological similarity; Incorporate knowledge from the field of Traditional Chinese Medicine to optimize the graph structure; Based on the aforementioned knowledge graph of tongue coating syndrome types, personalized diagnostic reasoning is achieved; wherein, the implementation includes: Map the tongue images of new patients to the tongue coating syndrome association knowledge graph; Identify similar evidence types based on topological path reasoning; Personalized diagnostic results are generated based on the individual patient's condition.

2. The method for constructing a knowledge graph of tongue coating-syndrome association according to claim 1, characterized in that, The extraction of topological features from the tongue image data specifically includes: The basic visual features of the tongue image data are extracted using an efficient convolutional neural network. A topology-sensing layer is constructed to calculate the topological invariants of the tongue image data at different scales. The topological invariants include connected components, holes, and high-dimensional cavities. Design a multi-scale filter to capture the topological features of the tongue image data under multiple scale parameters, and generate a scale-topological feature matrix; By integrating topological features from various levels through an attention mechanism, a unified topological feature representation of the tongue image is generated.

3. The method for constructing a knowledge graph of tongue coating-syndrome association according to claim 1, characterized in that, The extraction of semantic topological features from the TCM syndrome description data specifically includes: The TCM syndrome description data is transformed into semantic vectors using a bidirectional long short-term memory network. A neighborhood relationship graph of the TCM syndrome description data is constructed based on semantic similarity; The semantic topology space of the TCM syndrome description data is constructed by the neighborhood relationship graph, and the semantic topology space includes semantic connectivity components and semantic holes. By combining knowledge of syndrome correlation in Traditional Chinese Medicine, the structural information of the semantic topological space is enriched; Through topological persistence analysis, semantic topological features at different levels are integrated into a unified semantic topological representation.

4. The method for constructing a knowledge graph of tongue coating-syndrome association according to claim 1, characterized in that, The construction of the topological manifold alignment model specifically includes: The tongue topological features are considered as data points on a first high-dimensional manifold, and the semantic topological features are considered as data points on a second high-dimensional manifold. Two high-dimensional manifolds are represented in a low-dimensional parameterized manner by using local linear embedding techniques; Design an adaptive distance metric to calculate the geodesic distance between data points on the two high-dimensional manifolds; Based on the tongue image-symptom pair annotated by experts, the corresponding anchor points on the two high-dimensional manifolds are determined. The design maintains local structural constraints to ensure that local neighborhood relationships remain unchanged during alignment. The global optimization algorithm minimizes the overall structural difference between the two high-dimensional manifolds.

5. The method for constructing a knowledge graph of tongue coating-syndrome association according to claim 1, characterized in that, The optimization of feature representations in the common embedding space through topology-aware contrastive learning specifically includes: Calculate the homology groups of each dimension of the tongue image topological feature space and the semantic topological feature space; Design a mapping function that preserves the homology group structure to ensure that the topology remains unchanged during the mapping process; Tongue-symptom pairs with similar topological structures were selected as positive samples, and samples with similar topological structures but different semantics were selected as difficult negative samples. Design a contrastive loss function that incorporates topological consistency constraints; Dynamically adjust the gradient scale based on topological complexity to avoid overfitting simple samples; Construct a bidirectional topology-preserving mapping network to ensure the cyclic consistency of the mapping.

6. The method for constructing a knowledge graph of tongue coating-syndrome association according to claim 1, characterized in that, The construction of a tongue coating syndrome association knowledge graph based on feature representations in the public embedding space specifically includes: Define entity types such as tongue appearance features, syndrome descriptions, and syndrome types as nodes in the knowledge graph; The topological features in the common embedding space are embedded as the initial feature representation of the nodes; Define a corresponding set of attributes for each type of node, and initialize the node attribute values ​​based on domain knowledge and data analysis; The edges of a knowledge graph are defined as being represented by the types of relationships: belonging to, related to, or associated with. The strength of the association between entities is determined based on topological similarity, and the weight of the edges is determined accordingly. The rationality of automatically constructed relationships is verified through expert knowledge. Based on topological redundancy analysis, redundant relationships in the graph are removed; Strengthen key semantic paths in the graph to improve reasoning efficiency.

7. The method for constructing a knowledge graph of tongue coating-syndrome association according to claim 1, characterized in that, The personalized diagnostic reasoning based on the tongue coating syndrome association knowledge graph specifically includes: Extracting topological features from new patients' tongue images; Map the topological feature representation of the new patient's tongue image to the common embedding space; Similar cases in the knowledge graph of tongue coating syndrome types were retrieved based on topological similarity. A set of possible reasoning paths is generated based on the knowledge graph structure; Design a path scoring function that incorporates topology information; Select the reasoning path with the highest score for evidence identification; By comprehensively analyzing the distribution of syndrome types in similar cases, personalized diagnostic results are generated. Based on the patient's individual circumstances, personalized treatment plan suggestions are generated.

8. The method for constructing a knowledge graph of tongue coating-syndrome association according to claim 1, characterized in that, It also includes the step of dynamically updating the knowledge graph of tongue coating syndrome types: Collect newly added tongue image data and syndrome description data; Extract the topological features of the newly added data; Map the topological features of the newly added data to the common embedding space; Based on expert feedback, the relationships between the newly added data in the knowledge graph were verified and adjusted. Based on the time decay mechanism, the weights of new and old knowledge are balanced; The overall structure of the knowledge graph relating tongue coating syndromes is periodically re-optimized.

9. The method for constructing a knowledge graph of tongue coating-syndrome association according to claim 1, characterized in that, Before acquiring tongue image data and TCM syndrome description data, a data preprocessing step is also included: The tongue image data is standardized to adjust the image size to a uniform specification; The tongue region is extracted using a tongue localization algorithm; The tongue image data is subjected to quality assessment, and low-quality images are filtered out; Text cleaning was performed on the TCM syndrome description data to remove noise and redundant information; The TCM syndrome description data were standardized according to the TCM terminology standard. Data augmentation is performed on the tongue image data and the TCM syndrome description data to expand the training samples.

10. A tongue coating-syndrome association knowledge graph construction system, employing the tongue coating-syndrome association knowledge graph construction method according to any one of claims 1-9, characterized in that, include: The data acquisition and preprocessing module is used to acquire tongue image data and TCM syndrome description data, and to standardize the data. A topological feature extraction module is used to extract the topological features of the tongue image data and the semantic topological features of the TCM syndrome description data; wherein, the topological feature extraction module includes: The tongue image feature extraction submodule is used to extract multi-level tongue image topological features from the tongue image data using a feature extraction network. The syndrome feature extraction submodule is used to extract multi-level semantic topological features from the TCM syndrome description data using a semantic analysis network. A multimodal feature alignment module is used to construct a topological manifold alignment model, mapping the tongue image topological features and the semantic topological features to a common embedding space; wherein, the multimodal feature alignment module includes: The manifold construction submodule is used to construct manifold structures for the tongue topological features and the semantic topological features, respectively. The manifold alignment submodule is used to achieve alignment between two manifolds based on topology-preserving constraints. The contrastive learning submodule is used to optimize the feature representation in the common embedding space through topology-aware contrastive learning; A knowledge graph construction module is used to construct a tongue coating syndrome association knowledge graph based on feature representations in the public embedding space; wherein, the knowledge graph construction module includes: The graph construction submodule is used to take the feature representations in the public embedding space as node features of the knowledge graph and determine the association relationship between nodes based on topological similarity. The graph optimization submodule is used to incorporate knowledge from the field of Traditional Chinese Medicine to optimize the graph structure. A personalized dialectical reasoning module is used to implement personalized dialectical reasoning based on the tongue coating syndrome association knowledge graph; wherein, the personalized dialectical reasoning module includes: The feature mapping submodule is used to map the tongue image of a new patient to the tongue coating syndrome association knowledge graph. The path reasoning submodule is used to identify similar evidence types based on topological path reasoning; The results generation submodule is used to generate personalized diagnostic results based on the individual patient's situation.