Traditional Chinese medicine syndrome type intelligent identification system based on multi-source clinical data

By deeply integrating topological theory with TCM syndrome differentiation, a multi-source data acquisition and multi-dimensional syndrome relationship representation module was constructed, which solved the problems of high-dimensional complex relationships and uncertainties in TCM syndrome differentiation system, and achieved highly accurate, stable and efficient intelligent identification of TCM syndrome types.

CN121237383APending Publication Date: 2025-12-30YANAN UNIV AFFILIATED HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511431681.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-09
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing TCM diagnostic systems struggle to handle complex, high-dimensional relationships, adapt to uncertainty, lack multi-path parallel reasoning mechanisms, and effectively integrate multi-source clinical data, resulting in unstable diagnostic results and insufficient accuracy.

Method used

By deeply integrating topological theory with TCM syndrome differentiation practice, we construct modules for multi-source data acquisition, multi-dimensional representation of syndrome relationships, dynamic evolution, and syndrome type reasoning optimization. Through multi-dimensional topological space representation of symptom-syndrome-syndrome relationship, combined with homology topology, persistence theory, and homotopy equivalence principle, we achieve multi-path parallel reasoning and result fusion.

Benefits of technology

It improves the accuracy of diagnosis, enhances system stability and reasoning efficiency, enables personalized diagnosis, discovers new syndrome correlation patterns, and improves the accuracy of TCM diagnosis and the system's adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237383A_ABST
    Figure CN121237383A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and medical information processing, in particular to a traditional Chinese medicine syndrome type intelligent identification system based on multi-source clinical data, which comprises a multi-source data acquisition module, a syndrome relationship multi-dimensional representation module, a syndrome relationship dynamic evolution module, a syndrome type reasoning optimization module and a result output module, the system constructs a multi-dimensional topological space of symptoms, syndromes and syndrome types by adopting a homology theory in topology, analyzes the stability of a syndrome relationship through a topology persistence theory, simplifies an inference path based on a homotopy equivalence principle, realizes multi-path parallel inference, and improves the reliability of the system. The system solves the problems that an existing traditional Chinese medicine syndrome differentiation system is difficult to process complex interaction among symptoms, uncertainty of diagnosis knowledge, a single reasoning path, multi-source heterogeneous data integration and the like, and high-accuracy, high-stability and high-efficiency intelligent identification of traditional Chinese medicine syndrome types is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and medical information processing technology, specifically to a TCM syndrome identification system based on multi-source clinical data. This system is used to intelligently analyze multi-source clinical data such as patients' four diagnostic methods data and laboratory test data to achieve accurate identification of TCM syndromes. Background Technology

[0002] Traditional Chinese medicine (TCM) syndrome differentiation is a core component of TCM diagnosis and treatment. Traditional TCM syndrome differentiation relies heavily on the physician's experience and subjective judgment, inherently possessing a degree of subjectivity and uncertainty. While the development of artificial intelligence (AI) technology has made its application in TCM syndrome differentiation possible, it still faces the following challenges:

[0003] First, existing TCM syndrome differentiation systems mostly employ flat rule bases or simple statistical models, making it difficult to handle the complex interactions and nonlinear relationships between symptoms. TCM syndrome identification involves complex combinations of numerous symptoms, and traditional methods struggle to fully capture these high-dimensional associations.

[0004] Secondly, TCM diagnostic knowledge is characterized by a high degree of uncertainty and ambiguity, and existing systems lack effective mechanisms to handle this uncertainty, resulting in insufficient stability of diagnostic results. The system performs poorly when faced with noisy data and incomplete symptom information.

[0005] Third, existing systems mostly use a single reasoning path and lack a multi-path parallel reasoning mechanism, making it difficult to make a comprehensive judgment based on multiple pieces of evidence, resulting in low accuracy in cases of complex evidence types or multiple evidence types coexisting.

[0006] Fourth, the existing system has limited ability to integrate multi-source heterogeneous data, making it difficult to effectively integrate different types of clinical data such as diagnostic data, laboratory test data, etc., which affects the comprehensiveness and accuracy of diagnosis.

[0007] Therefore, there is an urgent need to develop a TCM syndrome identification system that can handle high-dimensional complex relationships, adapt to uncertainty, support multi-path reasoning, and effectively integrate multi-source clinical data. Summary of the Invention

[0008] The purpose of this invention is to provide an intelligent TCM syndrome identification system based on multi-source clinical data, which solves the problems existing in the prior art by introducing a deep integration of topological theory and TCM syndrome differentiation practice.

[0009] This invention proposes an intelligent TCM syndrome identification system based on multi-source clinical data, comprising:

[0010] The multi-source data acquisition module is used to collect patients' four diagnostic methods data, laboratory test data, and basic patient information;

[0011] The syndrome relationship multidimensional representation module is communicatively connected to the multi-source data acquisition module. It is used to receive patient data sent by the multi-source data acquisition module, construct a multidimensional topological space of symptoms, syndromes and syndrome types based on homology topology theory, map the patient data to the multidimensional topological space, and extract topological features.

[0012] The syndrome relationship dynamic evolution module is communicatively connected to the syndrome relationship multidimensional representation module. It is used to receive the topological features extracted by the syndrome relationship multidimensional representation module, analyze the stability of syndrome relationship under different intensity thresholds based on the topological persistence theory, identify stable syndrome relationship patterns, and construct a stable syndrome relationship core structure.

[0013] The syndrome type reasoning optimization module is communicatively connected to the syndrome relationship dynamic evolution module. It is used to receive the stable syndrome relationship core structure constructed by the syndrome relationship dynamic evolution module, simplify the reasoning path based on the homotopy equivalence principle while maintaining the diagnostic essence, and generate a syndrome type set and corresponding confidence set through multi-path parallel reasoning.

[0014] The result output module is communicatively connected to the syndrome type reasoning optimization module. It is used to receive the syndrome type set and the corresponding confidence set generated by the syndrome type reasoning optimization module, and to display the diagnostic results and the syndrome relationship visualization map.

[0015] Preferably, the multidimensional representation module of symptom relationships includes:

[0016] Entity coding unit is used to extract symptoms, syndromes and syndrome types from the four diagnostic methods data and knowledge base, and to assign a unique identifier to each entity;

[0017] The relation network construction unit is communicatively connected to the entity encoding unit and is used to calculate the initial association strength between entities based on co-occurrence frequency, set an association threshold to filter valid relationships, and construct an initial relation graph network.

[0018] The high-dimensional relationship extraction unit is communicatively connected to the relationship network construction unit and is used to identify frequently co-occurring symptom combinations, construct a high-dimensional simplex representing multiple symptom combinations, and calculate the weight and stability index of the simplex.

[0019] The topology feature calculation unit is communicatively connected to the high-dimensional relationship extraction unit and is used to identify connected components in the network, detect key loop structures, and extract feature vectors to represent the overall topology characteristics.

[0020] Preferably, the high-dimensional relationship extraction unit uses a multi-level simplex data structure to store medical entity relationships, the multi-level simplex data structure including:

[0021] A collection of nodes used to store all medical entities;

[0022] An edge set is used to represent direct associations between entities;

[0023] Higher-order simplex sets are used to represent multi-entity composition relationships;

[0024] The relation weight matrix is ​​used to record the strength of the association between entities;

[0025] Monomorphic complex structures are used to integrate relationships across all dimensions.

[0026] Preferably, the symptom relationship dynamic evolution module includes:

[0027] A multi-scale filtering sequence construction unit is used to set an incremental relational threshold sequence, construct the corresponding network state under each threshold, and record the network topology changes under different thresholds.

[0028] The persistent feature tracking unit is communicatively connected to the multi-scale filtering sequence construction unit and is used to detect the emergence of new topological features, track the disappearance of features, and calculate the persistence metric of features.

[0029] The stable structure identification unit is communicatively connected to the persistent feature tracking unit and is used to analyze the distribution of persistent intervals, screen long-lived topological features, and construct the core structure of stable syndrome relationships.

[0030] The dynamic update unit is communicatively connected to the stable structure identification unit and is used to monitor the impact of new clinical data, identify network regions that need to be updated, and locally reconstruct the network structure.

[0031] The noise filtering unit, which is communicatively connected to the persistent feature tracking unit, is used to identify short-lived unstable features, detect abnormal topological changes, and apply adaptive thresholds for noise suppression.

[0032] Preferably, the persistent feature tracking unit uses a persistent analysis data structure to store dynamic evolution information, and the persistent analysis data structure includes:

[0033] Filtered complex sequences are used to store network state sequences at different thresholds;

[0034] A persistent interval set is used to record the lifecycle of topological features;

[0035] Feature importance index, used to quantify the stability and significance of features;

[0036] Evolutionary trajectory diagrams are used to record changes in network structure over time.

[0037] Preferably, the proof-type reasoning optimization module includes:

[0038] The reasoning path planning unit is used to identify possible paths from symptoms to syndromes in the syndrome network, calculate the initial weight and reliability of each path, and construct an initial reasoning space containing all possible paths.

[0039] The path simplification and optimization unit is communicatively connected to the inference path planning unit and is used to identify topologically equivalent path classes, merge redundant paths, retain representative paths, and optimize the path set.

[0040] The multi-path parallel inference unit is communicatively connected to the path simplification and optimization unit. It is used to select the optimal multiple paths for parallel inference, apply the corresponding inference rules to each path, and independently calculate the proof result and confidence level of each path.

[0041] The result fusion and conflict resolution unit is communicatively connected to the multi-path parallel reasoning unit. It is used to perform weighted fusion of results based on path reliability, handle result conflicts between different paths, adjust results by applying the principles of traditional Chinese medicine syndrome differentiation, and generate the final syndrome set and confidence level.

[0042] The feedback and self-optimization unit is communicatively connected to the result fusion and conflict resolution unit, and is used to record the intermediate state of the reasoning process, evaluate the consistency between the reasoning results and clinical feedback, and dynamically adjust the path weights and reasoning strategies.

[0043] Preferably, the inference path planning unit uses an inference optimization data structure to store path information, and the inference optimization data structure includes:

[0044] The reasoning path set is used to store possible proof-type reasoning paths;

[0045] Path equivalence classes are used to classify inference paths that are topologically equivalent.

[0046] The path rating matrix is ​​used to record the reliability rating of different paths;

[0047] Proof-type reasoning trees are used to organize hierarchical reasoning logic.

[0048] Preferably, the multi-source data acquisition module includes:

[0049] The four diagnostic methods data acquisition unit is used to collect data on patients' inspection, auscultation and olfaction, inquiry and palpation.

[0050] The laboratory test data acquisition unit is used to collect patients' biochemical test, imaging and other auxiliary test data;

[0051] The patient basic information collection unit is used to collect information such as the patient's age, gender, physical type, and past medical history;

[0052] The data preprocessing unit is communicatively connected to the four diagnostic data acquisition unit, the laboratory test data acquisition unit, and the patient basic information acquisition unit. It is used to standardize the acquired data, remove outliers, and convert unstructured data into structured data format.

[0053] Preferably, the result output module includes:

[0054] The syndrome diagnosis result generation unit is used to generate the patient's main syndrome diagnosis result based on the syndrome set and the corresponding confidence set;

[0055] The syndrome relationship visualization unit is used to convert syndrome relationships in a multidimensional topological space into two-dimensional or three-dimensional visualization maps.

[0056] The diagnostic basis explanation unit is used to extract key symptoms and syndrome evidence that support the syndrome type diagnosis and generate a diagnostic basis explanation.

[0057] The treatment suggestion generation unit is used to provide corresponding TCM treatment suggestions based on the syndrome differentiation diagnosis results.

[0058] As a preferred option, it also includes:

[0059] The medical knowledge base management module is communicatively connected to the syndrome relationship multidimensional representation module, the syndrome relationship dynamic evolution module, and the syndrome type reasoning optimization module, and is used to store TCM theoretical knowledge, symptom-syndrome-syndrome relationship rules, and clinical case data;

[0060] The knowledge discovery module is communicatively connected to the syndrome relationship dynamic evolution module. It is used to discover new syndrome association patterns from clinical data, quantify the contribution of different symptoms to the syndrome type, and feed the discovered new knowledge back to the medical knowledge base management module.

[0061] The system learning and optimization module is connected in communication with the syndrome type reasoning optimization module. It is used to continuously optimize the reasoning strategy and parameter configuration based on clinical feedback to improve the system's diagnostic accuracy.

[0062] This invention employs homology theory, persistence theory, and the homotopy equivalence principle from topology to construct an innovative framework for representing and reasoning about symptom relationships, achieving intelligent dialectical processing throughout the entire process from multidimensional representation and dynamic evolution to optimized reasoning. Compared with existing technologies, this invention has the following advantages:

[0063] 1. Improved diagnostic accuracy: By representing the complex relationships between symptoms, syndromes, and syndrome types using a multidimensional topological space, the system can capture high-order patterns that traditional planar methods cannot identify, significantly improving the accuracy of identifying complex syndrome types. Experimental verification shows that the ability to identify complex syndrome types is enhanced by approximately 20%, and the accuracy of diagnosis in cases of multiple coexisting diseases is improved by approximately 15%.

[0064] 2. Enhanced system stability: The dynamic evolution mechanism of syndrome relationships based on topological persistence theory enables the system to identify stable syndrome relationship patterns from noisy data, improves robustness to incomplete symptom data by more than 50%, and enhances the adaptability of newly emerging symptom combinations by about 35%.

[0065] 3. Improved reasoning efficiency: Based on the homotopy equivalence-based syndrome reasoning optimization mechanism, the reasoning path is simplified while maintaining the diagnostic essence, and the reliability of results is improved through multi-path parallel reasoning. The reasoning speed is increased by approximately 3 times, and system resource consumption is reduced by approximately 40%, supporting real-time clinical decision-making.

[0066] 4. Knowledge discovery has been achieved: The system can automatically discover new syndrome correlation patterns, quantify the contribution of different symptoms to syndrome types, and promote the improvement and development of the TCM theoretical system. It has been verified that the system has automatically discovered more than 40 new symptom-syndrome correlation relationships.

[0067] 5. Enhanced personalized diagnosis: The system incorporates individualized factors such as the patient's constitution and age, achieving more accurate diagnosis. Personalized diagnostic capabilities have improved by approximately 30%, better adapting to individual patient differences.

[0068] In summary, this invention deeply integrates cutting-edge topological theories with TCM diagnostic knowledge, creatively proposing a syndrome relationship representation and reasoning framework based on topological persistence. This provides a new technical approach for the intelligentization of TCM syndrome differentiation and has significant clinical application value and theoretical innovation significance. Attached Figure Description

[0069] Figure 1 This is a diagram illustrating the overall architecture of the TCM syndrome identification system based on multi-source clinical data of this invention.

[0070] Figure 2 This is a schematic diagram of the structure of the multidimensional representation module of symptom relationships in this invention;

[0071] Figure 3 This is a schematic diagram of the structure of the dynamic evolution module of syndrome relationship in this invention;

[0072] Figure 4 This is a schematic diagram of the structure of the proof-type reasoning optimization module of the present invention;

[0073] Figure 5 This is a schematic diagram of the structure of the multi-source data acquisition module of the present invention;

[0074] Figure 6 This is a schematic diagram of the structure of the result output module of the present invention;

[0075] Figure 7 This is a schematic diagram of the structure of the system expansion module of the present invention. Detailed Implementation

[0076] Please refer to Figures 1-7 The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Those skilled in the art should understand that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.

[0077] Reference Figure 1 The TCM syndrome identification system based on multi-source clinical data provided by the present invention includes a multi-source data acquisition module 1, a multi-dimensional representation module of syndrome relationship 2, a dynamic evolution module of syndrome relationship 3, a syndrome reasoning optimization module 4, and a result output module 5.

[0078] The multi-source data acquisition module 1 is used to collect patients' four diagnostic methods data, laboratory test data, and basic patient information. The syndrome relationship multidimensional representation module 2 communicates with the multi-source data acquisition module 1, receiving patient data sent by the multi-source data acquisition module 1. Based on homology topology theory, it constructs a multidimensional topological space for symptoms, syndromes, and syndrome types, mapping patient data into this space and extracting topological features. The syndrome relationship dynamic evolution module 3 communicates with the syndrome relationship multidimensional representation module 2, receiving the extracted topological features. Based on topological persistence theory, it analyzes the stability of syndrome relationships under different intensity thresholds, identifies stable syndrome relationship patterns, and constructs a stable syndrome relationship core structure. The syndrome type reasoning optimization module 4 communicates with the syndrome relationship dynamic evolution module 3, receiving the stable syndrome relationship core structure constructed by the module. Based on the homotopy equivalence principle, it simplifies the reasoning path while maintaining the diagnostic essence, generating a syndrome type set and corresponding confidence set through multi-path parallel reasoning. The result output module 5 is connected to the syndrome type reasoning optimization module 4 and is used to receive the syndrome type set and corresponding confidence set generated by the syndrome type reasoning optimization module 4, and to display the diagnostic results and the syndrome relationship visualization map.

[0079] In embodiments of the present invention, a medical knowledge base management module 6, a knowledge discovery module 7, and a system learning and optimization module 8 may also be included for system knowledge management, knowledge discovery, and continuous optimization.

[0080] The following is a detailed description of each module of the system.

[0081] Reference Figure 5 The multi-source data acquisition module 1 includes a four diagnostic data acquisition unit 11, a laboratory test data acquisition unit 12, a patient basic information acquisition unit 13, and a data preprocessing unit 14.

[0082] The four diagnostic methods data acquisition unit 11 is used to collect data from the patient's inspection, auscultation, inquiry, and palpation. In a preferred embodiment of the present invention, the inspection data includes information such as appearance, expression, complexion, tongue coating, tongue body, and oral cavity; the auscultation data includes information such as odor and sound; the inquiry data includes chief complaints such as pain, chills and fever, diet, and sleep; and the palpation data includes information such as pulse and abdominal palpation.

[0083] The laboratory test data acquisition unit 12 is used to collect patients' biochemical test, imaging examination, and other auxiliary examination data. This examination data provides objective evidence for diagnosis and is an effective supplement to the traditional four diagnostic methods.

[0084] The patient basic information collection unit 13 is used to collect information such as the patient's age, gender, constitution type, and past medical history. This basic information is of great significance for individualized diagnosis. In this invention, the constitution type can include traditional Chinese medicine constitution classifications such as Yin deficiency, Yang deficiency, Qi deficiency, phlegm-dampness, damp-heat, Qi stagnation, and blood stasis.

[0085] The data preprocessing unit 14 is communicatively connected to the four diagnostic methods data acquisition unit 11, the laboratory test data acquisition unit 12, and the patient basic information acquisition unit 13. It is used to standardize the acquired data, remove outliers, and convert unstructured data into a structured data format. Preferably, Z-score normalization is used to normalize numerical data, ensuring a uniform scale for data from different sources and facilitating subsequent analysis. For text data, natural language processing techniques are used to extract key symptom descriptions and convert them into a structured representation.

[0086] Reference Figure 2 The multidimensional representation module 2 for syndrome relationships includes an entity encoding unit 21, a relationship network construction unit 22, a high-dimensional relationship extraction unit 23, and a topological feature calculation unit 24.

[0087] Entity coding unit 21 is used to extract symptoms, syndromes, and syndrome types from the four diagnostic methods data and knowledge base, and assign a unique identifier to each entity. In one embodiment of the present invention, the symptom coding adopts the S001-S999 format, the syndrome coding adopts the Z001-Z999 format, and the syndrome type coding adopts the T001-T999 format, ensuring that each medical entity in the system has a unique identifier. At the same time, according to the knowledge of traditional Chinese medicine theory, each entity is labeled with a type and initial weight.

[0088] The relationship network construction unit 22 is communicatively connected to the entity encoding unit 21, and is used to calculate the initial association strength between entities based on co-occurrence frequency, set an association threshold to filter effective relationships, and construct an initial relationship graph network. In one embodiment of the present invention, the co-occurrence frequency is calculated by analyzing the co-occurrence of symptoms, syndromes, and syndrome types in a large number of clinical cases. The association strength can be calculated using point mutual information (PMI).

[0089] ,

[0090] Where: PMI(x,y) is the point mutual information value between entity x and entity y, representing the association strength between the two entities; p(x,y) represents the probability of entity x and entity y co-occurring, calculated by dividing the co-occurrence frequency by the total number of samples; p(x) represents the probability of entity x appearing alone; p(y) represents the probability of entity y appearing alone; log represents the natural logarithm. The association threshold is preferably set to 0.5, that is, only associations with PMI values ​​greater than 0.5 are retained. This threshold setting can filter out weak associations while retaining sufficient effective information.

[0091] The high-dimensional relation extraction unit 23 is communicatively connected to the relation network construction unit 22, and is used to identify frequently co-occurring symptom combinations, construct a high-dimensional simplicium representing multiple symptom combinations, and calculate the weight and stability index of the simplicium. One of the core innovations of this invention is that it uses a multi-level simplex complex data structure to store medical entity relations. This data structure includes a set of nodes, a set of edges, a set of high-order simplicium, a relation weight matrix, and a simplicium complex structure.

[0092] The node set V is used to store all medical entities and can be represented as:

[0093] ,

[0094] Where: V is the set of nodes; v_{i} represents the i-th medical entity, which can be a symptom, syndrome, or syndrome type; n is the total number of medical entities.

[0095] edge set To represent a direct relationship between entities, it can be represented as:

[0096] ,

[0097] in: Let it be the set of edges; Representing entities and entity There are weights between them. The connection; Representing entities and entity Belongs to the set of nodes ; For entities and entity The correlation weights between them are calculated using PMI values; The association threshold is preferably set to 0.5 to filter valid relationships.

[0098] Higher-order simplex sets To represent a multi-entity composition relationship, it can be represented as:

[0099] ,

[0100] in: It is a set of higher-order simplexes; Indicates by Composed of entities Single-dimensional form; Indicates the constituent simplex One medical entity; This means that each entity constituting a simplex belongs to the set of nodes. ; This represents the weight of the simplex, which is calculated as the geometric mean of the weights of all edges in the simplex. The threshold for higher-order relationships is preferably set to 0.6 to filter valid higher-order relationships. For example, when three symptoms—headache, fever, and thirst—occur frequently together, they can form a 2D simplex (triangle).

[0101] Relationship weight matrix Used to record the strength of association between entities, it can be represented as:

[0102] ,

[0103] in: This is the relation weight matrix; Representing entities and entity The strength of the association between them, i.e., the PMI value; This represents the total number of medical entities; if there is no direct relationship between two entities, then... .

[0104] Monomorphic complex structure Relationships that integrate all dimensions can be represented as:

[0105] ,

[0106] in: It is a monomorphic complex structure; A set of nodes; Let it be the set of edges; It is a set of high-order simplexes. This data structure can capture high-dimensional complex relationships that traditional planar networks cannot represent, providing a foundation for subsequent topology analysis.

[0107] The topology feature calculation unit 24 is communicatively connected to the high-dimensional relationship extraction unit 23, and is used to identify connected components in the network, detect key loop structures, and extract feature vectors to represent the overall topology characteristics. In one embodiment of the present invention, homology group calculation is used to extract topology features. Conflict groups It can reveal complex shapes In These "hollow" structures reflect higher-order relational patterns in the symptom system. For example, Conflict groups This indicates a loop structure, which may correspond to a cyclical relationship between symptoms; Conflict groups This indicates a hollow structure, which may correspond to the syndrome characteristics formed by the combination of multiple symptoms.

[0108] Betti number It is a homology group The rank of represents The number of “holes”. For example, Indicates the number of connected components. Indicates the number of loops. This represents the number of cavities. These topological invariants constitute the eigenvectors of the symptom network:

[0109] ,

[0110] in: The feature vectors of the syndrome network; For the first Dimension Betti number, representing The number of “voids”; As the highest dimension to consider, it is preferably set to Because in traditional Chinese medicine diagnosis, it is usually considered that no more than The combination of symptoms.

[0111] Reference Figure 3 The dynamic evolution module 3 for syndrome relationships includes a multi-scale filtered sequence construction unit 31, a persistent feature tracking unit 32, a stable structure identification unit 33, a dynamic update unit 34, and a noise filtering unit 35.

[0112] The multi-scale filtering sequence construction unit 31 is used to set an incremental relation threshold sequence, construct the corresponding network state under each threshold, and record the network topology changes under different thresholds. In one embodiment of the present invention, the relation threshold sequence can be set as follows:

[0113] ,

[0114] in: For relational threshold sequences; This represents the threshold of the i-th relation; The threshold value represents the total number of thresholds, with a value of 17; the threshold interval. This represents the difference between adjacent thresholds. At each threshold... Below, only retain weights greater than 100%. The relationship forms the corresponding filtered complex. This forms a filtered complex sequence:

[0115] ,

[0116] in: To filter complex sequences; To be at the threshold The filtered complex constructed below; This indicates an inclusion relationship, meaning that a complex constructed at a smaller threshold is contained within a complex constructed at a larger threshold.

[0117] The persistent feature tracking unit 32 is communicatively connected to the multi-scale filtered sequence construction unit 31, and is used to detect the emergence of new topological features, track the disappearance of features, and calculate the persistence metric of features. This invention employs a persistence analysis data structure to store dynamic evolution information, which includes a filtered complex sequence, a persistent interval set, a feature importance index, and an evolutionary trajectory diagram.

[0118] Filtering complex sequences The network state sequences used for storing different thresholds have been defined previously.

[0119] Persistent interval set The lifecycle used to record topological features can be represented as:

[0120] ,

[0121] in: For persistent interval sets; This represents the lifecycle information of the i-th topological feature; The birth value of the feature represents the threshold at which the feature first appears; The death value of the feature represents the threshold at which the feature disappears; The characteristic type can be a connected component (H0), a loop (H1), or a cavity (H2), etc. This represents the total number of topological features. The durability of a feature is defined as... The greater the persistence, the more stable the feature.

[0122] The feature importance index I is used to quantify the stability and significance of features, and can be expressed as:

[0123] ,

[0124] in: It is a set of feature importance indicators; For the first Importance indicators of each feature; This is a function for calculating importance. The persistence of the feature; The birth value is a characteristic; The mortality value is a characteristic; Index for features; This represents the total number of topological features. The importance calculation function can be set as follows: This approach considers both the persistence of features and the timing of their appearance. Features that appear early and are persistent are given higher importance.

[0125] The evolutionary trajectory diagram T is used to record the changes in network structure over time, including the time series of network states and key evolutionary events.

[0126] The stable structure identification unit 33 is communicatively connected to the persistent feature tracking unit 32, and is used to analyze the distribution of persistent intervals, screen long-lived topological features, and construct a core structure of stable syndrome relationships. In one embodiment of the present invention, the persistence threshold β is set to 0.75, that is, only features with a persistence greater than 0.75 are retained as stable features. These stable features constitute the core structure of syndrome relationships, reflecting the most reliable symptom-syndrome-syndrome association pattern in traditional Chinese medicine diagnosis.

[0127] The dynamic update unit 34 is communicatively connected to the stable structure identification unit 33, and is used to monitor the impact of new clinical data, identify network regions that need to be updated, and locally reconstruct the network structure. In one embodiment of the present invention, when the deviation between the new data and the existing model exceeds a preset threshold (e.g., 20%), a local update of the network is triggered. This incremental update mechanism avoids the computational overhead of global reconstruction and improves the system's response speed and adaptability.

[0128] The noise filtering unit 35 is communicatively connected to the persistent feature tracking unit 32, and is used to identify short-lived unstable features, detect abnormal topological changes, and apply an adaptive threshold for noise suppression. In one embodiment of the invention, features with a persistence of less than 0.2 are considered noise features and filtered out. Furthermore, an adaptive threshold mechanism based on statistical anomaly detection is designed, which can dynamically adjust the stringency of noise filtering according to data distribution characteristics.

[0129] Reference Figure 4 The proof-type reasoning optimization module 4 includes a reasoning path planning unit 41, a path simplification and optimization unit 42, a multi-path parallel reasoning unit 43, a result fusion and conflict resolution unit 44, and a feedback and self-optimization unit 45.

[0130] The inference path planning unit 41 is used to identify possible paths from symptoms to syndromes in the syndrome network, calculate the initial weight and reliability of each path, and construct an initial inference space containing all possible paths. In one embodiment of the present invention, an inference optimization data structure is used to store path information. This data structure includes an inference path set, path equivalence classes, a path scoring matrix, and a syndrome inference tree.

[0131] The set of reasoning paths R is used to store possible proof-type reasoning paths, and can be represented as:

[0132] ,

[0133] in: A set of reasoning paths; Indicates the first One reasoning path; Indicates the first in the path One node; Indicates the starting point of the path, which is the symptom node; This indicates the endpoint of the path and is a proof-type node; Indicates the path length; Indicates a set of symptoms; Let represent the set of proof types.

[0134] Path equivalence class The reasoning path used to classify topological equivalence can be represented as:

[0135] ,

[0136] in: It is a set of path equivalence classes; Indicates the first One equivalence class; Indicates the first The th equivalence class Path; Indicates the first The number of paths in each equivalence class; Representing a path and path They are topologically equivalent, meaning they can be transformed into each other through continuous deformation; This indicates that any two paths in an equivalence class are equivalent.

[0137] Path rating matrix The reliability score used to record different paths can be expressed as:

[0138] ,

[0139] in: The path rating matrix; Indicates the first The path to the first Support for individual certificate types; This represents the number of paths. The number of certificate types.

[0140] Proof-of-concept reasoning tree The hierarchical reasoning logic used to organize is a tree structure, with the root node being the initial symptom set, the leaf nodes being the possible syndrome types, and the intermediate nodes being the intermediate reasoning states.

[0141] The path simplification and optimization unit 42 is communicatively connected to the inference path planning unit 41, and is used to identify topologically equivalent path classes, merge redundant paths, retain representative paths, and optimize the path set.

[0142] One of the core innovations of this invention is that it simplifies the reasoning path by using the homotopy equivalence principle, thereby reducing computational complexity while maintaining the diagnostic essence.

[0143] The determination of homotopy equivalence is based on the connectivity principle in topology. The necessary and sufficient condition for two reasoning paths to be homotopic equivalence is that their starting and ending points belong to the same connected component in the symptom relation network; that is, the starting point can reach the ending point, and the ending point can also reach the starting point. The specific determination rule is as follows:

[0144] ,

[0145] in, Representing a path and Determined to be homotopic equivalence Representing a path The starting node, Representing a path The starting node, Representing a path End point ( For path (length) Representing a path End point ( For path (length) Indicates the first Connected components Indicates the first Connected components Indicates a relationship of belonging. This represents the logical AND operation. This indicates a necessary and sufficient condition.

[0146] The topological meaning of this rule is: if the starting points and ending points of two paths belong to the same connected component, then these two paths are topologically equivalent. In proof-based reasoning scenarios, this means that although the specific intermediate nodes traversed by the two paths may be different, because their starting and ending points are mutually reachable in the network structure, the paths are equivalent. The destination is equivalent to the path to reach it. The endpoint can be reached by merging the two paths to simplify the process. Connected component identification employs a depth-first search algorithm for symptom relationship networks. The set of connected components is defined as follows:

[0147] ,

[0148] in, Represents the set of all connected components. Indicates the first A connected component, consisting of a set of interconnected nodes. This represents the total number of connected components. Each connected component satisfies: For Any two nodes in and ,existing from arrive The path also exists from arrive The path.

[0149] Path simplification based on connected components ensures that redundant computations are reduced while maintaining the diagnostic essence. If the start and end points of multiple paths belong to the same connected component, they convey the same diagnostic information in a topological sense. The path with the highest weight can be selected as the representative, and the other paths can be merged.

[0150] The multi-path parallel inference unit 43 is communicatively connected to the path simplification and optimization unit 42. It is used to select the optimal multiple paths for parallel inference, apply corresponding inference rules to each path, and independently calculate the proof result and confidence level of each path. In one embodiment of the invention, the m paths with the highest scores are selected for parallel inference, where m is preferably set to 5. For each path... Calculate its support for each certificate type:

[0151] ,

[0152] in: Representing a path For the certificate type Support level; This represents a chain multiplication operation, which multiplies the weights of all adjacent node pairs in the path. Representing a path Adjacent node pairs in; Indicates adjacent nodes and The association weights between them; this calculation method reflects the impact of each association in the path on the final result.

[0153] The result fusion and conflict resolution unit 44 is communicatively connected to the multi-path parallel inference unit 43. It is used to perform weighted fusion of results based on path reliability, handle result conflicts between different paths, adjust results using traditional Chinese medicine diagnostic principles, and generate a final set of syndrome types and confidence levels. In one embodiment of the invention, a weighted voting mechanism is used to fuse multi-path inference results:

[0154] ,

[0155] in: Indication of certificate type The final confidence level; This represents the summation operation; Representing a path For the certificate type Support level; Indicates the first The weight of a path can be set as the reliability score of that path; The denominator represents the number of paths participating in parallel inference, preferably set to 5; This represents the sum of all path weights, used for normalization to ensure the final confidence level is between 0 and 1.

[0156] When there is a conflict in the results, the primary symptom (the syndrome corresponding to the main symptom) should be given priority over the secondary symptom (the syndrome corresponding to the minor symptom), which is in line with the principle of TCM syndrome differentiation that the primary symptom is primary and the secondary symptom is secondary. In addition, the compatibility between syndromes should also be considered. For example, some syndromes can coexist (such as Qi deficiency syndrome and blood stasis syndrome), while some syndromes are mutually exclusive (such as cold syndrome and heat syndrome).

[0157] The feedback and self-optimization unit 45 is communicatively connected to the result fusion and conflict resolution unit 44, used to record intermediate states of the reasoning process, evaluate the consistency between the reasoning result and clinical feedback, and dynamically adjust path weights and reasoning strategies. In one embodiment of the invention, a self-optimization mechanism based on reinforcement learning is designed. The system adjusts the reasoning parameters according to clinical feedback to gradually improve diagnostic accuracy. Specifically, for each diagnosis, the consistency between the system's diagnostic result and the clinician's confirmation result is recorded. If they are consistent, the weight of the corresponding reasoning path is increased; if they are inconsistent, their weight is decreased, allowing the system to gradually learn the most effective reasoning strategy.

[0158] Reference Figure 6The output module 5 includes a syndrome diagnosis result generation unit 51, a syndrome relationship visualization unit 52, a diagnostic basis explanation unit 53, and a treatment suggestion generation unit 54.

[0159] The syndrome differentiation diagnosis result generation unit 51 is used to generate the patient's main syndrome differentiation diagnosis result based on the syndrome differentiation set and the corresponding confidence level set. In one embodiment of the present invention, the syndrome differentiations are arranged in descending order of confidence level, the syndrome differentiation with the highest confidence level is selected as the main syndrome differentiation, and the secondary syndrome differentiations with confidence levels exceeding a threshold (e.g., 0.3) are listed to form a complete syndrome differentiation result.

[0160] The syndrome relationship visualization unit 52 is used to convert syndrome relationships in a multidimensional topological space into two-dimensional or three-dimensional visualization maps. In one embodiment of the present invention, a dimensionality reduction technique (such as t-SNE or UMAP) is used to project the high-dimensional topological structure onto the visualization space, and different types of entities and relationships are represented by different colors, shapes and lines, so that physicians can intuitively understand the correspondence between patient symptoms and syndrome types.

[0161] The diagnostic basis description unit 53 is used to extract key symptoms and syndrome evidence supporting the syndrome diagnosis and generate a diagnostic basis description. In one embodiment of the present invention, the symptoms and syndromes that contribute most to the final syndrome are extracted from the reasoning path and sorted by contribution to form a structured diagnostic basis report.

[0162] The treatment suggestion generation unit 54 is used to provide corresponding traditional Chinese medicine treatment suggestions based on the syndrome differentiation diagnosis results. In one embodiment of the present invention, the system retrieves corresponding treatment methods, prescriptions, and acupoint suggestions from the knowledge base according to the identified syndrome differentiation as clinical reference. These treatment suggestions do not directly replace the physician's prescription, but rather provide auxiliary decision support.

[0163] Reference Figure 7 The present invention also includes a medical knowledge base management module 6, a knowledge discovery module 7, and a system learning and optimization module 8.

[0164] The medical knowledge base management module 6 is communicatively connected to the syndrome relationship multidimensional representation module 2, the syndrome relationship dynamic evolution module 3, and the syndrome type reasoning optimization module 4, and is used to store traditional Chinese medicine theoretical knowledge, symptom-syndrome-syndrome type relationship rules, and clinical case data. In one embodiment of the present invention, the knowledge base is stored using a graph database and includes basic traditional Chinese medicine theories, syndrome theory, classic medical cases, etc., providing basic knowledge support for the system.

[0165] The knowledge discovery module 7 is communicatively connected to the syndrome relationship dynamic evolution module 3. It is used to discover new syndrome association patterns from clinical data, quantify the contribution of different symptoms to the syndrome type, and feed the discovered new knowledge back to the medical knowledge base management module 6. In one embodiment of the invention, a knowledge discovery algorithm based on topological data analysis is designed, which can automatically identify hidden high-order relationship patterns in the data and compare them with existing knowledge to discover potential new knowledge.

[0166] The system learning and optimization module 8 is communicatively connected to the syndrome type reasoning optimization module 4, and is used to continuously optimize the reasoning strategy and parameter configuration based on clinical feedback to improve the system's diagnostic accuracy. In one embodiment of the invention, an incremental learning method is adopted, enabling the system to continuously learn from new clinical cases, adjust model parameters, and adapt to the dynamic changes in medical practice.

[0167] The workflow of the TCM syndrome identification system based on multi-source clinical data of the present invention is as follows:

[0168] 1. Multi-source data acquisition module 1 collects the patient's four diagnostic methods data, laboratory test data, and basic patient information, and forms structured patient data after standardization processing.

[0169] 2. The multidimensional representation module 2 of syndrome relationship receives patient data, constructs a multidimensional topological space of symptoms, syndromes and syndrome types based on homology topology theory, maps patient data into the multidimensional topological space, and extracts topological features.

[0170] 3. The dynamic evolution module 3 of syndrome relationship receives topological features, analyzes the stability of syndrome relationship under different intensity thresholds based on the theory of topological persistence, identifies stable syndrome relationship patterns, and constructs the core structure of stable syndrome relationship.

[0171] 4. The evidence type reasoning optimization module 4 receives the core structure of stable evidence relationship, simplifies the reasoning path based on the homotopy equivalence principle, and generates evidence type set and corresponding confidence set through multi-path parallel reasoning.

[0172] 5. The results output module 5 receives the set of syndrome types and the corresponding set of confidence levels, displays the diagnostic results and a visualization of the syndrome relationship, and provides explanations of the diagnostic basis and treatment suggestions.

[0173] 6. The Medical Knowledge Base Management Module 6 provides basic knowledge support for the system, the Knowledge Discovery Module 7 discovers new syndrome correlation patterns from clinical data, and the System Learning and Optimization Module 8 continuously optimizes system performance based on clinical feedback.

[0174] This invention presents a TCM syndrome identification system based on multi-source clinical data. By deeply integrating topological theory with TCM syndrome differentiation practice, it achieves highly accurate, stable, and efficient intelligent identification of TCM syndromes. Compared with existing technologies, this invention has the following significant advantages:

[0175] 1. Multidimensional representation capability: The multidimensional topological space constructed through homology topology theory can effectively capture the complex high-dimensional relationships between symptoms, syndromes and syndrome types, overcoming the limitations of traditional planar network representation.

[0176] 2. Noise resistance: Based on the dynamic evolution mechanism of topological persistence theory, it can identify stable syndrome relationship patterns from noisy data, which improves the robustness of the system when faced with incomplete or uncertain symptom information.

[0177] 3. Reasoning efficiency: Based on homotopy equivalence path simplification and multi-path parallel reasoning mechanism, the computational complexity is optimized while maintaining the diagnostic essence, thereby improving reasoning efficiency and result reliability.

[0178] 4. Personalized diagnosis: The system takes into account individual factors such as the patient's constitution and age, achieving more accurate diagnosis based on syndrome differentiation, which is in line with the TCM concept of individualized treatment.

[0179] 5. Knowledge discovery capability: The system can automatically discover new syndrome correlation patterns from clinical data, promoting the improvement and development of the TCM theoretical system.

[0180] 6. Multi-source data fusion: The system effectively integrates multi-source heterogeneous data such as diagnostic data, laboratory test data, etc., improving the comprehensiveness and accuracy of diagnosis.

[0181] This invention deeply integrates cutting-edge topological theories with TCM diagnostic knowledge, providing a new technical approach for the intelligentization of TCM syndrome differentiation, and has significant clinical application value and theoretical innovation significance.

[0182] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A traditional Chinese medicine syndrome type intelligent identification system based on multi-source clinical data, characterized in that, The method comprises the following steps: A multi-source data acquisition module is used to collect four diagnostic data, laboratory examination data and patient basic information of a patient; A syndrome relationship multi-dimensional representation module is in communication connection with the multi-source data acquisition module, and is used to receive patient data sent by the multi-source data acquisition module, construct a multi-dimensional topological space of symptoms, syndromes and syndrome types based on homotopy theory, map the patient data into the multi-dimensional topological space, and extract topological features; A syndrome relationship dynamic evolution module is in communication connection with the syndrome relationship multi-dimensional representation module, and is used to receive topological features extracted by the syndrome relationship multi-dimensional representation module, analyze the stability of the syndrome relationship under different intensity thresholds based on topological persistence theory, identify stable syndrome relationship patterns, and construct a stable syndrome relationship core structure; A syndrome type reasoning optimization module is in communication connection with the syndrome relationship dynamic evolution module, and is used to receive the stable syndrome relationship core structure constructed by the syndrome relationship dynamic evolution module, simplify reasoning paths on the premise of maintaining the essence of diagnosis based on homotopy equivalence principle, and generate a syndrome type set and a corresponding confidence set through multi-path parallel reasoning; A result output module is in communication connection with the syndrome type reasoning optimization module, and is used to receive the syndrome type set and the corresponding confidence set generated by the syndrome type reasoning optimization module, and display a diagnosis result and a syndrome relationship visualization graph.

2. The system of claim 1, wherein, The syndrome relationship multi-dimensional representation module comprises: An entity coding unit is used to extract symptoms, syndromes and syndrome types from four diagnostic data and a knowledge base, and assign a unique identifier to each entity; A relationship network construction unit is in communication connection with the entity coding unit, and is used to calculate the initial association strength between entities based on co-occurrence frequency, set an association threshold to screen effective relationships, and construct an initial relationship graph network; A high-dimensional relationship extraction unit is in communication connection with the relationship network construction unit, and is used to identify frequently co-occurring symptom combinations, construct high-dimensional simplices representing multi-symptom combinations, and calculate the weight and stability index of the simplices; A topological feature calculation unit is in communication connection with the high-dimensional relationship extraction unit, and is used to identify connected components in the network, detect key loop structures, and extract a feature vector to represent the overall topological properties.

3. The system of claim 2, wherein, The high-dimensional relationship extraction unit stores medical entity relationships in a multi-level simplex data structure, which comprises: A node set is used to store all medical entities; An edge set is used to represent the direct association between entities; A high-order simplex set is used to represent multi-entity combination relationships; A relationship weight matrix is used to record the association strength between entities; A simplex complex structure is used to integrate relationships in all dimensions.

4. The system of claim 1, wherein, The syndrome relationship dynamic evolution module comprises: A multi-scale filtering sequence construction unit is used to set an increasing relationship threshold sequence, construct a corresponding network state at each threshold, and record the network topological changes under different thresholds; A persistence feature tracking unit is in communication connection with the multi-scale filtering sequence construction unit, and is used to detect the appearance of new topological features, track the disappearance of features, and calculate the persistence measure of the features; A stable structure identification unit, in communication with the persistence feature tracking unit, is configured to analyze the persistence interval distribution, filter long-life topological features, and construct a stable syndrome relationship core structure; A dynamic update unit, in communication with the stable structure identification unit, is configured to monitor the influence of new clinical data, identify network regions that need to be updated, and locally reconstruct the network structure; A noise filtering unit, in communication with the persistence feature tracking unit, is configured to identify short-life unstable features, detect abnormal topological structure changes, and apply an adaptive threshold for noise suppression.

5. The system of claim 4, wherein, The persistence feature tracking unit stores dynamic evolution information using a persistence analysis data structure, which includes: A filtered complex sequence for storing network state sequences at different thresholds; A persistence interval set for recording the life cycle of topological features; A feature importance index for quantifying the stability and significance of features; An evolution trajectory graph for recording changes in network structure over time.

6. The system of claim 1, wherein, The syndrome reasoning optimization module includes: A reasoning path planning unit for identifying possible paths from symptoms to syndromes in the syndrome network, calculating the initial weight and reliability of each path, and constructing an initial reasoning space containing all possible paths; A path simplification and optimization unit, in communication with the reasoning path planning unit, for identifying topologically equivalent path classes, merging redundant paths, retaining representative paths, and optimizing the path set; A multi-path parallel reasoning unit, in communication with the path simplification and optimization unit, for selecting the optimal multiple paths for parallel reasoning, applying corresponding reasoning rules to each path, and independently calculating the syndrome results and confidence of each path; A result fusion and conflict resolution unit, in communication with the multi-path parallel reasoning unit, for weighting and fusing results based on path reliability, handling conflicts between different paths, applying TCM syndrome differentiation principles for result adjustment, and generating a final syndrome set and confidence; A feedback and self-optimization unit, in communication with the result fusion and conflict resolution unit, for recording the intermediate state of the reasoning process, evaluating the consistency of the reasoning results and clinical feedback, and dynamically adjusting path weights and reasoning strategies.

7. The system of claim 6, wherein, The reasoning path planning unit stores path information using a reasoning optimization data structure, which includes: A reasoning path set for storing possible syndrome reasoning paths; A path equivalence class for classifying topologically equivalent reasoning paths; A path scoring matrix for recording the reliability scores of different paths; A syndrome reasoning tree for organizing hierarchical reasoning logic.

8. The system of claim 1, wherein, The multi-source data acquisition module includes: A four-diagnosis data acquisition unit for collecting patient's pulse, tongue, smell, and touch data; A laboratory examination data acquisition unit for collecting patient's biochemical test, imaging examination, and other auxiliary examination data; A patient basic information acquisition unit for collecting patient's age, gender, constitution type, and past medical history information; A data preprocessing unit is in communication connection with the four diagnostic data acquisition unit, the laboratory examination data acquisition unit and the patient basic information acquisition unit, configured to standardize the collected data, remove abnormal values, and convert unstructured data into structured data format.

9. The system of claim 1, wherein, The result output module includes: A syndrome type diagnosis result generation unit configured to generate a main syndrome type diagnosis result of the patient based on the syndrome type set and the corresponding confidence set; A syndrome relationship visualization unit configured to convert the syndrome relationship in the multi-dimensional topological space into a two-dimensional or three-dimensional visual atlas; A diagnosis basis explanation unit configured to extract key symptoms and syndrome evidence supporting the syndrome type diagnosis, and generate a diagnosis basis explanation; A treatment suggestion generation unit configured to provide corresponding TCM treatment suggestions based on the syndrome type diagnosis result.

10. The system of claim 1, wherein, Further comprising: A medical knowledge base management module in communication connection with the syndrome relationship multi-dimensional representation module, the syndrome relationship dynamic evolution module and the syndrome type reasoning optimization module, configured to store TCM theoretical knowledge, symptom-syndrome-syndrome type relationship rules and clinical case data; A knowledge discovery module in communication connection with the syndrome relationship dynamic evolution module, configured to discover new syndrome correlation patterns from clinical data, quantify the contribution of different symptoms to the syndrome type, and feed the discovered new knowledge back to the medical knowledge base management module; A system learning and optimization module in communication connection with the syndrome type reasoning optimization module, configured to continuously optimize the reasoning strategy and parameter configuration based on clinical feedback, and improve the system diagnosis accuracy.