Traditional Chinese medicine service query system based on knowledge graph

The knowledge graph-based TCM service query system addresses the issues of fragmented TCM data and the shortcomings of traditional retrieval methods. It achieves unified symbolic processing and cross-dimensional semantic reasoning of TCM knowledge, thereby optimizing the systematic nature and accuracy of query results.

CN121636714APending Publication Date: 2026-03-10ZHEJIANG CANCER HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing TCM data is scattered, lacks unified knowledge organization and cross-source data fusion, and traditional retrieval methods are unable to handle the complex compatibility relationships and syndrome coupling characteristics among medicinal materials. Existing knowledge graph research has failed to effectively model the dynamic weights of medicinal material compatibility, the nonlinear correlation of syndrome coupling, and the overall rationality of the pathogenesis link, resulting in insufficient relevance or excessive redundancy in query results.

Method used

The TCM service query system based on knowledge graphs generates symbolic coded entities through a data collection and symbolization module, establishes semantic relationships between syndromes and medicinal materials through a syndrome coupling module, analyzes the compatibility relationships of medicinal materials through a prescription fusion module, scores multi-hop reasoning paths through a link tension module, generates relevance scores through a query matching module, and sorts and generates the final query results through a joint scoring module. The system optimizes candidate results through multi-source knowledge representation and pathogenesis link tension factors.

Benefits of technology

It achieves unified symbolic processing and semantic fusion of TCM knowledge, supports cross-dimensional semantic reasoning, optimizes candidate result screening, improves the systematicness and accuracy of query results, and ensures the rationality and relevance of the results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121636714A_ABST
    Figure CN121636714A_ABST
Patent Text Reader

Abstract

The invention discloses a traditional Chinese medicine service query system based on a knowledge graph, and relates to the technical field of knowledge graph and traditional Chinese medicine informatization. The system comprises a data acquisition and symbolization module for acquiring medicinal material and prescription data, generating a multi-dimensional medicinal property characteristic entity through symbolization processing, a syndrome coupling module for constructing and verifying a syndrome coupling relationship, a prescription fusion module for fusing prescription compatibility weights to generate multi-source knowledge representation, and a prescription matching module for matching the multi-source knowledge representation with the prescription compatibility weights. The link tension module completes multi-hop reasoning based on pathogenesis link tension factors, the query matching module generates candidate entity relevancy scores in combination with user query conditions, the joint scoring module obtains a preliminary candidate set through joint scoring, and the curative effect redundancy module finally performs comprehensive sorting on the candidate set through curative effect coverage and redundancy suppression optimization. And generating a query result. Through unified organization of traditional Chinese medicine knowledge, enhancement of semantic reasoning ability and optimization of candidate result screening, integrated, precise and efficient query oriented to diagnosis and treatment logic is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of knowledge graph and information technology in traditional Chinese medicine, specifically a traditional Chinese medicine service query system based on knowledge graph. Background Technology

[0002] Traditional Chinese medicine (TCM), as an important component of traditional Chinese medicine, encompasses multiple research directions including formulary, pharmacology, and clinical diagnosis and treatment, boasting a long history and abundant literature. With the systematic organization of medical literature and the digital storage of modern clinical cases, a large amount of TCM-related data is gradually being stored electronically. However, existing data is largely scattered across medical databases, modern medicinal material databases, formulary databases, and clinical case records, lacking unified knowledge organization and cross-source data fusion, resulting in semantic fragmentation and structural inconsistencies among the data.

[0003] In the knowledge system of Traditional Chinese Medicine (TCM), the properties and meridian tropism of medicinal materials, the compatibility rules of prescriptions, and the correspondence between syndromes and pathogenesis constitute a complex multidimensional network of relationships. Traditional TCM retrieval methods rely heavily on keyword matching and manual experience-based judgment, making it difficult to handle the complex compatibility relationships between medicinal materials and the coupling characteristics of syndromes, and lacking a systematic expression of cross-dimensional knowledge. Furthermore, existing knowledge organization methods are mostly limited to linear or tabular structures, making it difficult to support multi-hop reasoning and semantic-level logical deduction, and failing to meet the higher-order needs for diagnostic assistance and knowledge discovery.

[0004] With the development of knowledge graph technology, it has become possible to model entities, relationships, and attributes in a unified graph structure. Knowledge graphs can establish semantic relationships between medicinal materials, prescriptions, syndromes, and pathogenesis, supporting complex semantic queries and multi-hop path reasoning. However, most existing TCM research based on knowledge graphs remains at the level of static representation of entity relationships, lacking systematic modeling of the dynamic weights of medicinal material compatibility, the nonlinear associations of syndrome coupling, and the overall rationality of pathogenesis links. Furthermore, insufficient redundancy suppression and efficacy coverage control of candidate entities during the query process lead to problems such as insufficient relevance or excessive redundancy in the result set.

[0005] Therefore, there is an urgent need for a TCM service query system based on knowledge graphs to achieve unified symbolic processing and semantic fusion of multi-source heterogeneous TCM data, construct a multi-source knowledge representation that can comprehensively characterize the compatibility of medicinal materials, syndrome coupling, and pathogenesis links, and on this basis, realize joint scoring of candidate results, efficacy coverage control, and ranking of final query results, so as to improve the systematicness and accuracy of TCM knowledge retrieval. Summary of the Invention

[0006] Based on the shortcomings of the prior art described above, the purpose of this invention is to provide a knowledge graph-based traditional Chinese medicine service query system to solve the aforementioned technical problems.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a traditional Chinese medicine service query system based on knowledge graphs, comprising:

[0008] The data acquisition and symbolization module collects data related to traditional Chinese medicine, processes the collected data to generate symbolic coded entities, extracts multidimensional medicinal property features, and generates a set of symbolic entities.

[0009] The syndrome coupling module calculates the syndrome coupling factor between entities based on multidimensional medicinal properties, establishes the semantic relationship between syndromes and medicinal materials, and generates a preliminary set of triplet relationships.

[0010] The formula fusion module analyzes the compatibility relationships between medicinal materials within a formula, generates formula combination rule weights, and fuses them to generate multi-source knowledge representations.

[0011] The link tension module generates pathogenesis link tension factors based on multi-source knowledge representation, scores the rationality of multi-hop reasoning paths between entities, and generates candidate entity scoring information.

[0012] The query matching module maps user query conditions to entity feature space, and generates query matching relevance score by combining the weight of prescription combination rules, syndrome coupling factor and pathogenesis link tension factor.

[0013] The joint scoring module calculates a joint score based on query matching relevance score, multi-source knowledge representation, and pathogenesis link tension factor, and sorts the candidate entities to generate a preliminary candidate set.

[0014] The efficacy redundancy module processes the initial candidate set by analyzing the differences in the medicinal properties and syndrome coupling factors of the candidate entities to generate an optimized candidate set.

[0015] The final ranking module comprehensively ranks the optimized candidate set based on the combined score and efficacy coverage, generating the final query result set.

[0016] The present invention is further configured such that the data acquisition and symbolization module includes:

[0017] The medicinal materials database and prescription information are cleaned, deduplicated, and field-normalized to generate a preliminary entity set;

[0018] Each entity in the initial entity set is transformed into a symbolic encoded entity through mapping;

[0019] Calculate multidimensional pharmacological features for each symbolic coded entity;

[0020] The symbolic encoded entities are integrated with multidimensional pharmacological features to generate a complete set of symbolic entities.

[0021] The present invention is further configured such that the syndrome coupling module includes:

[0022] Calculate the syndrome coupling factor between entities based on multidimensional drug property characteristics;

[0023] Semantic relationships between entities are established based on syndrome coupling factors, and semantic relationship scores are generated by combining syndrome weights and attribute difference adjustment factors.

[0024] The semantic relationship scores are filtered and consistency is verified to remove entity relationships below the threshold and generate filtered and verified semantic relationships.

[0025] The filtered and verified semantic relations are converted into a preliminary set of triple relations, which consists of entity, relation and corresponding symptom dimensions.

[0026] The present invention is further configured such that the formula fusion module includes:

[0027] The compatibility strength of the medicinal materials in each prescription is calculated to quantify the synergistic effect and nonlinear correlation between the medicinal materials in the prescription.

[0028] Based on the compatibility relationship of medicinal materials and the coupling factor of syndromes between entities, the weight of the formula combination rule is generated to reflect the overall medicinal properties and compatibility of the formula.

[0029] By integrating the weights of formula combination rules with multidimensional drug properties and inter-entity syndrome coupling factors, a multi-source knowledge representation is constructed, generating a comprehensive knowledge vector for each entity in terms of formula and syndrome dimensions.

[0030] The present invention is further configured such that the link tension module includes:

[0031] Construct a set of multi-hop reasoning paths for each entity, which includes all reasonable paths from that entity to other candidate entities. The path length is limited by a preset maximum number of hops.

[0032] For each multi-hop reasoning path, generate a pathogenesis link tension factor, and integrate the multi-source knowledge representation of each entity in the path and the symptom coupling factor between entities;

[0033] The path tension factor is fused with the combination rule weights of the formulas containing each entity in the path to generate candidate entity scoring information.

[0034] The present invention is further configured such that the query matching module includes:

[0035] Map user query conditions to entity feature space to generate query feature representation;

[0036] A preliminary relevance under difference constraints is constructed based on query feature representation and multi-source knowledge representation of candidate entities;

[0037] By combining the inter-entity syndrome coupling factor and the weight of the prescription combination rule, the preliminary correlation is adjusted to generate the intermediate correlation.

[0038] Based on the relevant path set of query conditions, a pathogenesis link tension factor is introduced to comprehensively correct the intermediate relevance through multiple paths and generate a query matching relevance score.

[0039] The present invention is further configured such that the joint scoring module includes:

[0040] Construct a multi-hop reasoning path for each candidate entity, and aggregate the formula combination rule weights and pathogenesis link tension factors of each entity in the path to generate the path tension aggregation quantity;

[0041] The neighborhood consistency multiplicative penalty is generated based on the difference in multi-source knowledge representation between candidate entities and their neighboring entities and the symptom coupling factor.

[0042] The query matching relevance score, multi-source knowledge representation, path tension aggregation amount and neighborhood consistency multiplicative penalty amount are fused to generate a joint score for candidate entities;

[0043] Candidate entities are sorted based on joint scores, and a preliminary candidate set is generated according to preset truncation rules.

[0044] The present invention is further configured such that the processing of the preliminary candidate set includes:

[0045] By integrating multi-source knowledge representation with neighborhood symptom coupling factors, an attribute matrix of candidate entities is constructed.

[0046] Based on the differences in the physical properties of the drug and the coupling factors of syndromes, an efficacy coverage index is generated.

[0047] Redundancy assessment is performed by considering the differences in drug properties and syndrome coupling between entities, and a similarity matrix is ​​constructed among candidate entities to identify highly similar candidate entities.

[0048] The present invention is further configured such that generating the optimization candidate set includes:

[0049] The initial candidate set is filtered based on the redundancy determination threshold to generate a de-redundancy candidate set;

[0050] The redundant candidate set is optimized by weighting the efficacy coverage to generate an optimized candidate efficacy score.

[0051] The candidate entities are sorted according to their efficacy scores, a preset number of candidate entities are extracted, and the final set of candidate entities is generated.

[0052] The present invention is further configured such that the final sorting module includes:

[0053] A comprehensive score for candidate entities is constructed by combining joint scoring, optimized candidate efficacy coverage, neighborhood multi-source knowledge representation, and syndrome coupling factors, which serves as the basis for final ranking.

[0054] The candidate set for optimization is sorted in descending order according to the comprehensive score of the candidate entities to generate a sorted candidate set.

[0055] The predetermined number of candidate entities are extracted from the sorted set based on the predetermined number of queries, and the final query result set is generated.

[0056] This invention provides a Traditional Chinese Medicine (TCM) service query system based on knowledge graphs. The system employs a data acquisition and symbolization module to collect TCM-related data, process the collected data to generate symbolic coded entities, extract multi-dimensional medicinal property features, and generate a set of symbolic entities. A syndrome coupling module calculates syndrome coupling factors between entities based on multi-dimensional medicinal property features, establishes semantic relationships between syndromes and medicinal materials, and generates a preliminary set of triplet relationships. A prescription fusion module analyzes the compatibility relationships between medicinal materials within a prescription, generates prescription combination rule weights, and fuses them to generate a multi-source knowledge representation. A link tension module generates pathogenesis link tension factors based on the multi-source knowledge representation, scores the rationality of multi-hop reasoning paths between entities, and generates syndromes. The system selects entity scoring information; the query matching module maps user query conditions to entity feature space, combines formula combination rule weights, syndrome coupling factors, and pathogenesis link tension factors to generate query matching relevance scores; the joint scoring module calculates joint scores based on query matching relevance scores, multi-source knowledge representation, and pathogenesis link tension factors, and sorts candidate entities to generate a preliminary candidate set; the efficacy redundancy module processes the preliminary candidate set, analyzes the differences in drug properties and syndrome coupling factors of candidate entities to generate an optimized candidate set; the final ranking module comprehensively ranks the optimized candidate set based on the joint scores and efficacy coverage to generate the final query result set. The beneficial effects include:

[0057] 1. Unified organization of TCM knowledge: Through symbolic processing of multi-source heterogeneous data and knowledge graph modeling, a systematic semantic association between medicinal materials, prescriptions, syndromes and pathogenesis is generated, avoiding the problems of data dispersion and structural fragmentation in traditional methods, and realizing the integrated expression of TCM knowledge;

[0058] 2. Precise semantic reasoning ability: Through hierarchical modeling of compatibility relationships, syndrome coupling and pathogenesis links, it supports cross-dimensional semantic reasoning and multi-hop queries, enabling logical deduction of complex diagnostic and treatment knowledge and improving the consistency between query results and clinical logic;

[0059] 3. Optimized candidate result screening: By combining the combined score and efficacy coverage, the candidate results are redundancy suppressed and rationality controlled to ensure the systematicness and relevance of the final query result set and improve the effectiveness of knowledge retrieval.

[0060] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0061] Figure 1 The flowchart illustrates a knowledge graph-based traditional Chinese medicine service query system as an exemplary embodiment of the present invention. Detailed Implementation

[0062] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.

[0063] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0064] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.

[0065] A knowledge graph-based traditional Chinese medicine service query system, such as Figure 1 As shown, it includes:

[0066] The data acquisition and symbolization module collects data related to traditional Chinese medicine, processes the collected data to generate symbolic coded entities, extracts multidimensional medicinal property features, and generates a set of symbolic entities.

[0067] The syndrome coupling module calculates the syndrome coupling factor between entities based on multidimensional medicinal properties, establishes the semantic relationship between syndromes and medicinal materials, and generates a preliminary set of triplet relationships.

[0068] The formula fusion module analyzes the compatibility relationships between medicinal materials within a formula, generates formula combination rule weights, and fuses them to generate multi-source knowledge representations.

[0069] The link tension module generates pathogenesis link tension factors based on multi-source knowledge representation, scores the rationality of multi-hop reasoning paths between entities, and generates candidate entity scoring information.

[0070] The query matching module maps user query conditions to entity feature space, and generates query matching relevance score by combining the weight of prescription combination rules, syndrome coupling factor and pathogenesis link tension factor.

[0071] The joint scoring module calculates a joint score based on query matching relevance score, multi-source knowledge representation, and pathogenesis link tension factor, and sorts the candidate entities to generate a preliminary candidate set.

[0072] The efficacy redundancy module processes the initial candidate set by analyzing the differences in the medicinal properties and syndrome coupling factors of the candidate entities to generate an optimized candidate set.

[0073] The final ranking module comprehensively ranks the optimized candidate set based on the combined score and efficacy coverage, generating the final query result set.

[0074] The present invention is further configured such that the data acquisition and symbolization module includes:

[0075] The medicinal herb database and prescription information are cleaned, deduplicated, and their fields are normalized to generate a preliminary entity set. Specifically, raw records are collected from the medicinal herb database and prescription information, and text cleaning, duplicate detection, field normalization, and terminology synonym mapping are performed to generate a preliminary entity set. Text cleaning includes removing scanning noise, formatting marks, and annotation symbols; standardizing character encoding and sentence boundaries; duplicate detection includes similarity assessment and merging of synonymous entries, bibliographical variations, and version differences; field standardization includes unifying field names, unit normalization, and dosage format standardization for structured records; and terminology synonym mapping includes mapping ancient and modern synonyms, different names for prescriptions, and alternative names for medicinal materials to a unified terminology table. These are raw data records, sourced from medicine, medical records, or databases. To record the total number, a unified set of symbolic entities is constructed through standardized collection, synonym mapping, and symbolic encoding, providing basic data support for cross-source semantic fusion and traceability verification.

[0076] Each entity in the initial entity set is transformed into a symbolic encoded entity through mapping; specifically, for the initial entity set... Apply named entity recognition and field extraction rules to each record Mapped to symbolic encoded entities At the same time, a basic feature set is constructed at the entity level. As primitive values ​​represented symbolically, mapping relations are represented as mapping operators. The calculation formula is: ,in, To record the raw data Transform into symbolic entities The mapping function merges the structured fields, named entity recognition results, and semantic mapping results into a unique symbol code, generating... For symbolic encoded entities, For the original entity in the first The feature values ​​for each attribute dimension are obtained through text parsing and structured field extraction, and their values ​​range from [0, 10]. As symbolic moderating factors, the effects of quadratic and cubic terms on the encoding are amplified respectively. The quadratic term captures moderate-intensity effects, while the cubic term captures higher-order coupling or threshold behavior. These two factors are initialized based on historical statistics or expert experience and can be calibrated offline, with values ​​ranging from [0.1, 2.0], derived from historical literature statistics and pharmacological knowledge bases. The number of entity feature dimensions depends on the set of attributes of medicinal materials and prescriptions, such as properties, meridian tropism, efficacy, and dosage.

[0077] Calculate multidimensional pharmacological features for each symbolic coded entity; specifically, based on symbolic coded entities... Basic attribute values Overall potential indicators of the entity Calculate the multidimensional pharmacological characteristics of each entity. The calculation formula is: , ,in, For entities The multidimensional medicinal properties of medicinal materials or prescriptions can be used to quantify the intensity and extensibility of their effects across different attribute dimensions. For entity number 1 eigenvalues, with values ​​ranging from [0, 10]. This represents the overall pharmacological potential value of the entity, ranging from [0, 5], and strengthens the coupling relationship of internal attributes of the entity. The weights for drug properties are used to adjust the contribution of different dimensions to drug properties, with values ​​ranging from [0.1, 2.0]; the quadratic term... This indicates amplification of the strength of a single-dimensional attribute, a cubic term. This indicates higher-order coupling effects of attributes; cross terms. This represents the coupling relationship between single-dimensional attributes and overall pharmacological potential, ensuring that multidimensional pharmacological characteristics can reflect the overall pharmacological potential of the entity; the overall pharmacological potential value of the entity. The calculation formula is: ,in For the first Evidence factors, The number of evidence factors, such as the normalized number of clinically effective cases, the intensity of modern pharmacological experiments, and the frequency of commonly used doses; coefficients. To correspond to the weights, historical statistics are initialized, and the nonlinear suppression of neighborhood evidence is achieved through the product term of the denominator, thereby enhancing robustness. By combining the entity's overall potential coupling term with quadratic and cubic polynomial mapping, the generated multidimensional pharmacological features can characterize higher-order pharmacological relationships, thereby improving the expressive power of subsequent syndrome coupling and compatibility analysis.

[0078] The symbolic coded entities are integrated with multidimensional pharmacological features to generate a complete set of symbolic entities. Specifically, the coded scalars are... With multidimensional vectors Combined into a complete symbolic entity representation The set is represented as The combination is expressed as A complete set representation It facilitates direct input relationship extraction, prescription compatibility weight calculation, and multi-hop path reasoning, shortens the engineering integration link, and improves the efficiency of knowledge graph construction.

[0079] The present invention is further configured such that the syndrome coupling module includes:

[0080] Calculate the syndrome coupling factor between entities based on multidimensional drug property characteristics; specifically, based on the symbolic entity set. Multidimensional eigenvectors Higher-order coupling operators are constructed between entity pairs according to feature dimensions to generate coupling matrices. The formula for quantifying the coupling strength between any two entities in the symptom semantic space is as follows: ,in, For entities and The symptom coupling factor quantifies the coupling strength between two entities in terms of symptom characteristics. This represents the potential synergy coefficient between entities, with a value range of [0, 5], which enhances higher-order coupling effects between entities. This is the coupling weight factor, with a value range of [0.1, 2.0] to ensure that the contribution of features of different dimensions to the coupling factor is adjustable. To enhance the asymmetric coupling effect in high-value dimensions, it is ensured that when one feature is extremely strong in a certain dimension, the influence of that dimension on the coupling is superlinearly amplified. Introducing potential synergistic factors of entity pairs Enhance the coupling value of entity pairs with clear evidence of historical compatibility or clinical co-occurrence; achieve cross-dimensional accumulation through dimensional summation but avoid simple linear superposition; and present non-linear interactions and sensitivity differences between dimensions through higher-order powers.

[0081] Semantic relationships between entities are established based on syndrome coupling factors across various syndrome dimensions, and semantic relationship scores are generated by combining syndrome weights and attribute difference moderating factors; specifically, based on syndrome coupling factors... Based on the difference in single-dimensional attributes, along each symptom dimension Calculate semantic relation scores This means that the coupling strength and dimensional importance jointly determine the semantic relation value, while attribute differences exert nonlinear suppression on the relation value. The calculation formula is as follows: ,in, For entities and In the Semantic relation scoring on each symptom dimension This is a syndrome weighting factor, with a value range of [0.1, 2.0], which adjusts the importance of each syndrome dimension. The symptom-sensitive moderating factor, with a value range of [0.05, 1.0], controls the suppression of semantic relationships by attribute differences. The attribute difference cube term enhances the nonlinear suppression effect of dimensional differences. For the number of symptom dimensions, Multiplying the overall coupling strength by the importance weight of the symptom dimension reflects the contribution of that dimension to the coupling semantics. Employing a difference-based cubic suppression structure, when the difference between two entities increases along a certain dimension, the corresponding semantic relationship rapidly decays, preventing erroneous judgments of local dimensional correlation due to strong overall coupling. The result... For non-negative real numbers, the quantization is at the th... Semantic affinity across the symptom dimension;

[0082] The semantic relationship scores are filtered and consistency checked to remove entity relationships below a threshold, generating filtered and checked semantic relationship scores; specifically, the semantic relationship score for each pair of entities is calculated for each symptom dimension. Application consistency threshold The system determines and filters out high-confidence relations while removing low-confidence or inconsistent relations, thereby generating a robust set of semantic relations. The calculation formula is: ,in, , To score the semantic relationships after filtering and verification, This is a relation consistency threshold, with a value range of [0.1, 5.0], used to eliminate weak relations. The consistency determination function ensures that semantic relations are retained only if they meet a minimum threshold. The threshold value is either 0 or 1, implementing hard threshold filtering to ensure that only relationships with acceptable confidence are retained. The settings can be adaptively configured based on entity type or symptom dimension differences, or an empirical threshold range obtained from offline calibration can be used to output the results. This represents the filtered semantic relationships; a value of zero indicates that the relationship in that dimension has been removed.

[0083] The filtered and validated semantic relations are converted into a preliminary set of triple relations, consisting of entity, relation, and corresponding symptom dimensions. Specifically, the filtered semantic relations... Mapped to triples Among them, relationship tags Corresponding to the Based on the semantic category or relation type of the syndrome, generate a preliminary set of triples. The expression is: Among them, relationship tags The mapping is determined by the symptom ontology mapping rules, which can be implemented based on a predefined symptom category table and semantic type set. Each triple simultaneously records the confidence value. Source evidence, citation count, and timestamps, combined Imported entity-relationship-entity records from the graph database are stored as knowledge graph triples.

[0084] The present invention is further configured such that the formula fusion module includes:

[0085] The compatibility strength of medicinal material combinations within each prescription is calculated to quantify the synergistic and nonlinear relationships between medicinal materials within the prescription. Specifically, in traditional Chinese medicine prescriptions, synergistic or antagonistic effects often exist between medicinal materials, requiring quantification of the combination relationship between any two medicinal materials within the prescription. By introducing medicinal property feature vectors and nonlinear higher-order operations, the compatibility strength between medicinal materials is characterized, generating a medicinal material compatibility matrix. The calculation formula is as follows: ,in, Formula Internal medicinal materials and The strength of the compatibility relationship, characterizing the synergistic effect between medicinal materials. , medicinal materials , The Uyghur medicinal properties, with values ​​ranging from [0, 10]. , This is the matching weighting factor, with a value range of [0.1, 2.0]. The number of dimensions of medicinal properties. Enhance the higher-order synergistic effect of drug properties. By capturing the nonlinear inhibitory effect of differences in medicinal properties on compatibility, the compatibility matrix of medicinal materials enables a high-order nonlinear characterization of the compatibility relationship between medicinal materials, which can ensure that the rationality of compatibility and the inhibition of differences are balanced in quantitative calculation.

[0086] Based on the compatibility relationships of medicinal materials and the syndrome coupling factors between entities, a formula combination rule weight is generated to reflect the overall medicinal properties and compatibility rationality of the formula. Specifically, after obtaining the strength of the medicinal material compatibility relationships, it is necessary to comprehensively consider the syndrome coupling factors to generate the overall formula combination rule weight. By introducing a numerator to enhance the medicinal material compatibility effect and introducing a syndrome inhibition term in the denominator, the overall formula weight reflects both the synergy between medicinal materials and avoids excessively high scores caused by mismatched medicinal properties. The formula for calculating the overall formula combination rule weight is as follows: ,in, Formula The weighting of the combination rules comprehensively reflects the synergistic effect of medicinal materials within a prescription and the coupling of syndromes. These are elements of the medicinal herb compatibility matrix. These are elements of the inter-entity symptom coupling factor matrix. This is the compatibility adjustment factor, with a value range of [0.1, 2.0]. This is a syndrome inhibition factor, with a value range of [0.05, 1.0], which reduces the weight of highly coupled but drug-property mismatched factors. The compatibility strength between polymerized medicinal materials, and through compatibility regulating factors Adjust contribution level Introducing the square term of the syndrome coupling factor, and through the syndrome inhibition factor Control its inhibitory effect to avoid excessive weighting caused by mismatch in drug properties; achieve a dual balance between the synergistic effect of medicinal materials and the rationality of syndrome differentiation, so that the overall weight of the prescription is closer to the theory and practical application of traditional Chinese medicine compatibility.

[0087] By fusing the weights of formula combination rules with multidimensional medicinal property characteristics and inter-entity syndrome coupling factors, a multi-source knowledge representation is constructed, generating a comprehensive knowledge vector for each entity across the formula and syndrome dimensions. Specifically, the weights of formula combination rules are combined with medicinal material characteristics and syndrome coupling factors to generate a multi-source knowledge representation for each entity. Higher-order nonlinear effects are introduced by squaring the syndrome coupling factor term to enhance the entity's knowledge representation ability across different dimensions. The calculation formula for the multi-source knowledge representation is as follows: ,in, For entities Multi-source knowledge representation, integrating the weights of prescription combination rules, symbolic entity features, and syndrome coupling factors. For containing entities A collection of prescriptions, As the weight of the formula combination principle, For entities Multidimensional pharmacological characteristics As a symptom coupling factor between entities, multi-source knowledge representation realizes a high-order fusion of drug properties, prescription rules and symptom logic, enabling entities to have multi-dimensional knowledge support in querying and reasoning, and ensuring the integrity and accuracy of the system's knowledge expression.

[0088] The present invention is further configured such that the link tension module includes:

[0089] Construct a multi-hop inference path set for each entity, containing all reasonable paths from that entity to other candidate entities, with path length limited by a preset maximum number of hops; specifically, for each symbolized entity... Construct a set of multi-hop inference paths from this entity to other candidate entities. The path length is affected by the maximum number of hops. To constrain inference depth and suppress combinatorial explosion, the formula is constructed as follows: ,in, For entities The set of multi-hop inference paths, including those from... Departure to any candidate entity The path, From arrive The length is Path sequence, The maximum number of hops, with a value in the range [1, 5]. For intermediate entities in a path, a set of paths Generated using a graph traversal algorithm, it is recommended to use a limited width-first approach or heuristic pruning to control the size;

[0090] For each multi-hop inference path, a pathogenesis link tension factor is generated, integrating the multi-source knowledge representations of each entity in the path and the symptom coupling factors between entities; specifically, for the path set... Calculate the pathogenesis link tension factor for each path. This factor integrates the multi-source knowledge strength of continuous nodes along the path and the symptom coupling between nodes, and applies higher-order suppression to the knowledge differences between nodes. Finally, it aggregates the contributions of each segment along the path in a multiplicative structure, reflecting the overall coherence and strength of the path. The formula for calculating the pathogenesis link tension factor is as follows: ,in, and , , For path The pathogenesis linkage tension factor was used to quantify the overall rationality and potential intensity of the pathway. For multi-source knowledge representation, corresponding to continuous entities in the path, As a symptom coupling factor between entities, This is a higher-order difference suppression term, which reduces the tension value of high-difference pathways. This represents a continuous entity index along a path. The product terms accumulate the contributions of each segment of the path, and the multiplicative structure ensures that even if any segment is extremely weak, it will reduce the overall tension of the path. The knowledge strength of adjacent nodes is coupled with the symptom coupling to form a segment-level contribution. This is a high-order difference suppressor; the greater the difference, the more the contribution of that segment is suppressed. A cubic approach is used to reflect strong nonlinear decay. Each contribution segment is wrapped in a 1+ form to prevent the product from reaching zero, while ensuring the contribution is positive and can be amplified or reduced. (Pathogenesis link tension factor) It is a positive real number; the larger the value, the more coherent the path is and the stronger the potential connection.

[0091] The path tension factor is fused with the combination rule weights of the formulas involved in each entity within the path to generate candidate entity scoring information. Specifically, the path tension factor is combined with the combination rule weights of the formulas involved in the path, and the contribution is accumulated on a path-by-path basis to obtain the score for each candidate entity. Candidate entity scores The scoring comprehensively considers the rationality of the path and the importance of the compatibility of the Chinese medicinal materials in the prescription along the path, reflecting the potential explanatory power or relevance of the entity to the starting entity through multiple paths. The candidate entity scoring formula is as follows: ,in, Candidate entities The candidate entity score is determined by the combined path tension and the weighting of the formula combination rule. For path tension factor, The combination rule weight for each entity in the path is the formula. The product of the weights of the Chinese medicinal herbs in the pathway reflects the synergistic importance of the herbs in the pathway. The contribution of each pathway is determined by the pathway tension. The weight of the combination rules of the prescriptions belonging to each node on the path The product of the weights within a path determines the overall importance of the medicinal herb compatibility chain within that path. The multiplicative structure emphasizes that the absence of any formula weight within a path will significantly reduce the path contribution. Summing all paths accumulates the explanatory power brought about by path diversity, allowing different paths to provide complementary information.

[0092] The present invention is further configured such that the query matching module includes:

[0093] The process maps user query conditions to the entity feature space to generate query feature representations. Specifically, the user's natural language query is decomposed into several entries in a predefined vocabulary, mapped to a vector representation isomorphic to the entity feature space using a symbolic function, and then a multidimensional representation of the query in the entity feature space is generated using a nonlinear aggregation operator. Let the set of query items be... Symbolic vector mapping is The weight of the query item is The dimension-by-dimensional definition query is represented as: , ,in, , This is the set of user query conditions, containing all the query terms entered by the user. For the first One query condition, The number of query conditions, with a value range of [1, 10]. For symbolic functions, natural language conditions The mapping is a symbolic representation with values ​​ranging from [0, 1]. It represents user queries in the entity feature space; through symbolic mapping, it maps user queries to the same feature space as the entities, enabling direct semantic comparison and complex matching.

[0094] A preliminary relevance score under difference constraints is constructed based on query feature representation and multi-source knowledge representation of candidate entities; specifically, based on query vectors... Multi-source knowledge representation vectors of candidate entities The original matching strength is calculated, and nonlinear suppression is applied using a complex difference function to obtain the preliminary correlation under difference constraints. The difference function is defined as the cumulative sum of higher-order ratios in each dimension: Preliminary correlation is defined as the combination of inner product and differential suppression: ,in, inner product Difference function captures the amount of direct match between the query and the entity across multiple dimensions. By employing the ratio of cubic differences in the numerator to the sum of squares in the denominator, the contribution of a dimension to the overall difference increases sharply when there is a significant difference between the two in a certain dimension. However, when the sum of the overall components is large, the denominator buffers the difference amplification effect, achieving robust control over extreme components. Ultimately, the impact of differences on matching is reflected in nonlinear suppression in the denominator form, thus... Achieving a balance between similarity and local differences For entities In the Multi-source knowledge components; nonlinear suppression of inner product similarity is implemented through a higher-order difference function, taking into account both local differences and overall similarity, thereby improving the robustness of matching judgment;

[0095] The preliminary correlation is adjusted by combining the inter-entity syndrome coupling factor and the weighting of the prescription combination rule to generate an intermediate correlation; specifically, the inter-entity syndrome coupling factor is used as the basis for the adjustment. The combination rule weight of the formula containing entities in the path or neighborhood The combined effects on the initial correlation score are characterized by a multiplicative agglomerator to depict the nonlinear amplification or suppression effect of the syndrome-prescription dual constraint, generating an intermediate correlation score. Define the neighborhood set as A multiplicative aggregator is used: Intermediate relevance is defined as: Among them, the index Controlling multiplicative sensitivity, Indicates and A set of neighborhood entities with significant symptom coupling, which acts as a sample set for local coupling effects; It is a syndrome-prescription multiplicative aggregation factor, which acts as a nonlinear amplification or inhibition factor on the initial correlation. The intermediate relevance is used as the fusion score before the introduction of path tension. The minimum threshold for determining symptom coupling is defined, with a value ranging from [0.1, 1.0]. The sensitivity is the multiplicative power sensitivity, with a value range of [0.5, 3.0]. As for the weight of the prescription combination rule, if the neighborhood size is too large, it will affect the weight of the prescription combination rule. Implement a size cap or only retain the previous size. The largest The term is used to control the amount of computation, multiplicative aggregator By analyzing each entity within the neighborhood The method of exponentiation followed by product ensures that a significant increase in any term will significantly amplify the overall adjustment factor, reflecting the amplification effect of strong local coupling. The subtraction operation ensures that when there is no significant coupling in the neighborhood... Near zero, avoid to Meaningless amplification; intermediate correlation The query matching basis is determined by multiplication. The combined influence of syndrome-prescription coupling and integration is quantified into a holistic measure.

[0096] Based on the relevant path set of query conditions, a pathogenesis link tension factor is introduced to comprehensively correct the intermediate relevance through multiple paths, generating a query matching relevance score. Specifically, the relevant path set of query conditions is... Pathogenesis Link Tension Factors of Each Pathway Accumulated into a comprehensive metric at the path level, and correlated with intermediate correlation. Non-linear coupling to generate a final query matching relevance score for candidate entities. Define path-level cumulative quantities: The final relevance is defined as: ,in, To query To the entity The set of related multi-hop paths; The path tension factor is a measure of path rationality. , These represent the additive and multiplicative aggregation quantities of path tension, used to reflect the overall strength and cumulative diversity of the path, respectively. The final query match relevance score is used as one of the inputs for candidate ranking. The path tension power controls the nonlinear effect of a single path in the product term, with a value range of [0.5, 2.0]. The cumulative amplification factor for the path is [0.5, 3.0]. Let be the path multiplicative suppression power, with a value range of [0.5, 3.0]. When the value is too large, the path is first considered when determining the value. Sort and truncate To limit the computational load, the power parameter... Control path accumulation amplification and suppression behavior. Cumulative path tension reflects the overall contribution of the number and intensity of paths, using a power-law approach. Amplify the effect of strong paths. For path tension multiplicative accumulation, after exponential... and the power of the denominator After processing Implement nonlinear suppression to prevent a large number of medium-intensity paths from causing inflated scores. Achieving a balance between "path amplification and diversity suppression": when a few strong paths exist Significant, small denominator; when a large number of intermediate paths exist. Increase, inhibit excessive accumulation, ultimately The matching of candidate entities in this query is based on a combination of semantic matching, syndrome-prescription coupling, and the rationality of the pathogenesis link.

[0097] The present invention is further configured such that the joint scoring module includes:

[0098] Construct a multi-hop reasoning path for each candidate entity, and aggregate the formula combination rule weights and pathogenesis link tension factors of each entity in the path to generate the path tension aggregation quantity; specifically, along each candidate entity... Associative multi-hop inference path set Path-level information is aggregated to generate path tension aggregation. Path aggregation simultaneously incorporates the formula combination rule weights of each node within the path and the pathogenesis link tension of the path. A nonlinear expression combining multiplicative amplification and path length suppression is used to quantify the comprehensive strength of entities manifested in the graph through multiple inference links. The formula for calculating path tension polymerization is: ,in, For entities This is a set of multi-hop inference paths for endpoints, serving as a path-level information source. For nodes Prescription The combination rule weights serve as a local measure reflecting the contribution of prescription compatibility to the pathway. This is the pathway tension factor, which measures the rationality and intensity of the pathogenesis of a pathway. This refers to the path length, used for path length suppression. For the calculation of prescription weights, use non-negative real numbers, with an initial field of . , This is the power exponent of the formula weights. The path tension power, and the formula weight within the path. Power of power After transformation and addition, the product is taken. The multiplicative structure significantly amplifies the contribution of any single ingredient along the path when its weight is large, reflecting the cumulative effect of the compatibility chain and path tension. Power of power To enhance its differential characterization, paths with stronger tension account for a larger proportion in aggregation, and the path length term... Applying secondary suppression to long paths in the denominator to avoid numerical inflation of excessively long paths also reflects the intuition that long chains reduce reliability in clinical reasoning. Summing all paths preserves the complementary information brought about by path diversity, yielding the entity-level aggregation quantity. Path tension polymerization amount By jointly quantifying the compatibility of prescriptions and the rationality of the pathway, the pathway contribution of an entity reflects both the importance of compatibility and is subject to tension constraints, thereby improving the explanatory power of candidate entities for disease or syndrome pathways.

[0099] A neighborhood consistency multiplicative penalty is generated based on the differences in multi-source knowledge representations between candidate entities and their neighboring entities and the symptom coupling factor; specifically, it is based on candidate entities. Its significant coupling neighborhood Entities in The differences in multi-source knowledge representations and the strength of symptom coupling are taken as inputs, and multiplicative penalty quantities are applied. The effect of amplifying local inconsistencies on joint scoring is to prevent entities with significant inconsistencies in the neighborhood context from being incorrectly prioritized. The calculation formula is as follows: ,in, Representation and entity There exists a set of neighboring entities with significant symptom coupling. For multi-source knowledge representation The third difference term is used to emphasize the higher-order effects representing the differences. Using the intensity of syndrome coupling as an amplification factor for the difference effect, the index... The sensitivity of the coupling to the penalty can be adjusted; the multiplicative structure ensures that large differences in a single neighborhood can significantly amplify the penalty value, avoiding maskable inconsistencies caused by simple linear accumulation; ultimately... This is a multiplicative penalty for local consistency, used to reduce the priority of candidates that are inconsistent with their neighborhood in the joint scoring.

[0100] The query matching relevance score, multi-source knowledge representation, path tension aggregation, and neighborhood consistency multiplicative penalty are fused to generate a joint score for candidate entities; specifically, the query matching relevance score is used as the basis for the joint score. Multi-source knowledge representation Path tension aggregation amount Neighborhood penalty As input, a hybrid power-multiplication-division structure is used to generate a joint score. This structure allows for the use of nonlinear coupling to amplify the complementary effects of query semantics, knowledge depth, and path rationality, while simultaneously suppressing local inconsistencies through nonlinear penalty terms. The joint scoring formula is as follows: ,in, A query relevance score is assigned to candidate entities, representing the strength of the semantic match between the query and the entity. The strength of multi-source knowledge representation of an entity reflects its aggregate value across multiple sources of information, such as prescriptions, syndromes, and medicinal properties. This represents the path tension aggregation level, reflecting the link support based on the inference path. This is a neighborhood consistency penalty, reflecting local inconsistency. For joint scoring, as a basis for ranking candidate entities, To calibrate the query effect by setting the relevance exponent, the value range is [0.5, 3.0]. To represent exponentiation in knowledge, Path aggregation exponentiation To penalize the exponentiation, the value range is [0.5, 3.0], and the first term... The relevance of query matching can be enhanced or smoothed to ensure that the importance of query semantics is adjustable. By introducing a multiplicative coupling between knowledge depth and path strength, the multiplicative nature of which promotes the complementary amplification of the two, and by adding a term to ensure that... or The baseline value is still retained even when it is zero. By controlling the penalty intensity with a power factor and using a multiplicative inhibition mechanism to make local inconsistencies decay non-linearly in the scoring, this hybrid structure enables the joint scoring to both improve entities with strong knowledge support and path support and query relevance, and to effectively suppress inconsistencies in the neighborhood.

[0101] Candidate entities are ranked based on their joint scores, and a preliminary candidate set is generated according to preset truncation rules. Specifically, this is based on the joint scores of all candidate entities. Sort in descending order and truncate according to the preset truncation strategy (first half). Item or threshold Generate a preliminary candidate set This transforms the joint score into a deliverable candidate set, ensuring that subsequent optimization steps are run on a smaller, higher-quality candidate set, thereby improving overall efficiency and result quality.

[0102] The present invention is further configured such that the processing of the preliminary candidate set includes:

[0103] By integrating multi-source knowledge representations with neighborhood symptom coupling factors, an attribute matrix for candidate entities is constructed. Specifically, the multi-source knowledge representation of each preliminary candidate entity is concatenated column-wise with the symptom coupling factors of its salient neighborhood to form an attribute matrix that simultaneously reflects the entity's own knowledge and neighborhood coupling characteristics. The calculation formula is: , where the matrix Contains solids in the horizontal direction Multi-source knowledge representation vector and several coupling factors arranged in neighborhood order This enables the co-presentation of entity-neighborhood information, and the neighborhood set. Based on the selection of the symptom coupling threshold, only neighbors that are significantly associated with the entity are included. For entities The multi-source knowledge representation vector reflects the composite information of prescription association, drug properties, and syndrome coupling. As a symptom coupling factor between entities, quantifying entities and Coupling strength in the symptom dimension For entities The set of saliently coupled neighborhoods is determined by filtering based on a threshold. For the attribute matrix, construct an attribute matrix that takes into account both the internal knowledge of the entity and the coupling of the neighborhood, so as to provide a complete and divisible data foundation for subsequent efficacy coverage calculation and redundancy identification.

[0104] Efficacy coverage index is generated based on the differences in the physical properties of the medicine and the coupling factors of syndromes; specifically, it is based on the attribute matrix. The nonlinear coverage contribution is calculated for each dimension of the multi-source knowledge representation, and the cumulative effect of neighborhood symptom coupling is used as an amplification factor. At the same time, the repetitive contribution is suppressed by multiplicatively using the differences between entities in the neighborhood in this dimension, and finally, an entity-level efficacy coverage index is obtained. The calculation formula is: ,in, For entities In the Multi-source knowledge components of dimensions. The number of dimensions in the multi-source knowledge representation. For entities The efficacy coverage index measures an entity's ability to cover different syndrome / efficacy dimensions. Power-laws for the entity dimension control the baseline contribution sensitivity. To demonstrate the coupling power, the degree of amplification of neighborhood coupling is controlled. For neighborhood differences, the power of the difference is used to control the difference suppression sensitivity, and each dimension... The baseline contribution is Instead of simple linear addition, power-law amplification is used to amplify the effects of higher-value dimensions, leading to the accumulation of neighborhood coupling within the molecule. Using the overall symptom strength of the neighborhood as a positive amplification factor enhances the dimensional weights supported by the neighborhood, and the denominator adopts a multiplicative term. By strongly suppressing dimensions with high similarity in the neighborhood, the multiplicative structure ensures that significant differences in any neighborhood are not masked by the linear average, and the overall efficacy coverage is obtained by summing all dimensions. This indicator takes into account the strength of the entity itself, neighborhood support, and redundancy suppression.

[0105] Redundancy assessment is performed by considering the differences in pharmacological properties and syndrome coupling between entities. A similarity matrix is ​​constructed among candidate entities to identify highly similar candidate entities. Specifically, a high-order nonlinear similarity function matrix is ​​constructed based on the differences between candidate entities in various dimensions of multi-source knowledge and their syndrome coupling strength. To measure redundancy, the function matrix uses high powers for significant feature differences to enhance discriminative power and suppresses absolute value inflation with a candidate population scale term. The calculation formula is as follows: ,in, Candidate entities and The similarity function matrix, To demonstrate the power of the symptom, we need to control the amplification effect of coupling on similarity. To control the suppression power of the scale, the numerator uses a fourth power to control the strength of the suppression of the overall characteristic scale by the denominator. Strengthening the identification of significant feature differences makes it less likely that smaller differences will be misjudged as high redundancy under this metric, multiplied by the syndrome coupling term. Using semantic or clinical coupling as amplification or reduction factor for difference discrimination, if Larger differences are more noteworthy, with the denominator being the characteristic scale suppression term. To prevent the overall eigenvalues ​​from being much larger than the differences, thus distorting the numerator / score, the obtained... It is a non-negative quantity, and the larger the value, the more significant the difference and coupling between the two entities under the joint measure of feature differences.

[0106] The present invention is further configured such that generating the optimization candidate set includes:

[0107] The initial candidate set is filtered based on a redundancy threshold to generate a de-redundancy candidate set; specifically, a preset or verified similarity threshold is used. Perform pairwise judgment on the candidate set: an entity is considered only if it is paired with any other entity in the set. Only values ​​below a threshold are retained, thus generating a deduplicated candidate set. Decision formula and set definition: , To eliminate redundant candidate sets, a pairwise decision is used to ensure that the retained entities maintain relatively low redundancy in the global candidate set. The selection of the threshold should balance coverage and redundancy suppression: a low threshold tends to retain distinct entities, while a high threshold allows more similar entities to enter the de-redundancy set. In practice, a stepwise greedy strategy can be adopted: iteratively select and eliminate candidates similar to the already selected entities according to the joint score to avoid a perfect match. Complexity;

[0108] The redundant candidate set is then optimized by weighting based on efficacy coverage to generate optimized candidate efficacy scores; specifically, , where factor Indicates neighboring entities and The complementarity is such that when the similarity is low, this term approaches 1, and vice versa. The complementary contribution is nonlinearly amplified or attenuated to ensure that the coverage contribution of highly similar entities is not repeatedly accumulated. The neighborhood complementarity is summed and then incremented by one before being combined with the summation. Multiplication improves the scores of entities with strong complementarity, thereby maximizing coverage across the overall candidate set. To optimize the candidate efficacy coverage score, which serves as the ranking criterion, The neighborhood nonlinear enhancement index is used to adjust the nonlinear sensitivity of complementary contributions. By using neighborhood complementarity weighting, the efficacy coverage of the redundancy removal candidates is further optimized, and entities that are complementary in coverage and make significant contributions are preferentially retained, thereby improving the efficacy breadth and representativeness of the final candidate set.

[0109] Based on the optimized candidate efficacy scores, candidate entities are sorted, a predetermined number of candidate entities are extracted, and a final optimized candidate set is generated. Specifically, this is based on the optimized candidate efficacy scores. For the redundancy removal set Sort the data and extract the desired number of segments. Before selection The names form the final optimization candidate set. This is for use in the subsequent final sorting and result output stage. The calculation formula is: , To ultimately optimize the candidate set, the generated optimized candidate set has excluded highly redundant entities while ensuring the breadth of efficacy coverage.

[0110] The present invention is further configured such that the final sorting module includes:

[0111] A comprehensive score for candidate entities is constructed by combining joint scores, optimized candidate efficacy coverage, neighborhood multi-source knowledge representation, and syndrome coupling factors, and serves as the final ranking criterion. Specifically, a comprehensive score is constructed based on the joint scores of candidate entities, optimized efficacy coverage, neighborhood multi-source knowledge differences, and neighborhood syndrome coupling through exponential amplification and neighborhood multiplicative aggregation. This score numerically amplifies both query relevance and efficacy coverage, and further enhances the priority of the best candidate through neighborhood-level complementarity or difference, thus serving as the final ranking criterion. The comprehensive score calculation formula for candidate entities is as follows: ,in, Candidate entities The joint score characterizes the combined support strength of the query and the knowledge network. Candidate entities The optimized efficacy coverage score characterizes the breadth and complementarity of efficacy coverage. Candidate entities Multi-source knowledge represents scalar or aggregated strength values. As a symptom coupling factor between entities, Candidate entities The neighborhood set is determined according to the syndrome coupling threshold. The candidate entities are comprehensively scored, which serves as the final sorting key. For joint scoring index, control Nonlinear amplification, For efficacy coverage index, control The nonlinear effects, For knowledge difference power, control Sensitivity to neighborhood product For coupling powers, control syndrome coupling Weights in neighborhood amplification Power Moderated Joint Score The strength of the influence of query matching and joint evidence allows for non-linear amplification or smoothing in the scoring. Power Strengthen and optimize treatment coverage The contribution of this factor biases towards entities that both match the query and cover a wide range of therapeutic dimensions, with neighborhood multiplicative terms. Each neighborhood entity right The complementarity or difference of each term is added to the product: the absolute difference within each term. powers Amplifying the differences in multi-source knowledge representation, coefficients Treating the intensity of symptom coupling as an amplifier of the effect of difference, the multiplicative structure ensures that significant complementarity exists in any neighborhood (large difference and significant symptom coupling), thus enhancing... Conversely, when the neighborhood is highly consistent, the product approaches 1, avoiding repeated accumulation. This hybrid structure takes into account query relevance, efficacy coverage and neighborhood complementarity, and evaluates candidate entities from multiple dimensions.

[0112] The candidate set for optimization is sorted in descending order based on the comprehensive score of the candidate entities, generating a sorted candidate set; specifically, based on the comprehensive score... Using a unique sorting key, optimize the candidate set. The entities in the list are sorted in descending order to generate a sorted candidate set. The calculation formula is: Sort by The numerical values ​​are sorted from high to low to ensure that entities with high comprehensive scores appear first; multi-source information is compared at the entity level using a unified metric by sorting by comprehensive score, generating a priority sequence that takes into account semantics, efficacy and link support, which facilitates the generation of clinically and search-friendly result sets;

[0113] Based on a preset query count, a preset number of candidate entities are extracted from the sorted set to generate the final query result set. Specifically, based on the preset number of results... From the sorted set Before the middle cut Each entity generates the final query result set. The calculation formula is: By capturing high-quality front Each entity ensures that the output results take into account efficacy coverage, query matching, and link rationality, providing an interpretable and verifiable set of candidates for user or clinical decision-making.

[0114] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A traditional Chinese medicine service query system based on a knowledge graph, characterized in that, Comprise: Data acquisition and symbolization module, in the collection of medical related data, the collected data processing generates symbolization coding entity, extraction of multi-dimensional drug property characteristics, generating symbolization entity set; Syndrome coupling module, according to the multi-dimensional drug property characteristics, calculate the coupling factor between entities, establish the semantic relationship between syndrome and medicinal materials, generate the preliminary triad relationship set; Prescription fusion module, analyze the compatibility relationship between medicinal materials in prescription, generate the weight of prescription combination rule, fusion generate multi-source knowledge representation; Link tension module, based on multi-source knowledge representation generates pathogenic link tension factor, scores the rationality of multi-hop reasoning path between entities, generates candidate entity score information; Query matching module, the user query condition is mapped to the entity feature space, combined with the weight of prescription combination rule, syndrome coupling factor and pathogenic link tension factor, generate query matching correlation degree score; Joint scoring module, according to the query matching correlation degree score, multi-source knowledge representation and pathogenic link tension factor, calculate the joint score, sort the candidate entity to generate the preliminary candidate set; Efficacy redundancy module, processing the preliminary candidate set, through the analysis of the difference of drug property characteristics and syndrome coupling factor of candidate entity, generate the optimized candidate set; Final sorting module, according to the joint score and efficacy coverage, comprehensive sorting the optimized candidate set, generate the final query result set. 2.The knowledge graph-based traditional Chinese medicine service query system according to claim 1, characterized in that, Data acquisition and symbolization module includes: The medicinal material database and prescription information are cleaned, de-duplicated and field standardized to generate a preliminary entity set; Each entity in the preliminary entity set is converted into a symbolized coded entity through mapping; The multi-dimensional drug property characteristics of each symbolized coded entity are calculated; The symbolized coded entity and multi-dimensional drug property characteristics are integrated to generate a complete symbolized entity set. 3.The knowledge graph-based traditional Chinese medicine service query system according to claim 2, characterized in that, Syndrome coupling module includes: According to the multi-dimensional drug property characteristics, calculate the coupling factor between entities; Based on the syndrome coupling factor, establish the semantic relationship between entities in each syndrome dimension, generate semantic relationship score combined with syndrome weight and attribute difference adjustment factor; Screen and consistency check the semantic relationship score, screen out the relationship between entities below the threshold, generate screened and checked semantic relationship; The screened and checked semantic relationship is converted into a preliminary triad relationship set, which is composed of entities, relationships and corresponding syndrome dimensions.

4. The knowledge graph-based traditional Chinese medicine service query system according to claim 1, characterized in that, Prescription fusion module includes: Calculate the compatibility relationship strength of medicinal material combination in each prescription; Based on the compatibility relationship of medicinal materials and the syndrome coupling factor between entities, generate the weight of prescription combination rule; Fuse the weight of prescription combination rule, multi-dimensional drug property characteristics and syndrome coupling factor between entities, construct multi-source knowledge representation, generate comprehensive knowledge vector of each entity in the dimensions of prescription and syndrome. 5.The knowledge graph-based traditional Chinese medicine service query system according to claim 4, characterized in that, Link tension module includes: Construct a multi-hop reasoning path set for each entity, including all reasonable paths from the entity to other candidate entities, the path length is limited by the preset maximum number of hops; Generate pathogenic link tension factor for each multi-hop reasoning path, integrate the multi-source knowledge representation of each entity in the path and the syndrome coupling factor between entities; Fuse the path tension factor and the combination rule weight of the prescription of each entity in the path to generate candidate entity score information. 6.The knowledge graph-based traditional Chinese medicine service query system according to claim 1, characterized in that, Query matching module includes: mapping the user query condition to an entity feature space to generate a query feature representation; constructing a preliminary relevance degree under a difference constraint based on the query feature representation and a multi-source knowledge representation of the candidate entity; adjusting the preliminary relevance degree in combination with an inter-entity syndrome coupling factor and a prescription combination rule weight to generate an intermediate relevance degree; introducing a path tension factor of pathogenesis link according to a query condition related path set to perform multi-path comprehensive correction on the intermediate relevance degree to generate a query matching relevance degree score. 7.The knowledge graph-based traditional Chinese medicine service query system according to claim 1, characterized in that, The joint scoring module includes: constructing a multi-hop reasoning path of each candidate entity, aggregating the prescription combination rule weight and the path tension factor of pathogenesis link of each entity in the path to generate a path tension aggregation amount; generating a neighborhood consistency multiplicative penalty amount based on the difference between the multi-source knowledge representation of the candidate entity and its neighborhood entity and the syndrome coupling factor; fusing the query matching relevance degree score, the multi-source knowledge representation, the path tension aggregation amount, and the neighborhood consistency multiplicative penalty amount to generate a joint score of the candidate entity; sorting the candidate entities according to the joint score and generating a preliminary candidate set according to a preset pruning rule. 8.The knowledge graph-based traditional Chinese medicine service query system according to claim 1, characterized in that, The processing of the preliminary candidate set includes: integrating the multi-source knowledge representation and the neighborhood syndrome coupling factor to construct an attribute matrix of the candidate entity; generating an efficacy coverage index according to the difference between the entity drug property feature and the syndrome coupling factor; performing redundancy evaluation through the difference between the inter-entity drug property feature and the syndrome coupling factor to construct a similarity matrix between the candidate entities and identify highly similar candidate entities. 9.The knowledge graph-based traditional Chinese medicine service query system according to claim 8, characterized in that, Generating an optimized candidate set includes: screening the preliminary candidate set according to a redundancy determination threshold to generate a de-redundancy candidate set; optimizing the efficacy coverage weight of the de-redundancy candidate set to generate an optimized candidate efficacy score; sorting according to the optimized candidate efficacy score, and intercepting a preset number of candidate entities to generate a final optimized candidate set. 10.The knowledge graph-based traditional Chinese medicine service query system according to claim 1, characterized in that, The final sorting module includes: constructing a comprehensive score of the candidate entity in combination with the joint score, the optimized candidate efficacy coverage, the neighborhood multi-source knowledge representation, and the syndrome coupling factor as the final sorting basis; sorting the optimized candidate set in descending order according to the comprehensive score of the candidate entity to generate a sorted candidate set; intercepting a preset number of candidate entities in the sorted set according to a preset query number to generate a final query result set.