A traditional chinese medicine ancient and modern semantic alignment system based on multi-dimensional feature fusion

CN122842973APending Publication Date: 2026-09-29INSTITUTE OF CHINESE MATERIA MEDICA CHINA ACADEMY OF CHINESE MEDICAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611281160.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-24
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

这种常规方法在处理中医古今术语时面临着缺陷:第一,古今行文习惯差异巨大,单一的序列匹配容易产生语义漂移与匹配失效,导致古籍实体与现代标准实体无法准确对齐;第二,现有聚类算法在进行术语归一化时,缺乏中医先验知识的强逻辑约束,容易将字面高度相似但含义截然不同的实体(如“心气虚”与“肺气虚”)错误合并;第三,当这些未经精准归一化的异构知识库直接对接前端真实临床病历时,由于临床文本夹杂大量否定词、修饰语及西医干扰项,传统的实体匹配机制往往面临跨域误报,并在位置重叠的实体提取中产生严重的匹配冲突与语义丢失风险

Benefits of technology

第一,本发明首创四维特征融合语义相似度计算体系,从结构、短语、要素、深层语义全方位量化术语相似度,彻底突破传统单一字符串算法的局限性,完美适配中医古今术语语序灵活、修饰冗余、异形同义的表达特征,有效解决古今语义鸿沟、匹配失效、语义漂移等核心问题,术语对齐精度远高于传统通用算法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842973A_ABST
    Figure CN122842973A_ABST
Patent Text Reader

Abstract

This invention provides a semantic alignment system for ancient and modern Chinese medicine based on multi-dimensional feature fusion, belonging to the fields of natural language processing, knowledge graph, and traditional Chinese medicine informatics. It includes a knowledge base construction, a multi-dimensional fusion similarity calculation model construction, a controlled thesaurus generation, complex negation logic parsing and semantic denoising for clinical texts, a three-level cascade matching and resolution output module, and the following: A multi-source knowledge graph and a standard terminology lexicon of ancient and modern Chinese medicine are constructed. Through the multi-dimensional fusion similarity calculation model, edit distance, core phrase retention, deduplication character overlap, and core stem semantic similarity are comprehensively considered, while incorporating organ compatibility constraints. A disjoint-set data structure algorithm is used to generate synonym clusters and align them with national standard terms. Complex negation logic parsing and semantic denoising are performed on clinical medical record texts. The text is mapped to the knowledge base language through a three-level cascade matching architecture. Overlapping candidate entities are resolved based on a utility function, and a normalized entity sequence is output, achieving highly accurate alignment of ancient and modern terms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of natural language processing, knowledge graph and traditional Chinese medicine informatics, and specifically refers to a semantic alignment system for ancient and modern Chinese medicine based on multi-dimensional feature fusion. Background Technology

[0002] Classical Chinese medicine texts contain rich wisdom on syndrome differentiation and treatment, serving as a crucial source of knowledge for constructing intelligent TCM diagnostic and treatment systems. However, there are significant temporal and individual differences in the expression methods of ancient TCM texts, modern textbooks, and real-world clinical cases. Ancient texts are mostly highly condensed, unstructured natural language descriptions, resulting in a vast number of synonyms and variant names between them and modern national standard TCM terminology.

[0003] Currently, existing technologies for entity alignment and clinical semantic transformation in the field of Traditional Chinese Medicine (TCM) largely rely on single string matching algorithms or simple text similarity calculations. This conventional approach faces several shortcomings when dealing with ancient and modern TCM terminology: First, the significant differences in writing styles between ancient and modern times make single sequence matching prone to semantic drift and matching failures, leading to inaccurate alignment between ancient entities and modern standard entities. Second, existing clustering algorithms, when normalizing terms, lack strong logical constraints from TCM prior knowledge, easily merging entities with highly similar literal meanings but drastically different connotations (such as "heart qi deficiency" and "lung qi deficiency"). Third, when these heterogeneous knowledge bases, lacking precise normalization, are directly integrated with real clinical medical records, traditional entity matching mechanisms often face cross-domain false alarms due to the presence of numerous negative words, modifiers, and Western medical interference in the clinical text. Furthermore, they generate severe matching conflicts and semantic loss risks in extracting entities with overlapping positions. The fragmentation of underlying ancient and modern knowledge and the failure of clinical semantic alignment directly result in a lack of reliable, objective prior fact support for the upper-level reasoning model. Summary of the Invention

[0004] To address the technical problems existing in the prior art, this invention provides a semantic alignment system for ancient and modern Chinese medicine based on multi-dimensional feature fusion. This system includes: The knowledge base construction module is used to build a multi-source knowledge graph and a standard terminology glossary for TCM diagnosis and treatment; The multidimensional fusion similarity calculation model construction module is used to build a multidimensional fusion similarity calculation model. It comprehensively calculates the edit distance based on the longest continuous matching block of the text, the core phrase retention degree based on the longest common subsequence, the overlap of deduplicated characters, and the semantic similarity of the core stem after noise removal based on the TCM-specific modifier word list. It also introduces the organ compatibility constraint to hard filter and block the risk of erroneous normalization of ectopic pathogenesis entities. While preserving the core medical semantics, it bridges the gap between ancient and modern expression methods and achieves highly accurate alignment of ancient and modern terms. The controlled thesaurus generation module is used to generate thesaurus clusters by combining high-confidence nodes obtained through multi-dimensional fusion similarity calculation with the disjoint-set data structure algorithm, and to align national standard terms to generate a controlled thesaurus that guides the alignment of ancient and modern semantics. The complex negation logic parsing and semantic noise reduction module for clinical text is used to perform context delimitation based on semantic tags to prevent cross-domain false alarms on unstructured clinical medical record text, and to transform it into structured data blocks. It constructs a deep negation word detection algorithm to accurately handle pre-negation, post-negation, double exclusion and specific exception words, and filters out various semantic noise through a built-in semantic noise filter. The three-level cascaded matching module is used to map the clinical case text processed by the complex negation logic parsing and semantic noise reduction module for clinical text to the knowledge base language inside the model based on the three-level cascaded matching architecture of "precise mapping-similarity retrieval-generative alignment". The resolution output module is used to construct a utility function based on entity baseline confidence, covered text length and semantic integrity when there are matching candidate entities with overlapping positions. It follows the principle of maximum utility to perform resolution and outputs a normalized entity sequence.

[0005] The beneficial effects of the technical solution provided by this invention include at least the following: First, this invention pioneers a four-dimensional feature fusion semantic similarity calculation system, which quantifies term similarity from the perspectives of structure, phrase, element, and deep semantics. It completely breaks through the limitations of traditional single-string algorithms, perfectly adapts to the flexible word order, redundant modification, and heteronymous expression characteristics of ancient and modern Chinese medicine terms, and effectively solves core problems such as the semantic gap between ancient and modern times, matching failure, and semantic drift. The term alignment accuracy is far higher than that of traditional general algorithms.

[0006] Second, this invention innovatively introduces a priori constraint mechanism for the compatibility of organs in traditional Chinese medicine, embedding professional medical knowledge into the algorithm's underlying layer. This forcefully blocks erroneous clustering of entities with similar literal meanings but different organs and semantics, fundamentally solving the semantic confusion problem of traditional unsupervised clustering. This greatly improves the professionalism, accuracy, and stability of terminology normalization, and adapts to the core medical logic of TCM syndrome differentiation.

[0007] Third, this invention constructs a complex negation logic parsing and multi-layer noise filtering system adapted to the specific scenarios of traditional Chinese medicine, and specifically solves industry problems such as semantic reversal of negation in clinical medical records, cross-domain interference from Western medicine, and redundant invalid information, which greatly improves the purity and effectiveness of input text and enhances the robustness of entity recognition and semantic matching from the data source.

[0008] Fourth, this invention designs a three-level cascaded hierarchical matching architecture, which integrates the advantages of rule-based precise matching, data similarity retrieval, and large model generalization alignment. This ensures not only the ultra-high matching accuracy of common terms, but also the generalization ability of long-tail heterogeneous terms and obscure terms from ancient books, enabling full-scenario adaptation of multi-source heterogeneous texts and compatibility with various data sources such as ancient books, textbooks, and clinical medical records.

[0009] Fifth, this invention constructs a multi-dimensional utility maximization conflict resolution mechanism, which comprehensively considers multiple indicators such as confidence level, text length, and semantic integrity. It effectively solves the problem of entity overlap and matching conflict in clinical texts, avoids the defects of fragmented semantic loss, and outputs a unified, standardized, complete, and accurate standardized entity sequence. This provides solid and reliable underlying technical support for TCM intelligent diagnosis and treatment, the modernization of ancient book knowledge, and the development of TCM artificial intelligence models. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a block diagram of a semantic alignment system for ancient and modern Chinese medicine based on multi-dimensional feature fusion provided in an embodiment of the present invention; Figure 2 This is the overall architecture diagram of TCM underlying knowledge base construction and clinical semantic alignment provided in the embodiments of the present invention. Detailed Implementation

[0012] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0013] The specific implementation of this invention mainly relies on a computer system to execute the above-mentioned technical solution. The following is an example of coronary heart disease.

[0014] like Figure 1 and 2 As shown, this embodiment of the invention provides a semantic alignment system for ancient and modern Chinese medicine based on multi-dimensional feature fusion. The system includes: The knowledge base construction module is used to build a multi-source knowledge graph and a standard terminology glossary for TCM diagnosis and treatment; This module serves as the foundational data support module for the entire invention (underlying knowledge base construction). Its core purpose is to build a comprehensive, authoritative, compliant, and structurally unified TCM knowledge base specifically for coronary heart disease, addressing the problems of existing technologies such as single knowledge sources, fragmented knowledge between ancient and modern times, and non-standard terminology. This invention abandons the single data source construction model and adopts a multi-source fusion, human-machine collaboration, and national standard verification approach to ensure the comprehensiveness, authority, and professionalism of the knowledge.

[0015] At the multi-source knowledge acquisition level, targeted collection of heterogeneous data related to coronary heart disease in Traditional Chinese Medicine (TCM) from ancient and modern times was conducted, covering three core data sources: first, ancient classic medical texts, medical records and diagnostic discussions related to coronary heart disease by renowned physicians throughout history, covering ancient variant terminology and empirical expressions; second, modern TCM undergraduate textbooks and specialized monographs, covering modern commonly used academic terminology; and third, national and industry authoritative TCM guidelines and expert consensus on coronary heart disease, covering standardized clinical diagnostic expressions. The collected raw data underwent unified cleaning, deduplication, invalid text removal, entity extraction, relationship analysis, and attribute annotation. Finally, a structured TCM knowledge graph for coronary heart disease integrating ancient and modern knowledge was constructed, enabling the interconnection and interoperability of ancient empirical knowledge and modern academic knowledge.

[0016] At the level of standard terminology benchmark construction, 17 national standard documents were collected, including "Classification and Code of Basic Symptom Information in Traditional Chinese Medicine Clinical Practice," "Clinical Terminology in Traditional Chinese Medicine Part 3: Treatment Methods," "Clinical Terminology in Traditional Chinese Medicine Part 2: Syndromes," "Clinical Terminology in Traditional Chinese Medicine Part 1: Diseases," and "Classification and Code of Basic Symptom Information in Traditional Chinese Medicine Clinical Practice," covering all dimensions of TCM disease, syndrome, symptom, etiology, pathogenesis, treatment methods, and Chinese medicine terminology standards. Keyword entries were extracted in batches by machine, and professional manual verification was used to eliminate colloquial, non-standard, and redundant expressions, unifying term definitions, categories, and attributions to construct a comprehensive standard terminology thesaurus. This thesaurus serves as the sole authoritative benchmark for all subsequent semantic alignment and terminology normalization, ensuring that all variant terms can ultimately converge to the national standard system, achieving unified and standardized terminology.

[0017] The multidimensional fusion similarity calculation model construction module is used to build a multidimensional fusion similarity calculation model. It comprehensively calculates the edit distance based on the longest continuous matching block of the text, the core phrase retention degree based on the longest common subsequence, the overlap of deduplicated characters, and the semantic similarity of the core stem after noise removal based on the TCM-specific modifier word list. It also introduces the organ compatibility constraint to hard filter and block the risk of erroneous normalization of ectopic pathogenesis entities. While preserving the core medical semantics, it bridges the gap between ancient and modern expression methods and achieves highly accurate alignment of ancient and modern terms. To address the complex mapping relationship between ancient Chinese medicine texts and modern standard terminology, where literal differences exist but semantic cognition is present, simple string matching algorithms are ineffective, resulting in poor matching of ancient and modern terms. This invention employs a multi-dimensional feature fusion-based approach to calculate the similarity between ancient and modern terms. It comprehensively quantifies terminology similarity from four independent and complementary dimensions: text structure, core phrases, character elements, and deep semantics. This fully covers both the surface features and deep medical connotations of Chinese medicine terms, completely resolving the semantic drift problem inherent in single algorithms.

[0018] Optionally, define entity nodes corresponding to any two terms in the ancient and modern multi-source knowledge graph. The calculation formula for the multidimensional fusion similarity calculation model is as follows: (1) in, The weighting coefficients (can be fine-tuned according to the TCM scenario to adapt to the matching needs of different types of terms such as symptoms, syndromes, pathogenesis, and treatment methods, achieving multi-dimensional complementary and comprehensive accurate semantic calculation). This is a sequence structure similarity based on edit distance, used to capture the similarity of the overall structure; The similarity based on the longest common subsequence is used to extract the retention status of the core phrase. This is a similarity measure based on character overlap rate, used to calculate the degree of pure character overlap without considering character order. Based on the similarity of core semantics, it is used to eliminate domain noise and overcome the differences in writing habits between ancient and modern times.

[0019] Optionally, the The calculation method is as follows: A recursive matching algorithm based on the longest consecutive matching block of text is used to traverse and extract the first term. With the second term Calculate the total number of matched characters for all consecutive blocks of identical characters between them. Calculate the overall sequence structure similarity: in, and These represent the total character length of the two terminology texts, respectively.

[0020] This dimension quantifies the similarity of the overall character arrangement and word order of two terms, specifically adapting to scenarios where the word order of TCM terms is rearranged or where there are minor additions or subtractions of characters. Traditional edit distance algorithms only count the number of added, deleted, or modified characters, failing to identify the integrity of consecutive matching blocks. This invention optimizes the matching logic by employing a recursive matching mechanism for the longest consecutive matching block. It traverses all non-overlapping consecutive identical character blocks of the two terms, accumulates the total number of all valid matching characters, and obtains a similarity score through bidirectional length normalization. The core function of this dimension is to capture the overall structural similarity of terms, avoiding matching failures caused by minor additions or subtractions of modifiers or minor word order adjustments, thus adapting to the flexible word order expression characteristics of ancient Chinese texts.

[0021] Optionally, the The calculation method is as follows: Using dynamic programming, construct a structure of size... Given a two-dimensional state transition matrix dp, if during the iterative traversal, it satisfies... Then the state transition equation is dp[i][j] = dp[i-1][j-1] + 1; if this condition is not met, the current position count is interrupted and reset to zero; finally, the maximum value in the entire matrix is ​​obtained, which is the length of the longest common continuous substring. : .

[0022] This dimension is a core feature dimension specific to Traditional Chinese Medicine (TCM) terminology, with the core purpose of preserving the integrity of the core phrases of the terms. The medical connotation of TCM terms is entirely determined by the core phrases; modifiers do not alter the core diagnostic meaning. This dimension constructs a two-dimensional state matrix using a dynamic programming algorithm to accurately calculate the length of the longest common continuous substring of two terms, representing the degree of overlap and completeness of the core phrases. For example, for "heart qi deficiency" and "heart qi insufficiency," the longest common core phrase is the core part of "heart qi," accurately identifying the consistency of the core pathogenesis and effectively avoiding matching omissions caused by differences in the last character, ensuring that the core medical semantics are not lost.

[0023] Optionally, the The calculation method is as follows: First, let's discuss the first terminology. Second term Convert to an unordered and deduplicated set of characters. and Then, find the intersection of the two sets and calculate the number of elements in the intersection; finally, divide by the number of elements in the set with the smaller cardinality. ; This dimension eliminates the interference of word order and only counts the coincidence ratio of core character elements of the two terms, which is used for fallback verification of the consistency of the basic composition of terms. In some ancient Chinese medical books, the word order of terms is completely reversed but the semantics are consistent, and word order-based algorithms will determine that they have low similarity. This dimension, through character deduplication and set intersection statistics, ignores word order differences and accurately matches homologous terms with consistent elements. This dimension can effectively make up for the shortcomings of word order structure algorithms, adapts to the expression characteristics of flexible word order and no fixed format of terms in ancient books, and ensures that homologous heterogeneous terms are not missed in matching.

[0024] said the calculation method is: construct a traditional Chinese medicine (TCM)-specific modifier vocabulary (including TCM general synonym vocabulary, TCM word-formation vocabulary, viscera and position vocabulary, etc.), which covers common degree adverbs, temporal auxiliaries and function words in TCM clinical expressions (e.g., 'Cu', 'Bao', 'Cu', 'Tu Ran', 'Shen', 'Ju', 'Ji', etc.); using the said modifier vocabulary, perform text denoising and stemming on the original terms and respectively, remove all modifiers contained therein to obtain the corresponding first core stem and the second core stem ; for the extracted core stems and , apply the recursive matching algorithm based on the longest continuous matching block again to calculate the similarity score at the stem level, which is taken as the core semantic similarity.

[0025] This dimension is specially designed for the difference in modification redundancy between ancient and modern terms. Terms in ancient books often have degree and tense modifiers such as "Cu, Bao, Cu, Shen, Ju", while modern standard terms do not have such modifiers and only retain the core pathogenesis. Traditional algorithms will include modifiers in the calculation, resulting in distorted semantic scores. The present invention pre-constructs a TCM-specific modifier vocabulary, performs noise reduction and modifier removal on original terms, extracts pure pathogenesis stems, only calculates similarity for core medical stems, completely eliminates the interference of invalid modifiers, and focuses on the essential syndrome differentiation semantics.

[0026] optionally, said introducing viscera compatibility constraint specifically includes: for entity nodes corresponding to pathogenesis terms, in combination with the viscera and position vocabulary, only allow node pairs whose viscera sets are completely consistent or one of which is a general reference to be merged, and define a viscera feature extraction function to identify and extract the viscera positions contained in the current entity. Assuming that the viscera set of the extracted node is , and the viscera set of node is , the binary constraint function : The judgment logic is as follows: Its output can only be (Mergers are allowed) or (Merge Prohibited), where 1 indicates merge is allowed, and 0 indicates merge is prohibited. If the output is... The following two conditions must be met: Condition 1 In other words, if at least one of the two entities does not explicitly point to a specific organ, the set is empty. ; For example, "Qi deficiency" (without specific organs) and "heart Qi deficiency" (with the heart as the organ) do not conflict in terms of organs and can be considered together in subsequent analysis.

[0027] Condition 2 That is, the two entities contain completely identical internal organs; For example, "insufficient heart qi" and "deficient heart qi" both refer to the "heart" organ, so they can be combined.

[0028] Through logical multiplication This achieves a hard filter of ectopic pathogenesis entities (effectively blocking the risk of misclassification of terms such as "heart qi deficiency" and "lung qi deficiency" due to their high similarity in wording), among which This represents the four-dimensional basic similarity score calculated at the literal or textual level between two pathogenesis entities, combined with... The judgment logic at the algorithm's underlying level prevents the risk of incorrect normalization of entities from different organ parts.

[0029] The controlled thesaurus generation module is used to generate thesaurus clusters by combining high-confidence nodes obtained through multi-dimensional fusion similarity calculation with the disjoint-set data structure algorithm, and to align national standard terms to generate a controlled thesaurus that guides the alignment of ancient and modern semantics. Optionally, the controlled thesaurus generation module is specifically used for: Define the set of entity nodes in the graph as If node pair satisfy , To set a threshold, edges are established. ; If node and An edge was built between them. and An edge has also been established between them, so based on the logic of transitive closure, it can be automatically deduced that... and They also belong to the same set; By continuously searching for and merging nodes with connectivity, all words directly or indirectly connected by edges are eventually packaged into an independent synonym cluster. For any satisfy: in This represents the nodes in the cluster. Derived from the total set of entity nodes in the graph ; This is a core condition for becoming part of the same cluster. For path compression lookup operations, ensure that all nodes within the same cluster point to a unique root node. Any two nodes inside and ,go through After the search operation, the root nodes they point to must be exactly the same. This ensures from the bottom layer of the algorithm that no matter how these similar nodes are connected by edges in the graph, as long as they belong to the same cluster, they will eventually converge and point to a unique representative root node, thus ensuring that the words are strictly classified into the same category semantically. Subsequently, the central terms of each cluster are aligned with the compiled national standard terminology glossary to generate a controlled thesaurus: Determine each cluster Standard mapping terms Define the set of national standard terms as Alignment mapping function The definition is as follows: in This represents the alignment threshold, appearing in synonym clusters. With the National Standard Terminology Collection Piecewise functions for mapping alignment In this process, it plays a role in quality control.

[0030] The final generated controlled thesaurus (the controlled thesaurus is used for entity disambiguation and auxiliary clinical text recognition within the graph, and its content includes the correspondence of similar nodes within the graph, the correspondence between entities within the graph and national standard words, in a one-to-one correspondence format, without a central word, and includes symptoms, etiology, pathogenesis and treatment) serves as the subsequent semantic benchmark.

[0031] The complex negation logic parsing and semantic noise reduction module for clinical text is used to perform context delimitation based on semantic tags to prevent cross-domain false alarms on unstructured clinical medical record text, and to transform it into structured data blocks. It constructs a deep negation word detection algorithm to accurately handle pre-negation, post-negation, double exclusion and specific exception words, and filters out various semantic noise through a built-in semantic noise filter. Clinical medical record texts (referring to the input information of a single patient, which must include symptoms and may include etiology, pathogenesis, and treatment methods, but do not contain prescriptions; they can be understood as the information collected after a patient comes in, including the four diagnostic methods and analysis of pathogenesis and symptoms) not only contain core diagnostic and treatment information, but also contain a large number of non-symptom descriptions. Furthermore, the use of "negative words" in TCM symptom descriptions can easily lead to semantic reversal, which can cause the model to misidentify and affect the accuracy of inference. Therefore, the input natural language text is cleaned and preprocessed.

[0032] Optionally, the complex negation logic parsing and semantic noise reduction module for clinical text is specifically used for: (1) For the semi-structured features of some clinical case texts, the module first performs context delimitation based on semantic tags: Define clinical case text as Using regular expression patterns The scan is performed, dividing it into four independent semantic subspaces: in Represents the symptom subspace; Represents the pathogenic factor space; Represents the pathogenesis subspace; Represents the treatment method subspace; Analytical functions The input natural language case studies are accurately converted into structured data blocks: This mechanism ensures contextual constraints for entity extraction (e.g., preventing "phlegm-resolving" in the treatment section from being misidentified as "phlegm-turbidity" in the pathogenesis section), significantly reducing the cross-domain false alarm rate. (2) Construct a negative word detection algorithm to handle various negative word types such as pre-negation ("no heart pain"), post-negation ("thirst is not obvious"), and double exclusion ("not without discomfort"). For words such as "insomnia", "insensitive", and "loss of appetite" that contain negative words but are positive symptoms, define an exception word set. Perform an exception word check: For compound negative expressions such as "dry mouth but not thirsty" and "abdominal distension but not pain", we define a set of compound negative patterns. Using regular expression mode Perform clause splitting, extract the affirmative symptoms from the first half and add them to the matching queue, discard the negative descriptions from the second half: For a complete negation, define the set of negation prefixes and suffixes. Perform a negative judgment: in Indicates whether the affirmative portion can be extracted; For double exclusion negations, a logical exception rule is introduced to ensure that double negation expressions are not incorrectly cleaned up; Final negation handling strategy: (3) The module has a built-in semantic noise filter to clean the following three types of non-entity text: Normal state description: The script explicitly defines the set of ignored phrases. (Such as compound negative expressions like "dry mouth but no thirst" or "abdominal distension but no pain"), if If so, discard it directly; Interference from Western medicine and drugs: Use regular expressions to remove Western medicine diagnosis / examination text (such as "coronary heart disease", "ST segment depression", "50% stenosis") and Western medicine drug names ("aspirin", "metformin", etc.). Irrelevant modifier cleanup: Remove time adverbs that do not have substantial referential meaning ("3 years ago", "postoperative", "frequent", etc.), but retain degree adverbs with differential diagnostic value ("paroxysmal", "severe", etc.).

[0033] The three-level cascaded matching module is used to map the clinical case text processed by the complex negation logic parsing and semantic noise reduction module for clinical text to the knowledge base language inside the model based on the three-level cascaded matching architecture of "precise mapping-similarity retrieval-generative alignment". Optionally, the three-level cascaded matching module is specifically used for: The first level is based on a direct mapping between the controlled thesaurus and a general rule base in the field of Traditional Chinese Medicine: As the highest priority matching layer, this layer directly calls the controlled thesaurus and rule base built in the previous section to achieve rapid normalization of high-frequency and specific pattern terms; Second level, mixed similarity calculation: When the first-level matching fails, multi-dimensional feature fusion similarity calculation is initiated: for the terms to be matched... and candidate terms The same calculation method as formula (1) is used, and a knowledge base rule enhancement mechanism is introduced on this basis: Among them, the bonus items This includes bonuses for known synonyms (+0.50), standard words (+0.30), synonym replacement (+0.25), synonym patterns (+0.15), and core stemming matching (+0.10); penalty items. Includes a penalty for incompatibility with internal organs (-0.35); Only keep Candidate words, For the matching threshold (e.g., it can be set to...), ); Level 3, LLM-based generative semantic alignment: when When this is triggered, an LLM call is made, which provides the LLM with a candidate list and matching guidance by editing a specific type of Prompt template (solving the zero-sample matching problem for long-tail non-standard terms).

[0034] The resolution output module is used to construct a utility function based on entity baseline confidence, covered text length and semantic integrity when there are matching candidate entities with overlapping positions. It follows the principle of maximum utility to perform resolution and outputs a normalized entity sequence.

[0035] Optionally, the resolution output module is specifically used for: For each candidate matching entity Its final confidence level Based on baseline confidence level Determined together with the context correction factor, where, The value is assigned differently based on the different sources of the matching strategy: for matches that hit the controlled thesaurus, the confidence score directly inherits the comprehensive similarity score calculated in the thesaurus mining stage. For long-tail entities recalled via LLM, the discrimination probability output by LLM is directly used as the baseline confidence level; for matching based on regularity rules or edit distance, a dynamic threshold range is set, and weighting is applied according to the specificity of the rule. When there are multiple candidate entity sets with overlapping positions in the text At that time, the principle of maximizing utility should be followed to resolve the issue: Define entity utility function as follows: in, The length of the text covered by the entity; This is a semantic integrity indicator function; Final choice of utility value The largest entity is used as the output. While ensuring high confidence, the semantic integrity of TCM terms is preserved to the maximum extent, effectively avoiding the risk of semantic loss caused by fragmented matching, and finally outputting a normalized entity sequence.

[0036] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A semantic alignment system for ancient and modern Chinese medicine based on multi-dimensional feature fusion, characterized in that, The system includes: The knowledge base construction module is used to build a multi-source knowledge graph and a standard terminology glossary for TCM diagnosis and treatment; The multidimensional fusion similarity calculation model construction module is used to build a multidimensional fusion similarity calculation model. It comprehensively calculates the edit distance based on the longest continuous matching block of the text, the core phrase retention degree based on the longest common subsequence, the overlap of deduplicated characters, and the semantic similarity of the core stem after noise removal based on the TCM-specific modifier word list. It also introduces the organ compatibility constraint to hard filter and block the risk of erroneous normalization of ectopic pathogenesis entities. While preserving the core medical semantics, it bridges the gap between ancient and modern expression methods and achieves highly accurate alignment of ancient and modern terms. The controlled thesaurus generation module is used to generate thesaurus clusters by combining high-confidence nodes obtained through multi-dimensional fusion similarity calculation with the disjoint-set data structure algorithm, and to align national standard terms to generate a controlled thesaurus that guides the alignment of ancient and modern semantics. The complex negation logic parsing and semantic noise reduction module for clinical text is used to perform context delimitation based on semantic tags to prevent cross-domain false alarms on unstructured clinical medical record text, and to transform it into structured data blocks. It constructs a deep negation word detection algorithm to accurately handle pre-negation, post-negation, double exclusion and specific exception words, and filters out various semantic noise through a built-in semantic noise filter. The three-level cascaded matching module is used to map the clinical case text processed by the complex negation logic parsing and semantic noise reduction module for clinical text to the knowledge base language inside the model based on the three-level cascaded matching architecture of "precise mapping-similarity retrieval-generative alignment". The resolution output module is used to construct a utility function based on entity baseline confidence, covered text length and semantic integrity when there are matching candidate entities with overlapping positions. It follows the principle of maximum utility to perform resolution and outputs a normalized entity sequence.

2. The system according to claim 1, characterized in that, Defining an entity node corresponding to any two terms in the ancient and modern multi-source knowledge graph The calculation formula of the multi-dimensional fusion similarity calculation model is: (1) wherein, is a weight coefficient; is an edit distance-based sequence structure similarity, used to capture the overall structural similarity; is a longest common subsequence-based similarity, used to extract the preservation of core phrases, is a character overlap rate-based similarity, used to calculate the degree of pure character overlap without considering word order; is a core semantic-based similarity, used to eliminate domain noise and bridge the gap between ancient and modern writing habits.

3. The system according to claim 2, characterized in that, The The calculation method is as follows: A recursive matching algorithm based on the longest consecutive matching block of text is used to traverse and extract the first term. With the second term Calculate the total number of matched characters for all consecutive blocks of identical characters between them. Calculate the overall sequence structure similarity: in, and These represent the total character length of the two terminology texts, respectively.

4. The system according to claim 2, characterized in that, The The calculation method is as follows: Using dynamic programming, construct a system of size ... Given a two-dimensional state transition matrix dp, if during the iterative traversal, it satisfies... The state transition equation is dp[i][j] = dp[i-1][j-1] + 1; if this condition is not met, the current position count is interrupted and reset to zero; finally, the maximum value in the entire matrix is ​​obtained, which is the length of the longest common continuous substring. : 。 5. The system according to claim 2, characterized in that, The The calculation method is as follows: First, let's discuss the first terminology. Second term Convert to an unordered and deduplicated set of characters. and Then, find the intersection of the two sets and calculate the number of elements in the intersection; finally, divide by the number of elements in the set with the smaller cardinality. ; The The calculation method is as follows: A dictionary of modifiers specifically for Traditional Chinese Medicine (TCM) was constructed, covering common degree adverbs, time particles, and function words in TCM clinical expressions. Using the aforementioned modifier list, the original terms are respectively... and Text denoising and stemming are performed to remove all modifiers and obtain the corresponding first core stem. Second core stem ; Extracted core stems and Then, the recursive matching algorithm based on the longest continuous matching block is applied again to calculate the similarity score at the stem level, which is used as the core semantic similarity.

6. The system according to claim 2, characterized in that, The introduction of organ compatibility constraints specifically includes: For entity nodes corresponding to pathogenesis terms, and in conjunction with a lexicon of viscera and viscera, only node pairs with completely identical viscera and viscera sets or where one of them is a generic term are allowed to be merged. A viscera and viscera feature extraction function is defined. To identify and extract the internal organs contained in the current entity, assuming the extracted nodes The organs are a collection of ,node The organs are a collection of Binary constraint function : The judgment logic is as follows: Its output can only be (Mergers are allowed) or (Merge Prohibited), where 1 indicates merge is allowed, and 0 indicates merge is prohibited. If the output is... The following two conditions must be met: Condition 1 In other words, if at least one of the two entities does not explicitly point to a specific organ, the set is empty. ; Condition 2 That is, the two entities contain completely identical internal organs; Through logical multiplication This achieves a hard filtering of ectopic pathogenesis entities, among which... This represents the four-dimensional basic similarity score calculated at the literal or textual level between two pathogenesis entities, combined with... The judgment logic at the algorithm's underlying level prevents the risk of incorrect normalization of entities from different organ parts.

7. The system according to claim 1, characterized in that, The controlled thesaurus generation module is specifically used for: Define the set of entity nodes in the graph as If node pair satisfy , To set a threshold, edges are established. ; If node and An edge was built between them. and An edge has also been established between them, so based on the logic of transitive closure, it can be automatically deduced that... and They also belong to the same set; By continuously searching for and merging nodes with connectivity, all words directly or indirectly connected by edges are eventually packaged into an independent synonym cluster. For any satisfy: in This represents the nodes in the cluster. Derived from the total set of entity nodes in the graph ; This is a core condition for becoming part of the same cluster. For path compression lookup operations, ensure that all nodes within the same cluster point to a unique root node. Any two nodes inside and ,go through After the search operation, the root nodes they point to must be exactly the same. This ensures from the bottom layer of the algorithm that no matter how these similar nodes are connected by edges in the graph, as long as they belong to the same cluster, they will eventually converge and point to a unique representative root node, thus ensuring that the words are strictly classified into the same category semantically. Subsequently, the central terms of each cluster are aligned with the compiled national standard terminology glossary to generate a controlled thesaurus: Determine each cluster Standard mapping terms Define the set of national standard terms as Alignment mapping function The definition is as follows: in This represents the alignment threshold, appearing in synonym clusters. With the National Standard Terminology Collection Piecewise functions for mapping alignment In this process, it plays a role in quality control.

8. The system according to claim 1, characterized in that, The complex negation logic parsing and semantic noise reduction module for clinical texts is specifically used for: (1) For the semi-structured features of some clinical case texts, the module first performs context delimitation based on semantic tags: Define clinical case text as Using regular expression patterns The scan is performed, dividing it into four independent semantic subspaces: in Represents the symptom subspace; Represents the pathogenic factor space; Represents the pathogenesis subspace; Represents the treatment method subspace; Analytical functions The input natural language case studies are accurately converted into structured data blocks: This mechanism ensures contextual constraints for entity extraction, significantly reducing the false positive rate across domains; (2) Construct a negative word detection algorithm to handle pre-negation, post-negation, and double exclusion, and define an exception word set for words that contain negative words but are positive in nature. Perform an exception word check: For compound negation expressions, define a set of compound negation patterns. Using regular expression mode Perform clause splitting, extract the affirmative symptoms from the first half and add them to the matching queue, discard the negative descriptions from the second half: For a complete negation, define the set of negation prefixes and suffixes. Perform a negative judgment: in Indicates whether the affirmative portion can be extracted; For double exclusion negations, a logical exception rule is introduced to ensure that double negation expressions are not incorrectly cleaned up; Final negation handling strategy: (3) The module has a built-in semantic noise filter to clean the following three types of non-entity text: Normal state description: The script explicitly defines the set of ignored phrases. ,like If so, discard it directly; Western medicine and drug interference: Use regular expressions to remove Western medicine diagnosis / examination text and Western medicine drug names; Irrelevant modifier cleansing: cleans time adverbs that have no substantial referential meaning, but retains degree adverbs that have diagnostic value.

9. The system according to claim 1, characterized in that, The three-level cascaded matching module is specifically used for: The first level is based on a direct mapping between the controlled thesaurus and a general rule base in the field of Traditional Chinese Medicine: As the highest priority matching layer, this layer directly calls the controlled thesaurus and rule base built in the previous section to achieve rapid normalization of high-frequency and specific pattern terms; Second level, mixed similarity calculation: When the first-level matching fails, multi-dimensional feature fusion similarity calculation is initiated: for the terms to be matched... and candidate terms The same calculation method as formula (1) is used, and a knowledge base rule enhancement mechanism is introduced on this basis: Among them, the bonus items This includes bonuses for known synonyms, standard words, synonym replacement, synonym patterns, and core stem matching; penalty items. This includes punishments for incompatibility between internal organs and other parts of the body; Only keep Candidate words, The matching threshold; Level 3, LLM-based generative semantic alignment: when When this occurs, an LLM call is triggered, providing the LLM with a candidate list and matching guidance by editing a specific type of Prompt template.

10. The system according to claim 1, characterized in that, The digestion output module is specifically used for: For each candidate matching entity Its final confidence level Based on baseline confidence level Determined together with the context correction factor, where, The value is assigned differently based on the different sources of the matching strategy: for matches that hit the controlled thesaurus, the confidence score directly inherits the comprehensive similarity score calculated in the thesaurus mining stage. For long-tail entities recalled via LLM, the discrimination probability output by LLM is directly used as the baseline confidence level; for matching based on regularity rules or edit distance, a dynamic threshold range is set, and weighting is applied according to the specificity of the rule. When there are multiple candidate entity sets with overlapping positions in the text At that time, the principle of maximizing utility should be followed to resolve the issue: Define entity utility function as follows: in, The length of the text covered by the entity; This is a semantic integrity indicator function; Final choice of utility value The largest entity is used as the output. While ensuring high confidence, the semantic integrity of TCM terms is preserved to the maximum extent, effectively avoiding the risk of semantic loss caused by fragmented matching, and finally outputting a normalized entity sequence.