A training method of a traditional Chinese medicine vertical domain large model based on a knowledge graph

By constructing a symptom keyword database and generating easily confused groups through clustering, combined with dynamic disambiguation mechanisms and iterative optimization, the accuracy and applicability of TCM knowledge graphs in complex contexts were solved, and the model training effect was improved.

CN121436110BActive Publication Date: 2026-07-31DAAN HEALTH TECH (BEIJING) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DAAN HEALTH TECH (BEIJING) CO LTD
Filing Date
2025-09-22
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, TCM knowledge graphs lack accuracy and applicability in complex contexts, lack dynamic recognition and context disambiguation mechanisms for easily confused concepts, and lack quality verification and iterative optimization mechanisms, resulting in poor model training performance.

Method used

A symptom keyword library associated with syndrome keywords is constructed. A library of easily confused groups is formed through clustering. Corpus labeling and disambiguation are performed, substitution relationships are recorded, and a substitution keyword library is formed to realize the verification and optimization of easily confused groups.

Benefits of technology

It improves the semantic consistency and applicability of knowledge graphs, provides high-quality training data, solves the semantic ambiguity problem in the field of traditional Chinese medicine, and enhances the accuracy and adaptability of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121436110B_ABST
    Figure CN121436110B_ABST
Patent Text Reader

Abstract

This invention relates to the field of data processing, and more particularly to a training method for a large-scale TCM vertical domain model based on knowledge graphs. The invention systematically constructs a symptom keyword library associated with syndrome keywords, and clusters these symptom keyword libraries to form a library of easily confused groups. During the traversal of sample corpora, syndrome keywords are labeled and context-disambiguated based on these easily confused groups. Labeled keywords are replaced according to the matching results, and the easily confused groups are iteratively verified and optimized using a replacement keyword library. Finally, high-quality, highly consistent corpora are obtained as pre-training data. This invention achieves automated processing and semantic optimization of TCM corpora, effectively solving semantic ambiguity problems such as polysemy and synonymy in TCM texts, improving the semantic consistency between knowledge graph and language model training, and providing a high-quality training foundation for large-scale TCM domain models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and in particular to a training method for a large-scale TCM vertical domain model based on knowledge graphs. Background Technology

[0002] Traditional Chinese medicine (TCM) possesses a profound and complex knowledge system, containing a vast array of concepts, entities, and relationships, such as syndromes, symptoms, Chinese herbal medicines, and prescriptions. This knowledge is typically scattered in unstructured or semi-structured forms across ancient texts, clinical records, and modern research. A key challenge currently facing the intelligentization of TCM lies in how to systematically represent TCM knowledge and enable it to support the training of large-scale artificial intelligence models. Knowledge graphs, as a type of semantic network, can describe entities and their relationships in the real world in a structured form, providing a technical path for the effective organization and computation of knowledge in the TCM field. Constructing a TCM knowledge graph can serve as the foundation for training high-quality, highly semantically consistent large-scale TCM models, effectively improving the professional accuracy and logical reasoning capabilities of domain models.

[0003] Chinese Patent Publication No. CN117494811A discloses a method and system for constructing a knowledge graph of traditional Chinese medicine (TCM) classics, relating to the field of knowledge graph construction technology. The method includes: obtaining a first classic knowledge category from a classic knowledge database; retrieving the first classic directory of the first classic in a preset classic set and crawling to construct a first keyword set; obtaining the first page range corresponding to the first keyword, where the first keyword is any keyword in the first keyword set; when the first keyword exists in a first professional vocabulary set, extracting the first page range from the first classic and storing it in the first classic knowledge category to form a first graph relationship; constructing a first graph branch based on the first graph relationship, and constructing a target knowledge graph based on the first graph branch. This invention solves the technical problem of traditional methods being difficult to organize and associate the large amount of TCM medical knowledge contained in TCM classics, lacking structured information, and having low efficiency in searching for specific information.

[0004] However, the following problems still exist in the existing technology. 1. In the existing technology, the confusion of some syndrome keywords in the field of traditional Chinese medicine is not fully considered, and there is a lack of dynamic identification and context disambiguation mechanism for easily confused concepts, resulting in insufficient accuracy and applicability of the constructed knowledge graph in complex contexts; 2. In existing technologies, there is a lack of quality verification and iterative optimization mechanisms for knowledge graphs themselves, especially a lack of continuous evaluation of entity relationships and concept differentiation based on actual corpora and model feedback, making it difficult for the graph to evolve and improve synchronously with the model training process. Summary of the Invention

[0005] To address this, the present invention provides a training method for a large-scale TCM vertical domain model based on knowledge graphs. This method overcomes the problems in existing technologies that fail to fully consider the complexity of multiple descriptive terms that may describe different diseases commonly found in the field of TCM, resulting in insufficient accuracy and applicability of the constructed knowledge graph in complex contexts, and a lack of quality verification and iterative optimization mechanisms for the knowledge graph itself.

[0006] To achieve the above objectives, this invention provides a training method for a large-scale TCM vertical domain model based on knowledge graphs, comprising: Construct a symptom keyword library that associates syndrome keywords with syndrome keywords, wherein the symptom keyword library stores a number of symptom keywords that are associated with the syndrome keywords; Clustering is performed based on the symptom keyword database associated with each syndrome keyword to construct a database of easily confused groups, which contains easily confused groups. In response to the collected sample corpus, the syndrome keywords and symptom keywords in the sample corpus are traversed, and the syndrome keywords in the sample corpus are marked based on the easily confused tuple library; Disambiguation processing of the marked syndrome keywords is performed based on the easily confused groups corresponding to the marked syndrome keywords. This includes calling the syndrome keywords in the easily confused groups respectively, matching the symptom keyword library associated with each syndrome keyword with the text segment where the marked syndrome keyword is located, and determining whether to replace the marked syndrome keyword based on the matching result. Record the replacement keyword groups that form a replacement relationship to form a replacement keyword library. Verify the easily confused group based on the replacement keyword library, including determining the probability of each syndrome keyword in the easily confused group appearing as the replaced keyword in the replacement keyword library and determining whether to remove the syndrome keywords in the easily confused group. Obtain sample corpora as training data for pre-training TCM models; The easily confused tuple contains at least two syndrome keywords.

[0007] Furthermore, the process of constructing a symptom keyword database associated with syndrome keywords includes, Predetermine several key symptoms; Select several symptom keywords associated with the syndrome keywords, and store all the symptom keywords in the same thesaurus to obtain the symptom keyword library.

[0008] Furthermore, the symptom keyword database associated with each syndrome keyword is clustered to construct a database of easily confused groups, including... For any two syndrome keywords, construct several syndrome keyword groups; Obtain the symptom keyword library associated with each syndrome keyword in the syndrome keyword group, and calculate the semantic similarity between each symptom keyword in the symptom keyword library and each symptom keyword in another symptom keyword library; Calculate the average of each semantic similarity and use the average semantic similarity as the clustering index value of the syndrome keyword group; Based on the clustering index value and clustering threshold, the keywords of each syndrome are clustered to obtain several easily confused groups. The easily confused groups are stored in the same database to obtain an easily confused group library.

[0009] Further, clustering is performed on each of the syndrome keywords based on the clustering index value and the clustering threshold, including: If each syndrome keyword satisfies the clustering condition, then the syndrome keywords are grouped into easily confused groups. The clustering condition is that the clustering index values ​​among the keywords of each syndrome are all greater than the clustering threshold.

[0010] Furthermore, the process of traversing the syndrome keywords and symptom keywords in the sample corpus, and labeling the syndrome keywords in the sample corpus based on a tuple of easily confused tuples, includes: Read the collected sample corpus, perform word segmentation and keyword recognition on the sample corpus, and extract all syndrome keywords and symptom keywords; The extracted syndrome keywords are compared one by one with the syndrome keywords of each easily confused group in the easily confused group library. If the syndrome keyword is the same as the syndrome keyword in any easily confused set, then the syndrome keyword in the sample corpus is marked.

[0011] Furthermore, the syndrome keywords in the easily confused tuples are called respectively, and the symptom keyword library associated with each syndrome keyword is matched with the text segment where the syndrome keyword is located, including: Semantic analysis is performed on the text segments containing the marked syndrome keywords to extract the symptom keywords contained therein; The extracted symptom keywords are compared with the symptom keyword database associated with each syndrome keyword in the easily confused group, excluding the labeled syndrome keywords. Determine the average semantic relevance between each symptom keyword database and the extracted symptom keywords.

[0012] Furthermore, the process of obtaining matching results includes, Based on the mean semantic relevance, the keywords of each syndrome in the easily confused groups are sorted in descending order to determine the first syndrome keyword in the sorting.

[0013] Further, based on the matching results, determining whether to replace the marked syndrome keywords includes, Determine whether the syndrome keyword meets the replacement condition. If the replacement condition is met, replace the marked syndrome keyword with the syndrome keyword. The replacement condition is that the average semantic relevance of the syndrome keywords is greater than the predetermined semantic relevance replacement threshold, and the syndrome keywords are the first syndrome keywords in the ranking.

[0014] Furthermore, the replacement keyword groups that form replacement relationships are recorded to form a replacement keyword library, including, The labeled syndrome keywords generated in each disambiguation process and their replaced priority syndrome keywords are recorded as replacement keyword groups to form a replacement keyword library.

[0015] Furthermore, determining whether to remove syndrome keywords from the easily confused group is based on the probability of each syndrome keyword appearing as a replacement keyword in the replacement keyword library. From the replacement keyword library, the probability of each syndrome keyword in the easily confused group to be verified appearing as the replacement keyword is calculated. If the probability of a symptom keyword appearing as a replacement keyword is lower than a preset probability threshold, then the semantic distinguishability of the symptom keyword is determined to be high, and the symptom keyword is removed from the easily confused tuple.

[0016] Compared with existing technologies, this invention systematically constructs a symptom keyword library associated with syndrome keywords, and clusters these symptom keyword libraries to form a library of easily confused groups. During the traversal of sample corpora, syndrome keywords are labeled and contextually disambiguated based on these easily confused groups. Labeled keywords are replaced according to the matching results, and the easily confused groups are iteratively verified and optimized using a replacement keyword library. Ultimately, high-quality, highly consistent corpora are obtained as pre-training data. This invention achieves automated processing and semantic optimization of corpora in the field of Traditional Chinese Medicine (TCM), effectively solving semantic ambiguity problems such as polysemy and synonymy in TCM texts, improving the semantic consistency between knowledge graph and language model training, and providing a high-quality training foundation for large-scale TCM models.

[0017] In particular, this invention addresses the semantic ambiguity and conceptual overlap issues that are common in texts within the field of Traditional Chinese Medicine (TCM). In practice, the same symptom description may correspond to multiple different syndromes, and the same syndrome is often characterized by a combination of symptoms. This complex many-to-many mapping makes it difficult for traditional text processing methods to accurately capture semantic intent, leading to noise accumulation and semantic bias in subsequent model training. This invention constructs a symptom keyword library and further clusters it to generate easily confused tuples. It systematically identifies and structurally manages these easily confused concepts, thereby achieving accurate identification and differentiation of polysemous phenomena during the corpus cleaning stage. This provides the model with high-quality training samples that are semantically clear and consistently labeled.

[0018] In particular, this invention addresses the semantic inconsistencies and changing usage of concepts between the knowledge structure of Traditional Chinese Medicine (TCM) and actual corpora. In practice, TCM concepts often exhibit inconsistent naming and usage across different classic texts, schools of thought, and clinical records. Furthermore, language evolves over time, making simple string matching or dictionary-based methods ill-suited to the complex contexts of real-world texts. This invention introduces a disambiguation mechanism based on dynamic matching and context awareness. By identifying easily confused tuples and combining this with all symptom keywords in the current text segment for comprehensive matching and judgment, it achieves more accurate semantic alignment, improves the quality and reliability of corpus annotation, and provides a more reliable training data foundation for pre-training TCM models.

[0019] In particular, this invention addresses the lack of coordination and dynamic updates between knowledge graphs and corpus processing. In practice, knowledge graphs are often static after construction, making timely adjustments based on actual corpus distribution and model training feedback difficult, thus limiting their adaptability and effectiveness in real-world applications. This invention creates a replacement keyword library by recording replacement operations and verifies and filters easily confused tuples based on their occurrence probabilities. This forms a complete iterative optimization mechanism from corpus processing to knowledge verification and then to graph optimization, enabling the knowledge graph to continuously improve itself based on replacement feedback accumulated in actual use, thereby enhancing its representation quality of complex semantic relationships in Traditional Chinese Medicine and its support for model training.

[0020] In particular, this invention addresses the scarcity and high cost of acquiring high-quality labeled corpora for training large-scale TCM models. In practice, TCM corpora require manual annotation and verification by experts, which is inefficient, limited in scale, and susceptible to subjective biases, severely restricting the data scale and quality required for large-scale model training. This invention, through automated, algorithm-driven corpus disambiguation and enhancement processing, enables semantic annotation of large-scale raw corpora under unsupervised or weakly supervised conditions. This effectively expands the sources of high-quality training samples, reduces reliance on manual annotation, and lays a data foundation for training high-performance TCM models. Attached Figure Description

[0021] Figure 1 This is a schematic diagram illustrating the training steps of a knowledge graph-based large-scale model for traditional Chinese medicine vertical domains, as an embodiment of the invention. Figure 2 This is a logic block diagram for constructing easily confused units in an embodiment of the invention; Figure 3 This is a logic block diagram illustrating the replacement of syndrome keywords in an embodiment of the invention. Figure 4 This is a logic block diagram for removing syndrome keywords from easily confused groups in an embodiment of the invention. Detailed Implementation

[0022] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0023] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0024] Please see Figure 1 As shown, Figure 1 This is a schematic diagram illustrating the steps of a training method for a knowledge graph-based large-scale TCM vertical domain model according to an embodiment of the invention. The training method for a knowledge graph-based large-scale TCM vertical domain model according to an embodiment of the invention includes: Step S1: Construct a symptom keyword library associated with syndrome keywords, wherein the symptom keyword library stores several symptom keywords associated with the syndrome keywords; Step S2: Cluster the symptom keyword library associated with each syndrome keyword to construct a library of easily confused groups, wherein the library of easily confused groups contains easily confused groups; Step S3: In response to the collection of sample corpus, traverse the syndrome keywords and symptom keywords in the sample corpus, and mark the syndrome keywords in the sample corpus based on the easily confused tuple library; Step S4: Disambiguation processing is performed on the marked syndrome keywords based on the easily confused tuples corresponding to the marked syndrome keywords. This includes calling the syndrome keywords in the easily confused tuples respectively, matching the symptom keyword library associated with each syndrome keyword with the text segment where the marked syndrome keyword is located, and determining whether to replace the marked syndrome keyword based on the matching result. Step S5: Record the replacement keyword groups that form a replacement relationship to form a replacement keyword library. Based on the replacement keyword library, verify the easily confused group, including determining the probability of each syndrome keyword in the easily confused group appearing as the replaced keyword in the replacement keyword library and determining whether to remove the syndrome keywords in the easily confused group. Step S6: Obtain sample corpus as training data for pre-training the TCM model; The easily confused tuple contains at least two syndrome keywords.

[0025] Specifically, the process of constructing a symptom keyword library that associates syndrome keywords includes, Predetermine several key symptoms; Select several symptom keywords associated with the syndrome keywords, and store all the symptom keywords in the same thesaurus to obtain the symptom keyword library.

[0026] Specifically, in advance, technical personnel in the field of traditional Chinese medicine select several clear syndrome keywords and corresponding symptom keywords based on classic TCM works, clinical guidelines and authoritative medical records to construct a symptom keyword database associated with syndrome keywords. There are no restrictions on the type of syndrome keywords, such as "internal heat", "yin deficiency", "yang deficiency", etc., which will not be elaborated here.

[0027] Symptom keywords are keywords that describe the relevant symptoms under the corresponding syndrome, which will not be elaborated further here.

[0028] Specifically, the clustering based on the symptom keyword database associated with each syndrome keyword to construct a database of easily confused groups includes... For any two syndrome keywords, construct several syndrome keyword groups; Obtain the symptom keyword library associated with each syndrome keyword in the syndrome keyword group, and calculate the semantic similarity between each symptom keyword in the symptom keyword library and each symptom keyword in another symptom keyword library; Calculate the average of each semantic similarity and use the average semantic similarity as the clustering index value of the syndrome keyword group; Based on the clustering index value and clustering threshold, the keywords of each syndrome are clustered to obtain several easily confused groups. The easily confused groups are stored in the same database to obtain an easily confused group library.

[0029] Specifically, there are no restrictions on the method for solving semantic similarity. For example, keywords can be vectorized and then cosine similarity can be calculated, and the cosine similarity can be used as semantic similarity. Of course, other methods can also be used, which will not be elaborated here.

[0030] Specifically, this invention addresses the semantic ambiguity and conceptual overlap issues that are common in texts within the field of Traditional Chinese Medicine (TCM). In practice, the same symptom description may correspond to multiple different syndromes, and the same syndrome is often characterized by a combination of symptoms. This complex many-to-many mapping makes it difficult for traditional text processing methods to accurately capture semantic intent, leading to noise accumulation and semantic bias in subsequent model training. This invention constructs a symptom keyword library and further clusters it to generate easily confused tuples. It systematically identifies and structurally manages these easily confused concepts, thereby achieving accurate identification and differentiation of polysemous phenomena during the corpus cleaning stage. This provides the model with high-quality training samples that are semantically clear and consistently labeled.

[0031] This invention addresses the semantic inconsistencies and evolving usage of concepts between the knowledge structure of Traditional Chinese Medicine (TCM) and actual textual corpora. In practice, TCM concepts often exhibit inconsistent naming and usage across different classic texts, schools of thought, and clinical records. Furthermore, language evolves over time, making simple string matching or dictionary-based methods inadequate for the complex contexts of real-world texts. This invention introduces a disambiguation mechanism based on dynamic matching and context awareness. By identifying easily confused tuples and combining this with all symptom keywords in the current text segment for comprehensive matching and judgment, it achieves more accurate semantic alignment, improves the quality and reliability of corpus annotation, and provides a more reliable training data foundation for pre-training TCM models.

[0032] Specifically, please refer to Figure 2 As shown, Figure 2 This is a logical block diagram for constructing easily confused tuples according to an embodiment of the invention. Clustering is performed on each of the syndrome keywords based on the clustering index value and the clustering threshold, including: If each syndrome keyword satisfies the clustering condition, then the syndrome keywords are grouped into easily confused groups. The clustering condition is that the clustering index values ​​among the keywords of each syndrome are all greater than the clustering threshold.

[0033] Understandably, under clustering conditions, the symptom keyword databases associated with each keyword of a syndrome are highly correlated, indicating that the keywords of a syndrome are easily confused.

[0034] Specifically, the clustering threshold is determined in advance by those skilled in the art, with the aim of characterizing the high correlation between the symptom keyword databases associated with syndrome keywords. In practice, the closer the clustering index obtained by the cosine similarity calculation method is to 1, the higher the similarity. Based on this, the clustering threshold is selected in the range [0.65, 0.85] to characterize the case of high similarity. In practice, 0.75 is preferred.

[0035] Specifically, the process of traversing the syndrome keywords and symptom keywords in the sample corpus, and marking the syndrome keywords in the sample corpus based on a tuple of easily confused groups, includes: Read the collected sample corpus, perform word segmentation and keyword recognition on the sample corpus, and extract all syndrome keywords and symptom keywords; The extracted syndrome keywords are compared one by one with the syndrome keywords of each easily confused group in the easily confused group library. If the syndrome keyword is the same as the syndrome keyword in any easily confused set, then the syndrome keyword in the sample corpus is marked.

[0036] Specifically, the syndrome keywords in the easily confused tuples are called respectively, and the symptom keyword library associated with each syndrome keyword is matched with the text segment where the syndrome keyword is located, including: Semantic analysis is performed on the text segments containing the marked syndrome keywords to extract the symptom keywords contained therein; The extracted symptom keywords are compared with the symptom keyword database associated with each syndrome keyword in the easily confused group, excluding the labeled syndrome keywords. Determine the average semantic relevance between each symptom keyword database and the extracted symptom keywords.

[0037] It is understandable that the mean semantic relevance is obtained by calculating the semantic relevance between the symptom keyword and each symptom keyword in the symptom keyword library, and then calculating the mean.

[0038] Specifically, the process of obtaining matching results includes, Based on the mean semantic relevance, the keywords of each syndrome in the easily confused groups are sorted in descending order to determine the first syndrome keyword in the sorting.

[0039] It is understandable that the mean semantic relevance corresponds to a symptom keyword library, and then to syndrome keywords. Based on this, the syndrome keywords of easily confused groups can be sorted in descending order according to the mean semantic relevance, which will not be elaborated further.

[0040] Specifically, please refer to Figure 3 As shown, Figure 3 This is a logic block diagram of replacing syndrome keywords according to an embodiment of the invention. Determining whether to replace the marked syndrome keywords based on the matching results includes: Determine whether the syndrome keyword meets the replacement condition. If the replacement condition is met, replace the marked syndrome keyword with the syndrome keyword. The replacement condition is that the average semantic relevance of the syndrome keywords is greater than the predetermined semantic relevance replacement threshold, and the syndrome keywords are the first syndrome keywords in the ranking.

[0041] Specifically, the replacement conditions consist of two dimensions: the first symptom keyword in the sorting must be the most relevant symptom keyword among the symptom keywords in the text segment, and the average semantic relevance must be greater than the predetermined semantic relevance replacement threshold to indicate that the symptom keyword is extremely close to the symptoms described in the text segment, thus indicating a possible ambiguity.

[0042] The semantic relevance replacement threshold is determined based on the clustering threshold and is set as the product of the clustering threshold and the accuracy coefficient. This invention addresses the lack of coordination and dynamic updates between knowledge graphs and corpus processing. In practice, knowledge graphs are often static after construction, making timely adjustments based on actual corpus distribution and model training feedback difficult, thus limiting their adaptability and effectiveness in real-world applications. This invention creates a replacement keyword library by recording replacement operations and verifies and filters easily confused tuples based on their occurrence probabilities. This forms a complete iterative optimization mechanism from corpus processing to knowledge verification and then to graph optimization, enabling the knowledge graph to continuously improve itself based on replacement feedback accumulated in actual use, thereby enhancing its representation quality of complex semantic relationships in Traditional Chinese Medicine and its support for model training.

[0043] Specifically, it records the replacement keyword groups that form substitution relationships, forming a replacement keyword library, including... The labeled syndrome keywords generated in each disambiguation process and their replaced priority syndrome keywords are recorded as replacement keyword groups to form a replacement keyword library.

[0044] Please see Figure 4 As shown, Figure 4 The following is a logical block diagram illustrating the process of removing syndrome keywords from easily confused groups according to an embodiment of the invention. The process includes determining the probability of each syndrome keyword in the easily confused group appearing as a replaced keyword in the replacement keyword library, and then deciding whether to remove the syndrome keywords from the easily confused group. From the replacement keyword library, the probability of each syndrome keyword in the easily confused group to be verified appearing as the replacement keyword is calculated. If the probability of a symptom keyword appearing as a replacement keyword is lower than a preset probability threshold, then the semantic distinguishability of the symptom keyword is determined to be high, and the symptom keyword is removed from the easily confused tuple.

[0045] Specifically, the preset probability threshold is determined based on a statistical distribution method. The probability of several easily confused symptom keywords appearing as replacement keywords is pre-statistically calculated, and based on the statistical distribution characteristics of the probability values, a specific quantile value is selected as the threshold. For example, selecting 20% ​​as the threshold means that when the "replacement probability" of a symptom keyword is lower than 80% of other similar keywords, it is considered to have "high semantic distinguishability" and is eliminated. The reference range for selecting the specific quantile in this embodiment of the invention is [10%, 25%]. The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

Claims

1. A training method for a large-scale TCM vertical domain model based on knowledge graphs, characterized in that, include: Construct a symptom keyword library that associates syndrome keywords with syndrome keywords, wherein the symptom keyword library stores a number of symptom keywords that are associated with the syndrome keywords; Clustering is performed based on the symptom keyword database associated with each syndrome keyword to construct a database of easily confused groups, which contains easily confused groups. In response to the collected sample corpus, the syndrome keywords and symptom keywords in the sample corpus are traversed, and the syndrome keywords in the sample corpus are marked based on the easily confused tuple library; Disambiguation processing of the marked syndrome keywords is performed based on the easily confused groups corresponding to the marked syndrome keywords. This includes calling the syndrome keywords in the easily confused groups respectively, matching the symptom keyword library associated with each syndrome keyword with the text segment where the marked syndrome keyword is located, and determining whether to replace the marked syndrome keyword based on the matching result. Record the replacement keyword groups that form a replacement relationship to form a replacement keyword library. Verify the easily confused group based on the replacement keyword library, including determining the probability of each syndrome keyword in the easily confused group appearing as the replaced keyword in the replacement keyword library and determining whether to remove the syndrome keywords in the easily confused group. Obtain sample corpora as training data for pre-training TCM models; The easily confused tuples contain at least two syndrome keywords; Determining whether to replace the marked syndrome keywords based on the matching results includes determining whether the syndrome keywords meet the replacement conditions. If the replacement conditions are met, the marked syndrome keywords are replaced with the syndrome keywords. The replacement condition is that the average semantic relevance of the syndrome keywords is greater than the predetermined semantic relevance replacement threshold, and the syndrome keywords are the first syndrome keywords in the sorting.

2. The training method for a large-scale TCM vertical domain model based on knowledge graphs according to claim 1, characterized in that, The process of constructing a symptom keyword database associated with syndrome keywords includes, Predetermine several key symptoms; Select several symptom keywords associated with the syndrome keywords, and store all the symptom keywords in the same thesaurus to obtain the symptom keyword library.

3. The training method for a large-scale TCM vertical domain model based on knowledge graphs according to claim 1, characterized in that, The symptom keyword database associated with each syndrome keyword is clustered to construct a library of easily confused groups, including... For any two syndrome keywords, construct several syndrome keyword groups; Obtain the symptom keyword library associated with each syndrome keyword in the syndrome keyword group, and calculate the semantic similarity between each symptom keyword in the symptom keyword library and each symptom keyword in another symptom keyword library; Calculate the average of each semantic similarity and use the average semantic similarity as the clustering index value of the syndrome keyword group; Based on the clustering index value and clustering threshold, the keywords of each syndrome are clustered to obtain several easily confused groups. The easily confused groups are stored in the same database to obtain an easily confused group library.

4. The training method for a large-scale TCM vertical domain model based on knowledge graphs according to claim 3, characterized in that, Clustering is performed on each of the syndrome keywords based on the clustering index value and clustering threshold, including: If each syndrome keyword satisfies the clustering condition, then the syndrome keywords are grouped into easily confused groups. The clustering condition is that the clustering index values ​​among the keywords of each syndrome are all greater than the clustering threshold.

5. The training method for a large-scale TCM vertical domain model based on knowledge graphs according to claim 1, characterized in that, The process involves traversing the syndrome keywords and symptom keywords in the sample corpus, and then tagging the syndrome keywords in the sample corpus based on a tuple of easily confused groups. include, Read the collected sample corpus, perform word segmentation and keyword recognition on the sample corpus, and extract all syndrome keywords and symptom keywords; The extracted syndrome keywords are compared one by one with the syndrome keywords of each easily confused group in the easily confused group library. If the syndrome keyword is the same as the syndrome keyword in any easily confused set, then the syndrome keyword in the sample corpus is marked.

6. The training method for a large-scale TCM vertical domain model based on knowledge graphs according to claim 1, characterized in that, The syndrome keywords in the easily confused tuples are called respectively, and the symptom keyword library associated with each syndrome keyword is matched with the text segment where the syndrome keyword is located, including: Semantic analysis is performed on the text segments containing the marked syndrome keywords to extract the symptom keywords contained therein; The extracted symptom keywords are compared with the symptom keyword database associated with each syndrome keyword in the easily confused group, excluding the labeled syndrome keywords. Determine the average semantic relevance between each symptom keyword database and the extracted symptom keywords.

7. The training method for a large-scale TCM vertical domain model based on knowledge graphs according to claim 6, characterized in that, The process of obtaining matching results includes, Based on the mean semantic relevance, the keywords of each syndrome in the easily confused groups are sorted in descending order to determine the first syndrome keyword in the sorting.

8. The training method for a large-scale TCM vertical domain model based on knowledge graphs according to claim 1, characterized in that, Record the replacement keyword groups that form substitution relationships to form a replacement keyword library, including: The labeled syndrome keywords generated in each disambiguation process and their replaced priority syndrome keywords are recorded as replacement keyword groups to form a replacement keyword library.

9. The training method for a large-scale TCM vertical domain model based on knowledge graphs according to claim 1, characterized in that, Determine the probability of each syndrome keyword in the easily confused group appearing as the replacement keyword in the replacement keyword library to decide whether to remove syndrome keywords from the easily confused group, including: From the replacement keyword library, the probability of each syndrome keyword in the easily confused group to be verified appearing as the replacement keyword is calculated. If the probability of a symptom keyword appearing as a replacement keyword is lower than a preset probability threshold, then the semantic distinguishability of the symptom keyword is determined to be high, and the symptom keyword is removed from the easily confused tuple.