A rose acid c-based epilepsy medical knowledge graph construction method and system
By preprocessing and terminology conversion of initial medical text information, and constructing a knowledge graph using a medical terminology relationship knowledge base, the problem of low recognition accuracy caused by differences in medical habits and language expression in different regions is solved, achieving high-quality knowledge graph construction and new drug development support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- THE FIRST AFFILIATED HOSPITAL OF WENZHOU MEDICAL UNIV
- Filing Date
- 2026-03-13
- Publication Date
- 2026-06-26
AI Technical Summary
Existing systems struggle to accurately integrate medical data from different regions with varying medical practices and language expressions, resulting in low accuracy in knowledge graph recognition and hindering updates and expansion.
By acquiring initial medical text information, performing preprocessing and terminology conversion, and using a medical terminology relationship knowledge base to convert non-standard medical text information into standardized information, and extracting medical information related to the mechanism of action of rose acid C molecules and epilepsy diagnosis and treatment, a knowledge graph is constructed.
It improves the accuracy of data identification and the quality of knowledge graphs, significantly enhances the updating and expansion capabilities of knowledge graphs, and supports new drug development and clinical decision-making.
Smart Images

Figure CN121835852B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of knowledge graph construction technology, and in particular to a method and system for constructing a medical knowledge graph for epilepsy based on rose acid C. Background Technology
[0002] In cutting-edge fields such as epilepsy treatment and new drug development, medical information is rapidly increasing. Existing systems aim to aggregate scattered diagnostic and treatment information by learning fixed collocations, contextual meanings, and sentence structures of words in medical materials, thus building a knowledge graph around roseolic acid C and epilepsy treatment. However, due to differences in medical practices, research cultures, and language expressions across regions, medical materials often use many regionally specific and non-standard medical expressions when describing medical concepts. Existing systems struggle to accurately integrate this medical information, resulting in low data recognition accuracy and hindering the updating and expansion of the knowledge graph, ultimately leading to low-quality knowledge graphs.
[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention
[0004] The main objective of this invention is to propose a method and system for constructing a medical knowledge graph for epilepsy based on rose acid C. This method can convert non-standard medical text information into standardized medical text information by combining a medical terminology relationship knowledge base, thereby constructing a knowledge graph and improving the accuracy of data recognition and the quality of the knowledge graph.
[0005] On one hand, embodiments of the present invention provide a method for constructing a medical knowledge graph for epilepsy based on rose acid C, including the following steps:
[0006] Obtain initial medical text information;
[0007] The initial medical text information is preprocessed to obtain non-standard medical text information;
[0008] Based on a medical terminology knowledge base, the non-standard medical text information is converted into standardized medical text information.
[0009] Extract medical information related to the mechanism of action of rose acid C molecules and the diagnosis and treatment of epilepsy from the standardized medical text information;
[0010] The medical information will be added to the knowledge graph database.
[0011] On the other hand, embodiments of the present invention provide a knowledge graph construction system for epilepsy treatment based on rose acid C, including:
[0012] The information acquisition module is used to acquire initial medical text information;
[0013] The preprocessing module is used to preprocess the initial medical text information to obtain non-standard medical text information;
[0014] The terminology conversion module is used to convert the non-standard medical text information into standardized medical text information based on a medical terminology relationship knowledge base.
[0015] The medical information extraction module is used to extract medical information related to the mechanism of action of rose acid C molecules and the diagnosis and treatment of epilepsy from the standardized medical text information;
[0016] The knowledge graph augmentation module is used to add the medical information to the knowledge graph database.
[0017] The embodiments of this application include at least the following beneficial effects: First, the initial medical text information is obtained. Then, the initial medical text information is preprocessed to obtain non-standard medical text information. Next, based on the medical terminology relationship knowledge base, the non-standard medical text information is converted into standardized medical text information. Finally, medical information related to the mechanism of action of rose acid C molecules and epilepsy diagnosis and treatment is extracted from the standardized medical text information, and the medical information is added to the knowledge graph database. Thus, it is possible to combine the medical terminology relationship knowledge base to convert non-standard medical text information into standardized medical text information to construct a knowledge graph, thereby improving the accuracy of data recognition and the quality of the knowledge graph.
[0018] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the description and the drawings. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below.
[0020] Figure 1 This is a flowchart illustrating a method for constructing a medical knowledge graph for epilepsy based on rose acid C, according to an embodiment of the present invention.
[0021] Figure 2 This is a schematic diagram of a knowledge graph construction system for epilepsy based on rose acid C, according to an embodiment of the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments.
[0023] In related technologies, particularly in cutting-edge fields such as epilepsy treatment and new drug development, medical information is rapidly increasing. Existing systems aim to aggregate scattered diagnostic and treatment information by learning fixed collocations, contextual meanings, and sentence structures of words in medical materials, thus building a knowledge graph around roseolic acid C and epilepsy treatment. However, due to differences in medical practices, research cultures, and language expressions across regions, medical materials often use many regionally specific and non-standard medical expressions when describing medical concepts. Existing systems struggle to accurately integrate this medical information, resulting in low data recognition accuracy and hindering the updating and expansion of the knowledge graph, ultimately leading to low-quality knowledge graphs.
[0024] For example, when constructing a knowledge graph for epilepsy based on rosehip C, the main goal is to integrate fragmented knowledge about epilepsy diagnosis and treatment, and to improve the data utilization efficiency of rosehip C in new drug development. Existing systems automatically identify medical terms and specific information related to the molecular mechanism of action of rosehip C, the pathogenesis of epilepsy, and related treatment plans from a large number of research articles, clinical reports, and patent documents to construct the knowledge graph. When the system is initially running, it can find most of the necessary information when processing mainstream and standardized medical literature.
[0025] However, as knowledge graphs are constantly updated and new content is added, things become more complicated, often by incorporating the latest clinical trial data from different research institutions. This new data typically details the pathway of action of rosmarinic acid C, its efficacy in epilepsy models, and the related pathogenesis. However, due to differences in medical practices, research cultures, and language expressions across regions, these data often use many regionally specific and non-standard medical expressions when describing medical concepts. For example, the standardized literature might uniformly use "epileptic seizure" for the concept of "epileptic seizure," but some local reports might use terms like "abnormal brain electrical discharge event," "paroxysmal neuronal hyperexcitability," or even some local clinical slang to describe it. These expressions are semantically similar or highly related, but differ significantly in word choice and sentence structure.
[0026] Faced with such non-standard data, existing systems show a significant decline in performance when processing localized and non-standard expressions. This is because, during the design and learning phases, they primarily rely on mainstream, rigorously standardized medical literature. By learning the fixed collocations, contextual meanings, and sentence structures of words in these materials, a deep recognition pattern for standardized medical language is formed. When encountering expressions like "abnormal brain electrical discharge events," which differ greatly in form from "epileptic seizure" in the learning materials, even if the underlying meanings are similar, existing systems struggle to accurately identify them. This leads to a significant decrease in recognition rate and accuracy when identifying information closely related to the mechanism of action of rosehip C molecules, such as specific metabolites, receptor types, key proteins in signaling pathways, and drug interactions, thus affecting the effective extraction of new knowledge and the integrity of the knowledge graph.
[0027] The embodiments of this application will be explained in detail below with reference to the accompanying drawings:
[0028] Figure 1 This is an optional flowchart of a method for constructing a medical knowledge graph for epilepsy based on rose acid C, provided in an embodiment of this application. Figure 1 The method may include, but is not limited to, steps S101 to S105.
[0029] Step S101: Obtain initial medical text information;
[0030] Step S102: Preprocess the initial medical text information to obtain non-standard medical text information;
[0031] Step S103: Based on the medical terminology relationship knowledge base, perform terminology conversion on non-standard medical text information to obtain standardized medical text information;
[0032] Step S104: Extract medical information related to the mechanism of action of rose acid C molecules and the diagnosis and treatment of epilepsy from standardized medical text information;
[0033] Step S105: Add medical information to the knowledge graph database.
[0034] Steps S101 to S105, as illustrated in the embodiments of this application, can convert non-standard medical text information into standardized medical text information by combining a medical terminology relationship knowledge base to construct a knowledge graph, thereby improving the accuracy of data identification and the quality of the knowledge graph. Rosehipic acid C is an organic compound with potential medicinal value, playing an important role in the treatment of neurological diseases (especially in the field of anti-epileptic drugs). In the prevention and treatment of epilepsy, rosehipic acid C can restore abnormal electroencephalographic activity during epileptic seizures, significantly reducing the frequency and severity of seizures. Its mechanism of action includes: inhibiting the activity of the excitatory neurotransmitter receptor NMDA (N-methyl-D-aspartate receptor) in the hippocampus; stimulating the activity of the inhibitory neurotransmitter receptor GABA (γ-aminobutyric acid); and reducing excessive excitation of the cerebral cortex, thereby exerting an anti-epileptic effect.
[0035] The systems or pathways potentially affected by rosehip acid C include the GABAergic system, the glutamatergic system, and inflammatory pathways. Within the GABAergic system, γ-aminobutyric acid (GABA) is the primary inhibitory neurotransmitter in the central nervous system, and dysfunction of the GABAergic system is closely related to the occurrence of epilepsy. Rosehip acid C may affect the GABAergic system, exerting its anti-epileptic effect by enhancing inhibitory neurotransmission. Within the glutamatergic system, glutamate is the primary excitatory neurotransmitter in the central nervous system, and overexcitation of the glutamatergic system is also an important mechanism in the pathogenesis of epilepsy. Rosehip acid C may affect the glutamatergic system, exerting its effect by inhibiting excitatory neurotransmission. In the inflammatory pathway, neuroinflammation plays a crucial role in the occurrence and development of epilepsy. Rosehip acid C may affect the inflammatory pathway, exerting its anti-epileptic effect by reducing neuroinflammation.
[0036] In some embodiments, steps S101-S105 may involve first acquiring initial medical text information. For example, text data can be periodically retrieved from preset web links or file directories using a simple automated script. Alternatively, text content can be manually copied by operators from sources such as medical journals, clinical reports, and online databases. This initial medical text information may include various formats, such as plain text, PDF documents, and web page content. It is understood that initial medical text information refers to unstructured or semi-structured text data such as original medical literature, clinical records, and research reports obtained from various sources.
[0037] The initial medical text is then preprocessed to obtain non-standard medical text, aiming to remove noise and inconsistencies from the original text. For example, rule-based text cleaning methods can be used, employing regular expressions to identify and remove Hypertext Markup Language tags, special symbols, redundant spaces, and line breaks. Alternatively, simple string replacement operations can be used to uniformly convert all letters in the text to lowercase or uppercase, achieving case normalization. For sentence boundary analysis, simple segmentation methods based on punctuation marks (such as periods, question marks, and exclamation marks) can be used to divide the text into independent sentences.
[0038] Then, based on a medical terminology relation knowledge base, non-standard medical text information is converted into standardized medical text information, aiming to solve the problem of inconsistent medical terminology expression. For example, a dictionary-based matching method can be used to precisely match words in non-standard medical text information with the medical terminology relation knowledge base. If the match is successful, the standardized term in the knowledge base is directly used for replacement. If the precise match fails, a string similarity-based method, such as Levenshtein distance or Jaccard similarity, can be used to perform fuzzy matching between non-standard medical terms and terms in the knowledge base, identifying the most similar standardized term for replacement. It is understandable that the medical terminology relation knowledge base is a structured database that stores a large number of medical terms and their relationships (such as synonyms, hyponyms, and related words) to support the standardized conversion of terminology.
[0039] This involves extracting medical information related to the mechanism of action of rosmarinic acid C and the diagnosis and treatment of epilepsy from standardized medical text. For example, keyword matching can be used, pre-setting a series of keywords related to rosmarinic acid C and epilepsy treatment. By scanning standardized text, sentences or paragraphs containing these keywords are identified and treated as potential medical information. Alternatively, rule-based entity recognition methods can be employed, defining simple grammatical rules or patterns to identify entities such as diseases, drugs, symptoms, and treatments in the text. Medical information refers to entities, relationships, or events identified from standardized text that are directly related to the mechanism of action of rosmarinic acid C and the diagnosis and treatment of epilepsy. This medical information may include pathways of action, pharmacodynamics, metabolites, receptor types, key proteins in signaling pathways, drug interactions, and the pathogenesis of epilepsy. Among them, the pathway of action refers to the biological route by which rosehip acid C exerts its effects in vivo; pharmacodynamic performance refers to the therapeutic effects of rosehip acid C in various epilepsy experimental models (such as animal models and cell models); metabolites refer to the biologically active substances produced during the metabolism of rosehip acid C in vivo; receptor type refers to the specific receptors on the cell surface or inside the cell affected by rosehip acid C; key proteins in signaling pathways refer to important proteins in the cell signal transduction pathways regulated by rosehip acid C; drug interactions refer to the possible changes in pharmacodynamics or toxicity when rosehip acid C is used concurrently with other drugs; and epilepsy pathogenesis refers to information on the epilepsy pathogenesis mechanism related to the mode of action of rosehip acid C, such as the influence of rosehip acid C on the occurrence and development of epilepsy.
[0040] Finally, medical information is added to a knowledge graph database to store the extracted structured information. For example, extracted entities and relations can be stored in a relational database as triples (subject-verb-object), with each triple representing a knowledge fact. Alternatively, this information can be stored in a graph database as nodes and edges, with entities as nodes and relations as edges.
[0041] This embodiment effectively transforms raw, non-standardized medical text information into structured medical information usable in a knowledge graph. First, it acquires initial medical text information, providing a data foundation for subsequent processing. Then, it preprocesses the initial medical text information, removing noise and inconsistencies, laying the groundwork for terminology conversion. Subsequently, it performs terminology conversion based on a medical terminology relationship knowledge base, resolving inconsistencies in medical terminology expression and ensuring information standardization. Further, it extracts medical information related to the mechanism of action of rosmarinic acid C and epilepsy diagnosis and treatment from the standardized medical text information, accurately identifying the target knowledge. Finally, it adds the extracted medical information to the knowledge graph database, achieving knowledge storage and management. This embodiment effectively aggregates scattered diagnostic and treatment information and better utilizes data on rosmarinic acid C in new drug development, overcoming the limitations of existing entity recognition programs when processing non-standard medical data, and significantly improving the updating and expansion capabilities of the knowledge graph.
[0042] Through the above technical solutions, this embodiment significantly improves the ability to extract accurate and standardized knowledge from complex medical texts by introducing preprocessing, terminology conversion based on a medical terminology relationship knowledge base, and extraction of medical information for specific fields. For example, in the terminology conversion stage, not only is exact matching considered, but non-standard terms are also handled through a fuzzy matching mechanism, which is often a challenge for entity recognition in existing technologies. Furthermore, this embodiment focuses on the field of roselic acid C and epilepsy diagnosis and treatment, making the construction of the knowledge graph more targeted and practical, and more effectively supporting new drug development and clinical decision-making. This method not only improves the efficiency and accuracy of knowledge graph construction but also provides a new approach for the effective utilization of medical knowledge.
[0043] In some embodiments, step S102 involves preprocessing the initial medical text information to obtain non-standard medical text information, which may include, but is not limited to, the following steps:
[0044] The initial medical text information is cleaned to remove Hypertext Markup Language tags, special symbols, redundant spaces, and line breaks.
[0045] Perform case normalization on the initial medical text information after text cleaning;
[0046] Sentence boundary analysis is performed on the initial medical text information after case normalization to obtain non-standard medical text information.
[0047] In some embodiments, the initial medical text information can be pre-processed with text cleaning to remove Hypertext Markup Language (HTML) tags, special symbols, redundant spaces, and line breaks. This initial cleaning operation can identify and remove unnecessary elements that may interfere with subsequent analysis, such as HTML tags in web pages, various non-textual special symbols, and redundant spaces and line breaks caused by typesetting or data entry errors. The aim is to eliminate text noise, improve text purity and readability, and provide high-quality input for subsequent natural language processing tasks.
[0048] Then, the initial medical text information after text cleaning is subjected to case normalization. All characters in the initial medical text information can be uniformly converted to uppercase or lowercase. For example, all English characters can be uniformly converted to lowercase to ensure that "Epilepsy" and "epilepsy" are recognized as the same word. The purpose is to eliminate inconsistencies in vocabulary caused by case differences, thereby reducing redundant word representations and improving the accuracy of term matching and information extraction.
[0049] Then, sentence boundary analysis is performed on the initial medical text information after case normalization to obtain non-standard medical text information. This case normalization process can be further processed to accurately identify sentence boundaries within the text. Specifically, the start and end positions of each sentence can be determined by recognizing punctuation marks such as periods, question marks, and exclamation marks, as well as other linguistic rules. The aim is to segment the continuous text stream into independent, meaningful sentence units, which is crucial for subsequent tasks such as word segmentation, entity recognition, and relation extraction, ensuring accurate understanding and processing of medical information.
[0050] This embodiment presents a systematic preprocessing workflow by sequentially performing text cleaning, case normalization, and sentence boundary analysis on the initial medical text information. Text cleaning effectively removes irrelevant noise and formatting information from the original data, ensuring the purity of the text content. Secondly, case normalization unifies words with different capitalization forms in the text, greatly improving the consistency of word recognition and avoiding misjudgments or omissions caused by capitalization differences. Finally, sentence boundary analysis precisely segments the cleaned and normalized text into independent sentences, providing clear contextual units for subsequent fine-grained language analysis (such as word segmentation and entity recognition). This embodiment enables the effective transformation of raw and complex medical text information into more clearly structured and higher-quality non-standard medical text information, laying a solid foundation for subsequent terminology conversion and medical information extraction.
[0051] Through the above technical solutions, this embodiment can significantly improve the quality and usability of initial medical text information. Text cleaning effectively reduces data noise, allowing subsequent processing to focus on core medical content. Case normalization ensures the accuracy and consistency of medical terminology recognition, reducing errors caused by inconsistent formatting. Sentence boundary analysis provides structured input for accurate semantic understanding and information extraction, avoiding the complexity and potential errors of cross-sentence analysis. The implementation of these preprocessing steps makes extracting medical information related to the mechanism of action of roselic acid C molecules and epilepsy diagnosis and treatment from massive amounts of medical text more efficient and reliable, thus providing a solid data foundation for constructing a high-quality epilepsy medical knowledge graph.
[0052] In some embodiments, in step S103, the non-standard medical text information is converted into standardized medical text information based on a medical terminology relation knowledge base. This may include, but is not limited to, the following steps:
[0053] Step S201: Segment the non-standard medical text information into words and identify multiple non-standard medical terms;
[0054] Step S202: Perform precise matching between non-standard medical terms and the medical terminology relationship knowledge base to obtain precise matching results and precise matching values;
[0055] Step S203: If the exact match result is a successful match, then the exact match value is used as a standardized medical term.
[0056] Step S204: If the exact matching result is a failure, perform a fuzzy match between the non-standard medical terms and the medical terminology relationship knowledge base to identify the standardized medical terms.
[0057] Step S205: Based on standardized medical terminology, replace the words in the non-standard medical text information to obtain standardized medical text information.
[0058] In some embodiments, since non-standard medical text information may contain a large number of synonyms, near-synonyms, abbreviations, spelling errors, or non-standard terms, if only a simple matching method is used for term conversion, a large number of valuable non-standard medical terms may not be accurately identified and standardized, thereby affecting the completeness and accuracy of subsequent medical information extraction.
[0059] To address this, word segmentation can be performed on non-standard medical text information to identify multiple non-standard medical terms. Continuous non-standard medical text information can be divided into word units with independent semantics to facilitate subsequent term identification and matching. For example, dictionary-based methods, statistical model-based methods, or hybrid methods can be used for word segmentation. Through word segmentation, multiple potential non-standard medical terms can be identified from non-standard medical text information.
[0060] Then, a precise match is performed between the non-standard medical terms and the medical terminology relationship knowledge base to obtain precise match results and precise match values. The purpose is to quickly and accurately identify non-standard terms that are completely consistent with the standard terms in the knowledge base. The precise match result can be a Boolean value (success or failure), while the precise match value is the corresponding standardized term in the knowledge base when a match is successful.
[0061] If an exact match is successful, the exact match value is used as the standardized medical term. This ensures efficient standardization of standardized terms. If an exact match fails, it indicates that the non-standard medical term does not have a completely identical counterpart in the knowledge base. In this case, to improve the coverage and robustness of terminology conversion, fuzzy matching is needed between the non-standard medical term and the medical terminology relationship knowledge base to identify standardized medical terms. Fuzzy matching can employ various techniques, such as methods based on edit distance (Levenshtein distance), Jaccard similarity, TF-IDF similarity, or word vector similarity, aiming to find standard terms that are semantically or spelledly close to the non-standard term.
[0062] Finally, based on standardized medical terminology, the non-standard medical text information is replaced with its corresponding standardized medical terms to obtain standardized medical text information. This means replacing all identified non-standard medical terms in the text with their corresponding standardized medical terms, thereby obtaining the final standardized medical text information.
[0063] This embodiment effectively addresses the limitations of processing complex and varied non-standard medical text information by introducing a strategy combining word segmentation, exact matching, and fuzzy matching. Word segmentation breaks down continuous text into independent terminology units, laying the foundation for subsequent matching operations. Secondly, the exact matching mechanism efficiently handles standardized expressions that are completely consistent with standard terms in the knowledge base, ensuring the accuracy and efficiency of the conversion. When exact matching fails, the introduced fuzzy matching mechanism identifies standard terms that are semantically or spelledly similar to the non-standard terms by calculating similarity, thus effectively handling non-standard expressions such as synonyms, near-synonyms, abbreviations, and spelling errors. Therefore, by replacing the identified standardized medical terms back into the original text, comprehensive, accurate, and robust standardization processing of non-standard medical text information is ultimately achieved.
[0064] To illustrate this technical solution more clearly, a specific example is used below. Suppose there is a non-standard medical text: "The patient experienced an epileptic seizure, accompanied by myoclonus and absence." First, this non-standard medical text is segmented, identifying multiple non-standard medical terms such as "epilepsy seizure," "seizure," "myoclonus," "absence," and "absence." Next, these non-standard medical terms are precisely matched against a medical terminology knowledge base. For example, "epilepsy seizure" might precisely match "epilepsy seizure" in the knowledge base, resulting in a successful match with the value "epilepsy seizure." Similarly, "myoclonus" precisely matches "myoclonus," and "absence" precisely matches "absence." However, "seizure" and "myoclonus" might not have completely identical Chinese equivalents in the knowledge base, or their English forms might not be directly included as standard terms; in this case, the precise match fails. For "seizure," which fails the precise match, the system performs a fuzzy match against the medical terminology knowledge base. By calculating similarity, the system identified a high similarity between "seizure" and "epilepsy seizure" in the knowledge graph, thus recognizing "epilepsy seizure" as a standardized medical term. Similarly, for "myoclonus," the system identified "myoclonus" as its standardized medical term through fuzzy matching. Finally, based on the identified standardized medical terms, the original non-standard medical text information was replaced with words. For example, "seizure" was replaced with "epilepsy seizure," "myoclonus" with "myoclonus," and "absence" with "absence." Thus, the original text was standardized to: "The patient experienced an epileptic seizure, accompanied by myoclonus and absence." In this way, even if there are multiple expressions in the original text, they can be effectively standardized, ensuring the accuracy and consistency of medical information.
[0065] Through the above technical solution, this embodiment can significantly improve the accuracy and coverage of medical text terminology conversion. The combination of exact matching and fuzzy matching enables the system to not only efficiently process standard terms, but also effectively identify and standardize non-standard medical terms that cannot be identified by exact matching due to non-standard expression, variations, or spelling errors. This avoids information loss caused by incomplete terminology recognition, ensuring that the medical information extracted from the original medical text is more complete and accurate, thus providing a high-quality standardized data foundation for subsequent knowledge graph construction, and greatly enhancing the practical value and reliability of the knowledge graph.
[0066] In some embodiments, step S204, performing fuzzy matching between non-standard medical terms and the medical terminology relationship knowledge base to identify standardized medical terms, may include, but is not limited to, the following steps:
[0067] Step S301: Extract initial contextual information of non-standard medical terms from non-standard medical text information. The initial contextual information includes adjacent keywords, syntactic structure, and semantic category.
[0068] Step S302: Generate a contextual semantic fingerprint based on the initial context information;
[0069] Step S303: Calculate the similarity between the contextual semantic fingerprint and each matching key in the medical terminology relationship knowledge base using the cosine similarity algorithm;
[0070] Step S304: Based on the medical terminology relationship knowledge base, the fuzzy matching value corresponding to the matching key with the highest similarity is taken as the standardized medical term.
[0071] In some embodiments, initial contextual information of non-standard medical terms can be extracted from non-standard medical text information. Initial contextual information refers to the contextual information surrounding the non-standard medical term, aiming to capture the semantics of the term in a specific textual environment. This initial contextual information includes neighboring keywords, syntactic structure, and semantic category. Neighboring keywords are words that appear adjacent to the non-standard medical term in the text; they provide direct local semantic cues. Syntactic structure refers to the grammatical role of the non-standard medical term in a sentence and its syntactic relationships with other words, such as subject-verb-object structures, which helps in understanding the function and meaning of the term. Semantic category refers to the macro-conceptual category to which the non-standard medical term belongs, such as disease name, drug ingredient, treatment method, or symptom description, aiming to provide a high-level classification basis for subsequent semantic analysis.
[0072] Then, based on the initial context information, a contextual semantic fingerprint is generated. A contextual semantic fingerprint refers to the vectorized or numerical representation of non-standard medical terms and their contexts. For example, word embedding models (such as Word2Vec, GloVe, or BERT) encode neighboring keywords, syntactic structures, and semantic categories into high-dimensional vectors, thereby representing the meaning of the non-standard medical term in semantic space. This contextual semantic fingerprint can capture the deep semantic features of non-standard medical terms, and can match them through semantic similarity even if their surface form differs from standardized terms.
[0073] The cosine similarity algorithm is then used to calculate the similarity between the contextual semantic fingerprint and each matching key in the medical terminology knowledge base. Cosine similarity is a commonly used method to measure the similarity between two non-zero vectors; its value ranges from -1 to 1, with higher values indicating higher similarity. The medical terminology knowledge base stores a large number of standardized medical terms and their corresponding semantic fingerprints or feature representations, with each standardized term serving as a matching key. By calculating the cosine similarity between the contextual semantic fingerprint of a non-standard medical term and the semantic fingerprints of all matching keys in the knowledge base, their semantic proximity can be quantified.
[0074] Finally, based on the medical terminology relation knowledge base, the fuzzy matching value corresponding to the matching key with the highest similarity is taken as the standardized medical term. This means that among all standardized terms compared with non-standard medical terms, the one that is semantically closest is selected as its standardized form. The fuzzy matching value is the standardized medical term associated with the matching key with the highest similarity in the knowledge base.
[0075] This embodiment extracts the initial contextual information of non-standard medical terms and transforms it into contextual semantic fingerprints, thereby enabling a semantic understanding of the true meaning of the non-standard terms. This semantic understanding allows for the identification of the semantically closest standardized term even if the non-standard term and the standardized term in the knowledge base are not literally identical, by calculating the cosine similarity between the contextual semantic fingerprint and the matching key in the knowledge base. This fuzzy matching mechanism based on contextual semantics effectively compensates for the shortcomings of exact matching, improving the accuracy and robustness of terminology conversion.
[0076] Through the above technical solution, this embodiment can effectively handle problems such as terminology variations, synonyms, spelling errors, or non-standard expressions in medical texts, significantly improving the accuracy and coverage of converting non-standard medical text information to standardized medical text information. Therefore, it can ensure that even when faced with complex and varied medical text data, the correct standardized medical terms can be identified, thus laying a solid foundation for subsequent medical information extraction and knowledge graph construction, and improving the quality and completeness of the knowledge graph.
[0077] In some embodiments, in step S302, generating a contextual semantic fingerprint based on the initial context information may include, but is not limited to, the following steps:
[0078] Step S401: Perform structured analysis on non-standard medical text information to determine the region to which the non-standard medical terms belong;
[0079] Step S402: If the region to which it belongs is a region with sparse contextual information, then according to the semantic category, the non-standard medical text information is re-searched to extract supplementary contextual information;
[0080] Step S403: Integrate the initial context information and the supplementary context information to obtain the target context information;
[0081] Step S404: Generate a contextual semantic fingerprint based on the target context information.
[0082] In some embodiments, the initial contextual information extracted from non-standard medical text may sometimes suffer from sparsity. When the contextual information of non-standard medical terms is insufficient, the contextual semantic fingerprint generated based solely on limited initial contextual information may not adequately capture the complete semantic connotation of the term, thus affecting the accuracy and recall rate of fuzzy matching. If this problem is not addressed, some non-standard medical terms may not be accurately standardized, thereby affecting the comprehensiveness and accuracy of knowledge graph construction.
[0083] To address this, a structured analysis of non-standard medical text information can be performed first to determine the regions to which non-standard medical terms belong. This aims to identify the location of non-standard medical terms within the text and the richness of their surrounding information. For example, the sentence length, paragraph structure, and the presence of related definitions or explanatory statements can be analyzed to determine whether the region to which the term belongs is a region with sparse contextual information. A region with sparse contextual information refers to a region where the initial contextual information, such as adjacent keywords, syntactic structures, or semantic categories, is insufficient to fully and accurately represent the semantics of the non-standard medical term.
[0084] If the region is identified as having sparse contextual information, the non-standard medical text information is re-searched based on semantic category to extract supplementary contextual information. When a region is identified as having sparse contextual information, to compensate for the lack of initial contextual information, the non-standard medical text information can be re-searched based on the semantic category of the non-standard medical term. For example, if the semantic category of the non-standard medical term is "disease name," other text fragments related to the disease name can be retrieved from the entire text database based on this category to extract richer and more comprehensive supplementary contextual information. This supplementary contextual information can include descriptions of the disease's symptoms, treatment plans, related complications, etc., with the aim of obtaining semantic clues related to the term from a broader context.
[0085] The initial and supplementary contextual information are then integrated to form a more complete and accurate target contextual information. This aims to fuse contextual information from different sources, ensuring that the target contextual information more comprehensively reflects the semantics of non-standard medical terms. For example, techniques such as weighted averaging, feature splicing, or semantic fusion can be used to effectively combine the two types of contextual information. Based on the integrated target contextual information, a more accurate and robust contextual semantic fingerprint is generated.
[0086] To illustrate this technical solution more clearly, a specific example is used below. Suppose that when processing non-standard medical text about epilepsy, a non-standard medical term, such as "petit mal seizure," is encountered. During initial contextual information extraction, it is found that this term only appears in a short title or list, and its information such as adjacent keywords, syntactic structure, and semantic category is very limited, identified by structured analysis as a sparse region of contextual information. At this point, the system will re-search the entire non-standard medical text database based on the semantic category of "petit mal seizure" (e.g., type of epileptic seizure). Through this re-search, the system may find other paragraphs describing the symptoms, diagnostic criteria, or treatment methods of "petit mal seizure," thereby extracting supplementary contextual information such as "brief loss of consciousness," "gaze," and "automatisms." Subsequently, this supplementary contextual information is integrated with the initially sparse contextual information to form a target contextual information containing richer semantic cues. For example, the integrated target contextual information may include content such as "a petit mal seizure is a brief disturbance of consciousness, often manifested as a gaze, blinking, or mild automatisms, lasting from several seconds to tens of seconds." Ultimately, based on this comprehensive and accurate target context information, a contextual semantic fingerprint of "petit mal seizures" is generated. This fingerprint is more accurate than that generated solely based on sparse initial context information, thus enabling more accurate identification and standardization of it into standard terms such as "absence seizures" or "typical absence seizures" when performing fuzzy matching with a medical terminology knowledge base.
[0087] Through the above technical solution, this embodiment effectively addresses the challenge of sparse contextual information in non-standard medical terminology within medical texts. By intelligently identifying sparse regions of contextual information and performing targeted supplementary contextual information re-retrieval, this embodiment significantly enhances the generation quality and robustness of contextual semantic fingerprints. Therefore, even when faced with complex and varied medical texts, it ensures that the semantics of non-standard medical terms are understood and represented more comprehensively and accurately, thereby greatly improving the success rate and accuracy of fuzzy matching, ultimately promoting the construction quality and application value of the epilepsy medical knowledge graph based on rosehip C.
[0088] In some embodiments, step S403, integrating the initial context information and the supplementary context information to obtain the target context information, may include, but is not limited to, the following steps:
[0089] Step S501: Compare the initial context information and the supplementary context information to identify semantic conflict information;
[0090] Step S502: Perform fine-grained semantic analysis on the semantic conflict information to obtain fine-grained semantic information;
[0091] Step S503: Send the fine-grained semantic information to the expert adjudication terminal. The expert adjudication terminal is used to adjudicate the fine-grained semantic information and generate contextual information correction suggestions.
[0092] Step S504: Based on the context information correction suggestions, correct the initial context information and supplementary context information;
[0093] Step S505: Integrate the corrected initial context information and supplementary context information to obtain the target context information.
[0094] In some embodiments, simply integrating initial and supplementary contextual information may lead to semantic conflicts or inconsistencies due to differences in information sources, expression methods, or contexts. If these issues are not addressed, the generated target contextual information may fail to accurately reflect the true context of non-standard medical terms, thus affecting the accuracy of contextual semantic fingerprints and ultimately reducing the precision of standardized medical terminology recognition.
[0095] To address this, the initial and supplementary contextual information can be compared to identify semantically conflicting information. Cross-checking and comparing the two types of contextual information can reveal inconsistencies or contradictions. Semantically conflicting information refers to different or even contradictory descriptions or attributes of the same concept or entity in two different contextual pieces of information. For example, the initial contextual information might describe a drug's side effect as "mild nausea," while the supplementary contextual information describes it as "severe vomiting," thus constituting a semantic conflict. The purpose of identification is to enable targeted processing subsequently, ensuring the accuracy of the final integrated information.
[0096] Then, fine-grained semantic analysis is performed on the semantic conflict information to obtain fine-grained semantic information. This allows for in-depth analysis of the identified semantic conflicts, going beyond mere surface-level textual comparison to delve into their deeper meanings, contextual relationships, and potential semantic differences, thus yielding fine-grained semantic information. The aim is to precisely pinpoint the root cause and specific manifestation of the conflict; for example, analyzing whether the conflict stems from different expressions of synonyms, ambiguity of polysemous words, contextual dependence, or factual errors. Through this analysis, more detailed and precise fine-grained semantic information can be obtained, providing a sufficient basis for subsequent adjudication and correction.
[0097] The fine-grained semantic information is then sent to an expert adjudication terminal, which adjudicates the information and generates contextual information correction suggestions. This terminal can be a team of medical experts or a decision support system with specialized knowledge. Its role is to professionally judge and evaluate the semantically conflicting information after fine-grained analysis, providing authoritative opinions or correction schemes based on its medical knowledge and experience. For example, experts can determine which side's information is more accurate, or how to synthesize the two sides to form a more comprehensive description. The contextual information correction suggestions are the result of the expert adjudication, guiding subsequent adjustments to the initial and supplementary contextual information to eliminate conflicts and improve information quality.
[0098] Based on the contextual information correction suggestions, the initial and supplementary contextual information are revised. Conflicting parts of the original initial and supplementary contextual information can be modified, supplemented, or deleted based on the contextual information correction suggestions generated by the expert adjudication terminal. For example, if the expert recommends adopting a description from the supplementary contextual information, the corresponding conflicting part in the initial contextual information will be replaced or adjusted. The aim is to eliminate identified semantic conflicts, ensuring semantic consistency and accuracy between the two types of contextual information, thus preparing for final integration.
[0099] Finally, the revised initial context information and supplementary context information are integrated to obtain the target context information. Based on this, the revised initial and supplementary context information are merged again to form the final target context information. Because potential semantic conflicts have been resolved through comparison, fine-grained analysis, expert adjudication, and revision before integration, the resulting target context information will be high-quality, conflict-free, and semantically consistent, more accurately representing the context of non-standard medical terminology.
[0100] To illustrate this technical solution more clearly, a specific example is used below. Suppose that when processing non-standard medical terminology regarding "epileptic seizures," the initial contextual information extracted from the initial medical text is described as "the patient experienced mild convulsions lasting approximately 30 seconds," while the supplementary contextual information obtained through re-retrieval describes it as "the patient presented with a generalized tonic-clonic seizure accompanied by loss of consciousness." First, the system compares these two sets of contextual information, immediately identifying a semantic conflict between "mild convulsions" and "generalized tonic-clonic seizure," as well as inconsistencies between "lasting approximately 30 seconds" and "accompanied by loss of consciousness" in describing the severity and characteristics of the seizure. Next, fine-grained semantic analysis is performed on these semantically conflicting pieces of information. The analysis results may show that "mild convulsions" may refer to a focal seizure, while "generalized tonic-clonic seizure" refers to a more severe generalized seizure, indicating a significant difference in medical classification. Simultaneously, the combination of "lasting approximately 30 seconds" and "accompanied by loss of consciousness" also suggests a different seizure type. This yields fine-grained semantic information about the type, severity, and accompanying symptoms of the attack.
[0101] Subsequently, this fine-grained semantic information is sent to an expert adjudication terminal. Medical experts, based on their expertise, determine whether the description "generalized tonic-clonic seizure with loss of consciousness" is more accurate or clinically significant, or whether the initial contextual information may be misdiagnosed or incomplete. The expert adjudication terminal then generates contextual information correction suggestions, such as recommending the adoption of the core description from supplementary contextual information and supplementing or correcting the initial contextual information. Based on the expert adjudication terminal's contextual information correction suggestions, the initial contextual information is corrected to ensure semantic consistency with the supplementary contextual information; for example, "mild convulsions" is corrected to a more detailed description of "generalized tonic-clonic seizure," and the feature "with loss of consciousness" is added. Finally, the corrected initial contextual information and supplementary contextual information are integrated to obtain the final target contextual information, such as "the patient presented with a generalized tonic-clonic seizure with loss of consciousness, lasting approximately 30 seconds." This embodiment ensures that the final integrated target contextual information is accurate, consistent, and of high quality, thus providing a reliable foundation for subsequently generating accurate contextual semantic fingerprints.
[0102] Through the above technical solution, this embodiment effectively avoids semantic conflicts and inconsistencies arising from simply integrating contextual information from different sources. By introducing semantic conflict identification, fine-grained semantic analysis, and the intervention of expert adjudication terminals, it ensures that all potential semantic contradictions can be discovered, analyzed in depth, and professionally corrected before generating target contextual information. This significantly improves the quality and accuracy of the target contextual information, enabling it to more accurately reflect the real context of non-standard medical terms. This embodiment can effectively handle complex and ever-changing medical text information, thereby providing a more solid and accurate data foundation for subsequent contextual semantic fingerprint generation and standardized medical terminology recognition, ultimately improving the accuracy and practicality of the entire knowledge graph construction method.
[0103] In some embodiments, step S502, performing fine-grained semantic analysis on the semantic conflict information to obtain fine-grained semantic information, may include, but is not limited to, the following steps:
[0104] Step S601: Update the semantic conflict information according to the preset semantic mapping rule set;
[0105] Step S602: Identify multiple core entities related to the mechanism of action of rose acid C molecules and epilepsy diagnosis and treatment from the updated semantic conflict information;
[0106] Step S603: Analyze the entity relationships between multiple core entities;
[0107] Step S604: Based on entity associations, construct a preliminary association network, which contains multiple entity pairs;
[0108] Step S605: Based on the preset association rule set, identify the association paths in the preliminary association network to obtain the implicit association paths between entity pairs;
[0109] Step S606: Extract implicit association information from implicit association paths;
[0110] Step S607: Update the preliminary association network based on the implicit association information to obtain the enhanced association network;
[0111] Step S608: Based on the enhanced association network, perform fine-grained semantic analysis on the updated semantic conflict information to obtain fine-grained semantic information.
[0112] In some embodiments, since semantic conflicts may involve multi-layered entity relationships and implicit associations, it may be difficult to fully and accurately reveal the nature of these deep conflicts if only simple fine-grained semantic analysis is performed, thereby affecting the accuracy of subsequent expert decisions and the effectiveness of contextual information correction.
[0113] To address this, semantic conflict information can be updated first based on a predefined set of semantic mapping rules. This set refers to a predefined set of rules used to standardize and unify the semantics of medical terminology. These rules aim to resolve semantic inconsistencies in medical terms from different sources or with different expressions, for example, mapping synonyms, near-synonyms, or terms with hierarchical relationships to a unified, standardized representation. By applying these rules, semantic conflict information can be initially cleaned and standardized, laying the foundation for subsequent entity identification and association analysis. Several core entities related to the mechanism of action of rosmarinic acid C and the diagnosis and treatment of epilepsy can then be identified from the updated semantic conflict information. Core entities refer to medical concepts or objects that are crucial to the semantic conflict information, especially those directly related to the mechanism of action of rosmarinic acid C and the diagnosis and treatment of epilepsy. For example, these may include drug names (such as rosmarinic acid C), disease names (such as epilepsy), targets (such as receptors, enzymes), biological processes, symptoms, and treatment methods. These core entities form the basis for constructing the association network.
[0114] Then, the entity relationships between multiple core entities are analyzed. Entity relationships refer to various semantic relationships between core entities, such as "acts on," "causes," "inhibits," "treats," "is part of," etc. These relationships can be extracted from the text using natural language processing techniques, such as dependency parsing and semantic role labeling. Based on the entity relationships, a preliminary relationship network is constructed, which contains multiple entity pairs. The preliminary relationship network is a graph structure consisting of core entities as nodes and entity relationships as edges, where each entity pair represents a direct or indirect relationship between two core entities.
[0115] Next, based on a pre-defined set of association rules, the preliminary association network is used to identify association paths, revealing implicit association paths between entity pairs. The pre-defined set of association rules refers to a set of rules used to guide association path identification; these rules define possible valid association patterns or path types between different entity types. For example, it can be stipulated that there might be a "treatment" relationship between the "drug" entity and the "disease" entity, while there might be an "acting on" relationship between the "drug" entity and the "target" entity. Through path search algorithms, combined with these rules, implicit association paths between entity pairs that are not directly expressed but logically exist can be discovered in the preliminary association network.
[0116] Finally, implicit association information is extracted from the implicit association paths. This implicit association information reveals deep semantic relationships. Based on this implicit association information, the preliminary association network is updated to obtain an enhanced association network. The enhanced association network not only includes direct entity associations but also incorporates implicit associations obtained through rule-based reasoning, thus providing a more comprehensive and accurate view of entity relationships. Based on the enhanced association network, fine-grained semantic analysis is performed on the updated semantic conflict information to obtain fine-grained semantic information. This allows for a deeper understanding of the root causes and nature of conflicts, leading to more precise and valuable fine-grained semantic information.
[0117] To illustrate this technical solution more clearly, a specific example is used below. Suppose that during the integration of initial and supplementary contextual information, a semantically conflicting piece of information is identified, such as: "Rose acid C has an inhibitory effect on epileptic seizures, but its specific mechanism is still unclear; some studies suggest it may be achieved by regulating GABA receptor activity" and "GABA receptors play a key role in the regulation of neuronal excitability and are associated with various neurological diseases, including epilepsy." First, according to a pre-defined semantic mapping rule set, this semantically conflicting information is updated; for example, "epilepsy seizures" and "epilepsy" are uniformly mapped to "epilepsy." Next, core entities are identified from the updated semantically conflicting information, such as: "rose acid C," "epilepsy," "GABA receptor," and "regulation of neuronal excitability." Then, the entity relationships between these core entities are analyzed; for example, there is an association of "inhibitory effect" between "rose acid C" and "epilepsy," an association of "key role" between "GABA receptor" and "regulation of neuronal excitability," and an association of "correlation" between "regulation of neuronal excitability" and "epilepsy." A preliminary association network is thus constructed.
[0118] Furthermore, based on a pre-defined set of association rules, the preliminary association network is used to identify association paths. For example, through a path search algorithm, a latent association path can be discovered between "rosolic acid C" and "GABA receptor": "rosolic acid C" -> "inhibitory effect" -> "epilepsy" -> "related" -> "regulation of neuronal excitability" -> "key role" -> "GABA receptor", or a more direct path: "rosolic acid C" -> "regulation" -> "GABA receptor" -> "influence" -> "epilepsy". Latent association information is extracted from these paths, such as "rosolic acid C may affect epilepsy by regulating GABA receptor activity". Based on this latent association information, the preliminary association network is updated to obtain an enhanced association network. This enhanced association network not only includes direct textual associations but also the latent mechanism obtained through reasoning: "rosolic acid C affects epilepsy by regulating GABA receptor activity". Finally, based on this enhanced association network, fine-grained semantic analysis of the updated semantic conflict information can yield more precise fine-grained semantic information, such as: "Rose acid C inhibits epilepsy by regulating GABA receptor activity, thereby affecting neuronal excitability regulation; the specific mechanism requires further investigation." This fine-grained semantic information can more clearly reveal the underlying causes of semantic conflicts and the complex mechanisms of interaction between entities, providing a more comprehensive perspective for expert adjudication.
[0119] Through the above technical solution, this embodiment can significantly improve the depth and accuracy of fine-grained semantic analysis of semantic conflict information. By introducing a series of steps, including semantic mapping rules, core entity identification, association network construction, and implicit association path discovery, this embodiment can more comprehensively capture the complex entity relationships and potential semantic connections in medical texts. Therefore, the obtained fine-grained semantic information is not only more accurate but also reveals the underlying causes of semantic conflicts, thus providing a more reliable and valuable basis for subsequent expert adjudication and contextual information correction. This is crucial for constructing a high-quality epilepsy medical knowledge graph based on rosehip C, effectively reducing errors or information omissions in the knowledge graph due to inaccurate semantic understanding, thereby improving the practicality and reliability of the knowledge graph.
[0120] In some embodiments, in step S601, updating the semantic conflict information according to a preset semantic mapping rule set may include, but is not limited to, the following steps:
[0121] By classifying semantic conflict information, we obtain experimental model information, observation index information, and terminology system information;
[0122] Based on the preset semantic mapping rule set, the experimental model information, observation index information and terminology system information are semantically normalized.
[0123] Based on the semantically normalized experimental model information, observation index information, and terminology system information, update the semantic conflict information.
[0124] In some embodiments, if the update process lacks refined classification and normalization, the updated semantic conflict information may still exhibit heterogeneity or inconsistency, thereby affecting the accuracy and reliability of subsequent core entity identification, entity association analysis, and preliminary association network construction. If these issues are not addressed, the constructed knowledge graph may suffer from semantic bias, reducing its application value in the field of epilepsy treatment.
[0125] To address this, semantic conflict information can be categorized into experimental model information, observational indicator information, and terminology information. Experimental model information refers to conflicting or inconsistent descriptions related to experimental design, animal models, and cell lines; observational indicator information refers to conflicts related to quantitative or qualitative indicators such as disease symptoms, biomarkers, and treatment effects; and terminology information involves differences in the expression of medical terms, disease names, and drug names across different sources or contexts. The purpose of this categorization is to employ more targeted processing strategies for different types of semantic conflicts.
[0126] Then, based on a pre-defined semantic mapping rule set, the experimental model information, observation index information, and terminology system information are semantically normalized. Semantic normalization refers to mapping entities or concepts with different expressions but the same semantics to a unified standard representation. For example, for experimental model information, "rat epilepsy model" and "SD rat epilepsy induction model" can be normalized to "rat epilepsy model"; for observation index information, "seizure frequency" and "number of epileptic seizures" can be normalized to "epilepsy seizure frequency," and the values in different units can be unified to a standard unit; for terminology system information, "epilepsy" and "epilepsy" can be normalized to "epilepsy." The pre-defined semantic mapping rule set can be defined in advance by domain experts or learned from a large corpus through machine learning methods. Its purpose is to eliminate semantic heterogeneity and ensure the standardization and consistency of information.
[0127] Then, based on the experimental model information, observation index information, and terminology system information after semantic normalization, the semantic conflict information is updated. This means that the more accurate and consistent information after classification and normalization will replace or supplement the original semantic conflict information, thereby providing high-quality input for subsequent fine-grained semantic analysis.
[0128] To illustrate the technical solution more clearly, a specific example is used below. Suppose that when processing medical texts on epilepsy, the system identifies the following semantic conflict information: (1) Experimental model information conflict: The first document describes "Wistar rat epilepsy model", while the second document describes "Wistar rat hippocampal sclerosis-induced epilepsy model". (2) Observational indicator information conflict: The first document records "the frequency of seizures is 5 times / hour", while the second document records "the number of epileptic seizures is 5 times per hour". (3) Terminology information conflict: The first document uses "epilepsy", while the second document uses "epilepsy". This embodiment first classifies these semantic conflict information. "Wistar rat epilepsy model" and "Wistar rat hippocampal sclerosis-induced epilepsy model" are classified as experimental model information. "The frequency of seizures is 5 times / hour" and "the number of epileptic seizures is 5 times per hour" are classified as observational indicator information. "epilepsy" and "epilepsy" are classified as terminology information.
[0129] Subsequently, these categorized information are semantically normalized according to a pre-defined set of semantic mapping rules. For experimental model information, the pre-defined rules may normalize "Wistar rat hippocampal sclerosis-induced epilepsy model" to the more general "Wistar rat epilepsy model" to unify the description of the experimental model. For observational indicator information, the pre-defined rules normalize "5 seizures per hour" to "seizure frequency of 5 times / hour" to ensure consistency in indicator description. For terminology information, the pre-defined rules normalize "epilepsy" to the standard medical term "epilepsy". Finally, based on this semantically normalized information, the original semantic conflict information is updated. The updated semantic conflict information will be more standardized and consistent. This refined semantic conflict information will greatly improve the accuracy of subsequent identification of core entities (such as "rosolic acid C", "epilepsy", "Wistar rat") and analysis of entity associations (such as the influence of "rosolic acid C" on the "seizure frequency" of the "Wistar rat epilepsy model"), thereby constructing a high-quality knowledge graph.
[0130] Through the above technical solutions, this embodiment can significantly improve the precision and accuracy of semantic conflict information processing. Specifically, by classifying semantic conflict information, the heterogeneity problem of different types of conflicts can be addressed in a targeted manner, avoiding blind information processing. Furthermore, semantic normalization processing ensures the consistency of all relevant information at the semantic level, effectively reducing misunderstandings and errors caused by differences in expression, thereby providing high-quality, unambiguous data input for subsequent entity recognition and association analysis. Thus, this embodiment can construct a more accurate and reliable preliminary association network and enhanced association network, ultimately forming a high-quality epilepsy medical knowledge graph based on roselic acid C, whose application value in clinical decision support, drug development, and other fields will be significantly enhanced.
[0131] In some embodiments, step S605 involves identifying association paths in the preliminary association network based on a preset association rule set to obtain implicit association paths between entity pairs. This may include, but is not limited to, the following steps:
[0132] Multiple core entities are classified to obtain various entity types;
[0133] Based on multiple entity types and a preset set of association rules, determine the conversion cost corresponding to each entity type;
[0134] Adjust the priority and direction of path search based on various entity types and transformation costs;
[0135] Based on the preset set of association rules, the priority and direction of path search, implicit association paths are found through path search algorithms.
[0136] In some embodiments, since the initial association network may contain various types and properties of core entities, indiscriminate path identification may lead to inefficient path search or insufficient accuracy and relevance of identified implicit association paths, thereby affecting the quality of fine-grained semantic analysis. Without addressing these issues, it will be difficult to efficiently and accurately reveal deep associations within complex medical information.
[0137] To this end, multiple core entities can be classified first, resulting in various entity types. These core entities can be categorized into different entity types based on their inherent attributes and roles in the medical field. For example, core entities can be classified as drug entities, disease entities, gene entities, symptom entities, protein entities, etc. The purpose is to distinguish the semantic characteristics and connection patterns of different entities in the network, providing a refined foundation for subsequent path searching.
[0138] Then, based on multiple entity types and a pre-defined set of association rules, the conversion cost for each entity type is determined. Different "costs" or "weights" can be assigned to conversions or traversals between different entity types. For example, the conversion cost from one "drug" entity to another "disease" entity may differ from the conversion cost from one "symptom" entity to another "gene" entity. The conversion cost can be set based on expert knowledge, statistical analysis, or machine learning models, reflecting the strength, reliability, or importance of the association between different entity types. The purpose is to guide the algorithm to prioritize exploring paths with higher semantic relevance or greater medical significance during path search, avoiding meaningless path exploration.
[0139] Then, based on various entity types and transformation costs, the priority and direction of path search are adjusted. When executing the path search algorithm, the search strategy can be dynamically adjusted according to the current entity type, the target entity type, and the transformation costs between them. For example, for entity type transformations with lower transformation costs, their priority in the search queue can be increased; for certain combinations of entity types, the search direction can be restricted or guided to better align with medical logic. The aim is to optimize the search process, making it more efficient in focusing on valuable related paths and reducing unnecessary computational overhead.
[0140] Finally, based on the preset association rule set, the priority and direction of path search, implicit association paths are found using a path search algorithm. Graph traversal algorithms (such as Depth-First Search (DFS), Breadth-First Search (BFS), A*) can be used. Algorithms (such as Dijkstra's algorithm) search for paths connecting entity pairs in an initial association network. During the search process, the algorithm comprehensively considers a pre-set set of association rules (e.g., certain entities must be connected through specific relationships), adjusted path search priorities (e.g., prioritizing the exploration of drug-target-disease pathways), and direction (e.g., unidirectional influence from drug to disease), thereby efficiently and accurately identifying potential, indirect association paths between entity pairs.
[0141] To illustrate this technical solution more clearly, a specific example is used below. Assume that in the initial association network, there are core entities "rosolic acid C," "epilepsy," "GABA receptor," and "neuron." These entities are initially classified as: drug (rosolic acid C), disease (epilepsy), protein (GABA receptor), and cell (neuron). Based on a pre-defined set of association rules, the conversion cost between different entity types can be determined. For example, the conversion cost from "drug" to "protein" may be low (indicating a direct effect), the conversion cost from "protein" to "disease" is also low (indicating a mechanism association), while the conversion cost from "drug" directly to "cell" may be high (indicating an indirect effect). When searching for implicit association paths between "rosolic acid C" and "epilepsy," the path search algorithm adjusts its priority and direction based on these conversion costs. For example, the algorithm will prioritize exploring paths like "rosolic acid C" → "GABA receptor" → "epilepsy" because its conversion cost is low and conforms to known pharmacological mechanisms. For paths like "rosolic acid C" → "neuron" → "epilepsy," if the transformation cost is high or does not conform to preset rules, their search priority will be reduced, or they may even be excluded. In this way, path search algorithms (e.g., using A...) The algorithm can efficiently find hidden pathways connecting "rosolic acid C" and "epilepsy" in the initial association network. For example, "rosolic acid C" inhibits "GABA receptor" activity, thereby alleviating "epilepsy" seizures. This method ensures that the identified pathways not only exist but are also medically plausible and highly relevant.
[0142] Through the above technical solution, this embodiment can significantly improve the efficiency and accuracy of identifying implicit association paths between entity pairs in the initial association network. By classifying core entities and introducing transformation costs, as well as adjusting the priority and direction of path search, the path search algorithm can more intelligently understand and utilize the semantic differences and association strengths between different entity types. This embodiment can more effectively filter out low-value or irrelevant paths, thereby more accurately revealing the deep medical associations related to the mechanism of action of roselic acid C molecules and epilepsy diagnosis and treatment, providing high-quality implicit association information for subsequent fine-grained semantic analysis and knowledge graph expansion.
[0143] The beneficial effects of implementing the embodiments of the present invention include: the embodiments of this application first obtain initial medical text information, then preprocess the initial medical text information to obtain non-standard medical text information, then perform terminology conversion on the non-standard medical text information according to the medical terminology relationship knowledge base to obtain standardized medical text information, and finally extract medical information related to the mechanism of action of rose acid C molecules and epilepsy diagnosis and treatment from the standardized medical text information, and add the medical information to the knowledge graph database, thereby enabling the conversion of non-standard medical text information into standardized medical text information in combination with the medical terminology relationship knowledge base to construct a knowledge graph, thereby improving the accuracy of data recognition and the quality of the knowledge graph.
[0144] like Figure 2 As shown, this embodiment of the invention also provides a system for constructing a medical knowledge graph for epilepsy based on rose acid C, including:
[0145] Information acquisition module 701 is used to acquire initial medical text information;
[0146] Preprocessing module 702 is used to preprocess the initial medical text information to obtain non-standard medical text information;
[0147] The terminology conversion module 703 is used to convert non-standard medical text information into standardized medical text information based on a medical terminology relationship knowledge base.
[0148] The medical information extraction module 704 is used to extract medical information related to the mechanism of action of rose acid C molecules and the diagnosis and treatment of epilepsy from standardized medical text information.
[0149] The knowledge graph augmentation module 705 is used to add medical information to the knowledge graph database.
[0150] The content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0151] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
Claims
1. A method for constructing a medical knowledge graph for epilepsy based on rose acid C, characterized in that, Includes the following steps: Obtain initial medical text information; The initial medical text information is preprocessed to obtain non-standard medical text information; Based on a medical terminology knowledge base, the non-standard medical text information is converted into standardized medical text information. Extract medical information related to the mechanism of action of rose acid C molecules and the diagnosis and treatment of epilepsy from the standardized medical text information; Add the medical information to the knowledge graph database; The step of converting the non-standard medical text information into standardized medical text information based on a medical terminology knowledge base includes: The non-standard medical text information is segmented into words to identify multiple non-standard medical terms; The non-standard medical terms are precisely matched with the medical terminology relationship knowledge base to obtain precise matching results and precise matching values. If the exact match result is a successful match, then the exact match value will be used as a standardized medical term. If the exact matching result is a failure, then a fuzzy match is performed between the non-standard medical terminology and the medical terminology relationship knowledge base to identify the standardized medical terminology; Based on the standardized medical terminology, the non-standard medical text information is replaced with words to obtain the standardized medical text information; The step of performing fuzzy matching between the non-standard medical terms and the medical terminology relationship knowledge base to identify standardized medical terms includes: The initial context information of the non-standard medical terms is extracted from the non-standard medical text information, and the initial context information includes neighboring keywords, syntactic structure and semantic category; Based on the initial context information, a contextual semantic fingerprint is generated; The similarity between the contextual semantic fingerprint and each matching key in the medical terminology relation knowledge base is calculated using a cosine similarity algorithm. Based on the medical terminology relationship knowledge base, the fuzzy matching value corresponding to the matching key with the highest similarity is taken as the standardized medical term; The step of generating a contextual semantic fingerprint based on the initial context information includes: The non-standard medical text information is subjected to structured analysis to determine the region to which the non-standard medical terms belong; If the region to which it belongs is a region with sparse contextual information, then according to the semantic category, the non-standard medical text information is re-searched to extract supplementary contextual information; The initial context information and the supplementary context information are integrated to obtain the target context information; Based on the target context information, the contextual semantic fingerprint is generated; The process of integrating the initial context information and the supplementary context information to obtain the target context information includes: The initial context information and the supplementary context information are compared to identify semantic conflict information; Fine-grained semantic analysis is performed on the semantic conflict information to obtain fine-grained semantic information; The fine-grained semantic information is sent to an expert adjudication terminal, which adjudicates the fine-grained semantic information and generates contextual information correction suggestions. Based on the proposed contextual information correction, the initial contextual information and the supplementary contextual information are corrected. The corrected initial context information and the supplementary context information are integrated to obtain the target context information.
2. The method according to claim 1, characterized in that, The preprocessing of the initial medical text information to obtain non-standard medical text information includes: The initial medical text information is subjected to text cleaning processing, which is used to remove Hypertext Markup Language tags, special symbols, redundant spaces and line breaks; Perform case normalization on the initial medical text information after text cleaning; Sentence boundary analysis is performed on the initial medical text information after case normalization to obtain the non-standard medical text information.
3. The method according to claim 1, characterized in that, The fine-grained semantic analysis of the semantic conflict information to obtain fine-grained semantic information includes: The semantic conflict information is updated according to a preset set of semantic mapping rules; Identify multiple core entities related to the mechanism of action of rose acid C molecules and epilepsy diagnosis and treatment from the updated semantic conflict information; Analyze the entity relationships between the multiple core entities; Based on the entity associations, a preliminary association network is constructed, which contains multiple entity pairs; Based on a preset set of association rules, the preliminary association network is used to identify association paths to obtain implicit association paths between entity pairs. Extract implicit association information from the implicit association paths; Based on the implicit association information, the preliminary association network is updated to obtain an enhanced association network; Based on the enhanced association network, fine-grained semantic analysis is performed on the updated semantic conflict information to obtain the fine-grained semantic information.
4. The method according to claim 3, characterized in that, The step of updating the semantic conflict information according to a preset semantic mapping rule set includes: The semantic conflict information is classified to obtain experimental model information, observation index information, and terminology system information; Based on the preset semantic mapping rule set, semantic normalization processing is performed on experimental model information, observation index information, and terminology system information; The semantic conflict information is updated based on the experimental model information, observation index information, and terminology system information after semantic normalization.
5. The method according to claim 3, characterized in that, The step of identifying association paths in the preliminary association network according to a preset association rule set to obtain implicit association paths between entity pairs includes: The multiple core entities are classified to obtain various entity types; Based on the various entity types and the preset association rule set, determine the conversion cost corresponding to each entity type; Based on the various entity types and the conversion cost, adjust the priority and direction of path search; Based on the preset association rule set, the priority and direction of the path search, the implicit association path is found through the path search algorithm.
6. A knowledge graph construction system for epilepsy treatment based on rose acid C, characterized in that, include: The information acquisition module is used to acquire initial medical text information; The preprocessing module is used to preprocess the initial medical text information to obtain non-standard medical text information; The terminology conversion module is used to convert the non-standard medical text information into standardized medical text information based on a medical terminology relationship knowledge base. The medical information extraction module is used to extract medical information related to the mechanism of action of rose acid C molecules and the diagnosis and treatment of epilepsy from the standardized medical text information; The knowledge graph augmentation module is used to add the medical information to the knowledge graph database; The step of converting the non-standard medical text information into standardized medical text information based on a medical terminology knowledge base includes: The non-standard medical text information is segmented into words to identify multiple non-standard medical terms; The non-standard medical terms are precisely matched with the medical terminology relationship knowledge base to obtain precise matching results and precise matching values. If the exact match result is a successful match, then the exact match value will be used as a standardized medical term. If the exact matching result is a failure, then a fuzzy match is performed between the non-standard medical terminology and the medical terminology relationship knowledge base to identify the standardized medical terminology; Based on the standardized medical terminology, the non-standard medical text information is replaced with words to obtain the standardized medical text information; The step of performing fuzzy matching between the non-standard medical terms and the medical terminology relationship knowledge base to identify standardized medical terms includes: The initial context information of the non-standard medical terms is extracted from the non-standard medical text information, and the initial context information includes neighboring keywords, syntactic structure and semantic category; Based on the initial context information, a contextual semantic fingerprint is generated; The similarity between the contextual semantic fingerprint and each matching key in the medical terminology relation knowledge base is calculated using a cosine similarity algorithm. Based on the medical terminology relationship knowledge base, the fuzzy matching value corresponding to the matching key with the highest similarity is taken as the standardized medical term; The step of generating a contextual semantic fingerprint based on the initial context information includes: The non-standard medical text information is subjected to structured analysis to determine the region to which the non-standard medical terms belong; If the region to which it belongs is a region with sparse contextual information, then according to the semantic category, the non-standard medical text information is re-searched to extract supplementary contextual information; The initial context information and the supplementary context information are integrated to obtain the target context information; Based on the target context information, the contextual semantic fingerprint is generated; The process of integrating the initial context information and the supplementary context information to obtain the target context information includes: The initial context information and the supplementary context information are compared to identify semantic conflict information; Fine-grained semantic analysis is performed on the semantic conflict information to obtain fine-grained semantic information; The fine-grained semantic information is sent to an expert adjudication terminal, which adjudicates the fine-grained semantic information and generates contextual information correction suggestions. Based on the proposed contextual information correction, the initial contextual information and the supplementary contextual information are corrected. The corrected initial context information and the supplementary context information are integrated to obtain the target context information.
Citation Information
Patent Citations
Mapping processing system and method for solving problem of standard code control of medical data
CN104156415A
Epilepsy detection method, system and equipment based on mapping knowledge domain and electroencephalogram
CN115019950A
Credit granting processing and knowledge graph processing method and device, equipment, medium and product
CN118657212A
Intelligent auxiliary decision-making method and system for multi-mode data fusion of multiple myeloma
CN121215230A
Drug safety multi-center joint evaluation method and system based on graph neural network and federated learning
CN121281870A