Traditional Chinese medicine ancient book database management system and method
By performing multidimensional metadata annotation and deep correlation mining on ancient Chinese medicine books, a dynamic and scalable knowledge graph of ancient Chinese medicine advantageous diseases is constructed, which solves the problem of multi-source data fusion in Chinese medicine and realizes the structured transformation and efficient utilization of data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- EYE HOSPITAL CHINA ACAD OF CHINESE MEDICAL SCI
- Filing Date
- 2026-03-02
- Publication Date
- 2026-06-02
AI Technical Summary
In existing technologies, it is difficult to integrate multi-source data of traditional Chinese medicine, making it impossible to use directly. There is a lack of systematic governance, and the standards of multi-source heterogeneous data are inconsistent and the quality varies. The complex diagnostic logic in unstructured medical records has not been thoroughly and automatically cleaned and structured, resulting in case data that cannot be directly calculated and used by machines.
By acquiring ancient texts on diseases, performing multidimensional metadata annotation, deep correlation mining and visualization analysis, a dynamic and scalable knowledge graph of ancient texts on diseases with advantages in traditional Chinese medicine is constructed. Multidimensional encoding is used to map the data to a low-dimensional vector space, realizing deep correlation and structured transformation of multi-source data.
It achieves deep correlation of multi-source data, constructs a knowledge graph of ancient books on TCM advantageous diseases that is easy for users to access and analyze, improves the efficiency and accuracy of data utilization, and supports personalized knowledge services.
Smart Images

Figure CN122136027A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data management technology, specifically to a method and system for managing a database of ancient Chinese medical texts. Background Technology
[0002] Currently, with the deepening of healthcare informatization, traditional Chinese medicine (TCM) medical institutions have widely deployed hospital information systems and electronic medical record systems, accumulating massive amounts of clinical diagnosis and treatment data. This data provides valuable resources for TCM clinical research, efficacy evaluation, and experience transmission. Existing technologies mainly use standard interfaces of hospital information systems or customized data extraction tools to collect structured or semi-structured data such as patient basic information, diagnoses, prescriptions, and examination and test reports. Some advanced institutions or research projects have begun to attempt to integrate multi-source data, such as digitizing ancient books and records of famous doctors' experiences, and using natural language processing technology for preliminary information extraction. Existing technologies mostly focus on single-dimensional statistics, such as retrospectively analyzing the effectiveness of specific prescriptions in treating a certain disease, or using data mining algorithms to discover the compatibility patterns between drugs. Some exploratory studies attempt to construct TCM ontology or small knowledge graphs for use in teaching or literature retrieval.
[0003] Despite the progress made by existing technologies, the following problems still exist:
[0004] First, existing data collection methods lack systematic governance, resulting in inconsistent standards and varying quality among multi-source heterogeneous data, making effective integration difficult. Differences in Chinese medicine aliases, dosage units, and syndrome descriptions, as well as the complex diagnostic logic hidden in unstructured medical records, have not yet been thoroughly and automatically cleaned and structured, resulting in a large amount of case data that cannot be directly calculated and utilized by machines. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a method and system for managing a database of ancient Chinese medicine books, so as to solve the problem that multi-source data of Chinese medicine is difficult to integrate and cannot be directly utilized in the prior art.
[0006] According to a first aspect of the present invention, a method for managing a database of ancient Chinese medical texts is provided, comprising: Obtain ancient texts data for diseases, select modern diseases with advantages in TCM diagnosis and treatment from the ancient texts data for the diseases, perform multidimensional metadata annotation on the ancient texts data for each selected modern disease, and then perform deep correlation mining and visualization analysis to obtain a database of advantageous diseases. For each modern disease in the advantageous disease database, supplementary data is obtained from multiple sources, and the supplementary data is preprocessed and transformed into structured data entries. Each structured data entry is accompanied by its original source, version, and confidence label as a knowledge unit feature. A unique and semantically meaningful multidimensional code is assigned to each structured data entry to form a knowledge unit; each knowledge unit is then mapped to a low-dimensional vector space based on the multidimensional code. Based on knowledge units in a low-dimensional vector space, and following the principle of one modern disease corresponding to one independent subgraph, a dynamic and scalable knowledge graph of ancient Chinese medicine advantageous diseases is constructed.
[0007] Preferably, modern diseases with advantages in TCM diagnosis and treatment are selected from the ancient medical records of the aforementioned diseases, including: Based on a pre-established network of precise mappings between modern disease diagnosis names and descriptive terms in ancient Chinese medicine texts, modern diseases with advantages in TCM diagnosis and treatment are screened from the ancient text data of the aforementioned diseases. For each selected modern disease, a knowledge ontology framework is constructed, which includes disease information, symptom information, treatment information, prescription information, and drug information.
[0008] Preferably, multidimensional metadata annotation is performed on the ancient texts corresponding to each selected modern disease, followed by in-depth correlation mining and visualization analysis to derive a database of advantageous diseases, including: The ancient texts corresponding to the selected modern diseases are annotated with multi-dimensional metadata, including time information, geographical information, school of thought information, and physician information; The labeled modern diseases are forcibly associated with specific disease units to form a multi-dimensional metadata framework covering time, space, academic origins and subject dimensions. Based on a multidimensional metadata framework, we conduct in-depth correlation mining and visualization analysis, cross-analyze ancient book data under the same modern disease from multiple dimensions, and integrate and encapsulate the analysis results into a dominant disease database.
[0009] Preferably, for each modern disease in the advantageous disease database, supplementary data is obtained from multiple sources, including: Customize differentiated acquisition strategies for each channel in a multi-source channel; Based on the collection strategy corresponding to each channel, supplementary data corresponding to modern diseases are extracted from the channels. The multi-source channels include digital ancient book databases, ancient book image resources, published annotated editions, and related academic papers.
[0010] Preferably, the supplementary data is preprocessed and converted into structured data entries, including: A natural language processing model trained on ancient Chinese medicine texts is used to perform named entity recognition on each piece of supplementary data, automatically extracting key disease information; The application of rule engines and deep learning algorithms identifies and labels the semantic relationships between key information of diseases; Each modern disease is treated as an entity, its corresponding key disease information is treated as an attribute, and semantic relationships are treated as relations, which are transformed into structured data entries. The structured data entries are in the form of entity-attribute-relationship triples.
[0011] Preferably, the multidimensional encoding includes: Disease category code is the core classification identifier used to locate the modern disease to which a knowledge unit belongs; Knowledge type code is a content nature tag used to distinguish the system category of knowledge units; Historical period codes are time-series positioning tags used to mark the dynasty or stage of medical development in which a knowledge unit was generated; Evidence rating codes are quality assessment labels used to algorithmically evaluate the reliability of documents for knowledge units. Association strength code is a relation quantification label used to characterize the strength of logical associations between knowledge units.
[0012] Preferably, a dynamic and scalable knowledge graph of ancient books on diseases with advantages in traditional Chinese medicine is constructed, including: Based on all knowledge units in the low-dimensional vector space, and using the native graph database as the storage and computing engine, each knowledge unit is abstracted into a node in the graph, and the semantic relationship between nodes is abstracted into an edge with type and weight. The nodes and edges constitute the network basic framework that maps the cognitive logic of traditional Chinese medicine. Based on the disease-specific dimension code, all nodes and edges are automatically assigned to corresponding independent subgraphs; each subgraph is stored and managed independently, and flexible associations between subgraphs are achieved through shared public nodes; Once new ancient text data is processed and new knowledge units are obtained, new nodes and edges are automatically constructed and incorporated into the corresponding independent subgraphs.
[0013] Preferably, the method for managing the database of ancient Chinese medical books further includes: Based on the constructed knowledge graph of ancient books on diseases with advantages in traditional Chinese medicine, personalized knowledge service interfaces and dynamic outputs are provided for different application scenarios.
[0014] Preferably, the method for managing the database of ancient Chinese medical books further includes: The knowledge discovery engine based on graph algorithms is run on the knowledge graph of ancient books on the diseases with advantages in traditional Chinese medicine to automatically identify clusters of different prescription combinations targeting the same core pathogenesis or main symptoms, and automatically summarize the diagnosis and treatment rules.
[0015] According to a second aspect of the present invention, a database management system for ancient Chinese medical texts is provided, comprising: The advantageous disease database construction module is used to acquire ancient book data of diseases, select modern diseases with advantages in TCM diagnosis and treatment from the ancient book data of the diseases, perform multi-dimensional metadata annotation on the ancient book data of each selected modern disease, and then perform deep correlation mining and visualization analysis to obtain the advantageous disease database. The data acquisition and preprocessing module is used to acquire supplementary data from multiple sources for each modern disease in the advantageous disease database, and to preprocess the supplementary data to convert it into structured data entries. Each structured data entry is accompanied by its original source, version, and confidence label as a knowledge unit feature. The knowledge encoding and indexing module is used to assign unique and semantically meaningful multidimensional codes to structured data entries to form knowledge units; and to map each knowledge unit to a low-dimensional vector space according to the multidimensional codes. The knowledge graph construction module is used to construct a dynamic and scalable knowledge graph of ancient Chinese medicine advantageous diseases based on knowledge units in a low-dimensional vector space, with one independent subgraph corresponding to one modern disease.
[0016] The technical solution provided by this invention may include the following beneficial effects: It is understood that the technical solution presented in this invention can acquire ancient text data on diseases, select modern diseases with advantages in TCM diagnosis and treatment, and annotate them with multi-dimensional metadata to obtain a database of advantageous diseases; it acquires supplementary data for each modern disease from multiple sources, transforms it into structured data entries, assigns unique and semantically meaningful multi-dimensional codes, constitutes knowledge units, and maps them to a low-dimensional vector space; and it constructs a dynamically expandable knowledge graph of ancient texts on TCM advantageous diseases, with one independent subgraph corresponding to one modern disease. It is understood that this technical solution integrates multi-source data, enabling interconnection between them and transforming ancient text data from isolated entries into a deeply interconnected system. The constructed knowledge graph of ancient texts on TCM advantageous diseases is easy for users to access and analyze, and has high practicality.
[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0019] Figure 1 This is a schematic diagram illustrating the steps of a method for managing a database of ancient Chinese medical books according to an exemplary embodiment; Figure 2 This is a schematic block diagram illustrating a database management system for ancient Chinese medical books according to an exemplary embodiment. Detailed Implementation
[0020] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.
[0021] In one embodiment, Figure 1 This is a schematic diagram illustrating the steps of a method for managing a database of ancient Chinese medical books according to an exemplary embodiment. See also... Figure 1 This provides a method for managing a database of ancient Chinese medicine texts, including: Step S11: Obtain ancient book data of diseases, select modern diseases with advantages in TCM diagnosis and treatment from the ancient book data of the diseases, perform multidimensional metadata annotation on the ancient book data of each selected modern disease, and then perform deep association mining and visualization analysis to obtain the advantageous disease database.
[0022] In practice, this step receives disease-related ancient text data transmitted from the host computer, selects modern diseases with clear advantages in TCM diagnosis and treatment, and incorporates metadata dimensions such as time, region, school of thought, and physician for these modern diseases. This allows for in-depth correlation and multi-dimensional analysis of the modern diseases, establishing a database of advantageous diseases. Understandably, this step enables the focus on and systematic reconstruction of the clinical value of scattered ancient text resources, establishing a clear target framework and classification system for subsequent precise and structured knowledge management, fundamentally solving the problem of the disconnect between ancient text knowledge and application scenarios.
[0023] In a preferred embodiment, selecting modern diseases with advantages in TCM diagnosis and treatment from the ancient medical texts of the diseases includes: screening modern diseases with advantages in TCM diagnosis and treatment from the ancient medical texts of the diseases based on a pre-established network of precise mapping relationships between modern disease diagnosis names and descriptive terms in ancient TCM texts; and constructing a knowledge ontology framework for each screened modern disease, wherein the knowledge ontology framework includes disease information, symptom information, treatment method information, prescription information, and drug information.
[0024] This embodiment enables experts in the field to conduct historical literature review for each disease in the advantageous disease database, establishing a precise mapping network from modern disease diagnosis names to related disease, syndrome, symptom, and symptom descriptions in ancient Chinese medicine texts, thus completing the selection of advantageous diseases. Understandably, this technical solution can build a bridge between ancient and modern medical terminology, ensuring the accuracy and relevance of knowledge review and avoiding retrieval omissions or misinterpretations caused by terminology changes.
[0025] Furthermore, a dedicated, structured knowledge ontology framework is constructed for each selected dominant disease. This framework uses disease-symptom-treatment-prescription-drug as its core logical axis, clearly defining the attribute characteristics of each dominant disease and their semantic relationships. This technical solution provides a unified, machine-understandable data model for massive amounts of unstructured text, forcing logical organization and laying a core foundation for automated data processing and deep correlation.
[0026] In a preferred embodiment, multidimensional metadata annotation is performed on the ancient book data corresponding to each selected modern disease, followed by deep correlation mining and visualization analysis to obtain a dominant disease database, including: For the ancient texts corresponding to the selected modern diseases, multidimensional metadata annotations are performed on time information, regional information, school of thought information, and physician information. The annotated modern diseases are then forcibly associated with specific disease units to form a multidimensional metadata framework covering time, space, academic origins, and subject dimensions. Based on the multidimensional metadata framework, in-depth association mining and visualization analysis are carried out to cross-analyze the ancient texts under the same modern disease from multiple dimensions. The analysis results are then integrated and packaged into a superior disease database.
[0027] In practice, this technical solution targets selected modern diseases and organizes experts in the field to conduct research on ancient books to accurately annotate the time information such as the dynasty and historical period of each relevant discussion. This generates attribute information such as region, area of dissemination, school of thought, and background information. After standardized annotation, it is forcibly associated with specific disease units to form a multi-dimensional metadata framework covering time, space, academic origin, and subject dimensions.
[0028] Subsequently, based on a multidimensional metadata framework, in-depth correlation mining and visualization analysis were carried out to cross-analyze the ancient knowledge of the same dominant disease from multiple dimensions, revealing the historical evolution, regional characteristics and academic preferences of schools of thought behind the knowledge, and systematically integrating and encapsulating the knowledge clusters after in-depth analysis into a dominant disease database.
[0029] Step S12: For each modern disease in the advantageous disease database, supplementary data is obtained from multiple sources, and the supplementary data is preprocessed and transformed into structured data entries. Each structured data entry is accompanied by its original source, version, and confidence label as a knowledge unit feature.
[0030] In practice, this step, based on an established database of advantageous diseases, involves collecting and verifying unstructured or semi-structured raw data related to the target advantageous diseases from multiple sources to obtain supplementary data. This supplementary data then undergoes a multi-level structured preprocessing workflow, ultimately transforming it into structured data entries. This step enables the transformation from raw documents to standardized, computable knowledge atoms, ensuring data traceability and the possibility of quality assessment, and is the cornerstone of building high-quality knowledge assets.
[0031] In a preferred embodiment, for each modern disease in the dominant disease database, supplementary data is obtained from multiple sources, including: Differentiated acquisition strategies are customized for each channel in the multi-source channels; supplementary data corresponding to modern diseases are extracted from the channels according to the acquisition strategy corresponding to each channel; the multi-source channels include digital ancient book databases, ancient book image resources, published annotated editions and related professional academic papers.
[0032] In practice, this technical solution plans and connects to four major data channels for each modern disease in the established advantageous disease database: authoritative digital ancient book databases, high-precision scanned ancient book image resources, published annotated editions compiled by modern scholars, and relevant professional academic papers. This technical solution can ensure the comprehensiveness, authority, and complementarity of data sources, and maximize the coverage of all relevant records and research on the target disease.
[0033] Subsequently, differentiated data collection strategies were developed for each channel based on the knowledge ontology. For example, the main text and annotations were extracted from annotated editions, and ancient texts cited in academic papers were extracted to ensure comprehensive coverage. This technical solution enables intelligent and targeted data collection, effectively extracting core and valid information from different sources and improving data collection efficiency and accuracy.
[0034] In a preferred embodiment, the supplementary data is preprocessed to transform it into structured data entries, including: A natural language processing model trained on ancient Chinese medicine texts is used to perform named entity recognition on each supplementary data entry, automatically extracting key disease information such as physician, prescription name, drug name, symptoms, and pathogenesis. A rule engine and deep learning algorithm are applied to identify and label the semantic relationships between key disease information, and to automatically standardize and replace variant characters, alternative characters, and taboo characters unique to ancient texts. Each modern disease is treated as an entity, its corresponding key disease information as an attribute, and semantic relationships as relations, transforming them into machine-readable structured data entries. These structured data entries are in the form of entity-attribute-relation triples.
[0035] Step S13: Assign a unique and semantically meaningful multidimensional code to each structured data entry to form a knowledge unit; map each knowledge unit to a low-dimensional vector space according to the multidimensional code.
[0036] In practice, this step assigns multidimensional encoding to structured data entries. Through knowledge graph embedding technology, the features and relational information of knowledge units are mapped to a low-dimensional vector space, so that semantically similar or closely related knowledge units are arranged close to each other in the vector space. This enables any prescription or discussion recorded in an ancient book to be accurately located to a specific coordinate in the disease knowledge network to which it belongs, and related knowledge content in the knowledge unit can be quickly indexed through vector calculation.
[0037] This technical solution can transform discrete knowledge items into semantic vectors with mathematical expressions, enabling computers to understand and reason about the relationships between knowledge by calculating "distance," thus achieving a leap from symbol matching to semantic computation.
[0038] It should be noted that the multidimensional encoding includes: The disease-specific dimension code serves as the core classification identifier, used to locate the modern disease to which a knowledge unit belongs. This code enables primary-level classification management of knowledge, ensuring that all operations and queries are conducted within clearly defined disease boundaries, thus guaranteeing the systematic nature of knowledge and the relevance of its application scenarios.
[0039] Knowledge type codes are content-related tags used to distinguish the system categories of knowledge units such as pathogenesis, treatment methods, and prescriptions. Knowledge type codes can finely classify knowledge based on logical type and support targeted retrieval, statistics, and analysis by knowledge type.
[0040] Historical period codes are time-series positioning tags used to mark the dynasty or medical development stage in which a knowledge unit was generated. This technical solution can timestamp knowledge units, giving the knowledge system a diachronic perspective and supporting trend analysis and evolution research along the timeline.
[0041] Evidence rating codes are quality assessment labels used to algorithmically evaluate the reliability of literature for knowledge units. This technical solution can introduce the concept of evidence-based medicine to quantify the credibility of knowledge units themselves, providing an objective quality weight basis for subsequent recommendations and rankings.
[0042] Association strength codes are quantitative labels used to characterize the strength of logical connections between knowledge units. Association strength codes concretize the abstract concept of "relevance" into comparable numerical values, giving weight to connections in knowledge networks and better reflecting the reality that knowledge connections vary in strength in the real world.
[0043] Step S14: Based on the knowledge units in the low-dimensional vector space, construct a dynamic and scalable knowledge graph of ancient books on traditional Chinese medicine advantageous diseases, in the manner of one modern disease corresponding to one independent subgraph.
[0044] In a preferred embodiment, a dynamically scalable knowledge graph of ancient Chinese medicine texts on diseases with advantages in traditional Chinese medicine is constructed, including: Based on all knowledge units in the low-dimensional vector space, and using the native graph database as the storage and computing engine, each knowledge unit of disease, syndrome, symptom, prescription and drug is abstracted into a node in the graph, and the semantic relationship between nodes is abstracted into an edge with type and weight. The nodes and edges constitute the network basic framework that maps the cognitive logic of traditional Chinese medicine.
[0045] Based on the disease dimension code, all nodes and edges are automatically divided into corresponding independent subgraphs; each subgraph is stored and managed independently, and flexible associations between subgraphs are achieved through shared public nodes.
[0046] The knowledge graph of ancient books on diseases with advantages in traditional Chinese medicine supports real-time expansion. When new ancient book data is processed and new knowledge units are obtained and indexed, new nodes and edges are automatically constructed and incorporated into the corresponding independent subgraphs, so as to realize the continuous growth and updating of the knowledge system.
[0047] In a preferred embodiment, the method for managing the database of ancient Chinese medical books further includes: A knowledge discovery engine based on graph algorithms is applied to the knowledge graph of ancient Chinese medicine texts on the aforementioned advantageous diseases. This automatically identifies clusters of different prescription combinations targeting the same core pathogenesis or main symptom, and automatically summarizes diagnostic and treatment rules. These rules can be implicit and have high confidence. This technical solution can activate a static knowledge base into a dynamic, complex system with growth and reasoning capabilities, achieving a qualitative leap from "storing the known" to "discovering the unknown," and uncovering deep-seated patterns that are difficult for humans to summarize directly.
[0048] In a preferred embodiment, the method for managing the database of ancient Chinese medical books further includes: Based on the constructed knowledge graph of ancient books on diseases with advantages in traditional Chinese medicine, personalized knowledge service interfaces and dynamic outputs are provided for different application scenarios.
[0049] In practical applications, all output results support interactive exploration. Users can click on any recommended item or statistical node to view the complete knowledge background and related networks. This technical solution encapsulates complex underlying knowledge computing capabilities into easy-to-use high-level applications, directly creating value for end users and realizing a complete closed loop from data resources to knowledge services.
[0050] For example, in clinical decision support scenarios, after receiving patient symptoms, signs, and tongue and pulse information input by the user, this information is mapped to nodes in a knowledge graph. Through graph traversal and similarity calculation, relevant syndrome types, treatment methods, and historical prescriptions recorded in ancient books are quickly retrieved. Recommendations are then sorted and ranked according to evidence level and correlation strength. Simultaneously, the literature sources, historical commentaries, and evolutionary paths of the recommended prescriptions are visually displayed in the form of a source map. This technical solution provides clinicians with a traceable intelligent auxiliary reference based on massive historical experience, integrating ancient book experience into modern diagnostic and treatment processes in a structured and intelligent manner, thus helping to improve the accuracy and rationality of syndrome differentiation and treatment.
[0051] For scientific research analysis scenarios, multidimensional statistical analysis tools are provided, enabling comparative analysis of the evolution of treatment methods and prescriptions for a specific disease across multiple dimensions, including dynasties, physicians, and regions. Alternatively, network pharmacology calculations can be used to correlate high-frequency drug combinations with modern pharmacological research. This technical solution provides TCM researchers with powerful data mining and visualization analysis tools, supporting the discovery of macroscopic academic patterns and interdisciplinary research combining TCM and Western medicine, significantly improving research efficiency and depth.
[0052] Understandably, this technical solution constructs a knowledge ontology framework based on modern advantageous diseases and incorporates multi-dimensional metadata such as time, region, school of thought, and physician. It anchors and reorganizes scattered ancient texts into a structured modern clinical context. By utilizing deep learning-based named entity recognition and relation extraction technology, it transforms unstructured text into machine-readable "entity-attribute-relationship" triples. Combined with knowledge graph embedding technology, it generates semantically rich multi-dimensional intelligent codes. This breaks down the barriers between ancient text knowledge units at the underlying data level, achieving a fundamental transformation of knowledge from isolated entries to a deeply interconnected, precisely located, and computationally calculable networked system. By constructing a dynamic knowledge graph with dominant diseases as independent subgraphs and running graph algorithms for community discovery and rule mining, the system can automatically identify the patterns of prescription combinations and summarize high-confidence treatment rules, upgrading the static knowledge base into a "living" knowledge system with autonomous discovery and reasoning capabilities. Personalized service interfaces for clinical and research scenarios rely on this graph to achieve intelligent retrieval, evidence-based recommendation, and comprehensive traceability based on semantic similarity. This not only significantly improves the efficiency and depth of ancient book knowledge retrieval but also provides powerful intelligent tools to support the precision of TCM clinical decision-making and the paradigm innovation of academic research in a visualized and interactive way.
[0053] In another embodiment, see Figure 2 This provides a database management system for ancient Chinese medicine texts, including: The advantageous disease database construction module is used to acquire ancient book data of diseases, select modern diseases with advantages in TCM diagnosis and treatment from the ancient book data of the diseases, perform multi-dimensional metadata annotation on the ancient book data of each selected modern disease, and then perform deep correlation mining and visualization analysis to obtain the advantageous disease database. The data acquisition and preprocessing module is used to acquire supplementary data from multiple sources for each modern disease in the advantageous disease database, and to preprocess the supplementary data to convert it into structured data entries. Each structured data entry is accompanied by its original source, version, and confidence label as a knowledge unit feature. The knowledge encoding and indexing module is used to assign unique and semantically meaningful multidimensional codes to structured data entries to form knowledge units; and to map each knowledge unit to a low-dimensional vector space according to the multidimensional codes. The knowledge graph construction module is used to construct a dynamic and scalable knowledge graph of ancient Chinese medicine advantageous diseases based on knowledge units in a low-dimensional vector space, with one independent subgraph corresponding to one modern disease.
[0054] It is understood that the same or similar parts in the above embodiments can be referred to each other, and the contents not described in detail in some embodiments can be referred to the same or similar contents in other embodiments.
[0055] It should be noted that in the description of this invention, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this invention, unless otherwise stated, "a plurality of" means at least two.
[0056] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.
[0057] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0058] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0059] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0060] The storage media mentioned above can be read-only memory, disk, or optical disk, etc.
[0061] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0062] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for managing a database of ancient Chinese medical texts, characterized in that, include: Obtain ancient texts data for diseases, select modern diseases with advantages in TCM diagnosis and treatment from the ancient texts data for the diseases, perform multidimensional metadata annotation on the ancient texts data for each selected modern disease, and then perform deep correlation mining and visualization analysis to obtain a database of advantageous diseases. For each modern disease in the advantageous disease database, supplementary data is obtained from multiple sources, and the supplementary data is preprocessed and transformed into structured data entries. Each structured data entry is accompanied by its original source, version, and confidence label as a knowledge unit feature. A unique and semantically meaningful multidimensional code is assigned to each structured data entry to form a knowledge unit; each knowledge unit is then mapped to a low-dimensional vector space based on the multidimensional code. Based on knowledge units in a low-dimensional vector space, and following the principle of one modern disease corresponding to one independent subgraph, a dynamic and scalable knowledge graph of ancient Chinese medicine advantageous diseases is constructed.
2. The method for managing a database of ancient Chinese medical books according to claim 1, characterized in that, Modern diseases with advantages in TCM diagnosis and treatment were selected from the ancient medical records of the aforementioned diseases, including: Based on a pre-established network of precise mappings between modern disease diagnosis names and descriptive terms in ancient Chinese medicine texts, modern diseases with advantages in TCM diagnosis and treatment are screened from the ancient text data of the aforementioned diseases. For each selected modern disease, a knowledge ontology framework is constructed, which includes disease information, symptom information, treatment information, prescription information, and drug information.
3. The method for managing a database of ancient Chinese medical books according to claim 2, characterized in that, For each selected modern disease, the corresponding ancient text data is annotated with multidimensional metadata, followed by deep correlation mining and visualization analysis to derive a dominant disease database, including: The ancient texts corresponding to the selected modern diseases are annotated with multi-dimensional metadata, including time information, geographical information, school of thought information, and physician information; The labeled modern diseases are forcibly associated with specific disease units to form a multi-dimensional metadata framework covering time, space, academic origins and subject dimensions. Based on a multidimensional metadata framework, we conduct in-depth correlation mining and visualization analysis, cross-analyze ancient book data under the same modern disease from multiple dimensions, and integrate and encapsulate the analysis results into a dominant disease database.
4. The method for managing a database of ancient Chinese medical books according to claim 1, characterized in that, For each modern disease in the advantageous disease database, supplementary data is obtained from multiple sources, including: Customize differentiated acquisition strategies for each channel in a multi-source channel; Based on the collection strategy corresponding to each channel, supplementary data corresponding to modern diseases are extracted from the channels. The multi-source channels include digital ancient book databases, ancient book image resources, published annotated editions, and related academic papers.
5. The method for managing a database of ancient Chinese medical books according to claim 4, characterized in that, The supplementary data is preprocessed and transformed into structured data entries, including: A natural language processing model trained on ancient Chinese medicine texts is used to perform named entity recognition on each piece of supplementary data, automatically extracting key disease information; The application of rule engines and deep learning algorithms identifies and labels the semantic relationships between key information of diseases; Each modern disease is treated as an entity, its corresponding key disease information is treated as an attribute, and semantic relationships are treated as relations, which are transformed into structured data entries. The structured data entries are in the form of entity-attribute-relationship triples.
6. The method for managing a database of ancient Chinese medical books according to claim 1, characterized in that, The multidimensional encoding includes: Disease category code is the core classification identifier used to locate the modern disease to which a knowledge unit belongs; Knowledge type code is a content nature tag used to distinguish the system category of knowledge units; Historical period codes are time-series positioning tags used to mark the dynasty or stage of medical development in which a knowledge unit was generated; Evidence rating codes are quality assessment labels used to algorithmically evaluate the reliability of documents for knowledge units. Association strength code is a relation quantification label used to characterize the strength of logical associations between knowledge units.
7. The method for managing a database of ancient Chinese medical books according to claim 6, characterized in that, Constructing a dynamic and scalable knowledge graph of ancient Chinese medicine texts on diseases with advantages in traditional Chinese medicine, including: Based on all knowledge units in the low-dimensional vector space, and using the native graph database as the storage and computing engine, each knowledge unit is abstracted into a node in the graph, and the semantic relationship between nodes is abstracted into an edge with type and weight. The nodes and edges constitute the network basic framework that maps the cognitive logic of traditional Chinese medicine. Based on the disease-specific dimension code, all nodes and edges are automatically assigned to corresponding independent subgraphs; each subgraph is stored and managed independently, and flexible associations between subgraphs are achieved through shared public nodes; Once new ancient text data is processed and new knowledge units are obtained, new nodes and edges are automatically constructed and incorporated into the corresponding independent subgraphs.
8. The method for managing a database of ancient Chinese medical books according to claim 1, characterized in that, Also includes: Based on the constructed knowledge graph of ancient books on diseases with advantages in traditional Chinese medicine, personalized knowledge service interfaces and dynamic outputs are provided for different application scenarios.
9. The method for managing a database of ancient Chinese medical books according to claim 1, characterized in that, Also includes: The knowledge discovery engine based on graph algorithms is run on the knowledge graph of ancient books on the diseases with advantages in traditional Chinese medicine to automatically identify clusters of different prescription combinations targeting the same core pathogenesis or main symptoms, and automatically summarize the diagnosis and treatment rules.
10. A database management system for ancient Chinese medical texts, characterized in that, include: The advantageous disease database construction module is used to acquire ancient book data of diseases, select modern diseases with advantages in TCM diagnosis and treatment from the ancient book data of the diseases, perform multi-dimensional metadata annotation on the ancient book data of each selected modern disease, and then perform deep correlation mining and visualization analysis to obtain the advantageous disease database. The data acquisition and preprocessing module is used to acquire supplementary data from multiple sources for each modern disease in the advantageous disease database, and to preprocess the supplementary data to convert it into structured data entries. Each structured data entry is accompanied by its original source, version, and confidence label as a knowledge unit feature. The knowledge encoding and indexing module is used to assign unique and semantically meaningful multidimensional codes to structured data entries to form knowledge units; and to map each knowledge unit to a low-dimensional vector space according to the multidimensional codes. The knowledge graph construction module is used to construct a dynamic and scalable knowledge graph of ancient Chinese medicine advantageous diseases based on knowledge units in a low-dimensional vector space, with one independent subgraph corresponding to one modern disease.