Information mining methods for traditional Chinese medicine texts
Through word segmentation processing and information network diagram establishment, the problems of low efficiency and poor accuracy in TCM text information mining are solved, and efficient and accurate extraction and integration of TCM text information is achieved.
Patent Information
- Application Number
- CN202410628182.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-21
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-05-21
AI Technical Summary
Existing technologies make it difficult to extract information from traditional Chinese medicine texts efficiently and accurately, especially because traditional Chinese medicine texts use a large amount of classical Chinese and have ambiguous semantics, resulting in low information mining efficiency and poor accuracy.
Through word segmentation, keyword retrieval and establishment of information network diagram, keywords and their associations are determined, similar associations are inferred, and keywords with similar associations are merged to form a concrete information network diagram.
It improves the efficiency and accuracy of TCM text information mining, achieves the standardization and repeatability of information, and can effectively integrate and compare information in TCM texts.
Smart Images

Figure CN118585620B_ABST
Abstract
Claims
1. A method for information mining of traditional Chinese medicine texts, characterized in that: The method comprises the following steps: Performing word segmentation processing on the traditional Chinese medicine text; Based on the TCM text after word segmentation processing, a text information set is determined, wherein the text information set includes keywords and association relationships between the keywords, the association relationships include at least one of similarity association relationships and context association relationships of the keywords, and the context association relationship refers to a logical association relationship between the keywords in the context of the TCM text. The text information set is determined by keyword retrieval, and the keyword retrieval includes the following steps: Determine the source keywords based on the target information to be mined. The source keyword is used as a search base word to perform a keyword search to determine a first-level keyword associated with the source keyword, wherein the keyword search is to search for sentences including the search base word in the text, extract associated words from the found sentences, split and / or merge the extracted associated words, perform cluster analysis on the split and / or merged associated words, and select the same part of the associated words from the cluster analysis results as keywords, Repeatedly use the previous level keyword as the search base word to perform keyword search to determine the next level keyword that is related to the previous level keyword. When no new keywords can be selected from the cluster analysis results, the search for the next-level keywords is terminated. storing the search base words, the selected keywords and the association relationship in each keyword search in the text information set; and An information network diagram is established based on the text information set, wherein the information network diagram has the keywords as nodes and the association relationships as edges. Based on the information network diagram, similarity associations of the keywords are inferred. The inferred similarity associations are similarity associations of keywords that have no original textual basis but objectively have semantic similarity. The inferred similarity associations are marked in the information network diagram. The information mining method includes performing graphic cluster analysis on an information network graph, finding and merging similar graphic structures in the information network graph, and making the merged graphic structure inherit the contextual association relationship of the pre-merged graphic structure.
2. The information mining method of traditional Chinese medicine text according to claim 1, characterized in that: Keywords with similar association relationships in the information network graph are merged, and the merged keywords are made to inherit the context association relationships of the keywords before the merger.
3. The information mining method of traditional Chinese medicine text according to claim 1, characterized in that: The information network diagram is used for information comparison between traditional Chinese medicine texts.
4. The information mining method of traditional Chinese medicine text according to claim 1, characterized in that: The information network diagram is used for information comparison between traditional Chinese medicine texts and modern medical texts.
Citation Information
Patent Citations
Data mining method and device for network pharmacology
CN115148374A
Text data information mining method, device and equipment
CN115374781A
Potential customer mining method and device
CN116860981A