An intelligent retrieval method and system for authentic medicinal materials based on a spectrum

Through the intelligent search method of authentic medicinal materials based on the graph, the problem that the Chinese medicinal materials information query system cannot understand the semantics of user input is solved, and intelligent matching with high accuracy and flexibility is achieved, potential medicinal materials nodes are discovered, and the semantic understanding and reasoning capabilities of the retrieval system are improved.

CN120216741BActive Publication Date: 2025-08-01HUNAN VOCATIONAL COLLEGE OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510691794.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-01
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

The existing Chinese herbal medicine information query system cannot understand the semantics of user input, the search results have low coverage and insufficient accuracy, and cannot achieve intelligent matching and semantic retrieval.

Method used

The intelligent search method of authentic medicinal materials based on graphs is adopted to generate hierarchical semantic edges through semantic particle size processing, build hierarchical semantic graphs, and perform semantic annotation of relational edges. Combining multi-conditional intention processing and multi-hop path combination, the edge reasoning completion mechanism is used to achieve multi-dimensional semantic analysis and intelligent matching.

Benefits of technology

It significantly improves the accuracy, flexibility and intelligence of authentic medicinal materials retrieval, can discover potential medicinal materials nodes, improves the semantic understanding depth and inference accuracy of the search, and supports intelligent matching of medicinal materials under various conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216741B_ABST
    Figure CN120216741B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of knowledge graph retrieval, and particularly to an intelligent retrieval method and system for authentic medicinal materials based on a graph. The method includes the following steps: performing semantic granularity parsing on a retrieval request input by a user, constructing multi-level semantic edges, and generating a hierarchical semantic graph structure. Subsequently, semantic enhancement is performed on the graph relationships through semantic annotation, and multi-condition intents are extracted in combination with dimensions such as region, medicinal property, and meridian tropism. Further, the system executes multi-hop path combination, mines structured semantic paths that meet the composite intent, and performs reasoning and complementation on the edge nodes in the paths to generate a retrieval result of authentic medicinal materials with a closed structure and complete semantics. Compared with traditional keyword matching and static field retrieval methods, the present invention has stronger semantic perception ability and reasoning intelligence, and significantly improves the accuracy and adaptability of the system in processing retrieval scenarios such as ambiguity, complexity, and incomplete paths.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graph retrieval, and in particular to an intelligent retrieval method and system for genuine regional drugs based on a graph. Background Art

[0002] Genuine regional drugs refer to Chinese medicinal materials produced under specific natural ecological environments and specific cultivation, harvesting and processing technical conditions, with excellent, stable quality and curative effects, and they play an important role in the clinical application of traditional Chinese medicine and the development of the traditional Chinese medicine industry. With the improvement of the digital and intelligent level of traditional Chinese medicine, how to extract the core attribute information of genuine regional drugs from the vast traditional Chinese medicine literature, standard pharmacopoeias, local chronicles and modern scientific research achievements, and realize intelligent matching and semantic retrieval based on user needs has become a key problem urgently to be solved in the field of traditional Chinese medicine information services. Most of the existing Chinese medicinal material information query systems are based on traditional relational databases or keyword matching models, and mainly use static field retrieval methods for query, such as screening by drug name, function and indications, place of origin, etc., and there are problems such as inability to understand the semantics of user input, low coverage rate of retrieval results, and insufficient accuracy. Summary of the Invention

[0003] In order to solve the above technical problems, the present invention proposes an intelligent retrieval method and system for genuine regional drugs based on a graph to solve at least one of the above technical problems.

[0004] The present application provides an intelligent retrieval method for genuine regional drugs based on a graph, and the method includes:

[0005] S1. Obtain the genuine regional drug retrieval request data, perform semantic granularity processing on the genuine regional drug retrieval request data to obtain semantic granularity data; generate hierarchical semantic edges according to the semantic granularity data to obtain hierarchical semantic edge data;

[0006] S2. Generate a hierarchical semantic graph for a preset genuine regional drug knowledge graph according to the hierarchical semantic edge data to obtain hierarchical semantic graph data;

[0007] S3. Perform semantic annotation on the relationship edges of the hierarchical semantic graph data to obtain semantic annotation data;

[0008] S4. Perform multi-condition intention processing according to the genuine regional drug retrieval request data to obtain multi-condition intention data; perform multi-hop path combination on the semantic annotation data according to the multi-condition intention data to obtain preliminary retrieval data; perform edge reasoning and complementation on the preliminary retrieval data to obtain genuine regional drug retrieval data.

[0009] In the present invention, by performing semantic granularity processing on the data of the retrieval request for genuine regional drugs and generating hierarchical semantic edge data, the multi-dimensional semantic parsing of the user's intention is effectively realized; further, by constructing a hierarchical semantic graph and semantic annotation operations, the inferability of the graph structure and the semantic expression ability are improved. By using multi-condition intention processing and multi-hop path combination, it not only supports the intelligent matching of drugs under various conditions, but also can discover potential drug nodes that are not explicitly represented in the original graph but have semantic rationality through the edge reasoning completion mechanism, thereby significantly improving the accuracy, flexibility of the retrieval and the intelligence of the recommendation, and providing an intelligent knowledge retrieval method with both deep semantic understanding and structural reasoning ability for the field of traditional Chinese medicine.

[0010] Optionally, S1 includes:

[0011] Obtain the data of the retrieval request for genuine regional drugs;

[0012] Perform rough semantic granularity processing on the data of the retrieval request for genuine regional drugs to obtain rough semantic granularity data;

[0013] Perform fine semantic granularity processing on the data of the retrieval request for genuine regional drugs to obtain fine semantic granularity data, where the semantic granularity processing is carried out according to the preset semantic ontology of traditional Chinese medicine and the granularity stratification rule. The rough semantic granularity is the drug category, regional attribute and basic efficacy, and the fine semantic granularity is the meridian tropism, drug nature, collection season and classical compatibility; <s

[0014] Perform semantic path construction and multi-level edge weight calculation on the rough semantic granularity data and the fine semantic granularity data to obtain semantic path data and multi-level edge weight data;

[0015] Generate a multi-granularity semantic hierarchy graph according to the semantic path data and the multi-level edge weight data to obtain multi-granularity semantic hierarchy graph data;

[0016] Perform semantic edge classification annotation according to the multi-granularity semantic hierarchy graph data to obtain hierarchical semantic edge data.

[0017] In the present invention, by dividing the authentic medicinal material retrieval request data into two levels of coarse semantic granularity and fine semantic granularity, and processing it in combination with the traditional Chinese medicine semantic ontology and the granularity stratification rule, the user's query intention can be comprehensively expressed from different semantic depth dimensions. Among them, the coarse-grained processing can quickly locate the medicinal material category, region, and basic efficacy, while the fine-grained processing further extracts fine features such as meridian tropism, medicinal properties, collection season, and classic compatibility, improving the hierarchy and accuracy of semantic expression. At the same time, by constructing a semantic path and performing multi-layer edge weight calculation, a quantitative modeling of the structured semantic path is realized, and further a multi-granularity semantic hierarchy graph and semantic edge classification annotation data are generated, providing a high-quality structural basis for graph reasoning. Compared with the prior art, this step not only supports the multi-granularity fusion modeling of user intentions, but also enhances the hierarchical expression ability of the graph structure and the context semantic computability, significantly improving the semantic understanding depth and reasoning accuracy of authentic medicinal material retrieval, and having stronger flexibility and intelligent adaptability.

[0018] Optionally, S2 includes:

[0019] Performing a function attribution semantic perception mapping and a regional co-occurrence semantic perception mapping on a preset authentic medicinal material knowledge graph according to the hierarchical semantic edge data, respectively obtaining function attribution mapping data and regional co-occurrence mapping data;

[0020] Performing node twin expansion according to the function attribution mapping data and the regional co-occurrence mapping data to obtain semantic completion data;

[0021] Performing structure clustering according to the semantic incomplete data to obtain hierarchical semantic graph data.

[0022] In the present invention, based on the hierarchical semantic edge data, a multi-dimensional semantic structure perception of the "medicinal property-efficacy-meridian tropism" path and the "region-season-medicinal material" path in the authentic medicinal material graph is realized by performing a function attribution semantic perception mapping and a regional co-occurrence semantic perception mapping respectively, effectively constructing a context semantic view for reasoning. Through the node twin expansion mechanism, combined with graph structure similarity, attribute feature consistency, and path co-occurrence rules, potential but not explicitly expressed medicinal material nodes are automatically identified to generate semantic completion data, thereby improving the coverage ability of the graph in the knowledge sparse area. Subsequently, by performing structure clustering processing on the completion data, aggregating subgraphs are constructed according to semantic similarity, forming a hierarchical semantic graph structure with classification, scalability, and reasoning consistency. Compared with the single static path matching method in the prior art, the present invention introduces a dual-channel semantic perception, node twin intelligent expansion, and structure hierarchical reconstruction mechanism, significantly enhancing the semantic relevance, structure reasoning ability, and dynamic completion ability of the knowledge graph, making the authentic medicinal material retrieval more intelligent and generalized.

[0023] Optionally, the function attribution semantic perception mapping is specifically:

[0024] Filter the functional semantic trigger edges according to the hierarchical semantic edge data to obtain the functional semantic edge data;

[0025] Construct the functional semantic path of the preset authentic medicinal materials knowledge graph according to the functional semantic edge data to obtain the functional semantic path sub-graph data;

[0026] Perform semantic path context verification according to the hierarchical semantic edge data and the functional semantic path sub-graph data to obtain the functional context verification data;

[0027] Detect the semantic coverage conflict of the functional context verification data to obtain the functional semantic coverage conflict data;

[0028] Compare the mapping accuracy according to the functional semantic coverage conflict data to obtain the functional attribution mapping data.

[0029] In the present invention, by constructing a multi-stage and multi-level functional semantic reasoning path processing mechanism, the semantic matching depth and reasoning accuracy between medicinal materials and functional attribution are significantly improved. Filter the functional semantic trigger edges according to the hierarchical semantic edge data, effectively filter out the key paths related to the core functional attributes such as medicinal properties, main indications, and meridian tropism, and the constructed functional semantic path sub-graph data retains the structural context characteristics of functional attribution. By performing context verification and semantic consistency analysis on the path, the path deviation caused by semantic ambiguity or context ambiguity can be avoided. The proposed semantic coverage conflict detection mechanism can identify the semantic conflicts in functional attribution between multiple paths. Combining with the mapping accuracy comparison mechanism, the optimal functional attribution path mapping result under the current query intention is finally selected. Compared with the existing method of mapping and matching only based on keywords or single-path scoring, the present invention adopts structure triggering and path context verification to realize the fusion mechanism of conflict detection and multi-path accuracy competition, improves the system's understanding ability and semantic discrimination ability of functional attribution intention, has stronger anti-interference, interpretability and semantic discrimination ability, and has significant advantages in the reasoning accuracy and scenario adaptability of the knowledge graph.

[0030] Optionally, the geographical co-occurrence semantic perception mapping is specifically as follows:

[0031] Extract the region-related trigger nodes according to the hierarchical semantic edge data to obtain the region node data;

[0032] Perform the medicinal material geographical co-occurrence path mapping on the preset authentic medicinal materials knowledge graph according to the region node data to obtain the medicinal material geographical co-occurrence data;

[0033] Expand the semantic path of the medicinal material geographical co-occurrence data to obtain the region co-occurrence candidate path data;

[0034] Calculate the path co-occurrence weight for the region co-occurrence candidate path data to obtain the path co-occurrence weight data;

[0035] Generate the regional co-occurrence mapping data based on the path co-occurrence weight data and the region co-occurrence candidate path data to obtain the regional co-occurrence mapping data.

[0036] In the present invention, by constructing a semantic co-occurrence reasoning mechanism based on the multi-dimensional path of geography-medicinal material-function, a systematic modeling of the distribution relationship, harvesting characteristics, and functional attribution of genuine medicinal materials under different regional environments and time conditions is realized. Extracting region-related trigger nodes based on the hierarchical semantic edge data enables the system to perceive the geographical regions or administrative divisions involved in the user's intention and construct a region-dominated reasoning entry. By constructing the medicinal material region co-occurrence path and unfolding the semantic path, the medicinal material information related to the region can be comprehensively captured, including the composite paths related to the harvesting season, ecological environment, local uses, etc. Through the calculation of path co-occurrence weights, combined with factors such as historical literature frequency, co-occurrence relationship of classics, and regional suitability, semantic intensity scores are given to each path. The generated regional co-occurrence mapping data not only has a clear spatial positioning ability but also has the characteristic of multi-hop path semantic fusion. Different from the static matching method of single query by "production area field" or regional label in the prior art, the present invention significantly improves the semantic accuracy and retrieval coverage of the regional attribution of medicinal materials by introducing the region trigger, semantic path unfolding, and co-occurrence weight modeling mechanism, realizing dynamic, interpretable, and structure-aware geographical semantic reasoning ability, and having stronger practicality and novelty in the scenario of genuine medicinal materials.

[0037] Optionally, the node twin expansion is specifically as follows:

[0038] Extract the structurally similar node data based on the functional attribution mapping data and the regional co-occurrence mapping data to obtain the structurally similar node data;

[0039] Calculate the similarity of graph nodes based on the structurally similar node data to obtain the graph node similarity data;

[0040] Perform twin node path inversion based on the graph node similarity data to obtain the twin node path data;

[0041] Calculate the path confidence based on the twin node path data to obtain the path confidence data;

[0042] Generate semantic completion data for the twin node path data, functional attribution mapping data, and regional co-occurrence mapping data based on the path confidence data, the semantic completion data.

[0043] In the present invention, by introducing a structural similarity analysis and path inversion reasoning mechanism, semantic completion of potential but not explicitly associated medicinal material nodes is achieved on the basis of the existing knowledge graph, significantly improving the knowledge integrity and structural generalization ability of the genuine medicinal material knowledge graph. Structural similar nodes are extracted based on the function attribution mapping data and the geographical co-occurrence mapping data, and the similarity calculation of the knowledge graph nodes is carried out. Starting from multi-dimensional features such as structural adjacency relationships, attribute similarities, and context co-occurrence frequencies, the potential semantic consistency between different nodes is quantified. Through twin node path inversion, taking the highly similar nodes as the starting point, a reverse semantic path similar to the known path structure is constructed to form a set of semantic approximate paths. After confidence scoring of the inverted paths, the rationality of using them as missing paths can be effectively evaluated, and based on this, combined with the original function attribution and geographical co-occurrence results, semantic completion data is generated. Different from the existing knowledge graph technologies that only rely on static entity relationships or node similarities to complete the missing edges, the present invention not only improves the semantic rationality and path consistency of the completion results, but also has stronger context adaptation ability and knowledge graph reasoning transparency, and has stronger practical value in semantic prediction and discovery of unknown knowledge.

[0044] Optionally, S3 includes:

[0045] Extract the node semantic nested context of the hierarchical semantic graph data to obtain edge context semantic data;

[0046] Generate multi-source semantic feature tensors for the hierarchical semantic graph data according to the edge context semantic data to obtain edge feature data to be labeled;

[0047] Calculate the path probability according to the edge feature data to be labeled to obtain path probability data;

[0048] Perform functional analogy annotation on the edge feature data to be labeled according to the path probability data to obtain preliminary semantic annotation data;

[0049] Detect edge semantic conflicts in the preliminary semantic annotation data to obtain semantic conflict detection data;

[0050] Complete the relationships in the preliminary semantic annotation data according to the semantic conflict detection data to obtain semantic annotation data.

[0051] In the present invention, by constructing a context nested semantic expression and multi-source feature fusion mechanism for the edges of the graph, the depth of edge semantic understanding and the accuracy of edge relationship annotation in the structural reasoning process of the hierarchical semantic graph are improved. Through node semantic nested context extraction, the implicit semantic collaboration relationship between nodes in the path is obtained and transformed into edge context semantic data, providing structured semantic support for annotation. A semantic feature tensor is constructed by fusing multi-source information such as node attributes, path patterns, and historical co-occurrence frequencies, obtaining edge feature data to be annotated with high-dimensional expression, enhancing the expression ability of edge features. Subsequently, a path probability calculation module is introduced. Based on the co-occurrence probability of the edge in the functional path and combined with the semantic intention similarity, functional analogy annotation is performed to automatically infer the functional role of the edge, realizing the automatic annotation ability driven by path semantics. To ensure structural consistency, the system further performs edge semantic conflict detection, identifies semantic ambiguity or mutually exclusive relationships, and uniformly optimizes and corrects the edge semantics through a relationship completion strategy to form a complete semantic graph annotation result. Different from the static rules or single-sided attribute annotation methods commonly used in the prior art, the present invention significantly improves the structural interpretability of edge annotation, the consistency of functional reasoning, and the semantic error correction ability, has stronger intelligence and automatic expansion ability, and achieves a key breakthrough in the quality of knowledge graph annotation and the reasoning efficiency.

[0052] Optionally, S4 includes:

[0053] Retrieve request feature extraction is performed on the genuine regional drug retrieval request data to obtain retrieve request feature data, where the retrieve request feature extraction includes regional preference feature extraction, meridian tropism feature extraction, and medicinal property restriction feature extraction;

[0054] Intent attention calculation is performed on the retrieve request feature data to obtain multi-condition intent data;

[0055] An intent path candidate space is constructed for the semantic graph annotation data based on the multi-condition intent data to obtain intent path candidate space data;

[0056] Graph traversal is performed based on the intent path candidate space data to obtain path candidate graph data;

[0057] Semantic coverage calculation is performed based on the path candidate graph data and the retrieve request feature data to obtain semantic coverage data;

[0058] The path candidate graph data is filtered based on the semantic coverage data to obtain path filtering data;

[0059] Edge node semantic filling processing is performed on the path filtering data to obtain genuine regional drug retrieval data.

[0060] In the present invention, a linkage mechanism of intention recognition - path construction - semantic completion driven by user retrieval features is constructed, realizing a closed-loop reasoning process from natural language retrieval intention to structured knowledge path matching and complete result generation. For the retrieval request data of authentic Chinese medicinal materials, multi-dimensional features such as regional preference, meridian tropism, and property limitation are extracted, and a multi-angle feature modeling framework for personalized semantic understanding is established. On this basis, a multi-condition intention representation is formed through intention attention calculation, effectively improving the context focusing ability during semantic graph matching. Subsequently, a path candidate space is constructed according to the intention features, and path candidate graph data is generated through a structure traversal method, ensuring that the retrieval is not limited to direct explicit paths but also covers combinable paths. Further, by combining the retrieval intention and the graph structure to calculate the semantic coverage rate, the semantic satisfaction degree of the candidate path for the user's needs can be accurately measured, and a quantitative evaluation of the intention alignment degree is realized. After screening by the semantic coverage rate, high-matching paths are retained, and semantic filling is performed on the edge nodes to automatically complete missing information such as the harvesting season and regional attribution, generating a retrieval result of authentic Chinese medicinal materials with complete semantics and a closed structure. Different from the prior art which mainly relies on keyword indexing or shallow entity matching methods, the present invention proposes a structure path construction and completion strategy driven by intention features, integrating the attention mechanism, path coverage calculation, and edge semantic completion, significantly improving the intelligent matching accuracy, semantic adaptation ability, and retrieval generalization ability of the system.

[0061] Optionally, the semantic filling process of the edge nodes is specifically as follows:

[0062] Perform boundary node analysis on the path screening data to obtain half-edge connection node data;

[0063] Generate edge node missing data according to the half-edge connection node data;

[0064] Perform semantic path missing type mapping according to the edge node missing data to obtain semantic path missing type data;

[0065] Perform attribute near neighbor matching according to the semantic path missing type data to obtain supplementary node data;

[0066] Perform similarity screening on the supplementary node data according to the retrieval request feature data to obtain supplementary screening data;

[0067] Integrate the supplementary screening data and the path screening data to obtain the retrieval data of authentic Chinese medicinal materials.

[0068] In the present invention, the system identifies half-connected nodes with structural breaks at the end or start of a path through boundary node analysis, then generates edge missing nodes based on the path structure and semantic tags, and performs semantic path missing type mapping to clarify that the missing part belongs to types such as origin missing, meridian tropism missing, or harvesting season missing. By combining the node data existing under similar paths in the atlas for attribute near-neighbor matching, candidate nodes for supplementation are mined from dimensions such as medicinal material category, medicinal property, and meridian tropism, and semantic similarity is calculated based on the retrieval request features to screen out the supplementary nodes closest to the user's needs. The supplementary nodes are fused with the original path to generate a complete, semantically rich, and target-aligned retrieval result for authentic medicinal materials. Different from the prior art where path missing is usually processed by manual annotation or static templates, the present invention constructs an automatic completion mechanism based on semantic type mapping and near-neighbor reasoning, significantly improving the system's fault tolerance, reasoning ability, and output integrity for incomplete paths, and ensuring that even in the case of holes in the knowledge graph, intelligent recommendation results with reasonable structures and accurate semantics can still be provided.

[0069] Optionally, the present application also provides a semantic-graph-based intelligent retrieval system for authentic medicinal materials, which is used to execute the semantic-graph-based intelligent retrieval method for authentic medicinal materials as described above. The semantic-graph-based intelligent retrieval system for authentic medicinal materials includes:

[0070] A semantic granularity analysis module, which is used to obtain the retrieval request data for authentic medicinal materials, perform semantic granularity processing on the retrieval request data for authentic medicinal materials to obtain semantic granularity data, and generate hierarchical semantic edges based on the semantic granularity data to obtain hierarchical semantic edge data;

[0071] A hierarchical semantic graph construction module, which is used to generate a hierarchical semantic graph based on the hierarchical semantic edge data for a preset knowledge graph of authentic medicinal materials to obtain hierarchical semantic graph data;

[0072] A semantic edge annotation module, which is used to perform relational edge semantic annotation on the hierarchical semantic graph data to obtain semantically annotated data;

[0073] An intelligent semantic-driven retrieval module, which is used to perform multi-condition intention processing based on the retrieval request data for authentic medicinal materials to obtain multi-condition intention data, perform multi-hop path combination on the semantically annotated data according to the multi-condition intention data to obtain preliminary retrieval data, and perform edge reasoning and completion on the preliminary retrieval data to obtain retrieval data for authentic medicinal materials.

[0074] The object of the present invention is to perform semantic granularity processing on the retrieval request for genuine regional drugs, and generate hierarchical semantic edge data based on the ontology of traditional Chinese medicine knowledge. It not only realizes the multi-granularity understanding of dimensions such as region, drug property, and meridian tropism, but also enhances the expression ability of the semantic structure in the knowledge graph. Step S2 constructs a hierarchical semantic graph, enabling the knowledge graph to have the flexible combination ability of multi-level structure representation and reasoning path. Step S3 improves the interpretability and structural integrity of the semantic knowledge graph through context nesting, semantic tensor fusion, and edge relationship annotation driven by path probability. Step S4 combines the multi-condition intention input of the user, constructs multi-hop path combinations, and implements an edge node complementation strategy in the reasoning results to fill the semantic void area and achieve more complete and accurate drug recommendation. Compared with the prior art, the overall method no longer relies on single entity matching or shallow rule retrieval, but constructs a dynamic retrieval framework with multi-dimensional semantic understanding, graph structure driving, and reasoning and complementation capabilities, significantly improving the adaptability, intelligence, and retrieval accuracy of the system when facing fuzzy, complex, or incomplete retrieval intentions. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Other features, objects, and advantages of the present application will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:

[0076] Figure 1 FIG. shows a flowchart of the steps of a method for intelligent retrieval of genuine regional drugs based on a knowledge graph in an embodiment;

[0077] Figure 2 FIG. shows a flowchart of the steps of a method for semantic granularity analysis in an embodiment;

[0078] Figure 3 FIG. shows a flowchart of the steps of a method for constructing a hierarchical semantic graph in an embodiment;

[0079] Figure 4 FIG. shows a flowchart of the steps of a method for semantic edge annotation in an embodiment;

[0080] Figure 5 FIG. shows a flowchart of the steps of a method for intelligent semantic-driven retrieval in an embodiment;

[0081] The implementation, functional features, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0082] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0083] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The functional entities may be implemented in the form of software, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor methods and / or microcontroller methods.

[0084] It should be understood that although terms such as "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly, the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.

[0085] Please refer to Figures 1 to 5 , this application provides a method for intelligent retrieval of authentic Chinese medicinal materials based on a spectrum, and the method includes:

[0086] S1. Obtain the authentic Chinese medicinal material retrieval request data, perform semantic granularity processing on the authentic Chinese medicinal material retrieval request data to obtain semantic granularity data; generate hierarchical semantic edges according to the semantic granularity data to obtain hierarchical semantic edge data;

[0087] In one embodiment, the acquisition methods of the authentic medicinal materials retrieval request data include, but are not limited to, keywords input by users, query statements, or retrieval questions expressed in natural language form. Taking "What are the similar medicinal materials to Fritillaria cirrhosa for treating cough?" as an example, this request belongs to a natural language retrieval statement. When performing semantic granularity processing on the retrieval request data, a natural language processing model can be used for word vector encoding, such as the BERT or ERNIE model or other preset large text classification models, to enhance the expression ability of semantic features. After completing the encoding, combined with part-of-speech tagging and named entity recognition methods, the key semantic entities contained therein are identified, including but not limited to medicinal material names, medicinal material properties, symptom names, usage intentions, etc. The above entity types can be classified and summarized according to a preset traditional Chinese medicine semantic tag system to form structured semantic granularity data. For example, after processing the above example statement, the obtained semantic granularity data includes: the medicinal material name is "Fritillaria cirrhosa", the symptom name is "cough", the functional intention is "treatment", and the content of the rest that does not belong to the predefined tag system can be treated as other items. Further, a hierarchical semantic edge is constructed according to the semantic granularity data. This hierarchical structure is set based on the semantic relationship of the traditional Chinese medicine knowledge graph, setting a multi-level semantic hierarchy from "symptom" to "functional intention", then to "medicinal material entity", and extending to "similar medicinal materials". Specifically, the semantic path constructed according to the above example can be expressed as: the "cough" node is connected to the "treatment" node through the "has_effect" edge, the latter is connected to the "Fritillaria cirrhosa" node through the "suitable medicinal material" edge, and then is connected to the "Fritillaria thunbergii" node through the "similar_to" edge. The above edges can be specifically labeled as semantic edge types such as "has_symptom" (associated symptom), "has_effect" (has effect), "similar_to" (similar medicinal materials) in the semantic graph. The edges have directions to represent the semantic flow between entities and can be given different semantic labels and weights. The "medicinal property" information of medicinal materials can not only come from the records in traditional pharmacopoeias, but also be comprehensively modeled by combining the microbial components, endophytic fungi, or actinomycete metabolites of medicinal materials obtained in modern pharmaceutical research. For example, in the study of the authentic medicinal material Magnolia officinalis, through pure culture, gene sequence alignment, and potential metabolic function analysis of the endophytic actinomycetes isolated from its roots, stems, bark and other tissues, its modern pharmacological properties such as antibacterial and anti-inflammatory can be deduced. These pharmacological properties can be further mapped to the functional medicinal property expressions such as "clearing heat", "detoxifying", and "resolving dampness" under the traditional Chinese medicine theory system, and then used as the support source for the "medicinal property" node of the medicinal material in the graph.

[0088] S2. Generate a hierarchical semantic graph for the preset authentic medicinal materials knowledge graph according to the hierarchical semantic edge data to obtain hierarchical semantic graph data;

[0089] In one embodiment, the authentic medicinal materials knowledge graph is a preset semantic structure graph constructed based on the traditional Chinese medicine field, and its basic structure is in the form of triples, expressed as the structure form of "entity 1 - relationship - entity 2". For example, it may include semantic triples such as "there is an 'applicable' relationship between the symptom 'cough' and the medicinal material 'Fritillaria cirrhosa'", "the medicinal material 'Fritillaria cirrhosa' belongs to the 'Sichuan-produced medicinal materials' region", and "there is a 'homogeneous medicinal materials' relationship between the medicinal material 'Fritillaria cirrhosa' and 'Fritillaria thunbergii'". Each pair of head and tail entities in the above triples is connected through a specific edge relationship, forming the edge structure in the graph, which is convenient for subsequent entity connection and information dissemination in the semantic dimension. In the process of generating the hierarchical semantic graph, preferably, a semantic graph construction tool that supports knowledge graph processing can be called to perform graph structure generation operations on the matched semantic edges and their corresponding entities. During the construction process, semantic attribute annotation can be performed on the entity nodes connected by each semantic edge. The semantic attributes of the entity nodes may include, but are not limited to, the origin information, ingredient composition, historical usage frequency, and usage literature coverage of the medicinal material; each semantic edge can also carry edge attribute information, such as edge weight (representing the connection strength or co-occurrence probability) and semantic confidence (representing the credibility of the relationship). Further, the hierarchical semantic graph is organized according to the functional semantic levels. Preferably, according to the semantic category division in the authentic medicinal materials knowledge graph, the graph structure can be divided into multiple semantic levels, including but not limited to the symptom layer, the function layer (representing the efficacy attribution), the medicinal material layer (representing the candidate traditional Chinese medicine entities), and the region / homogeneous expansion layer (for expanding regional information and homogeneous medicinal material relationships). The connections between the above semantic levels have directions. For example, the symptom layer is connected to the function layer, the function layer is connected to the medicinal material layer, and the medicinal material layer is then connected to the regional information or homogeneous medicinal material information, forming a hierarchical graph structure that can be used for subsequent path reasoning and node annotation. Preferably, during the graph traversal process, the generation depth of the graph can be controlled according to the set path jump number threshold. For example, when performing multi-hop path reasoning, the maximum traversal depth is limited to no more than three layers to avoid information drift problems during semantic expansion and ensure that the constructed hierarchical semantic graph has good domain focus and reasoning controllability.

[0090] S3. Perform semantic annotation on the relationship edges of the hierarchical semantic graph data to obtain semantic annotation data of the graph;

[0091] In one embodiment, all edge relationships in the semantic graph can be semantically annotated based on a preset edge semantic mapping dictionary. The edge semantic mapping dictionary is used to define the correspondence between edge types in the knowledge graph and natural language semantic labels, so as to facilitate the subsequent visual expression of the graph structure and semantic reasoning processing. The specific mapping relationships may include, but are not limited to, mapping "has_symptom" to "can be associated with symptoms"; "from_region" to "origin of production"; "similar_to" to "similar in medicinal properties"; "used_with" to "commonly paired medicinal materials"; other edge types can also be flexibly extended according to the requirements of knowledge graph construction. When performing annotation, according to the edge type of each edge, the corresponding natural language description can be found in the semantic mapping dictionary, and the semantic label can be assigned to the corresponding edge to form structured semantic graph annotation data. The annotation result of each edge includes the starting entity, the ending entity, and their corresponding semantic labels. For example, if the "Fritillaria cirrhosa" node points to the "cough" node and the edge type is "has_symptom", then this edge can be annotated as "can be associated with symptoms". Further, optionally, semantic intensity weight calculation is performed on the already annotated edges. The edge weight is used to measure the importance of this edge in the semantic graph structure, and thus serves as a sorting basis in subsequent path selection and multi-hop path reasoning. The calculation method of the weight may include, but is not limited to, using the term frequency-inverse document frequency (TF-IDF) method, and calculating the relative importance of the edge according to the co-occurrence frequency of the entities connected by the edge in the knowledge base or the corpus of prescription texts; based on the above calculation results, a normalized edge weight value can be assigned to each edge to represent its semantic association strength.

[0092] S4. Perform multi-condition intention processing according to the genuine medicinal materials retrieval request data to obtain multi-condition intention data; perform multi-hop path combination on the semantic graph annotation data according to the multi-condition intention data to obtain preliminary retrieval data; perform edge reasoning and complementation on the preliminary retrieval data to obtain genuine medicinal materials retrieval data.

[0093] In one embodiment, the multi-condition intention processing includes extracting intention dimensions from the retrieval request input by the user. The user request usually contains multiple semantic dimensions, and typical dimensions include, but are not limited to, target symptoms, functional intentions, regional preferences, common medicinal material collocation relationships, and the five-taste and meridian tropism attributes of medicinal properties. To achieve the above intention extraction, an intention classification model based on natural language understanding can be preset, and a multi-label classification mechanism can be used to perform structured parsing on the original retrieval statement to obtain a set of semantic labels covering multiple dimensions. For example, for a retrieval statement containing information such as "cough", "Sichuan", and "clearing heat and resolving phlegm", after parsing, it can be extracted that the "symptom" is "cough", the "regional preference" is "Sichuan", and the "functional preference" is "clearing heat and resolving phlegm". After obtaining the multi-condition intention data, path combination operations can be performed based on the semantic graph annotation data. Preferably, a weighted graph search strategy is adopted, such as A An algorithm or a breadth-first search (BFS) strategy with a limited depth is used to perform multi-hop path traversal on the nodes and edges in the semantic graph to search for a path sequence that meets multiple intention constraints of the user. The nodes in the path need to meet the intention constraint conditions. For example, the starting node is "cough", the intermediate nodes are related to "clearing heat", and the end node is "Fritillaria cirrhosa" or its similar medicinal materials. During the path combination process, preferably, path screening rules are set to preferentially retain the paths with higher intention hit rates.

[0094] Perform a semantic edge completion operation on the above-mentioned preliminary retrieval paths to repair the path breakpoints or relationship missing areas in the semantic graph. The completion operation can be implemented through a graph neural network model (such as a relational graph convolutional network R-GCN) or a rule-based graph structure inference engine. The completion logic can include the following two categories. If two medicinal materials frequently co-occur in multiple prescriptions, it can be inferred that they are similar medicinal materials. If a medicinal material A has been marked as having a semantic association with a certain semantic entity (such as "cough"), and there is a high structural similarity or co-occurrence degree between medicinal material B and A, it can be inferred that medicinal material B also has a potential association with this semantic entity. For example, in the original semantic graph, if there is no direct semantic connection established between "Thunberg Fritillary Bulb" and "cough", but there is a semantic edge from "Fritillaria cirrhosa" to "cough", and there is a semantic edge of "similar medicinal materials" between "Fritillaria cirrhosa" and "Thunberg Fritillary Bulb", then based on the above rules, the edge relationship between "Thunberg Fritillary Bulb" and "cough" can be inferred. The confidence score of the inferred edge can be evaluated by a similarity model. If the score is 0.76, then this relationship can be added as a completed edge to the graph structure. Combining the preliminary retrieval paths with the inference completion results, structured retrieval data of genuine regional medicinal materials including fields such as the name of the target medicinal material, origin attribution, functional attributes, and recommended paths is output.

[0095] Optionally, S1 includes:

[0096] S11. Obtain the retrieval request data of genuine regional medicinal materials;

[0097] In one embodiment, a Chinese text tokenization tool (such as Jieba tokenizer or a custom tokenization module built based on traditional Chinese medicine professional corpus) is used to perform preliminary tokenization on the retrieval request text to obtain a sequence of basic lexical units. The system applies a named entity recognition model (NER) to the above-mentioned lexical unit sequence. This model can be constructed based on deep learning or rule-driven methods and is used to identify semantic entities related to the field of traditional Chinese medicine. The recognized entity types include at least, but are not limited to, the following categories: symptom entities, which represent physiological or perceptual characteristics mentioned in the user input, such as "cough", "dry mouth", etc.; medicinal material entities, which represent the names of traditional Chinese medicines with geographical attributes or efficacy attributes, such as "Fritillaria cirrhosa", "Astragalus membranaceus", etc.; function entities, which represent the common traditional usage functions of medicinal materials, such as "clearing heat", "tonifying qi", "removing dampness", etc.; geographical entities, which represent the names of geographical regions related to the genuine producing areas, such as "Yunnan", "Daozhou", "Bozhou", etc. The named entity recognition model can be trained based on structures such as BERT, BiLSTM-CRF, etc., or constructed by means of matching a rule base with a dictionary, and its purpose is to map the user input into structured information.

[0098] S12. Coarsely process the retrieval request data of genuine medicinal materials to obtain coarsely grained semantic data;

[0099] In one embodiment, the system uses a preset dictionary of genuine medicinal material attributes to match the words in the user input text. This dictionary includes, but is not limited to, the names of medicinal material categories, geographical attribute words, and common efficacy description phrases. The matching process can be completed by means of string similarity, stemming matching, etc., and the terms belonging to specific semantic fields in the text are initially identified. The system introduces a vector representation mechanism based on a pre-trained language model (such as the BERT model) to calculate the semantic similarity between the terms in the text and the standard field terms. The similarity calculation is based on the cosine similarity of word vectors. Based on the above matching and calculation processes, the system can classify the keywords or phrases in the user input into the following three categories of coarsely grained semantic fields: medicinal material category fields, such as "rhizome category", "flower category", "fruit category", etc., which represent the basic plant part attribution of traditional Chinese medicines; geographical attribute fields, such as "produced in Sichuan", "produced in Zhejiang", "Lingnan", etc., which represent the geographical origin characteristics of genuine medicinal materials; basic efficacy fields, such as "clearing heat", "moistening the lungs", "tonifying deficiency", etc., which represent the general function labels widely used in traditional Chinese medicine to describe the common application directions of medicinal materials (not referring to specific diseases or treatment behaviors).

[0100] S13. Finely process the retrieval request data of genuine medicinal materials to obtain finely grained semantic data, where the semantic granularity processing is carried out according to the preset traditional Chinese medicine semantic ontology and granularity stratification rules. The coarsely grained semantic granularity is the medicinal material category, geographical attribute, and basic efficacy, and the finely grained semantic granularity is meridian tropism, medicinal property, collection season, and classical compatibility;

[0101] In one embodiment, the fine-grained semantic processing is based on a preset traditional Chinese medicine semantic ontology structure and semantic granularity layering rules. The semantic granularity levels include coarse-grained semantic items, including medicinal material categories, regional attributes, and basic efficacy, etc.; fine-grained semantic items, including meridian attribution (such as belonging to the heart meridian, lung meridian, kidney meridian, etc.), medicinal property characteristics (such as the four natures of cold, heat, warm, and cool, and the five flavors of pungent, sweet, bitter, sour, and salty), collection seasons (such as "autumn", "before winter", "after frost's descent", etc.), and classical compatibility information (such as "compatible with almonds", "used in the formula of Qingfei Decoction", etc.). The system performs fine-grained semantic entity extraction on the retrieved content based on the constructed traditional Chinese medicine semantic ontology structure and granularity layering rule table. The extracted semantic items include meridian information, medicinal property characteristics, collection seasons, and classical compatibility, etc. The system realizes the intelligent recognition of classical compatibility relationships by constructing a traditional Chinese medicine knowledge graph and its supporting semantic inference rules. For example, when the co-occurrence frequency of a certain medicinal material and a specific formula in the pharmacopoeia literature is higher than the set threshold, it can be inferred that there is a typical compatibility relationship. In addition, by introducing traditional Chinese medicine literature such as "Compendium of Materia Medica", a vectorized text index structure is constructed, which improves the accuracy and coverage of semantic extraction and reasoning.

[0102] S14. Construct semantic paths and calculate multi-level edge weights for the coarse semantic granularity data and the fine semantic granularity data to obtain semantic path data and multi-level edge weight data;

[0103] In one embodiment, construct a semantic path starting from a query entity (such as "moistening the lungs"); each path is in the form of a triple link: [moistening the lungs]—can be attributed to→[lung meridian]—medicinal property→[sweet, warm]—origin→[Sichuan]—→[Fritillaria cirrhosa]. Each semantic edge is assigned a weight, and the weight of each edge is defined as: , is the weight coefficient of the ontology structure distance or cosine similarity between entities, with a value of 0.4, is the ontology structure distance or cosine similarity between entities, is the weight of whether the edge relationship comes from an authoritative knowledge ontology, with a value of 0.3, is whether the edge relationship comes from an authoritative knowledge ontology, is the weight of the co-occurrence frequency in historical used prescriptions, with a value of 0.3, is the co-occurrence frequency in historical used prescriptions.

[0104] S15. Generate a semantic level graph according to the semantic path data and the multi-level edge weight data to obtain multi-granularity semantic level graph data;

[0105] In one embodiment, the graph nodes in the semantic hierarchy graph are divided into multiple levels according to semantic granularity. The typical structure includes but is not limited to the following three levels. The first level (function category) represents the general functional directions of medicinal materials, such as "moistening the lungs", "tonifying deficiency", "drying dampness", etc., which belong to coarse-grained semantic entities. The second level (traditional Chinese medicine attributes) includes the meridians that the medicinal materials enter (such as "lung meridian", "spleen meridian"), the medicinal properties (such as "sweet", "warm", "cold"), and the application seasons (such as "spring", "autumn"), etc., serving as an intermediate layer between the function category and the specific medicinal materials. The third level (specific medicinal material entities) represents specific traditional Chinese medicine examples, such as "Fritillaria cirrhosa", "Astragalus membranaceus", "Ophiopogon japonicus", etc., which belong to fine-grained semantic entities. The above levels are connected by relationship edges in the semantic path data to establish a graph structure. Each edge is weighted and modeled according to its semantic relationship type and edge weight data. For example, an indirect connection is established between "moistening the lungs" and "Fritillaria cirrhosa" through the path of "lung meridian → warm nature → autumn", and each jump edge has a corresponding semantic label and edge weight value.

[0106] S16. Perform semantic edge classification and annotation based on the multi-granularity semantic hierarchy graph data to obtain hierarchical semantic edge data.

[0107] In one embodiment, the semantic types of the edges may include but are not limited to the following categories: efficacy edge (function attribution relationship), medicinal property edge (taste and property attribute relationship), region edge (production area attribution relationship), or compatibility edge (common collocation relationship). Define a set of matching rules for known semantic patterns, and perform category annotation on the edges that meet the patterns. For example, the following is an example description of the matching logic of a rule template. If the starting node type of the edge connection is "region", the ending node type is "medicinal material", and the edge relationship is "source_of" or a synonymous relationship, then the edge is labeled as "production area attribution". Similarly, rules templates such as "function node → function node" corresponding to "efficacy association" and "taste and property attribute → medicinal material entity" corresponding to "medicinal property pointing" can also be set. The rule templates can be stored as a structured matching rule set, and combined with entity type judgment, path structure verification, etc. to complete the semantic annotation operation of the edges.

[0108] Optionally, S2 includes:

[0109] S21. Perform function attribution semantic perception mapping and regional co-occurrence semantic perception mapping on the preset genuine medicinal materials knowledge graph according to the hierarchical semantic edge data, and obtain function attribution mapping data and regional co-occurrence mapping data respectively;

[0110] In one embodiment, the connection relationship between the medicinal material and its "attributed functional layer" is identified. For example, the functional semantic landing point of Fritillaria cirrhosa is confirmed from the path of "moistening the lungs → lung meridian → Fritillaria cirrhosa". Extract the path of "function → meridian tropism → medicinal material" from the hierarchical semantic edge data; map the medicinal material node to its corresponding functional attribution label, such as "medicinal materials for moistening the lungs"; introduce a similarity function to support fuzzy matching: , is, is the functional semantic keyword, is the name or descriptive text of the medicinal material node, is the functional word vector, is the medicinal material-related word vector.

[0111] S22. Perform node twin expansion according to the functional attribution mapping data and the regional co-occurrence mapping data to obtain semantic completion data;

[0112] In one embodiment, by jointly analyzing the functional labels and regional attributes of the medicinal material nodes, the system discovers that there is a phenomenon of functional convergence of medicinal materials within a specific regional scope. For example, "medicinal materials produced in Sichuan often co-occur in prescriptions for moistening the lungs". Statistically analyze the co-occurrence frequency in the medicinal material-region-function triple, and construct a co-occurrence matrix : is the medicinal material and the region co-occurrence count value, is the medicinal material and the region co-occur in the functional semantic path times, is the th medicinal material node, , where is the co-occurrence probability of the two, is the marginal probability of the medicinal material, is the marginal probability of the region, is the medicinal material, is the region.

[0113] Map nodes that are semantically similar or co-occur frequently as "twins", and expand hidden medicinal materials or functional entities that are not explicitly connected. Define that the twin nodes meet the conditions of belonging to the same functional label; the regional co-occurrence intensity > 0.6; the similarity of the medicinal property structure (taste and nature, meridian tropism) > 0.75. There is no connection of "medicinal material A function F" in the original graph; if there exists a twin node B and there is a clear connection of "B F"; then construct a "soft connection": . There is no "Flos farfara "Moistening the lungs", but "Tussilago farfara ~ Fritillaria cirrhosa" (i.e., there is a semantic similarity relationship) and "Fritillaria cirrhosa" "Moistening the lungs", then it is completed. The twin relationship can be represented by "virtual edges" in the graph structure, with edge weights < 0.5, but participating in path reasoning.

[0114] S23. Perform structure clustering on the semantically incomplete data to obtain hierarchical semantic graph data.

[0115] In one embodiment, by calculating the Laplacian matrix of the graph and performing eigenvalue decomposition, the spectral embedding representation of the nodes is obtained, and clustering methods (such as K-Means) are used in the spectral space to group the nodes to obtain hierarchical semantic graph data.

[0116] In one embodiment, the system combines the multi-dimensional attribute similarity (such as Euclidean distance, cosine similarity) between node vectors with the structural connectivity relationship between nodes, calculates the similarity between the vector representations of any two nodes; based on the preliminary clustering results, perform connectivity constraint optimization. For nodes that are assigned to the same class but lack sufficient structural edges (such as node pairs with less than the specified strong connection threshold), perform reclassification or isolation operations; for node pairs with high structural weight edges, give priority to maintaining aggregation to ensure that the clustering results take into account both attribute similarity and topological connectivity. The clustering results are mapped to the hierarchical structure of the semantic graph, forming the following layers: The first layer: functional theme nodes, representing abstract categories of the uses of medicinal materials, such as "moistening the lungs", "clearing heat", "expelling wind", etc.; The second layer: regional attribution nodes, indicating the geographical distribution attribution of specific medicinal materials in the semantic space, such as "produced in Sichuan", "produced in Zhejiang", "Lingnan", etc.; The third layer: entity medicinal material nodes, specifically corresponding to Chinese medicinal material instances, such as "Fritillaria cirrhosa", "Scutellaria baicalensis", "Lonicera japonica", etc.

[0117] Optionally, the functional attribution semantic perception mapping is specifically:

[0118] Filter the functional semantic trigger edges according to the hierarchical semantic edge data to obtain the functional semantic edge data;

[0119] In one embodiment, the system identifies several types of functional semantic connections from the hierarchical semantic edge data, specifically including connection relationships representing treatment goals (such as "having a certain efficacy"), connection relationships representing the meridian attribution direction (such as "acting on a certain meridian"), connection relationships representing the usage in typical prescriptions (such as "applied in a certain type of prescription"), and connection relationships for relieving specific symptoms, etc. The system presets a set of functional semantic edge types, including the following semantic types: hasEffectOn representing efficacy; targetsMeridian representing meridian attribution; usedInPrescriptionFor representing the prescription application relationship; relievesSymptom representing the symptom relief effect. The system traverses the entire set of candidate semantic edges and filters them according to whether the semantic type of the edge belongs to the above set of functional semantic types. If a semantic edge meets the conditions, it is included in the set of functional semantic edges as the starting or intermediate link node for triggering the functional semantic path.

[0120] Construct a functional semantic path for the preset authentic medicinal materials knowledge graph based on the functional semantic edge data to obtain functional semantic path sub-graph data;

[0121] In one embodiment, starting from a functional semantic node, a functional association path graph is generated in the preset graph for analyzing the medicinal material attribution link. In the authentic medicinal materials knowledge graph, a functional node is selected as the starting point; the path hop count is restricted to ≤3, and the path relationship is restricted to functional-related semantic edges; breadth-first search (BFS) or graph traversal is used. Each hop edge type needs to be a functional semantic trigger edge or its subtype; skip paths (such as directly jumping from function to property without an intermediate node) are excluded. During the traversal process, it is restricted that each hop path passes through an edge belonging to the "functional semantic trigger edge" set or its predefined subclass relationship edge. The semantic types of such edges include, but are not limited to, efficacy inheritance edges, function pointing edges, function attribution edges, etc. For example, starting from the functional node "clearing heat", it can point to "moistening the lungs" through an edge of the "has_function" type, and then point to the lung meridian through an "affects_meridian" edge, and then connect to a specific medicinal material node.

[0122] Perform semantic path context verification based on the hierarchical semantic edge data and the functional semantic path sub-graph data to obtain functional context verification data;

[0123] In one embodiment, the query context refers to a set of key semantic elements extracted from a user's retrieval request, which usually includes information in multiple dimensions, such as functional keywords (e.g., "moistening the lungs"); meridian attribution (e.g., "lung meridian"); applicable solar terms (e.g., "autumn"); medicinal property characteristics (e.g., "sweet and warm"), etc. This context set constitutes the benchmark standard for semantic path determination. The functional semantic path is represented by a node sequence. For each node in the path, the system determines whether it matches a certain element in the current query context according to its attribute value. If there is a match, it is counted as a context match item of the path. Define the path context match degree: , if ContextMatch(P) < 0.5, it is marked as a path with inconsistent context.

[0124] Perform semantic coverage conflict detection on the functional context verification data to obtain functional semantic coverage conflict data;

[0125] In one embodiment, when a medicinal material belongs to two functionally conflicting categories (such as "warming yang" and "clearing heat"); or multiple paths point to different meridian attributions (such as "heart meridian" and "lung meridian") and there is no homologous logic support. The system calls the functional mutual exclusion rule table and the meridian independence judgment table set in the traditional Chinese medicine semantic ontology to perform pairwise verification on all path pairs. If a certain medicinal material involves multiple functional semantic paths, the system will construct a conflict judgment matrix to identify whether there is a semantic conflict between any two paths: . The conflict path set is marked for the next accuracy comparison process.

[0126] Compare the mapping accuracy according to the functional semantic coverage conflict data to obtain functional attribution mapping data.

[0127] In one embodiment, the system detects the structural integrity of each functional path. The path should meet the following structural requirements: the starting point is a functional node, the ending point is a medicinal material node, and the path should include a meridian attribution node and a medicinal property node in the middle. The system introduces a context match score to measure the matching degree between the candidate path and the user's current query semantics. The calculation of this score combines the TF-IDF weight of the keywords in the path and the cosine similarity between the path semantic representation and the query vector, that is, weighted calculation, with the former weight being 0.6 and the latter being 0.4. To characterize the semantic credibility of the path, the system also assigns an inference confidence score to each path. If the path comes from an explicit relationship in the knowledge graph ontology, its confidence is relatively high; if the path is a soft connection path inferred and complemented by means such as twin mapping, a lower confidence is given.

[0128] Optionally, the regional co-occurrence semantic perception mapping is specifically as follows:

[0129] Extract region-related trigger nodes according to the hierarchical semantic edge data to obtain region node data;

[0130] In one embodiment, a type of "geographical semantic edge type" can be predefined to identify the edge relationships between entities related to geographical attribution, geographical use, geographical origin, etc. Typical geographical edge types include, but are not limited to, fromRegion: indicating that a certain medicinal material or entity originates from a certain geographical region; originProvince: indicating the provincial annotation of an entity; recordedInArea: indicating that a certain medicinal material or entity is recorded in a specific local chronicle; usedInLocalPrescription: indicating that a certain medicinal material appears or is applied in a local prescription. During the actual extraction process, the edge set in the hierarchical semantic graph is traversed to filter out all edges with geographical pointing significance. All endpoint nodes connected to geographical entities are extracted from the obtained geographical edge set, that is, all geographical nodes. The geographical node is the starting or ending end of the edge, which is specifically determined according to the semantic directionality of the edge.

[0131] Map the medicinal material geographical co-occurrence paths to the preset authentic medicinal materials knowledge graph according to the geographical node data to obtain the medicinal material geographical co-occurrence data;

[0132] In one embodiment, starting from the geographical node, the system searches for the prescription nodes associated with this region, and then further searches for the medicinal material nodes included in this prescription, that is, the following three-hop path is formed: , or the system starts from the medicinal material node and searches for the directly connected geographical node, forming the following two-hop path, . The system sets a maximum path length limit, not exceeding three hops. The system can execute path query tasks based on a graph database engine (such as Neo4j), and starting from the geographical node through the graph traversal strategy, map all reachable medicinal material nodes and their path structures according to the above mode.

[0133] Expand the semantic paths of the medicinal material geographical co-occurrence data to obtain the region co-occurrence candidate path data;

[0134] In one embodiment, the system starts from the medicinal material node and performs a reverse path tracing operation in the semantic graph structure to gradually identify all paths that are semantically connected to the geographical entity. In addition to the direct "medicinal material - production area" path, the system also identifies a type of indirect geographical semantic path. For example, if a certain medicinal material participates in a certain prescription and the symptoms targeted by this prescription have epidemic characteristics in a specific region, then a path can be constructed: "medicinal material" → "prescription" → "symptom" → "region", and this type of path is classified as a soft co-occurrence path. The system performs semantic label annotation on all co-occurrence paths, specifically including the direct production area path, which represents that there is a clear production area pointing relationship between the medicinal material and the region; the formula co-occurrence path, which represents that the prescriptions participated in by the medicinal material are widely used in a specific region; and the symptom transmission path, which represents that the symptoms treated by the medicinal material are epidemic in a specific region.

[0135] Calculate the path co-occurrence weights for the regional co-occurrence candidate path data to obtain path co-occurrence weight data;

[0136] In one embodiment, the credibility or importance of each candidate path in the regional co-occurrence semantics is calculated. Path co-occurrence weight calculation , For path co-occurrence weight labeling data, label the region co-occurrence candidate path data to form path co-occurrence weight data, is the number of occurrences of the path Chinese herbal medicine-region in the corpus / atlas, The weight of the number of occurrences of the Chinese herbal medicine-region in the corpus / atlas is set to 0.35. It is the path type, and the value is set according to the path semantic structure. For example, the direct origin path is greater than the indirect recipe path, which is determined by the preset path structure type value mapping table. is the weight of the path type, set to 0.35, is the path length penalty term. The larger the path length, the more the overall weight should decay (e.g., calculated as ,in is the natural index, For path hop length), is the path length penalty weight, set to 0.3.

[0137] Regional co-occurrence mapping data is generated based on the path co-occurrence weight data and the regional co-occurrence candidate path data to obtain the regional co-occurrence mapping data.

[0138] In one embodiment, for each medicinal material node, all co-occurrence paths and their corresponding weights are summarized. A weight threshold (e.g., 0.7) is set to filter trusted paths. For each medicinal material node, the system summarizes all regional nodes connected by paths that meet the weight threshold, forming a set of candidate regions and including the path sources. If a medicinal material co-occurs with multiple regions, the regions are sorted by score, retaining the top 5-10 candidate locations. The primary origin and secondary co-occurrence locations are annotated.

[0139] Optionally, the node twin extension is specifically:

[0140] Extract structurally similar nodes based on functional attribution mapping data and regional co-occurrence mapping data to obtain structurally similar node data;

[0141] In one embodiment, the system sets the trigger conditions for structural similarity according to the following rules, identifies the medicinal material nodes that show consistency or similarity in multiple semantic dimensions, and determines the medicinal property similarity, including that the medicinal materials have the same or similar nature and flavor attributes, such as "pungent and warm" and "pungent and hot"; determines the compatibility co-occurrence, including that the medicinal materials co-occur with the same high-frequency compatible medicinal material (such as "licorice") in multiple prescriptions; determines the meridian tropism similarity, including that the medicinal materials belong to the same meridian or have similar meridian tropism functions (such as "lung meridian" and "simultaneously treating the spleen and lung"); determines the co-occurrence of region and prescription, including that the producing areas of the medicinal materials are the same, or there are co-occurrence records in historical classics or modern compound prescriptions. To efficiently perform the above similarity extraction, the system uses the subgraph pattern matching algorithm to compare the adjacency structures of the medicinal material nodes in the knowledge graph, extracts the node groups with high structural similarity in the functional path and regional path structures, and labels them as structurally similar node groups.

[0142] Calculate the graph node similarity based on the structurally similar node data to obtain the graph node similarity data;

[0143] In one embodiment, the calculation of the graph node similarity includes the graph adjacency similarity , , where is the neighbor set of node u (considering the edge direction + semantic type); the attribute embedding similarity , encodes the node attributes (such as medicinal properties, meridian tropisms, producing areas) into multi-dimensional feature vectors and calculates using the cosine similarity; the functional path structure similarity , and the judgment basis includes the matching degree of the path topological structure and the similarity of the semantic edge weights in the path. The specific calculation can combine the path edit distance and the semantic edge weight difference index to calculate the path structure similarity between two nodes from the functional level to the entity level. Based on the foregoing, calculate the graph node similarity data : , where the empirical parameters are set as is the weight of the graph adjacency similarity, with a value of 0.3, is the weight of the attribute embedding similarity, with a value of 0.3, is the weight of the functional path structure similarity, with a value of 0.4.

[0144] Perform twin node path inversion based on the graph node similarity data to obtain the twin node path data;

[0145] In one embodiment, the system retrieves a node B that has a high degree of similarity with node A in the structure attributes and semantic features in an existing functional path P. If the twin score (graph node similarity data) of A and B exceeds a set threshold, then the latter B is regarded as a potential twin node of the former A. Using the original path P as a template, the system attempts to construct a path P' for the twin node B that is consistent with P in both structure and semantics, setting constraint conditions such as keeping the semantic edges consistent; the replaceable nodes must appear in the twin mapping table; there are no semantic conflicts in the replacement path (such as inconsistent meridians, contradictory medicinal properties, time festival logic conflicts, etc.).

[0146] Calculate the path confidence based on the twin node path data to obtain the path confidence data;

[0147] In one embodiment, score each inversion path to indicate the credibility of the path. The path confidence score consists of the following four core dimensions: the structural consistency score , indicating whether it is isomorphic to the original path structure; the edge weight stability score , indicating whether the semantic weight of the edge after migration is not lower than that of the original path; the entity semantic compatibility score , indicating the attribute consistency of the replacement nodes in the path; the path context co-occurrence degree , indicating the co-occurrence frequency of the path in actual prescriptions or clinical corpora. , where is the weight of the structural consistency score, with a value of 0.3, is the weight of the edge weight stability score, with a value weight of 0.2, is the weight of the entity semantic compatibility score, with a value of 0.3, is the weight of the path context co-occurrence degree score, with a value of 0.2.

[0148] Generate semantic completion data for the twin node path data, functional attribution mapping data, and regional co-occurrence mapping data according to the path confidence data, semantic completion data.

[0149] In one embodiment, the system calculates the path confidence for all inversion paths generated by the inference mechanism. If the confidence score of a certain path is greater than or equal to a preset threshold (such as 0.7), then include this path in the semantic completion result, and label it according to its connection direction and semantic type. The completion methods include: if the path represents a potential functional attribution relationship, then add a "soft functional attribution edge" to the graph; if the path shows a potential regional co-occurrence relationship, then add it as a "candidate regional edge". To identify that this type of completed path is not a direct connection between original entities, the system attaches an inference source marker field to the edge structure, for example, set isInferred=true, to distinguish the original edge from the completed edge during graph management or path inference.

[0150] Optionally, S3 includes:

[0151] S31. Extract the node semantic nested context of the hierarchical semantic graph data to obtain edge context semantic data;

[0152] In one embodiment, in the hierarchical semantic graph, extract the context semantic information of each edge where the node is located to provide a semantic background for the construction of edge features. The extraction dimensions include the graph structure position, whether the node is in the function layer, the geographical layer, or the medicinal material layer; the adjacency relationship, the type of semantic edge directly connected to the current node and the attributes of adjacent nodes; the node label, such as function attribution, geographical co-occurrence, medicinal property meridian entry, whether it is a twin complement, etc.; the edge context. For example, if the edge is "moistening the lungs → Fritillaria cirrhosa", the edge context includes "lung meridian", "sweet and warm", "collected in autumn".

[0153] S32. Generate multi-source semantic feature tensors for the hierarchical semantic graph data according to the edge context semantic data to obtain edge feature data to be labeled;

[0154] In one embodiment, the so-called edge context semantic data refers to the set of information such as the node semantics connected by the edge itself, the position of the edge in the path, the semantic level where the edge is located, and the semantic labels of other adjacent edges. The system expresses the above context semantic information in multiple dimensions and constructs a three-dimensional semantic feature tensor , where is the number of edges; is the number of feature channels (function features, geographical features, structural features, etc.); is the vector dimension of each channel (for example, using BERT embedding as 768 dimensions). The feature channels include function vectors (function name semantic embedding); medicinal material vectors (the fusion representation of node medicinal properties + meridian entry attributes); regional co-occurrence weight vectors; path attribution confidence vectors; adjacent edge label embeddings (upstream and downstream semantic influences).

[0155] S33. Calculate the path probability according to the edge feature data to be labeled to obtain path probability data;

[0156] In one embodiment, the system evaluates each candidate path based on path probability modeling technology, with the goal of inferring whether it has semantic validity. The estimation of path probability can be achieved through the following two types of methods. For example, by calculating using a Bayesian graph model, constructing a conditional probability graph based on a graph structure, modeling and reasoning about the joint distribution of each semantic edge in the path, so as to calculate the posterior probability that the path is a valid functional path; or, by calculating using a neural path scoring model, encoding the semantic features of all edges in the path (such as edge type, directionality, co-occurrence frequency, semantic weight, etc.) into a tensor form as the input of the neural network model. The system performs a linear combination of the preset weight parameters and the feature tensor, adds a bias term, and uses the Sigmoid activation function to normalize the output result to obtain the probability score that the path is a valid functional path.

[0157] S34. Perform functional analogy annotation on the edge feature data to be annotated according to the path probability data to obtain preliminary semantic icon annotation data;

[0158] In one embodiment, according to the path probability result, perform automatic preliminary functional semantic annotation on the edges. If ( is the semantic credibility of the candidate path, is the determination threshold of path credibility, which can be set by the system according to the application scenario (such as set to 0.7)), then label the edge as the corresponding semantic relationship; for example, if the path is "moistening the lungs → Fritillaria cirrhosa", and P = 0.93, then label the edge type as hasEffectOn, indicating that the function "moistening the lungs" acts on the medicinal material "Fritillaria cirrhosa"; if the path is obtained by twin reasoning, then append the flag inferred=true;

[0159] S35. Perform edge semantic conflict detection on the preliminary semantic icon annotation data to obtain semantic conflict detection data;

[0160] In one embodiment, an edge label coexistence matrix C(i, j) is constructed by statistically analyzing the edge label distribution of all entity nodes in a semantic graph. This matrix represents the co-occurrence of label i and label j within the same node or path context. If logically mutually exclusive label pairs appear simultaneously on an entity node or path structure, a semantic conflict is considered to exist. Within the annotated path structure, the semantic direction and attribution logic between functional paths are checked for conflicts. If an entity node participates in multiple paths, and the functional attribution, meridian system, and medicinal property hierarchy of each path have inconsistent semantic attribution relationships, the path logic is considered to be inconsistent. A rule engine is constructed using preset parameters (preset parameters are set based on an expert knowledge base) to define constraint rules for edge label combinations and path structure logic. The rule engine determines the pair based on preset semantic conflict sets, meridian logic rules, and medicinal property compatibility matrices. For example, the rule system defines that if any entity node has both the "clearing heat" and "warming and nourishing" labels, it is considered a conflicting pair. The system can automatically traverse all edge annotations to determine whether the conflict triggering conditions are met, and record the conflicting edge pairs, conflict types and related node information, and output them as structured semantic conflict detection data.

[0161] S36. Perform relationship completion on the preliminary semantic diagram annotation data according to the semantic conflict detection data to obtain the semantic diagram annotation data.

[0162] In one embodiment, for semantically conflicting edges, the system selects the semantic label corresponding to the path with higher confidence as the main semantic edge of the node pair based on the path confidence data or semantic context consistency measure; the remaining labels can be marked as "secondary candidates" or eliminated. For semantically missing edges, if the system has identified reliable twin paths (i.e., paths with high structural similarity) in the graph or has stable co-occurrence support (such as the same type of medicinal materials showing consistent semantic paths in different regions), then a completion edge is added to the current graph and treated as an inferred edge, and the edge is marked with inferred=true to indicate that it is generated by reasoning rather than directly labeled. After the completion and correction operations are completed, the system uses a graph neural network model to perform unified type prediction on the updated edge set. The edge type prediction model is constructed by combining a graph convolutional neural network and a multi-label classifier, and the prediction function is defined as follows: , is the predicted semantic label of the edge, is a multi-layer perceptron classifier, is the graph embedding representation of edge node u, is the graph embedding representation of edge node v, is the contextual semantic feature vector of the edge.

[0163] Optionally, S4 includes:

[0164] S41. Extract retrieval request feature data according to the authentic medicinal materials retrieval request data, where the retrieval request feature extraction includes regional preference feature extraction, meridian tropism feature extraction, and medicinal property restriction feature extraction;

[0165] In one embodiment, the system combines natural language processing technology, keyword recognition rules, and the traditional Chinese medicine knowledge ontology library to perform structured parsing on the user input statement. For example, in regional preference feature extraction, regional preference information is extracted by matching geographical related keywords (such as "produced in Sichuan", "Lingnan", etc.) in the user input. If the user does not explicitly specify a region, the system will supplement and infer the default regional preference in combination with the common regional usage preferences or medicinal material production habits in the field of traditional Chinese medicine. For meridian tropism feature extraction, based on the syntactic analysis results and the traditional Chinese medicine meridian tropism ontology library, the meridian attribution related to the user input semantics is extracted. For example, for a statement involving specific symptom or function descriptions, the system will annotate according to the associated meridian direction inferred from the semantics (such as "lung meridian", "heart meridian") to form meridian tropism features. For medicinal property restriction feature extraction, based on the preset mapping rules of traditional Chinese medicine terms, the medicinal property description information involved in the user input is identified, including the "four natures" category (cold, heat, warm, cool) and the "five flavors" category (pungent, sweet, bitter, sour, salty).

[0166] S42. Calculate the intention attention of the retrieval request feature data to obtain multi-condition intention data;

[0167] In one embodiment, weight modeling is performed on multiple intention dimensions to clarify the attribute dimension that the user is most concerned about. The user request contains semantic intention dimensions, and an attention mechanism is established for the attention weight of each dimension: where , is the attention score of the th dimension, indicating the importance of this dimension, is the exponential function, is the dimension order, is the weight parameter of the th dimension, and the parameter can be a trainable item or a system setting value, is the th dimension feature embedding vector, generated by combining the context text content and the user portrait information.

[0168] S43. Construct an intention path candidate space for the semantic annotation data according to the multi-condition intention data to obtain intention path candidate space data;

[0169] In one embodiment, based on the above weighted intention information, a path space that meets the user's multi-dimensional conditions is constructed. In the labeled semantic graph, multi-dimensional filters are set, such as must include meridian nodes → meet meridian_pref (meridian constraint); the medicinal property nodes need to meet property_constraints (medicinal property restrictions); the geographical attribution path should at least include any region in region_pref (geographical preference). Use graph filtering to construct an intention constraint subgraph; the combination logic can include "AND", "OR", "EXCLUDE" semantic settings.

[0170] S44. Perform graph traversal based on the candidate space data of the intention path to obtain path candidate graph data;

[0171] In one embodiment, perform structural traversal within the intention path space to obtain a collection of candidate paths. The graph traversal uses restricted depth-first search (DFS) or breadth-first search (BFS); the restriction condition is that the path length ≤ 4; it should at least include one functional starting point and one medicinal material entity; all edge types ∈ the labeled semantic set (such as hasEffectOn, usedWith, fromRegion);

[0172] S45. Calculate the semantic coverage rate based on the path candidate graph data and the retrieval request feature data to obtain semantic coverage rate data;

[0173] In one embodiment, let the total number of intention dimensions included in the user's retrieval request be , and for a certain path , the number of intention dimensions covered is , then the semantic coverage rate of this path is denoted as , or, further, introduce a weighted semantic coverage rate function based on the attention mechanism, , is the weighted semantic coverage rate, which is a representation of the semantic coverage rate data, is the intention dimension index, is the total number of intention dimensions, If the path meets the -dimensional intention.

[0174] S46. Screen the path candidate graph data according to the semantic coverage rate data to obtain path screening data;

[0175] In one embodiment, select the path with the highest semantic coverage rate and the strongest confidence from all candidate paths. First, select the path set with a coverage rate ≥ 0.8; if there are less than 5 paths, appropriately relax it to ≥ 0.6; the sorting logic is to perform weighted combination according to the coverage rate and the path confidence to perform sorting: , where is the path comprehensive score of is the path semantic coverage of is the path confidence of

[0176] S47. Perform edge node semantic filling processing on the path screening data to obtain the authentic medicinal material retrieval data.

[0177] In one embodiment, the system identifies the edge nodes in the path structure. If the node lacks geographical attributes, medicinal property descriptions, or meridian tropism information, it is regarded as a semantic breakpoint. The system will use the entity twin reasoning mechanism in the path context information and the knowledge graph to intelligently fill the missing attributes of the node. For example, when a certain path represents "moistening the lungs → Aster tataricus", and "Aster tataricus" is not labeled with geographical attributes, the system will retrieve other medicinal material entities with the same efficacy in the knowledge graph through its function label "moistening the lungs", and analyze their main production area information, so as to infer the geographical attribution of "Aster tataricus". The inference result is attached to the target node in the form of a soft label, including but not limited to inference attributes such as "presumed origin = Hubei". The system can add this complemented information as a node attribute field, or add a semantic edge in the graph to represent the inference association, and set the "inference source" identifier to indicate that this complemented information comes from the semantic reasoning process rather than explicit definition.

[0178] Optionally, the edge node semantic filling processing is specifically as follows:

[0179] Perform boundary node analysis on the path screening data to obtain half-edge connection node data;

[0180] In one embodiment, identify the nodes that are only connected by one side, lack attribute annotations or context relationships in the path screening results, that is, semantic edge nodes. The system determines the edge of the node through the following rules: the node degree deg() = 1 or the only connected edge does not contain semantic annotations; the node lacks core attribute fields (such as medicinal properties, regions, meridian tropisms); the node does not participate in a closed triple or lacks upstream and downstream path information.

[0181] Generate edge node missing based on the half-edge connection node data to obtain edge node missing data;

[0182] In one embodiment, identify which core semantic paths or attributes are missing from the edge nodes, and establish a set to be complemented. For each edge node, list the attributes that should be in the graph model but are missing in the current path, such as geographical attribution (fromRegion), medicinal property (hasProperty), meridian tropism (targetsMeridian), efficacy path (hasEffectOn).

[0183] Perform semantic path missing type mapping based on the missing data of edge nodes to obtain semantic path missing type data;

[0184] In one embodiment, map the above missing information to the completable types. The system divides all missing situations into the following three types and makes explicit annotations. The missing type definitions are as follows: TypeA: Missing attribute class nodes (such as lack of drug property / region); TypeB: Missing edge class relationships (such as lack of functional connections); TypeC: Missing upstream and downstream structure jump points (such as path faults).

[0185] Perform attribute neighbor matching based on the semantic path missing type data to obtain supplementary node data;

[0186] In one embodiment, find the node with the most similar structural attributes to the edge node from the knowledge graph for migrating its missing attributes. Construct feature vectors (drug property, meridian tropism, function, compatibility frequency, etc.); use cosine similarity or graph embedding distance to measure candidate nodes: , is the semantic similarity score of the node, representing the node to be completed and the candidate node The overall semantic similarity of, which is equivalent to Or the result of the distance function based on the graph embedding space, is the cosine similarity scoring function, used to calculate the cosine angle similarity between the semantic vectors of two nodes. The numerical range is [0,1], and the larger the value, the higher the similarity. is the feature vector (node to be completed), is the feature vector (candidate entity node), is the similarity threshold. For each missing item, extract the most similar (the top 5 - 10 after sorting from large to small) entities from the completion candidate set; extract their complete attributes and construct transferable supplementary items (such as "Aster tataricus → Region: Hubei").

[0187] Perform similarity screening on the supplementary node data according to the retrieval request feature data to obtain supplementary screening data;

[0188] In one embodiment, the system constructs a matching scoring model for all candidate supplementary nodes according to the semantic preference information input by the user, including region preference, meridian tropism direction, and drug property characteristics, etc., to evaluate the semantic fitting degree between them and the user's retrieval target: , is the matching score of the candidate node , is the region weight coefficient, with a value of 0.2, is the region matching degree index, for example, RegionMatch = 1 (fully matched) or 0.5 (same provincial region), is the meridian tropism weight coefficient, with a value of 0.5, is the meridian tropism matching degree index, indicating the consistency between the meridian tropism attribute of the candidate node and the meridians preferred by the user. If they match, it takes 1; if they are partially relevant, it takes 0.5, etc. is the medicinal property weight coefficient, with a value of 0.3, is the medicinal property matching degree index, indicating the degree of conformity between the medicinal properties (four natures and five flavors) of the candidate node and the user's preferences. It can be calculated through label comparison or vector cosine. The system will sort according to the scoring results and preferentially retain the node with the highest score as the semantic supplement item. If there are multiple candidate nodes with similar scores, the system can also perform the node semantic fusion operation to integrate the information of multiple nodes into a virtual node.

[0189] Integrate the supplementary screening data and the path screening data to obtain the genuine medicinal materials retrieval data.

[0190] In one embodiment, the semantic path information generated by the complementation mechanism is embedded into the original path structure, while the expression format of the attribute fields is unified, and the source information of the inference or complementation content is recorded, so that the retrieval output result has both integrity and traceability. The integration methods include but are not limited to the following aspects. For example, path information field merging, where the path sequence obtained in the original path screening stage is fused with the path information generated in the complementation inference stage, and the original path structure order and semantics are maintained. Attribute field completion, for each medicinal material entity node, supplement its missing semantic attribute fields, including geographical attribution (such as "source area: Sichuan"), medicinal property labels (such as "nature and flavor: sweet and warm"), meridian tropism information (such as "pertaining to the lung meridian"), etc., to ensure that all core semantic labels are clearly marked in the final result. Complemented path identification, for the path or attribute information generated by the semantic inference or structure complementation mechanism, explicit identification is given. Preferably, an "inference flag" field can be set to indicate whether this part of the information is generated by inference or twin nodes. For example, set the field inferred = true or the complementation source type as "structural similarity inference". Inference source record, for each complemented piece of information, record its source path or inference basis. For example, if the meridian tropism information of a certain medicinal material comes from the co-occurrence path of a medicinal material with a similar structure, it is noted in the result as "information source: inferred from the medicinal materials with the same path as Fritillaria cirrhosa". The output genuine medicinal materials retrieval data can be organized in a structured data form, such as a structured key-value pair list, a graph structure JSON expression, or a visual recommendation list, etc. The content includes fields such as medicinal material name, path structure, geographical information, medicinal property label, meridian tropism attribute, inference source, path score, etc., constituting an intelligent retrieval response result with multi-dimensional fusion.

[0191] Optionally, the present application also provides a geographic authentic medicinal materials intelligent retrieval system based on a knowledge graph for performing the geographic authentic medicinal materials intelligent retrieval method as described above. The geographic authentic medicinal materials intelligent retrieval system based on a knowledge graph includes:

[0192] A semantic granularity analysis module, configured to obtain geographic authentic medicinal materials retrieval request data, perform semantic granularity processing on the geographic authentic medicinal materials retrieval request data to obtain semantic granularity data, and generate hierarchical semantic edges based on the semantic granularity data to obtain hierarchical semantic edge data;

[0193] A hierarchical semantic graph construction module, configured to generate a hierarchical semantic graph based on the hierarchical semantic edge data for a preset geographic authentic medicinal materials knowledge graph to obtain hierarchical semantic graph data;

[0194] A semantic edge annotation module, configured to perform relational edge semantic annotation on the hierarchical semantic graph data to obtain semantic graph annotation data;

[0195] An intelligent semantic-driven retrieval module, configured to perform multi-condition intention processing based on the geographic authentic medicinal materials retrieval request data to obtain multi-condition intention data, perform multi-hop path combination on the semantic graph annotation data according to the multi-condition intention data to obtain preliminary retrieval data, and perform edge reasoning and completion on the preliminary retrieval data to obtain geographic authentic medicinal materials retrieval data.

[0196] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended application documents rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be encompassed within the present invention.

[0197] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for intelligent retrieval of authentic medicinal materials based on graphs, characterized in that: The method includes: S1. Obtain the retrieval request data of genuine regional drugs, perform semantic granularity processing on the retrieval request data of genuine regional drugs to obtain semantic granularity data; generate hierarchical semantic edges according to the semantic granularity data to obtain hierarchical semantic edge data; S2. Generate a hierarchical semantic graph for the preset genuine regional drug knowledge graph according to the hierarchical semantic edge data to obtain hierarchical semantic graph data; S3. Perform semantic annotation of relationship edges on the hierarchical semantic graph data to obtain semantically annotated data; S4. Perform multi-condition intention processing according to the retrieval request data of genuine regional drugs to obtain multi-condition intention data; perform multi-hop path combination on the semantically annotated data according to the multi-condition intention data to obtain preliminary retrieval data; perform edge reasoning and complementation on the preliminary retrieval data to obtain the retrieval data of genuine regional drugs.

2. The method according to claim 1, characterized in that, S1 includes: Obtain the retrieval request data of genuine regional drugs; Perform coarse semantic granularity processing on the retrieval request data of genuine regional drugs to obtain coarse semantic granularity data; Perform fine semantic granularity processing on the retrieval request data of genuine regional drugs to obtain fine semantic granularity data, where the semantic granularity processing is carried out according to the preset traditional Chinese medicine semantic ontology and granularity stratification rules, the coarse semantic granularity is the drug category, regional attribute and basic efficacy, and the fine semantic granularity is the meridian tropism, drug nature, collection season and classical compatibility; Construct semantic paths and calculate multi-level edge weights for the coarse semantic granularity data and the fine semantic granularity data to obtain semantic path data and multi-level edge weight data; Generate a semantic hierarchical graph according to the semantic path data and the multi-level edge weight data to obtain multi-granularity semantic hierarchical graph data; Perform semantic edge classification annotation on the multi-granularity semantic hierarchical graph data to obtain hierarchical semantic edge data.

3. The method according to claim 1, wherein S2 Includes: Perform functional attribution semantic perception mapping and regional co-occurrence semantic perception mapping on the preset genuine regional drug knowledge graph according to the hierarchical semantic edge data to obtain functional attribution mapping data and regional co-occurrence mapping data respectively; Perform node twin expansion according to the functional attribution mapping data and the regional co-occurrence mapping data to obtain semantic complementation data; 4. The method according to claim 3, wherein Perform structure clustering according to the semantic incomplete data to obtain hierarchical semantic graph data. Among them, the functional attribution semantic perception mapping is specifically: Screen functional semantic trigger edges according to the hierarchical semantic edge data to obtain functional semantic edge data; Construct a functional semantic path for the preset genuine regional drug knowledge graph according to the functional semantic edge data to obtain functional semantic path sub-graph data; Perform semantic path context verification according to the hierarchical semantic edge data and the functional semantic path sub-graph data to obtain functional context verification data; Perform semantic coverage conflict detection on the functional context verification data to obtain functional semantic coverage conflict data; 5. The method according to claim 3, characterized in that, Compare the mapping accuracy according to the functional semantic coverage conflict data to obtain functional attribution mapping data. Among them, the regional co-occurrence semantic perception mapping is specifically: Extract region-related trigger nodes according to the hierarchical semantic edge data to obtain region node data; Perform drug region co-occurrence path mapping on the preset genuine regional drug knowledge graph according to the region node data to obtain drug region co-occurrence data; Expand the semantic path of the drug region co-occurrence data to obtain region co-occurrence candidate path data; Calculate the path co-occurrence weight for the region co-occurrence candidate path data to obtain the path co-occurrence weight data; Generate the geographical co-occurrence mapping data based on the path co-occurrence weight data and the region co-occurrence candidate path data to obtain the geographical co-occurrence mapping data.

6. The method according to claim 3, wherein Among them, the node twin expansion is specifically as follows: Extract the structurally similar node data based on the function attribution mapping data and the geographical co-occurrence mapping data; Calculate the similarity of graph nodes based on the structurally similar node data to obtain the graph node similarity data; Perform the twin node path inversion based on the graph node similarity data to obtain the twin node path data; Calculate the path confidence based on the twin node path data to obtain the path confidence data; Generate the semantic completion data for the twin node path data, the function attribution mapping data, and the geographical co-occurrence mapping data based on the path confidence data.

7. The method according to claim 1, wherein S3 Including: Extract the edge context semantics data by performing node semantic nested context extraction on the hierarchical semantic graph data; Generate the multi-source semantic feature tensors for the hierarchical semantic graph data based on the edge context semantics data to obtain the edge feature data to be labeled; Calculate the path probability based on the edge feature data to be labeled to obtain the path probability data; Perform function analogy annotation on the edge feature data to be labeled based on the path probability data to obtain the preliminary semantic graph annotation data; Detect the edge semantic conflicts in the preliminary semantic graph annotation data to obtain the semantic conflict detection data; Perform relationship completion on the preliminary semantic graph annotation data based on the semantic conflict detection data to obtain the semantic graph annotation data.

8. The method according to claim 1, wherein S4 Including: Extract the retrieval request feature data based on the authentic medicinal materials retrieval request data, where the retrieval request feature extraction includes geographical preference feature extraction, meridian tropism feature extraction, and medicinal property restriction feature extraction; Calculate the intention attention for the retrieval request feature data to obtain the multi-condition intention data; Construct the intention path candidate space for the semantic graph annotation data based on the multi-condition intention data to obtain the intention path candidate space data; Perform graph traversal based on the intention path candidate space data to obtain the path candidate graph data; Calculate the semantic coverage rate based on the path candidate graph data and the retrieval request feature data to obtain the semantic coverage rate data; Filter the path candidate graph data based on the semantic coverage rate data to obtain the path filtering data; Perform edge node semantic filling processing on the path filtering data to obtain the authentic medicinal materials retrieval data.

9. The method according to claim 8, wherein Among them, the edge node semantic filling processing is specifically as follows: Analyze the boundary nodes of the path filtering data to obtain the half-edge connected node data; Generate the edge node missing data based on the half-edge connected node data; Map the semantic path missing type based on the edge node missing data to obtain the semantic path missing type data; Perform attribute neighbor matching based on the semantic path missing type data to obtain the supplementary node data; Perform similarity filtering on the supplementary node data based on the retrieval request feature data to obtain the supplementary filtering data; Integrate the supplementary filtering data and the path filtering data to obtain the authentic medicinal materials retrieval data.

10. An intelligent retrieval system for genuine regional drugs based on a spectrum, characterized in that, For executing the atlas-based intelligent retrieval method of authentic Chinese medicinal materials as described in claim 1, the atlas-based intelligent retrieval system of authentic Chinese medicinal materials includes: A semantic granularity analysis module, configured to obtain authentic Chinese medicinal material retrieval request data, perform semantic granularity processing on the authentic Chinese medicinal material retrieval request data to obtain semantic granularity data; generate hierarchical semantic edges according to the semantic granularity data to obtain hierarchical semantic edge data; A hierarchical semantic graph construction module, configured to generate a hierarchical semantic graph for a preset authentic Chinese medicinal material knowledge graph according to the hierarchical semantic edge data to obtain hierarchical semantic graph data; A semantic edge annotation module, configured to perform relational edge semantic annotation on the hierarchical semantic graph data to obtain semantic graph annotation data; An intelligent semantic-driven retrieval module, configured to perform multi-condition intention processing according to the authentic Chinese medicinal material retrieval request data to obtain multi-condition intention data; perform multi-hop path combination on the semantic graph annotation data according to the multi-condition intention data to obtain preliminary retrieval data; perform edge reasoning and completion on the preliminary retrieval data to obtain authentic Chinese medicinal material retrieval data.

Citation Information

Patent Citations

  • Potential relation reasoning-based medical knowledge graph retrieval system and method

    CN119739867A

  • System and / or method for an autonomous linked managed semantic model based knowledge graph generation framework

    US20220121964A1