Intelligent retrieval method and system for genuine medicinal materials based on atlas
Through the graph-based method, hierarchical semantic graphs are constructed and semantic annotated, the problem of inaccurate user semantics and search results in the prior art is solved, and intelligent search of authentic medicinal materials with high accuracy and flexibility is achieved.
Patent Information
- Application Number
- CN202510691794.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The existing Chinese herbal medicine information query system cannot understand the semantics of user input, the search results have low coverage and insufficient accuracy, making it difficult to achieve intelligent matching and semantic retrieval based on user needs.
A map-based intelligent search method for authentic medicinal materials is proposed. Through semantic particle size processing and hierarchical semantic edge generation, a hierarchical semantic graph is constructed, and relationship edge semantic annotation is performed, supporting multi-conditional intention processing and multi-hop path combination, and edge inference completion is performed.
It realizes multi-dimensional semantic analysis of authentic medicinal materials, improves the reasonability and semantic expression ability of the graph structure, and significantly improves the accuracy, flexibility and recommendation intelligence of retrieval.
Smart Images

Figure CN120216741A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge graph retrieval, and particularly to an intelligent retrieval method and system for authentic Chinese medicinal materials based on a graph. Background Art
[0002] Authentic Chinese medicinal materials refer to Chinese medicinal materials produced under specific natural ecological environments and specific cultivation, harvesting, and processing technical conditions, which have excellent, stable quality and curative effects and play an important role in the clinical application of traditional Chinese medicine and the development of the traditional Chinese medicine industry. With the improvement of the digital and intelligent levels of traditional Chinese medicine, how to extract the core attribute information of authentic Chinese medicinal materials from the vast traditional Chinese medicine literature, standard pharmacopoeias, local chronicles, and modern scientific research results, and achieve intelligent matching and semantic retrieval based on user needs has become a key problem urgently to be solved in the field of traditional Chinese medicine information services. Most of the existing Chinese medicinal material information query systems are based on traditional relational databases or keyword matching models, and mainly use static field retrieval methods for querying, such as screening by medicine name, functions and indications, place of origin, etc., and there are problems such as being unable to understand the semantics of user input, low coverage rate of retrieval results, and insufficient accuracy. Summary of the Invention
[0003] In order to solve the above technical problems, the present invention proposes an intelligent retrieval method and system for authentic Chinese medicinal materials based on a graph to solve at least one of the above technical problems.
[0004] The present application provides an intelligent retrieval method for authentic Chinese medicinal materials based on a graph, and the method includes: S1. Obtain authentic Chinese medicinal material retrieval request data, perform semantic granularity processing on the authentic Chinese medicinal material retrieval request data to obtain semantic granularity data; generate hierarchical semantic edges according to the semantic granularity data to obtain hierarchical semantic edge data; S2. Generate a hierarchical semantic graph for a preset authentic Chinese medicinal material knowledge graph according to the hierarchical semantic edge data to obtain hierarchical semantic graph data; S3. Perform semantic annotation on the relationship edges of the hierarchical semantic graph data to obtain semantically annotated data; S4. Perform multi-condition intention processing according to the authentic Chinese medicinal material retrieval request data to obtain multi-condition intention data; perform multi-hop path combination on the semantically annotated data according to the multi-condition intention data to obtain preliminary retrieval data; perform edge reasoning and completion on the preliminary retrieval data to obtain authentic Chinese medicinal material retrieval data.
[0005] In the present invention, by performing semantic granularity processing on the data of the retrieval request for genuine regional drugs and generating hierarchical semantic edge data, the multi-dimensional semantic parsing of user intentions is effectively realized; further, by constructing a hierarchical semantic graph and performing semantic annotation operations, the inferability of the graph structure and the semantic expression ability are improved. Using multi-condition intention processing and multi-hop path combination, it not only supports the intelligent matching of drugs under various conditions, but also can discover potential drug nodes that are not explicitly represented in the original graph but have semantic rationality through the edge reasoning completion mechanism, thus significantly improving the accuracy, flexibility of retrieval and the intelligence of recommendation, and providing an intelligent knowledge retrieval method with both deep semantic understanding and structural reasoning ability for the field of traditional Chinese medicine.
[0006] Optionally, S1 includes: Obtain the data of the retrieval request for genuine regional drugs; Perform coarse semantic granularity processing on the data of the retrieval request for genuine regional drugs to obtain coarse semantic granularity data; Perform fine semantic granularity processing on the data of the retrieval request for genuine regional drugs to obtain fine semantic granularity data, where the semantic granularity processing is carried out according to the preset semantic ontology of traditional Chinese medicine and the granularity stratification rules. The coarse semantic granularity is the drug category, regional attribute and basic efficacy, and the fine semantic granularity is the meridian tropism, drug nature, collection season and classical compatibility; Perform semantic path construction and multi-level edge weight calculation on the coarse semantic granularity data and the fine semantic granularity data to obtain semantic path data and multi-level edge weight data; Generate a multi-granularity semantic hierarchical graph according to the semantic path data and the multi-level edge weight data to obtain multi-granularity semantic hierarchical graph data; Perform semantic edge classification annotation according to the multi-granularity semantic hierarchical graph data to obtain hierarchical semantic edge data.
[0007] In the present invention, by dividing the data of the retrieval request for genuine regional drugs into two levels of coarse semantic granularity and fine semantic granularity, and processing them in combination with the semantic ontology of traditional Chinese medicine and the granularity stratification rules, the user's query intention can be comprehensively expressed from different semantic depth dimensions. Among them, the coarse-grained processing can quickly locate the drug category, region and basic efficacy, and the fine-grained processing further extracts fine features such as meridian tropism, drug nature, collection season and classical compatibility, improving the hierarchy and accuracy of semantic expression. At the same time, by constructing semantic paths and performing multi-level edge weight calculation, the quantitative modeling of structured semantic paths is realized, and further multi-granularity semantic hierarchical graphs and semantic edge classification annotation data are generated, providing a high-quality structural basis for graph reasoning. Compared with the prior art, this step not only supports the multi-granularity fusion modeling of user intentions, but also enhances the hierarchical expression ability of the graph structure and the context semantic computability, significantly improving the depth of semantic understanding and the reasoning accuracy of the retrieval of genuine regional drugs, and having stronger flexibility and intelligent adaptability.
[0008] Optionally, S2 includes: Performing functional attribution semantic perception mapping and regional co-occurrence semantic perception mapping on the preset genuine medicinal materials knowledge graph according to the hierarchical semantic edge data to obtain functional attribution mapping data and regional co-occurrence mapping data respectively; Performing node twin expansion based on the functional attribution mapping data and the regional co-occurrence mapping data to obtain semantic completion data; Performing structure clustering on the semantic incomplete data to obtain hierarchical semantic graph data.
[0009] In the present invention, based on the hierarchical semantic edge data, functional attribution semantic perception mapping and regional co-occurrence semantic perception mapping are respectively performed, realizing the multi-dimensional semantic structure perception of the "medicinal property - efficacy - meridian tropism" path and the "region - season - medicinal material" path in the genuine medicinal materials atlas, effectively constructing a context semantic view for reasoning. Through the node twin expansion mechanism, combined with graph structure similarity, attribute feature consistency and path co-occurrence rules, potential but not explicitly expressed medicinal material nodes are automatically identified to generate semantic completion data, thereby improving the coverage ability of the atlas in the knowledge sparse area. Subsequently, by performing structure clustering processing on the completion data, aggregating subgraphs are constructed according to semantic similarity, forming a hierarchical semantic graph structure with classification, scalability and reasoning consistency. Compared with the single static path matching method in the prior art, the present invention introduces a dual-channel semantic perception, node twin intelligent expansion and structure hierarchical reconstruction mechanism, significantly enhancing the semantic relevance, structure reasoning ability and dynamic completion ability of the knowledge graph, making the retrieval of genuine medicinal materials more intelligent and generalized.
[0010] Optionally, the functional attribution semantic perception mapping is specifically: Screening functional semantic trigger edges according to the hierarchical semantic edge data to obtain functional semantic edge data; Constructing a functional semantic path on the preset genuine medicinal materials knowledge graph according to the functional semantic edge data to obtain functional semantic path subgraph data; Performing semantic path context verification according to the hierarchical semantic edge data and the functional semantic path subgraph data to obtain functional context verification data; Performing semantic coverage conflict detection on the functional context verification data to obtain functional semantic coverage conflict data; Comparing the mapping accuracy according to the functional semantic coverage conflict data to obtain functional attribution mapping data.
[0011] In the present invention, by constructing a multi-stage and multi-level functional semantic reasoning path processing mechanism, the semantic matching depth and reasoning accuracy between medicinal materials and functional attribution are significantly improved. According to the hierarchical semantic edge data, functional semantic trigger edges are screened, effectively filtering out the key paths related to core functional attributes such as medicinal properties, main indications, and meridians entered. The constructed functional semantic path subgraph data retains the structural context characteristics of functional attribution. By performing context verification and semantic consistency analysis on the paths, path deviations caused by semantic ambiguity or context ambiguity can be avoided. The proposed semantic coverage conflict detection mechanism can identify semantic conflicts in functional attribution among multiple paths. Combining with the mapping accuracy comparison mechanism, the optimal functional attribution path mapping result under the current query intention is finally selected. Compared with the method in the prior art that only performs mapping matching based on keywords or single-path scoring, the present invention adopts structure-triggered and path context verification to realize a fusion mechanism for conflict detection and multi-path accuracy competition, improving the system's understanding ability and semantic discrimination ability for functional attribution intentions, with stronger anti-interference, interpretability, and semantic discrimination abilities, and having significant advantages in graph reasoning accuracy and scenario adaptability.
[0012] Optionally, the regional co-occurrence semantic perception mapping is specifically as follows: Extract regional-related trigger nodes according to the hierarchical semantic edge data to obtain regional node data; Perform medicinal material regional co-occurrence path mapping on the preset genuine medicinal materials knowledge graph according to the regional node data to obtain medicinal material regional co-occurrence data; Expand the semantic path of the medicinal material regional co-occurrence data to obtain regional co-occurrence candidate path data; Calculate the path co-occurrence weight for the regional co-occurrence candidate path data to obtain path co-occurrence weight data; Generate regional co-occurrence mapping data according to the path co-occurrence weight data and the regional co-occurrence candidate path data to obtain regional co-occurrence mapping data.
[0013] In the present invention, by constructing a semantic co-occurrence reasoning mechanism based on the multi-dimensional path of geography-medicinal material-function, a systematic modeling of the distribution relationship, harvesting characteristics and functional attribution of genuine medicinal materials under different regional environments and time conditions is realized. Based on the hierarchical semantic edge data, region-related trigger nodes are extracted, enabling the system to perceive the geographical regions or administrative divisions involved in the user's intention and constructing a region-dominated reasoning entry. By constructing the medicinal material-region co-occurrence path and expanding the semantic path, the medicinal material information related to the region can be comprehensively captured, including the composite paths related to the harvesting season, ecological environment, local uses, etc. Through the calculation of the path co-occurrence weight, combined with factors such as the frequency of historical documents, the co-occurrence relationship of classics, and regional suitability, a semantic intensity score is given to each path. The generated region co-occurrence mapping data not only has a clear spatial positioning ability but also has the characteristics of multi-hop path semantic fusion. Different from the static matching method of single query by "production area field" or regional label in the prior art, the present invention significantly improves the semantic accuracy and retrieval coverage of the regional attribution of medicinal materials by introducing the region trigger, semantic path expansion and co-occurrence weight modeling mechanism, realizing dynamic, interpretable and structure-aware geographical semantic reasoning ability, and having stronger practicability and novelty in the scenario of genuine medicinal materials.
[0014] Optionally, the node twin expansion specifically includes: Extracting structure-similar nodes according to the functional attribution mapping data and the region co-occurrence mapping data to obtain structure-similar node data; Calculating the similarity of graph nodes according to the structure-similar node data to obtain graph node similarity data; Performing twin node path inversion according to the graph node similarity data to obtain twin node path data; Calculating the path confidence according to the twin node path data to obtain path confidence data; Generating semantic completion data for the twin node path data, the functional attribution mapping data and the region co-occurrence mapping data according to the path confidence data, the semantic completion data.
[0015] In the present invention, by introducing a structural similarity analysis and path inversion reasoning mechanism, semantic completion of potential but not explicitly associated medicinal material nodes is achieved on the basis of the existing atlas, significantly improving the knowledge integrity and structural generalization ability of the genuine medicinal material atlas. Structural similar nodes are extracted based on the function attribution mapping data and the geographical co-occurrence mapping data, and the similarity of the atlas nodes is calculated. Starting from multi-dimensional features such as the structural adjacency relationship, attribute similarity, and context co-occurrence frequency, the potential semantic consistency between different nodes is quantified. Through twin node path inversion, taking the highly similar nodes as the starting point, a reverse semantic path similar to the known path structure is constructed to form a set of semantic approximate paths. After confidence scoring of the inverted paths, the rationality of using them as missing paths can be effectively evaluated, and based on this, combined with the original function attribution and geographical co-occurrence results, semantic completion data is generated. Different from the existing atlas technologies that only rely on static entity relationships or node similarities to complete the missing edges, the present invention not only improves the semantic rationality and path consistency of the completion results, but also has stronger context adaptation ability and atlas reasoning transparency, and has stronger practical value in semantic prediction and unknown knowledge discovery.
[0016] Optionally, S3 includes: Extract the node semantic nested context of the hierarchical semantic graph data to obtain edge context semantic data; Generate multi-source semantic feature tensors for the hierarchical semantic graph data according to the edge context semantic data to obtain edge feature data to be annotated; Calculate the path probability according to the edge feature data to be annotated to obtain path probability data; Perform functional analogy annotation on the edge feature data to be annotated according to the path probability data to obtain preliminary semantic graph annotation data; Detect the edge semantic conflict of the preliminary semantic graph annotation data to obtain semantic conflict detection data; Complete the relationship of the preliminary semantic graph annotation data according to the semantic conflict detection data to obtain semantic graph annotation data.
[0017] In the present invention, by constructing a context nested semantic expression and multi-source feature fusion mechanism for the edges of the graph, the depth of edge semantic understanding and the accuracy of edge relationship annotation in the structural reasoning process of the hierarchical semantic graph are improved. Through the extraction of node semantic nested context, the implicit semantic collaboration relationship between nodes in the path is obtained and transformed into edge context semantic data, providing structured semantic support for annotation. A semantic feature tensor is constructed by fusing multi-source information such as node attributes, path patterns, and historical co-occurrence frequencies, obtaining edge feature data to be annotated with high-dimensional expression, enhancing the expression ability of edge features. Subsequently, a path probability calculation module is introduced. Based on the co-occurrence probability of the edge in the functional path and combined with the semantic intention similarity, functional analogy annotation is performed to automatically infer the functional role of the edge, realizing the automatic annotation ability driven by path semantics. To ensure structural consistency, the system further performs edge semantic conflict detection, identifies semantic ambiguity or mutually exclusive relationships, and uniformly optimizes and corrects the edge semantics through a relationship completion strategy to form a complete semantic graph annotation result. Different from the static rules or single-sided attribute annotation methods commonly used in the prior art, the present invention significantly improves the structural interpretability of edge annotation, the consistency of functional reasoning, and the semantic error correction ability, has stronger intelligent and automatic expansion capabilities, and achieves a key breakthrough in the quality of knowledge graph annotation and the reasoning efficiency.
[0018] Optionally, S4 includes: Retrieve request feature extraction is performed on the genuine medicinal materials retrieval request data to obtain retrieve request feature data, where the retrieve request feature extraction includes geographical preference feature extraction, meridian tropism feature extraction, and medicinal property restriction feature extraction; Intent attention calculation is performed on the retrieve request feature data to obtain multi-condition intent data; An intent path candidate space is constructed for the semantic graph annotation data according to the multi-condition intent data to obtain intent path candidate space data; Graph traversal is performed according to the intent path candidate space data to obtain path candidate graph data; Semantic coverage calculation is performed according to the path candidate graph data and the retrieve request feature data to obtain semantic coverage data; The path candidate graph data is filtered according to the semantic coverage data to obtain path filtering data; Edge node semantic filling processing is performed on the path filtering data to obtain genuine medicinal materials retrieval data.
[0019] In the present invention, by constructing an intention recognition-path construction-semantic completion linkage mechanism driven by user retrieval features, a closed-loop reasoning process from natural language retrieval intention to structured knowledge path matching and complete result generation is realized. For the retrieval request data of genuine medicinal materials, multi-dimensional features such as regional preference, meridian tropism and property limitation are extracted, and a multi-angle feature modeling framework for personalized semantic understanding is established. On this basis, a multi-condition intention representation is formed through intention attention calculation, effectively improving the context focusing ability during semantic graph matching. Subsequently, a path candidate space is constructed according to the intention features, and path candidate graph data is generated through a structure traversal method, ensuring that the retrieval is not limited to direct explicit paths but also covers combinable paths. Further, by combining the retrieval intention and the graph structure to calculate the semantic coverage rate, the semantic satisfaction degree of the candidate path for the user's needs can be accurately measured, and the quantitative evaluation of the intention alignment degree is realized. After screening by the semantic coverage rate, paths with high matching degrees are retained, and semantic filling is performed on the edge nodes to automatically complete missing information such as harvesting season and regional attribution, generating a semantic-complete and structure-closed retrieval result of genuine medicinal materials. Different from the prior art which mainly relies on keyword indexing or shallow entity matching methods, the present invention proposes an intention feature-driven structure path construction and completion strategy, integrating the attention mechanism, path coverage calculation and edge semantic completion, significantly improving the intelligent matching accuracy, semantic adaptation ability and retrieval generalization ability of the system.
[0020] Optionally, the semantic filling process of the edge nodes is specifically as follows: Perform boundary node analysis on the path screening data to obtain half-edge connection node data; Generate edge node missing data according to the half-edge connection node data; Perform semantic path missing type mapping according to the edge node missing data to obtain semantic path missing type data; Perform attribute near neighbor matching according to the semantic path missing type data to obtain supplementary node data; Perform similarity screening on the supplementary node data according to the retrieval request feature data to obtain supplementary screening data; Integrate the supplementary screening data and the path screening data to obtain the retrieval data of genuine medicinal materials.
[0021] In the present invention, the system identifies half-connected nodes with structural breaks at the end or beginning of a path through boundary node analysis, then generates edge missing nodes based on the path structure and semantic tags, and performs semantic path missing type mapping to clarify that the missing part belongs to types such as origin missing, meridian tropism missing, or harvesting season missing. By combining the node data existing under similar paths in the atlas for attribute near-neighbor matching, candidate nodes for supplementation are mined from dimensions such as medicinal material category, medicinal property, and meridian tropism, and the semantic similarity is calculated based on the retrieval request features to screen out the supplementary nodes that are closest to the user's needs. The supplementary nodes are fused with the original path to generate a complete-structured, semantically rich, and target-aligned retrieval result for genuine medicinal materials. Different from the prior art where path missing is usually processed by manual annotation or static templates, the present invention constructs an automatic completion mechanism based on semantic type mapping and near-neighbor reasoning, significantly improving the system's fault tolerance, reasoning ability, and output integrity for incomplete paths, and ensuring that even when there are holes in the knowledge graph, intelligent recommendation results with reasonable structures and accurate semantics can still be provided.
[0022] Optionally, the present application further provides a smart retrieval system for genuine medicinal materials based on an atlas, which is used to execute the smart retrieval method for genuine medicinal materials based on an atlas as described above. The smart retrieval system for genuine medicinal materials based on an atlas includes: A semantic granularity analysis module, which is used to obtain the retrieval request data for genuine medicinal materials, perform semantic granularity processing on the retrieval request data for genuine medicinal materials to obtain semantic granularity data; and generate hierarchical semantic edges based on the semantic granularity data to obtain hierarchical semantic edge data; A hierarchical semantic graph construction module, which is used to generate a hierarchical semantic graph for a preset knowledge graph of genuine medicinal materials based on the hierarchical semantic edge data to obtain hierarchical semantic graph data; A semantic edge annotation module, which is used to perform semantic annotation on the relationship edges of the hierarchical semantic graph data to obtain semantically annotated graph data; A smart semantic-driven retrieval module, which is used to perform multi-condition intention processing based on the retrieval request data for genuine medicinal materials to obtain multi-condition intention data; perform multi-hop path combination on the semantically annotated graph data based on the multi-condition intention data to obtain preliminary retrieval data; and perform edge reasoning and completion on the preliminary retrieval data to obtain retrieval data for genuine medicinal materials.
[0023] The objective of the present invention is to perform semantic granularity processing on the retrieval request for genuine regional drugs, and generate hierarchical semantic edge data based on the ontology of traditional Chinese medicine knowledge. This not only realizes multi-granularity understanding of dimensions such as region, drug properties, and meridians entered, but also enhances the expressive ability of the semantic structure in the graph. Step S2 constructs a hierarchical semantic graph, enabling the graph to have the flexible combination ability of multi-level structure representation and reasoning path. Step S3 improves the interpretability and structural integrity of the semantic graph through context nesting, semantic tensor fusion, and edge relationship annotation driven by path probability. Step S4 combines the multi-condition intention input of the user, constructs a multi-hop path combination, and implements an edge node complementation strategy in the reasoning result to fill the semantic void area and achieve more complete and accurate drug recommendation. Compared with the prior art, the overall method no longer relies on single entity matching or shallow rule retrieval, but constructs a dynamic retrieval framework with multi-dimensional semantic understanding, graph structure driving, and reasoning complementation capabilities, significantly improving the adaptability, intelligence, and retrieval accuracy of the system when facing fuzzy, complex, or incomplete retrieval intentions. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Other features, objectives, and advantages of this application will become more apparent by reading the detailed description of the non-restrictive embodiments with reference to the following drawings: Figure 1 FIG. shows a flowchart of the steps of an intelligent retrieval method for genuine regional drugs based on a graph in one embodiment; Figure 2 FIG. shows a flowchart of the steps of a semantic granularity analysis method in one embodiment; Figure 3 FIG. shows a flowchart of the steps of a hierarchical semantic graph construction method in one embodiment; Figure 4 FIG. shows a flowchart of the steps of a semantic edge annotation method in one embodiment; Figure 5 FIG. shows a flowchart of the steps of an intelligent semantic-driven retrieval method in one embodiment; The realization, functional characteristics, and advantages of the objective of the present invention will be further described with reference to the embodiments and the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The technical method of the present invention will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0026] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0027] It should be understood that although terms such as "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.
[0028] Please refer to Figures 1 to 5 , this application provides a method for intelligent retrieval of genuine regional drugs based on a map, and the method includes: S1. Obtain the genuine regional drug retrieval request data, perform semantic granularity processing on the genuine regional drug retrieval request data to obtain semantic granularity data; generate hierarchical semantic edges according to the semantic granularity data to obtain hierarchical semantic edge data; In one embodiment, the acquisition methods of the authentic medicinal materials retrieval request data include, but are not limited to, keywords input by users, query statements, or retrieval questions expressed in natural language form. Taking "What are the similar medicinal materials to Fritillaria cirrhosa for treating cough?" as an example, this request belongs to a natural language retrieval statement. When performing semantic granularity processing on the retrieval request data, a natural language processing model can be used for word vector encoding, such as the BERT or ERNIE model or other pre-set large text classification models, to enhance the expression ability of semantic features. After completing the encoding, combined with part-of-speech tagging and named entity recognition methods, the key semantic entities included are identified, including but not limited to medicinal material names, medicinal material properties, symptom names, usage intentions, etc. The above entity types can be classified and summarized according to the pre-set traditional Chinese medicine semantic label system to form structured semantic granularity data. For example, after processing the above example statement, the obtained semantic granularity data includes: the medicinal material name is "Fritillaria cirrhosa", the symptom name is "cough", the functional intention is "treatment", and the remaining content that does not belong to the predefined label system can be treated as other items. Further, a hierarchical semantic edge is constructed according to the semantic granularity data. This hierarchical structure is set based on the semantic relationship of the traditional Chinese medicine knowledge graph, setting a multi-level semantic hierarchy from "symptom" to "functional intention", then to "medicinal material entity", and extending to "similar medicinal materials". Specifically, the semantic path constructed according to the above example can be expressed as: the "cough" node is connected to the "treatment" node through the "has_effect" edge (has_effect), the latter is connected to the "Fritillaria cirrhosa" node through the "applicable medicinal material" edge, and then is connected to the "Fritillaria thunbergii" node through the "similar_to" edge (similar_to). The above edges can be specifically marked as semantic edge types such as "has_symptom" (associated symptom), "has_effect" (has effect), "similar_to" (similar medicinal materials) in the semantic graph. The edges have directions to represent the semantic flow between entities and can be assigned different semantic labels and weights. The "medicinal property" information of medicinal materials can not only come from the records in traditional pharmacopoeias, but also be comprehensively modeled by combining the microbial components, endophytic fungi, or actinomycete metabolites of medicinal materials obtained in modern pharmaceutical research. For example, in the research on the authentic medicinal material Magnolia officinalis, through the pure culture, gene sequence alignment, and potential metabolic function analysis of the endophytic actinomycetes isolated from its roots, stems, bark and other tissues, its modern pharmacological properties such as antibacterial and anti-inflammatory can be deduced. These pharmacological properties can be further mapped to the functional medicinal property expressions such as "clearing heat", "detoxifying", and "removing dampness" under the traditional Chinese medicine theory system, and then used as the supporting source of the "medicinal property" node of the medicinal material in the graph.
[0029] S2. Generate a hierarchical semantic graph for the pre-set authentic medicinal materials knowledge graph according to the hierarchical semantic edge data to obtain hierarchical semantic graph data; In one embodiment, the authentic Chinese medicinal materials knowledge graph is a preset semantic structure graph constructed based on the field of traditional Chinese medicine, and its basic structure is in the form of triples, represented as the structure form of "entity 1 - relationship - entity 2". For example, it may include semantic triples such as "there is an 'applicable' relationship between the symptom 'cough' and the medicinal material 'Fritillaria cirrhosa'", "the medicinal material 'Fritillaria cirrhosa' belongs to the 'Sichuan-produced medicinal materials' region", and "there is a 'kindred medicinal materials' relationship between the medicinal material 'Fritillaria cirrhosa' and 'Fritillaria thunbergii'". Each pair of head and tail entities in the above triples are connected through specific edge relationships, forming the edge structure in the graph, which is convenient for subsequent entity connection and information dissemination in the semantic dimension. During the process of generating the hierarchical semantic graph, preferably, a semantic graph construction tool that supports knowledge graph processing can be called to perform graph structure generation operations on the matched semantic edges and their corresponding entities. During the construction process, semantic attribute annotation can be performed on the entity nodes connected by each semantic edge. The semantic attributes of the entity nodes may include, but are not limited to, the origin information, component composition, historical usage frequency, and usage literature coverage of the medicinal material; each semantic edge can also carry edge attribute information, such as edge weight (representing the connection strength or co-occurrence probability) and semantic confidence (representing the credibility of the relationship). Further, the hierarchical semantic graph is organized according to the functional semantic levels. Preferably, according to the semantic category division in the authentic Chinese medicinal materials knowledge graph, the graph structure can be divided into multiple semantic levels, including but not limited to the symptom layer, the function layer (representing the efficacy attribution), the medicinal material layer (representing the candidate traditional Chinese medicine entity), and the region / similar kindred expansion layer (for expanding regional information and similar kindred medicinal material relationships). The connections between the above semantic layer levels are directional. For example, the symptom layer is connected to the function layer, the function layer is connected to the medicinal material layer, and the medicinal material layer is then connected to the regional information or similar kindred medicinal material information, forming a hierarchical graph structure that can be used for subsequent path reasoning and node annotation. Preferably, during the graph traversal process, the generation depth of the graph can be controlled according to the set path hop threshold. For example, when performing multi-hop path reasoning, the maximum traversal depth is limited to no more than three layers to avoid information drift problems during the semantic expansion process and ensure that the constructed hierarchical semantic graph has good domain focus and reasoning controllability.
[0030] S3. Perform semantic annotation on the relationship edges of the hierarchical semantic graph data to obtain semantic annotation data of the graph; In one embodiment, all edge relationships in the semantic graph can be semantically annotated based on a preset edge semantic mapping dictionary. The edge semantic mapping dictionary is used to define the correspondence between edge types in the knowledge graph and natural language semantic labels, facilitating subsequent visualization of the graph structure and semantic reasoning processing. The specific mapping relationships may include, but are not limited to, mapping "has_symptom" to "can be associated with symptoms"; "from_region" to "origin of production"; "similar_to" to "similar in medicinal properties"; "used_with" to "commonly paired medicinal materials"; other edge types can also be flexibly extended according to the requirements of knowledge graph construction. When performing annotation, according to the edge type of each edge, the corresponding natural language description can be found in the semantic mapping dictionary, and this semantic label is assigned to the corresponding edge to form structured semantic graph annotation data. The annotation result of each edge includes the starting entity, the ending entity, and their corresponding semantic labels. For example, if the "Fritillaria cirrhosa" node points to the "cough" node and the edge type is "has_symptom", then this edge can be annotated as "can be associated with symptoms". Further, optionally, semantic intensity weight calculation is performed on the already annotated edges. The edge weight is used to measure the importance of this edge in the semantic graph structure and is used as a sorting basis in subsequent path selection and multi-hop path reasoning. The calculation method of the weight may include, but is not limited to, using the term frequency-inverse document frequency (TF-IDF) method, and calculating the relative importance of the edge according to the co-occurrence frequency of the entities connected by the edge in the knowledge base or the corpus of prescription texts; based on the above calculation results, a normalized edge weight value can be assigned to each edge to represent its semantic association strength.
[0031] S4. Perform multi-condition intention processing based on the genuine medicinal materials retrieval request data to obtain multi-condition intention data; perform multi-hop path combination on the semantic graph annotation data according to the multi-condition intention data to obtain preliminary retrieval data; perform edge reasoning and complementation on the preliminary retrieval data to obtain genuine medicinal materials retrieval data.
[0032] In one embodiment, the multi-condition intention processing includes extracting intention dimensions from the retrieval request input by the user. The user request usually contains multiple semantic dimensions, and typical dimensions include, but are not limited to, target symptoms, functional intentions, regional preferences, common medicinal material collocation relationships, and the properties of the nature and flavor of medicinal properties, etc. To achieve the above intention extraction, an intention classification model based on natural language understanding can be preset, and a multi-label classification mechanism is used to perform structured parsing on the original retrieval statement to obtain a set of semantic labels covering multiple dimensions. For example, for a retrieval statement containing information such as "cough", "Sichuan", and "clearing heat and resolving phlegm", after parsing, it can be extracted that the "symptom" is "cough", the "regional preference" is "Sichuan", and the "functional preference" is "clearing heat and resolving phlegm". After obtaining the multi-condition intention data, path combination operations can be performed based on the semantic graph annotation data. Preferably, a weighted graph search strategy is adopted, such as A An algorithm or a breadth-first search (BFS) strategy with a limited depth is used to perform multi-hop path traversal on the nodes and edges in the semantic graph to search for a path sequence that satisfies multiple intent constraints of the user. The nodes in the path need to meet the intent constraint conditions. For example, the starting node is "cough", the intermediate nodes are related to "clearing heat", and the end node is "Fritillaria cirrhosa" or its similar medicinal materials. During the path combination process, preferably, path screening rules are set to preferentially retain paths with a higher intent hit rate.
[0033] Perform a semantic edge completion operation on the above-mentioned preliminary retrieval paths to repair path breakpoints or relationship missing areas in the semantic graph. The completion operation can be implemented through a graph neural network model (such as a relational graph convolutional network R-GCN) or a rule-based graph structure inference engine. The completion logic can include the following two categories. If two medicinal materials co-occur frequently in multiple prescriptions, it can be inferred that they are similar medicinal materials. If a medicinal material A has been marked as having a semantic association with a certain semantic entity (such as "cough"), and there is a high structural similarity or co-occurrence degree between medicinal material B and A, it can be inferred that medicinal material B also has a potential association with this semantic entity. For example, in the original semantic graph, if there is no direct semantic connection established between "Thunberg fritillary bulb" and "cough", but there is a semantic edge from "Fritillaria cirrhosa" to "cough", and there is a "similar medicinal materials" semantic edge between "Fritillaria cirrhosa" and "Thunberg fritillary bulb", then based on the above rules, the edge relationship between "Thunberg fritillary bulb" and "cough" can be inferred. The confidence score of the inferred edge can be evaluated by a similarity model. If the score is 0.76, this relationship can be added as a completed edge to the graph structure. Combining the preliminary retrieval path and the inference completion result, structured retrieval data of genuine regional medicinal materials including fields such as the name of the target medicinal material, origin attribution, functional attributes, and recommended paths is output.
[0034] Optionally, S1 includes: S11. Obtain the retrieval request data of genuine regional medicinal materials; In one embodiment, a Chinese text tokenization tool (such as Jieba tokenizer or a custom tokenization module built based on traditional Chinese medicine professional corpus) is used to perform preliminary tokenization on the retrieval request text to obtain a sequence of basic lexical units. The system applies a named entity recognition model (NER) to the above-mentioned lexical unit sequence. This model can be constructed based on deep learning or rule-driven methods and is used to identify semantic entities related to the field of traditional Chinese medicine. The recognized entity types include at least, but are not limited to, the following categories: symptom entities, which represent physiological or perceptual characteristics mentioned in the user input, such as "cough", "dry mouth", etc.; medicinal material entities, which represent the names of traditional Chinese medicines with geographical attributes or efficacy attributes, such as "Fritillaria cirrhosa", "Astragalus membranaceus", etc.; function entities, which represent the common traditional usage functions of medicinal materials, such as "clearing heat", "tonifying qi", "removing dampness", etc.; geographical entities, which represent the names of geographical regions related to the authentic producing areas, such as "Yunnan", "Daozhou", "Bozhou", etc. The named entity recognition model can be trained based on structures such as BERT, BiLSTM-CRF, etc., or constructed by matching a rule base with a dictionary, with the aim of mapping the user input into structured information.
[0035] S12. Coarsely process the retrieval request data for authentic medicinal materials to obtain coarsely grained semantic data; In one embodiment, the system uses a preset dictionary of authentic medicinal material attributes to match the words in the user input text. This dictionary includes, but is not limited to, the names of medicinal material categories, geographical attribute words, and common efficacy description phrases. The matching process can be completed by means such as string similarity and stem matching to initially identify the terms in the text that belong to specific semantic fields. The system introduces a vector representation mechanism based on a pre-trained language model (such as the BERT model) to calculate the semantic similarity between the terms in the text and the standard field terms. The similarity calculation is based on the cosine similarity of word vectors. Based on the above matching and calculation processes, the system can classify the keywords or phrases in the user input into the following three categories of coarsely grained semantic fields: the medicinal material category field, such as "rhizome class", "flower class", "fruit class", etc., which represents the basic plant part attribution of traditional Chinese medicines; the geographical attribute field, such as "produced in Sichuan", "produced in Zhejiang", "Lingnan", etc., which represents the geographical origin characteristics of authentic medicinal materials; the basic efficacy field, such as "clearing heat", "moistening the lungs", "tonifying deficiency", etc., which represents the general function labels widely used in traditional Chinese medicine to describe the common application directions of medicinal materials (not referring to specific diseases or treatment behaviors).
[0036] S13. Fine-grainedly process the retrieval request data for authentic medicinal materials to obtain fine-grained semantic data, where the semantic granularity processing is carried out according to the preset traditional Chinese medicine semantic ontology and granularity stratification rules. The coarse semantic granularity is the medicinal material category, geographical attribute, and basic efficacy, and the fine semantic granularity is the meridian tropism, medicinal property, collection season, and classical compatibility; In one embodiment, the fine-grained semantic processing is based on a preset traditional Chinese medicine semantic ontology structure and semantic granularity stratification rules. The semantic granularity levels include coarse-grained semantic items, including medicinal material categories, regional attributes, and basic efficacy, etc.; fine-grained semantic items, including meridian attribution (such as belonging to the heart meridian, lung meridian, kidney meridian, etc.), medicinal property characteristics (such as the four natures of cold, heat, warm, and cool, and the five flavors of pungent, sweet, bitter, sour, and salty), collection seasons (such as "autumn", "before winter", "after frost's descent", etc.), and classical compatibility information (such as "compatible with almond", "used in the formula of Qingfei Decoction", etc.). The system performs fine-grained semantic entity extraction on the retrieved content based on the constructed traditional Chinese medicine semantic ontology structure and granularity stratification rule table. The extracted semantic items include meridian information, medicinal property characteristics, collection seasons, and classical compatibility, etc. The system realizes the intelligent recognition of classical compatibility relationships by constructing a traditional Chinese medicine knowledge graph and its supporting semantic reasoning rules. For example, when the co-occurrence frequency of a certain medicinal material and a specific formula in the pharmacopoeia literature is higher than the set threshold, it can be inferred that there is a typical compatibility relationship. In addition, by introducing traditional Chinese medicine literature such as "Compendium of Materia Medica", a vectorized text index structure is constructed, which improves the accuracy and coverage ability of semantic extraction and reasoning.
[0037] S14. Construct semantic paths and calculate multi-level edge weights for the coarse semantic granularity data and the fine semantic granularity data to obtain semantic path data and multi-level edge weight data; In one embodiment, construct a semantic path starting from a query entity (such as "moistening the lungs"); each path is in the form of a triple link: [moistening the lungs] — can be attributed to → [lung meridian] — medicinal property → [sweet, warm] — origin → [Sichuan] —→ [Fritillaria cirrhosa]. Each semantic edge is assigned a weight, and the weight of each edge is defined as: , is the weight coefficient of the ontological structure distance or cosine similarity between entities, with a value of 0.4, is the ontological structure distance or cosine similarity between entities, is the weight of whether the edge relationship comes from an authoritative knowledge ontology, with a value of 0.3, is whether the edge relationship comes from an authoritative knowledge ontology, is the weight of the co-occurrence frequency in historical prescription medications, with a value of 0.3, is the co-occurrence frequency in historical prescription medications.
[0038] S15. Generate a semantic hierarchy graph based on the semantic path data and the multi-level edge weight data to obtain multi-granularity semantic hierarchy graph data; In one embodiment, the graph nodes in the semantic hierarchy graph are divided into multiple levels according to semantic granularity. The typical structure includes but is not limited to the following three levels. The first level (function category) represents the general functional directions of medicinal materials, such as "moistening the lungs", "tonifying deficiency", "removing dampness", etc., which belong to coarse-grained semantic entities. The second level (traditional Chinese medicine attributes) includes the meridians that the medicinal materials belong to (such as "lung meridian", "spleen meridian"), the properties of the medicinal materials (such as "sweet", "warm", "cold"), and the application seasons (such as "spring", "autumn"), etc., serving as an intermediate layer between the function category and the specific medicinal materials. The third level (specific medicinal material entity) represents specific traditional Chinese medicine instances, such as "Fritillaria cirrhosa", "Astragalus membranaceus", "Ophiopogon japonicus", etc., which belong to fine-grained semantic entities. The above levels are connected by relationship edges in the semantic path data to establish a graph structure connection, and each edge is weighted and modeled according to its semantic relationship type and edge weight data. For example, an indirect connection is established between "moistening the lungs" and "Fritillaria cirrhosa" through the path of "lung meridian → warm nature → autumn", and each jump edge has a corresponding semantic label and edge weight value.
[0039] S16. Perform semantic edge classification and annotation according to the multi-granularity semantic hierarchy graph data to obtain hierarchical semantic edge data.
[0040] In one embodiment, the semantic types of the edges may include but are not limited to the following categories: efficacy edge (function attribution relationship), property edge (taste and property attribute relationship), geographical edge (origin attribution relationship), or compatibility edge (common collocation relationship). Define a set of matching rules for known semantic patterns and perform category annotation on the edges that meet the patterns. For example, the following is an example description of the matching logic of a rule template. If the starting node type of the edge connection is "geographical area", the ending node type is "medicinal material", and the edge relationship is "source_of" or a synonymous relationship, then the edge is labeled as "origin attribution". Similarly, rules such as "function node → function node" corresponding to "efficacy association" and "taste and property attribute → medicinal material entity" corresponding to "property direction" can also be set. The rule templates can be stored as a structured matching rule set, and combined with entity type judgment, path structure verification, etc., to complete the semantic annotation operation of the edges.
[0041] Optionally, S2 includes: S21. Perform function attribution semantic perception mapping and geographical co-occurrence semantic perception mapping on the preset genuine medicinal materials knowledge graph according to the hierarchical semantic edge data to obtain function attribution mapping data and geographical co-occurrence mapping data respectively; In one embodiment, identify the connection relationship between the medicinal material and its "attributed function layer". For example, confirm the functional semantic landing point of Fritillaria cirrhosa from the path of "moistening the lungs → lung meridian → Fritillaria cirrhosa". Extract the path of "function → meridian → medicinal material" from the hierarchical semantic edge data; map the medicinal material node to its corresponding function attribution label, such as "medicinal materials for moistening the lungs"; introduce a similarity function to support fuzzy matching: , be be the functional semantic keyword be the name of the medicinal material node or descriptive text be the functional word vector be the word vector related to the medicinal material
[0042] S22. Perform node twin expansion based on the function attribution mapping data and the regional co-occurrence mapping data to obtain semantic completion data; In one embodiment, the system jointly analyzes the function labels and regional attributes of the medicinal material nodes and finds that there is a phenomenon of functional convergence of medicinal materials within a specific regional scope. For example, "medicinal materials produced in Sichuan often co-occur in the formulas for moistening the lungs". The co-occurrence frequency in the medicinal material-region-function triple is statistically calculated to construct a co-occurrence matrix : be the medicinal material and the region co-occurrence count value be the medicinal material and the region times of co-occurrence in the functional semantic path be the th medicinal material node be the th regional node. Use PMI to calculate the regional co-occurrence intensity: , where be the co-occurrence probability of the two be the marginal probability of the medicinal material be the marginal probability of the region be the medicinal material be the region
[0043] Map the nodes that are semantically similar or co-occur frequently as "twins" to expand hidden medicinal materials or functional entities that are not explicitly connected. Define that the twin nodes meet the conditions of belonging to the same function label; regional co-occurrence intensity > 0.6; similarity of the medicinal property structure (taste, meridian tropism) > 0.75. There is no connection of "medicinal material A function F" in the original graph; if there exists a twin node B and there is a clear connection of "B F"; then construct a "soft connection": . There is no "Tussilago farfara moistening the lungs" in the original graph, but "Tussilago farfara ~ Fritillaria cirrhosa" (i.e., there is a semantic similarity relationship) and "Fritillaria cirrhosa moistening the lungs", then it is supplemented. The twin relationship can be represented by a "virtual edge" in the graph structure, with the edge weight < 0.5, but it participates in path reasoning.
[0044] S23. Perform structure clustering based on the semantic incomplete data to obtain hierarchical semantic graph data.
[0045] In one embodiment, the spectral embedding representation of nodes is obtained by calculating the Laplacian matrix of the graph and performing eigenvalue decomposition, and the nodes are grouped using a clustering method (such as K-Means) in the spectral space to obtain hierarchical semantic graph data.
[0046] In one embodiment, the system combines the multi-dimensional attribute similarity between node vectors (such as Euclidean distance, cosine similarity) and the structural connectivity relationship between nodes, calculates the similarity between the vector representations of any two nodes; based on the preliminary clustering results, connectivity constraint optimization is performed. For nodes that are assigned to the same class but lack sufficient structural edges (such as node pairs with less than the specified strong connection threshold), reclassification or isolation operations are performed; for node pairs with high structural weight edges, aggregation is preferentially maintained to ensure that the clustering results take into account both attribute similarity and topological connectivity. The clustering results are mapped to the hierarchical structure of the semantic graph, forming the following layers: The first layer: functional topic nodes, representing abstract categories of the uses of traditional Chinese medicines, such as "moistening the lungs", "clearing heat", "dispelling wind" and so on; The second layer: regional attribution nodes, indicating the geographical distribution attribution of specific traditional Chinese medicines in the semantic space, such as "produced in Sichuan", "produced in Zhejiang", "Lingnan" and so on; The third layer: entity traditional Chinese medicine nodes, specifically corresponding to traditional Chinese medicine instances, such as "Fritillaria cirrhosa", "Scutellaria baicalensis", "Honeysuckle" and so on.
[0047] Optionally, the functional attribution semantic perception mapping is specifically: Filter the functional semantic trigger edges according to the hierarchical semantic edge data to obtain the functional semantic edge data; In one embodiment, the system identifies several types of functional semantic connections from the hierarchical semantic edge data, specifically including connection relationships representing treatment goals (such as "having a certain efficacy"), connection relationships representing meridian attribution directions (such as "acting on a certain meridian"), connection relationships representing the usage in typical prescriptions (such as "applied in a certain type of prescription") and connection relationships for relieving specific symptoms, etc. The system presets a set of functional semantic edge types, including the following semantic types: hasEffectOn representing efficacy; targetsMeridian representing meridian attribution; usedInPrescriptionFor representing prescription application relationship; relievesSymptom representing symptom relief effect. The system traverses the entire set of candidate semantic edges and filters them according to whether the semantic type of the edge belongs to the above set of functional semantic types. If a semantic edge meets the conditions, it is included in the set of functional semantic edges as the starting or intermediate link node for triggering the functional semantic path.
[0048] Construct a functional semantic path for the preset genuine regional traditional Chinese medicine knowledge graph according to the functional semantic edge data to obtain the functional semantic path subgraph data; In one embodiment, starting from the functional semantic nodes, a functional association path graph is generated in the preset knowledge graph for analyzing the attribution link of medicinal materials. In the authentic medicinal materials knowledge graph, a functional node is selected as the starting point; the number of path hops is restricted to ≤3, and the path relationship is limited to the functional-related semantic edges; breadth-first search (BFS) or graph traversal is used. Each hop edge type needs to be a functional semantic trigger edge or its subtype; jump paths (such as directly jumping from function to property without an intermediate node) are excluded. During the traversal process, it is restricted that each hop path passes through edges belonging to the "functional semantic trigger edge" set or its predefined subclass relationship edges. The semantic types of such edges include, but are not limited to, efficacy inheritance edges, function pointing edges, function attribution edges, etc. For example, starting from the functional node "clearing heat", it can point to "moistening the lungs" through the "has_function" type edge, then point to the lung meridian through the "affects_meridian" edge, and further connect to the specific medicinal material node.
[0049] Semantic path context verification is performed according to the hierarchical semantic edge data and the functional semantic path subgraph data to obtain functional context verification data; In one embodiment, the query context refers to the set of key semantic elements extracted from the user's retrieval request, usually including information in multiple dimensions, such as functional keywords (such as "moistening the lungs"); meridian attribution (such as "lung meridian"); application solar terms (such as "autumn"); property characteristics (such as "sweet and warm"), etc. This context set constitutes the benchmark standard for semantic path determination. The functional semantic path is represented by a node sequence. For each node in the path, the system determines whether it matches a certain element in the current query context according to its attribute value. If there is a match, it is included in the context match item of the path. Define the path context match degree: , if ContextMatch(P)<0.5, it is marked as a path with inconsistent context.
[0050] Semantic coverage conflict detection is performed on the functional context verification data to obtain functional semantic coverage conflict data; In one embodiment, when a medicinal material belongs to two functional categories with semantic conflicts (such as "warming yang" and "clearing heat"); or multiple paths point to different meridians (such as "heart meridian" and "lung meridian") and there is no homologous logic support. The system calls the functional mutual exclusion rule table and the meridian independence judgment table set in the traditional Chinese medicine semantic ontology to perform pairwise verification on all path pairs. If a certain medicinal material involves multiple functional semantic paths, the system will construct a conflict judgment matrix to identify whether there is a semantic conflict between any two paths: . The conflict path set is marked for the next precision comparison process.
[0051] Mapping accuracy comparison is performed according to the functional semantic coverage conflict data to obtain functional attribution mapping data.
[0052] In one embodiment, the system detects the structural integrity of each functional path, and the path should meet the following structural requirements: the starting point is a functional node, the ending point is a medicinal material node, and the middle of the path should include the meridian tropism node and the medicinal property node. The system introduces a context matching score to measure the matching degree between the candidate path and the user's current query semantics. The calculation of this score combines the TF-IDF weights of the keywords in the path and the cosine similarity between the path semantic representation and the query vector, that is, weighted calculation, with the weight of the former being 0.6 and the latter being 0.4. To characterize the semantic credibility of the path, the system also assigns an inference confidence score to each path. If the path comes from an explicit relationship in the knowledge graph ontology, its confidence is relatively high; if the path is a soft connection path inferred and complemented by means such as twin mapping, a lower confidence is given.
[0053] Optionally, the regional co-occurrence semantic perception mapping is specifically as follows: Extract regional-related trigger nodes according to the hierarchical semantic edge data to obtain regional node data; In one embodiment, a type of "regional semantic edge type" can be predefined to identify the edge relationships between entities related to regional attribution, regional use, regional origin, etc. Typical regional edge types include but are not limited to fromRegion: indicating that a certain medicinal material or entity comes from a certain geographical region; originProvince: indicating the provincial annotation of the entity; recordedInArea: indicating that a certain medicinal material or entity is recorded in a specific local chronicles; usedInLocalPrescription: indicating that a certain medicinal material appears or is applied in a local prescription. During the actual extraction process, traverse the edge set in the hierarchical semantic graph and filter out all edges with regional pointing significance. Extract all endpoint nodes connected to regional entities from the above-mentioned filtered regional edge set, that is, all regional nodes. The regional node is the starting end or the ending end of the edge, which is specifically determined according to the semantic directionality of the edge.
[0054] Perform medicinal material regional co-occurrence path mapping on the preset genuine medicinal materials knowledge graph according to the regional node data to obtain medicinal material regional co-occurrence data; In one embodiment, starting from the regional node, the system searches for the prescription nodes associated with this region, and then further searches for the medicinal material nodes included in this prescription, that is, the following three-hop path is formed: , or the system starts from the medicinal material node and searches for the directly connected regional node, forming the following two-hop path, . The system sets a maximum path length limit, which does not exceed three hops. The system can execute path query tasks based on a graph database engine (such as Neo4j), and starting from the regional node through the graph traversal strategy, map all reachable medicinal material nodes and their path structures according to the above mode.
[0055] Semantically expand the co-occurrence data of medicinal materials by region to obtain candidate path data for regional co-occurrence; In one embodiment, the system starts from the medicinal material node and performs a reverse path tracing operation in the semantic graph structure to gradually identify all paths that are semantically connected to the regional entity. In addition to the direct "medicinal material - place of origin" path, the system also identifies a type of indirect regional semantic path. For example, if a medicinal material is involved in a certain prescription and the symptoms targeted by this prescription are prevalent in a specific region, then a path can be constructed: "medicinal material" → "prescription" → "symptom" → "region", and this type of path is classified as a soft co-occurrence path. The system performs semantic label annotation on all co-occurrence paths, specifically including the direct place of origin path, which represents that there is a clear place of origin pointing relationship between the medicinal material and the region; the formula co-occurrence path, which represents that the prescriptions in which the medicinal material participates are widely used in a specific region; and the symptom transmission path, which represents that the symptoms treated by the medicinal material are prevalent in a specific region.
[0056] Calculate the path co-occurrence weight for the candidate path data of regional co-occurrence to obtain path co-occurrence weight data; In one embodiment, calculate the credibility or importance of each candidate path in the regional co-occurrence semantics. Path co-occurrence weight calculation , is the path co-occurrence weight marking data, marking the candidate path data of regional co-occurrence to form path co-occurrence weight data, is the number of occurrences of the medicinal material-region in the corpus / graph for the path, is the weight of the number of occurrences of the medicinal material-region in the corpus / graph for the path, set to 0.35, is the path type, with values set according to the path semantic structure. For example, the direct place of origin path is greater than the indirect formula path, which is determined by a preset numerical mapping table of path structure types, is the weight of the path type, set to 0.35, is the path length penalty term. The greater the path length, the overall weight should decay (such as calculated as , where is the natural exponent, is the path hop length), is the weight of the path length penalty term, set to 0.3.
[0057] Generate regional co-occurrence mapping data based on the path co-occurrence weight data and the candidate path data of regional co-occurrence to obtain regional co-occurrence mapping data.
[0058] In one embodiment, for each medicinal material node, all its co-occurrence paths and corresponding weights are summarized; a weight threshold (such as 0.7) is set to filter out reliable paths; for each medicinal material node, the system summarizes all the regional nodes connected by the paths that meet the weight threshold to form a candidate region set and attaches the path sources. If a medicinal material co-occurs with multiple regions simultaneously, they are sorted by score, and the top 5-10 attribution candidates are retained; the main production area and secondary co-occurrence areas are marked.
[0059] Optionally, the node twin expansion is specifically as follows: Extract structurally similar node data according to the function attribution mapping data and the regional co-occurrence mapping data; In one embodiment, the system sets the trigger conditions for structural similarity according to the following rules to identify medicinal material nodes that show consistency or similarity in multiple semantic dimensions. The determination of medicinal property similarity includes that the medicinal materials have the same or similar nature and flavor attributes, such as "pungent and warm" and "pungent and hot"; the determination of compatibility co-occurrence includes that the medicinal materials co-occur with the same high-frequency compatible medicinal material (such as "licorice") in multiple prescriptions; the determination of meridian similarity includes that the medicinal materials belong to the same meridian or have similar meridian functions (such as "lung meridian" and "simultaneously treating the lung and spleen"); the determination of regional and prescription co-occurrence includes that the production areas of the medicinal materials are the same, or there are co-occurrence records in historical classics or modern compound prescriptions. To efficiently perform the above similarity extraction, the system uses a subgraph pattern matching algorithm to compare the adjacency structures of medicinal material nodes in the knowledge graph, extracts node groups with high structural similarity in the functional path and regional path structures, and labels them as structurally similar node groups.
[0060] Calculate the similarity of graph nodes according to the structurally similar node data to obtain the similarity data of graph nodes; In one embodiment, the calculation of the similarity of graph nodes includes graph adjacency similarity , , where is the neighbor set of node u (considering edge direction + semantic type); attribute embedding similarity , encode the node attributes (such as medicinal properties, meridians, production areas) into multi-dimensional feature vectors and calculate using cosine similarity; functional path structure similarity , the judgment basis includes the matching degree of the path topological structure and the similarity of the semantic edge weights in the path. The specific calculation can combine the path edit distance and the semantic edge weight difference index to calculate the path structure similarity of two nodes from the functional level to the entity level. Based on the foregoing, calculate the similarity data of graph nodes : , where the empirical parameter is set as is the weight of the graph adjacency similarity, with a value of 0.3, is the weight of the attribute embedding similarity, with a value of 0.3, is the weight of the functional path structure similarity, with a value of 0.4.
[0061] Perform twin node path inversion based on the graph node similarity data to obtain twin node path data; In one embodiment, the system retrieves a node B that has a high structural similarity and semantic feature with node A in an existing functional path P. If the twin scores (graph node similarity data) of A and B exceed a set threshold, then the latter B is regarded as a potential twin node of the former A. The system uses the original path P as a template and attempts to construct a path P' that is both structurally and semantically consistent with P for the twin node B, setting constraints such as keeping the semantic edges consistent; the replaceable nodes must appear in the twin mapping table; there are no semantic conflicts in the replacement path (such as inconsistent meridians, contradictory medicinal properties, time and season logic conflicts, etc.).
[0062] Calculate the path confidence based on the twin node path data to obtain path confidence data; In one embodiment, score each inversion path to indicate the credibility of the path. The path confidence score consists of the following four core dimensions: the structural consistency score , indicating whether it is isomorphic to the original path structure; the edge weight stability score , indicating whether the semantic weight of the edge after migration is not lower than that of the original path; the entity semantic compatibility score , indicating the attribute consistency of the replacement nodes in the path; the path context co-occurrence degree , indicating the co-occurrence frequency of the path in actual prescriptions or clinical corpora. , where is the weight of the structural consistency score, with a value of 0.3, is the weight of the edge weight stability score, with a value of 0.2, is the weight of the entity semantic compatibility score, with a value of 0.3, is the weight of the path context co-occurrence degree score, with a value of 0.2.
[0063] Generate semantic completion data for the twin node path data, functional attribution mapping data, and regional co-occurrence mapping data based on the path confidence data, semantic completion data.
[0064] In one embodiment, the system calculates the path confidence for all inversion paths generated by the inference mechanism. If the confidence score of a certain path is greater than or equal to a preset threshold (such as 0.7), then this path is included in the semantic completion result and is labeled according to its connection direction and semantic type. The completion methods include: if the path represents a potential functional attribution relationship, then a "soft functional attribution edge" is added to the graph; if the path exhibits a potential geographical co-occurrence relationship, then it is added as a "candidate geographical edge". To identify that such completed paths are not direct connections between original entities, the system attaches an inference source marker field to the edge structure, for example, setting isInferred=true, which is used to distinguish original edges from completed edges during graph management or path inference processes.
[0065] Optionally, S3 includes: S31. Extract the node semantic nested context from the hierarchical semantic graph data to obtain edge context semantic data; In one embodiment, in the hierarchical semantic graph, the context semantic information of each edge where a node is located is extracted to provide a semantic background for the construction of edge features. The extraction dimensions include the graph structure position, whether the node is in the function layer, geographical layer, or medicinal material layer; the adjacency relationship, the types of semantic edges directly connected to the current node and the attributes of adjacent nodes; the node label, functional attribution, geographical co-occurrence, medicinal property meridian entry, whether it is a twin completion, etc.; the edge context. For example, if the edge is "moistening the lungs → Fritillaria cirrhosa", then the edge context includes "lung meridian", "sweet and warm", "collected in autumn".
[0066] S32. Generate multi-source semantic feature tensors for the hierarchical semantic graph data according to the edge context semantic data to obtain edge feature data to be labeled; In one embodiment, the so-called edge context semantic data refers to the set of information such as the semantic of the nodes connected by the edge itself, the position of the edge in the path, the semantic level where the edge is located, and the semantic labels of other adjacent edges. The system expresses the above context semantic information in multiple dimensions and constructs a three-dimensional semantic feature tensor , where is the number of edges; is the number of feature channels (functional features, geographical features, structural features, etc.); is the vector dimension of each channel (for example, using BERT embedding as 768 dimensions). The feature channels include functional vectors (semantic embeddings of functional names); medicinal material vectors (fusion representations of node medicinal properties + meridian entry attributes); regional co-occurrence weight vectors; path attribution confidence vectors; adjacent edge label embeddings (upstream and downstream semantic influences).
[0067] S33. Calculate the path probability according to the edge feature data to be labeled to obtain path probability data; In one embodiment, the system evaluates each candidate path based on path probability modeling technology, with the goal of inferring whether it has semantic validity. The estimation of path probability can be achieved through the following two types of methods. For example, by using a Bayesian graph model to calculate, constructing a conditional probability graph based on the graph structure, modeling and reasoning about the joint distribution of each semantic edge in the path, so as to calculate the posterior probability that the path is a valid functional path; or, by using a neural path scoring model to calculate, encoding the semantic features of all edges in the path (such as edge type, directionality, co-occurrence frequency, semantic weight, etc.) into a tensor form as the input of the neural network model. The system performs a linear combination of the preset weight parameters and the feature tensor, adds a bias term, and uses the Sigmoid activation function to normalize the output result to obtain the probability score that the path is a valid functional path.
[0068] S34. Perform functional analogy annotation on the edge feature data to be annotated according to the path probability data to obtain preliminary semantic icon annotation data; In one embodiment, according to the path probability result, perform automatic preliminary functional semantic annotation on the edge. If ( is the semantic credibility of the candidate path, is the decision threshold of path credibility, which can be set according to the application scenario (such as set to 0.7)), then label the edge as the corresponding semantic relationship; for example, if the path is "moistening the lungs → Fritillaria cirrhosa", and P = 0.93, then label the edge type as hasEffectOn, indicating that the function "moistening the lungs" acts on the medicinal material "Fritillaria cirrhosa"; if the path is obtained by twin reasoning, then add the flag inferred=true; S35. Perform edge semantic conflict detection on the preliminary semantic icon annotation data to obtain semantic conflict detection data; In one embodiment, by counting the edge label distribution of all entity nodes in the semantic graph, an edge label co-occurrence matrix C(i,j) is constructed to represent the co-occurrence of label i and label j in the same node or the same path context. If a pair of logically exclusive labels appears simultaneously on a certain entity node or path structure, it is regarded as a semantic conflict. In the annotation path structure, it is detected whether there are conflicts in the semantic direction and attribution logic between each functional path. If an entity node participates in multiple paths, and there are inconsistent semantic attribution relationships in the functional attribution, meridian tropism system, drug property hierarchy, etc. to which each path belongs, it is considered that the path logic is inconsistent. A rule engine is constructed through preset parameters (preset parameters are set based on an expert knowledge base) to define constraint rules for edge label combinations and path structure logics. The rule engine makes judgments based on a preset set of semantic conflicts, meridian tropism logic rules, and drug property compatibility matrices. For example, if it is defined in the rule system that if any entity node has both the labels of "clearing heat" and "warming and tonifying" simultaneously, it is a conflict pair. The system can automatically traverse all edge annotations, determine whether the conflict trigger condition is met, record the conflict edge pairs, conflict types, and the information of the nodes to which they belong, and output structured semantic conflict detection data.
[0069] S36. Perform relationship completion on the preliminary semantic graph annotation data according to the semantic conflict detection data to obtain semantic graph annotation data.
[0070] In one embodiment, for semantic conflict edges, the system selects the semantic label corresponding to the path with higher confidence according to the path confidence data or semantic context consistency metric as the main semantic edge of the node pair; the remaining labels can be marked as "secondary candidates" or removed. For semantic missing edges, if the system has identified reliable twin paths (i.e., paths with high structural similarity) or has stable co-occurrence support basis (such as the same type of medicinal materials showing consistent semantic paths in different regions) in the graph, a complementary edge is added in the current graph, processed as an inference edge, and the identifier inferred=true is added to this edge to indicate that it is generated by inference rather than directly annotated. After the completion of the complementation and correction operations, the system uses a graph neural network model to perform unified type prediction on the updated edge set. The edge type prediction model is constructed by combining a graph convolutional neural network and a multi-label classifier, and the prediction function is defined as follows: , is the predicted semantic label of this edge, is a multi-layer perceptron classifier, is the graph embedding representation of edge node u, is the graph embedding representation of edge node v, is the context semantic feature vector of the edge.
[0071] Optionally, S4 includes: S41. Extract retrieval request feature data based on the retrieval request data of authentic Chinese medicinal materials. The retrieval request feature extraction includes regional preference feature extraction, meridian tropism feature extraction, and medicinal property restriction feature extraction; In one embodiment, the system combines natural language processing technology, keyword recognition rules, and the Chinese medicine knowledge ontology library to perform structured parsing on the user input statement. For example, for regional preference feature extraction, regional preference information is extracted by matching geographical related keywords (such as "produced in Sichuan", "Lingnan", etc.) in the user input. If the user does not explicitly specify a region, the system will supplement and infer the default regional preference in combination with the commonly used regional usage preferences or the medicinal material origin habits in the Chinese medicine field. For meridian tropism feature extraction, based on the syntax analysis result and the Chinese medicine meridian tropism ontology library, the meridian attribution related to the user input semantics is extracted. For example, for a statement involving specific symptom or function description, the system will mark the associated meridian direction (such as "lung meridian", "heart meridian") according to semantic reasoning to form the meridian tropism feature. For medicinal property restriction feature extraction, based on the preset Chinese medicine noun mapping rules, the medicinal property description information involved in the user input is identified, including the "four natures" category (cold, heat, warm, cool) and the "five flavors" category (pungent, sweet, bitter, sour, salty).
[0072] S42. Calculate the intention attention of the retrieval request feature data to obtain multi-condition intention data; In one embodiment, weight modeling is performed on multiple intention dimensions to clarify the attribute dimension that the user is most concerned about. The user request contains semantic intention dimensions, and an attention mechanism is established for the attention weight of each dimension : , is the attention score of the th dimension, indicating the importance of this dimension, is the exponential function, is the dimension order, is the weight parameter of the th dimension. The parameter can be a trainable item or a system setting value, is the th dimension feature embedding vector, generated by combining the context text content and the user portrait information.
[0073] S43. Construct the intention path candidate space for the semantic annotation data according to the multi-condition intention data to obtain the intention path candidate space data; In one embodiment, based on the above weighted intention information, a path space that meets the user's multi-dimensional conditions is constructed. In the labeled semantic graph, multi-dimensional filters are set, such as must include meridian nodes → satisfy meridian_pref (meridian constraint); the property nodes need to satisfy property_constraints (property restrictions); the geographical attribution path should at least include any region in region_pref (geographical preference). Use graph filtering to construct an intention constraint subgraph; the combination logic can include "and", "or", "exclude" semantic settings.
[0074] S44. Perform graph traversal according to the intention path candidate space data to obtain path candidate graph data; In one embodiment, perform structural traversal within the intention path space to obtain a set of candidate paths. The graph traversal uses restricted depth-first search (DFS) or breadth-first search (BFS); the restriction conditions are: path length ≤ 4; it should include at least one functional starting point and one medicinal material entity; all edge types ∈ the labeled semantic set (such as hasEffectOn, usedWith, fromRegion); S45. Calculate the semantic coverage rate according to the path candidate graph data and the retrieval request feature data to obtain semantic coverage rate data; In one embodiment, let the total number of intention dimensions included in the user's retrieval request be , and for a certain path , the number of intention dimensions covered is , then the semantic coverage rate of this path is denoted as , or, further, introduce a weighted semantic coverage rate function based on the attention mechanism, , is the weighted semantic coverage rate, which is a representation of the semantic coverage rate data, is the intention dimension index, is the total number of intention dimensions, If the path meets the -dimensional intention.
[0075] S46. Screen the path candidate graph data according to the semantic coverage rate data to obtain path screening data; In one embodiment, select the path with the highest semantic coverage rate and the strongest confidence from all candidate paths. First, select the path set with a coverage rate ≥ 0.8; if there are less than 5 paths, appropriately relax it to ≥ 0.6; the sorting logic is to perform weighted combination according to the coverage rate and the path confidence to perform sorting: , where is the comprehensive score of the path , is the semantic coverage rate of the path , is the path confidence level.
[0076] S47. Perform edge node semantic filling processing on the path screening data to obtain the authentic medicinal material retrieval data.
[0077] In one embodiment, the system identifies the edge nodes in the path structure. If a node lacks geographical attributes, medicinal property descriptions, or meridian tropism information, it is regarded as a semantic breakpoint. The system will use the entity twin reasoning mechanism in the path context information and the knowledge graph to intelligently fill in the missing attributes of the node. For example, when a path represents "moistening the lungs → Aster tataricus", and "Aster tataricus" is not labeled with geographical attributes, the system will retrieve other medicinal material entities with the same efficacy in the knowledge graph through its function label "moistening the lungs", and analyze their main production areas, so as to infer the geographical attribution of "Aster tataricus". The inference result is attached to the target node in the form of a soft label, including but not limited to inference attributes such as "presumed origin = Hubei". The system can add the complemented information as a node attribute field, or add a semantic edge in the graph to represent the inference association, and set an "inference source" identifier to indicate that the complemented information comes from the semantic reasoning process rather than explicit definition.
[0078] Optionally, the edge node semantic filling processing is specifically as follows: Perform boundary node analysis on the path screening data to obtain half-edge connection node data; In one embodiment, nodes that are only connected by one side, lack attribute annotations, or context relationships are identified in the path screening results, that is, semantic edge nodes. The system makes edge determination on the nodes through the following rules: the node degree deg() = 1 or the only connected edge does not contain semantic annotations; the node lacks core attribute fields (such as medicinal properties, regions, meridian tropisms); the node does not participate in a closed triple or lacks upstream and downstream path information.
[0079] Generate edge node missing data according to the half-edge connection node data; In one embodiment, identify which core semantic paths or attributes are missing from the edge nodes, and establish a set to be complemented. For each edge node, list the attributes that should be in the graph model but are missing in the current path, such as geographical attribution (fromRegion), medicinal properties (hasProperty), meridian tropism (targetsMeridian), and efficacy path (hasEffectOn).
[0080] Perform semantic path missing type mapping according to the edge node missing data to obtain semantic path missing type data; In one embodiment, the above missing information is mapped to a completable type. The system divides all missing cases into the following three types and makes explicit annotations. The missing type definitions are as follows: TypeA: missing attribute class nodes (such as lack of drug property / region); TypeB: missing edge class relationships (such as lack of functional connections); TypeC: missing upstream and downstream structural jump points (such as path faults).
[0081] Perform attribute near-neighbor matching based on the semantic path missing type data to obtain supplementary node data; In one embodiment, find the node in the knowledge graph that is most similar to the edge node structure attribute to migrate its missing attribute. Construct feature vectors (drug property, meridian tropism, function, compatibility frequency, etc.); use cosine similarity or graph embedding distance to measure candidate nodes: , is the semantic similarity score of the node, representing the node to be completed and the candidate node The overall semantic similarity, equivalent to Or the result of the distance function based on the graph embedding space, is the cosine similarity scoring function, used to calculate the cosine angle similarity between the semantic vectors of two nodes. The numerical range is [0,1], and the larger the value, the higher the similarity. is the feature vector (node to be completed), is the feature vector (candidate entity node), is the similarity threshold. For each missing item, extract the most similar (the top 5-10 after sorting from large to small) entities from the completion candidate set; extract their complete attributes and construct transferable supplementary items (such as "Aster tataricus → Region: Hubei").
[0082] Perform similarity screening on the supplementary node data according to the retrieval request feature data to obtain supplementary screening data; In one embodiment, the system constructs a matching scoring model for all candidate supplementary nodes according to the semantic preference information input by the user, including regional preference, meridian tropism direction, and drug property characteristics, etc., to evaluate the semantic fit degree between them and the user's retrieval target: , is the matching score of the candidate node , is the regional weight coefficient, with a value of 0.2, is the regional matching degree index, for example, RegionMatch = 1 (fully matched) or 0.5 (same provincial region), is the meridian tropism weight coefficient, with a value of 0.5, is the meridian tropism matching degree index, indicating the consistency degree between the meridian tropism attribute of the candidate node and the meridian preferred by the user. If they match, take 1; if they are partially relevant, take 0.5, etc., is the medicinal property weight coefficient, with a value of 0.3, is the medicinal property matching degree index, indicating the degree of conformity between the medicinal properties (four natures and five flavors) of the candidate nodes and the user's preferences, which can be calculated through label comparison or vector cosine. The system will sort according to the scoring results and preferentially retain the nodes with the highest scores as semantic supplement items. If there are multiple candidate nodes with close scores, the system can also perform node semantic fusion operations to integrate the information of multiple nodes into a virtual node.
[0083] Integrate the supplementary screening data and the path screening data to obtain the genuine medicinal materials retrieval data.
[0084] In one embodiment, the semantic path information generated by the complement mechanism is embedded into the original path structure, while the expression format of the attribute fields is unified, and the source information of the reasoning or complement content is recorded, so that the retrieval output result has both integrity and traceability. The integration methods include but are not limited to the following aspects. For example, path information field merging, where the path sequence obtained in the original path screening stage is fused with the path information generated in the complement reasoning stage, while maintaining the order and semantic consistency of the original path structure. Attribute field completion, for each medicinal material entity node, supplement its missing semantic attribute fields, including geographical attribution (such as "source area: Sichuan"), medicinal property labels (such as "nature and flavor: sweet and warm"), meridian tropism information (such as "attributed to the lung meridian"), etc., to ensure that all core semantic labels are clearly marked in the final result. Complemented path identification, explicitly identify the paths or attribute information generated by the semantic reasoning or structure complement mechanism. Preferably, a "reasoning flag" field can be set to indicate whether this part of the information is generated by reasoning or twin nodes, for example, setting the field inferred=true or the complement source type as "structural similarity inference". Reasoning source record, for each piece of complemented information, record its source path or reasoning basis. For example, if the meridian tropism information of a certain medicinal material comes from the co-occurrence path of medicinal materials with similar structures, it is noted in the result as "information source: inferred from medicinal materials with the same path as Fritillaria cirrhosa". The output genuine medicinal materials retrieval data can be organized in the form of structured data, such as a list of structured key-value pairs, a graph structure JSON expression, or a visual recommendation list, etc. The content includes fields such as medicinal material name, path structure, geographical information, medicinal property label, meridian tropism attribute, reasoning source, path score, etc., constituting a multi-dimensional integrated intelligent retrieval response result.
[0085] Optionally, the present application also provides a graph-based genuine medicinal materials intelligent retrieval system for performing the graph-based genuine medicinal materials intelligent retrieval method described above. The graph-based genuine medicinal materials intelligent retrieval system includes: A semantic granularity analysis module, configured to obtain the data of the authentic medicinal materials retrieval request, perform semantic granularity processing on the data of the authentic medicinal materials retrieval request to obtain semantic granularity data; generate hierarchical semantic edges according to the semantic granularity data to obtain hierarchical semantic edge data; A hierarchical semantic graph construction module, configured to generate a hierarchical semantic graph for a preset authentic medicinal materials knowledge graph according to the hierarchical semantic edge data to obtain hierarchical semantic graph data; A semantic edge annotation module, configured to perform relational edge semantic annotation on the hierarchical semantic graph data to obtain semantic graph annotation data; An intelligent semantic-driven retrieval module, configured to perform multi-condition intention processing according to the data of the authentic medicinal materials retrieval request to obtain multi-condition intention data; perform multi-hop path combination on the semantic graph annotation data according to the multi-condition intention data to obtain preliminary retrieval data; perform edge reasoning and completion on the preliminary retrieval data to obtain the authentic medicinal materials retrieval data.
[0086] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended application documents rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be encompassed within the present invention.
[0087] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features invented herein.
Claims
1. An intelligent retrieval method for genuine medicinal materials based on a spectrum, characterized in that, The method includes: S1. Obtain the data of the retrieval request for authentic Chinese medicinal materials, perform semantic granularity processing on the data of the retrieval request for authentic Chinese medicinal materials to obtain semantic granularity data; generate hierarchical semantic edges according to the semantic granularity data to obtain hierarchical semantic edge data; S2. Generate a hierarchical semantic graph for the preset authentic Chinese medicinal materials knowledge graph according to the hierarchical semantic edge data to obtain hierarchical semantic graph data; S3. Perform semantic annotation on the relationship edges of the hierarchical semantic graph data to obtain semantic annotation data; S4. Perform multi-condition intention processing according to the data of the retrieval request for authentic Chinese medicinal materials to obtain multi-condition intention data; perform multi-hop path combination on the semantic annotation data according to the multi-condition intention data to obtain preliminary retrieval data; perform edge inference completion on the preliminary retrieval data to obtain the retrieval data of authentic Chinese medicinal materials.
2. The method according to claim 1, wherein S1 includes: Obtain the data of the retrieval request for authentic Chinese medicinal materials; Perform coarse semantic granularity processing on the data of the retrieval request for authentic Chinese medicinal materials to obtain coarse semantic granularity data; Perform fine semantic granularity processing on the data of the retrieval request for authentic Chinese medicinal materials to obtain fine semantic granularity data, where the semantic granularity processing is carried out according to the preset traditional Chinese medicine semantic ontology and granularity stratification rules, the coarse semantic granularity is the medicinal material category, regional attribute and basic efficacy, and the fine semantic granularity is the meridian tropism, drug nature, collection season and classical compatibility; Construct semantic paths and calculate multi-level edge weights for the coarse semantic granularity data and the fine semantic granularity data to obtain semantic path data and multi-level edge weight data; Generate a semantic hierarchical graph according to the semantic path data and the multi-level edge weight data to obtain multi-granularity semantic hierarchical graph data; Perform semantic edge classification annotation on the multi-granularity semantic hierarchical graph data to obtain hierarchical semantic edge data.
3. The method according to claim 1, characterized in that S2 Includes: Perform functional attribution semantic perception mapping and regional co-occurrence semantic perception mapping on the preset authentic Chinese medicinal materials knowledge graph according to the hierarchical semantic edge data to obtain functional attribution mapping data and regional co-occurrence mapping data respectively; Perform node twin expansion according to the functional attribution mapping data and the regional co-occurrence mapping data to obtain semantic completion data; Perform structure clustering according to the semantic incomplete data to obtain hierarchical semantic graph data.
4. The method according to claim 3, characterized in that, Among them, the functional attribution semantic perception mapping is specifically: Screen the functional semantic trigger edges according to the hierarchical semantic edge data to obtain functional semantic edge data; Construct a functional semantic path for the preset authentic Chinese medicinal materials knowledge graph according to the functional semantic edge data to obtain functional semantic path sub-graph data; Perform semantic path context verification according to the hierarchical semantic edge data and the functional semantic path sub-graph data to obtain functional context verification data; Perform semantic coverage conflict detection on the functional context verification data to obtain functional semantic coverage conflict data; Compare the mapping accuracy according to the functional semantic coverage conflict data to obtain functional attribution mapping data.
5. The method according to claim 3, characterized in that, Among them, the regional co-occurrence semantic perception mapping is specifically: Extract region-related trigger nodes according to the hierarchical semantic edge data to obtain region node data; Perform medicinal material region co-occurrence path mapping on the preset authentic Chinese medicinal materials knowledge graph according to the region node data to obtain medicinal material region co-occurrence data; Expand the semantic path of the medicinal material region co-occurrence data to obtain region co-occurrence candidate path data; Calculate the path co-occurrence weight for the regional co-occurrence candidate path data to obtain the path co-occurrence weight data; Generate the geographical co-occurrence mapping data based on the path co-occurrence weight data and the regional co-occurrence candidate path data to obtain the geographical co-occurrence mapping data.
6. The method according to claim 3, wherein Among them, the node twin expansion is specifically as follows: Extract structurally similar node data based on the functional attribution mapping data and the geographical co-occurrence mapping data; Calculate the similarity of graph nodes based on the structurally similar node data to obtain the graph node similarity data; Perform twin node path inversion based on the graph node similarity data to obtain the twin node path data; Calculate the path confidence based on the twin node path data to obtain the path confidence data; Generate semantic completion data for the twin node path data, the functional attribution mapping data, and the geographical co-occurrence mapping data based on the path confidence data, which is the semantic completion data.
7. The method according to claim 1, wherein S3 Include: Extract the edge context semantics data by extracting the node semantic nested context from the hierarchical semantic graph data; Generate the multi-source semantic feature tensors for the hierarchical semantic graph data based on the edge context semantics data to obtain the edge feature data to be annotated; Calculate the path probability based on the edge feature data to be annotated to obtain the path probability data; Perform functional analogy annotation on the edge feature data to be annotated based on the path probability data to obtain the preliminary semantic graph annotation data; Detect the edge semantic conflicts in the preliminary semantic graph annotation data to obtain the semantic conflict detection data; Perform relationship completion on the preliminary semantic graph annotation data based on the semantic conflict detection data to obtain the semantic graph annotation data.
8. The method according to claim 1, characterized in that S4 Include: Extract the retrieval request feature data according to the genuine regional drug retrieval request data, where the retrieval request feature extraction includes geographical preference feature extraction, meridian tropism feature extraction, and medicinal property restriction feature extraction; Calculate the intention attention for the retrieval request feature data to obtain the multi-condition intention data; Construct the intention path candidate space for the semantic graph annotation data based on the multi-condition intention data to obtain the intention path candidate space data; Perform graph traversal based on the intention path candidate space data to obtain the path candidate graph data; Calculate the semantic coverage rate based on the path candidate graph data and the retrieval request feature data to obtain the semantic coverage rate data; Filter the path candidate graph data based on the semantic coverage rate data to obtain the path filtering data; Perform edge node semantic filling processing on the path filtering data to obtain the genuine regional drug retrieval data.
9. The method according to claim 8, characterized in that Among them, the edge node semantic filling processing is specifically as follows: Analyze the boundary nodes of the path filtering data to obtain the half-edge connected node data; Generate edge node missing data based on the half-edge connected node data; Map the semantic path missing type based on the edge node missing data to obtain the semantic path missing type data; Perform attribute near neighbor matching based on the semantic path missing type data to obtain the supplementary node data; Perform similarity filtering on the supplementary node data based on the retrieval request feature data to obtain the supplementary filtering data; Integrate the supplementary filtering data and the path filtering data to obtain the genuine regional drug retrieval data.
10. An intelligent retrieval system for authentic Chinese medicinal materials based on a spectrum, characterized in that, For implementing the atlas-based intelligent retrieval method of authentic Chinese medicinal materials as described in claim 1, the atlas-based intelligent retrieval system of authentic Chinese medicinal materials includes: A semantic granularity analysis module, configured to obtain the retrieval request data of authentic Chinese medicinal materials, perform semantic granularity processing on the retrieval request data of authentic Chinese medicinal materials to obtain semantic granularity data; generate hierarchical semantic edges according to the semantic granularity data to obtain hierarchical semantic edge data; A hierarchical semantic graph construction module, configured to generate a hierarchical semantic graph for a preset knowledge graph of authentic Chinese medicinal materials according to the hierarchical semantic edge data to obtain hierarchical semantic graph data; A semantic edge annotation module, configured to perform relational edge semantic annotation on the hierarchical semantic graph data to obtain semantic graph annotation data; An intelligent semantic-driven retrieval module, configured to perform multi-condition intention processing on the retrieval request data of authentic Chinese medicinal materials to obtain multi-condition intention data; perform multi-hop path combination on the semantic graph annotation data according to the multi-condition intention data to obtain preliminary retrieval data; perform edge reasoning and complementation on the preliminary retrieval data to obtain the retrieval data of authentic Chinese medicinal materials.
Citation Information
Patent Citations
Short video recommendation method, system and equipment
CN114117203A
Intelligent vibratory digital twinning system and method for industrial environments
CN115039045A
Potential relation reasoning-based medical knowledge graph retrieval system and method
CN119739867A
System and / or method for an autonomous linked managed semantic model based knowledge graph generation framework
US20220121964A1
Knowledge graph implementation
US20240135391A1
Cited By
Security code automatic generation method and system based on full information theory and mechanism ambiguity
CN120447881A
Method and system for automatic generation of secure code based on full information theory and mechanism theory
CN120447881B
Neurological disease rehabilitation period medication intelligent management system and digital prescription generation method
CN120564952A
Intelligent agent high-order relation modeling method based on side attention weight
CN120611644A
Session recommendation method based on mixed intention and double constraints
CN120705278A