A data retrieval system and method for Chinese medicinal materials data
By constructing semantically annotated enhanced data and an extended knowledge graph for Chinese medicinal materials, and using logical rule matching and implicit relationship reasoning to generate a standardized graph query language, the problems of single data structure and insufficient understanding of user query intentions in the Chinese medicinal materials data management system are solved, and efficient multi-dimensional medicinal materials knowledge retrieval and accurate response are achieved.
Patent Information
- Application Number
- CN202510499353.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The existing Chinese medicinal materials data management system has a single data structure and weak semantic understanding ability, making it difficult to support in-depth mining and intelligent retrieval of complex medicinal materials knowledge. It also lacks a deep semantic understanding and precise matching mechanism for user query intentions.
By constructing semantically annotated enhanced data and extended knowledge graphs for Chinese medicinal materials, adopting logical rule matching and implicit relationship reasoning, a standardized graph query language is generated to realize graph structure retrieval. Combining vector semantic encoding with the set of Chinese medicinal materials semantic vectors, multi-dimensional expression and precise matching are performed. Finally, through the fusion and comparison of graph query results and semantic retrieval results, the retrieval accuracy and context consistency are improved.
It significantly improves the organizational efficiency and semantic expression capabilities of Chinese medicinal materials data, enhances the system's response to multi-dimensional medicinal material relationships, realizes multi-dimensional expression and precise matching of user search intentions, solves the problem of poor relevance of search results in traditional systems, and meets the needs of modern medical research and development and digital research of Chinese medicine.
Smart Images

Figure CN120045744B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a data retrieval and management method for Chinese herbal medicine data. Background Art
[0002] In the field of traditional Chinese medicinal materials, data management mainly relies on manual compilation and static document records, which cannot support the efficient organization and retrieval of data required by modern pharmaceutical research and development and precision medicine. With the development of informatization, Chinese medicinal materials management systems based on relational databases have been gradually introduced, enabling electronic storage and simple retrieval of basic medicinal materials properties, classifications, and sources. However, such systems generally suffer from simple data structures, weak semantic understanding capabilities, and a lack of knowledge association modeling, making it difficult to support in-depth mining and intelligent retrieval of complex medicinal material knowledge. In recent years, some studies have attempted to introduce semantic modeling methods to improve the organization and expression capabilities of Chinese medicinal materials data. However, existing methods still have many shortcomings: on the one hand, semantic construction mostly relies on manual rules, resulting in low system intelligence and poor scalability; on the other hand, they lack deep semantic understanding of user query intent and precise matching mechanisms, making it difficult to achieve retrieval responses in complex contexts. Summary of the Invention
[0003] Based on this, it is necessary for the present invention to provide a data retrieval system and method for Chinese medicinal materials data management to solve at least one of the above technical problems.
[0004] To achieve the above-mentioned purpose, a data retrieval method for Chinese herbal medicine data management includes the following steps:
[0005] Step S1: Acquire Chinese medicinal material sample morphological data; perform grade assessment based on the Chinese medicinal material morphological data to generate Chinese medicinal material grade label data; acquire images of the Chinese medicinal material to be tested; preprocess the images of the Chinese medicinal material to be tested, and perform entity recognition and relationship extraction to obtain Chinese medicinal material semantic annotation data; fuse the Chinese medicinal material grade label data and the Chinese medicinal material semantic annotation data to obtain Chinese medicinal material semantic annotation enhanced data;
[0006] Step S2: Using the semantic annotation enhancement data of Chinese medicinal materials to build an initial knowledge graph, perform logical rule matching and implicit relationship reasoning on the initial knowledge graph to obtain an extended knowledge graph;
[0007] Step S3: Map the extended knowledge graph to the graph database structure and generate standardized graph query language data;
[0008] Step S4: constructing a graph query based on the graph query language data by combining conditions, traversing paths, and linking nodes to obtain graph structure query result data;
[0009] Step S5: Obtain user search request data; perform semantic vector encoding on the user search request data to obtain user query semantic vector data; construct a set of Chinese medicinal material description semantic vectors based on the Chinese medicinal material semantic annotation enhanced data, and perform matching calculation on the user query semantic vector data to obtain vector semantic search result data;
[0010] Step S6: Fusing and comparing the graph structure query result data and the vector semantic retrieval result data to obtain the final Chinese medicinal material semantic retrieval response data.
[0011] The present invention achieves deep fusion and semantic structured expression of Chinese medicinal material images, grades, and text information by constructing semantically annotated enhanced data and an extended knowledge graph for Chinese medicinal materials, significantly improving the organizational efficiency and semantic expression ability of Chinese medicinal material data. Based on the graph, logical rule matching and implicit relationship reasoning are adopted to establish a knowledge network with linkage reasoning ability, breaking through the limitations of static associations in traditional databases. Through graph database mapping and standardized graph query language generation, a graph structure retrieval mechanism supporting complex query modes such as path traversal, condition combination, and node linkage is constructed, enhancing the system's responsiveness to multidimensional medicinal material relationships. The fusion of vector semantic coding and Chinese medicinal material semantic vector set construction achieves multidimensional expression and precise matching of user search intent, solving the problems of traditional systems' insufficient understanding of user input and poor relevance of search results. Finally, through the fusion comparison of graph query results and semantic search results, taking into account both structural relationships and semantic similarity, the accuracy and contextual consistency of search results are improved, so that the Chinese medicinal material data management system not only has highly intelligent data organization capabilities, but also supports accurate response to complex semantic problems and deep semantic retrieval, meeting the core needs of modern pharmaceutical research and development and digital research on Chinese medicine.
[0012] Preferably, the present invention further provides a data retrieval Chinese medicinal material data management system for executing the above-mentioned data retrieval Chinese medicinal material data management method, wherein the data retrieval Chinese medicinal material data management system comprises:
[0013] The Chinese medicine information semantic annotation module is used to obtain the morphological data of Chinese medicinal materials; perform grade assessment based on the Chinese medicinal materials morphological data to generate Chinese medicinal materials grade label data; obtain images of the Chinese medicinal materials to be tested; pre-process the images of the Chinese medicinal materials to be tested, and perform entity recognition and relationship extraction to obtain Chinese medicinal materials semantic annotation data; and fuse the Chinese medicinal materials grade label data and the Chinese medicinal materials semantic annotation data to obtain Chinese medicinal materials semantic annotation enhanced data.
[0014] The medicinal material knowledge graph construction and reasoning module is used to use the semantic annotation of traditional Chinese medicine to enhance data to build an initial knowledge graph, perform logical rule matching and implicit relationship reasoning on the initial knowledge graph, and obtain an extended knowledge graph;
[0015] The graph database mapping module is used to map the extended knowledge graph to the graph database structure and generate standardized graph query language data;
[0016] The graph structure query module is used to construct graph queries based on graph query language data by combining conditions, traversing paths, and linking nodes to obtain graph structure query result data;
[0017] The user semantic matching retrieval module is used to obtain user search request data; perform semantic vector encoding on the user search request data to obtain user query semantic vector data; construct a set of Chinese medicinal material description semantic vectors based on the Chinese medicinal material semantic annotation enhanced data, and perform matching calculation on the user query semantic vector data to obtain vector semantic retrieval result data;
[0018] The retrieval result fusion response module is used to fuse and compare the graph structure query result data with the vector semantic retrieval result data to obtain the final Chinese medicinal material semantic retrieval response data.
[0019] Through the coordinated operation of various functional modules, the present invention has built a complete closed-loop data management mechanism in key links such as Chinese medicinal materials data collection, semantic expression, knowledge association, and intelligent retrieval, which has significantly improved the structuralization, standardization, and semantic processing capabilities of Chinese medicinal materials information; in the data acquisition and processing stage, it has realized the multimodal fusion processing of Chinese medicinal materials morphological characteristics, grade labels and image semantic information, and improved the recognition accuracy of original data and the accuracy of label expression; in the knowledge modeling stage, the structured modeling and logical extension between semantic entities, attributes and relationships are realized through the graph construction and reasoning mechanism, which enhances the system's organization and reasoning capabilities for the complex knowledge system of Chinese medicinal materials; through the graph database mapping With graph query construction, the system has the ability to efficiently operate graph structure data, supports flexible multi-dimensional condition combinations and associated path discovery, and provides a graph structure-level response basis for professional problems; at the user interaction layer, based on semantic vectorization representation and matching mechanism, it significantly enhances the system's ability to identify, understand and retrieve user intentions, and realizes deep semantic expression and multi-angle matching of query requests; finally, through the joint comparison and fusion ranking of graph structure results and semantic vector results, it realizes the dual guarantee of structural logical relevance and semantic relevance, improves the system's retrieval accuracy, result relevance and response intelligence level in complex contexts, and fully meets the practical needs of Chinese medicinal materials knowledge management and intelligent services. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments thereof made with reference to the following drawings:
[0021] Figure 1 A schematic flow chart of the steps of a data retrieval method for managing Chinese medicinal materials data according to the present invention;
[0022] Figure 2 for Figure 1 Detailed step flow diagram of step S1;
[0023] Figure 3 for Figure 1 Detailed step flow chart of step S2 in FIG. DETAILED DESCRIPTION
[0024] The following is a clear and complete description of the technical method of the present invention in conjunction with the accompanying drawings. It is obvious that the embodiments described are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of the present invention.
[0025] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor and / or microcontroller approaches.
[0026] It should be understood that although the terms "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. The term "and / or" as used herein includes any and all combinations of one or more of the listed associated items.
[0027] To achieve this, please refer to Figures 1 to 3 The present invention provides a data retrieval method for managing Chinese medicinal materials data, the method comprising the following steps:
[0028] Step S1: Acquire Chinese medicinal material sample morphological data; perform grade assessment based on the Chinese medicinal material morphological data to generate Chinese medicinal material grade label data; acquire images of the Chinese medicinal material to be tested; preprocess the images of the Chinese medicinal material to be tested, and perform entity recognition and relationship extraction to obtain Chinese medicinal material semantic annotation data; fuse the Chinese medicinal material grade label data and the Chinese medicinal material semantic annotation data to obtain Chinese medicinal material semantic annotation enhanced data;
[0029] In an embodiment of the present invention, a microscopic camera is used to collect surface images of each batch of Chinese medicinal material samples, and the image resolution is set to no less than 4096×2160 pixels; at the same time, a hyperspectral imager with a wavelength coverage range of 400nm to 1000nm is used to obtain fiber structure images and surface texture maps, and the image sampling frequency is set to 20 pixels per millimeter, and the obtained image data are integrated to form the morphological data of the Chinese medicinal material samples; then, a texture analyzer is used to collect component characteristic data of the Chinese medicinal materials at a compression speed of 1mm / s, and the collected data include hardness, elasticity, adhesion and cohesion, and the units are N, mm and g. The morphological data and the component characteristic data are fused through a feature cascade algorithm, and a redundant removal method based on principal component analysis (PCA) is adopted to retain the principal components with a cumulative variance contribution rate of no less than 95%, and to construct a Chinese medicinal material morphology-component joint feature matrix; then, according to the manually preset classification rules, the rules include grade judgment based on indicators such as fiber density, surface texture clarity, and texture strength, and each row in the joint feature matrix is input into the rule logic judgment process, and the results are analyzed according to the results of the analysis. The Chinese medicinal materials grade label data is generated according to the grade standard, with the grade classification accuracy of not less than grade 5. Then, a high-definition imaging device fixedly installed on the conveyor belt is used to obtain images of the Chinese medicinal materials to be tested. The image size is 1920×1080 pixels, and the image acquisition rate is not less than 30 frames per second. The obtained image is input into the image preprocessing module, and operations including image denoising (Gaussian filter with a kernel size of 5×5), edge enhancement (Sobel operator), and standardized grayscale transformation (normalized range [0,1]) are performed. Subsequently, an image entity recognition algorithm based on morphological feature matching is used to identify the boundaries of the entity areas in the Chinese medicinal materials images, with the recognition accuracy limited to not less than 92%. Then, a relationship extraction algorithm based on dictionary matching and co-occurrence frequency analysis is used to construct a set of attribute-relationship-value triples between entities to form Chinese medicinal materials semantic annotation data. Finally, the Chinese medicinal materials grade label data generated above are aligned with the Chinese medicinal materials semantic annotation data based on the sample unique number, and the dimensions are fused by the feature connection method to construct an enhanced semantic feature vector set. This fused data is named Chinese medicinal materials semantic annotation enhanced data.
[0030] Step S2: Using the semantic annotation enhancement data of Chinese medicinal materials to build an initial knowledge graph, perform logical rule matching and implicit relationship reasoning on the initial knowledge graph to obtain an extended knowledge graph;
[0031] In the embodiment of the present invention, the semantic annotation enhanced data of Chinese medicinal materials is used as input, and the field parsing algorithm is first used to perform field segmentation and attribute recognition on the text data in the semantic annotation enhanced data. The field segmentation adopts a linear traversal method based on character boundary position matching, and the attribute recognition rules include fixed keyword label recognition and part-of-speech dependency structure judgment. The field parsing result is stored in a field-value pair structure; then, the entity classification process is executed, and the field parsing result is input into the entity classification dictionary table. The dictionary table is constructed with a standard Chinese medicinal material terminology set, and contains entity types such as "medicinal material name", "origin", "morphological characteristics", "grade label", "processing method", etc. The classification process adopts regular expression matching and hash dictionary indexing. The graph node candidate data is generated by combining the two methods; then, semantic relationship mapping between nodes is performed based on the graph node candidate data. The mapping rules are based on the attribute co-occurrence frequency, semantic dictionary mapping pairs and hyponym judgment. The defined edge types include "belong to", "include", "derive from", "have characteristics" and other relationship types. Each type of edge is assigned a direction and a unique identification code. The generated relationship data is constructed into structured knowledge graph data; after the graph is constructed, the data deduplication process is executed, and the uniqueness of nodes and edges is judged by the hash check code (the length is set to 32 bits), and duplicate nodes and redundant edges are deleted; then semantic consistency processing is performed, and the method is word vector cosine similarity calculation (the threshold is set to 0 .9) Combined filtering with semantic merge rule judgment (based on attribute value similarity greater than 0.85) to obtain semantically consistent graph data; then map the structured graph data to a preset standard graph model. The standard model is built based on the Neo4j graph structure template. During mapping, the node attribute format is unified as a "key-value" pair structure, and the edge attribute structure is a triple format; finally, the initial knowledge graph is generated; based on the initial knowledge graph, the graph structure is traversed and the rule node filtering operations are performed. The graph traversal uses the depth-first traversal (DFS) method. The node filtering conditions include node type matching of not less than 95%, edge weight greater than 0.6, and semantic coverage greater than 80%; then, the logical inference is called. The rule set is stored as rule entries. Each rule contains a premise node, a logical relationship, and a result node triple. The symbolic rule matching process is executed, and the inference path is triggered by conditional judgment to generate rule inference path data. The rule inference path data is input into the graph embedding processing module, which initializes the embedding vectors of the nodes in the path (the vector dimension is set to 128). The vector semantics is expanded based on the path sequential propagation mechanism, and the attention mechanism weights are updated. The initial value of the attention weight is set to a mean of 0.1 and a maximum value of 1.0, resulting in the graph inference embedding data. The embedded data is then clustered based on similarity using a vector clustering method based on cosine similarity, with the clustering similarity threshold set to 0.92, and performs upstream and downstream path prediction on the clustering results, limiting the predicted path to no more than three hops, and outputs the inference relationship generation data. Finally, the initial knowledge graph and the inference relationship generation data are structurally merged. The merging process adopts a two-layer structural integration method of node alignment + edge merging. After the merging is completed, the weight value of each edge in the new graph is reassigned. The edge weight is calculated based on the semantic strength and path frequency. The formula is weight value = 0.6 × semantic similarity + 0.4 × path occurrence frequency normalization value, and the expanded knowledge graph data is finally obtained.
[0032] Step S3: Map the extended knowledge graph to the graph database structure and generate standardized graph query language data;
[0033] In an embodiment of the present invention, an extended knowledge graph is used as input data, and a graph database structure mapping module is used to convert the extended knowledge graph into graph query language data. The graph database structure mapping module is deployed in a multi-core parallel computing node, and the graph database used is Neo4j Community Edition 5.12, which supports graph data interaction based on the Cypher language. During the mapping process, the node type identification submodule is first called to perform a structural scan on the entity nodes in the knowledge graph. The identified node types include four categories: "medicinal material node", "function node", "meridian node" and "disease node". Each type of node is accompanied by an attribute field, where the text field uniformly adopts the UTF-8 encoding format, and the maximum length is limited to 128 characters. Then, the relationship edge in the triple structure is extracted through the edge attribute parsing submodule, and the edge attributes are organized in a key-value pair format. The edge weight field is limited to a floating point type, and the precision is controlled to three decimal places. After structure recognition, the hierarchical data of the graph structure is written into the graph database through the graph database structure dynamic loading module. The loading process enables the memory cache mechanism, and the cache size is limited to 30% of the total physical memory. The cache replacement strategy adopts the least recently used (LRU) algorithm. The graph database schema definition is constructed based on the Schema-first method, and entity labels, attribute constraints and relationship types are defined through Cypher statements. After writing is completed, the syntax reconstruction sub-module performs path reachability verification and relationship constraint analysis, and the path expressions that do not conform to the grammatical rules are structurally adjusted. The connector between nodes in the path expression is uniformly used as "--", and the attribute filtering conditions are uniformly expressed using the "WHERE" clause. All expressions must be verified for grammatical compliance through a regular matching function before they are finally converted into graph query language data. The final output standardized graph query language data is encapsulated in JSON format and stored in a high-speed read-write storage module.
[0034] Step S4: constructing a graph query based on the graph query language data by combining conditions, traversing paths, and linking nodes to obtain graph structure query result data;
[0035] In an embodiment of the present invention, a graph query construction operation is performed based on standardized graph query language data. First, a semantic analysis is performed on the attribute screening clauses in the standardized graph query language data through a condition combination module, and a Boolean operation priority parsing algorithm is used to perform left-associative priority merging on the "AND", "OR", and "NOT" logical symbols to form query condition template data. The length limit of each field value in the query condition is not more than 64 characters, and the field types include string type, Boolean type, and enumeration type. Then, a semantic path expansion process is performed on the query condition template data through a path traversal module. The path expansion adopts a breadth-first search (BFS) algorithm, which limits the search level to no more than 5 layers to avoid the problem of path combination explosion. The node and edge data used in the search are derived from the graph database mapping data loaded into the Neo4j graph database in step S3, and are expressed through the graph database native Cypher syntax. The starting and ending nodes of each path are verified, and only path data that meets the grammatical conditions and has consistent edge directions is retained; then the node linkage module is called to perform linkage calculation on the node set in the valid path set. This process is based on the attribute matching degree of the nodes, and the calculation method adopts the weighted cosine similarity method. The linkage strength score ranges from 0 to 1, and node pairs with a score lower than 0.3 are excluded from the current path combination; after completing the node linkage, a graph structure query expression is formed and submitted to the Neo4j graph database for query operation through the query expression execution engine, and the query result node set, path set and edge set data are returned. During the query operation execution process, four threads are used to submit tasks concurrently. The maximum execution time of each query statement does not exceed 3000 milliseconds. If it exceeds, it will be forced to interrupt and recorded in the exception log. All query results are encapsulated into graph structure query result data in a unified format through structured encapsulation components.
[0036] Step S5: Obtain user search request data; perform semantic vector encoding on the user search request data to obtain user query semantic vector data; construct a set of Chinese medicinal material description semantic vectors based on the Chinese medicinal material semantic annotation data, and perform matching calculation on the user query semantic vector data to obtain vector semantic search result data;
[0037] In the embodiment of the present invention, first, the user search request data is obtained through the semantic encoding module. The source of the search request data is the form field of the standard HTTP POST request, the encoding format is UTF-8, and the total length of the field does not exceed 512 characters; then, the user search request data is segmented by a forward maximum matching word segmentation algorithm based on dictionary mapping. The Chinese herbal medicine professional terminology dictionary used contains more than 10,000 entries, and the length of each entry does not exceed 10 Chinese characters. All stop words are excluded in the word segmentation process, and the stop word list is limited to 300 commonly used function words; after the word segmentation is completed, the user search request data is semantically encoded by constructing a word vector. The encoding method adopts a bag-of-words model combined with TF-IDF weights, sets the word frequency threshold to 1, the inverse document frequency lower limit to 0.1, and the dimension generated after encoding is 300. In parallel, a Chinese medicinal material description semantic vector set is constructed based on the Chinese medicinal material semantic annotation data constructed in step S1. The set construction process is based on the standard description fields of each Chinese medicinal material, such as "properties", "functions and indications", "sources" and "meridians". TF-IDF feature extraction is performed, and the extraction dimension remains consistent at 300 dimensions. The total length of characters in the description field is limited to no more than 2048 characters. After completing the construction of the above two vector sets, the vector matching module is called to perform matching calculations. The matching method adopts the Euclidean distance method, and the distance threshold is limited to 2.5. The distance less than the threshold is considered a successful match. The data are sorted from small to large by distance, and the top 10 are taken to constitute the vector semantic retrieval result data.
[0038] Step S6: Fusing and comparing the graph structure query result data and the vector semantic retrieval result data to obtain the final Chinese medicinal material semantic retrieval response data.
[0039] In the embodiment of the present invention, the graph structure query result data and the vector semantic retrieval result data are loaded into independent data structures respectively. The graph structure query result data is stored in a relational table structure, and each record contains an entity ID, an entity name, a node attribute set, and a path identifier; the vector semantic retrieval result data is organized in a dense vector array structure, and each vector dimension is unified as 192 dimensions, and is represented in the floating point format float32. The information retention rate reaches 92% to ensure the accuracy of the expression semantics; in the implementation process, first, a hash index mapping is performed on the entity ID fields in the two data structures to establish a one-to-one corresponding entity index mapping table, and then an entity alignment operation is performed based on the mapping table; after the alignment is completed, attribute vectorization processing is performed on each pair of corresponding entities, and the attribute field adopts the Term Frequency-Inverse Document Frequency (TF-IDF) weighted average method is used to convert it into an attribute vector, with the dimension set to 64 dimensions and the value range normalized to the interval [0,1]. After the attribute vector is prepared, the cosine similarity is calculated with the corresponding semantic vector, and the similarity result is set as the semantic similarity label. Based on the semantic similarity label and the graph structure weight parameter of the node, a weighted fusion calculation is performed, and the fusion formula is: final score = 0.6 × graph structure weight score + 0.4 × semantic similarity label score, where the score range is limited to [0,1]. The fusion score is then sorted in descending order and filtered with a confidence threshold of 0.7. The results above the threshold are retained to form the semantic sorting result data of Chinese medicinal materials. Next, the Chinese medicinal materials in the sorting results are sorted. The medicinal material node performs attribute information extraction, association path analysis and similar node reasoning. The path analysis adopts the shortest path search based on the Dijkstra algorithm. The node similarity is determined based on the weighted calculation of the attribute intersection ratio and the structural proximity, and the similarity threshold is set to 0.65. The above-mentioned extracted information is encapsulated in a structured data format, and the fields include node ID, main attributes, path sequence, similar node ID set and similarity score. Finally, the structured data is encoded in a unified JSON format and encapsulated into a response data structure through a predefined interface template. At the same time, 10% of the total memory space is allocated in the cache module for response data caching. The cache adopts the FIFO strategy for elimination control to obtain the final Chinese medicinal material semantic retrieval response data.
[0040] The present invention achieves deep fusion and semantic structured expression of Chinese medicinal material images, grades, and text information by constructing semantically annotated enhanced data and an extended knowledge graph for Chinese medicinal materials, significantly improving the organizational efficiency and semantic expression ability of Chinese medicinal material data. Based on the graph, logical rule matching and implicit relationship reasoning are adopted to establish a knowledge network with linkage reasoning ability, breaking through the limitations of static associations in traditional databases. Through graph database mapping and standardized graph query language generation, a graph structure retrieval mechanism supporting complex query modes such as path traversal, condition combination, and node linkage is constructed, enhancing the system's responsiveness to multidimensional medicinal material relationships. The fusion of vector semantic coding and Chinese medicinal material semantic vector set construction achieves multidimensional expression and precise matching of user search intent, solving the problems of traditional systems' insufficient understanding of user input and poor relevance of search results. Finally, through the fusion comparison of graph query results and semantic search results, taking into account both structural relationships and semantic similarity, the accuracy and contextual consistency of search results are improved, so that the Chinese medicinal material data management system not only has highly intelligent data organization capabilities, but also supports accurate response to complex semantic problems and deep semantic retrieval, meeting the core needs of modern pharmaceutical research and development and digital research on Chinese medicine.
[0041] Preferably, step S1 includes the following steps:
[0042] Step S11: collecting appearance images of the Chinese medicinal materials through microscopic photography, collecting fiber structure images and surface texture atlas data through high-definition multispectral imaging, and recording the appearance images, fiber structure images and surface texture atlas data as Chinese medicinal materials sample morphology data; collecting intrinsic component characteristic data of the Chinese medicinal materials through texture analysis;
[0043] Step S12: performing feature fusion on the Chinese medicinal material sample morphology data and the Chinese medicinal material internal component feature data, and performing redundancy removal to construct a Chinese medicinal material morphology-component joint feature matrix;
[0044] Step S13: Classifying the Chinese medicinal materials using classification rules based on the Chinese medicinal materials morphology-ingredient joint feature matrix to generate Chinese medicinal materials grade label data;
[0045] Step S14: obtaining a handwritten image of Chinese medicinal materials; performing image preprocessing, image enhancement, and text recognition on the handwritten image of Chinese medicinal materials to obtain structured Chinese medicinal materials data;
[0046] Step S15: performing expected recognition and content structured annotation on the structured Chinese medicinal material data to obtain Chinese medicinal material text parsing data;
[0047] Step S16: performing named entity recognition and term boundary segmentation on the Chinese herbal medicine text parsing data to obtain Chinese herbal medicine semantic entity data;
[0048] Step S17: performing context-aware analysis based on the Chinese medicinal material semantic entity data to obtain Chinese medicinal material semantic relationship data;
[0049] Step S18: Merging the data structure of the Chinese medicinal material semantic entity data and the Chinese medicinal material semantic relationship data, and uploading them to the edge server to obtain the Chinese medicinal material semantic annotation data;
[0050] Step S19: Fusing the Chinese medicinal material grade label data and the Chinese medicinal material semantic annotation data to obtain Chinese medicinal material semantic annotation enhanced data.
[0051] In the embodiment of the present invention, an industrial-grade microscopic camera with a resolution of not less than 0.5 microns is used to perform tiled image acquisition on a Chinese medicinal material sample. The image acquisition light source adopts an LED cold light source, and the brightness is constantly controlled at 4000 lux. The output image is saved in an uncompressed TIFF format with a size of 4096×2160 pixels. At the same time, a hyperspectral imager with a wavelength range of 400nm to 1000nm and a spectral resolution of not more than 5nm is used to perform multi-angle scanning on the Chinese medicinal material sample to obtain fiber structure images and surface texture atlas data. The image acquisition angle is 45° incident light, and the scanning speed is set to 10mm / s. The three types of image data are named and uniformly classified as Chinese medicinal material sample morphology data. Further, a texture analyzer (Texture Analyzer) uses an indenter with a diameter of 5mm and a test speed of 1mm / s to collect parameters such as hardness, elasticity and adhesion of Chinese medicinal materials. Each indicator is sampled no less than 10 times, and the average value is extracted as the characteristic data of the intrinsic components of the Chinese medicinal materials; the above-mentioned morphological data and component data are feature fused by feature cascade. The fusion method is feature cascade. First, texture features are extracted from image data. Then, the texture data is standardized and spliced to form a unified numerical vector. The obtained data is input into the redundancy removal algorithm. The principal component analysis method (PCA) is used to retain the principal components with a cumulative variance contribution rate of no less than 95%, and finally a Chinese medicinal material morphology-component joint feature matrix with a dimension of no more than 64 is formed; according to each record in the joint feature matrix, a set of manually formulated classification rules is applied in sequence. The rules include texture clarity distribution threshold, elasticity range definition, and texture periodicity judgment. The number of rules is no less than 15, and each rule corresponds to a first-level grade division logic. Finally, the Chinese medicinal material grade label data is generated. The labels are "L1" to "L5" to represent 5 grade levels; paper records of Chinese medicinal materials are collected, Handwritten images were acquired using a high-precision image acquisition device with a resolution of 300 dpi. The images were then fed into an image preprocessing module, where they were subjected to denoising (using a median filter with a window size of 3×3), image enhancement (stretching the contrast range to [0, 255]), and binarization (using the OTSU thresholding method). After image enhancement, the images were fed into a text recognition engine. The recognition method was based on character outline extraction and template matching. The extracted text structure was archived by row and column position to obtain structured Chinese medicinal material data. The structured Chinese medicinal material data was segmented into semantic units using a sentence recognition method based on a regular rule set. The number of rules was no less than 30, covering content types such as origin, properties, and processing methods. The segmented units were marked as data fields and embedded with content tags. The output was Chinese medicinal material text parsing data. Named entity recognition was performed on the text parsing data. Entity localization methods based on dictionary matching were used to locate entity boundaries. Term boundary segmentation was completed through contextual part-of-speech consistency judgment. The segmented unit length was limited to between 2 and 10 characters to generate Chinese medicinal material semantic entity data.Context-aware analysis is performed based on a context window mechanism. The left and right extended window widths of each entity are set to no more than five words. The co-occurrence probability, dependency grammatical structure, and semantic direction relationships between words are analyzed. Contextual logical dependencies are extracted to form a semantic relationship set, generating Chinese medicinal material semantic relationship data. Semantic entity data and semantic relationship data are structurally merged based on the entity's unique identifier. The merging method is to construct a triple structure centered on the main entity. The structure includes a "main entity-relationship-slave entity" format. After the structural data is completed, it is uploaded to the edge server via a 5G edge gateway device. The server configuration requirements are a CPU main frequency of no less than 2.6GHz and a memory capacity of no less than 64GB. The output data is the Chinese medicinal material semantic annotation data. Based on the Chinese medicinal material sample number and collection timestamp, the Chinese medicinal material grade label data and the Chinese medicinal material semantic annotation data are mapped one-to-one. The two types of data are dimensionally fused using feature field splicing to generate Chinese medicinal material semantic annotation enhanced data and write it to the local cache database.
[0052] The present invention realizes the synchronous collection and coordinated expression of the morphological characteristics and intrinsic components of Chinese medicinal materials by constructing a collection and fusion mechanism of multi-source heterogeneous data, thereby improving the scientificity and objectivity of the evaluation of the grade of Chinese medicinal materials; introduces microscopic imaging and multispectral imaging methods to realize the fine-grained collection of surface structure, fiber characteristics and texture maps, and combines the quantitative processing of texture parameters to effectively support the standardization and automation of grade classification; through image preprocessing and handwriting recognition processes, it further supplements historical records and manual annotation information, and enhances the source breadth and expression dimension of structured data; in the text processing chain, it completes the process from The complete process from image recognition to structured annotation, and then to semantic entity extraction and contextual relationship construction, strengthens the semantic organization ability of original text information; in the semantic data aggregation stage, it is merged and uploaded through the edge computing platform, which optimizes data processing efficiency and deployment flexibility; finally, through the fusion processing of grade labels and semantic annotation results, the unified expression of quantitative grade information and semantic description information is achieved, providing a unified, standardized, high-dimensional semantic input data foundation for subsequent knowledge modeling and retrieval calculations, and effectively supporting the system's ability to build semantic modeling, knowledge linkage and intelligent recognition of Chinese medicinal materials throughout the entire process.
[0053] Preferably, step S2 includes the following steps:
[0054] Step S21: performing field parsing and entity classification on the Chinese medicinal material semantic annotation enhanced data to obtain candidate data of Chinese medicinal material graph nodes;
[0055] Step S22: Based on the candidate data of the Chinese medicinal material graph nodes, semantic relationship mapping and edge type definition are performed between nodes to obtain structured knowledge graph data;
[0056] Step S23: Deduplication, semantic consistency processing and standard graph model mapping are performed on the structured knowledge graph data to obtain an initial knowledge graph;
[0057] Step S24: traverse the graph structure and filter the rule nodes of the initial knowledge graph, and use the preset logical reasoning rule set to perform symbolic rule matching to obtain rule reasoning path data;
[0058] Step S25: Based on the rule-based reasoning path data, perform semantic propagation of the graph node embedding vector and update the attention mechanism weight to obtain graph reasoning embedding data;
[0059] Step S26: perform similarity clustering on the graph reasoning embedding data and perform upstream and downstream path prediction to obtain reasoning relationship generation data;
[0060] Step S27: Structurally merge the initial knowledge graph and the inference relationship generation data, and redistribute the weights to obtain extended knowledge graph data.
[0061] In the embodiment of the present invention, a field parsing operation is performed on the semantic annotation enhanced data of Chinese medicinal materials, and a field extraction rule based on attribute name matching is adopted to group and store the entity information in the semantic annotation data according to the four fields of "name", "category", "attribute value" and "position code", and classify the entities according to the categories to which the entity attributes belong. The categories include "medicinal material entity", "efficacy entity", "indication entity", "meridian entity" and "application site entity". The classification standard is determined by the field "belonging category" in the semantic annotation data. After the classification is completed, candidate data of Chinese medicinal material atlas nodes are generated, and each candidate node is recorded in the node registration table with a unique number; based on the atlas node candidate data, Semantic relationship mapping between row nodes. The relationship mapping logic adopts the co-occurrence position relationship matching method of entity pairs in the original text. By setting the maximum window span to 120 characters, it is judged whether there is a valid semantic connection between the node pairs. If so, the edge type is defined according to its category and semantic content. The edge type includes four fixed semantic relationships: "having efficacy", "used for treatment", "belonging to the meridians", and "acting on the part". The starting node, end node, relationship type and semantic score of each edge are recorded to form structured knowledge graph data; data deduplication operation is performed on the structured knowledge graph data. The deduplication strategy is that if there is a duplicate relationship for the same node pair, the one with the highest semantic score is retained to ensure semantic consistency. The processing is based on the preset synonym normalization table and relationship category redirection table. The standard graph model is organized in the form of Triplet "entity A-relationship-entity B". The semantic score of each triple is standardized to the interval [0,1] to form an initial knowledge graph; a structural traversal operation is performed on the initial knowledge graph, and a breadth-first search algorithm is used to explore the graph hierarchy. The exploration depth is limited to 5 layers, and rule node filtering is performed in combination with node label rules. The filtering conditions are nodes with edge connectivity less than 2 or whose categories are marked as "non-main entities". After completing the structural filtering, the preset logical reasoning rule set is used to perform symbolic rule matching on the retained path. The rule format is "If A and B exist, In relation R1, and B and C have relation R2, then A and C have relation R3. Each rule matching must satisfy path closure and entity category continuity, and output rule reasoning path data. Based on the rule reasoning path data, semantic propagation operations are performed on the nodes involved in the graph. The propagation method is the adjacent node weight average diffusion method. The number of propagation iterations per round is limited to 5 times, and the propagation range is all nodes on the path. The attention mechanism weight update process uses the importance coefficient of the relationship between nodes for adjustment. The importance coefficient is calculated based on the weighted edge semantic score and connectivity, and finally the graph reasoning embedding data is formed. The data structure includes node number, semantic vector after propagation, attention distribution vector and propagation path sequence.Based on the graph reasoning embedded data, similarity clustering is performed using the K-Means clustering algorithm. The number of clusters is set to 12, and similarity is calculated using the Euclidean distance sum-squared minimization criterion. After clustering, upstream and downstream paths are predicted for each cluster center node. The path prediction method is a probabilistic path enumeration method based on the graph structure, with a path depth of 3 hops. The semantic distance and structural distance between each node and the center node in the predicted path are recorded and combined to generate the reasoning relationship generation data. The initial knowledge graph and the reasoning relationship generation data are structurally merged. The merging strategy is to add triplets in the reasoning relationship that do not exist in the initial graph to the graph structure. If there are duplicates, the relationship weights are redistributed. The redistribution formula is: final weight = 0.7 × original weight + 0.3 × reasoning weight. The weight value is constrained to the range [0, 1]. This completes the construction of the extended knowledge graph data.
[0062] The present invention realizes the refined extraction and classification of Chinese medicinal materials semantic annotation enhancement data into graph nodes, so that various types of Chinese medicinal materials related information can be uniformly expressed in the form of structured nodes; by constructing clear semantic relationships and edge type definitions between nodes, the graph has a semantic network structure with strong interpretability and high traversability; through data deduplication and semantic consistency processing, redundant and ambiguous data are eliminated, and the overall accuracy and standardization of the graph are improved; through the graph structure traversal and rule matching method driven by logical rules, implicit Chinese medicinal materials relationship knowledge can be derived based on existing data, enhancing the knowledge expansion capability of the knowledge graph; the introduction of graph node embedding and attention mechanism improves the accuracy of semantic feature propagation and the effectiveness of reasoning paths; similarity clustering and path prediction operations based on reasoning embedding realize the completion from local node relationships to overall network semantics, effectively bridging knowledge blind spots; through the merging and weight redistribution of reasoning generated data and original graph structure, an extended knowledge graph with more complete semantic coverage, stronger reasoning ability and faster retrieval response is formed, which comprehensively improves the practicality and intelligence level of Chinese medicinal materials data in data retrieval, data analysis and intelligent applications.
[0063] Preferably, step S3 includes the following steps:
[0064] Step S31: Identify node types and parse edge attributes on the extended knowledge graph to obtain graph structure hierarchical data;
[0065] Step S32: dynamically load and configure the graph database structure based on the graph structure hierarchical data to obtain graph database schema definition data, where the cache size is limited to 30% of the total memory;
[0066] Step S33: writing entity nodes and edge relationships into the extended knowledge graph based on the graph database schema definition data to obtain graph database mapping data;
[0067] Step S34: Reconstruct the syntax of path reachability, node attribute key values, and relationship constraint structure based on the graph database mapping data to obtain graph query language original expression data;
[0068] Step S35: Perform language conversion, alias mapping, and path semantic decoupling on the original graph query language expression data to obtain standardized graph query language data.
[0069] In an embodiment of the present invention, node type identification and edge attribute parsing operations are performed on the extended knowledge graph. Node type identification performs fixed type mapping based on the category field to which the entity belongs, and maps the categories "herb entity", "efficacy entity", "indication entity", "meridian entity", and "application site entity" to "HerbNode", "EffectNode", "IndicationNode", "MeridianNode", and "PartNode" respectively. Edge attribute parsing is organized into edge attribute sets in a triplet manner based on the "relationship type", "semantic score", and "source path" of the edge. Finally, graph structure hierarchical data is generated according to the graph depth structure. The graph structure hierarchical data includes the category of each layer of nodes, the number of nodes, and the edge connectivity with the upper and lower layer nodes. Dynamic loading and configuration of the graph database structure is performed based on the graph structure hierarchical data, and the database type is limited to Neo4j that supports graph structure indexing and multi-label nodes. Enterprise version, the configuration process includes automatic registration of node labels, initialization of edge attribute templates and loading of index rules. A unified index mapping table is generated for node labels and edge types during the loading process. The database memory management strategy is set to a cache size of 30% of the total memory capacity. If the total memory is 64GB, the cache size is limited to 19.2GB. After loading, the graph database schema definition data is generated; based on the graph database schema definition data, all entity nodes and edge relationships in the extended knowledge graph are written. Node writing is performed in batch mode, with each batch processing quantity being 10,000 nodes and 20,000 edges. The transaction isolation level used in the writing process is READ_COMMITTED. When writing edges, the relationship type, connection node ID, and edge ID are all processed. Weights and edge attribute values are fully entered. Once written, graph database mapping data is generated. This mapping data structure contains each node's label, unique ID, attribute key-value set, edge type, start and end node IDs, and attribute list. Syntactic reconstruction expressions for path reachability, node attribute key-value mapping, and relationship constraint structures are constructed based on the graph database mapping data. Path reachability is determined using a graph traversal method, with a maximum path length of 5 hops. If a path exists between nodes that meets the edge type and edge weight requirements, a path expression is generated. Node attribute key-value construction rules are based on a mapping table of field names corresponding to node labels. Edge relationship constraint structures are defined based on semantic template rules. For example, "HerbNode-[has effect]->EffectNode" must contain a semantic score greater than 0.75 edge relationships are reconstructed to generate graph query language raw expression data. This expression is organized using the Cypher language syntax structure, and the statement structure includes a MATCH clause, a WHERE filter condition, and a RETURN output field. The graph query language raw expression data undergoes language conversion, alias mapping, and path semantic decoupling. The language conversion process uniformly replaces Chinese field names in the expression with English abbreviations, such as "herb name" with "herb_name." Alias mapping sets abbreviation rules based on node type, such as mapping HerbNode to "H" and EffectNode to "E." Path semantic decoupling splits nested query paths into independent sub-paths and introduces intermediate node variables, making the query logic independent and clear. Finally, standardized graph query language data is generated. Each statement in this data conforms to the Neo4j Cypher language execution syntax. Variable naming rules uniformly use lowercase letters followed by numbers. Each statement limits the maximum query path length to no more than four hops. A semantic decoupling description and a field alias mapping table are included.
[0070] The present invention constructs clear and layered graph structure data by identifying node types and parsing edge attributes of the extended knowledge graph, so that the graph has a high degree of organization and hierarchy in the subsequent storage and query process, which is convenient for efficiently locating target information; the dynamic loading and configuration process of the graph database structure driven by graph structure layered data improves the automation of database initialization and the adaptability of graph data, and at the same time ensures the rational utilization of system memory resources and the stability of query performance by limiting the cache size to 30% of the total memory; during the graph database writing process, standardized node and edge structure mapping is used to ensure data consistency, transaction integrity and writing efficiency; path reachability analysis and relationship constraint reconstruction mechanism give query logic precise structural semantic support, ensuring that graph query statements can be strictly constructed based on actual semantic paths, avoiding invalid path calculations; finally, through language conversion, alias mapping and path semantic decoupling operations, the original query expression is standardized into a unified, clearly structured and easy-to-execute graph query language format, thereby significantly improving the graph retrieval efficiency, semantic matching ability and query interaction flexibility of Chinese medicinal materials data, and providing stable support for accurate access and knowledge discovery of large-scale Chinese medicinal materials semantic data.
[0071] Of particular importance is the grammatical reconstruction of path reachability, node attribute key values, and relationship constraint structures based on graph database mapping data, including:
[0072] Perform path structure topology analysis on the graph database mapping data to obtain node connection vector data;
[0073] Perform path reachability evaluation on the graph database mapping data based on the node connection vector data to obtain path reachability label data;
[0074] Based on the path reachability label data, the constraint logic rules of the edge relationship attributes are matched to obtain the relationship constraint structure data;
[0075] Perform node attribute key value extraction and semantic mapping on the relational constraint structure data to obtain attribute semantic key value pair data;
[0076] A path structure syntax tree is constructed based on the attribute semantic key values to obtain the original expression data of the graph query language.
[0077] In an embodiment of the present invention, the Cypher language based on the Neo4j graph database is first used to perform path structure topology analysis on the constructed Chinese medicinal materials knowledge graph, and the path traversal API interface of the graph database is called to obtain all path structures within three hops in the graph, and the node IDs and edge relationships in the path are sequentially combined to form vectorized node connection data. Each connection vector is represented by a fixed format as [starting node ID, relationship type 1, intermediate node ID, relationship type 2, target node ID]. Then, the depth-first search (DFS) algorithm is used to calculate the path connectivity of the node connection vector. There must be a continuous edge connection between any two nodes in the path. If connected, the path is marked as "reachable" to generate path reachability label data. The label Boolean values are placed on the corresponding path structure; then, based on the generated reachable path, logical rule matching is performed on each edge relationship attribute through the pre-defined Chinese medicinal material relationship constraint rule library. The rule library is stored in JSON format, and each rule contains at least three items: relationship type, entity category pair, and attribute constraint. For example, if there is an edge relationship of "medicinal material-containing ingredients-chemical composition" in a path, and the ingredient has "molecular formula" as the attribute key, it is necessary to verify whether its value matches the regular expression ^[AZ][az]?[0-9]*$. If it meets the requirements, the edge is retained as a legal relationship to form a structured relationship constraint dataset; then, the attribute key value extraction operation is performed on each node in the dataset, and the Cypher query statement format is called MATCH (n) WHERE ID(n)=node ID RETURN n.attribute name. The extraction result is a key-value pair structure, where the attribute name is uniformly named in the graph, such as "function and indication" and "origin distribution". Then, based on the extracted key-value pairs, the forward dictionary mapping technology is used to standardize the terms to the ontology vocabulary set. For example, "antipyretic" is mapped to "pharmacological effect / antipyretic class" to form attribute semantic key-value pair data; finally, the grammar rule tree construction method is adopted, and the recursive descent method is used to combine the path structure and attribute semantic data to generate a graph query language expression. The root node of each syntax tree is the target query entity, and the child nodes are nested in sequence. The relationship edge and the target attribute node. The expression construction format is (MATCH (n1)-[:relationship1]->(n2)-[:relationship2]->(n3) WHERE n3.attribute='standard value'), completing the generation of the original expression data of the graph query language.
[0078] The path structure topology analysis in the present invention clarifies the connection relationship between nodes, and provides basic support for path-level graph structure analysis by constructing node connection vector data; the path reachability evaluation is based on the node connection vector data, accurately determining the set of reachable paths between any two nodes in the graph, thereby avoiding invalid paths from participating in the query logic construction and improving the computational efficiency of the grammatical structure; the constraint logic rule matching operation of the edge relationship attribute establishes a matching relationship between the edge type and the node semantics, and enhances the semantic constraint of the edge attribute in the graph structure by binding strict semantic constraint rules to each edge; the node attribute key value extraction and semantic mapping process parses clear key-value pair data from the relationship constraint structure, and combines the graph semantic information to construct an attribute mapping system with semantic identification capabilities; finally, through the construction of a path structure syntax tree based on semantic key values, a graph query language original expression with clear structure, complete semantics and consistent rules is formed, which provides a structural basis and logical support for the subsequent standardization, execution and optimization of the graph query language, and comprehensively improves the performance of the Chinese medicinal materials knowledge graph in terms of retrieval structure accuracy, query response speed and semantic recognition ability.
[0079] Preferably, step S4 includes the following steps:
[0080] Step S41: performing parameterized condition definition and logical expression construction on graph query language data to obtain query condition template data;
[0081] Step S42: combining the query condition template data and sorting the priority to obtain an optimized query condition set;
[0082] Step S43: performing graph data path calculation and shortest path analysis on the optimized query condition set to obtain a set of Chinese medicinal material knowledge paths;
[0083] Step S44: Calculate the node association degree of the Chinese medicinal material knowledge path set and evaluate the impact factor to obtain node linkage relationship data;
[0084] Step S45: Structuring the node linkage relationship data and performing weight calculation to obtain graph structure query result data.
[0085] In an embodiment of the present invention, parameterized condition definitions are performed on the Cypher language query statements generated by the graph database mapping, and hard-coded rules are used to replace constant values in the query with placeholders in the format of "$parameter name". For example, "function and indication = clearing heat and detoxifying" is converted into "function and indication = $function and indication", and data type restrictions (such as string, integer, Boolean type) and value ranges are set for the placeholder parameters. By setting a key-value pair structure, it is uniformly saved in the parameter configuration file, and then a logical expression tree bound to the parameters is constructed. The logical operators are restricted to three types: AND, OR, and NOT. The infix expression method is used to combine the conditions so that each condition node and its logical dependency relationship are clearly presented, and query condition template data is generated; based on the subset expansion principle in Boolean algebra, multiple query condition templates are combined and traversed to generate all legal condition permutation groups The conditions are then sorted according to their importance priority. The priority is set with a weight value based on the domain knowledge experience rule base. For example, "name of medicinal material" is set to a weight of 10, "chemical composition" is set to a weight of 8, and "place of origin" is set to a weight of 6. The optimized query condition set is obtained by weighting and sorting in all combinations; the Dijkstra algorithm is used to calculate the path distance of each path corresponding to the node pair in the graph in the optimized query condition set, limiting the search depth to no more than 5 hops, and ensuring that all edges in the path meet the entity category constraints. The shortest legal path is screened and summarized into a set of Chinese medicinal material knowledge paths, where the path data structure is represented by an ordered node sequence and an edge sequence, and the path length is measured by the number of hops and recorded in the path metadata; the co-occurrence frequency method is used to calculate the correlation value of any two nodes in the knowledge path set, and the correlation value = co-occurrence number / The total number of paths is calculated by limiting the calculation scope to only node pairs with adjacent hops less than or equal to 2. Then, based on the position of each node in the path, the type of connection relationship, and the path length, a weighted scoring function is used to evaluate the influence factor value of the node in the query structure. The influence factor scoring formula is: influence factor = α × connectivity + β × relationship diversity + γ × path centrality, with parameters α, β, and γ fixed at 0.4, 0.3, and 0.3 respectively. The node linkage relationship data is structured in a triple format as (starting node ID, ending node ID, weight value). All weight values are normalized according to the influence factor. The Min-Max normalization method is used to normalize all weight values to the range [0, 1]. Finally, the graph structure query result data is generated. The result data format is unified as a JSON array. Each record contains the node ID, node name, path details, edge type, and its weight value.
[0086] The parameterized condition definition and logical expression construction in the present invention provide a universal and reusable query template structure, support condition configuration for multiple types of Chinese medicinal materials information, and improve the flexibility and adaptability of query rules; condition combination and priority sorting operations realize the orderly management of complex query conditions, so that the query execution process can efficiently parse and schedule resources according to logical priority, thereby improving query efficiency and accuracy; graph data path calculation and shortest path analysis effectively avoid redundant paths through a path selection mechanism based on a graph traversal algorithm, thereby improving the path focusing ability in knowledge graph queries; node correlation calculation combined with the impact factor evaluation method constructs a semantic strength perception mechanism, quantifies the actual semantic linkage relationship between nodes in the graph, and provides a data basis for graph structure reasoning and recommendation; finally, the structured expression and weight calculation of the node linkage relationship provide an interpretable sorting basis for the graph query return results, ensuring that the graph structure query results are highly stable and targeted in terms of structural integrity, semantic consistency and availability, thereby achieving high-quality retrieval and intelligent decision-making support for Chinese medicinal materials data.
[0087] Of particular importance is that the graph data path calculation and shortest path analysis for the optimized query condition set includes:
[0088] Perform path node matching and edge weight initialization on the optimized query condition set to obtain path candidate structure data;
[0089] Perform graph traversal strategy selection and path expansion based on the path candidate structure data to obtain path traversal result data;
[0090] Perform path reachability screening and loop removal on the path traversal result data to obtain legal path set data;
[0091] Based on the legal path set data, path length calculation and multi-path scoring are performed to obtain the shortest path candidate set data;
[0092] The shortest path candidate set data is subjected to knowledge entity correlation analysis and semantic similarity scoring to obtain a set of Chinese medicinal materials knowledge paths.
[0093] In the embodiment of the present invention, the path node matching and edge weight initialization operations are first performed. For each parameter in the optimized query condition set, the node label matching function of the Neo4j database is called, and the entity node with the query attribute is identified based on the MATCH statement. The matching node set is obtained by using the attribute value precise matching method. Each node is set as the path starting point or end point according to the attribute. Then, the edge weight is initialized based on the edge attribute information established in the knowledge graph. The edge weight is calculated in a fixed way: edge weight = base value 1.0 ÷ relationship weight factor, where the relationship weight factor is set in descending order according to the relationship semantic strength as 1.0 (direct ingredient), 1.2 (function and indication), 1.5 (origin), 2.0 (drug property), etc. class), all nodes and edge structures constitute the path candidate structure data; then, based on the path candidate structure data, the heuristic A* algorithm is selected as the graph traversal strategy to implement the path extension operation in the graph database. During the extension process, the Euclidean heuristic function is used as the basis for path cost estimation, and the maximum extension depth is limited to 5 hops, the maximum number of extended paths does not exceed 100, and each hop extension operation must comply with the entity category connection constraint rules. For example, the "Chinese herbal medicine" entity can only point to the "chemical component" entity through the "containing ingredients" edge to complete the extraction of path traversal result data; then, all the paths obtained by the traversal are screened for path reachability and loop removal. The reachability is based on the node continuous connectivity and direction consistency. Legality determination: a breadth-first search algorithm is used to check whether each node in the path has both incoming and outgoing edges, which is consistent with the path definition. The loop removal strategy adopts the node access status marking method. Once a node is repeatedly visited during the traversal process, the path is removed to obtain the set of loop-free legal paths. Subsequently, path length calculation and multi-path scoring are performed on each path in the legal path set data. The path length is recorded in units of hops, and the scoring is sorted according to the path score function: path score = ∑(node confidence × edge weight ⁻¹), where the node confidence is set by the node type, such as "core medicinal materials" is set to 0.9 and "auxiliary ingredients" is set to 0.7. Finally, the top 10 shortest paths are included in the shortest path candidate set. Finally, the shortest path candidate set data is subjected to knowledge entity correlation analysis and semantic similarity scoring. The correlation analysis is based on the node co-occurrence frequency matrix. The higher the number of co-occurrences in the path, the stronger the correlation. The similarity scoring adopts the average cosine similarity method of term vectors. The word vector library (limited to the ontology vocabulary vector table) is called for the entity name in each path, and the path-level average vector is constructed respectively. The semantic similarity between the path and the query condition is calculated. Finally, the paths with high correlation and high semantic similarity are screened and summarized into a set of Chinese medicinal materials knowledge paths. The set is expressed in JSON structure. Each path record includes path ID, entity sequence, relationship sequence, number of path hops, average edge weight, semantic similarity value and correlation score.
[0094] The path node matching and edge weight initialization in the present invention ensure the accurate correspondence between the query conditions and the entity nodes and relationship edges in the graph structure, and set the initial weight value according to the edge type and semantic strength, providing a quantitative basis for subsequent path evaluation; the graph traversal strategy selection and path expansion adopt a graph search algorithm based on breadth-first or depth-first, combined with the semantic characteristics of Chinese medicinal materials to gradually expand feasible paths, improve the path coverage while controlling the expansion depth to avoid resource redundancy; path reachability screening and loop removal constrain the traversal results and perform structural norms, eliminate unreachable paths and loop structures, and ensure the logical coherence and semantic validity of the path; path length Degree calculation and multi-path scoring and sorting set the maximum path depth to 6 layers, and combine the path coverage, edge weight accumulation value and semantic integrity scoring mechanism to quantitatively score and sort the path candidate set; finally, the knowledge entity relevance analysis and semantic similarity scoring are integrated based on the calculation method of word vector cosine similarity and context semantic relevance, and the knowledge entity pairs contained in the path are bidirectionally semantically scored and weighted averaged to screen out a set of Chinese medicinal materials knowledge paths with wide coverage, high correlation strength and excellent semantic consistency, thereby providing knowledge pathway support with accurate path, clear structure and clear logic for subsequent knowledge retrieval and graph reasoning.
[0095] Preferably, the step S5 of encoding the user search request with a semantic vector includes:
[0096] Perform word segmentation and stop word filtering on user search request data to obtain a standardized search term set;
[0097] The word segmentation granularity in the word segmentation is limited to two-tuple to four-tuple phrases, and the length of the reserved words is limited to 2-10 characters;
[0098] Perform semantic classification on the normalized search term set to obtain user search intent label data, where the classification confidence threshold of the semantic classification is set to be greater than or equal to 0.75;
[0099] Contextualize the user's search intent label data and supplement it with information to obtain enhanced search request data. The amount of supplemented information is limited to 50%-100% of the original request.
[0100] Initialize the neural network model for the enhanced retrieval request data to obtain calculation-ready state data;
[0101] Perform deep learning model inference on the computing readiness state data and perform high-dimensional vector mapping to obtain the initial user query semantic vector data;
[0102] Perform dimensionality reduction and feature extraction on the initial user query semantic vector data to obtain user query semantic vector data;
[0103] The dimension of the reduced vector in the dimensionality reduction process is limited to 128-256 dimensions, and the information retention rate threshold is set to 90%.
[0104] In an embodiment of the present invention, user search request data is first received, and word segmentation and stop word filtering operations are performed on the input text. The word segmentation process adopts the custom dictionary mode of the Jieba word segmentation tool, and the word segmentation granularity range is set to two-tuple to four-tuple phrases. The length of each phrase is limited to 2 to 10 characters. After segmentation, common words, invalid conjunctions and meaningless auxiliary words preset in the stop word list are removed to obtain a normalized search word set; then, the normalized search word set is semantically classified, and a keyword and semantic label mapping rule set is called. The rule set is constructed based on the Chinese medicinal material field ontology, including four types of semantic labels: "medicinal material name", "pharmacological effect", "applied subject" and "adaptive symptoms". A keyword matching scoring method is used to assign a confidence value to each word. The confidence is determined by the word frequency and the overlap of the label keywords. The classification label confidence threshold is limited to greater than or equal to 0.75. Only labels with confidence that meet the requirements are retained to generate user search intention label data; then, a contextual semantic association operation is performed on the search intention label data, and a rule engine is used to call the Chinese medicinal material field knowledge base for information supplement according to the label type. The amount of information supplementation The number of characters in the original request is controlled between 50% and 100%. The supplementary information includes synonyms of the terms under the label, common compound words and typical associated words, which are combined to form enhanced retrieval request data; then the neural network model initialization operation is performed on the enhanced retrieval request data, and the preloaded BERT embedding model is called and preheated and loaded into the calculation graph. After completing the graph construction and parameter weight loading, the data enters the calculation ready state; then the deep learning model inference process is performed on the loaded model, and the enhanced retrieval request data text is input. The BERT encoder outputs the embedding representation of each token, and then it is aggregated into a fixed-dimensional high-dimensional vector through the pooling layer to obtain the initial user query semantic vector data, and the vector dimension is set to 768 dimensions; then the vector is subjected to dimensionality reduction processing, and the principal component analysis (PCA) algorithm is used to extract the main axis of features. The target dimension of dimensionality reduction is limited to between 128 and 256 dimensions, and the lower limit of the information retention rate is set to 90% to retain the main component information and contextual semantic features in the input semantic structure. Finally, the user query semantic vector data that meets the dimensionality reduction requirements is output.
[0105] The word segmentation and stop word filtering in the present invention ensure that invalid words in the original text are eliminated, and the word segmentation granularity setting of two-tuple to four-tuple phrases and the word length range limitation of 2 to 10 characters make the keyword extraction more stable and semantically complete; the semantic classification operation classifies the normalized search word set into semantic labels, and improves the accuracy and representativeness of the classification results by setting the classification confidence threshold to be no less than 0.75; the context association and information supplement mechanism enables the enhanced search request data to expand its information volume without exceeding the original semantic boundary, and the information supplement ratio is controlled in the range of 50% to 100%, effectively enhancing the context background and improving the subsequent semantics The integrity of the expression; the neural network model initialization operation ensures that the system reasoning conditions are consistent with the data input dimensions, providing an environmental guarantee for efficient reasoning; the high-dimensional vector mapping stage adopts a deep learning reasoning method to construct the initial semantic vector expression, and extracts potential semantic features through a multi-layer network structure to realize the vector encoding of the user request semantics; finally, by setting the dimension of the vector after dimensionality reduction to 128 to 256 dimensions and the information retention rate to be no less than 90%, it is ensured that the dimensionality reduction result retains the original semantic information to the maximum extent while compressing the dimension, so that the final user query semantic vector data in the subsequent graph matching and path calculation has significantly improved the computational efficiency and matching accuracy.
[0106] Preferably, the step S5 of constructing a set of medicinal material description semantic vectors based on the Chinese medicinal material semantic annotation enhanced data includes:
[0107] Extract feature attributes from the semantically enhanced data of Chinese medicinal materials and screen key descriptions to obtain the core description data of Chinese medicinal materials;
[0108] Perform word segmentation and semantic unit splitting on the core description data of Chinese medicinal materials to obtain standardized Chinese medicinal materials text units;
[0109] The maximum length of a semantic unit in the semantic unit splitting is limited to 20 characters, and the minimum length of a semantic unit is limited to 2 characters;
[0110] Perform parallel vector conversion on the standardized Chinese herbal medicine text units to obtain the original Chinese herbal medicine description vector set, where the vector model batch size is limited to 64-128;
[0111] The original Chinese herbal medicine description vector set is subjected to dimensionality reduction and noise filtering to obtain a set of Chinese herbal medicine description semantic vectors, where the singular value decomposition (SVD) energy retention ratio is set to 85%-95%;
[0112] The semantic vector set of Chinese herbal medicine descriptions is used to perform matching calculations on the user query semantic vector data to obtain vector semantic retrieval result data.
[0113] In an embodiment of the present invention, feature attributes are first extracted from the constructed Chinese medicinal material semantic annotation enhanced data, and the feature attributes include five structured fields: medicinal material functions and indications, medicinal properties, morphological description, harvesting and processing, and nature, flavor and meridians. Regular rules and part-of-speech tagging algorithms are called to extract high-frequency keywords and phrases in each field, and the top 10% keywords of each field are extracted to form a key description set. After merging, redundant phrases are removed to obtain the core description data of Chinese medicinal materials; then, word segmentation and semantic unit splitting operations are performed on the core description data of Chinese medicinal materials. In the word segmentation stage, a maximum matching method based on a domain dictionary is used to segment sentences into basic semantic blocks according to part-of-speech rules. The length of the semantic units after splitting is strictly controlled between 2 and 20 characters. For parts exceeding the maximum length, sentence segmentation is performed according to punctuation and grammatical dependencies to obtain normalized Chinese medicinal material text units; then, a parallel vector conversion operation is performed on the text units, and a BERT embedding engine or an equivalent static embedding mapping function is called to batch input normalized text data. The batch size of the vector model is limited to between 64 and 128 items, and each text unit After encoding, an initial description vector with a dimension of 768 is generated and summarized as the original Chinese medicinal material description vector set. Then, the original Chinese medicinal material description vector set is subjected to dimensionality reduction and noise filtering. The singular value decomposition (SVD) algorithm is used to perform principal component decomposition according to the vector column dimension. The energy retention ratio is set between 85% and 95%. The dimensions corresponding to low-weight singular values are removed and the vector set is reconstructed. At the same time, vector dimensions with variance below a set threshold are zeroed to achieve noise filtering, resulting in a set of Chinese medicinal material description semantic vectors. Finally, the above semantic vector set is matched with the generated user query semantic vector data. The matching algorithm adopts the cosine similarity matching mechanism. Under the premise of consistent feature dimensions, each Chinese medicinal material description vector and user query vector are normalized separately. Then, a vector-by-vector dot product operation is performed and the cosine value of the angle is calculated. Matching pairs with a similarity threshold of 0.65 or above are retained as matching records. Finally, the vector semantic retrieval result data is outputted in descending order according to the matching score, which includes the Chinese medicinal material identification, description text fragment, matching score and corresponding semantic label.
[0114] The present invention can accurately extract representative semantic content from the medicinal material body by extracting feature attributes and screening key descriptions of semantic annotation enhancement data of traditional Chinese medicine; it can refine the medicinal material description text into basic units with clear semantic granularity through word segmentation and semantic unit splitting operations, and limit the length of each semantic unit to between 2 and 20 characters during splitting, thereby enhancing the stability and consistency of text representation; when performing parallel vector conversion, the number of samples processed in each batch is set between 64 and 128 to ensure computing resource control and operation efficiency in the vector generation stage; then the original Chinese medicinal material description vector set is reduced in dimension by singular value decomposition, and the SVD retained energy ratio is controlled within the range of 85% to 95%, effectively compressing the vector dimension while retaining the semantic principal component information, reducing high-dimensional redundancy and vector noise interference; finally, matching calculation is performed between the user query semantic vector data and the Chinese medicinal material description semantic vector set at the vector level, which can achieve accurate alignment of user intention and medicinal material information in the high-dimensional semantic space, and significantly improve the relevance and response accuracy of semantic retrieval.
[0115] Preferably, the matching calculation of the user query semantic vector data in step S5 includes:
[0116] Set retrieval parameters for the user query semantic vector data to obtain retrieval-ready vector data, where the similarity calculation accuracy is limited to float32;
[0117] Calculate the cosine similarity of the retrieval-ready vector data and the set of semantic vectors describing Chinese medicinal materials. Perform cosine similarity calculation and Euclidean distance evaluation. After linear weighting, the two are used to obtain the comprehensive vector similarity and form a vector similarity matrix.
[0118] The vector similarity matrix is sorted in descending order and threshold filtered to obtain a preliminary matching result set, where the similarity threshold is set to 0.65;
[0119] Calculate the semantic association strength of the preliminary matching result set, and perform context consistency evaluation on the semantic association strength calculation results to obtain the semantic weighted matching results;
[0120] Perform semantic relevance clustering on the semantic weighted matching results to obtain classified retrieval result data;
[0121] Remove redundancy from the classification retrieval result data to obtain optimized vector semantic retrieval results;
[0122] The optimized vector semantic retrieval results are structurally encapsulated and metadata supplemented to obtain vector semantic retrieval result data.
[0123] In an embodiment of the present invention, the generated user query semantic vector data is first input into the matching engine, and the search parameters are set, including limiting the vector data precision type to float32 and the vector dimension to 128 to 256 dimensions, to obtain the search-ready vector data; then the search-ready vector data is compared one by one with the Chinese medicinal material description semantic vector set, and cosine similarity calculation and Euclidean distance evaluation are performed. The Euclidean distance is normalized by an exponential function and converted into a similarity score. The weight coefficient of the cosine similarity is 0.7, and the weight coefficient of the Euclidean distance similarity is 0.3. After linear weighting, the comprehensive vector similarity is obtained to form a vector similarity matrix; the vector similarity matrix is sorted in descending order according to the similarity value, and the sorting result is threshold filtered, and the similarity lower limit is set to 0.65. Only records with similarity values greater than or equal to 0.65 are retained to obtain a preliminary matching result set; the semantic association strength of each group of vectors in the preliminary matching result set is calculated, and the semantic association strength is determined by the matching keyword overlap, part of speech consistency rate, and the semantic association strength. , and the angle between the vector directions constitute a weighted scoring system, with the basic weight set to 0.6; after the calculation is completed, the context consistency evaluation is performed based on the context co-occurrence frequency of the semantic tag, the context consistency score weight is set to 0.4, and it is weighted and fused with the semantic association strength to obtain the final semantic weighted matching result; then the semantic weighted matching result is subjected to a semantic relevance clustering operation, and the clustering is based on the cosine similarity, the context keyword co-occurrence frequency and the number of keyword matches. The weighting coefficient of the keyword matching part is set to 1.5 to increase the influence of keyword overlap on the selection of clustering centers, and finally the classified retrieval result data is obtained; in the classification results, the redundant or similar description content is removed, and the duplicates are removed by comparing the semantic label and keyword coverage, and the most representative description vector is retained to obtain the optimized vector semantic retrieval result; finally, the optimized vector semantic retrieval result is added with the original identification of the Chinese medicinal materials, the description source, and the semantic label metadata, encapsulated into a structured record format, and the vector semantic retrieval result data is output.
[0124] The present invention ensures the balance between computing resources and numerical accuracy in the retrieval calculation process by setting the data type format of float32 precision; adopts a combined scoring mechanism of cosine similarity and Euclidean distance, and sets the cosine similarity weight coefficient to 0.7, the Euclidean distance weight coefficient to 0.3, and the Euclidean distance conversion coefficient to 5.0, so that the directionality and amplitude information can be measured simultaneously in the semantic space, thereby improving the discrimination and stability of the similarity evaluation; after the vector similarity matrix is generated, the similarity threshold lower limit is set to 0.65, which effectively eliminates low-correlation data and improves the semantic accuracy of the preliminary matching result set; the semantic association strength of the preliminary matching results is evaluated. Calculation and context consistency evaluation enhance the responsiveness of matching results to actual user search intentions, and realize multi-level fusion between semantic judgment and language expression through multi-factor control of association strength basic weight (0.6), context consistency weight (0.4) and keyword matching weight coefficient (1.5); further semantic relevance clustering of semantic weighted matching results is performed to realize classification structure presentation, and strengthen the logical aggregation ability of similar medicinal material results; finally, through redundant data removal, structured encapsulation and metadata supplementation, the expression quality and application effect of the final vector semantic retrieval results in terms of structural integrity, content readability and system compatibility are improved.
[0125] Preferably, step S6 includes the following steps:
[0126] Step S61: aligning the graph structure query result data with the vector semantic retrieval result data by entity index to obtain result alignment mapping data;
[0127] Step S62: performing a cross calculation between the node attribute labels in the graph structure query result and the high-dimensional semantic labels in the vector semantic retrieval result based on the result alignment mapping data to obtain cross semantic association data;
[0128] Step S63: fusing the graph structure relationship weight and the semantic similarity score of the cross-semantic association data to obtain fused score data;
[0129] Step S64: performing multi-factor sorting, i.e., confidence threshold filtering, on the fused scoring data to obtain semantic sorting result data of Chinese medicinal materials;
[0130] Step S65: based on the Chinese medicinal material semantic ranking result data, the attribute information, path relationship and similar Chinese medicinal material nodes related to the Chinese medicinal material nodes are extracted to obtain semantic retrieval result structure data;
[0131] Step S66: formatting and interface encapsulating the semantic search result structure data, and performing data caching processing to obtain the final Chinese medicinal material semantic search response data.
[0132] In an embodiment of the present invention, first, entity index alignment is performed on the Chinese medicinal material node identifier, path structure and attribute information in the graph structure query result data and the Chinese medicinal material description vector index field in the vector semantic retrieval result data, and a hash mapping table is constructed based on the matching method of the primary key field (such as the unique code of the Chinese medicinal material or the standard naming field) to generate result alignment mapping data; based on the aligned mapping data, a cross-similarity calculation is performed on the node attribute label in the graph structure query result and the high-dimensional semantic label in the vector semantic retrieval result, and the vector dot product method is used to calculate the cosine similarity between the attribute keyword vector and the description semantic vector, and the semantic matching confidence threshold is set to 0.75 to generate cross-semantic association data; the cross-semantic association data is weightedly fused with the node relationship weight data recorded in the graph structure, and the fusion method is to multiply the semantic similarity by 0.6 and the structure weight value by 0.4, and normalize them to the [0,1] range. The fusion scoring data is formed; the fusion scoring data is sorted in descending order by score, and the confidence threshold is set to be greater than or equal to 0.7 for filtering to obtain the semantic sorting result data of Chinese medicinal materials; based on the high-matching nodes in the semantic sorting results, the attribute fields, path-related nodes and edge relationship data connected to them in the graph structure are extracted, and the adjacent Chinese medicinal material nodes with cosine similarity values greater than 0.8 are extracted as semantically similar nodes based on the vector distance ranking, and the semantic retrieval result structure data including the target Chinese medicinal material node, attribute set, path set and similar Chinese medicinal material set is constructed; the format of the semantic retrieval result structure data is normalized, including field naming unification, field type formatting, timestamp standardization, and structure encapsulation in JSON format. At the same time, the structured data is cached in the Redis database according to the access frequency strategy, and the cache validity period is limited to 300 seconds to generate the final Chinese medicinal material semantic retrieval response data.
[0133] The present invention achieves a unique mapping relationship between entity nodes in the graph structure and semantic vectors in the vector space through entity index alignment operations, ensuring the consistency and correspondence of cross-modal data. The calculation of cross-semantic association data integrates graph structure node attributes and semantic similarity labels, providing a joint scoring basis for subsequent sorting. In the fusion scoring stage, the weights of the edge relationships between nodes in the graph structure are weighted and synthesized with the semantic vector matching scores, enhancing the structural dependence and context perception capabilities of the semantic scoring, effectively improving the interpretability and stability of the sorting process. A confidence threshold filtering mechanism is introduced into the multi-factor sorting. By setting a lower limit for the sorting confidence, non-relevant items with weak structural associations and low semantic scores are filtered out, further enhancing the accuracy of the sorting results. Based on the sorting results, the attribute information, path relationships, and similar nodes of the Chinese medicinal material nodes are extracted, which not only enriches the semantic output content but also enhances the reasoning and recommendation capabilities of the retrieval system. Finally, the semantic retrieval result structure data is formatted and encapsulated in an interface, improving the integration efficiency of the output results during the system call process. Combined with the data caching processing strategy, it can effectively reduce the computing resource overhead and response time of repeated queries, improving the system operation efficiency and service quality.
[0134] Preferably, the present invention further provides a data retrieval Chinese medicinal material data management system for executing the above-mentioned data retrieval Chinese medicinal material data management method, wherein the data retrieval Chinese medicinal material data management system comprises:
[0135] The Chinese medicine information semantic annotation module is used to obtain the morphological data of Chinese medicinal materials; perform grade assessment based on the Chinese medicinal materials morphological data to generate Chinese medicinal materials grade label data; obtain images of the Chinese medicinal materials to be tested; pre-process the images of the Chinese medicinal materials to be tested, and perform entity recognition and relationship extraction to obtain Chinese medicinal materials semantic annotation data; and fuse the Chinese medicinal materials grade label data and the Chinese medicinal materials semantic annotation data to obtain Chinese medicinal materials semantic annotation enhanced data.
[0136] The medicinal material knowledge graph construction and reasoning module is used to use the semantic annotation of traditional Chinese medicine to enhance data to build an initial knowledge graph, perform logical rule matching and implicit relationship reasoning on the initial knowledge graph, and obtain an extended knowledge graph;
[0137] The graph database mapping module is used to map the extended knowledge graph to the graph database structure and generate standardized graph query language data;
[0138] The graph structure query module is used to construct graph queries based on graph query language data by combining conditions, traversing paths, and linking nodes to obtain graph structure query result data;
[0139] The user semantic matching retrieval module is used to obtain user search request data; perform semantic vector encoding on the user search request data to obtain user query semantic vector data; construct a set of Chinese medicinal material description semantic vectors based on the Chinese medicinal material semantic annotation enhanced data, and perform matching calculation on the user query semantic vector data to obtain vector semantic retrieval result data;
[0140] The retrieval result fusion response module is used to fuse and compare the graph structure query result data with the vector semantic retrieval result data to obtain the final Chinese medicinal material semantic retrieval response data.
[0141] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present invention is not limited by the above description. Therefore, it is intended that all changes that fall within the meaning and scope of the equivalent elements of the application documents are included in the present invention.
[0142] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.
Claims
1. A method for managing Chinese medicinal materials data based on data retrieval, characterized in that: The following steps are involved: Step S1: Acquire Chinese medicinal material sample morphological data; perform grade assessment based on the Chinese medicinal material morphological data to generate Chinese medicinal material grade label data; acquire images of the Chinese medicinal material to be tested; preprocess the images of the Chinese medicinal material to be tested, and perform entity recognition and relationship extraction to obtain Chinese medicinal material semantic annotation data; fuse the Chinese medicinal material grade label data and the Chinese medicinal material semantic annotation data to obtain Chinese medicinal material semantic annotation enhanced data; Step S2: Using the semantically-annotated data of Chinese medicinal materials to build an initial knowledge graph, performing logical rule matching and implicit relationship reasoning on the initial knowledge graph to obtain an extended knowledge graph; wherein, step S2 includes the following steps: Step S21: performing field parsing and entity classification on the Chinese medicinal material semantic annotation enhanced data to obtain candidate data of Chinese medicinal material graph nodes; Step S22: Based on the candidate data of the Chinese medicinal material graph nodes, semantic relationship mapping and edge type definition are performed between nodes to obtain structured knowledge graph data; Step S23: Deduplication, semantic consistency processing and standard graph model mapping are performed on the structured knowledge graph data to obtain an initial knowledge graph; Step S24: traverse the graph structure and filter the rule nodes of the initial knowledge graph, and use the preset logical reasoning rule set to perform symbolic rule matching to obtain rule reasoning path data; Step S25: Based on the rule-based reasoning path data, perform semantic propagation of the graph node embedding vector and update the attention mechanism weight to obtain graph reasoning embedding data; Step S26: perform similarity clustering on the graph reasoning embedding data and perform upstream and downstream path prediction to obtain reasoning relationship generation data; Step S27: Structurally merge the initial knowledge graph with the inference relationship generation data and redistribute the weights to obtain the extended knowledge graph data; Step S3: Map the extended knowledge graph to the graph database structure and generate standardized graph query language data; Step S4: constructing a graph query based on the graph query language data by combining conditions, traversing paths, and linking nodes to obtain graph structure query result data; Step S5: Obtain user search request data; perform semantic vector encoding on the user search request data to obtain user query semantic vector data; construct a set of Chinese medicinal material description semantic vectors based on the Chinese medicinal material semantic annotation enhanced data, and perform matching calculation on the user query semantic vector data to obtain vector semantic search result data; Step S6: Fusing and comparing the graph structure query result data and the vector semantic retrieval result data to obtain the final Chinese medicinal material semantic retrieval response data.
2. The Chinese medicinal material data management method based on data retrieval according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: collecting appearance images of the Chinese medicinal materials through microscopic photography, collecting fiber structure images and surface texture atlas data through high-definition multispectral imaging, and recording the appearance images, fiber structure images and surface texture atlas data as Chinese medicinal materials sample morphology data; collecting intrinsic component characteristic data of the Chinese medicinal materials through texture analysis; Step S12: performing feature fusion on the Chinese medicinal material sample morphology data and the Chinese medicinal material internal component feature data, and performing redundancy removal to construct a Chinese medicinal material morphology-component joint feature matrix; Step S13: Classifying the Chinese medicinal materials using classification rules based on the Chinese medicinal materials morphology-ingredient joint feature matrix to generate Chinese medicinal materials grade label data; Step S14: obtaining a handwritten image of Chinese medicinal materials; performing image preprocessing, image enhancement, and text recognition on the handwritten image of Chinese medicinal materials to obtain structured Chinese medicinal materials data; Step S15: performing corpus recognition and content structured annotation on the structured Chinese medicinal material data to obtain Chinese medicinal material text parsing data, wherein the structured Chinese medicinal material data is segmented into semantic units using a sentence recognition method based on a regular rule set, with the number of rules being no less than 30, including content types such as origin, properties, and processing methods. The segmented units are marked as data fields and embedded with content tags to obtain Chinese medicinal material text parsing data; Step S16: performing named entity recognition and term boundary segmentation on the Chinese herbal medicine text parsing data to obtain Chinese herbal medicine semantic entity data; Step S17: performing context-aware analysis based on the Chinese medicinal material semantic entity data to obtain Chinese medicinal material semantic relationship data; Step S18: Merging the data structure of the Chinese medicinal material semantic entity data and the Chinese medicinal material semantic relationship data, and uploading them to the edge server to obtain the Chinese medicinal material semantic annotation data; Step S19: Fusing the Chinese medicinal material grade label data and the Chinese medicinal material semantic annotation data to obtain Chinese medicinal material semantic annotation enhanced data.
3. The Chinese medicinal material data management method based on data retrieval according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: Identify node types and parse edge attributes on the extended knowledge graph to obtain graph structure hierarchical data; Step S32: dynamically load and configure the graph database structure based on the graph structure hierarchical data to obtain graph database schema definition data, where the cache size is limited to 30% of the total memory; Step S33: writing entity nodes and edge relationships into the extended knowledge graph based on the graph database schema definition data to obtain graph database mapping data; Step S34: Reconstruct the syntax of path reachability, node attribute key values, and relationship constraint structure based on the graph database mapping data to obtain graph query language original expression data; Step S35: Perform language conversion, alias mapping, and path semantic decoupling on the original graph query language expression data to obtain standardized graph query language data.
4. The Chinese medicinal material data management method based on data retrieval according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: performing parameterized condition definition and logical expression construction on graph query language data to obtain query condition template data; Step S42: combining the query condition template data and sorting the priority to obtain an optimized query condition set; Step S43: performing graph data path calculation and shortest path analysis on the optimized query condition set to obtain a set of Chinese medicinal material knowledge paths; Step S44: Calculate the node association degree of the Chinese medicinal material knowledge path set and evaluate the impact factor to obtain node linkage relationship data; Step S45: Structuring the node linkage relationship data and performing weight calculation to obtain graph structure query result data.
5. The Chinese medicinal material data management method based on data retrieval according to claim 1, characterized in that: The semantic vector encoding of the user search request in step S5 includes: Perform word segmentation and stop word filtering on user search request data to obtain a standardized search term set; The word segmentation granularity in the word segmentation is limited to two-tuple to four-tuple phrases, and the length of the reserved words is limited to 2-10 characters; Perform semantic classification on the normalized search term set to obtain user search intent label data, where the classification confidence threshold of the semantic classification is set to be greater than or equal to 0.75; Contextualize the user's search intent label data and supplement it with information to obtain enhanced search request data. The amount of supplemented information is limited to 50%-100% of the original request. Initialize the neural network model for the enhanced retrieval request data to obtain calculation-ready state data; Perform deep learning model inference on the computing readiness state data and perform high-dimensional vector mapping to obtain the initial user query semantic vector data; Perform dimensionality reduction and feature extraction on the initial user query semantic vector data to obtain user query semantic vector data; The dimension of the reduced vector in the dimensionality reduction process is limited to 128-256 dimensions, and the information retention rate threshold is set to 90%.
6. The Chinese medicinal material data management method based on data retrieval according to claim 1, characterized in that: The step S5 of constructing a set of Chinese medicinal material description semantic vectors based on the Chinese medicinal material semantic annotation enhanced data includes: Extract feature attributes from the semantically enhanced data of Chinese medicinal materials and screen key descriptions to obtain the core description data of Chinese medicinal materials; Perform word segmentation and semantic unit splitting on the core description data of Chinese medicinal materials to obtain standardized Chinese medicinal materials text units; The maximum length of a semantic unit in the semantic unit splitting is limited to 20 characters, and the minimum length of a semantic unit is limited to 2 characters; Perform parallel vector conversion on the standardized Chinese herbal medicine text units to obtain the original Chinese herbal medicine description vector set, where the vector model batch size is limited to 64-128; The original Chinese herbal medicine description vector set is subjected to dimensionality reduction and noise filtering to obtain a set of Chinese herbal medicine description semantic vectors, where the singular value decomposition (SVD) energy retention ratio is set to 85%-95%; The semantic vector set of Chinese herbal medicine descriptions is used to perform matching calculations on the user query semantic vector data to obtain vector semantic retrieval result data.
7. The Chinese medicinal material data management method based on data retrieval according to claim 1, characterized in that: The matching calculation of the user query semantic vector data in step S5 includes: Set retrieval parameters for the user query semantic vector data to obtain retrieval-ready vector data, where the similarity calculation accuracy is limited to float32; Calculate the cosine similarity of the retrieval-ready vector data and the set of semantic vectors describing Chinese medicinal materials. Perform cosine similarity calculation and Euclidean distance evaluation. After linear weighting, the two are used to obtain the comprehensive vector similarity and form a vector similarity matrix. The vector similarity matrix is sorted in descending order and threshold filtered to obtain a preliminary matching result set, where the similarity threshold is set to 0.65; Calculate the semantic association strength of the preliminary matching result set, and perform context consistency evaluation on the semantic association strength calculation results to obtain the semantic weighted matching results; Perform semantic relevance clustering on the semantic weighted matching results to obtain classified retrieval result data; Remove redundancy from the classification retrieval result data to obtain optimized vector semantic retrieval results; The optimized vector semantic retrieval results are structurally encapsulated and metadata supplemented to obtain vector semantic retrieval result data.
8. The data retrieval and management method for Chinese medicinal materials according to claim 1, characterized in that: Step S6 includes the following steps: Step S61: aligning the graph structure query result data with the vector semantic retrieval result data by entity index to obtain result alignment mapping data; Step S62: performing a cross calculation between the node attribute labels in the graph structure query result and the high-dimensional semantic labels in the vector semantic retrieval result based on the result alignment mapping data to obtain cross semantic association data; Step S63: fusing the graph structure relationship weight and the semantic similarity score of the cross-semantic association data to obtain fused score data; Step S64: performing multi-factor sorting, i.e., confidence threshold filtering, on the fused scoring data to obtain semantic sorting result data of Chinese medicinal materials; Step S65: based on the Chinese medicinal material semantic ranking result data, the attribute information, path relationship and similar Chinese medicinal material nodes related to the Chinese medicinal material nodes are extracted to obtain semantic retrieval result structure data; Step S66: formatting and interface encapsulating the semantic search result structure data, and performing data caching processing to obtain the final Chinese medicinal material semantic search response data.
9. A data retrieval system for Chinese herbal medicine data, characterized in that: A Chinese herbal medicine data management method for executing data retrieval according to claim 1, wherein the Chinese herbal medicine data management system for data retrieval comprises: The Chinese medicine information semantic annotation module is used to obtain the morphological data of Chinese medicinal materials; perform grade assessment based on the Chinese medicinal materials morphological data to generate Chinese medicinal materials grade label data; obtain images of the Chinese medicinal materials to be tested; pre-process the images of the Chinese medicinal materials to be tested, and perform entity recognition and relationship extraction to obtain Chinese medicinal materials semantic annotation data; and fuse the Chinese medicinal materials grade label data and the Chinese medicinal materials semantic annotation data to obtain Chinese medicinal materials semantic annotation enhanced data. The medicinal material knowledge graph construction and reasoning module is used to use the semantic annotation of traditional Chinese medicine to enhance data to build an initial knowledge graph, perform logical rule matching and implicit relationship reasoning on the initial knowledge graph, and obtain an extended knowledge graph; The graph database mapping module is used to map the extended knowledge graph to the graph database structure and generate standardized graph query language data; The graph structure query module is used to construct graph queries based on graph query language data by combining conditions, traversing paths, and linking nodes to obtain graph structure query result data; The user semantic matching retrieval module is used to obtain user search request data; perform semantic vector encoding on the user search request data to obtain user query semantic vector data; construct a set of Chinese medicinal material description semantic vectors based on the Chinese medicinal material semantic annotation enhanced data, and perform matching calculation on the user query semantic vector data to obtain vector semantic retrieval result data; The retrieval result fusion response module is used to fuse and compare the graph structure query result data with the vector semantic retrieval result data to obtain the final Chinese medicinal material semantic retrieval response data.
Citation Information
Patent Citations
Knowledge inference and fault diagnosis method based on knowledge graph
CN114756686A
Rule and path-based traditional Chinese medicine multi-modal knowledge graph reasoning method and device
CN116705338A