Traditional Chinese medicinal material data management system and method for data retrieval

By constructing a semantic labeling of Chinese medicinal materials, and combining graph database mapping and standardized graph query language generation, the problems of single data structure and weak semantic understanding ability in the existing system are solved, and the deep fusion and semantic structured expression of Chinese medicinal materials are realized, and the accuracy and context consistency of search results are improved.

CN120045744AActive Publication Date: 2025-05-27JINGZHOU QIANCAO BIOTECHNOLOGY CO LTD

Patent Information

Application Number
CN202510499353.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-27
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing Chinese medicinal materials data management system has problems such as single data structure, weak semantic understanding ability, and lack of knowledge correlation modeling, which is difficult to support the in-depth exploration and intelligent retrieval of complex medicinal materials knowledge.

Method used

By constructing a semantic labeling enhancement data and extended knowledge graph, combining graph database mapping and standardized graph query language generation, it realizes the deep fusion and semantic structured expression of Chinese medicinal materials images, levels, and text information, and uses logical rule matching and implicit relationship reasoning to establish a knowledge network with linkage reasoning capabilities.

Benefits of technology

It significantly improves the organizational efficiency and semantic expression ability of Chinese medicinal materials data, supports search response in complex contexts, improves the accuracy and context consistency of search results, and meets the needs of modern pharmaceutical research and development and digital research of traditional Chinese medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045744A_ABST
    Figure CN120045744A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electric digital data processing, in particular to a traditional Chinese medicine data management system and method for data retrieval. The method comprises the following steps: acquiring morphology data of a traditional Chinese medicinal material sample; performing grade evaluation based on the traditional Chinese medicinal material morphology data to generate traditional Chinese medicinal material grade label data; acquiring a to-be-detected traditional Chinese medicinal material image; preprocessing the to-be-detected traditional Chinese medicinal material image, and performing entity recognition and relation extraction to obtain traditional Chinese medicinal material semantic annotation data; and fusing the traditional Chinese medicinal material grade label data and the traditional Chinese medicinal material semantic annotation data to obtain traditional Chinese medicinal material semantic annotation enhanced data. According to the method, the information of the traditional Chinese medicinal materials is expressed through structured processing and semantic enhancement, the knowledge graph containing dominant and implicit semantic relationships is constructed, and the retrieval depth, flexibility and intelligent application of knowledge of the traditional Chinese medicinal materials are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a Chinese medicinal material data management system and method for data retrieval. Background Art

[0002] In the traditional field of Chinese medicinal materials, data management mainly relies on manual sorting and static literature records, which are difficult to support the efficient organization and retrieval requirements of data for modern medical research and precision medicine. With the development of informatization, a Chinese medicinal material management system based on a relational database has been gradually introduced, realizing the electronic storage and simple retrieval of the basic attributes, classification, and sources of medicinal materials. However, such systems generally have problems such as a single data structure, weak semantic understanding ability, and lack of knowledge association modeling, making it difficult to support the in-depth mining and intelligent retrieval of complex medicinal material knowledge. In recent years, some studies have attempted to introduce semantic modeling means to improve the organization and expression ability of Chinese medicinal material data. However, the existing methods still have many deficiencies: on the one hand, semantic construction mostly relies on manual rules, resulting in low system intelligence and poor scalability; on the other hand, there is a lack of in-depth semantic understanding and precise matching mechanism for user query intentions, making it difficult to achieve retrieval responses in complex contexts. Summary of the Invention

[0003] Based on this, it is necessary for the present invention to provide a Chinese medicinal material data management system and method for data retrieval to solve at least one of the above technical problems.

[0004] To achieve the above object, a Chinese medicinal material data management method for data retrieval includes the following steps: Step S1: Obtain the morphological data of Chinese medicinal material samples; perform grade evaluation based on the morphological data of Chinese medicinal materials to generate Chinese medicinal material grade label data; obtain the image of the Chinese medicinal material to be tested; preprocess the image of the Chinese medicinal material to be tested, and perform entity recognition and relationship extraction to obtain Chinese medicinal material semantic annotation data; fuse the Chinese medicinal material grade label data and the Chinese medicinal material semantic annotation data to obtain enhanced Chinese medicinal material semantic annotation data; Step S2: Use the enhanced Chinese medicinal material semantic annotation data to construct an initial knowledge graph, perform logical rule matching and implicit relationship reasoning on the initial knowledge graph to obtain an extended knowledge graph; Step S3: Map the extended knowledge graph to a graph database structure to generate standardized graph query language data; Step S4: Based on the graph query language data, perform graph query construction of condition combination, path traversal, and node linkage to obtain graph structure query result data; Step S5: Obtain user retrieval request data; perform semantic vector encoding on the user retrieval request data to obtain user query semantic vector data; construct a Chinese medicinal material description semantic vector set based on the Chinese medicinal material semantic annotation enhanced data, and perform matching calculation on the user query semantic vector data to obtain vector semantic retrieval result data; Step S6: Perform fusion comparison on the graph structure query result data and the vector semantic retrieval result data to obtain the final Chinese medicinal material semantic retrieval response data.

[0005] Through the construction of Chinese medicinal material semantic annotation enhanced data and an extended knowledge graph, the present invention realizes the deep fusion and semantic structured expression of Chinese medicinal material images, grades, and text information, significantly improving the organization efficiency and semantic expression ability of Chinese medicinal material data; on the basis of the knowledge graph, logical rule matching and implicit relationship reasoning are adopted to establish a knowledge network with linkage reasoning ability, breaking through the limitations of static associations in traditional databases; through graph database mapping and the generation of a standardized graph query language, a graph structure retrieval mechanism supporting complex query modes such as path traversal, condition combination, and node linkage is constructed, enhancing the system's response ability to multi-dimensional medicinal material relationships; the fusion of vector semantic encoding and the construction of a Chinese medicinal material semantic vector set realizes the multi-dimensional expression and precise matching of the user's retrieval intention, solving the problems of insufficient understanding of user input and poor relevance of retrieval results in traditional systems; finally, through the fusion comparison of graph query results and semantic retrieval results, taking into account both structural relationships and semantic similarity, the accuracy and context consistency of retrieval results are improved, enabling the Chinese medicinal material data management system to not only have a highly intelligent data organization ability but also support precise response to complex semantic problems and in-depth semantic retrieval, meeting the core needs of modern pharmaceutical research and Chinese medicine digital research.

[0006] Preferably, the present invention further provides a Chinese medicinal material data management system for data retrieval, which is used to execute the above-mentioned Chinese medicinal material data management method for data retrieval. The Chinese medicinal material data management system for data retrieval includes: A traditional Chinese medicine information semantic annotation module, which is used to obtain the morphological data of Chinese medicinal material samples; perform grade evaluation based on the morphological data of Chinese medicinal materials to generate Chinese medicinal material grade label data; obtain the image of the Chinese medicinal material to be tested; preprocess the image of the Chinese medicinal material to be tested, and perform entity recognition and relationship extraction to obtain Chinese medicinal material semantic annotation data; fuse the Chinese medicinal material grade label data and the Chinese medicinal material semantic annotation data to obtain Chinese medicinal material semantic annotation enhanced data; A medicinal material knowledge graph construction and reasoning module, which is used to construct an initial knowledge graph by using the Chinese medicinal material semantic annotation enhanced data, and perform logical rule matching and implicit relationship reasoning on the initial knowledge graph to obtain an extended knowledge graph; A graph database mapping module, which is used to perform graph database structure mapping on the extended knowledge graph to generate standardized graph query language data; The graph structure query module is used to construct a graph query for conditional combination, path traversal, and node linkage based on graph query language data, and obtain graph structure query result data; The user semantic matching and retrieval module is used to obtain user retrieval request data; perform semantic vector encoding on the user retrieval request data to obtain user query semantic vector data; construct a collection of Chinese herbal medicine description semantic vectors based on enhanced Chinese herbal medicine semantic annotation data, and perform matching calculations on the user query semantic vector data to obtain vector semantic retrieval result data; The retrieval result fusion and response module is used to fuse and compare the graph structure query result data and the vector semantic retrieval result data to obtain the final Chinese herbal medicine semantic retrieval response data.

[0007] Through the collaborative operation of each functional module, the present invention constructs a complete closed-loop data management mechanism in key links such as Chinese herbal medicine data collection, semantic expression, knowledge association, and intelligent retrieval, significantly improving the information structuring, standardization, and semantic processing capabilities of Chinese herbal medicine; in the data acquisition and processing stage, it realizes the multi-modal fusion processing of the morphological characteristics, grade labels, and image semantic information of Chinese herbal medicine, improving the recognition accuracy of raw data and the label expression accuracy; in the knowledge modeling stage, through graph construction and reasoning mechanisms, it realizes the structured modeling and logical extension between semantic entities, attributes, and relationships, enhancing the system's organization and reasoning capabilities for the complex knowledge system of Chinese herbal medicine; through graph database mapping and graph query construction, the system has efficient graph structure data operation capabilities, supports flexible multi-dimensional conditional combination and associated path discovery, and provides a graph structure-level response basis for professional problems; in the user interaction layer, based on semantic vector representation and matching mechanisms, it significantly enhances the system's ability to recognize, understand, and retrieve user intentions, realizing the deep semantic expression and multi-angle matching of query requests; finally, through the joint comparison and fusion sorting of graph structure results and semantic vector results, it realizes the dual guarantee of structural logical relevance and semantic relevance, improving the retrieval accuracy, result relevance, and response intelligence level of the system in complex contexts, and comprehensively meeting the practical needs of Chinese herbal medicine knowledge management and intelligent services. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objectives, and advantages of the present invention will become more apparent: Figure 1 It is a schematic flow chart of the steps of a Chinese herbal medicine data management method for data retrieval according to the present invention; Figure 2 is Figure 1 a detailed schematic flow chart of step S1 in Figure 3 is Figure 1 a detailed schematic flow chart of step S2 in Detailed implementation mode

[0009] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those skilled in the art within the scope of the present invention without creative work belong to the scope of protection of the present invention.

[0010] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.

[0011] It should be understood that although the terms "first", "second", etc. may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.

[0012] To achieve the above object, please refer to Figures 1 to 3 , the present invention provides a method for managing traditional Chinese medicine data for data retrieval, and the method includes the following steps: Step S1: Obtain the morphological data of traditional Chinese medicine samples; perform grade evaluation based on the morphological data of traditional Chinese medicine to generate grade label data of traditional Chinese medicine; obtain the image of the traditional Chinese medicine to be tested; preprocess the image of the traditional Chinese medicine to be tested, and perform entity recognition and relationship extraction to obtain semantic annotation data of traditional Chinese medicine; fuse the grade label data of traditional Chinese medicine and the semantic annotation data of traditional Chinese medicine to obtain enhanced semantic annotation data of traditional Chinese medicine; In the embodiments of the present invention, a microscopic camera is used to collect surface images of each batch of traditional Chinese medicine samples, and the image resolution is set to not less than 4096×2160 pixels; at the same time, a hyperspectral imager with a wavelength coverage range of 400nm to 1000nm is used to obtain fiber structure images and surface texture maps, and the image sampling frequency is set to 20 pixel points per millimeter. The obtained image data is integrated to form the morphology data of the traditional Chinese medicine samples; subsequently, a texture analyzer is used to collect component characteristic data of the traditional Chinese medicine at a compression speed of 1mm / s. The collected data includes hardness, elasticity, adhesiveness, and cohesiveness, with units of N, mm, and g. The morphology data and the component characteristic data are fused through a feature cascading algorithm, and a redundancy removal method based on principal component analysis (PCA) is adopted to retain the principal components with a cumulative variance contribution rate of not less than 95%, constructing a combined morphology-component feature matrix of traditional Chinese medicine; then, according to the artificially preset classification rules, which include making grade judgments based on indicators such as fiber density, surface texture clarity, and texture strength, each row in the combined feature matrix is input into the rule logic judgment process, and traditional Chinese medicine grade label data is generated according to the grade standard, with the grade division accuracy not less than 5 levels; immediately afterwards, an image of the traditional Chinese medicine to be tested is obtained through a high-definition imaging device fixedly installed on the conveyor belt. The image size is 1920×1080 pixels, and the image acquisition rate is not less than 30 frames per second. The obtained image is input into an image preprocessing module, which performs operations including image denoising (Gaussian filter, kernel size of 5×5), edge enhancement (Sobel operator), and normalized gray-scale transformation (normalization range [0,1]). Subsequently, an image entity recognition algorithm based on morphological feature matching is used to identify the boundaries of the entity regions in the traditional Chinese medicine image, with the recognition accuracy limit of not less than 92%. Then, an attribute-relation-value triple set between entities is constructed based on a relationship extraction algorithm of dictionary matching and co-occurrence frequency analysis to form traditional Chinese medicine semantic annotation data; finally, the above-generated traditional Chinese medicine grade label data and the traditional Chinese medicine semantic annotation data are aligned based on the sample unique number, and dimension fusion is performed through a feature connection method to construct an enhanced semantic feature vector set. This fused data is named enhanced traditional Chinese medicine semantic annotation data.

[0013] Step S2: Use the enhanced traditional Chinese medicine semantic annotation data to construct an initial knowledge graph, perform logical rule matching and implicit relationship reasoning on the initial knowledge graph to obtain an extended knowledge graph; In the embodiments of the present invention, taking the enhanced data of traditional Chinese medicine semantic annotation as input, first, a field parsing algorithm is used to perform field segmentation and attribute recognition on the text data in the enhanced data of semantic annotation. The field segmentation adopts a linear traversal method based on character boundary position matching, and the attribute recognition rules include fixed keyword tag recognition and part-of-speech dependency structure judgment. The field parsing results are saved in a field-value pair structure; then, the entity classification process is executed. The field parsing results are input into an entity classification dictionary table, which is constructed based on a standard traditional Chinese medicine terminology set and contains entity types such as "herb name", "origin", "morphological characteristics", "grade label", "processing method", etc. In the classification process, a method combining regular expression matching and hash dictionary indexing is used to generate candidate data for graph nodes; subsequently, based on the candidate data for graph nodes, semantic relationship mapping between nodes is performed. The mapping rules are based on attribute co-occurrence frequency, semantic dictionary mapping pairs, and hyponymy judgment. The defined edge types include relationship types such as "belongs to", "contains", "originates from", "has characteristics", etc. Each type of edge is specified with a direction and a unique identification code, and the generated relationship data is constructed into structured knowledge graph data; after the graph is constructed, a data deduplication process is executed. The uniqueness of nodes and edges is judged through a hash check code (with a length set to 32 bits), and duplicate nodes and redundant edges are deleted; then, semantic consistency processing is performed. The method is a combined filtering of word vector cosine similarity calculation (with a threshold set to 0.9) and semantic merging rule judgment (based on attribute value similarity greater than 0.85) to obtain the graph data after semantic consistency; then, the structured graph data is mapped to a preset standard graph model, which is constructed based on the Neo4j graph structure template. When mapping, the unified node attribute format is a "key-value" pair structure, and the edge attribute structure is a triple format; finally, an initial knowledge graph is generated; based on the initial knowledge graph, graph structure traversal and rule node filtering operations are performed. The graph traversal uses the depth-first traversal (DFS) method, and the node filtering conditions include a node type matching degree not lower than 95%, an edge weight greater than 0.6, and a semantic coverage rate greater than 80%; subsequently, a logical inference rule set is called. The rule set is stored in the form of rule entries, and each rule contains a triple of a premise node, a logical relationship, and a result node. The symbolic rule matching process is executed, and the inference path is triggered through condition judgment to generate rule inference path data; the rule inference path data is input into the graph embedding processing module to initialize the embedding vectors of the nodes in the path (with the vector dimension set to 128), and vector semantic extension is performed based on the path order propagation mechanism and the attention mechanism weights are updated. The initial value of the attention weight is set to a mean of 0.1, and the maximum value is limited to 1.0 to obtain graph inference embedding data; then, a vector clustering method based on cosine similarity is used to perform similarity clustering on the embedding data, and the clustering similarity threshold is set to 0.92, and perform upstream and downstream path prediction on the clustering results. The predicted path is limited to no more than 3 hops, and the inferred relationship generation data is output. Finally, the initial knowledge graph and the inferred relationship generation data are structurally merged. The merging process adopts a double-layer structure integration method of node alignment + edge merging. After the merging is completed, the weight value of each edge in the new graph is reallocated. The edge weight is calculated by weighting the semantic intensity and the path frequency. The formula is weight value = 0.6 × semantic similarity + 0.4 × path occurrence frequency normalization value. Finally, the extended knowledge graph data is obtained.

[0014] Step S3: Map the extended knowledge graph to a graph database structure to generate standardized graph query language data; In the embodiment of the present invention, the extended knowledge graph is used as input data, and the conversion of the extended knowledge graph to graph query language data is realized through a graph database structure mapping module. The graph database structure mapping module is deployed in a multi-core parallel computing node. The graph database used is Neo4j Community Edition 5.12, which supports graph data interaction based on the Cypher language. During the mapping process, the node type recognition sub-module is first called to perform a structural scan on the entity nodes in the knowledge graph. The identified node types include four types: "herb node", "function node", "meridian node", and "disease node". Each type of node is attached with an attribute field. Among them, the text fields are uniformly encoded in UTF-8 format, and the maximum length is limited to 128 characters. Then, the relationship edges in the triple structure are extracted through the edge attribute parsing sub-module. The edge attributes are organized in a key-value pair format, and the edge weight field is limited to a floating-point number type, with the precision controlled to three decimal places. After the structure recognition is completed, the graph structure hierarchical data is written into the graph database through the graph database structure dynamic loading module. The memory cache mechanism is enabled during the loading process, and the cache size limit is 30% of the total physical memory. The cache replacement policy adopts the least recently used (LRU) algorithm. The graph database schema definition is constructed based on the Schema-first method. The entity labels, attribute constraints, and relationship types are defined through Cypher statements. After the writing is completed, the path reachability verification and relationship constraint analysis are performed through the syntax reconstruction sub-module, and the structure of the path expression that does not conform to the syntax rules is adjusted. The node connection symbol in the path expression is uniformly used as "--", and the attribute filtering condition is uniformly expressed using the "WHERE" clause. All expressions need to pass the regular matching function to verify the syntax compliance before finally being converted into graph query language data. The finally output standardized graph query language data is encapsulated in JSON format and stored in the high-speed read-write storage module.

[0015] Step S4: Build a graph query based on the graph query language data for condition combination, path traversal, and node linkage to obtain graph structure query result data; In the embodiment of the present invention, for graph query construction operations based on standardized graph query language data, first, the conditional combination module performs semantic analysis on the attribute filtering clauses in the standardized graph query language data, and uses the Boolean operation precedence parsing algorithm to perform left-associative precedence merging on the logical symbols "AND", "OR", and "NOT" to form query condition template data. The length limit of each field value in the query condition does not exceed 64 characters, and the field types include three types: string type, Boolean type, and enumeration type. Then, the path traversal module performs semantic path expansion processing on the query condition template data. The path expansion uses the breadth-first search (BFS) algorithm, and the search level is limited to no more than 5 levels to avoid the problem of path combination explosion. The node and edge data used in the search are derived from the graph database mapping data loaded into the Neo4j graph database in step S3. The starting node and ending node of each path are verified through the native Cypher syntax of the graph database, and only the path data that meets the syntax conditions and has the same edge direction is retained. Subsequently, the node linkage module is called to perform linkage calculation on the node set in the effective path set. This process is based on the attribute matching degree of the nodes, and the calculation method uses the weighted cosine similarity method. The linkage strength scoring range is between 0 and 1, and the node pairs with a score lower than 0.3 are excluded from the current path combination. After the node linkage is completed, a graph structure query expression is formed and submitted to the Neo4j graph database through the query expression execution engine for query operations. The query result node set, path set, and edge set data are returned. During the execution of the query operation, tasks are submitted in a four-thread concurrent manner, and the maximum execution duration of each query statement does not exceed 3000 milliseconds. If it exceeds, it is forced to interrupt and recorded in the exception log. All query results are encapsulated into graph structure query result data in a unified format by the structured encapsulation component.

[0016] Step S5: Obtain user retrieval request data; perform semantic vector encoding on the user retrieval request data to obtain user query semantic vector data; construct a Chinese herbal medicine description semantic vector set based on the Chinese herbal medicine semantic annotation data, and perform matching calculation on the user query semantic vector data to obtain vector semantic retrieval result data; In the embodiments of the present invention, first, the semantic encoding module is used to obtain the user's retrieval request data. The retrieval request data is sourced from the form fields of a standard HTTP POST request, with an encoding format of UTF-8 and a total field length not exceeding 512 characters. Subsequently, the forward maximum matching word segmentation algorithm based on dictionary mapping is used to segment the user's retrieval request data. The traditional Chinese medicine professional term dictionary used contains more than 10,000 entries, with the length of each entry not exceeding 10 Chinese characters. All stop words are excluded during the word segmentation process, and the stop word list is limited to 300 common function words. After word segmentation, the user's retrieval request data is semantically vector encoded by constructing word vectors. The encoding method combines the bag-of-words model and TF-IDF weights, with a word frequency threshold set to 1 and an inverse document frequency lower limit set to 0.1. After encoding, user query semantic vector data with a dimension of 300 is generated. In parallel, based on the traditional Chinese medicine semantic annotation data constructed in step S1, a set of traditional Chinese medicine description semantic vectors is constructed. The construction process of this set is based on the standard description fields of each traditional Chinese medicine, such as "medicinal properties", "functions and indications", "source", and "meridians entered", for TF-IDF feature extraction, with the extraction dimension kept consistent at 300 dimensions, and the total character length in the description fields limited to not exceed 2048 characters. After the construction of the above two vector sets is completed, the vector matching module is called to perform matching calculations. The matching method uses the Euclidean distance method, and the distance threshold is set to 2.5. Those with a distance less than the threshold are considered to have a successful match. They are sorted from smallest to largest in terms of distance, and the top 10 are taken to form the vector semantic retrieval result data.

[0017] Step S6: Fuse and compare the graph structure query result data and the vector semantic retrieval result data to obtain the final traditional Chinese medicine semantic retrieval response data.

[0018] In the embodiments of the present invention, the graph structure query result data and the vector semantic retrieval result data are respectively loaded into independent data structures. The graph structure query result data is stored in a relational table structure, and each record contains an entity ID, an entity name, a set of node attributes, and a path identifier; the vector semantic retrieval result data is organized in a dense vector array structure, with each vector dimension uniformly set to 192 dimensions and represented in the floating-point format float32, and the information retention rate reaches 92% to ensure the accuracy of semantic expression; during the implementation process, first, a one-to-one entity index mapping table is established by performing a hash index mapping on the entity ID fields in the two data structures, and then entity alignment operations are performed based on this mapping table; after the alignment is completed, attribute vectorization processing is performed on each pair of corresponding entities. The attribute fields are transformed into attribute vectors by using the Term Frequency-Inverse Document Frequency (TF-IDF) weighted average method, with the dimension set to 64 dimensions and the value range normalized to the [0, 1] interval; after the attribute vectors are prepared, the cosine similarity is calculated with the corresponding semantic vectors, and the similarity result is set as the semantic similarity label; based on this semantic similarity label and the graph structure weight parameters of the nodes, weighted fusion calculation is performed, and the fusion formula is: Final score = 0.6 × graph structure weight score + 0.4 × semantic similarity label score, where the score range is restricted to [0, 1]; then the fusion scores are sorted in descending order and filtered in combination with a confidence threshold of 0.7, and the results higher than this threshold are retained to form the traditional Chinese medicine semantic sorting result data; next, attribute information extraction, associated path analysis, and similar node reasoning are performed on each traditional Chinese medicine node in the sorting result. The path analysis uses the shortest path search based on the Dijkstra algorithm, and the node similarity determination is weighted by the attribute intersection ratio and the structural adjacency degree, with the similarity threshold set to 0.65; the above-extracted information is encapsulated in a structured data format, and the fields include node ID, main attributes, path sequence, set of similar node IDs, and similarity score; finally, this structured data is encoded in a unified JSON format and encapsulated into a response data structure through a predefined interface template. At the same time, 10% of the total memory space is allocated in the cache module for response data caching, and the cache uses the FIFO strategy for elimination control to obtain the final traditional Chinese medicine semantic retrieval response data.

[0019] The present invention realizes the deep integration and semantic structured expression of Chinese herbal medicine image, grade and text information by constructing the semantic annotation enhanced data and extended knowledge graph of Chinese herbal medicine, significantly improving the organization efficiency and semantic expression ability of Chinese herbal medicine data; on the basis of the knowledge graph, logical rule matching and implicit relationship reasoning are adopted to establish a knowledge network with linkage reasoning ability, breaking through the limitation of the static association of traditional databases; through graph database mapping and standardized graph query language generation, a graph structure retrieval mechanism supporting complex query modes such as path traversal, condition combination and node linkage is constructed, enhancing the system's response ability to multi-dimensional medicinal material relationships; by integrating vector semantic encoding and the construction of a set of Chinese herbal medicine semantic vectors, the multi-dimensional expression and precise matching of the user's retrieval intention are realized, solving the problems of insufficient understanding of the user's input and poor relevance of retrieval results in traditional systems; finally, through the fusion and comparison of graph query results and semantic retrieval results, taking into account both structural relationships and semantic similarity, the accuracy and context consistency of retrieval results are improved, enabling the Chinese herbal medicine data management system to not only have a highly intelligent data organization ability, but also support precise response to complex semantic problems and in-depth semantic retrieval, meeting the core needs of modern medicine research and development and Chinese medicine digital research.

[0020] Preferably, step S1 includes the following steps: Step S11: Collect the appearance image of Chinese herbal medicine through microscopic photography, and collect the fiber structure image and surface texture atlas data through a high-definition multi-spectral imager. Denote the appearance image, fiber structure image and surface texture atlas data as the morphological data of Chinese herbal medicine samples; collect the internal component characteristic data of Chinese herbal medicine through a texture analyzer; Step S12: Perform feature fusion on the morphological data of Chinese herbal medicine samples and the internal component characteristic data of Chinese herbal medicine, and remove redundancy to construct a morphological-component joint feature matrix of Chinese herbal medicine; Step S13: Based on the morphological-component joint feature matrix of Chinese herbal medicine, use classification rules to classify the grades of Chinese herbal medicine and generate grade label data of Chinese herbal medicine; Step S14: Obtain the handwritten image of Chinese herbal medicine; perform image preprocessing, image enhancement and character recognition on the handwritten image of Chinese herbal medicine to obtain structured Chinese herbal medicine data; Step S15: Perform prediction recognition and content structured annotation on the structured Chinese herbal medicine data to obtain Chinese herbal medicine text parsing data; Step S16: Perform named entity recognition and term boundary segmentation on the Chinese herbal medicine text parsing data to obtain Chinese herbal medicine semantic entity data; Step S17: Perform context-aware analysis based on the Chinese herbal medicine semantic entity data to obtain Chinese herbal medicine semantic relationship data; Step S18: Merge the data structures of the Chinese herbal medicine semantic entity data and the Chinese herbal medicine semantic relationship data, and upload them to the edge server to obtain Chinese herbal medicine semantic annotation data; Step S19: Integrate the traditional Chinese medicine grade label data and the traditional Chinese medicine semantic annotation data to obtain enhanced traditional Chinese medicine semantic annotation data.

[0021] In the embodiments of the present invention, an industrial-grade microscopic imaging device with a resolution of not less than 0.5 microns is used to perform tiled image acquisition on Chinese herbal medicine samples. The image acquisition light source uses an LED cold light source, and the brightness is constantly controlled at 4000 lux. The output image is saved in the uncompressed TIFF format with a size of 4096×2160 pixels. At the same time, a hyperspectral imager with a wavelength range of 400 nm to 1000 nm and a spectral resolution of not higher than 5 nm is used to perform multi-angle scanning on Chinese herbal medicine samples to obtain fiber structure images and surface texture map data. The image acquisition angle is 45° incident light, and the scanning speed is set to 10 mm / s. The three types of image data are respectively named and uniformly classified as the morphological data of Chinese herbal medicine samples. Further, a texture analyzer is used to collect parameters such as hardness, elasticity, and adhesion of Chinese herbal medicine with a probe diameter of 5 mm and a test speed of 1 mm / s. Each index is sampled no less than 10 times, and its average value is extracted as the internal component characteristic data of Chinese herbal medicine. The above-mentioned morphological data and component data are subjected to feature fusion, and the fusion method is feature concatenation. First, texture features are extracted from the image data, and then the texture data is standardized and spliced to form a unified numerical vector. The obtained data is input into a redundancy removal algorithm, and the principal component analysis method (PCA) is used to retain the principal components with a cumulative variance contribution rate of not less than 95%. Finally, a Chinese herbal medicine morphology-component joint feature matrix with a dimension of not more than 64 is formed. According to each record in the joint feature matrix, a manually formulated classification rule set is applied in turn. The rules include texture clarity distribution thresholds, elasticity range definitions, and texture periodicity judgments. The number of rules is not less than 15, and each rule corresponds to a first-level classification logic. Finally, Chinese herbal medicine grade label data is generated, and the labels represent 5 grade levels from "L1" to "L5". The paper records of Chinese herbal medicine are collected, and the handwritten images are obtained through a high-precision image acquisition instrument. The image resolution is set to 300 dpi. The images are input into the image preprocessing module, and are successively denoised (using median filtering with a window size of 3×3), image enhanced (stretching the contrast range to [0,255]), and binarized (using the OTSU threshold method). After image enhancement, the images are sent to the character recognition engine. The recognition method is based on character contour extraction and template matching. The extracted text structure is archived according to row and column positions to obtain structured Chinese herbal medicine data. A sentence pattern recognition method based on a regular rule set is used to perform semantic unit segmentation on the structured Chinese herbal medicine data. The number of rules is not less than 30, covering content types such as origin, characteristics, and processing methods. The segmented units are marked as data fields and embedded with content labels, and the output is Chinese herbal medicine text parsing data. A named entity recognition operation is performed on the text parsing data. The entity boundaries are located using an entity positioning method based on dictionary matching, and the term boundaries are segmented through context part-of-speech consistency judgment. The length limit of the segmented units is between 2 and 10 characters, and Chinese herbal medicine semantic entity data is generated.Perform context-aware analysis based on the context window mechanism. The window width for expanding to the left and right of each entity is set to no more than 5 words. Analyze the co-occurrence probability, dependency grammar structure, and semantic direction relationship between words, extract the context logical dependencies to form a set of semantic relationships, and generate the semantic relationship data of traditional Chinese medicinal materials; Merge the semantic entity data and semantic relationship data according to the unique entity identifier. The merging method is to construct a triple structure centered on the main entity, and the structure includes the format of "main entity - relationship - subordinate entity". After the structure data is completed, it is uploaded to the edge server through a 5G edge gateway device. The server configuration requirements are that the CPU main frequency is not less than 2.6 GHz and the memory capacity is not less than 64 GB. The output data is the semantic annotation data of traditional Chinese medicinal materials; Based on the sample number and collection timestamp of traditional Chinese medicinal materials, map the grade label data and semantic annotation data of traditional Chinese medicinal materials one by one, and use the feature field splicing method to fuse the two types of data in dimension to generate enhanced semantic annotation data of traditional Chinese medicinal materials and write it into the local cache database.

[0022] Through the construction of a multi-source heterogeneous data acquisition and fusion mechanism, the present invention realizes the synchronous acquisition and collaborative expression of the morphological characteristics and internal components of traditional Chinese medicinal materials, and improves the scientificity and objectivity of the grade evaluation of traditional Chinese medicinal materials; By introducing microscopic imaging and hyperspectral imaging means, the fine-grained acquisition of surface structure, fiber characteristics, and texture maps is realized. Combined with the quantitative processing of texture parameters, it effectively supports the standardization and automation of grade division; Through the image preprocessing and handwritten recognition processes, historical records and manual annotation information are further supplemented, enhancing the source breadth and expression dimension of structured data; In the text processing chain, a complete process from image recognition to structured annotation, then to semantic entity extraction and context relationship construction is completed, strengthening the semantic organization ability of the original text information; In the semantic data aggregation stage, through the edge computing platform for merging and uploading, the data processing efficiency and deployment flexibility are optimized; Finally, through the fusion processing of grade labels and semantic annotation results, the unified expression of quantitative grade information and semantic description information is realized, providing a unified, standardized, and high-dimensional semantic input data basis for subsequent knowledge modeling and retrieval calculations, effectively supporting the construction of the system's full-process semantic modeling, knowledge linkage, and intelligent recognition capabilities for traditional Chinese medicinal materials.

[0023] Preferably, step S2 includes the following steps: Step S21: Parse the fields and classify the entities of the enhanced semantic annotation data of traditional Chinese medicinal materials to obtain the candidate data of the traditional Chinese medicinal material atlas nodes; Step S22: Based on the candidate data of the traditional Chinese medicinal material atlas nodes, perform semantic relationship mapping between nodes and edge type definition to obtain structured knowledge graph data; Step S23: Perform data deduplication, semantic consistency processing, and standard atlas model mapping on the structured knowledge graph data to obtain the initial knowledge graph; Step S24: Traverse the graph structure of the initial knowledge graph and filter the rule nodes, and perform symbolic rule matching using a preset logical inference rule set to obtain rule inference path data; Step S25: Based on the rule inference path data, perform semantic propagation of the graph node embedding vectors and update the attention mechanism weights to obtain graph inference embedding data; Step S26: Perform similarity clustering on the graph inference embedding data and predict the upstream and downstream paths to obtain inference relationship generation data; Step S27: Merge the structures of the initial knowledge graph and the inference relationship generation data, and perform weight reallocation to obtain extended knowledge graph data.

[0024] In the embodiments of the present invention, a field parsing operation is performed on the enhanced data of the semantic annotation of traditional Chinese medicinal materials. The field extraction rule based on attribute name matching is adopted, and the entity information in the semantic annotation data is grouped and stored according to four types of fields: "name", "category", "attribute value", and "location code". And the entities are classified according to the category to which the entity attributes belong. The categories include "medicinal material entity", "efficacy entity", "indication entity", "meridian tropism entity", and "application part entity". The classification standard is determined by the field "belonging category" in the semantic annotation data. After classification, candidate data for the nodes of the traditional Chinese medicinal material atlas is generated, and each candidate node is recorded in the node registration table with a unique number; based on the candidate data of the atlas nodes, semantic relationship mapping between nodes is performed. The relationship mapping logic adopts the co-occurrence position relationship matching method of entity pairs in the original text. By setting the maximum window span to 120 characters, it is judged whether there is an effective semantic connection between the node pairs. If so, the edge type is defined according to its category and semantic content. The edge types include four fixed semantic relationships: "has efficacy", "is used for treatment", "belongs to meridian", and "acts on part". And the starting node, ending node, relationship type, and semantic score of each edge are recorded to form structured knowledge graph data; a data deduplication operation is performed on the structured knowledge graph data. The deduplication strategy is that if there are duplicate relationships for the same node pair, only the one with the highest semantic score is retained. The semantic consistency processing is based on a preset synonym normalization table and relationship category redirection table. The standard atlas model is organized in the form of Triplet "Entity A - Relationship - Entity B". The semantic score of each triplet is normalized to the interval [0, 1], and an initial knowledge graph is formed; a structure traversal operation is performed on the initial knowledge graph. The breadth-first search algorithm is used to explore the atlas hierarchy. The exploration depth is limited to 5 layers. Combined with the node label rules, rule node filtering is performed. The filtering condition is nodes with an edge connection degree lower than 2 or nodes with the category label "non-primary entity". After the structure filtering is completed, the symbolic rule matching is performed on the remaining paths using a preset logical inference rule set. The rule format is "If there is a relationship R1 between A and B, and there is a relationship R2 between B and C, then there is a relationship R3 between A and C". Each rule match must satisfy the path closure and entity category continuity, and the rule inference path data is output; based on the rule inference path data, a semantic propagation operation is performed on the nodes involved in the atlas. The propagation method is the adjacent node weight average diffusion method. The number of propagation iterations per round is limited to 5 times. The propagation range is all nodes on the path. The attention mechanism weight update process is adjusted using the node relationship importance coefficient. The importance coefficient is weighted and calculated according to the edge semantic score and connection degree. Finally, atlas inference embedding data is formed. The data structure includes node number, propagated semantic vector, attention distribution vector, and propagation path sequence;Perform similarity clustering operations on the data embedded based on graph reasoning. Adopt the K-Means clustering algorithm, set the number of clusters to 12, use the minimization criterion of the sum of squared Euclidean distances as the similarity calculation method. After clustering, predict the upstream and downstream paths for each class center node. The path prediction method is the probabilistic path enumeration method based on the graph structure, limit the path depth to 3 hops, and record the semantic distance and structural distance between each node and the center node in the predicted path, and combine them to generate the data for generating inference relationships; perform structural merging on the initial knowledge graph and the data for generating inference relationships. The merging strategy is to add the triples that do not exist in the initial graph in the inference relationship to the graph structure. If there are duplicates, reassign the relationship weights. The reassigning formula is: Final weight = 0.7×Original weight + 0.3×Inference weight, and the weight value is restricted between [0,1] to complete the construction of the extended knowledge graph data.

[0025] The present invention realizes the refined extraction and classification of the enhanced data of the semantic annotation of traditional Chinese medicinal materials into graph nodes, enabling various types of information related to traditional Chinese medicinal materials to be uniformly expressed in the form of structured nodes; by constructing clear semantic relationships and edge type definitions between nodes, the graph has a semantic network structure with strong interpretability and high traversability; through data deduplication and semantic consistency processing, redundant and ambiguous data are eliminated, improving the overall accuracy and standardization of the graph; through the graph structure traversal and rule matching method driven by logical rules, implicit knowledge of the relationships between traditional Chinese medicinal materials can be deduced based on the existing data, enhancing the knowledge expansion ability of the knowledge graph; introducing graph node embedding and attention mechanism, improving the accuracy of semantic feature propagation and the effectiveness of the inference path; the similarity clustering and path prediction operations based on inference embedding realize the complementation from local node relationships to the overall network semantics, effectively bridging the knowledge blind spots; through the merging and weight reallocation of the inference-generated data and the original graph structure, an extended knowledge graph with more comprehensive semantic coverage, stronger inference ability, and faster retrieval response is formed, comprehensively improving the practicality and intelligent level of traditional Chinese medicinal materials data in data retrieval, data analysis, and intelligent applications.

[0026] Preferably, step S3 includes the following steps: Step S31: Identify the node types and parse the edge attributes of the extended knowledge graph to obtain the hierarchical data of the graph structure; Step S32: Dynamically load and configure the graph database structure based on the hierarchical data of the graph structure to obtain the graph database schema definition data, where the cache size limit is 30% of the total memory; Step S33: Write the entity nodes and edge relationships of the extended knowledge graph based on the graph database schema definition data to obtain the graph database mapping data; Step S34: Based on the graph database mapped data, perform syntax reconstruction of path reachability, node attribute key values, and relationship constraint structures to obtain the original expression data of the graph query language; Step S35: Perform language conversion, alias mapping, and path semantic decoupling on the original expression data of the graph query language to obtain the standardized graph query language data.

[0027] In the embodiments of the present invention, node type recognition and edge attribute parsing operations are performed on the extended knowledge graph. Node type recognition performs fixed type mapping based on the category fields to which entities belong, mapping the categories "herb entity", "efficacy entity", "indication entity", "meridian tropism entity", and "application part entity" to "HerbNode", "EffectNode", "IndicationNode", "MeridianNode", and "PartNode" respectively. Edge attribute parsing is based on the "relationship type", "semantic score", and "source path" of the edge, and is organized into an edge attribute set in a triple manner. Finally, graph structure hierarchical data is generated according to the graph depth structure. The graph structure hierarchical data includes the category of each layer of nodes, the number of nodes, and the edge connection degree with the upper and lower layer nodes. Based on the graph structure hierarchical data, dynamic loading and configuration of the graph database structure are performed. The database type is limited to the Neo4j Enterprise version that supports graph structure indexing and multi-label nodes. The configuration process includes automatic registration of node labels, initialization of edge attribute templates, and loading of index rules. A unified index mapping table is generated during the loading process for node labels and edge types. The database memory management policy is set such that the cache size is 30% of the total memory capacity. If the total memory is 64 GB, the cache size limit is 19.2 GB. After loading, graph database schema definition data is generated. Based on the graph database schema definition data, write operations are performed on all entity nodes and edge relationships in the extended knowledge graph. Node writing is performed in a batch processing manner, with 10,000 nodes and 20,000 edges processed in each batch. The transaction isolation level during the write process is READ_COMMITTED. When writing edges, the relationship type, connected node ID, edge weight, and edge attribute values are fully entered. After writing is completed, graph database mapping data is generated. The mapping data structure includes the label, unique ID, set of attribute key values of each node, and the type of the edge, start and end node IDs, and attribute list. Based on the graph database mapping data, syntax reconstruction expressions for path reachability, node attribute key value mapping, and relationship constraints are constructed. Path reachability is judged using the graph traversal method, and the maximum path length is set to 5 hops. If there is a path that meets the edge type and edge weight requirements between nodes, a path expression is generated. The node attribute key value construction rule is based on the field name mapping table corresponding to the node label. The edge relationship constraint structure is defined based on semantic template rules, such as "HerbNode - [has efficacy] -> EffectNode" must include a semantic score greater than 0.For the edge relationship of 75, after reconstruction, the original expression data of the graph query language is generated. This expression is organized using the syntax structure of the Cypher language, and the statement structure includes a MATCH clause, a WHERE filtering condition, and RETURN output fields; perform language conversion, alias mapping, and path semantic decoupling processing on the original expression data of the graph query language. In the language conversion part, the Chinese field names existing in the expression are uniformly replaced with English abbreviation fields, such as "herb name" replaced with "herb_name". The alias mapping sets abbreviation rules according to the node type. For example, HerbNode is mapped to "H", and EffectNode is mapped to "E". The path semantic decoupling operation splits the nested query path into independent sub-paths and introduces intermediate node variables, so that the query logic is independent and clear. Finally, the standardized graph query language data is generated. Each statement in this data conforms to the execution syntax of the Neo4j Cypher language. The variable naming rule uniformly adopts the format of lowercase letters plus numbers. Each statement limits the maximum query path length not to exceed 4 hops, and is accompanied by a semantic decoupling description and a field alias mapping table.

[0028] Through the recognition of node types and the analysis of edge attributes in the extended knowledge graph, the present invention constructs a clearly layered graph structure data, making the graph have high organization and hierarchy in the subsequent storage and query processes, facilitating the efficient positioning of target information; the dynamic loading and configuration process of the graph database structure driven by the graph structure layered data improves the automation of database initialization and the adaptation ability of graph data. At the same time, by limiting the cache size to 30% of the total memory, the reasonable utilization of system memory resources and the stability of query performance are ensured; during the writing process of the graph database, through the standardized mapping of node and edge structures, data consistency, transaction integrity, and writing efficiency are guaranteed; the path reachability analysis and relationship constraint reconstruction mechanism endow the query logic with precise structural semantic support, ensuring that graph query statements can be strictly constructed based on the actual semantic path, avoiding invalid path calculations; finally, through operations such as language conversion, alias mapping, and path semantic decoupling, the original query expression is standardized into a unified, clearly structured, and easy-to-execute graph query language format, thus significantly improving the graph retrieval efficiency, semantic matching ability, and query interaction flexibility of traditional Chinese medicine data, providing stable support for the accurate access and knowledge discovery of large-scale traditional Chinese medicine semantic data.

[0029] Especially importantly, the syntax reconstruction of path reachability, node attribute key values, and relationship constraint structures based on the graph database mapping data includes: Perform path structure topology parsing on the graph database mapping data to obtain node connection vector data; Based on the node connection vector data, perform path reachability evaluation on the graph database mapping data to obtain path reachability label data; Perform constraint logic rule matching on edge relation attributes based on path reachability label data to obtain relation constraint structure data; Extract node attribute key values and perform semantic mapping on the relation constraint structure data to obtain attribute semantic key value pair data; Construct a path structure syntax tree based on the attribute semantic keys to obtain the original expression data of the graph query language.

[0030] In the embodiment of the present invention, first, use the Cypher language based on the Neo4j graph database to perform path structure topology parsing on the constructed traditional Chinese medicine knowledge graph. Obtain all path structures within three hops in the graph by calling the path traversal API interface of the graph database, and combine the node IDs and edge relation orders in the path to form vectorized node connection data. Each connection vector is represented in a fixed format as [starting node ID, relation type 1, intermediate node ID, relation type 2, target node ID]; then use the depth-first search (DFS) algorithm to calculate the path connectivity of the node connection vector. There must be continuous edge connections between any two nodes in the path. If it is connected, mark the path as "reachable" to generate path reachability label data, and the label is set on the corresponding path structure using boolean values; subsequently, based on the generated reachable paths, perform logical rule matching on each edge relation attribute through a predefined traditional Chinese medicine relation constraint rule library. The rule library is stored in JSON format, and each rule contains at least three items: relation type, entity category pair, and attribute constraint. For example, if there is an edge relation of "medicinal material - contains ingredient - chemical component" in a certain path, and the component has "molecular formula" as the attribute key, it is necessary to verify whether its value matches the regular expression ^[A-Z][a-z]?[0-9]*$. Those that meet the requirements retain the edge as a legal relation to form a structured relation constraint data set; then, perform an operation to extract the attribute key values for each node in this data set. Call the Cypher query statement in the format of MATCH (n) WHERE ID(n)=node ID RETURN n.attribute name, and the extraction result is in the key value pair structure, where the attribute name is uniformly named in the graph, such as "functions and indications", "production area distribution", etc. Then, based on the extracted key value pairs, use the forward dictionary mapping technology to standardize the terms into the ontology vocabulary set. For example, "relieving the exterior" is mapped to "pharmacological action / relieving the exterior category" to form attribute semantic key value pair data; finally, adopt the syntax rule tree construction method, and use the recursive descent method to combine the path structure and attribute semantic data to generate a graph query language expression. The root node of each syntax tree is the target query entity, and the child nodes are nested with relation edges and target attribute nodes in sequence. The expression is constructed in the format of (MATCH (n1)-[:relation 1]->(n2)-[:relation 2]->(n3) WHERE n3.attribute='standard value') to complete the generation of the original expression data of the graph query language.

[0031] The topological analysis of the path structure in the present invention clarifies the connection relationship between nodes. By constructing node connection vector data, it provides a basic support for the graph structure analysis at the path level; the path reachability evaluation is based on the node connection vector data, accurately determining the reachable path set between any two nodes in the graph, thereby avoiding invalid paths from participating in the construction of the query logic and improving the calculation efficiency of the syntax structure; the constraint logic rule matching operation of the edge relationship attributes establishes a matching relationship between the edge type and the node semantics. By binding strict semantic constraint rules to each edge, the semantic constraint of the edge attributes in the graph structure is enhanced; the node attribute key-value extraction and semantic mapping process parses clear key-value pair data from the relationship constraint structure and constructs an attribute mapping system with semantic identification ability in combination with the semantic information of the knowledge graph; finally, through the construction of the path structure syntax tree based on semantic key-values, an original expression of the graph query language with clear structure, complete semantics, and consistent rules is formed, providing a structural basis and logical support for the standardization, execution, and optimization of the subsequent graph query language, and comprehensively improving the performance of the Chinese medicinal material knowledge graph in terms of the accuracy of the retrieval structure, the query response speed, and the semantic recognition ability.

[0032] Preferably, step S4 includes the following steps: Step S41: Define parametric conditions and construct logical expressions for the graph query language data to obtain query condition template data; Step S42: Combine the query condition template data, and perform priority sorting to obtain an optimized query condition set; Step S43: Perform graph data path calculation and shortest path analysis on the optimized query condition set to obtain the Chinese medicinal material knowledge path set; Step S44: Calculate the node correlation degree of the Chinese medicinal material knowledge path set, and perform influence factor evaluation to obtain node linkage relationship data; Step S45: Structure the node linkage relationship data and perform weight calculation to obtain graph structure query result data.

[0033] In the embodiments of the present invention, parameterized condition definitions are made for Cypher language query statements generated by mapping a graph database. The hard-coding rule is used to replace the constant values in the query with placeholders in the format of "$parameter name". For example, "function and indication = clearing heat and detoxifying" is transformed into "function and indication = $function and indication", and data type restrictions (such as string, integer, boolean type) and value ranges are set for the placeholder parameters. They are uniformly stored in the parameter configuration file by setting key-value pairs. Subsequently, a logical expression tree bound to the parameters is constructed. The logical operators are restricted to three categories: AND, OR, and NOT. The infix expression is used to combine the conditions, so that each condition node and its logical dependency relationship are clearly presented, and query condition template data is generated; based on the subset expansion principle in Boolean algebra, multiple query condition templates are combined and traversed to generate all legal condition permutations and combinations. Subsequently, they are sorted according to the importance priority of the conditions. The priority refers to the weight values set in the domain knowledge experience rule library. For example, "herb name" is set with a weight of 10, "chemical composition" is set with a weight of 8, and "origin" is set with a weight of 6. The optimized query condition set is obtained by sorting according to the weighted sum in all combinations; the Dijkstra algorithm is used to calculate the path distance between node pairs in the graph corresponding to each path in the optimized query condition set, and the search depth is limited to no more than 5 hops, and it is ensured that all edges in the path satisfy the entity category constraint conditions. The shortest legal paths are screened and summarized into the traditional Chinese medicine knowledge path set, where the path data structure is represented by an ordered node sequence and an edge sequence together, and the path length is measured by the number of hops and recorded in the path metadata; the co-occurrence frequency method is used to calculate the correlation degree value between any two nodes in the knowledge path set. The correlation degree value = co-occurrence times / total number of paths. The calculation range is limited to only cover node pairs with adjacent hops less than or equal to 2. Subsequently, according to the position of each node in the path, the type of connection relationship, and the path length, a weighted scoring function is used to evaluate the influence factor value of the node in the query structure. The influence factor scoring formula is: influence factor = α × connection degree + β × relationship diversity + γ × path centrality, and the parameters α, β, and γ are fixed at 0.4, 0.3, and 0.3; the node linkage relationship data is structured in a triple format as (starting node ID, ending node ID, weight value). All weight values are normalized according to the influence factor, and the Min-Max normalization method is used to normalize all weight values to the [0,1] interval. Finally, the graph structure query result data is generated, and the result data format is uniformly a JSON array. Each record contains the node ID, node name, path details, edge type, and its weight value.

[0034] The parameterized condition definition and logical expression construction in the present invention provide a general and reusable query template structure, support condition configuration for multiple types of Chinese herbal medicine information, and enhance the flexibility and adaptability of query rules; the condition combination and priority sorting operations achieve the orderly management of complex query conditions, enabling the query execution process to efficiently parse and schedule resources according to logical priorities, and improving query efficiency and accuracy; the graph data path calculation and shortest path analysis effectively avoid redundant paths and enhance the path focusing ability in knowledge graph queries through a path selection mechanism based on graph traversal algorithms; the node correlation degree calculation combines an influence factor evaluation method to construct a semantic intensity perception mechanism, quantifying the actual semantic linkage relationship between nodes in the graph and providing a data basis for graph structure reasoning and recommendation; finally, the structured expression and weight calculation of node linkage relationships provide an interpretable sorting basis for the graph query return results, ensuring that the graph structure query results have high stability and pertinence in terms of structural integrity, semantic consistency, and usability, thereby realizing high-quality retrieval of Chinese herbal medicine data and intelligent decision-making support.

[0035] Particularly importantly, the graph data path calculation and shortest path analysis for the optimized query condition set include: Performing path node matching and edge weight initialization on the optimized query condition set to obtain path candidate structure data; Selecting a graph traversal strategy and expanding paths based on the path candidate structure data to obtain path traversal result data; Performing path reachability screening and loop removal on the path traversal result data to obtain a legal path set data; Calculating the path length and performing multi-path scoring and sorting based on the legal path set data to obtain shortest path candidate set data; Performing knowledge entity relevance analysis and semantic similarity scoring on the shortest path candidate set data to obtain a Chinese herbal medicine knowledge path set.

[0036] In the embodiments of the present invention, first, path node matching and edge weight initialization operations are performed. For each parameter in the optimized query condition set, the node label matching function of the Neo4j database is called to identify entity nodes with query attributes based on the MATCH statement, and the matching node set is obtained by using the accurate matching method of attribute values. Each node is set as the starting point or ending point of the path according to the attributes. Subsequently, based on the edge attribute information established in the knowledge graph, the edge weights are initialized. The edge weights are calculated in a fixed manner: edge weight = base value 1.0 ÷ relationship weight factor, where the relationship weight factors are set as 1.0 (direct component), 1.2 (function and indication), 1.5 (place of origin), and 2.0 (drug property classification) in descending order according to the relationship semantic strength. All node and edge structures form path candidate structure data. Subsequently, based on the path candidate structure data, the heuristic A* algorithm is selected as the graph traversal strategy to implement path expansion operations in the graph database. During the expansion process, the Euclidean heuristic function is used as the basis for estimating the path cost, and the maximum expansion depth is limited to 5 hops and the maximum number of expanded paths does not exceed 100. Each hop expansion operation must comply with the entity category connection constraint rules. For example, the "Chinese medicinal material" entity can only point to the "chemical component" entity through the "contains component" edge, and the path traversal result data is extracted. Then, reachability screening and loop removal processing are performed on all the paths obtained by traversal. The reachability is determined based on the node continuous connectivity and direction legality. The breadth-first search algorithm is used to verify whether each node in the path has both incoming and outgoing edges, which is consistent with the path definition. The loop removal strategy uses the node access status marking method. Once a repeated access node appears during the traversal process, the path is removed, and the acyclic legal path set data is obtained. Subsequently, for each path in the legal path set data, path length calculation and multi-path scoring and sorting are performed. The path length is recorded in terms of the number of hops. The scoring and sorting are based on the path score function: path score = ∑(node confidence × edge weight⁻¹), where the node confidence is set by the node type. For example, "core medicinal material" is set to 0.9 and "auxiliary component" is set to 0.7. Finally, the top 10 shortest paths in terms of score are included in the shortest path candidate set data. Finally, knowledge entity relevance analysis and semantic similarity scoring are performed on the shortest path candidate set data. The relevance analysis is constructed based on the node co-occurrence frequency matrix. The higher the number of times they co-occur in the path, the stronger the relevance. The similarity scoring uses the average cosine similarity method of term word vectors. For the entity names in each path, the word vector library (limited to the ontology vocabulary table) is called to construct the average vector at the path level and calculate the semantic similarity between the path and the query condition. Finally, the paths with high relevance and high semantic similarity are screened and summarized into the Chinese medicinal material knowledge path set. This set is expressed in JSON structure, and each path record includes path ID, entity sequence, relationship sequence, path hop count, average edge weight, semantic similarity value, and relevance score.

[0037] The path node matching and edge weight initialization in the present invention ensure the exact correspondence between the query conditions and the entity nodes and relationship edges in the graph structure, and set the initial weight values according to the edge type and semantic intensity, providing a quantitative basis for subsequent path evaluation; the graph traversal strategy selection and path expansion adopt a graph search algorithm based on breadth-first or depth-first, and gradually expand the feasible paths in combination with the semantic characteristics of traditional Chinese medicines, improving the path coverage rate while controlling the expansion depth to avoid resource redundancy; the path reachability screening and loop removal perform constraint determination and structural specification on the traversal results, eliminating unreachable paths and cyclic structures to ensure the logical coherence and semantic validity of the paths; the path length calculation and multi-path scoring and sorting set the maximum path depth to 6 layers, and combine the path coverage, edge weight accumulation value and semantic integrity scoring mechanism to perform quantitative scoring and sorting screening on the path candidate set; finally, the knowledge entity relevance analysis and semantic similarity scoring integrate the calculation methods based on the cosine similarity of word vectors and the context semantic association degree, perform two-way semantic scoring on the knowledge entity pairs contained in the paths and weighted average, and screen out a set of traditional Chinese medicine knowledge paths with wide coverage, high association strength and excellent semantic consistency, so as to provide a knowledge path support with accurate paths, clear structures and clear logics for subsequent knowledge retrieval and graph reasoning.

[0038] Preferably, the semantic vector encoding of the user retrieval request in step S5 includes: Performing word segmentation and stop word filtering on the user retrieval request data to obtain a set of standardized retrieval terms; The word segmentation granularity limit in the word segmentation is a two-tuple to a four-tuple word group, and the reserved word length limit is 2-10 characters; Performing semantic classification on the set of standardized retrieval terms to obtain user retrieval intention label data, where the classification confidence threshold of the semantic classification is set to be greater than or equal to 0.75; Performing context association on the user retrieval intention label data and performing information supplementation to obtain enhanced retrieval request data, where the supplemented information amount limit is 50%-100% of the original request; Initializing a neural network model for the enhanced retrieval request data to obtain calculation-ready state data; Performing deep learning model inference on the calculation-ready state data and performing high-dimensional vector mapping to obtain initial user query semantic vector data; Performing dimensionality reduction processing and feature extraction on the initial user query semantic vector data to obtain user query semantic vector data; The dimensionality limit of the vector after dimensionality reduction in the dimensionality reduction processing is 128-256 dimensions, and the information retention rate threshold is set to 90%.

[0039] In the embodiments of the present invention, first, user retrieval request data is received, and word segmentation, cutting, and stop word filtering operations are performed on the input text. The word segmentation process adopts the custom dictionary mode of the Jieba word segmentation tool, and the word segmentation granularity range is set from bigrams to four-word phrases, with the length limit of each phrase being 2 to 10 characters. After segmentation, the common words, invalid conjunctions, and meaningless auxiliary words preset in the stop word list are removed to obtain a normalized retrieval term set; subsequently, semantic classification is performed on the normalized retrieval term set, and the keyword and semantic label mapping rule set is called. The rule set is constructed based on the ontology of the Chinese medicinal materials field and includes four types of semantic labels: "medicinal material name", "pharmacological effect", "application subject", and "adaptive symptom". The keyword matching scoring method is used to assign a confidence value to each word, and the confidence is jointly determined by the word frequency and the overlap degree of the label keyword. The classification label confidence threshold is limited to be greater than or equal to 0.75, and only the labels that meet the requirements are retained to generate user retrieval intention label data; then, context semantic association operations are performed on the retrieval intention label data, and the rule engine is used to call the knowledge base of the Chinese medicinal materials field for information supplementation according to the label type. The information supplementation amount is controlled between 50% and 100% of the original request character number. The supplemented information content includes synonyms of subordinate terms of the label, common combined words, and typical associated words, and an enhanced retrieval request data is formed by combination; then, neural network model initialization operations are performed on the enhanced retrieval request data, the pre-loaded BERT embedding model is called and preheated and loaded into the computational graph. After the graph construction and parameter weight loading are completed, the data enters the calculation-ready state; subsequently, the deep learning model inference process is performed on the loaded model, the enhanced retrieval request data text is input, the embedding representation of each Token is output through the BERT encoder, and then it is pooled into a fixed-dimensional high-dimensional vector through the pooling layer to obtain the initial user query semantic vector data, and the vector dimension is set to 768 dimensions; then, dimensionality reduction processing is performed on this vector, the principal component analysis (PCA) algorithm is used to extract the main axis of features, the dimensionality reduction target dimension is limited to be between 128 and 256 dimensions, and the lower limit of the information retention rate is set to 90% to retain the main component information and context semantic features in the input semantic structure, and finally, the user query semantic vector data that meets the dimensionality reduction requirements is output.

[0040] In the present invention, word segmentation and stop word filtering ensure that invalid terms in the original text are removed, and by setting the word segmentation granularity from bigrams to four-grams and restricting the word length range to 2 to 10 characters, the keyword extraction becomes more stable and semantically complete; the semantic classification operation classifies semantic tags for the set of normalized retrieval terms, and by setting the classification confidence threshold to be not less than 0.75, the accuracy and representativeness of the classification results are improved; the context association and information supplementation mechanism expands the information volume of the enhanced retrieval request data without exceeding the original semantic boundary, and the information supplementation ratio is controlled within the range of 50% to 100%, effectively enhancing the context background and improving the integrity of subsequent semantic expressions; the neural network model initialization operation ensures that the system inference conditions are consistent with the data input dimension, providing an environmental guarantee for efficient inference; the high-dimensional vector mapping stage uses a deep learning inference method to construct an initial semantic vector representation, extracts latent semantic features through a multi-layer network structure, and realizes the vector encoding of the user request semantics; finally, by setting the dimension of the vector after dimensionality reduction to be 128 to 256 dimensions and the information retention rate to be not less than 90%, it is ensured that the dimensionality reduction result maximally retains the original semantic information while compressing the dimension, significantly improving the operation efficiency and matching accuracy of the finally obtained user query semantic vector data in subsequent graph matching and path calculation.

[0041] Preferably, constructing the medicinal material description semantic vector set based on the enhanced data of traditional Chinese medicinal material semantic annotation in step S5 includes: Extracting the characteristic attributes of the enhanced data of traditional Chinese medicinal material semantic annotation, and performing key description screening to obtain the core description data of traditional Chinese medicinal materials; Performing word segmentation annotation and semantic unit splitting on the core description data of traditional Chinese medicinal materials to obtain normalized text units of traditional Chinese medicinal materials; The maximum length limit of the semantic unit in the semantic unit splitting is 20 characters, and the minimum length limit of the semantic unit is 2 characters; Performing parallel vector conversion on the normalized text units of traditional Chinese medicinal materials to obtain the original medicinal material description vector set, where the batch size limit of the vector model is 64 - 128; Performing dimensionality reduction processing and noise filtering on the original medicinal material description vector set to obtain the medicinal material description semantic vector set, where the energy retention ratio of singular value decomposition (SVD) is set to 85% - 95%; Using the medicinal material description semantic vector set to perform matching calculations on the user query semantic vector data to obtain vector semantic retrieval result data.

[0042] In the embodiments of the present invention, first, characteristic attributes are extracted from the enhanced semantic annotation data of traditional Chinese medicinal materials that have been constructed. The characteristic attributes include five structured fields: the functions and indications of the medicinal materials, the properties of the medicinal materials, the morphological description, the harvesting and processing, and the meridian tropism. The regular rules and part-of-speech tagging algorithms are called to extract the high-frequency keywords and phrases in each field, and the keywords with the top 10% word frequencies in each field are extracted to form a key description set. After merging, redundant phrases are removed to obtain the core description data of traditional Chinese medicinal materials. Subsequently, word segmentation annotation and semantic unit splitting operations are performed on the core description data of traditional Chinese medicinal materials. In the word segmentation stage, the maximum matching method based on the domain dictionary is used to segment the sentences into basic semantic blocks according to the part-of-speech rules. The length of the split semantic units is strictly controlled between 2 and 20 characters. For the part exceeding the maximum length, sentence breaking is performed according to punctuation and syntactic dependency relationships to obtain the standardized text units of traditional Chinese medicinal materials. Then, a parallel vector conversion operation is performed on the text units. The BERT embedding engine or an equivalent static embedding mapping function is called to batch input the standardized text data. The batch size of the vector model is limited between 64 and 128. Each text unit generates an initial description vector with a dimension of 768 after encoding, and they are summarized into the original description vector set of traditional Chinese medicinal materials. Next, dimensionality reduction processing and noise filtering are performed on the original description vector set of traditional Chinese medicinal materials. The singular value decomposition (SVD) algorithm is used to perform principal component decomposition according to the vector column dimension. The retained energy ratio is set between 85% and 95%. After removing the dimensions corresponding to the low-weight singular values, the vector set is reconstructed. At the same time, the vector dimensions with variances lower than the set threshold are zeroed to achieve noise filtering, and the semantic vector set of traditional Chinese medicinal materials descriptions is obtained. Finally, the above semantic vector set is matched and calculated with the user query semantic vector data that has been generated. The matching algorithm uses the cosine similarity matching mechanism. On the premise that the feature dimensions are the same, each traditional Chinese medicinal material description vector and the user query vector are normalized respectively, and then the dot product of each vector is performed and the cosine value of the included angle is calculated. The matching pairs with a similarity threshold set above 0.65 are retained as matching records. Finally, the sorted vector semantic retrieval result data is output in descending order of the matching score, including the traditional Chinese medicinal material identification, the description text segment, the matching score, and the affiliated semantic label.

[0043] Through the extraction of the characteristic attributes of the enhanced data of the semantic annotation of traditional Chinese medicinal materials and the screening of key descriptions, the present invention can accurately refine the representative semantic content in the medicinal material ontology; by means of word segmentation annotation and semantic unit splitting operations, the medicinal material description text is refined into basic units with clear semantic granularity, and the length of each semantic unit is restricted to be between 2 and 20 characters during splitting, so as to enhance the stability and consistency of text representation; when performing parallel vector conversion, the number of samples processed in each batch is set to be between 64 and 128 to ensure the control of computing resources and operation efficiency in the vector generation stage; subsequently, the singular value decomposition method is used to perform dimensionality reduction processing on the original traditional Chinese medicinal material description vector set, and the SVD retained energy ratio is controlled within the range of 85% to 95%, effectively compressing the vector dimension while retaining the semantic principal component information and reducing high-dimensional redundancy and vector noise interference; finally, matching calculations are performed between the user query semantic vector data and the traditional Chinese medicinal material description semantic vector set at the vector level, and the precise alignment of the user intention and the medicinal material information can be realized in the high-dimensional semantic space, significantly improving the relevance and response accuracy of semantic retrieval.

[0044] Preferably, the matching calculation of the user query semantic vector data in step S5 includes: Setting retrieval parameters for the user query semantic vector data to obtain retrieval-ready vector data, where the similarity calculation accuracy is limited to float32; Performing cosine similarity calculation on the retrieval-ready vector data and the traditional Chinese medicinal material description semantic vector set, and performing Euclidean distance evaluation based on the cosine similarity calculation result to obtain a vector similarity matrix, where the cosine similarity calculation weight coefficient is set to 0.7, the Euclidean distance weight coefficient is set to 0.3, and the Euclidean distance conversion coefficient is set to 5.0; The specific cosine similarity calculation is (q, di)=q·di / (||q|| 2 ×||di|| 2 )=q·di, the Euclidean distance is specifically (q, di)=||q - di|| 2 , the normalized Euclidean distance similarity = exp(-Euclidean distance(q, di) / Euclidean distance conversion coefficient), the vector similarity = cosine similarity(q, di)×cosine similarity weight coefficient + normalized Euclidean distance similarity×Euclidean distance weight coefficient, where q represents the retrieval-ready vector data, di represents the i-th vector in the traditional Chinese medicinal material description semantic vector set, and i is the vector index in the traditional Chinese medicinal material description semantic vector set; Performing a descending order arrangement on the vector similarity matrix and performing threshold filtering to obtain a preliminary matching result set, where the similarity threshold is set to 0.65; Performing semantic association strength calculation on the preliminary matching result set and performing context consistency evaluation on the semantic association strength calculation result to obtain a semantically weighted matching result; Perform semantic relevance clustering on the semantic weighted matching results to obtain classified retrieval result data, where the basic weight of the association strength is set to 0.6, the context consistency weight is set to 0.4, and the keyword matching weighting coefficient is set to 1.5; Remove redundancy from the classified retrieval result data to obtain an optimized vector semantic retrieval result; Perform structured encapsulation and metadata supplementation on the optimized vector semantic retrieval result to obtain vector semantic retrieval result data.

[0045] In the embodiment of the present invention, first, the generated user query semantic vector data is input into the matching engine, and retrieval parameters are set, including that the precision type of the vector data is limited to float32, and the vector dimension is limited to 128 to 256 dimensions, to obtain retrieval-ready vector data; then the retrieval-ready vector data is compared one by one with the set of Chinese herbal medicine description semantic vectors, and cosine similarity calculation and Euclidean distance evaluation are performed. The cosine similarity calculation method is to perform a dot product operation after normalizing the two vectors by the L2 norm respectively, and the expression is (q, di) = q·di / (||q|| 2 × ||di|| 2 ) = q·di, and the Euclidean distance calculation method is the L2 norm of the difference between the two vectors, that is, (q, di) = ||q - di|| 2, the Euclidean distance is normalized by the exponential function and converted into a similarity score. The formula is exp(-Euclidean distance(q,di) / 5.0). The weight coefficient of cosine similarity is 0.7, and the weight coefficient of Euclidean distance similarity is 0.3. The two are linearly weighted to obtain the comprehensive vector similarity, forming a vector similarity matrix; the vector similarity matrix is sorted in descending order according to the similarity value, and the sorting result is filtered by a threshold. The similarity lower limit is set to 0.65, and only the records with similarity values greater than or equal to 0.65 are retained to obtain the preliminary matching result set; for each group of vectors in the preliminary matching result set, the semantic association strength is calculated. The semantic association strength consists of a weighted scoring system composed of the matching keyword coincidence degree, the part-of-speech consistency rate, and the vector direction angle. The basic weight is set to 0.6; after the calculation, the context consistency is evaluated based on the context co-occurrence frequency of semantic tags. The context consistency scoring weight is set to 0.4, and it is weighted and fused with the semantic association strength to obtain the final semantic weighted matching result; then, the semantic weighted matching result is subjected to a semantic correlation clustering operation. The clustering is based on the cosine similarity proximity, the context keyword co-occurrence frequency, and the keyword matching quantity. The weighted coefficient of the keyword matching part is set to 1.5 to improve the influence of keyword coincidence on the selection of the clustering center, and finally the classified retrieval result data is obtained; for the repeated or similar description content in the classification result, the redundancy removal operation is performed. The duplication is removed by comparing the semantic tags and the keyword coverage rate, and the most representative description vector is retained to obtain the optimized vector semantic retrieval result; finally, the optimized vector semantic retrieval result is added with the original identification of traditional Chinese medicine materials, the description source, and the semantic tag metadata, and encapsulated into a structured record format to output the vector semantic retrieval result data.

[0046] The present invention ensures the balance between computing resources and numerical precision during the retrieval calculation process by setting the data type format of float32 precision; adopts a combined scoring mechanism of cosine similarity and Euclidean distance, and sets the cosine similarity weight coefficient to 0.7, the Euclidean distance weight coefficient to 0.3, and the Euclidean distance conversion coefficient to 5.0, enabling the simultaneous measurement of directional and amplitude information in the semantic space, thereby improving the discrimination and stability of similarity evaluation; after generating the vector similarity matrix, sets the lower limit of the similarity threshold to 0.65 to effectively eliminate low-correlation data and enhance the semantic accuracy of the preliminary matching result set; calculates the semantic association strength and evaluates the context consistency of the preliminary matching results, enhancing the response ability of the matching results to the actual user's retrieval intention, and achieving a multi-level fusion between semantic determination and language expression through the multi-factor control of the basic weight of association strength (0.6), the weight of context consistency (0.4), and the keyword matching weighting coefficient (1.5); further performs semantic relevance clustering on the semantically weighted matching results to achieve a classified structure display, strengthening the logical aggregation ability of similar medicinal material results; finally, through redundant data removal, structured encapsulation, and metadata supplementation, improves the expression quality and application effect of the final vector semantic retrieval results in terms of structural integrity, content readability, and system compatibility.

[0047] Preferably, step S6 includes the following steps: Step S61: Align the entity indexes of the graph structure query result data and the vector semantic retrieval result data to obtain result alignment mapping data; Step S62: Based on the result alignment mapping data, perform cross-calculation of the graph structure node attributes and the semantic vector similarity labels to obtain cross-semantic association degree data; Step S63: Perform fusion of the graph structure relationship weight and the semantic similarity score on the cross-semantic association degree data to obtain fusion score data; Step S64: Perform multi-factor sorting, i.e., confidence threshold filtering, on the fusion score data to obtain the Chinese medicinal material semantic sorting result data; Step S65: Based on the Chinese medicinal material semantic sorting result data, expand and extract the attribute information, path relationships, and similar Chinese medicinal material nodes related to the Chinese medicinal material nodes to obtain the semantic retrieval result structure data; Step S66: Standardize the format and encapsulate the interface of the semantic retrieval result structure data, and perform data caching processing to obtain the final Chinese medicinal material semantic retrieval response data.

[0048] In the embodiments of the present invention, first, the traditional Chinese medicine material node identifiers, path structures, and attribute information in the graph structure query result data are entity-index aligned with the traditional Chinese medicine material description vector index fields in the vector semantic retrieval result data. A hash mapping table is constructed by using a matching method based on the primary key field (such as the unique code of the traditional Chinese medicine material or the standard naming field) to generate result alignment mapping data; based on the alignment mapping data, the cross-similarity between the node attribute labels in the graph structure query result and the high-dimensional semantic labels in the vector semantic retrieval result is calculated. The cosine similarity between the attribute keyword vector and the description semantic vector is calculated by using the vector dot product method, and the semantic matching confidence threshold is set to 0.75 to generate cross-semantic association degree data; the cross-semantic association degree data is weighted and fused with the node relationship weight data recorded in the graph structure. The fusion method is that the semantic similarity is multiplied by 0.6 plus the structure weight value multiplied by 0.4, and it is uniformly normalized to the [0, 1] interval to form fusion score data; the fusion score data is sorted in descending order of scores, and at the same time, the confidence threshold limit is set to be greater than or equal to 0.7 for filtering to obtain the traditional Chinese medicine material semantic sorting result data; according to the high-matching nodes in the semantic sorting result, the attribute fields, path-associated nodes, and edge relationship data connected to them in the graph structure are extracted, and the adjacent traditional Chinese medicine material nodes with a cosine similarity value greater than 0.8 are extracted based on the vector distance ranking as semantic similar nodes to construct semantic retrieval result structure data including the target traditional Chinese medicine material node, attribute set, path set, and similar traditional Chinese medicine material set; the semantic retrieval result structure data is formatted and standardized, including unified field naming, formatted field types, and standardized timestamp processing. The structure is encapsulated in JSON format, and at the same time, according to the access frequency policy, the structured data is cached in the Redis database, and the cache expiration period is set to 300 seconds to generate the final traditional Chinese medicine material semantic retrieval response data.

[0049] Through entity index alignment operations, the present invention realizes the unique mapping relationship between entity nodes in the graph structure and semantic vectors in the vector space, ensuring the consistency and correspondence of cross-modal data; the calculation of cross-semantic correlation degree data integrates the node attributes and semantic similarity labels in the graph structure, providing a combined scoring basis for subsequent sorting; in the fusion scoring stage, the edge relationship weights between nodes in the graph structure and the semantic vector matching scores are weighted and synthesized, enhancing the structural dependence ability and context awareness ability of semantic scoring, and effectively improving the interpretability and stability of the sorting process; in multi-factor sorting, a confidence threshold filtering mechanism is introduced, and by setting a lower limit value for sorting confidence, irrelevant items with weak structural associations and low semantic scores are filtered out, further enhancing the accuracy of the sorting results; based on the sorting results, the attribute information, path relationships, and similar nodes of traditional Chinese medicine material nodes are expanded and extracted, not only enriching the semantic output content, but also enhancing the reasoning ability and recommendation ability of the retrieval system; finally, the semantic retrieval result structure data is processed for format specification and interface encapsulation, improving the integration efficiency of the output results during system calls. Combined with the data caching processing strategy, it can effectively reduce the computational resource overhead and response time of repeated queries, and improve the system operation efficiency and service quality.

[0050] Preferably, the present invention further provides a traditional Chinese medicine material data management system for data retrieval, which is used to execute the above-mentioned traditional Chinese medicine material data management method for data retrieval. The traditional Chinese medicine material data management system for data retrieval includes: A traditional Chinese medicine information semantic annotation module, which is used to obtain the morphological data of traditional Chinese medicine material samples; conduct grade evaluation based on the morphological data of traditional Chinese medicine materials to generate traditional Chinese medicine material grade label data; obtain the image of the traditional Chinese medicine material to be tested; preprocess the image of the traditional Chinese medicine material to be tested, and conduct entity recognition and relationship extraction to obtain traditional Chinese medicine material semantic annotation data; fuse the traditional Chinese medicine material grade label data and the traditional Chinese medicine material semantic annotation data to obtain enhanced traditional Chinese medicine material semantic annotation data; A medicinal material knowledge graph construction and reasoning module, which is used to construct an initial knowledge graph using the enhanced traditional Chinese medicine material semantic annotation data, and conduct logical rule matching and implicit relationship reasoning on the initial knowledge graph to obtain an extended knowledge graph; A graph database mapping module, which is used to map the extended knowledge graph to the graph database structure to generate standardized graph query language data; A graph structure query module, which is used to construct a graph query for conditional combination, path traversal, and node linkage based on the graph query language data to obtain graph structure query result data; A user semantic matching retrieval module, which is used to obtain user retrieval request data; encode the user retrieval request data into a semantic vector to obtain user query semantic vector data; construct a traditional Chinese medicine material description semantic vector set based on the enhanced traditional Chinese medicine material semantic annotation data, and conduct matching calculations on the user query semantic vector data to obtain vector semantic retrieval result data; The retrieval result fusion response module is used to fuse and compare the graph structure query result data and the vector semantic retrieval result data to obtain the final traditional Chinese medicine semantic retrieval response data.

[0051] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is not limited by the above description. Therefore, it is intended to cover all changes falling within the meaning and scope of the equivalent elements of the application documents within the present invention.

[0052] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will conform to the widest scope consistent with the principles and novel features invented herein.

Claims

1. A method for managing Chinese medicinal materials data based on data retrieval, characterized in that: The following steps are involved: Step S1: Acquire the morphological data of Chinese medicinal materials samples; perform grade assessment based on the Chinese medicinal materials morphological data to generate Chinese medicinal materials grade label data; acquire the Chinese medicinal materials images to be tested; pre-process the Chinese medicinal materials images to be tested, and perform entity recognition and relationship extraction to obtain Chinese medicinal materials semantic annotation data; fuse the Chinese medicinal materials grade label data and the Chinese medicinal materials semantic annotation data to obtain Chinese medicinal materials semantic annotation enhanced data; Step S2: Use the semantic annotation enhancement data of Chinese medicinal materials to construct an initial knowledge graph, perform logical rule matching and implicit relationship reasoning on the initial knowledge graph, and obtain an extended knowledge graph; Step S3: Map the extended knowledge graph to a graph database structure and generate standardized graph query language data; Step S4: constructing a graph query by condition combination, path traversal and node linkage based on the graph query language data to obtain graph structure query result data; Step S5: obtaining user search request data; performing semantic vector encoding on the user search request data to obtain user query semantic vector data; constructing a Chinese medicinal material description semantic vector set based on the Chinese medicinal material semantic annotation enhanced data, and performing matching calculation on the user query semantic vector data to obtain vector semantic search result data; Step S6: Fusion and comparison are performed on the graph structure query result data and the vector semantic retrieval result data to obtain the final Chinese medicinal material semantic retrieval response data.

2. The Chinese medicinal material data management method based on data retrieval according to claim 1 is characterized in that: Step S1 includes the following steps: Step S11: collecting the appearance image of the Chinese medicinal material through microscopic photography, collecting the fiber structure image and surface texture atlas data through a high-definition multispectral imager, and recording the appearance image, fiber structure image and surface texture atlas data as the Chinese medicinal material sample morphology data; collecting the intrinsic component characteristic data of the Chinese medicinal material through a texture analyzer; Step S12: performing feature fusion on the Chinese medicinal material sample morphology data and the Chinese medicinal material internal component feature data, and performing redundancy removal to construct a Chinese medicinal material morphology-component joint feature matrix; Step S13: Classify the Chinese medicinal materials using classification rules based on the Chinese medicinal materials morphology-ingredient joint feature matrix to generate Chinese medicinal materials grade label data; Step S14: obtaining a handwritten image of Chinese medicinal materials; performing image preprocessing, image enhancement and text recognition on the handwritten image of Chinese medicinal materials to obtain structured Chinese medicinal materials data; Step S15: performing prediction recognition and content structured annotation on the structured Chinese medicinal material data to obtain Chinese medicinal material text parsing data; Step S16: performing named entity recognition and term boundary segmentation on the Chinese medicinal materials text parsing data to obtain Chinese medicinal materials semantic entity data; Step S17: Performing context-aware analysis based on the Chinese medicinal material semantic entity data to obtain Chinese medicinal material semantic relationship data; Step S18: Merging the data structure of the Chinese medicinal material semantic entity data and the Chinese medicinal material semantic relationship data, and uploading them to the edge server to obtain the Chinese medicinal material semantic annotation data; Step S19: Fusing the Chinese medicinal material grade label data and the Chinese medicinal material semantic annotation data to obtain Chinese medicinal material semantic annotation enhanced data.

3. The Chinese medicinal material data management method based on data retrieval according to claim 1 is characterized in that: Step S2 includes the following steps: Step S21: performing field analysis and entity classification on the Chinese medicinal materials semantic annotation enhanced data to obtain candidate data of Chinese medicinal materials graph nodes; Step S22: Based on the candidate data of the Chinese medicinal material graph nodes, semantic relationship mapping and edge type definition are performed between nodes to obtain structured knowledge graph data; Step S23: performing data deduplication, semantic consistency processing and standard graph model mapping on the structured knowledge graph data to obtain an initial knowledge graph; Step S24: traverse the graph structure and filter the rule nodes of the initial knowledge graph, and use the preset logical reasoning rule set to match the symbolic rules to obtain the rule reasoning path data; Step S25: Based on the rule reasoning path data, the semantic propagation of the graph node embedding vector and the attention mechanism weight update are performed to obtain the graph reasoning embedding data; Step S26: perform similarity clustering on the graph reasoning embedding data, and perform upstream and downstream path prediction to obtain reasoning relationship generation data; Step S27: Structurally merge the initial knowledge graph and the inference relationship generation data, and redistribute the weights to obtain extended knowledge graph data.

4. The Chinese medicinal material data management method based on data retrieval according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: Identify the node types and analyze the edge attributes of the extended knowledge graph to obtain graph structure hierarchical data; Step S32: dynamically load and configure the graph database structure based on the graph structure hierarchical data to obtain graph database schema definition data, where the cache size is limited to 30% of the total memory; Step S33: writing entity nodes and edge relationships into the extended knowledge graph based on the graph database schema definition data to obtain graph database mapping data; Step S34: Perform grammatical reconstruction of path reachability, node attribute key values, and relationship constraint structures based on the graph database mapping data to obtain graph query language original expression data; Step S35: Perform language conversion, alias mapping and path semantic decoupling on the original graph query language expression data to obtain standardized graph query language data.

5. The Chinese medicinal material data management method based on data retrieval according to claim 1 is characterized in that: Step S4 includes the following steps: Step S41: performing parameterized condition definition and logical expression construction on graph query language data to obtain query condition template data; Step S42: combining the query condition template data and sorting the priority to obtain an optimized query condition set; Step S43: performing graph data path calculation and shortest path analysis on the optimized query condition set to obtain a Chinese medicinal material knowledge path set; Step S44: Calculate the node association degree of the Chinese medicinal material knowledge path set, and evaluate the impact factor to obtain node linkage relationship data; Step S45: Structuring the node linkage relationship data and performing weight calculation to obtain graph structure query result data.

6. The Chinese medicinal material data management method based on data retrieval according to claim 1 is characterized in that: The step S5 of encoding the user's search request with a semantic vector includes: Perform word segmentation and stop word filtering on user search request data to obtain a standardized search word set; The word segmentation granularity in the word segmentation is limited to two-tuple to four-tuple phrases, and the length of reserved words is limited to 2-10 characters; Perform semantic classification on the standardized search term set to obtain user search intent label data, where the classification confidence threshold of the semantic classification is set to be greater than or equal to 0.75; Contextualize the user's search intent label data and supplement the information to obtain enhanced search request data, where the amount of supplementary information is limited to 50%-100% of the original request; Initializing the neural network model for the enhanced retrieval request data to obtain computing ready state data; Perform deep learning model inference on the computing-ready state data and perform high-dimensional vector mapping to obtain the initial user query semantic vector data; Performing dimensionality reduction and feature extraction on the initial user query semantic vector data to obtain user query semantic vector data; The dimension of the vector after dimensionality reduction in the dimensionality reduction process is limited to 128-256 dimensions, and the information retention rate threshold is set to 90%.

7. The Chinese medicinal material data management method based on data retrieval according to claim 1 is characterized in that: The step S5 of constructing a set of Chinese medicinal material description semantic vectors based on Chinese medicinal material semantic annotation enhanced data includes: Extract feature attributes from the semantic annotation enhanced data of Chinese herbal medicines, and screen key descriptions to obtain the core description data of Chinese herbal medicines; The core description data of Chinese herbal medicines is segmented and annotated and the semantic units are split to obtain the standardized Chinese herbal medicine text units; The maximum length of the semantic unit in the semantic unit splitting is limited to 20 characters, and the minimum length of the semantic unit is limited to 2 characters; Perform parallel vector conversion on the standardized Chinese herbal medicine text units to obtain the original Chinese herbal medicine description vector set, where the vector model batch size is limited to 64-128; The original Chinese herbal medicine description vector set is subjected to dimensionality reduction and noise filtering to obtain a set of Chinese herbal medicine description semantic vectors, where the singular value decomposition (SVD) energy retention ratio is set to 85%-95%; The semantic vector set of Chinese herbal medicine description is used to perform matching calculations on the user query semantic vector data to obtain vector semantic retrieval result data.

8. The method for managing Chinese medicinal materials data based on data retrieval according to claim 1, characterized in that: The matching calculation of the user query semantic vector data in step S5 includes: Set the retrieval parameters for the user query semantic vector data to obtain the retrieval-ready vector data, where the similarity calculation accuracy is limited to float32; The cosine similarity calculation is performed on the retrieval-ready vector data and the set of semantic vectors of Chinese herbal medicine descriptions, and the Euclidean distance evaluation is performed based on the cosine similarity calculation results to obtain a vector similarity matrix, in which the cosine similarity calculation weight coefficient is set to 0.7, the Euclidean distance weight coefficient is set to 0.3, and the Euclidean distance conversion coefficient is set to 5.0; The cosine similarity calculation is specifically (q, d i ) = q·d i / (||q||2×||d i ||2)=q·d i , the Euclidean distance is specifically (q, d i )=||qd i ||2, normalized Euclidean distance similarity = exp(-Euclidean distance(q,d i ) / Euclidean distance conversion coefficient), vector similarity = cosine similarity (q, d i )×cosine similarity weight coefficient+normalized Euclidean distance similarity×Euclidean distance weight coefficient, where q represents the retrieval-ready vector data, d i represents the i-th vector in the set of semantic vectors describing Chinese medicinal materials, where i is the vector index in the set of semantic vectors describing Chinese medicinal materials; The vector similarity matrix is ​​sorted in descending order and threshold filtered to obtain a preliminary matching result set, where the similarity threshold is set to 0.65; Calculating the semantic association strength of the preliminary matching result set, and evaluating the context consistency of the semantic association strength calculation result to obtain a semantic weighted matching result; The semantic weighted matching results are clustered by semantic relevance to obtain classified retrieval result data, where the basic weight of association strength is set to 0.6, the context consistency weight is set to 0.4, and the keyword matching weight coefficient is set to 1.5; Remove redundancy from the classified search result data to obtain optimized vector semantic search results; The optimized vector semantic retrieval results are structured and encapsulated and metadata is supplemented to obtain vector semantic retrieval result data.

9. The data retrieval method for Chinese medicinal materials data management according to claim 1, characterized in that: Step S6 includes the following steps: Step S61: aligning the graph structure query result data with the vector semantic retrieval result data by entity index to obtain result alignment mapping data; Step S62: performing cross calculation of graph structure node attributes and semantic vector similarity labels based on the result alignment mapping data to obtain cross semantic association data; Step S63: fusing the graph structure relationship weight and the semantic similarity score of the cross-semantic association data to obtain fused score data; Step S64: performing multi-factor sorting, i.e., confidence threshold filtering, on the fused scoring data to obtain Chinese medicinal materials semantic sorting result data; Step S65: based on the Chinese medicinal materials semantic sorting result data, the attribute information, path relationship and similar Chinese medicinal materials nodes related to the Chinese medicinal materials nodes are extended and extracted to obtain the semantic search result structure data; Step S66: standardize the format and interface encapsulate the semantic search result structure data, and perform data cache processing to obtain the final Chinese medicinal material semantic search response data.

10. A data retrieval system for Chinese medicinal materials data, characterized in that: The Chinese medicinal material data management method for executing the data retrieval as claimed in claim 1, wherein the Chinese medicinal material data management system for data retrieval comprises: The Chinese medicine information semantic annotation module is used to obtain the morphological data of Chinese medicine samples; perform grade evaluation based on the Chinese medicine morphological data to generate Chinese medicine grade label data; obtain the image of the Chinese medicine to be tested; pre-process the image of the Chinese medicine to be tested, and perform entity recognition and relationship extraction to obtain the semantic annotation data of the Chinese medicine; fuse the Chinese medicine grade label data and the Chinese medicine semantic annotation data to obtain the Chinese medicine semantic annotation enhanced data; The medicinal material knowledge graph construction and reasoning module is used to build the initial knowledge graph using the Chinese medicinal material semantic annotation enhanced data, perform logical rule matching and implicit relationship reasoning on the initial knowledge graph, and obtain an extended knowledge graph; The graph database mapping module is used to map the extended knowledge graph to the graph database structure and generate standardized graph query language data; The graph structure query module is used to construct graph queries based on graph query language data by combining conditions, traversing paths, and linking nodes to obtain graph structure query result data; The user semantic matching retrieval module is used to obtain user search request data; perform semantic vector encoding on the user search request data to obtain user query semantic vector data; construct a Chinese medicinal material description semantic vector set based on the Chinese medicinal material semantic annotation enhanced data, and perform matching calculation on the user query semantic vector data to obtain vector semantic retrieval result data; The retrieval result fusion response module is used to fuse and compare the graph structure query result data with the vector semantic retrieval result data to obtain the final Chinese medicinal material semantic retrieval response data.

Citation Information

Patent Citations

  • Knowledge inference and fault diagnosis method based on knowledge graph

    CN114756686A

  • Multi-modal tourism information positioning type retrieval method based on tourism knowledge graph

    CN115827881A

  • Rule and path-based traditional Chinese medicine multi-modal knowledge graph reasoning method and device

    CN116705338A

  • Modeling method and system for diversified retrieval of traditional Chinese medicinal materials based on knowledge graph

    CN118820532A

  • Knowledge reasoning method, system and device for multivariate relation scene and medium

    CN119358689A

Cited By

  • Pharmaceutical quality traceability decision-making method and system based on dynamic knowledge graph

    CN120317535A

  • Enterprise information generation and retrieval method based on AI and knowledge graph

    CN120578742A

  • Target keyword-based JSON data dynamic addition and management method

    CN120744117A

  • Method for dynamically appending and managing JSON data based on target keywords

    CN120744117B

  • Knowledge base retrieval method fused with natural language large model

    CN121168677A