Traditional Chinese medicine whole industry chain big data construction method and system

By constructing a decentralized consortium blockchain network and graph neural network in the traditional Chinese medicine (TCM) industry chain, combined with knowledge graphs, the problem of data silos in the entire TCM industry chain has been solved. This has enabled cross-link data fusion and management, improved data transparency and credibility, and enabled full-chain digital traceability and risk warning, as well as optimization of process parameters.

CN120952701APending Publication Date: 2025-11-14XIAN KUNTENG PHARM CO LTD

Patent Information

Application Number
CN202511078257.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing methods and systems for constructing big data across the entire Chinese medicine industry chain suffer from data silos, making it impossible to achieve effective cross-link data fusion and analysis. Data storage and management efficiency is low, blockchain applications are insufficient, trust mechanisms and privacy protection cannot be addressed, and the level of intelligence is low, making it difficult to achieve precise process optimization and quality traceability.

Method used

By designing data structuring transformation rules, constructing a decentralized consortium blockchain network, and combining graph neural networks and knowledge graphs, cross-stage data association and fusion are achieved. A hybrid collaborative storage architecture is adopted for efficient management, and risk warning is based on LSTM time series analysis.

Benefits of technology

It has achieved full-chain digital traceability of the Chinese medicine industry chain, improved the transparency and credibility of data, enabled precise control of quality and safety, predicted potential risks in advance, and carried out effective process optimization and closed-loop management of quality feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952701A_ABST
    Figure CN120952701A_ABST
Patent Text Reader

Abstract

The invention discloses a traditional Chinese medicine whole industry chain big data construction method and system, and relates to the field of data links, and the method comprises the steps: generating a standardized metadata template for seven links of a traditional Chinese medicine industry chain; block chain nodes are deployed in each link of an industrial chain, timestamp fingerprints are generated for key process data, and the key process data are automatically chained for evidence storage; constructing a data fusion model based on a graph neural network, and generating cross-link associated data; constructing a traditional Chinese medicine ontology knowledge graph, and mapping the association relationship between the cross-link associated data in the graph; full-chain digital tracing from implantation to clinic is realized through a path reasoning algorithm; based on the hybrid collaborative storage architecture, access management is carried out through dynamic fragmentation strategy data points; and constructing a process prediction model, detecting process parameter deviation, and carrying out risk early warning. The method has the advantages that through block chain evidence storage and combination of the graph neural network and the knowledge graph, full-chain digital tracing and risk early warning from implantation to clinic are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data chains, and in particular to a method and system for constructing big data across the entire traditional Chinese medicine industry chain. Background Technology

[0002] With the rapid development of the traditional Chinese medicine (TCM) industry, traditional Chinese medicine faces numerous challenges in its modernization process. These include low levels of informatization, standardization, and regulation in the planting, processing, distribution, and sales of medicinal materials, leading to uneven resource allocation and inefficient data utilization within the industry chain. By collecting, integrating, and analyzing large amounts of data from various stages of the TCM industry, including information on the planting environment, active ingredients, quality standards, and supply chain flow of medicinal materials, enterprises can achieve precise management, optimize resource allocation, improve product quality, and promote the modernization and internationalization of TCM.

[0003] Current methods and systems for building big data across the entire TCM industry chain often face the problem of data silos, hindering effective interconnection between different stages. Many systems lack unified, standardized data processing workflows, resulting in inconsistent data formats, sources, and quality, making accurate cross-stage data fusion and analysis difficult. Furthermore, existing systems suffer from bottlenecks in data storage and management, struggling to efficiently process massive amounts of time-series, graph, and relational data, leading to slow data access speeds and impacting real-time decision-making and risk warning effectiveness. While some systems attempt to use blockchain technology to ensure data security, most blockchain applications in the TCM industry chain remain in their early stages, failing to effectively address issues such as trust mechanisms, data sharing, and privacy protection. Moreover, existing methods have relatively low levels of intelligence in process parameter prediction and closed-loop quality feedback management, failing to fully leverage the potential of big data and hindering precise process optimization and quality traceability. Summary of the Invention

[0004] To improve existing methods and systems, this method provides a big data construction method and system for the entire Chinese medicine industry chain. By combining data standardization, blockchain notarization, graph neural networks and knowledge graphs, this method realizes full-chain digital traceability and risk warning from planting to clinical practice, thereby improving the transparency, quality management and intelligence level of the Chinese medicine industry.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] The methodology for constructing big data across the entire traditional Chinese medicine industry chain includes:

[0007] For each of the seven links in the Chinese medicinal materials industry chain, structured data transformation rules were designed to generate standardized metadata templates;

[0008] Deploy blockchain nodes at each stage of the industrial chain to build a decentralized consortium blockchain network for evidence storage, generate timestamp fingerprints for key process data and automatically upload them to the blockchain for evidence storage via smart contracts;

[0009] Based on data from different stages on each blockchain node, a data fusion model based on graph neural networks is constructed through feature extraction to correlate data between different stages and generate cross-stage correlated data.

[0010] Based on local processing standards data and historical materia medica literature data, a knowledge graph of Chinese medicine is constructed to map the relationships between cross-linked data in the graph.

[0011] Based on the industrial chain data in the knowledge graph of traditional Chinese medicine, a path reasoning algorithm is used to achieve full-chain digital traceability from planting to clinical application.

[0012] Based on a hybrid collaborative storage architecture of time-series databases, graph databases and relational databases, it efficiently manages the access of billions of data points through a dynamic sharding strategy.

[0013] A process prediction model is built based on planting, climate, and sales data from the industry chain. LSTM time series analysis is used to detect process parameter deviations and provide risk warnings.

[0014] Preferably, the design of structured data transformation rules for each of the seven major links in the Chinese medicinal materials industry chain, and the generation of standardized metadata templates, specifically includes:

[0015] The seven major links in the Chinese medicinal materials industry chain include planting and breeding, harvesting and processing, processing and production, warehousing and logistics, preparation production, clinical use and market circulation.

[0016] Structured transformation rules were designed for sensor data, image and video data, process parameter data, quality inspection data, clinical electronic medical record data, and market transaction data based on seven key aspects.

[0017] The structured transformation rules include: extracting morphological features of medicinal materials from image and video data using the YOLOv5 model and converting them into 128-dimensional feature vectors; parsing equipment logs using the OPC UA protocol to generate quadruple data for process parameter data; and using ICD-11 standard coding for syndrome differentiation and classification of clinical data, and grading and quantifying efficacy feedback.

[0018] The data is transformed according to structured transformation rules to generate standardized metadata templates.

[0019] Preferably, the step of deploying blockchain nodes at each stage of the industry chain, constructing a decentralized consortium blockchain network for evidence storage, generating timestamp fingerprints for key process data, and automatically uploading the data to the blockchain for evidence storage via smart contracts specifically includes:

[0020] Blockchain nodes are deployed across all links of the Chinese medicinal herb industry chain, and trust relationships and access permissions are established between nodes.

[0021] By designing a consensus mechanism and data sharing rules for the consortium blockchain, data can be shared among the nodes.

[0022] The data of each key process in each link of the Chinese medicinal materials industry chain is timestamped, and the fingerprint of the data is generated by hash function to verify the integrity and immutability of the data.

[0023] Design smart contracts to define the conditions and processes for uploading data to the blockchain, and automatically package and upload the data to the blockchain for evidence storage based on the smart contracts.

[0024] Preferably, the step of constructing a data fusion model based on graph neural networks through feature extraction, based on data from different stages on each blockchain node, and generating cross-stage related data specifically includes:

[0025] Each link and data item in the industrial chain is regarded as a node in the graph. Edges are constructed based on the logical relationships and data correlations between the links in the industrial chain to build a graph model.

[0026] Feature extraction is performed on each node, and the graph model is trained using data from each stage based on a graph neural network model.

[0027] By using the node embedding results of graph neural networks, data from different stages are associated and fused, and associated data and the relationships between them are generated.

[0028] Preferably, the step of constructing a knowledge graph of traditional Chinese medicine based on local processing standards data and historical herbal literature data, and mapping the relationships between cross-process related data in the graph, specifically includes:

[0029] Collect local processing technology data, including processing methods, traditional processing techniques, time and temperature control, and label and classify the processing methods with corresponding medicinal materials, processing conditions, timing of use, and precautions;

[0030] We organize the descriptions of medicinal materials in historical herbal literature, including their properties, uses, effects, and compatibility. We extract information about medicinal materials from the literature, integrate different descriptions of medicinal materials, remove duplicates, and standardize them.

[0031] A knowledge graph of traditional Chinese medicine is constructed, with nodes including medicinal materials, processing methods, literature records, pharmacological properties, and usage routes, and edges representing the relationships between nodes;

[0032] The nodes in the cross-process related data are matched with the nodes in the Chinese medicine ontology knowledge graph, and the association relationships of the cross-process related data are mapped to the Chinese medicine ontology knowledge graph.

[0033] Preferably, the process of achieving full-chain digital traceability from planting to clinical application based on the industry chain data in the knowledge graph of traditional Chinese medicine ontology, through path reasoning algorithms, specifically includes:

[0034] Based on the industry chain data in the knowledge graph of traditional Chinese medicine, starting from the planting and breeding link and ending at the clinical use link, the data path and relationship are marked.

[0035] By using an improved Dijkstra algorithm, multiple weights are added to each edge, and the weights are selected according to different needs when solving the path.

[0036] Based on the dependencies between different stages, a complete tracing path is formed by using an improved Dijkstra algorithm, and the optimal path is displayed in the graph.

[0037] The improved Dijkstra algorithm dynamically adjusts the path based on new data and periodically updates the path reasoning model.

[0038] Preferably, the hybrid collaborative storage architecture based on time-series databases, graph databases, and relational databases, which efficiently manages access to billions of data points through a dynamic sharding strategy, specifically includes:

[0039] Separate databases are established for storing time-series data, graph data, and relational data, and a hybrid collaborative storage architecture is constructed.

[0040] Time series databases are sharded according to time range; graph databases are sharded by nodes or edges, and sharded according to the subgraphs of the graph; relational databases are sharded by primary key hash or range sharding.

[0041] Based on the access frequency and update time characteristics of the data, hot data and cold data are managed separately. Active time series data are stored in fast-access storage, while historical cold data is archived to low-cost storage.

[0042] By establishing connection bridges between different databases, cross-database joint queries can be constructed.

[0043] Preferably, the step of constructing a process prediction model based on planting data, climate data, and sales data from the industry chain data, and using LSTM time series analysis to detect process parameter deviations and conduct risk warnings specifically includes:

[0044] Feature extraction is performed based on planting data, climate data, and sales data in the industrial chain data, and composite features are constructed by combining different data sources for model training;

[0045] The model is trained using a long short-term memory network, with the input layer containing multiple features, including key node data in the phenological and planting processes.

[0046] Based on the trained LSTM model, real-time data is input into the model to predict process parameters.

[0047] By comparing actual process parameter data with predicted results, offset detection is performed, and risk warnings are issued for offsets exceeding thresholds.

[0048] Preferably, a closed-loop quality feedback mechanism is added: clinical adverse reaction data are collected, and the production batches of the formulation are traced back; Bayesian networks are used to calculate the probability of defects in each step, generate process improvement plans, and update the standard system.

[0049] Furthermore, a big data construction system for the entire traditional Chinese medicine industry chain is proposed, including:

[0050] Data structuring and transformation module: This module designs structuring and transformation rules and generates standardized metadata templates based on the data characteristics of each link in the industry chain;

[0051] Blockchain Evidence Preservation Module: This module is used to deploy blockchain nodes, build a decentralized consortium blockchain network for evidence preservation, and ensure the immutability and timestamp fingerprint evidence preservation of key process data.

[0052] Data fusion and association module: This module is based on a graph neural network model and uses feature extraction to achieve the association and fusion of cross-stage data, generating cross-stage associated data;

[0053] Knowledge Graph Module: This module is used to construct a knowledge graph of traditional Chinese medicine, mapping and integrating historical documents, processing standards, and cross-process related data.

[0054] Chain traceability reasoning module: The module is based on the industrial chain data in the knowledge graph of traditional Chinese medicine, and uses path reasoning algorithm to realize full-chain digital traceability from planting to clinical use;

[0055] Hybrid collaborative storage module: The module is used to establish a hybrid collaborative storage architecture for time series data, graph data, and relational databases, and to achieve efficient data management through a dynamic sharding strategy;

[0056] Process prediction and risk warning module: Based on industry chain data, the module predicts process parameters through LSTM time series analysis and performs offset detection and risk warning;

[0057] Processor: The processor is used to handle the calculation process of each formula and the construction calculation process of each model.

[0058] Compared with the prior art, the advantages of the present invention are:

[0059] By standardizing data at each stage through structured transformation rules, data silos are effectively broken down, enabling cross-stage data fusion and sharing. Secondly, the application of blockchain technology ensures data immutability and security, allowing for real-time monitoring and precise recording of each process, enhancing data credibility. Combined with innovative applications of graph neural networks and knowledge graphs, this method can uncover deep-seated relationships across the entire supply chain from cultivation to clinical application, facilitating precise control over the quality and safety of Chinese medicinal materials. Furthermore, through path reasoning and time-series analysis, potential risks can be predicted in advance, enabling effective process optimization and closed-loop quality feedback management. Attached Figure Description

[0060] Figure 1 This is a schematic diagram of the method proposed in this invention;

[0061] Figure 2 This is a schematic diagram of the standardized metadata template proposed in this invention;

[0062] Figure 3 This is a schematic diagram of the construction of a consortium blockchain network proposed in this invention;

[0063] Figure 4 This is a schematic diagram illustrating the generation of cross-stage related data proposed in this invention;

[0064] Figure 5 This is a schematic diagram illustrating the construction of a knowledge graph of traditional Chinese medicine as proposed in this invention;

[0065] Figure 6 This is a schematic diagram of the full-chain digital traceability proposed in this invention;

[0066] Figure 7 This is a schematic diagram of the data access management proposed in this invention;

[0067] Figure 8 This is a schematic diagram of the construction process prediction model proposed in this invention. Detailed Implementation

[0068] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.

[0069] The big data construction system for the entire Chinese medicine industry chain includes:

[0070] Data structuring and transformation module: This module designs structuring and transformation rules and generates standardized metadata templates based on the data characteristics of each link in the industry chain;

[0071] Blockchain Evidence Preservation Module: This module is used to deploy blockchain nodes, build a decentralized consortium blockchain network for evidence preservation, and ensure the immutability and timestamp fingerprint evidence preservation of key process data.

[0072] Data fusion and association module: This module is based on a graph neural network model and uses feature extraction to achieve the association and fusion of cross-stage data, generating cross-stage associated data;

[0073] Knowledge Graph Module: This module is used to construct a knowledge graph of traditional Chinese medicine, mapping and integrating historical documents, processing standards, and cross-process related data.

[0074] Chain traceability reasoning module: The module is based on the industrial chain data in the knowledge graph of traditional Chinese medicine, and uses path reasoning algorithm to realize full-chain digital traceability from planting to clinical use;

[0075] Hybrid collaborative storage module: The module is used to establish a hybrid collaborative storage architecture for time series data, graph data, and relational databases, and to achieve efficient data management through a dynamic sharding strategy;

[0076] Process prediction and risk warning module: Based on industry chain data, the module predicts process parameters through LSTM time series analysis and performs offset detection and risk warning;

[0077] Processor: The processor is used to handle the calculation process of each formula and the construction calculation process of each model.

[0078] See Figure 1 As shown, the method for constructing big data across the entire traditional Chinese medicine industry chain includes:

[0079] Step 1: Design structured data transformation rules for each of the seven links in the Chinese medicinal materials industry chain, and generate standardized metadata templates;

[0080] Step 2: Deploy blockchain nodes at each stage of the industry chain to build a decentralized consortium blockchain network for evidence storage, generate timestamp fingerprints for key process data and automatically upload them to the blockchain for evidence storage via smart contracts;

[0081] Step 3: Based on the data from different stages on each blockchain node, construct a data fusion model based on graph neural networks through feature extraction, and perform correlation processing on the data between different stages to generate cross-stage correlation data;

[0082] Step 4: Based on local processing standards data and historical herbal literature data, construct a knowledge graph of traditional Chinese medicine ontology, and map the relationships between cross-linked data in the graph;

[0083] Step 5: Based on the industry chain data in the knowledge graph of traditional Chinese medicine, realize the digital traceability of the entire chain from planting to clinical application through path reasoning algorithm;

[0084] Step Six: Based on a hybrid collaborative storage architecture of time-series databases, graph databases, and relational databases, efficient access management of billions of data points is achieved through dynamic sharding strategies;

[0085] Step 7: Based on planting data, climate data, and sales data in the industry chain data, construct a process prediction model, and use LSTM time series analysis to detect process parameter deviations and conduct risk warnings.

[0086] See Figure 2 As shown, structured data transformation rules are designed for each of the seven major links in the Chinese medicinal materials industry chain, generating standardized metadata templates, specifically including:

[0087] The seven major links in the Chinese medicinal materials industry chain include planting and breeding, harvesting and processing, processing and production, warehousing and logistics, preparation production, clinical use and market circulation.

[0088] Structured transformation rules were designed for sensor data, image and video data, process parameter data, quality inspection data, clinical electronic medical record data, and market transaction data based on seven key aspects.

[0089] The structured transformation rules include: extracting morphological features of medicinal materials from image and video data using the YOLOv5 model and converting them into 128-dimensional feature vectors; parsing equipment logs using the OPC UA protocol to generate quadruple data for process parameter data; and using ICD-11 standard coding for syndrome differentiation and classification of clinical data, and grading and quantifying efficacy feedback.

[0090] The data is transformed according to structured transformation rules to generate standardized metadata templates.

[0091] Specifically, the image and video data are used to identify the morphological features of medicinal materials using the YOLOv5 model, and converted into 128-dimensional feature vectors to represent the basic morphological features of medicinal materials, such as color and shape, and then converted into image feature fields in the standardized metadata template.

[0092] Process parameter data is parsed from equipment logs using the OPC UA protocol to extract relevant process data, and each equipment log is converted into a standardized four-tuple of data.

[0093] Clinical electronic medical record data uses the ICD-11 standard to encode symptoms, classifies them according to syndrome differentiation, and generates standardized clinical data templates.

[0094] Quality inspection data can be obtained through sensors to acquire indicators such as the freshness, humidity, moisture content, dryness, and effective ingredient content of medicinal materials. After being converted into standardized templates, it can be used for quality monitoring and screening of qualified medicinal materials.

[0095] Market transaction data includes the transaction price, quantity, supplier, and buyer information of medicinal materials. This data is converted into a standardized format for data traceability and market analysis in the distribution process.

[0096] See Figure 3 As shown, deploying blockchain nodes at each stage of the industry chain to build a decentralized consortium blockchain network for evidence storage, generating timestamp fingerprints for key process data, and automatically uploading and storing the data on the blockchain via smart contracts specifically includes:

[0097] Blockchain nodes are deployed across all links of the Chinese medicinal herb industry chain, and trust relationships and access permissions are established between nodes.

[0098] By designing a consensus mechanism and data sharing rules for the consortium blockchain, data can be shared among the nodes.

[0099] The data of each key process in each link of the Chinese medicinal materials industry chain is timestamped, and the fingerprint of the data is generated by hash function to verify the integrity and immutability of the data.

[0100] Design smart contracts to define the conditions and processes for uploading data to the blockchain, and automatically package and upload the data to the blockchain for evidence storage based on the smart contracts.

[0101] Specifically, blockchain nodes will be deployed in each link of the Chinese medicinal materials industry chain, including planting, harvesting, processing, warehousing, logistics, preparation production, market circulation, and clinical use. These nodes will be responsible for receiving, storing, and verifying data in each link.

[0102] Trust relationships are established between nodes through a consortium blockchain mechanism. Each node must verify the identity of other nodes to ensure the legality of data transmission and operations. Different permissions are set according to the roles and responsibilities of the nodes, and permissions are managed through smart contracts to ensure that only authorized nodes can perform data read and write operations.

[0103] By using the Byzantine fault tolerance algorithm as the consensus mechanism in the consortium blockchain, the consensus mechanism ensures that each node consistently verifies the validity of the data before writing it. In the Chinese medicine industry chain, nodes at each stage reach a consensus on the validity of each piece of data through the consensus mechanism.

[0104] By defining a unified data format and interface, seamless data exchange between different stages is ensured. Data at each stage must be encrypted during recording to protect data privacy. Once data is stored on the blockchain, all verified nodes can read it, but write operations are only permitted with authorization.

[0105] A timestamp is generated for each key process data point to record its creation time. These timestamps are marked using block timestamps in the blockchain, and each block contains this timestamp, indicating when the data occurred. A hash function is used to generate a unique fingerprint for the data. The hash function converts input data into a fixed-length unique fingerprint for data verification.

[0106] Smart contracts define which data meets the conditions for being uploaded to the blockchain and which data needs to be verified before it can be uploaded. Smart contracts define the data upload process, including data verification, timestamp generation, and hash value calculation. The smart contract will automatically execute these operations when the data meets the conditions.

[0107] See Figure 4 As shown, based on data from different stages on each blockchain node, a data fusion model based on graph neural networks is constructed through feature extraction. This model correlates data from different stages to generate cross-stage correlated data, specifically including:

[0108] Each link and data item in the industrial chain is regarded as a node in the graph. Edges are constructed based on the logical relationships and data correlations between the links in the industrial chain to build a graph model.

[0109] Feature extraction is performed on each node, and the graph model is trained using data from each stage based on a graph neural network model.

[0110] By using the node embedding results of graph neural networks, data from different stages are associated and fused, and associated data and the relationships between them are generated.

[0111] Specifically, each link in the industrial chain is regarded as a node in the graph. Each node contains data items generated by that link. Edges represent the logical relationships and data correlations between different links. The weight of the edges can be set according to the strength of the correlation to reflect the data flow and interaction between different links.

[0112] Each node has a set of features, which can come from various data types generated in that stage. During the training of a graph neural network, node features usually need to be normalized to ensure that the values ​​are within an appropriate range.

[0113] During the training of graph neural networks, node features typically need to be normalized to ensure that the values ​​are within an appropriate range. The core formula for graph convolution is:

[0114] H l+1 =σ(AH (l) W (l) )

[0115] Among them, H l Let W be the feature matrix of the nodes in the l-th layer. (l) The weight matrix of the l-th layer, A is the adjacency matrix of the graph, σ is the non-linear activation function, and H... l+1 This is the updated node feature matrix at layer (l+1).

[0116] After multiple layers of graph convolution, the final embedding vector of each node is obtained. These embedding vectors will serve as the high-dimensional feature representation of the node, reflecting the semantics of the node in the graph.

[0117] Node embedding vectors obtained through graph neural networks are used for cross-stage data fusion. Data from different stages are interconnected through node embedding vectors, forming richer relational information.

[0118] See Figure 5 As shown, based on local processing standards data and historical herbal literature data, a knowledge graph of traditional Chinese medicine is constructed, mapping the relationships between cross-process related data in the graph. Specifically, this includes:

[0119] Collect local processing technology data, including processing methods, traditional processing techniques, time and temperature control, and label and classify the processing methods with corresponding medicinal materials, processing conditions, timing of use, and precautions;

[0120] We organize the descriptions of medicinal materials in historical herbal literature, including their properties, uses, effects, and compatibility. We extract information about medicinal materials from the literature, integrate different descriptions of medicinal materials, remove duplicates, and standardize them.

[0121] A knowledge graph of traditional Chinese medicine is constructed, with nodes including medicinal materials, processing methods, literature records, pharmacological properties, and usage routes, and edges representing the relationships between nodes;

[0122] The nodes in the cross-process related data are matched with the nodes in the Chinese medicine ontology knowledge graph, and the association relationships of the cross-process related data are mapped to the Chinese medicine ontology knowledge graph.

[0123] Specifically, the nodes include: Chinese medicinal materials, processing methods, literature records, pharmacological properties, and usage routes. The edges represent the relationships between the nodes, such as the relationship between medicinal materials and their processing methods, the relationship between medicinal materials and literature records, and the relationship between the pharmacological properties of medicinal materials.

[0124] The collected local processing technology data and descriptions of medicinal materials in ancient herbal literature are matched with nodes in the Chinese medicine ontology knowledge graph. Based on the matching results, cross-process related data are mapped to the corresponding nodes and edges in the Chinese medicine ontology knowledge graph, ensuring that the mapped data and relationships can accurately reflect the characteristics of medicinal materials, processing procedures, and their status and uses in traditional Chinese medicine theory.

[0125] See Figure 6 As shown, based on the industry chain data in the knowledge graph of traditional Chinese medicine, the entire chain of digital traceability from planting to clinical application is achieved through path reasoning algorithms, specifically including:

[0126] Based on the industry chain data in the knowledge graph of traditional Chinese medicine, starting from the planting and breeding link and ending at the clinical use link, the data path and relationship are marked.

[0127] By using an improved Dijkstra algorithm, multiple weights are added to each edge, and the weights are selected according to different needs when solving the path.

[0128] Based on the dependencies between different stages, a complete tracing path is formed by using an improved Dijkstra algorithm, and the optimal path is displayed in the graph.

[0129] The improved Dijkstra algorithm dynamically adjusts the path based on new data and periodically updates the path reasoning model.

[0130] Specifically, the basic idea of ​​Dijkstra's algorithm is to use dynamic programming to gradually expand the path from the starting point and find the shortest path from the starting point to the ending point by continuously updating the path value. In the improved Dijkstra's algorithm, multiple weights are considered and selected to meet the needs of different stages.

[0131] Starting from the origin, select the node with the shortest distance (i.e., the node used for reasoning along the current path). For each neighbor node directly connected to the origin, calculate the total weight from the origin to the neighbor node using the following formula:

[0132] dist(v) = min(dist(v), ω uv (k)·dist(u))

[0133] Where u is the node with the smallest distance, v is the neighboring node, and ω uv (k) represents the k-th weight from node u to node v, such as time or cost, and dist represents the distance.

[0134] When selecting a path, each edge has multiple weights. Under different requirements, the most suitable weight is selected. At each step of reasoning, the optimal path of the current node is updated and the path is recorded.

[0135] See Figure 7 As shown, the hybrid collaborative storage architecture based on time-series databases, graph databases, and relational databases, through a dynamic sharding strategy, efficiently manages access to billions of data points, specifically including:

[0136] Separate databases are established for storing time-series data, graph data, and relational data, and a hybrid collaborative storage architecture is constructed.

[0137] Time series databases are sharded according to time range; graph databases are sharded by nodes or edges, and sharded according to the subgraphs of the graph; relational databases are sharded by primary key hash or range sharding.

[0138] Based on the access frequency and update time characteristics of the data, hot data and cold data are managed separately. Active time series data are stored in fast-access storage, while historical cold data is archived to low-cost storage.

[0139] By establishing connection bridges between different databases, cross-database joint queries can be constructed.

[0140] Specifically, time series data is suitable for storing massive amounts of data indexed by time, such as IoT sensor data and financial market data; graph data is suitable for storing complex relational data, such as social networks, recommendation systems, and knowledge graphs; relational data is suitable for structured, closely related data, such as traditional business data, orders, and customer information.

[0141] To enable cross-database join queries, a data access layer is designed. This layer allows collaborative queries between different databases in the following ways:

[0142] The data access layer enables cross-database queries through middleware or API gateways;

[0143] The data fusion layer provides the ability to aggregate data and query across databases, combining the query needs of time series data, graph data, and relational data;

[0144] Time series data is typically indexed by time, so it can be sharded by time range. Frequently accessed data (hot data) can be stored in high-performance storage devices, while infrequently accessed historical data (cold data) can be archived in low-cost storage.

[0145] The partitioning strategy for graph data mainly depends on the structure of the graph, and partitioning is performed by nodes or edges.

[0146] There are generally two sharding strategies for relational databases: hash sharding and range sharding.

[0147] See Figure 8 As shown, a process prediction model is constructed based on planting data, climate data, and sales data from the industry chain. LSTM time-series analysis is used to detect process parameter deviations and provide risk warnings. Specifically, this includes:

[0148] Feature extraction is performed based on planting data, climate data, and sales data in the industrial chain data, and composite features are constructed by combining different data sources for model training;

[0149] The model is trained using a long short-term memory network, with the input layer containing multiple features, including key node data in the phenological and planting processes.

[0150] Based on the trained LSTM model, real-time data is input into the model to predict process parameters.

[0151] By comparing actual process parameter data with predicted results, offset detection is performed, and risk warnings are issued for offsets exceeding thresholds.

[0152] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0153] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0154] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for constructing big data across the entire traditional Chinese medicine industry chain, characterized in that: include: For each of the seven links in the Chinese medicinal materials industry chain, structured data transformation rules were designed to generate standardized metadata templates; Deploy blockchain nodes at each stage of the industrial chain to build a decentralized consortium blockchain network for evidence storage, generate timestamp fingerprints for key process data and automatically upload them to the blockchain for evidence storage via smart contracts; Based on data from different stages on each blockchain node, a data fusion model based on graph neural networks is constructed through feature extraction to correlate data between different stages and generate cross-stage correlated data. Based on local processing standards data and historical materia medica literature data, a knowledge graph of Chinese medicine is constructed to map the relationships between cross-linked data in the graph. Based on the industrial chain data in the knowledge graph of traditional Chinese medicine, a path reasoning algorithm is used to achieve full-chain digital traceability from planting to clinical application. Based on a hybrid collaborative storage architecture of time-series databases, graph databases and relational databases, it efficiently manages the access of billions of data points through a dynamic sharding strategy. A process prediction model is built based on planting, climate, and sales data from the industry chain. LSTM time series analysis is used to detect process parameter deviations and provide risk warnings.

2. The method for constructing big data across the entire traditional Chinese medicine industry chain according to claim 1, characterized in that, The aforementioned design of structured data transformation rules for the seven key links in the Chinese medicinal materials industry chain, and the generation of standardized metadata templates, specifically includes: The seven major links in the Chinese medicinal materials industry chain include planting and breeding, harvesting and processing, processing and production, warehousing and logistics, preparation production, clinical use and market circulation. Structured transformation rules were designed for sensor data, image and video data, process parameter data, quality inspection data, clinical electronic medical record data, and market transaction data based on seven key aspects. The structured transformation rules include: extracting morphological features of medicinal materials from image and video data using the YOLOv5 model and converting them into 128-dimensional feature vectors; parsing equipment logs using the OPC UA protocol to generate quadruple data for process parameter data; and using ICD-11 standard coding for syndrome differentiation and classification of clinical data, and grading and quantifying efficacy feedback. The data is transformed according to structured transformation rules to generate standardized metadata templates.

3. The method for constructing big data across the entire traditional Chinese medicine industry chain according to claim 1, characterized in that, The deployment of blockchain nodes at various stages of the industry chain to build a decentralized consortium blockchain network for evidence storage, generating timestamp fingerprints for key process data and automatically uploading them to the blockchain for evidence storage via smart contracts, specifically includes: Blockchain nodes are deployed across all links of the Chinese medicinal herb industry chain, and trust relationships and access permissions are established between nodes. By designing a consensus mechanism and data sharing rules for the consortium blockchain, data can be shared among the nodes. The data of each key process in each link of the Chinese medicinal materials industry chain is timestamped, and the fingerprint of the data is generated by hash function to verify the integrity and immutability of the data. Design smart contracts to define the conditions and processes for uploading data to the blockchain, and automatically package and upload the data to the blockchain for evidence storage based on the smart contracts.

4. The method for constructing big data across the entire traditional Chinese medicine industry chain according to claim 1, characterized in that, The process of constructing a graph neural network-based data fusion model based on feature extraction, using data from different stages on each blockchain node, to correlate data between different stages and generate cross-stage correlated data specifically includes: Each link and data item in the industrial chain is regarded as a node in the graph. Edges are constructed based on the logical relationships and data correlations between the links in the industrial chain to build a graph model. Feature extraction is performed on each node, and the graph model is trained using data from each stage based on a graph neural network model. By using the node embedding results of graph neural networks, data from different stages are associated and fused, and associated data and the relationships between them are generated.

5. The method for constructing big data across the entire traditional Chinese medicine industry chain according to claim 1, characterized in that, The construction of a knowledge graph of traditional Chinese medicine based on local processing standards and historical materia medica literature, and the mapping of cross-linked data relationships within the graph, specifically includes: Collect local processing technology data, including processing methods, traditional processing techniques, time and temperature control, and label and classify the processing methods with corresponding medicinal materials, processing conditions, timing of use, and precautions; We organize the descriptions of medicinal materials in historical herbal literature, including their properties, uses, effects, and compatibility. We extract information about medicinal materials from the literature, integrate different descriptions of medicinal materials, remove duplicates, and standardize them. A knowledge graph of traditional Chinese medicine is constructed, with nodes including medicinal materials, processing methods, literature records, pharmacological properties, and usage routes, and edges representing the relationships between nodes; The nodes in the cross-process related data are matched with the nodes in the Chinese medicine ontology knowledge graph, and the association relationships of the cross-process related data are mapped to the Chinese medicine ontology knowledge graph.

6. The method for constructing big data across the entire traditional Chinese medicine industry chain according to claim 1, characterized in that, The aforementioned digital traceability of the entire chain from planting to clinical application, based on the industry chain data in the knowledge graph of traditional Chinese medicine ontology and through path reasoning algorithms, specifically includes: Based on the industry chain data in the knowledge graph of traditional Chinese medicine, starting from the planting and breeding link and ending at the clinical use link, the data path and relationship are marked. By using an improved Dijkstra algorithm, multiple weights are added to each edge, and the weights are selected according to different needs when solving the path. Based on the dependencies between different stages, a complete tracing path is formed by using an improved Dijkstra algorithm, and the optimal path is displayed in the graph. The improved Dijkstra algorithm dynamically adjusts the path based on new data and periodically updates the path reasoning model.

7. The method for constructing big data across the entire traditional Chinese medicine industry chain according to claim 1, characterized in that, The hybrid collaborative storage architecture based on time-series databases, graph databases, and relational databases, which efficiently manages access to billions of data points through a dynamic sharding strategy, specifically includes: Separate databases are established for storing time-series data, graph data, and relational data, and a hybrid collaborative storage architecture is constructed. Time series databases are sharded according to time range; graph databases are sharded by nodes or edges, and sharded according to the subgraphs of the graph; relational databases are sharded by primary key hash or range sharding. Based on the access frequency and update time characteristics of the data, hot data and cold data are managed separately. Active time series data are stored in fast-access storage, while historical cold data is archived to low-cost storage. By establishing connection bridges between different databases, cross-database joint queries can be constructed.

8. The method for constructing big data across the entire traditional Chinese medicine industry chain according to claim 1, characterized in that, The process prediction model, constructed based on planting, climate, and sales data from the industry chain, and the risk warning generated through LSTM time-series analysis to detect process parameter deviations, specifically includes: Feature extraction is performed based on planting data, climate data, and sales data in the industrial chain data, and composite features are constructed by combining different data sources for model training; The model is trained using a long short-term memory network, with the input layer containing multiple features, including key node data in the phenological and planting processes. Based on the trained LSTM model, real-time data is input into the model to predict process parameters. By comparing actual process parameter data with predicted results, offset detection is performed, and risk warnings are issued for offsets exceeding thresholds.

9. The method for constructing big data across the entire traditional Chinese medicine industry chain according to claim 1, characterized in that, Enhance the quality feedback loop: collect clinical adverse reaction data and trace the production batches of the formulation in reverse; use Bayesian networks to calculate the probability of defects in each link, generate process improvement plans, and update the standard system.

10. A big data construction system for the entire traditional Chinese medicine industry chain, used to implement the big data construction method for the entire traditional Chinese medicine industry chain as described in any one of claims 1-9, characterized in that, include: Data structuring and transformation module: This module designs structuring and transformation rules and generates standardized metadata templates based on the data characteristics of each link in the industry chain; Blockchain Evidence Preservation Module: This module is used to deploy blockchain nodes, build a decentralized consortium blockchain network for evidence preservation, and ensure the immutability and timestamp fingerprint evidence preservation of key process data. Data fusion and association module: This module is based on a graph neural network model and uses feature extraction to achieve the association and fusion of cross-stage data, generating cross-stage associated data; Knowledge Graph Module: This module is used to construct a knowledge graph of traditional Chinese medicine, mapping and integrating historical documents, processing standards, and cross-process related data. Chain traceability reasoning module: The module is based on the industrial chain data in the knowledge graph of traditional Chinese medicine, and uses path reasoning algorithm to realize full-chain digital traceability from planting to clinical use; Hybrid collaborative storage module: The module is used to establish a hybrid collaborative storage architecture for time series data, graph data, and relational databases, and to achieve efficient data management through a dynamic sharding strategy; Process prediction and risk warning module: Based on industry chain data, the module predicts process parameters through LSTM time series analysis and performs offset detection and risk warning; Processor: The processor is used to handle the calculation process of each formula and the construction calculation process of each model.

Citation Information

Patent Citations

  • Traditional Chinese medicine product traceability system based on block chain

    CN114819996A

  • Traditional Chinese medicine decoction piece quality traceability system and method based on block chain

    CN117788007A

  • Traditional Chinese medicinal material tracing method and system based on knowledge graph, medium and product

    CN119537614A

  • Medical quality management method and system based on big data and artificial intelligence

    CN119557779A

  • Knowledge graph construction method and system based on Chinese medicine classics

    CN120124733A

Cited By

  • Economic database construction method for algae industry

    CN121935414A