Blockchain-based methods, systems, and devices for the traceability management of Chinese medicinal materials.

By constructing a dynamic standard adaptation mechanism in the blockchain, the problem of consistency in identification conclusions after standard updates in the Chinese medicinal materials information traceability system was solved, enabling credible traceability and auditing of Chinese medicinal materials quality judgment, and improving the long-term credibility and self-consistency of the system.

CN122414941APending Publication Date: 2026-07-17GUIZHOU JIANYI MEASUREMENT TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUIZHOU JIANYI MEASUREMENT TECH CO LTD
Filing Date
2026-06-22
Publication Date
2026-07-17

Smart Images

  • Figure CN122414941A_ABST
    Figure CN122414941A_ABST
Patent Text Reader

Abstract

This invention discloses a blockchain-based method, system, and device for tracing and managing information on traditional Chinese medicine (TCM) materials. Specifically, it relates to the field of blockchain data verification and processing technology, addressing the technical problem of existing blockchain-based traceability systems struggling to reconcile the contradictions between fixed historical conclusions on the blockchain and evolving off-chain standards after updates to identification criteria. By acquiring original characteristic data of TCM materials, current standard identifiers, and historical conclusions, and binding them to the blockchain for evidence storage, the invention monitors standard version updates, constructs feature spaces based on both old and new standard identifiers, performs cluster analysis, establishes cluster mapping relationships between the old and new clustering results, determines a new conclusion generation strategy, and finally obtains a new identification conclusion based on the strategy. This conclusion is then bound to the new standard identifier and associated with the original hash value, stored on the blockchain. This achieves intelligent re-evaluation of historical conclusions under standard evolution, ensuring the long-term credibility and consistency of traceability information while maintaining the immutability of the blockchain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of blockchain data verification and processing technology, and more specifically, to a blockchain-based method, system, and device for tracing and managing information on traditional Chinese medicine materials. Background Technology

[0002] In the field of quality control of Chinese medicinal materials, using blockchain technology to achieve information traceability is an important development direction. Existing technical solutions usually store key information of Chinese medicinal materials in planting, processing and circulation, as well as the authenticity identification conclusions based on characteristic component analysis, on the blockchain after hash processing, so as to achieve data immutability and full traceability. The aim is to ensure the reliability of traceability information through technical means and serve the authenticity identification of Chinese medicinal materials.

[0003] However, the aforementioned solutions based on existing technologies present technical contradictions when dealing with the long-term validity of authenticity identification conclusions. The immutability of blockchain fixes the data state at a specific point in time, including the identification standards or model versions used at that time. When the scientific standards or algorithm models for authenticity identification are updated, the existing technical architecture struggles to reconcile the relationship between the fixed historical conclusions stored on the chain and the continuously evolving standard system off the chain. This results in the blockchain-based traceability system being unable to self-consistently verify the consistency between historical identification conclusions and current standards over a long timescale, thereby weakening its technical foundation as a continuous and trustworthy traceability platform. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides a method, system and device for information traceability management of Chinese medicinal materials based on blockchain to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: Blockchain-based methods for tracing and managing information on traditional Chinese medicinal materials include: S1: Obtain the original characteristic data set, current standard identifier and corresponding historical identification conclusions of the target batch of Chinese medicinal materials; S2: Bind the original feature data set, the current standard identifier, and the historical identification conclusion to generate the first data combination, and store its hash value in the blockchain; S3: When a version update event of the authenticity identification standard is detected, obtain the new standard identifier; S4: Construct feature spaces based on the current standard identifier and the new standard identifier respectively, and perform cluster analysis on the sample set containing the target batch of Chinese medicinal materials to obtain the first clustering result and the second clustering result; S5: Establish the cluster mapping relationship between the first clustering result and the second clustering result, and determine the new identification conclusion generation strategy for the target batch of Chinese medicinal materials based on the cluster mapping relationship; S6: Based on the new identification conclusion generation strategy, obtain a new identification conclusion corresponding to the new standard identifier, bind the new identification conclusion with the new standard identifier to generate a second data combination, and associate the hash value of the second data combination with the hash value of the first data combination and store it in the blockchain.

[0006] Furthermore, S1 includes: Obtain multiple candidate original feature data sets, multiple candidate standard identifiers, and multiple historical identification conclusions associated with the target batch of Chinese medicinal materials; The current standard identifier is determined from multiple candidate standard identifiers based on the version time information corresponding to the identifier; Based on the current standard identifier, select the historical identification conclusions that correspond to the current standard identifier from multiple historical identification conclusions; Based on the data timestamp associated with the historical identification conclusions corresponding to the current standard identifier, the original feature data set of the target batch of Chinese medicinal materials is determined from multiple candidate original feature data sets.

[0007] Furthermore, S2 includes: The original feature data set, the current standard identifier, and the historical identification conclusions are concatenated in a preset order to generate the first data combination; Perform a hash operation on the first data combination to obtain the first hash value; The first hash value and the batch identifier of the target batch of Chinese medicinal materials are used together to form the on-chain data unit; The data unit is sent to the blockchain network node for verification and written into the blockchain.

[0008] Furthermore, S3 includes: Periodically monitor multiple predefined standard publishing addresses; When the version information published at any standard publication address is inconsistent with the current standard identifier, record the version update event; Perform multi-source cross-validation on the recorded version update events; When a version update event passes multi-source cross-validation, the new standard identifier is obtained from the validated standard release address.

[0009] Furthermore, S4 includes: Analyze the feature weight vectors corresponding to the current standard identifier and the new standard identifier; Based on the feature weight vector, the original feature data sets of all samples containing the target batch of Chinese medicinal materials are weighted and transformed to form the first feature space data and the second feature space data. Density-based clustering analysis was performed on the first feature space data and the second feature space data, respectively. Obtain and record the first clustering result based on the first feature space data and the second clustering result based on the second feature space data.

[0010] Furthermore, S5 includes: Calculate the distance between the centroid of each cluster in the first clustering result and the centroid of each cluster in the second clustering result; Establish a cluster-to-cluster mapping relationship from the first clustering result to the second clustering result based on the minimum distance between the centroids; Identify the first cluster to which the target batch of Chinese medicinal materials belongs in the first clustering result; Find the second cluster corresponding to the first cluster in the cluster-to-cluster mapping relationship; Based on the distance between the center points of the first and second clusters and the category attributes of the second cluster, a new identification conclusion generation strategy for the target batch of Chinese medicinal materials is determined.

[0011] Furthermore, the strategies for generating new identification conclusions include inheritance strategies, recalculation strategies, and verification strategies; When the distance between the center points of the first cluster and the second cluster is less than the preset center point distance threshold and the category attribute of the second cluster is consistent with the historical identification conclusion, the inheritance strategy is activated and the historical identification conclusion is used as the new identification conclusion. When the distance between the center points of the first cluster and the second cluster is greater than or equal to the preset center point distance threshold, the recalculation strategy is activated, and the original feature data set is recalculated according to the identification rules corresponding to the new standard identifier to obtain a new identification conclusion. When the distance between the center points of the first cluster and the second cluster is less than the preset center point distance threshold, but the category attribute of the second cluster is inconsistent with the historical identification conclusion, the verification strategy is activated, triggering manual or higher-order models to verify the original feature data set and the new standard label to obtain a new identification conclusion.

[0012] Furthermore, S6 includes: Based on the new identification conclusion generation strategy, the original characteristic data set of the target batch of Chinese medicinal materials is standardized and logically judged to obtain a new identification conclusion corresponding to the new standard identifier; The new identification conclusion and the new standard identifier are encoded and bound according to a preset data format to generate a second data combination; A hash operation is performed on the second data combination to obtain a second hash value, and the second hash value, together with the first hash value and the batch identifier of the target batch of Chinese medicinal materials, are used to construct an associated transaction; Submitting related transactions to the blockchain network allows the blockchain network to record the related transactions in the same blockchain data block containing the first hash value or in adjacent blockchain data blocks with pointer association.

[0013] On the other hand, the present invention provides a blockchain-based information traceability management system for traditional Chinese medicine materials, including: The data acquisition module is used to acquire the original characteristic data set, current standard identifier and corresponding historical identification conclusions of the target batch of Chinese medicinal materials; The data storage module is used to bind the original feature data set, the current standard identifier and the historical identification conclusion to generate the first data combination, and store its hash value in the blockchain; The monitoring and acquisition module is used to acquire the new standard identifier when a version update event of the locality identification standard is detected. The clustering analysis module is used to construct feature spaces based on the current standard identifier and the new standard identifier, respectively, and to perform clustering analysis on the sample set containing the target batch of Chinese medicinal materials to obtain the first clustering result and the second clustering result; The strategy determination module is used to establish the cluster mapping relationship between the first clustering result and the second clustering result, and to determine the new identification conclusion generation strategy for the target batch of Chinese medicinal materials based on the cluster mapping relationship. The conclusion update module is used to obtain a new identification conclusion corresponding to the new standard identifier based on the new identification conclusion generation strategy, bind the new identification conclusion with the new standard identifier to generate a second data combination, and associate the hash value of the second data combination with the hash value of the first data combination and store it in the blockchain.

[0014] On the other hand, the present invention provides a blockchain-based traceability management device for traditional Chinese medicine materials, comprising: One or more processors; Storage device for storing one or more programs; When one or more programs are executed by one or more processors, they enable one or more processors to implement a blockchain-based method for tracing and managing information on Chinese medicinal materials.

[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. By introducing a dynamic standard adaptation and conclusion calculation mechanism into the blockchain evidence storage architecture, the credibility and consistency of Chinese medicinal material traceability information over a long period are effectively improved. When the authenticity identification standard is updated, the system does not simply overwrite the old record with the new conclusion. Instead, it constructs feature spaces under the old and new standards respectively and performs cluster analysis on the same batch of samples. The abstract standard differences are transformed into changes in the topological structure of data distribution in the feature space for quantitative evaluation. By establishing a mapping relationship between the old and new clusters and intelligently deciding the generation strategy of the new conclusion based on the position and distance of the target batch in the mapping, the re-evaluation of historical conclusions under the standard evolution has a calculable and verifiable technical basis. This systematically reconciles the inherent contradiction between the fixed evidence storage of blockchain and the dynamic updating of standards, ensuring the logical self-consistency and authority of the traceability chain in the long-term evolution.

[0016] 2. The system upgrades the traceability process from static record keeping to dynamic intelligent governance. Through fully automated processing of monitoring, analysis, decision-making, and associated storage, the system can proactively and accurately respond to standard update events. Based on the inheritance, recalculation, or verification strategies determined by cluster mapping relationships and preset thresholds, it not only ensures the scientific validity of new conclusions but also optimizes the allocation of computing resources, avoiding the overhead of indiscriminate full recalculation. The system binds new conclusions with new identifiers and forms a verifiable association structure with their historical records on the blockchain, constructing a complete evidence chain centered on data batches and accommodating multiple versions of standard conclusions. This allows the quality judgment history of any batch of medicinal materials to be clearly and reliably traced and audited, significantly enhancing the technical capabilities of the blockchain-based traceability platform to cope with knowledge evolution and long-term regulatory requirements. Attached Figure Description

[0017] Figure 1 This is a flowchart of the blockchain-based method for tracing and managing information on traditional Chinese medicinal materials according to the present invention; Figure 2 This is a schematic diagram of the structure of the blockchain-based Chinese medicinal herb information traceability management system of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Example 1: Figure 1 The present invention provides a blockchain-based method for traceability management of traditional Chinese medicinal materials, comprising: S1: Obtain the original characteristic data set, current standard identifier and corresponding historical identification conclusions of the target batch of Chinese medicinal materials; S2: Bind the original feature data set, the current standard identifier, and the historical identification conclusion to generate the first data combination, and store its hash value in the blockchain; S3: When a version update event of the authenticity identification standard is detected, obtain the new standard identifier; S4: Construct feature spaces based on the current standard identifier and the new standard identifier respectively, and perform cluster analysis on the sample set containing the target batch of Chinese medicinal materials to obtain the first clustering result and the second clustering result; S5: Establish the cluster mapping relationship between the first clustering result and the second clustering result, and determine the new identification conclusion generation strategy for the target batch of Chinese medicinal materials based on the cluster mapping relationship; S6: Based on the new identification conclusion generation strategy, obtain a new identification conclusion corresponding to the new standard identifier, bind the new identification conclusion with the new standard identifier to generate a second data combination, and associate the hash value of the second data combination with the hash value of the first data combination and store it in the blockchain.

[0020] S1: Obtain the original characteristic data set, current standard identifier, and corresponding historical identification conclusions of the target batch of Chinese medicinal materials, and implement the following: The system acquires multiple candidate original characteristic data sets associated with the target batch of Chinese medicinal materials. It retrieves all test report files marked with the same batch identifier as the target batch from the databases of one or more Chinese medicinal material testing institutions or internal quality control laboratories. The test report files exist in a structured data format, with each report file corresponding to a candidate original characteristic data set. A candidate original characteristic data set contains a series of quantitative or qualitative data of characteristic components obtained through instrumental analysis, such as the astragaloside A content determined by high-performance liquid chromatography (HPLC) in milligrams per gram (mg / g), the content of various mineral elements determined by inductively coupled plasma mass spectrometry (ICP-MS) in micrograms per gram (µg / g), and spectral feature vectors obtained and preprocessed by near-infrared spectroscopy. These characteristic data are generated at different stages of the Chinese medicinal material circulation process by different operators at different times, based on the then-current effective operating procedures and instrument calibration status; therefore, multiple versions of candidate original characteristic data sets may exist. The system also retrieves all standard documents related to the identification of the authenticity of this type of Chinese medicinal material from publicly available databases published through publicly available standard release channels. Each standard document is parsed to extract its core identification rules, list of characteristic components, and corresponding threshold requirements, generating a unique standard identifier. This identifier includes the issuing organization, standard name, standard code, and standard version number. All these identifiers constitute multiple candidate standard identifiers. The system queries and retrieves all identification conclusion records associated with the target batch of Chinese medicinal materials from the local historical conclusion database or from the on-chain evidence records. These records constitute multiple historical identification conclusions. Each historical identification conclusion clearly records the conclusion content, issuance time, issuing organization, and the then-valid standard identifier upon which it was based.

[0021] The system determines the current standard identifier based on the version time information corresponding to each identifier from multiple candidate standard identifiers. The system parses the standard version number information contained in each candidate standard identifier one by one. The standard version number information follows a clear naming or numbering rule, such as combining the release year with the revision number. The system extracts the implicit or explicit release date or version effective date from each candidate standard identifier and compares it as version time information. The system sorts all candidate standard identifiers in descending order according to their corresponding version time information, that is, the standard identifier with the latest release date or latest version number is placed at the top. The system selects the standard identifier at the top of the list, i.e., the one with the latest version time information, and determines it as the current standard identifier. For example, among multiple candidate standard identifiers for identifying the authenticity of Astragalus membranaceus, the system identifies one candidate standard identifier as the 2023 version and another as the 2015 version; therefore, the 2023 version standard identifier with the updated version time information is determined as the current standard identifier. This determination process is automated and does not rely on manual selection, ensuring that the determined standard identifier represents the latest effective identification basis in the time dimension.

[0022] Based on the current standard identifier, the system filters historical identification conclusions corresponding to the current standard identifier from multiple historical identification conclusions. The system precisely matches the standard identifier field recorded in each historical identification conclusion record with the identified current standard identifier. The matching is a strict equality comparison based on the unique code or complete characteristic string of the standard identifier. Only when the standard identifier recorded in a historical identification conclusion record is completely identical to the current standard identifier in terms of issuing organization, standard name, standard code, and standard version number, is that historical identification conclusion determined to correspond to the current standard identifier. For example, if the current standard identifier is Appendix XX of the 2020 edition of the Pharmacopoeia, only those historical identification conclusions explicitly stating that they were issued based on Appendix XX of the 2020 edition of the Pharmacopoeia will be filtered out, while historical identification conclusions issued based on the 2015 edition of the Pharmacopoeia or other technical specifications will be excluded. This filtering process ensures that the extracted historical identification conclusions have a strict correspondence with the currently selected latest standard in terms of technical basis, avoiding misuse of conclusions due to inconsistencies in standard versions.

[0023] Based on the data timestamps associated with historical identification conclusions corresponding to the current standard identifier, the system determines the original characteristic data set of the target batch of Chinese medicinal materials from multiple candidate original characteristic data sets. Each selected historical identification conclusion corresponding to the current standard identifier contains a data timestamp accurate to the second in its metadata. The data timestamp records the acquisition or generation time of the original test data used to generate this historical identification conclusion. The system reads this data timestamp and then searches through the previously acquired multiple candidate original characteristic data sets. Each candidate original characteristic data set also has its own generation timestamp. The system compares the data timestamp of the historical identification conclusion with the generation timestamp of each candidate original characteristic data set. The system calculates the absolute value of the time difference between the data timestamp of the historical identification conclusion and the generation timestamp of each candidate original characteristic data set. The system then finds the candidate original characteristic data set with the smallest absolute value of the time difference. When a candidate set of original feature data is found, and the absolute value of the time difference between its generation timestamp and the timestamp of the historical identification conclusion is within a preset allowable time error threshold (e.g., the absolute value of the time difference does not exceed 24 hours), the system determines this candidate set of original feature data as the original feature data set for the target batch of Chinese medicinal materials. This determined set of original feature data is the best-matching underlying feature data actually used to generate the historical identification conclusion corresponding to the latest standard, thus ensuring the accuracy and consistency of the feature data upon which subsequent analysis is based. The preset allowable time error threshold range is set according to the actual testing process and data reporting cycle, for example, it can be set to 12 hours, 24 hours, or 48 hours, and the specific value is adjusted according to the type of Chinese medicinal materials and the real-time requirements of the testing process. The system predefines this allowable time error threshold range in the configuration file, enabling the system to adapt to the time matching requirements in different scenarios.

[0024] S2: Bind the original feature data set, the current standard identifier, and historical identification conclusions to generate the first data combination, and store its hash value in the blockchain, as follows: The original feature data set, the current standard identifier, and historical identification conclusions are concatenated in a preset order to generate the first data combination. The original feature data set exists in binary serialization format. The current standard identifier and historical identification conclusions are converted into UTF-8 encoded strings. The preset order is to place the byte stream of the original feature data set first, followed by the string byte stream of the current standard identifier, and finally the string byte stream of the historical identification conclusions. The concatenation operation joins the byte streams of the original feature data set, the string byte stream of the current standard identifier, and the string byte stream of the historical identification conclusions sequentially to form a continuous byte array. This continuous byte array is the first data combination. The preset order is defined as a rule in the system configuration file. The system configuration file is a text file that contains the field order definitions. The system reads the rule in the system configuration file and performs the concatenation operation at runtime. This preset rule ensures the consistency of the generated sequence structure.

[0025] The first data combination is hashed to obtain the first hash value. The hash operation uses the SHA-256 algorithm to process the byte array of the first data combination. The system calls a cryptographic function library that implements the SHA-256 algorithm. The cryptographic function library is an existing, publicly available software library. The system passes the byte array of the first data combination as an input parameter to the entry function of the cryptographic function library. The cryptographic function library performs the standard SHA-256 algorithm calculation process. The SHA-256 algorithm calculation process includes padding the input byte array so that the length of the padded data satisfies the condition that modulo 512 equals 448. Padding involves adding a binary 1 and several binary 0s to the end of the original data. After padding, a 64-bit integer representing the bit length of the original data is appended to the end of the data. The padded data is then divided into multiple 512-bit data blocks. Eight 32-bit hash variables are initialized; these hash variables are preset fixed constants. For each 512-bit data block, it is expanded into 64 32-bit words. The algorithm performs 64 rounds of iterative computation, updating the hash variable in each round using a preset constant and a character generated from the data block. After processing the last data block, the final eight hash variables are concatenated to form a 256-bit sequence. This 256-bit sequence is then converted into a string of 64 hexadecimal characters. This string of 64 hexadecimal characters is the first hash value.

[0026] The first hash value and the batch identifier of the target batch of medicinal herbs are combined to form the on-chain data unit. The batch identifier of the target batch of medicinal herbs is a unique string. It consists of the abbreviation of the herb's name in pinyin, the administrative division code of the place of origin, a four-digit harvest year, and a six-digit production serial number, with each part connected by hyphens. The system organizes the first hash value string and the batch identifier string of the target batch of medicinal herbs together according to a predefined format: a JSON object. The system creates a JSON object. Within the JSON object, a key-value pair is created with the key named `hash` and the corresponding value being the first hash value string. Another key-value pair is created within the JSON object with the key named `batchId` and the corresponding value being the batch identifier string of the target batch of medicinal herbs. The system serializes this JSON object into a UTF-8 encoded string. This final generated string is the on-chain data unit.

[0027] The system sends the on-chain data unit to a blockchain network node for verification and writes it to the blockchain. The system connects to a pre-configured node address on the blockchain network. The blockchain network is a consortium blockchain employing a practical Byzantine fault-tolerant consensus mechanism. The system sends transactions through the application programming interface (API) provided by this node. The endpoint of the API is, for example, an HTTPS URL. The system constructs an HTTP POST request. The request body of the HTTP POST request is the string representing the on-chain data unit. The header of the HTTP POST request contains authentication information, such as an API key. The blockchain network node receives the HTTP POST request. The blockchain network node verifies the validity of the API key. The blockchain network node verifies the correct JSON format of the on-chain data unit. After successful verification, the blockchain network node places the transaction data into its local mempool of transactions to be packaged. When it is the node's turn to produce a block, the node packages multiple transactions from the mempool into a new block. The node broadcasts the new block to other consensus nodes in the consortium blockchain via a peer-to-peer network. Other consensus nodes verify and vote on the new block according to the consensus algorithm. When more than two-thirds of the nodes confirm the block, it is linked to the end of the blockchain. A transaction's location on the blockchain is uniquely determined by the block height and the transaction's index within the block. This process permanently stores the on-chain data unit on the blockchain.

[0028] An index pointer is established on the blockchain for the first hash value. This index pointer is implemented by deploying a smart contract. The smart contract is written in Solidity. A public state variable is defined within the smart contract; this state variable is a mapping. The key of the mapping is a string type, used to store the batch identifier of the target batch of medicinal materials. The value of the mapping is an array type, where each element is a structure. Each structure contains three fields: a hash value field, a timestamp field, and a standard identifier field. The smart contract includes a function, for example, named `createIndex`. The `createIndex` function takes two arguments: the batch identifier of the target batch of medicinal materials and the first hash value. The logic of the `createIndex` function is to create a new array in the mapping, using the passed-in batch identifier of the target batch of medicinal materials as the key. A new structure element is added to this newly created array. The hash value field of the newly added structure element is set to the passed-in first hash value, the timestamp field is set to the timestamp of the current block, and the standard identifier field is set to an empty string. After the smart contract is compiled into bytecode, it is deployed to the blockchain and obtains a unique contract address. The index pointer is the combination of the smart contract address and the batch identifier of the target batch of medicinal herbs. Once the transaction corresponding to the first hash value is recorded on the blockchain, an off-chain listener detects this transaction. The off-chain listener calls the `createIndex` function of the deployed smart contract, passing in the batch identifier of the target batch of medicinal herbs and the first hash value. This call is executed on the blockchain as a transaction, thus creating an index in the blockchain state.

[0029] When the second hash value is associated with the first hash value, the smart contract calls the index pointer and appends the second hash value, its corresponding timestamp, and the new standard identifier to the index pointer. After the system generates the second hash value, a transaction calling the smart contract needs to be constructed. Another function, such as `appendRecord`, is defined in the smart contract. The `appendRecord` function receives three parameters: the batch identifier of the target batch of medicinal materials, the second hash value, and the new standard identifier. The logic of the `appendRecord` function is to use the batch identifier of the target batch of medicinal materials as the key to look up the corresponding array in the mapping within the smart contract. A new structure element is appended to the end of the found array. The hash value field of the newly appended structure element is set to the passed-in second hash value, the timestamp field is set to the timestamp of the current block, and the standard identifier field is set to the passed-in new standard identifier. The system initiates a transaction through a blockchain node, calling the `appendRecord` function of the smart contract and passing in the batch identifier of the target batch of medicinal materials, the second hash value, and the new standard identifier. This transaction is packaged and uploaded to the blockchain by miners or consensus nodes. The global state of the blockchain has been updated, and new records have been added to the array pointed to by the index pointers. The timestamps are automatically generated by the blockchain network when packaging blocks and are therefore trustworthy. The new standard identifier is the standard version identifier string that triggered this update.

[0030] A traceable chain structure is formed, with an index pointer as the root and multiple versions of data combined as hash values, which are associated with leaf nodes. The index pointer, a combination of the smart contract address and the batch identifier of the target batch of medicinal materials, is the unique entry point for querying the entire traceability history. Anyone or any system can obtain the traceability chain by querying the state of the smart contract. The query is performed by calling a read-only function of the smart contract, such as a function named `getHistory`. The `getHistory` function takes the batch identifier of the target batch of medicinal materials as a parameter and returns the corresponding entire array. The structure elements in the returned array are arranged in the order they were added. The first structure element corresponds to the first hash value and its on-chain time and initial standard identifier state. The second and subsequent structure elements correspond to the second hash value, the third hash value, and subsequent version hash values, along with their respective on-chain timestamps and explicit new standard identifiers. Each hash value is a cryptographic commitment pointing to a specific transaction on the blockchain. Through the blockchain's public interface, the corresponding complete transaction content can be queried using the hash value, thereby verifying the original data. This structure forms a logical chain, with the index pointer as the root and each node of the chain being a structure containing a hash value, a timestamp, and a standard identifier. The entire chain clearly records the complete digital authentication history of the same batch of Chinese medicinal materials as identification standards have evolved. The verifier only needs the batch identifier and smart contract address of the target batch of Chinese medicinal materials to completely reproduce the traceability chain without relying on any centralized database.

[0031] S3: When a version update event of the authenticity identification standard is detected, obtain the new standard identifier and implement the following: The system periodically monitors multiple predefined standard publication addresses. These predefined standard publication addresses are multiple Internet-accessible Uniform Resource Locators (URLs). Each URL points to a webpage or file download endpoint where an official or authoritative institution publishes documents on the identification standards for the authenticity of traditional Chinese medicinal materials. The system stores these predefined standard publication addresses in a text-based address list configuration file. Each line in the address list configuration file records one standard publication address. Periodic monitoring is performed by a separate background scheduled task. The scheduled task is configured using the operating system's scheduled task tool. The configuration of the scheduled task tool includes the execution cycle and the start command. The execution cycle can be set to, for example, once a day. The start command points to the specific program script that executes the monitoring task. When the scheduled task starts, the program script reads the address list configuration file. The read operation retrieves all predefined standard publication addresses in the file. For each standard publication address in the address list, the system performs a network access operation. The network access operation initiates a Hypertext Transfer Protocol (HTTP) request. The HTTP request sets a connection timeout parameter, for example, 30 seconds. The system checks the HTTP request's response status code. If the response status code is 200, the system retrieves the response body. The response body may contain a Hypertext Markup Language (HTML) page. The system parses the retrieved HTML page content. This parsing is performed by analyzing the Document Object Model (DOM) structure of HTML. The system locates the HTML element at a specific position on the page and reads its text content. This read text content is used as the version information string extracted from the standard publication address. The version information string is a text sequence. The system associates the extracted version information string with the version information string stored from the last successful monitoring at the same standard publication address. The system uses a local database table to store monitoring records. The local database table contains fields for address number, standard publication address, latest version information, and last monitoring time. After each successful monitoring and extraction of the version information string, the system updates the record in the local database table corresponding to the standard publication address. The update operation sets the latest version information field to the extracted version information string and the last monitoring time field to the current time. The periodic monitoring process repeats according to a preset cycle.

[0032] A version update event is recorded when the version information published at any standard publishing address is inconsistent with the current standard identifier. The current standard identifier is the identification standard identifier currently used by the target Chinese medicinal material category recorded in the system. After completing periodic monitoring of all predefined standard publishing addresses, the system performs a consistency comparison operation. For each predefined standard publishing address, the system reads the latest version information string corresponding to that address from the local database table. Simultaneously, the system reads the current standard identifier string from the system configuration storage. The system uses a string exact match algorithm to compare the latest version information string and the current standard identifier string. The string exact match algorithm performs character-by-character equality checks. If the character sequences of the latest version information string and the current standard identifier string are not completely identical, the system determines that they are inconsistent. When a discrepancy is found between the latest version information string corresponding to a standard publishing address and the current standard identifier string, the system triggers the recording of a version update event. Recording a version update event involves creating a new database record. The new database record is inserted into a data table named the Version Update Event Table. The Version Update Event Table contains a unique event number field, a standard publishing address field, a new version information string field, a current standard identifier string field, an event trigger timestamp field, and a status field. The system will fill the corresponding field of the new record with the standard publication address that triggered the event, the detected new version information string, the current standard identifier string of the current system record, and the current timestamp. The initial value of the status field is set to unverified.

[0033] Multi-source cross-validation is performed on recorded version update events. Multi-source cross-validation uses multiple independent information sources to confirm the same version update event. The validation process begins when a new record with an unverified status is generated in the version update event table. The system reads the records with an unverified status from the version update event table. For each unverified record, the system retrieves the value of the standard publishing address field stored in that record. Simultaneously, the system retrieves other addresses from a predefined list of standard publishing addresses. The system selects at least two other addresses. The system immediately initiates immediate network access requests to these selected other standard publishing addresses. These immediate network access requests are identical to the Hypertext Transfer Protocol (HTTP) requests in periodic monitoring. The system extracts the currently published version information string from the response content of these other standard publishing addresses. The system performs a cross-comparison operation. The cross-comparison operation checks whether at least two version information strings extracted from the other standard publishing addresses are identical to each other. The cross-comparison operation also checks whether these version information strings are identical to the value of the new version information string field stored in the version update event record. The system sets a consistency matching threshold. The consistency matching threshold is the minimum number of version information strings from other independent information sources that must match the new version information string in the event record. The consistency matching threshold is set as a positive integer in the system configuration, such as the number 2. The threshold is set based on the minimum number of sources required to ensure information reliability. If the consistency matching threshold is reached or exceeded, the system determines that the version update event has passed multi-source cross-validation. The system updates the status field of the record in the version update event table, changing it from unvalidated to validated. If the consistency matching threshold is not reached, the system marks the record's status field as validation failed. The multi-source cross-validation process is then complete.

[0034] When a version update event passes multi-source cross-validation, a new standard identifier is obtained from the validated standard publication address. The new standard identifier is a complete standard description identifier string. The process of obtaining the new standard identifier begins after the version update event record's status is updated to validated. The system selects a validated standard publication address as the metadata retrieval source. The selection rule can be choosing the first validated standard publication address. The system sends a specific request to the selected standard publication address to retrieve detailed standard metadata. The specific request is constructed by appending a query parameter to the original standard publication address; the key-value pair of the query parameter is, for example, `action` equals `detail`. Detailed standard metadata is returned in JSON format. The JSON data packet contains the standard name field, standard code field, publishing organization field, publication date field, effective date field, and version number field. The system parses this JSON data packet. The system generates a new standard identifier string by combining elements according to a predefined string template. The string template is, for example, "Publishing Organization:Standard Name (Standard Code, Version Number)". The system parses the values ​​of the publishing organization, standard name, standard code, and version number fields and fills them into the corresponding placeholder positions in the string template. This filling operation generates the final string, which serves as the new standard identifier string. The system associates the new standard identifier string with the version update event record. The system stores the new standard identifier string in a new standard identifier database table. This table contains an event number field, a new standard identifier string field, an retrieval time field, and a source address field. The system fills the event number of the version update event record, the generated new standard identifier string, the current time, and the source address into the corresponding fields of the new record. The process of obtaining the new standard identifier ends, and the resulting new standard identifier string is used in subsequent processing steps.

[0035] S4: Construct feature spaces based on the current standard identifier and the new standard identifier respectively, and perform cluster analysis on the sample set containing the target batch of Chinese medicinal materials to obtain the first clustering result and the second clustering result, as follows: The system parses the feature weight vectors corresponding to the current standard identifier and the new standard identifier. It queries a local standard knowledge base for records that exactly match the current standard identifier string. The local standard knowledge base is a table stored in a relational database. The table contains a standard identifier string field and a feature weight vector field. The feature weight vector field stores a serialized numeric array string. The feature weight vector is a floating-point array, the length of which is equal to the number of types of characteristic components of the considered Chinese medicinal materials. For example, for Astragalus membranaceus, the number of characteristic components might be 10, so the feature weight vector would be an array containing 10 floating-point numbers. Each element in the array represents the importance weight of a specific characteristic component in identification, with a weight value between 0 and 1. The sum of all weight values ​​equals 1. The feature weight vectors are determined by domain experts during standard development. Experts assign an initial weight value to each characteristic component based on the emphasis and pharmacological importance of each indicator in the standard text, normalize the values ​​to make the sum equal to 1, and then store the result in the local standard knowledge base. The system uses the current standard identifier string as the query key, executes a Structured Query Language (SCL) query, and retrieves the string value of the corresponding feature weight vector field from the local standard knowledge base table. The system deserializes this string value into a floating-point array, which is the current feature weight vector. The system then uses the new standard identifier string as the query key, performs the same query operation, and retrieves the string value of another feature weight vector field from the local standard knowledge base table. The system deserializes this string value into another floating-point array, which is the new feature weight vector. If no record is found that exactly matches the new standard identifier string, the system performs a weight derivation process. This process reads the feature weight vector corresponding to the old version standard identifier that is closest to the new standard identifier string, and then, based on the revision information contained in the new standard identifier string, applies a predefined weight adjustment mapping table to modify the weight values. The weight adjustment mapping table is a configuration file that lists the adjustment amount of the corresponding weight when a specific type of change occurs to the threshold or description of a feature component in the standard; the adjustment amount is, for example, a positive or negative percentage. The system calculates new weight values ​​based on the mapping table and normalizes them so that their sum is 1, thus obtaining a new feature weight vector.

[0036] Based on the feature weight vector, a weighted transformation is performed on all original feature data sets in the sample set containing the target batch of Chinese medicinal materials to form first feature space data and second feature space data. The sample set includes the original feature data set of the target batch of Chinese medicinal materials and the original feature data sets of multiple reference batches of Chinese medicinal materials from other known producing areas. All original feature data sets have the same structure, that is, they all contain the measured values ​​of the same type and order of feature components. The system first performs data standardization preprocessing. For all original feature data sets in the sample set, for each feature component, the system calculates the arithmetic mean of the measured values ​​of that feature component in all samples. The system calculates the standard deviation of the measured values ​​of that feature component in all samples. For the original measured value of each feature component in each original feature data set, standardization calculation is performed. Standardization calculation is to subtract the arithmetic mean of the feature component from the original measured value, and then divide the difference by the standard deviation of the feature component. The result of the standardization calculation is a dimensionless numerical value, called the standardized value. After the above processing, each original feature data set is converted into an array of standardized values. Then, a weighted transformation is performed to form the first feature space data. The system iterates through each array of standardized values ​​in the sample set. For the standardized value of the u-th feature component in the array, the system multiplies it by the u-th weight value in the current feature weight vector, where u represents the ordinal index of the feature component in the array, an integer counting from 1. The multiplication operation yields the weighted feature values. The system combines the weighted feature values ​​of all feature components in their original order into a new array. This new array represents the feature space coordinates of the sample from the current standard perspective. The above multiplication and combination operations are repeated for each sample in the sample set, and all the resulting new arrays constitute a dataset, which is the first feature space data. The first feature space data is represented in memory as a two-dimensional matrix, where the rows of the matrix correspond to samples, and the columns correspond to the feature dimensions after weighted transformation. Completely independently, the system performs weighted transformations to form the second feature space data. The system iterates through each array of standardized values ​​in the same sample set. For the standardized value of the u-th feature component in the array, the system multiplies it by the u-th weight value in the new feature weight vector. The multiplication operation yields another set of weighted feature values. The system combines the weighted eigenvalues ​​of all feature components into a new array in their original order. This new array represents the feature space coordinates of the sample under the new standard perspective. The above multiplication and combination operations based on the new feature weight vector are repeated for each sample in the sample set. All the resulting new arrays constitute another data set, which is the second feature space data. The second feature space data is represented in memory as another two-dimensional matrix.

[0037] Density-based clustering analysis is performed on the first and second feature space data respectively. The density-based clustering analysis uses the DBSCAN algorithm. The DBSCAN algorithm requires setting a neighborhood radius parameter and a minimum number of points parameter. The neighborhood radius parameter is a distance value used to define the neighborhood range. The minimum number of points parameter is an integer used to define the minimum number of neighbors for a core point. The system sets the neighborhood radius parameter for the first feature space data. This is done by calculating the average Euclidean distance between all sample points in the two-dimensional matrix of the first feature space data. The Euclidean distance is calculated by taking the square root of the sum of the squares of the differences between two sample points in each feature dimension. The system multiplies the calculated average distance by a scaling factor, for example, 0.2, and uses the product as the neighborhood radius parameter value. The system also sets the minimum number of points parameter for the first feature space data. This is done by multiplying the total number of samples by a density factor, for example, 0.05, and then rounding the product up to obtain the minimum number of points parameter value. The system then runs the DBSCAN algorithm on the two-dimensional matrix of the first feature space data. During algorithm initialization, all sample points are marked as unvisited. The algorithm randomly selects an unvisited point. It then calculates the neighborhood of this point, finding all points in the first feature space data's two-dimensional matrix whose Euclidean distance to the point is less than the neighborhood radius parameter. If the number of points in the neighborhood is greater than or equal to the minimum number of points parameter, the point is marked as a core point, and a new cluster is created. The algorithm iterates through all points in the neighborhood of this core point. If these points are unvisited, they are marked as visited and added to the current cluster; if these points are also core points, their neighborhoods are recursively added to the current cluster. If the number of points in the neighborhood of the initially selected point is less than the minimum number of points parameter, the point is marked as a noise point. The algorithm repeats the process of selecting unvisited points and executing the above steps until all points are visited. After the algorithm finishes, the sample points in the first feature space data are divided into several clusters and a set of noise points. This partitioning result is called the clustering result based on the first feature space data. The system independently performs the same DBSCAN clustering process on the second feature space data's two-dimensional matrix. The system uses the same calculation method to set independent neighborhood radius and minimum point count parameters for the second feature space data. The system runs the DBSCAN algorithm to obtain the clustering results of the sample points in the second feature space data. This clustering result is called the clustering result based on the second feature space data. Density-based clustering analysis can identify sample clusters of arbitrary shapes.

[0038] Obtain and record the first clustering result based on the first feature space data and the second clustering result based on the second feature space data. The first clustering result is a structured data object. The first clustering result contains a list of cluster labels. The length of the cluster label list is equal to the number of samples. The nth element in the cluster label list is an integer representing the cluster number to which the nth sample in the sample set is assigned in the first feature space data clustering. n represents the sequential index position of the sample in the sample set, which is an integer counting from 1. The special number -1 is an integer label reserved in the clustering result to represent samples that are judged by the algorithm as noise points and do not belong to any core cluster. The first clustering result contains a list of core points. The core point list records the index positions of all samples marked as core points by the DBSCAN algorithm in the sample set. The first clustering result contains a list of cluster centroids. The method for generating the list of cluster centroids is to calculate the arithmetic mean of the coordinates of all core points belonging to each non-noise cluster for each non-noise cluster. The arithmetic mean is calculated independently on each feature dimension, and the result is a coordinate point, which is the centroid of the cluster. The system serializes the first clustering result into a JSON string. This JSON string contains an array of cluster labels, an array of core point indices, and an array of cluster centroids. The system constructs the second clustering result into another structured data object. This second clustering result contains a list of cluster labels with the same meaning as the first clustering result, but records the distribution of samples within the second feature space data clustering. The second clustering result contains its own list of core points and cluster centroids, obtained through analysis of the clustering results in the second feature space data. The system serializes the second clustering result into another JSON string. The system stores the JSON strings of the first and second clustering results into an analysis result database table. This table contains an analysis task identifier field, a first clustering result JSON field, a second clustering result JSON field, and a creation timestamp field. The system generates a unique analysis task identifier and then inserts the first clustering result JSON string, the second clustering result JSON string, and the current timestamp as a new record into the analysis result database table. The system simultaneously retains the first and second clustering result data objects in memory, allowing subsequent calculation steps to directly access the cluster labels and centroid information. This logging operation ensures the persistence and traceability of the clustering results.

[0039] JSON is a lightweight data exchange format used to store and transmit structured data objects; DBSCAN is a density-based spatial clustering algorithm that can discover clusters of arbitrary shapes and identify noise points; ID is an abbreviation for identifier, referring to the unique number of the analysis task in the database.

[0040] S5: Establish the cluster mapping relationship between the first clustering result and the second clustering result, and determine the strategy for generating new identification conclusions for the target batch of Chinese medicinal materials based on the cluster mapping relationship, as follows: The system calculates the distance between the centroids of each cluster in the first clustering result and each cluster in the second clustering result. The first clustering result contains a list of cluster centroids. The second clustering result contains another list of cluster centroids. Each element in the list of cluster centroids is a coordinate point. The number of dimensions of the coordinate points is the same as the number of dimensions of the feature space. The system sequentially reads the coordinates of the centroids of each cluster from the list of cluster centroids in the first clustering result. The system sequentially reads the coordinates of the centroids of each cluster from the list of cluster centroids in the second clustering result. The system uses the Euclidean distance formula to calculate the centroid distance between all possible cluster pairs. For the coordinates of a cluster centroid in the first clustering result and the coordinates of a cluster centroid in the second clustering result, the Euclidean distance is calculated by squared differences in each dimension between the two coordinate points, summing the squared differences in all dimensions, and then calculating the square root of the sum. The result of the square root operation is the centroid distance between the two cluster centroids. The centroid distance is a non-negative real number. The system initializes a two-dimensional array as a distance matrix. The number of rows in the distance matrix is ​​equal to the number of clusters in the first clustering result. The number of columns in the distance matrix equals the number of clusters in the second clustering result. The system uses a nested loop to iterate through all cluster combinations. The outer loop iterates through the index i of each cluster in the first clustering result. The inner loop iterates through the index j of each cluster in the second clustering result. In each iteration, the system calculates the Euclidean distance between the centroid of the cluster with index i in the first clustering result and the centroid of the cluster with index j in the second clustering result. The system assigns the calculated distance value to the element in the i-th row and j-th column of the distance matrix. After the loop completes, the distance matrix stores the centroid distances between all cluster pairs.

[0041] The system establishes a cluster-to-cluster mapping relationship from the first clustering result to the second clustering result based on the minimum centroid distance. For each cluster in the first clustering result, the system performs a lookup operation. The lookup operation is performed on a single cluster in the first clustering result, denoted as index i. The system reads all elements in the i-th row of the distance matrix. All elements in the i-th row form a vector containing the distances from the cluster at index i in the first clustering result to the centroids of each cluster in the second clustering result. The system searches for the element with the smallest value in this row vector. The method for finding the minimum value is to initialize a temporary variable to record the current minimum distance value, with an initial value set to a very large number, such as positive infinity. Simultaneously, a temporary variable is initialized to record the corresponding second cluster index. The system sequentially traverses each column index j of this row vector. During the traversal, the value of the element in the i-th row and j-th column of the distance matrix is ​​compared with the currently recorded minimum distance value. If the value of the element in the i-th row and j-th column is less than the currently recorded minimum distance value, the minimum distance value of the current record is updated with the value of that element, and the corresponding second cluster index of the record is updated with column index j. After traversal, the index of the second cluster in the current record is the index of the second cluster that is closest to the cluster center of index i in the first clustering result. The system creates a mapping record. The mapping record contains two fields: the first field stores the index i of the first cluster, and the second field stores the index of the found second cluster. The system adds this mapping record to a cluster mapping relationship list. The system repeats the above lookup and recording operation for each cluster in the first clustering result. Finally, the cluster mapping relationship list contains the mapping relationship from each first cluster to its nearest second cluster; this list is the cluster-to-cluster mapping relationship. The system serializes the cluster-to-cluster mapping relationship list into a JSON format string for persistent storage. The top level of the JSON string is an array. Each element in the array is an object. Each object has two properties: the first property is named fromClusterId, whose value is the index number of the first cluster; the second property is named toClusterId, whose value is the index number of the second cluster.

[0042] The system identifies the first cluster to which the target batch of Chinese medicinal herbs belongs in the first clustering result. The system reads the cluster label list from the first clustering result. The cluster label list is an integer array. The length of the array is equal to the total number of samples in the sample set. The index position of the array corresponds to the sequential position of the sample in the sample set. The system needs to determine the index position of the original feature data set of the target batch of Chinese medicinal herbs in the sample set. The index position of the target batch of Chinese medicinal herbs in the sample set is assigned and recorded in an index mapping table when the sample set is constructed. The system queries this index mapping table, using the batch identifier of the target batch of Chinese medicinal herbs as the query key, to obtain the corresponding sample index number. The system uses the obtained sample index number as the subscript to access the cluster label array. The system reads the integer value at that subscript position in the cluster label array. This integer value represents the cluster number to which the target batch of Chinese medicinal herbs belongs in the first clustering result. If the read integer value is -1, it indicates that the target batch of Chinese medicinal herbs is determined to be a noise point in the first clustering result. At this point, the system treats it as a conceptual noise cluster and assigns it a specific virtual number that does not conflict with any real cluster number, such as virtual number 0. The system records this identified cluster number, whether it is a real cluster number or a virtual number, and this number is recorded as the number of the first cluster.

[0043] The system searches for the second cluster corresponding to the first cluster in the cluster-to-cluster mapping relationship. It loads the previously established cluster-to-cluster mapping relationship list from storage. This list is a data structure containing multiple mapping records. The system iterates through each record in the list. For each record, the system compares the value of the first cluster number field stored in the record with the previously identified first cluster number. If they match, the system reads the value of the second cluster number field from this record. The system records this read value as the search result. This result is the number of the second cluster with the minimum distance mapping relationship to the first cluster. If no matching record is found in the entire cluster-to-cluster mapping relationship list, this may occur when the first cluster is a noisy cluster with a dummy number. In this case, the system executes an alternative nearest neighbor search logic. This alternative logic retrieves the coordinates of the target batch of medicinal herbs in the first feature space data. The coordinates are vectors. The system then reads the list of cluster centers from the second clustering result. The system calculates the Euclidean distance from the coordinates of the target batch of medicinal herbs to the center point of each cluster in the second clustering result. The system selects the minimum value among all calculated Euclidean distances. The number of the second cluster corresponding to the minimum value is the search result. The system records the second cluster number found through backup logic. Regardless of the method used, the final obtained second cluster number will be used for subsequent decision-making.

[0044] Based on the distance between the center points of the first and second clusters and the category attribute of the second cluster, a new identification conclusion generation strategy for the target batch of Chinese medicinal materials is determined. The new identification conclusion generation strategy includes inheritance, recalculation, and verification strategies. The system first obtains the distance between the center points of the first and second clusters. This is done by using the cluster number as the row index and the second cluster number as the column index to query the previously calculated and stored distance matrix. The system reads the element value at the corresponding row and column index positions from the distance matrix. This element value is the distance between the center points of the first and second clusters. The system then reads a preset center point distance threshold. This preset center point distance threshold is a real number pre-configured in the system parameter file. The preset center point distance threshold is set based on historical data analysis. The setting method involves collecting all valid cluster pair mapping records generated during multiple past standard updates. For each valid cluster pair mapping record, its corresponding center point distance value is extracted. The arithmetic mean of all these center point distance values ​​is calculated. The standard deviation of all these center point distance values ​​is calculated. The arithmetic mean is added to the standard deviation, and the sum is used as the suggested value for the preset center point distance threshold. The system administrator makes final settings and adjustments in the parameter file based on the suggested values. The system reads the category attributes of the second cluster. The category attribute refers to the authenticity classification label represented by the cluster. The determination of the category attribute depends on the labeling information of the sample set. When constructing the sample set, each reference batch of Chinese medicinal materials samples has a real origin category label, such as "authentic" or "non-authentic". After the second clustering result is formed, for each cluster in the second clustering result, the system finds all reference samples assigned to this cluster. The system counts the number of reference samples with the origin category label "authentic". The system counts the number of reference samples with the origin category label "non-authentic". If the number of authentic labels is greater than the number of non-authentic labels, the category attribute of the cluster is determined to be authentic; otherwise, the category attribute of the cluster is determined to be non-authentic. The system reads the historical identification conclusions of the target batch of Chinese medicinal materials. The historical identification conclusion is a string containing the content "authentic" or "non-authentic". The system executes the strategy judgment logic. The first step of the strategy judgment logic is numerical comparison: determining whether the distance between the center points of the first cluster and the second cluster is less than a preset center point distance threshold. The second step of the strategy judgment logic is string comparison: determining whether the category attribute string of the second cluster is completely identical to the historical identification conclusion string. If the result of the first step comparison is true and the result of the second step comparison is also true, the system determines to activate the inheritance strategy. The inheritance strategy means that the new identification conclusion will directly adopt the string value of the historical identification conclusion without recalculation. If the result of the first step comparison is false, that is, the center point distance is greater than or equal to the preset center point distance threshold, the system determines to activate the recalculation strategy. The recalculation strategy means that the system needs to call the complete identification rule function corresponding to the new standard identifier, take the original feature dataset of the target batch of Chinese medicinal materials as input, and perform a completely new identification calculation to obtain a new identification conclusion.If the result of the first step comparison is true but the result of the second step comparison is false (i.e., the centroid distance is less than the threshold but the category attributes are inconsistent), the system determines to activate the review strategy. The review strategy means the system will generate a review task. The review task is placed in a task queue. The review task includes the batch identifier of the target batch of Chinese medicinal materials, the original feature data set, and the new standard identifier. The review task can be handled by a human expert, who views the data through an interface and provides a conclusion; or it can be handled by a pre-trained high-order machine learning model, which is more complex than the basic identification rules, such as a deep neural network model that receives the same input and outputs an identification conclusion. The system records the final determined new identification conclusion generation strategy in a decision result data structure. The decision result data structure contains the following fields: strategy type field, first cluster number field, second cluster number field, centroid distance field, preset centroid distance threshold field, and category attribute consistency flag field. The decision result data structure serves as the output of this step.

[0045] S6: Based on the new authentication conclusion generation strategy, obtain a new authentication conclusion corresponding to the new standard identifier. Bind the new authentication conclusion and the new standard identifier to generate a second data combination. Then, associate the hash value of the second data combination with the hash value of the first data combination and store it in the blockchain. The implementation is as follows: Based on the new identification conclusion generation strategy, the original feature data set of the target batch of Chinese medicinal materials is standardized and logically judged to obtain a new identification conclusion corresponding to the new standard identifier. The new identification conclusion generation strategy can be an inheritance strategy, a recalculation strategy, or a verification strategy. The system reads the new identification conclusion generation strategy string. If the content of the new identification conclusion generation strategy string is completely consistent with the inheritance strategy string, the system performs the inheritance operation. The inheritance operation involves the system retrieving the historical identification conclusion strings of the target batch of Chinese medicinal materials from records stored in the local database or on the blockchain. The system directly assigns the retrieved historical identification conclusion strings to the new identification conclusion string. The new identification conclusion generation process ends. If the content of the new identification conclusion generation strategy string is completely consistent with the recalculation strategy string, the system performs the recalculation operation. The first step of the recalculation operation is standardization. The system obtains the original feature data set of the target batch of Chinese medicinal materials. The original feature data set is an array containing the measured values ​​of multiple feature components. The system reads the global mean and global standard deviation of each feature component from a historical feature statistical database. The historical feature statistical database is calculated by statistically analyzing the feature measurement values ​​of a large number of similar Chinese medicinal material samples. For each feature component's measured value in the original feature dataset, the system performs a standardization calculation. Standardization involves subtracting the global mean of that feature component from its measured value to obtain a first difference. This first difference is then divided by the global standard deviation of the feature component to obtain its standardized value. The system performs standardization calculations on all feature components in the original feature dataset, resulting in an array of standardized values. The second step of the recalculation operation is logical discrimination. The system queries a rule base for a set of discrimination rules associated with the new standard identifier string. The discrimination rule set contains multiple rules. Each rule consists of three core elements: a feature component identifier, a comparison operator, and a threshold value. The threshold value is a specific real number, and its source and determination method are as follows: during the development of the new standard, the standard-setting organization sets the threshold value for each feature component through statistical analysis of a large-scale, known-origin sample set of benchmark medicinal materials. The specific setting process is as follows: For a specific characteristic component, all samples clearly marked as belonging to the authentic producing area are selected from the benchmark medicinal material sample set. A specific quantile of the measured value of this characteristic component in these samples is calculated, such as the 5th percentile or the 95th percentile. This quantile is used as the threshold value for that characteristic component in the rules. For example, when establishing the authenticity standard for Astragalus membranaceus, the content of astragaloside A is measured from hundreds of authentic Astragalus membranaceus samples, and its 5th percentile is calculated to be 0.05%. Therefore, the threshold value for the rules regarding astragaloside A in the rule base is set to 0.05%. The system inputs the standardized value array into the identification rule set. The system evaluates each rule sequentially. For each rule, the system extracts the corresponding standardized value from the standardized value array based on the characteristic component identifier specified in the rule.The system uses comparison operators defined in the rules, such as "greater than" or "greater than or equal to," to compare the standardized value with a threshold value defined in the rules and determined by the aforementioned statistical methods. The comparison operation produces a Boolean result: true or false. The system combines the Boolean results from all rule evaluations according to the logical AND and OR relationships defined in the rule set. The final Boolean result of the combined logic determines the new identification conclusion. If the final Boolean value is true, the new identification conclusion string is "authentic." If the final Boolean value is false, the new identification conclusion string is "not authentic." If the content of the new identification conclusion generation strategy string is completely consistent with the review strategy string, the system performs a review operation. The review operation involves the system creating a review task record. The review task record contains a batch identifier field for the target batch of medicinal materials, an original feature data set field, and a new standard identifier string field. The system writes the review task record to a review task queue database table. A separate review processing service monitors the review task queue database table. When the review processing service detects a new review task record, it performs manual review or model review according to the system configuration. If configured for manual review, the review processing service pushes the review task record information to a graphical user interface (GUI). Experts review the information on the GUI and input a conclusion string. The GUI returns the expert's input conclusion string to the system, which uses it as the new identification conclusion string. If configured for model review, the review processing service loads a pre-trained high-order machine learning model. The high-order machine learning model is a neural network model. The review processing service transforms the original feature data set and the new standard label string into the input feature vector required by the model. The review processing service inputs the input feature vector into the high-order machine learning model. The high-order machine learning model performs forward propagation calculations and outputs a probability distribution vector. The review processing service selects the category label with the highest probability value as the new identification conclusion string. The system ultimately obtains a definitive new identification conclusion string.

[0046] The new identification conclusion and the new standard identifier are encoded and bound according to a preset data format to generate the second data combination. The preset data format is a structured data serialization format. The system creates a new data object. The data object is represented in memory as a dictionary structure. The system adds a key-value pair to the dictionary structure. The key is named `conclusion`, and the corresponding value is the new identification conclusion string. The system adds another key-value pair to the dictionary structure. The key is named `standardId`, and the corresponding value is the new standard identifier string. The system adds a third key-value pair to the dictionary structure. The key is named `generationTime`, and the corresponding value is the system timestamp string when the new identification conclusion was generated. The system uses a JSON serializer to convert the dictionary structure into a string. The JSON serializer is a library function that converts an in-memory object into JSON formatted text. The JSON serializer takes a dictionary structure as input and outputs a string that conforms to JSON syntax. This output string is the second data combination. The character encoding of the second data combination string is UTF-8.

[0047] The second data combination is hashed to obtain a second hash value, which is then combined with the first hash value and the batch identifier of the target batch of Chinese medicinal materials to construct an associated transaction. The hash operation uses the SHA-256 algorithm. The system calls a cryptographic library function that implements the SHA-256 algorithm. The system takes the UTF-8 encoded byte array of the second data combination string as input and passes it to the SHA-256 algorithm function. The SHA-256 algorithm function executes its internal calculation process. The internal calculation process includes padding the input byte array so that the length of the padded data is an integer multiple of 512 bits. The padding process involves adding a 1 bit to the end of the input data, then adding several 0 bits, and finally adding a 64-bit integer representing the original data length. The padded data is then divided into multiple 512-bit data blocks. Eight 32-bit hash initial value constants are initialized. For each 512-bit data block, it is expanded into 64 32-bit words. 64 rounds of loop operation are performed, with each round updating eight working variables using a preset constant and expansion words. After processing all data blocks, the final eight working variables are concatenated into a 256-bit hash value sequence. The system converts this 256-bit hash value sequence into a string of 64 hexadecimal digits. This string is the second hash value. The system then begins constructing the associated transaction data structure. The associated transaction data structure is a dictionary containing specific fields. The system adds a field to the associated transaction data structure dictionary: `transactionType`, with the value being the version association. The system adds another field: `firstHash`, with the value being the first hash value string. The system adds a third field: `secondHash`, with the value being the second hash value string. The system adds a fourth field: `batchId`, with the value being the batch identifier string of the target batch of medicinal herbs. The system adds a fifth field: `timestamp`, with the value being the system timestamp at the time the associated transaction was constructed. The system uses a JSON serializer to serialize the associated transaction data structure dictionary into a single string, called the transaction data string. The system then digitally signs the transaction data string using a private key stored in the security module. The digital signature algorithm used is the Elliptic Curve Digital Signature Algorithm (ECA). The signing process involves calculating the SHA-256 hash value of the transaction data string. This hash value is then encrypted using the ECA and the private key to generate a digital signature string. The system combines the transaction data string and the digital signature string to form the associated transaction data packet.

[0048] The associated transaction is submitted to the blockchain network so that the blockchain network records the associated transaction in the same blockchain data block containing the first hash value or in adjacent blockchain data blocks with pointer association. The system reads the Uniform Resource Locator (URL) of the blockchain node application programming interface specified in the configuration file. The system constructs a Hypertext Transfer Protocol (HTTP) POST request. The header of the POST request sets the content type to application / json. The body of the POST request is a JSON object. The JSON object contains two fields. The first field name is data, and its value is the string obtained by Base64 encoding the transaction data string. The second field name is signature, and its value is the string obtained by Base64 encoding the digital signature string. The system sends the POST request to the URL of the blockchain node application programming interface through the HTTP client. The blockchain network node receives the POST request. The blockchain network node verifies the request. Verification includes decoding the Base64 encoded data to obtain the transaction data string and the digital signature string. The blockchain network node uses the public key corresponding to the private key to verify the validity of the digital signature string. After successful verification, the blockchain network node parses the transaction data string and extracts the first hash value field. A blockchain network node queries its local blockchain state index for a transaction record containing the first hash value string. The query returns the block height number of the transaction record, denoted as block height H. The blockchain network node places the currently associated transaction into a pending transaction pool. When the blockchain network's consensus mechanism reaches the point of generating a new block, a miner or block-producing node selects multiple transactions, including the associated transaction, from the transaction pool and packages them into a new block. The block header of this new block contains a field pointing to the hash value of the previous block. The hash value of the previous block ensures the chained association between blocks. After being confirmed by network consensus, the new block is appended to the blockchain with a block height of H+1. Therefore, the associated transaction is recorded in the block at block height H+1, which is directly associated with the block containing the first hash value at block height H through a hash pointer. If the blockchain network supports and deploys indexed smart contracts, the system can use another approach. The system directly constructs a transaction that calls the smart contract. The target address of the call transaction is the address of the indexed smart contract. The input data of the call transaction encodes a function call. The function is named `appendRecord`. Its parameters include a batch identifier string representing the target batch of medicinal herbs and a second hash value string. The system signs this transaction using its private key and broadcasts it to the blockchain network. Miners package this transaction into a new block. Regardless of the block height of this new block, any state changes to the indexed smart contract are recorded on the blockchain.By querying the state of the indexed smart contract and using the batch identifier string as the key, the associated first and second hash values ​​can be retrieved, thus logically establishing a link between them. Both methods achieve a permanent and verifiable association between the hash values ​​of the second data combination and the hash values ​​of the first data combination on the blockchain.

[0049] Example 2: Figure 2 A schematic diagram of the blockchain-based traceability management system for traditional Chinese medicine materials is provided. The blockchain-based traceability management system for traditional Chinese medicine materials includes: The data acquisition module is used to acquire the original characteristic data set, current standard identifier and corresponding historical identification conclusions of the target batch of Chinese medicinal materials; The data storage module is used to bind the original feature data set, the current standard identifier and the historical identification conclusion to generate the first data combination, and store its hash value in the blockchain; The monitoring and acquisition module is used to acquire the new standard identifier when a version update event of the locality identification standard is detected. The clustering analysis module is used to construct feature spaces based on the current standard identifier and the new standard identifier, respectively, and to perform clustering analysis on the sample set containing the target batch of Chinese medicinal materials to obtain the first clustering result and the second clustering result; The strategy determination module is used to establish the cluster mapping relationship between the first clustering result and the second clustering result, and to determine the new identification conclusion generation strategy for the target batch of Chinese medicinal materials based on the cluster mapping relationship. The conclusion update module is used to obtain a new identification conclusion corresponding to the new standard identifier based on the new identification conclusion generation strategy, bind the new identification conclusion with the new standard identifier to generate a second data combination, and associate the hash value of the second data combination with the hash value of the first data combination and store it in the blockchain.

[0050] Example 3: This example introduces a blockchain-based Chinese medicinal herb information traceability management device. The blockchain-based Chinese medicinal herb information traceability management device includes one or more processors; it also includes a storage device for storing one or more programs; when one or more programs are executed by one or more processors, the one or more processors implement the blockchain-based Chinese medicinal herb information traceability management method.

[0051] All calculations involved in the embodiments are dimensionless numerical calculations, and the preset parameters and thresholds in the calculations are set by those skilled in the art according to the actual situation.

[0052] It should be noted that this invention can be deployed on the device itself to realize embedded applications, or it can run on a PC or other terminal with a user interface, thereby meeting various hardware environments and usage requirements.

[0053] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions according to the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wireless or wired transmission; wired transmission methods include optical fiber, twisted pair, coaxial cable, etc.; wireless transmission includes infrared, microwave, etc. Computer-readable storage media can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more sets of available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. Semiconductor media can be solid-state drives.

[0054] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0055] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0056] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0057] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0058] If a function is implemented as a software module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0059] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0060] In conclusion, the above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A blockchain-based method for traceability management of traditional Chinese medicinal materials, characterized in that, include: S1: Obtain the original characteristic data set, current standard identifier and corresponding historical identification conclusions of the target batch of Chinese medicinal materials; S2: Bind the original feature data set, the current standard identifier, and the historical identification conclusion to generate the first data combination, and store its hash value in the blockchain; S3: When a version update event of the authenticity identification standard is detected, obtain the new standard identifier; S4: Construct feature spaces based on the current standard identifier and the new standard identifier respectively, and perform cluster analysis on the sample set containing the target batch of Chinese medicinal materials to obtain the first clustering result and the second clustering result; S5: Establish the cluster mapping relationship between the first clustering result and the second clustering result, and determine the new identification conclusion generation strategy for the target batch of Chinese medicinal materials based on the cluster mapping relationship; S6: Based on the new identification conclusion generation strategy, obtain a new identification conclusion corresponding to the new standard identifier, bind the new identification conclusion with the new standard identifier to generate a second data combination, and associate the hash value of the second data combination with the hash value of the first data combination and store it in the blockchain.

2. The method for traceability management of traditional Chinese medicinal materials based on blockchain according to claim 1, characterized in that, S1 includes: Obtain multiple candidate original feature data sets, multiple candidate standard identifiers, and multiple historical identification conclusions associated with the target batch of Chinese medicinal materials; The current standard identifier is determined from multiple candidate standard identifiers based on the version time information corresponding to the identifier; Based on the current standard identifier, select the historical identification conclusions that correspond to the current standard identifier from multiple historical identification conclusions; Based on the data timestamp associated with the historical identification conclusions corresponding to the current standard identifier, the original feature data set of the target batch of Chinese medicinal materials is determined from multiple candidate original feature data sets.

3. The blockchain-based method for tracing and managing information on traditional Chinese medicinal materials according to claim 1, characterized in that, S2 include: The original feature data set, the current standard identifier, and the historical identification conclusions are concatenated in a preset order to generate the first data combination; Perform a hash operation on the first data combination to obtain the first hash value; The first hash value and the batch identifier of the target batch of Chinese medicinal materials are used together to form the on-chain data unit; The data unit is sent to the blockchain network node for verification and written into the blockchain.

4. The blockchain-based method for traceability management of traditional Chinese medicinal materials according to claim 1, characterized in that, S3 includes: Periodically monitor multiple predefined standard publishing addresses; When the version information published at any standard publication address is inconsistent with the current standard identifier, record the version update event; Perform multi-source cross-validation on the recorded version update events; When a version update event passes multi-source cross-validation, the new standard identifier is obtained from the validated standard release address.

5. The blockchain-based method for traceability management of Chinese medicinal materials according to claim 1, characterized in that, S4 include: Analyze the feature weight vectors corresponding to the current standard identifier and the new standard identifier; Based on the feature weight vector, the original feature data sets of all samples containing the target batch of Chinese medicinal materials are weighted and transformed to form the first feature space data and the second feature space data. Density-based clustering analysis was performed on the first feature space data and the second feature space data, respectively. Obtain and record the first clustering result based on the first feature space data and the second clustering result based on the second feature space data.

6. The blockchain-based method for traceability management of Chinese medicinal materials according to claim 1, characterized in that, S5 include: Calculate the distance between the centroid of each cluster in the first clustering result and the centroid of each cluster in the second clustering result; Establish a cluster-to-cluster mapping relationship from the first clustering result to the second clustering result based on the minimum distance between the centroids; Identify the first cluster to which the target batch of Chinese medicinal materials belongs in the first clustering result; Find the second cluster corresponding to the first cluster in the cluster-to-cluster mapping relationship; Based on the distance between the center points of the first and second clusters and the category attributes of the second cluster, a new identification conclusion generation strategy for the target batch of Chinese medicinal materials is determined.

7. The blockchain-based method for traceability management of traditional Chinese medicinal materials according to claim 6, characterized in that, New identification conclusion generation strategies include inheritance strategy, recalculation strategy, and verification strategy; When the distance between the center points of the first cluster and the second cluster is less than the preset center point distance threshold and the category attribute of the second cluster is consistent with the historical identification conclusion, the inheritance strategy is activated and the historical identification conclusion is used as the new identification conclusion. When the distance between the center points of the first cluster and the second cluster is greater than or equal to the preset center point distance threshold, the recalculation strategy is activated, and the original feature data set is recalculated according to the identification rules corresponding to the new standard identifier to obtain a new identification conclusion. When the distance between the center points of the first cluster and the second cluster is less than the preset center point distance threshold, but the category attribute of the second cluster is inconsistent with the historical identification conclusion, the verification strategy is activated, triggering manual or higher-order models to verify the original feature data set and the new standard label to obtain a new identification conclusion.

8. The method for traceability management of traditional Chinese medicinal materials based on blockchain according to claim 1, characterized in that, S6 include: Based on the new identification conclusion generation strategy, the original characteristic data set of the target batch of Chinese medicinal materials is standardized and logically judged to obtain a new identification conclusion corresponding to the new standard identifier; The new identification conclusion and the new standard identifier are encoded and bound according to a preset data format to generate a second data combination; A hash operation is performed on the second data combination to obtain a second hash value, and the second hash value, together with the first hash value and the batch identifier of the target batch of Chinese medicinal materials, are used to construct an associated transaction; Submitting related transactions to the blockchain network allows the blockchain network to record the related transactions in the same blockchain data block containing the first hash value or in adjacent blockchain data blocks with pointer association.

9. A blockchain-based traceability management system for traditional Chinese medicinal materials, used to implement the blockchain-based traceability management method for traditional Chinese medicinal materials as described in any one of claims 1-8, characterized in that, include: The data acquisition module is used to acquire the original characteristic data set, current standard identifier and corresponding historical identification conclusions of the target batch of Chinese medicinal materials; The data storage module is used to bind the original feature data set, the current standard identifier and the historical identification conclusion to generate the first data combination, and store its hash value in the blockchain; The monitoring and acquisition module is used to acquire the new standard identifier when a version update event of the locality identification standard is detected. The clustering analysis module is used to construct feature spaces based on the current standard identifier and the new standard identifier, respectively, and to perform clustering analysis on the sample set containing the target batch of Chinese medicinal materials to obtain the first clustering result and the second clustering result; The strategy determination module is used to establish the cluster mapping relationship between the first clustering result and the second clustering result, and to determine the new identification conclusion generation strategy for the target batch of Chinese medicinal materials based on the cluster mapping relationship. The conclusion update module is used to obtain a new identification conclusion corresponding to the new standard identifier based on the new identification conclusion generation strategy, bind the new identification conclusion with the new standard identifier to generate a second data combination, and associate the hash value of the second data combination with the hash value of the first data combination and store it in the blockchain.

10. A blockchain-based information traceability management device for traditional Chinese medicinal materials, characterized in that: include: One or more processors; Storage device for storing one or more programs; When one or more programs are executed by one or more processors, the one or more processors implement the blockchain-based Chinese medicinal materials information traceability management method as described in any one of claims 1-8.