Blockchain and internet of things based agricultural product safety traceability method and system

CN122736637APending Publication Date: 2026-09-11NINGXIA FUNING QING AGRICULTURAL PRODUCTS DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611044562.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0003]然而,现有技术侧重于记录数据内容和流转节点,未能完整反映数据形成过程中原始采样源、检测样本、数据生成设备、校准设备、网关路径及预处理程序之间的依赖关系

Benefits of technology

[0027]本发明通过提取原始采样源、检测样本、数据生成设备、校准设备、网关路径及预处理程序等因果血缘信息,并与数据摘要关联上链,提高了溯源数据形成过程的完整性和可核验性。通过构建证据依赖知识图谱,识别共同关键祖先节点并划分证据根,能够发现不同溯源数据之间的隐性同源关系,避免将同源重复数据误判为相互独立的证据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122736637A_ABST
    Figure CN122736637A_ABST
Patent Text Reader

Abstract

This invention provides a method and system for agricultural product safety traceability based on blockchain and the Internet of Things (IoT). The method includes: acquiring IoT sensing data and detection data from the production, storage, transportation, and sales stages of agricultural products; extracting the original sampling source, detection sample, data generation equipment, calibration equipment, gateway path, preprocessing program, and submission node; generating a causal lineage vector and writing it into the blockchain along with a associated data digest; constructing an evidence dependency knowledge graph based on the causal lineage vector, identifying common key ancestor nodes, dividing evidence roots, and calculating causal independence; allocating model influence based on causal independence, optimizing using the Bat algorithm, and forming a candidate training set for the main model and a dataset to be isolated through deletion-based counterfactual inference; using reinforcement learning to determine training control actions, updating the agricultural product risk identification model and shadow model, generating model influence fingerprints, and correcting the evidence dependency knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agricultural product quality safety and data processing technology, and in particular to a method and system for agricultural product safety traceability based on blockchain and the Internet of Things. Background Technology

[0002] With the application of IoT and blockchain technologies in agricultural product quality and safety management, environmental parameters, cold chain status, and circulation information generated during the production, testing, storage, transportation, and sales of agricultural products can be collected through sensors, testing equipment, and business terminals, and stored and traced on a batch-by-batch basis. Data summaries can also be written into the blockchain to improve the tamper-proof capability of traceability records.

[0003] However, existing technologies focus on recording data content and flow nodes, failing to fully reflect the dependencies between the original sampling source, test sample, data generation equipment, calibration equipment, gateway path, and preprocessing procedures during data formation. When multiple data points share the same sampling source, calibration equipment, or processing procedure, existing systems are prone to misinterpreting them as independent evidence, leading to the duplication and amplification of data from the same source, thus affecting the accuracy of agricultural product safety risk assessment.

[0004] Meanwhile, existing risk identification models often use source data based on data quantity or fixed weights, lacking constraints on the degree of causal independence of different pieces of evidence and the cumulative model impact. Once a common data source, gateway, or preprocessing procedure malfunctions, the relevant data may collectively drive the shift of model parameters and decision boundaries, making it difficult to discover missed evidentiary dependencies based on actual model changes.

[0005] When test samples, equipment, calibration records, or processing procedures are confirmed to be invalid during subsequent re-inspection or audit, existing systems typically only add abnormal or invalidation marks to the original records. It is difficult to accurately determine the derived data of the invalidity evidence and the model version in which it is involved, and it is also difficult to eliminate the historical impact of the invalidity data on subsequent model training and batch risk assessment.

[0006] Therefore, this invention proposes a method and system for agricultural product safety traceability based on blockchain and the Internet of Things. The information disclosed in the background section is only for enhancing understanding of the background of this disclosure and may therefore contain prior art information that is not common knowledge to those skilled in the art. Summary of the Invention

[0007] The purpose of this invention is to address the shortcomings of existing technologies by providing a method and system for tracing the safety of agricultural products based on blockchain and the Internet of Things, thereby solving the technical problems mentioned in the background section.

[0008] To achieve the above objectives, the present invention provides the following technical solution:

[0009] The method for tracing the safety of agricultural products based on blockchain and the Internet of Things includes the following steps:

[0010] S1. Acquire IoT sensing data and detection data from agricultural product production, testing, storage, transportation and sales, extract original sampling sources, test samples, data generation equipment, calibration equipment, clock sources, gateway paths, preprocessing programs and submission nodes, generate causal lineage vectors and associate them with data summaries to write to the blockchain;

[0011] S2. Construct an evidence dependency knowledge graph based on the causal lineage vector, identify common key ancestor nodes, group data affected by the same common key ancestor node into evidence roots, and calculate causal independence based on the number of independent paths.

[0012] S3. The impact of the model is allocated according to the degree of causal independence, and the Bat algorithm is used for optimization. The impact of evidence roots on the predicted comprehensive model of agricultural product risk identification model is calculated by deletion counterfactual inference, forming a candidate training set and a dataset to be isolated.

[0013] S4. Based on the impact of the optimized model and the estimated impact of the comprehensive model, use reinforcement learning to determine the training control actions, update the agricultural product risk identification model and shadow model, and generate the model impact fingerprint; when the actual comprehensive model impact exceeds the limit, reverse-correct the evidence dependency knowledge graph.

[0014] S5. When subsequent test results determine that the evidence root is invalid, the invalid derived data and the affected model version are determined along the corrected evidence dependency knowledge graph. Based on the model impact fingerprint, checkpoint reloading and effective training batch replay are performed to obtain the agricultural product risk identification model after the impact is revoked. The safety risk status of the affected batch is determined and the impact revocation proof is written into the blockchain.

[0015] S1 specifically includes: acquiring data from production, testing, warehousing, transportation, and sales processes; generating agricultural product batch identifiers; correcting equipment recording times; and aggregating data into traceability data unit sets according to batch identifiers, business processes, and time windows; extracting original sampling sources, test samples, data generation equipment, calibration equipment, clock sources, gateway paths, preprocessing programs, and submission nodes from sampling records and related logs to generate causal lineage vector sets; normalizing and serializing the traceability data unit sets and causal lineage vector sets and calculating a joint digest; associating the joint digest with batch identifiers, business process identifiers, and original data storage locations to form evidence lineage records; signing and writing the records to the blockchain to obtain an on-chain evidence lineage record set.

[0016] S2 specifically includes: reading the on-chain evidence lineage record set and corresponding logs; constructing an evidence dependency knowledge graph with source data units and dependent objects as nodes and collection, detection, calibration, time synchronization, forwarding, processing calls, and on-chain submission relationships as directed edges; searching for ancestor nodes in reverse along the evidence dependency knowledge graph, identifying common key ancestor nodes, merging them with affected source data units into evidence roots, and constructing an evidence dependency hypergraph representing multi-dimensional dependency relationships; searching for independent paths that do not share common key ancestor nodes based on the evidence dependency hypergraph, calculating the maximum number of independent paths, the degree of source difference, the proportion of common key ancestor nodes, and the missing proportion, to obtain the causal independence set of each evidence root.

[0017] S3 specifically includes: reading the causal independence set, dividing it into training and validation sets, and generating an initial model influence set based on causal independence, risk category distribution contribution, and the importance of business processes; using the initial model influence set as the search center, encoding the candidate model influence as individual bat positions, and using the bat algorithm to evaluate candidate solutions based on risk identification performance, decision boundary displacement, and degree of influence violation to obtain an optimized model influence set; performing deletion-based counterfactual inference on each evidence root, calculating the estimated influence of parameters, feature centers, and decision boundaries, generating an estimated comprehensive model influence, and comparing it with the optimized model influence to form the main model candidate training set and the dataset to be isolated.

[0018] S4 specifically includes: reading the causal independence set, the optimized model influence amount set, and the evidence root predicted influence set, forming a state vector with the historical training state, and using reinforcement learning to train the control model to output the training control action sequence; updating the agricultural product risk identification model and shadow model based on the training control action sequence, calculating the model parameter increment, feature center offset, and decision boundary displacement of the evidence root, and generating the comprehensive model influence and model influence fingerprint; comparing the comprehensive model influence with the optimized model influence amount, and when it exceeds the limit, identifying the implicit common source based on the causal lineage vector, equipment logs, and model influence fingerprint, correcting the evidence dependency knowledge graph, and writing the model influence fingerprint and correction results into the blockchain.

[0019] S5 specifically includes: obtaining subsequent detection results, matching object identifiers, version identifiers, and failure time periods with the evidence dependency knowledge graph to determine the failure evidence root, searching for failure-derived data along valid dependencies and querying the affected model version; determining the revocation start checkpoint based on the failure-derived data, affected model version, model impact fingerprint, and off-chain training logs, restoring the model and optimizer state, replaying valid training batches in the original order after excluding failure-derived data, and obtaining the agricultural product risk identification model after impact revocation; determining the affected batches based on the impact path and batch relationship of the failure-derived data, re-determining the security risk status using the agricultural product risk identification model, and writing the impact revocation proof to the blockchain.

[0020] A blockchain and IoT-based agricultural product safety traceability system includes:

[0021] The data acquisition and storage module is used to acquire IoT sensing data and testing data from the production, testing, storage, transportation and sales of agricultural products.

[0022] The evidence dependency analysis module is used to construct an evidence dependency knowledge graph based on causal lineage vectors, identify common key ancestor nodes, and group data influenced by the same common key ancestor node into evidence roots.

[0023] The model impact optimization module is used to allocate the model impact amount according to the degree of causal independence, optimize the model impact amount using the bat algorithm, and calculate the estimated comprehensive model impact of each evidence root on the agricultural product risk identification model through deletion counterfactual inference.

[0024] The training control correction module is used to determine training control actions based on the impact of the optimized model and the estimated impact of the comprehensive model, and to update the agricultural product risk identification model and shadow model using reinforcement learning.

[0025] The Failure Impact Revocation Module is used to determine the failure-derived data and affected model versions along the corrected evidence dependency knowledge graph when subsequent detection results determine that the evidence root is invalid.

[0026] The beneficial effects of this invention are as follows:

[0027] This invention improves the integrity and verifiability of the data formation process by extracting causal lineage information such as the original sampling source, test sample, data generation device, calibration device, gateway path, and preprocessing program, and associating it with the data summary on the blockchain. By constructing an evidence dependency knowledge graph, identifying common key ancestor nodes, and dividing evidence roots, it can discover implicit homology relationships between different source data, avoiding the misjudgment of homologous data as independent evidence.

[0028] This invention limits the cumulative impact of low-independence evidence in model training by calculating the causal independence of evidence roots and allocating model influence limits, thereby reducing the amplifying effect of anomalies in the same device, gateway, or process on risk identification results. By optimizing the model influence limits through the Bat Algorithm and combining it with deletion-based counterfactual reasoning to evaluate the estimated comprehensive model influence of evidence roots, the rationality of training data selection and influence constraints is improved.

[0029] This invention uses reinforcement learning to determine training control actions, and combines agricultural product risk identification models, shadow models, and model impact fingerprints to verify the impact of the actual comprehensive model. It can also reverse-engineer and correct any missing evidence dependencies. When evidence is subsequently confirmed as invalid, by locating the invalid derived data and the affected model version, performing checkpoint reloading and replaying of valid training batches, the historical model impact of invalid evidence can be revoked, and the safety risk status of the relevant agricultural product batches can be accurately updated. Attached Figure Description

[0030] Figure 1 This is a flowchart of the agricultural product safety traceability method based on blockchain and the Internet of Things according to the present invention;

[0031] Figure 2 This is a framework diagram of the agricultural product safety traceability system based on blockchain and the Internet of Things according to the present invention. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] Example 1: As Figure 1 As shown in the figure, this embodiment provides a method for tracing the safety of agricultural products based on blockchain and the Internet of Things, including the following steps:

[0034] S1. Acquire IoT sensing data and detection data from agricultural product production, testing, storage, transportation and sales, extract original sampling sources, test samples, data generation equipment, calibration equipment, clock sources, gateway paths, preprocessing programs and submission nodes, generate causal lineage vectors and associate them with data summaries to write to the blockchain;

[0035] S2. Construct an evidence dependency knowledge graph based on the causal lineage vector, identify common key ancestor nodes, group data affected by the same common key ancestor node into evidence roots, and calculate causal independence based on the number of independent paths.

[0036] S3. The impact of the model is allocated according to the degree of causal independence, and the Bat algorithm is used for optimization. The impact of evidence roots on the predicted comprehensive model of agricultural product risk identification model is calculated by deletion counterfactual inference, forming a candidate training set and a dataset to be isolated.

[0037] S4. Based on the impact of the optimized model and the estimated impact of the comprehensive model, use reinforcement learning to determine the training control actions, update the agricultural product risk identification model and shadow model, and generate the model impact fingerprint; when the actual comprehensive model impact exceeds the limit, reverse-correct the evidence dependency knowledge graph.

[0038] S5. When subsequent test results determine that the evidence root is invalid, the invalid derived data and the affected model version are determined along the corrected evidence dependency knowledge graph. Based on the model impact fingerprint, checkpoint reloading and effective training batch replay are performed to obtain the agricultural product risk identification model after the impact is revoked. The safety risk status of the affected batch is determined and the impact revocation proof is written into the blockchain.

[0039] S1 specifically includes the following sub-steps:

[0040] S110. By deploying IoT devices at agricultural production sites, testing sites, storage facilities, transportation vehicles, and sales terminals, the original data generated in the production, testing, storage, transportation, and sales processes are acquired, and a traceability data unit set is established.

[0041] The raw data in the production process is generated by environmental sensors, irrigation controllers, pesticide spraying equipment, and harvesting terminals; the raw data in the testing process is generated by pesticide residue testing equipment, microbial testing equipment, heavy metal testing equipment, and sample management terminals; and the raw data in the warehousing, transportation, and sales processes are generated by in-warehouse sensors, weighing equipment, vehicle-mounted sensors, positioning terminals, vehicle-mounted gateways, barcode scanning terminals, and handover confirmation terminals, respectively.

[0042] During the first harvest, the batch generation terminal generates agricultural product batch identifiers according to the production entity identifier, plot identifier, variety identifier, harvest date, and batch number. When batches are split, sub-batch identifiers are generated based on the parent batch identifier and sub-batch number. When batches are merged, merged batch identifiers are generated, and each parent batch identifier and its input quality are recorded.

[0043] A traceability data unit is defined as the smallest independently verifiable set of data generated by a specific device for a specific batch of agricultural products at a single collection time or within a preset collection period. Each traceability data unit includes at least the original data, the original time, the data generating device identifier, the agricultural product batch identifier, the business process identifier, and the device digital signature.

[0044] Perform time correction on the i-th source data unit, where i represents the sequence number of the source data unit and its corresponding kinship record:

[0045]

[0046] In the formula, This represents the standard time of the i-th traceability data unit; Indicates the original time recorded by the data generating device; This indicates the time correction amount of the data generating device relative to a trusted time server or satellite timing terminal.

[0047] Devices that can directly connect to a standard clock source determine the time correction amount based on the time synchronization log; devices that cannot connect directly determine the time correction amount based on the clock deviation recorded by their respective gateways. Data with the same agricultural product batch identifier, the same business process, or adjacent data in a preset business process, and whose standard time difference does not exceed the preset time window of the corresponding business process, are grouped into the same traceability event.

[0048] Data lacking agricultural product batch identification will be supplemented and matched based on equipment installation location, adjacent scanning records, and preceding and following traceability events; if the batch attribution still cannot be determined, it will be marked as data to be associated and will not be included in the traceability data unit set. Data that fails digital signature verification, exceeds the equipment's measurement range, or has an incorrect format will be marked as invalid data.

[0049] Taking strawberry harvesting as an example, if the plot temperature data, harvesting data and first weighing data of the same batch carry the same batch identifier and the standard time difference does not exceed 30 minutes, they are classified into the same traceability event; if the weighing equipment has read the batch label within 1 minute, the batch attribution of the weighing data can be supplemented accordingly.

[0050] S120. For each traceability data unit in the traceability data unit set, extract causal lineage information along the sampling, generation, time synchronization, transmission, preprocessing and blockchain submission process to generate a causal lineage vector set.

[0051] A causal lineage vector is a structured data set that records the objects on which traceable data units depend from the original sampling to the blockchain submission in a fixed field order. The fields are, in order, the original sampling source identifier, the test sample identifier, the data generation device identifier, the device firmware identifier, the calibration device identifier, the clock source identifier, the gateway path, the preprocessor identifier, and the submission node identifier.

[0052] The original sampling source identifier and the test sample identifier are extracted from the sampling task record, the test sampling record, and the sample code; the data generation device identifier, the device firmware identifier, the calibration device identifier, and the clock source identifier are extracted from the device digital certificate, the firmware version register, the calibration log, and the time synchronization log; the gateway path, the preprocessor identifier, and the submission node identifier are extracted from the data packet header, the gateway communication log, the program call log, and the blockchain transaction initiation record.

[0053] The original sampling source is the object under test, the sampling location, or the test sample. The data generation device is the device that converts the original signal into digital data, and the two are recorded separately. If a certain field does not exist objectively during the corresponding data generation process, a null value flag is written; if it should exist according to the device configuration but is not obtained, a missing value flag is written.

[0054] For traceability data units that share at least one identical element among the original sampling source identifier, test sample identifier, calibration device identifier, clock source identifier, gateway identifier in the gateway path, or preprocessor identifier, an initial association is established, and the type of the same field is recorded. If only the device firmware version is the same but the data generating device is different, no initial association is established.

[0055] Taking cold storage temperature data as an example, after the temperature probe T01 samples the data, it is converted into digital data by the acquisition terminal D01. D01 runs firmware F03 and uses the calibration result of calibration device C02. After the data is timed by clock source TS01, forwarded by gateway G01 and processed by preprocessing program P02, it is submitted by submission node N01. The corresponding causal lineage vector is recorded in sequence as T01, null value identifier, D01, F03, C02, TS01, D01-G01-N01, P02 and N01.

[0056] S130. Perform normalized serialization on the traceability data unit set and the causal lineage vector set to generate evidence lineage records and write them to the blockchain, resulting in an on-chain evidence lineage record set. An evidence lineage record is a blockchain record to be written, formed by associating the traceability data unit summary, the causal lineage vector summary, the agricultural product batch identifier, the business process identifier, the standard time, and the original data storage location. During normalized serialization, the field order, character encoding, time format, numerical precision, null value identifier, and missing value identifier are standardized.

[0057] The joint summary of the kinship record of the i-th piece of evidence is calculated according to the following formula:

[0058]

[0059] In the formula, Let H represent the joint digest of the i-th evidence lineage record; H represents the pre-defined cryptographic hash function. This represents the i-th source data unit after normalization and serialization; This represents the i-th causal lineage vector after normalization and serialization; Indicates the corresponding agricultural product batch identifier; Indicates the corresponding business process identifier; This indicates that data is concatenated in the order it appears.

[0060] Raw data and complete device logs are stored on enterprise servers or distributed file systems. The raw data storage location consists of a storage system identifier, a data object identifier, a data version identifier, and an access path digest. The submitting node encapsulates the evidence lineage record into a blockchain transaction and digitally signs it. After verifying the submitting node's identity, record format, and combined digest, the blockchain node writes the data into a block and returns the block identifier and transaction identifier.

[0061] If the write operation fails, resubmit using the same evidence record identifier and joint summary; if both the evidence record identifier and joint summary already exist, it is considered a duplicate submission; if the joint summary is the same but the evidence record identifier is different, it is marked as a suspected duplicate record and retained.

[0062] An index is created based on the evidence record identifier, block identifier, transaction identifier, and original data storage location to form an on-chain evidence lineage record set. When reading off-chain original data, the digest is recalculated and compared with the on-chain digest; if they match, the corresponding evidence lineage record is allowed to proceed to subsequent processing; if they do not match, they are marked as mismatched records and are no longer used.

[0063] S2 specifically includes the following sub-steps:

[0064] S210: Read the on-chain evidence lineage record set formed in step S130. Recalculate the digest for the off-chain original data corresponding to each evidence lineage record. Only use the evidence lineage records whose off-chain digests match the on-chain original data digests for graph construction. Read the agricultural product batch identifier, business process identifier, standard time, and identifiers of each dependent object from the evidence lineage records and causal lineage vectors; read the equipment installation location, firmware effective time, calibration validity period, and program call order from the equipment logs, calibration logs, and program call logs with consistent digest verification.

[0065] Using traceable data units, original sampling sources, test samples, data generation devices, device firmware, calibration devices, clock sources, gateways, preprocessing programs, submission nodes, and on-chain evidence lineage records as nodes, and using collection relationships, detection relationships, calibration relationships, time synchronization relationships, forwarding relationships, processing call relationships, and on-chain submission relationships as directed edges, an evidence dependency knowledge graph is constructed.

[0066] The evidence dependency knowledge graph is a directed attribute graph that uses source data units and their dependent objects as nodes, direct dependencies as directed edges, and stores node type, object identifier, version identifier, and validity period. Directed edges point from upstream dependent objects to downstream results, recording the relationship type, occurrence time, and corresponding evidence record identifier. The unique identifier of a node is determined by its node type, object identifier, version identifier, and validity period; when the same device uses different firmware versions, separate device status nodes are established.

[0067] Null value identifiers do not create nodes. Missing identifiers are only used as independent missing attributes for the corresponding evidence of lineage records to avoid multiple missing fields being incorrectly grouped into a common source. If the downstream event time recorded by a directed edge is earlier than the upstream event time, and cannot be supplemented by cache or batch submission for interpretation, then the directed edge is marked as a time conflict edge and will not participate in subsequent independent path calculations.

[0068] Taking the cold storage temperature record R01 as an example, when the temperature probe T01 forms R01 through the acquisition terminal D01, gateway G01 and preprocessing program P02, the acquisition, forwarding, processing call and on-chain submission relationship are established in sequence; when D01 forms record R02 after the firmware is changed, a new device status node should be established.

[0069] S220. Using the evidence dependency knowledge graph as input, perform a reverse search along the directed edges from each on-chain evidence lineage record node to obtain the set of ancestor nodes for the corresponding source data unit. Ancestor nodes are upstream nodes that can be reached by the reverse search from the on-chain evidence lineage record nodes; the intersection of the ancestor node sets of different source data units constitutes the common ancestor node set.

[0070] A common ancestor node is identified only if it meets the following conditions: it is jointly depended on by two or more traceability data units; it belongs to the original sampling source, detection sample, data generation device, calibration device, clock source, gateway, or preprocessing program; it is in the same valid state during the corresponding data generation; and when the node fails, it can simultaneously affect the data value, generation time, source identity, or processing result of multiple traceability data units.

[0071] Based on the affected objects, common critical ancestor nodes are categorized into numerical, temporal, source, and processing types. Different devices using only the same model or firmware version, different calibration batches using only the same brand of standard material, or independently signed data submitted only through the same submission node are not considered to share a common critical ancestor node.

[0072] An initial root of evidence is established for each common critical ancestor node and all source data units reachable along the forward path. A root of evidence is a set of evidence dependencies consisting of a common critical ancestor node and the source data units it can influence. Two initial roots of evidence with the same influence type and overlapping member data are merged into a single root of evidence; those with different influence types are retained separately.

[0073] Using the multivariate dependencies between a common critical ancestor node and multiple affected source data units as hyperedges, an evidence dependency hypergraph is constructed; each hyperedge records the common critical ancestor node identifier, impact type, member data identifier, and effective time.

[0074] When records R01 and R02 are generated by different sensors but are both processed by gateway G01 which has not performed the original signature verification, G01 is a common key ancestor node of the processing class; when record R03 has been signed by the original device before entering G01, G01 is not considered a common key ancestor node of the data value of R03.

[0075] S230. Based on the evidence dependency hypergraph, determine the common key ancestor node, member data and multivariate dependency relationships corresponding to each evidence root. Then, starting from different original sampling sources or different detection sample nodes, follow the collection relationship, detection relationship, calibration relationship, time synchronization relationship, forwarding relationship, processing call relationship and on-chain submission relationship to reach the corresponding on-chain evidence lineage record node. Search for directed valid paths with consistent time order, nodes in valid time and not sharing common key ancestor nodes, and define them as independent paths.

[0076] Duplicate upload paths and cached retransmission transactions from the same original sampling source are counted as only one path, and paths containing time-conflicting edges are not included in the calculation. Different original sampling sources are connected to the virtual starting point, and the corresponding on-chain evidence lineage records are connected to the virtual ending point. Each common key ancestor node is split into input nodes and output nodes, and a connection edge with a capacity of 1 is set between them. The maximum flow value from the virtual starting point to the virtual ending point is calculated, and it is determined as the maximum number of independent paths for the corresponding evidence root.

[0077] Calculate the proportion of each original sampling source, detection sample, calibration device, and gateway path to the total number of member data, and sum them according to preset weights to obtain the normalized source difference degree. The causal independence of the r-th evidence root is calculated according to the following formula, where r represents the evidence root number:

[0078]

[0079] In the formula, This represents the causal independence of the r-th evidence root, with a value ranging from 0 to 1; Indicates the maximum number of independent paths; Indicates the number of member data; Indicates the degree of difference in origin; This represents the proportion of common critical ancestor nodes to the total number of valid ancestor nodes. This indicates the proportion of missing fields to the total number of fields that should be recorded; , , and This indicates that the weights of the four indicators are all greater than or equal to 0 and their sum is 1.

[0080] When the evidence root contains only one source data unit, the causal independence degree does not exceed the preset single-source upper limit; null value identifiers are not included in the missing proportion. The output includes evidence root identifiers, member data identifiers, common key ancestor node identifiers, maximum number of independent paths, source difference degree, missing proportion, causal independence degree, and a set of causal independence degrees for the knowledge graph version.

[0081] For example, if two of the four test records come from the same test sample and the other two come from different test samples, then the maximum number of independent paths is three; when two of these paths share the same calibration equipment, the causal independence decreases accordingly.

[0082] S3 specifically includes the following sub-steps:

[0083] S310. Read the causal independence set and evidence dependency knowledge graph output in step S230, and obtain the identifier, member data identifier, causal independence, common key ancestor node identifier and knowledge graph version of each evidence root; extract environmental, production, testing, warehousing, transportation and sales features from the traceability data unit formed in step S130 and whose summary is consistent, and obtain training labels from the confirmation results of testing institutions, the results of regulatory spot checks, recall records and the reviewed risk event records.

[0084] The training set and a fixed validation set are divided according to the batch identifier of agricultural products, and data from the same batch are only included in one set. A neural network containing an input layer, a feature extraction layer, and a risk classification output layer is used as the agricultural product risk identification model.

[0085] Specifically, the input layer receives normalized and serialized traceability data units; the feature extraction layer uses a one-dimensional convolutional neural network (1D-CNN), specifically consisting of two one-dimensional convolutional layers with a kernel size of 3 and one fully connected layer, with ReLU activation function used between layers; the risk classification output layer uses the Softmax function. A baseline risk identification model is obtained using the training set; the agricultural product risk identification model takes traceability data units verified for integrity as input and outputs the agricultural product safety risk category or risk probability.

[0086] Based on the frequency of historical risk events, the contribution of business process characteristics, the level of regulatory risk, and the number of affected batches, the importance of business processes is determined to be normalized to a value between 0 and 1. The model impact limit is the maximum normalized comprehensive model impact value that all traceable data units within a given evidence root are allowed to generate within one model update cycle.

[0087] The initial model impact limit for the r-th evidence root is determined according to the following formula:

[0088]

[0089] In the formula, This indicates that the initial model affects the credit limit; Indicates the degree of causal independence; Indicates the contribution of risk category distribution; Indicates the importance of each business process; , and This indicates that the weights of the three indicators are all greater than or equal to 0 and their sum is 1; in a preferred embodiment, the following settings are provided. =0.5、 =0.3、 =0.2; and These represent the lower and upper limits of the credit limit, respectively; for example, setting... =0.1, =1.0; This indicates that the calculation results are restricted to the upper and lower limits.

[0090] The contribution of risk category distribution is calculated based on the proportion of valid samples in each risk category to the total number of samples in the corresponding category within the evidence root. For example, the initial model impact of 10 temperature records sharing the same gateway may be less than that of 3 pesticide residue detection records from independent sources.

[0091] S320. Using the initial set of model influence values ​​as the initial search center for the Bat Algorithm, each individual Bat is defined as a candidate solution for model influence value. The dimension of the position vector is equal to the number of evidence roots, and the r-th component represents the candidate model influence value of the corresponding evidence root. The Bat Algorithm searches for the fitness-optimal solution among the candidate solutions by adjusting the search frequency, movement direction, and local search range.

[0092] Suppose that the search frequency, velocity component, and position component of the k-th bat individual in the g-th iteration are updated according to the following formula, where k represents the bat individual index and g represents the iteration index:

[0093]

[0094]

[0095]

[0096] In the formula, Indicates search frequency; and Indicates the search frequency boundary; Represents a random number in the range of 0 to 1; This indicates the direction and magnitude of the adjustment of the quota by the candidate model; This indicates the impact of the candidate model on the credit limit; This represents the model influence of the r-th evidence root in the current optimal candidate solution.

[0097] Perform a local search around the current best candidate solution and evaluate the candidate solution under the same training set, fixed validation set, benchmark risk identification model, training rounds and random seeds.

[0098] The fitness of the k-th candidate solution is calculated according to the following formula:

[0099]

[0100] In the formula, This represents the fitness score; a smaller value indicates a better candidate solution. This represents the average F1 score of a fixed validation set macro. This represents the imbalance loss resulting from the risk category identification. Indicates the displacement of the decision boundary; This indicates the extent to which the historical actual combined model influence of the evidence base exceeds the influence of the candidate model in the previous model update cycle; to This represents four weights, all greater than or equal to 0, and their sum is 1. In a preferred embodiment, the following is set... =0.4、 =0.2、 =0.2、 =0.2. The search stops when the maximum number of iterations is reached, or when the improvement in fitness after a consecutive preset number of iterations is less than the convergence threshold. The candidate solution with the lowest fitness is output, and the components of the current optimal candidate solution are... The optimized model that determines the corresponding evidence root affects the credit limit. This forms an optimization model that influences the set of credit limits.

[0101] S330. Using the set of influence of the optimized model as a constraint, perform deletion-based counterfactual inference for each evidence root. Deletion-based counterfactual inference is a process in which, under the condition that the initial parameters of the model, the number of training rounds, and the random seed are consistent, local retraining is performed using a training set containing the target evidence root data and a training set excluding the target evidence root data, and the results of the two models are compared. All source data units within the same evidence root participate in the inference as a whole to prevent the splitting of the same source data from circumventing the quota constraint.

[0102] Calculate the normalized norm of the difference between the two model parameter vectors, the change in the center vector of the feature extraction layer, and the average difference of the predicted probability vector of the fixed validation samples to obtain the predicted parameter influence, the predicted feature center shift, and the predicted decision boundary displacement, respectively.

[0103] Normalize the three indicators to 0 to 1, and calculate the estimated composite model impact of the r-th evidence root:

[0104]

[0105] In the formula, This indicates the estimated impact of the integrated model; Indicates the impact of the estimated parameters; This indicates the estimated feature center offset; This indicates the estimated decision boundary displacement; , and This indicates that the weights of the three indicators are all greater than or equal to 0 and their sum is 1. In a preferred embodiment, the following is set: =0.3、 =0.3、 =0.4.

[0106] When the estimated impact of the integrated model does not exceed the corresponding impact limit of the optimized model, the evidence root data is included in the candidate training set of the main model; otherwise, it is included in the dataset to be isolated. The output includes the evidence root identifier, causal independence, optimized model impact limit, three estimated impacts, estimated integrated model impact, data partitioning results, model version, and knowledge graph version of the evidence root estimated impact set.

[0107] For example, if the optimized model impact of a certain evidence root is 0.30 and the estimated comprehensive model impact is 0.365, its data will be included in the dataset to be isolated.

[0108] S4 specifically includes the following sub-steps:

[0109] S410: Read the causal independence set formed in step S230, the optimized model influence set formed in step S320, and the evidence root prediction influence set formed in step S330. Combine the causal independence of each evidence root, the optimized model influence, the three predicted influences, the predicted comprehensive model influence, the data partitioning results, and the historical training states of the previous model update cycle into a state vector, and normalize each state index to 0 to 1.

[0110] Historical training states are extracted from the model impact fingerprints confirmed by the blockchain, including the training control actions used in the previous model update cycle, the actual integrated model impact, the excess model impact identifier, the shadow model verification results, and the number of knowledge graph corrections. When this method is executed for the first time, all indicators in the historical training states are set to 0 or a preset neutral baseline value.

[0111] The reinforcement learning referred to in this embodiment is a reinforcement learning mechanism that makes training control decisions based on environmental state, control actions, reward values, and state transition relationships. A reinforcement learning training control model is constructed, preferably using the Deep Q-Network (DQN) algorithm, which includes a current Q-Network and a target Q-Network for evaluating the value of actions. The network structure contains 2 to 3 fully connected layers. The evidence root state is used as the environmental state, and the four mutually exclusive discrete control actions output by the Deep Q-Network algorithm are: directly participating in the training of the agricultural product risk identification model, reducing training weights, switching to the shadow model, and pausing training and reviewing.

[0112] When reducing the training weights, the training weight coefficient of the r-th evidence root is calculated according to the following formula:

[0113]

[0114] In the formula, This represents the training weight coefficient of the r-th evidence root; This indicates the impact of the corresponding optimization model on the credit limit; This indicates the impact of the corresponding comprehensive model prediction; This indicates a preset positive number to prevent the denominator from being 0.

[0115] After executing the control action, the reward value is calculated using a fixed validation set:

[0116]

[0117] In the formula, This represents the reward value obtained after performing a control action on the r-th evidence root; This represents the macro-average F1 score of a fixed validation set. Indicates the stability of the decision boundary; Indicates the degree to which anomalous evidence is correctly isolated; Indicates the degree of misisolation of normal evidence; This indicates the extent to which the actual impact of the integrated model exceeds the impact of the optimized model. to This indicates that the weights of the five indicators are all greater than or equal to 0.

[0118] In a preferred embodiment, the following is provided: =1.0、 =0.5、 =2.0、 =2.0、 =1.5, by assigning an isolation penalty term ( ) and excess penalty items ( Higher weights guide the model towards a more conservative isolation strategy. Abnormal evidence labels originate from confirmed results of regulatory spot checks, re-inspections by testing institutions, metrological calibration verifications, or safety audits; normal evidence labels originate from evidence roots that have undergone re-inspection and found no abnormalities; unconfirmed evidence roots are not included in the isolation degree calculation. The state, control action, reward value, and next state are written into the empirical dataset, and the reinforcement learning training control model is updated using empirical replay; when the maximum number of training rounds is reached or the average reward improvement over consecutive preset training rounds is less than the convergence threshold, the training control action sequence is output.

[0119] The verification identifier is output along with the corresponding control action to trigger equipment calibration verification, test sample re-inspection, or gateway security audit, and to generate the subsequent test results used in step S510.

[0120] S420. Update the agricultural product risk identification model and the shadow model according to the training control action sequence. The shadow model is a control model that uses the same model structure, initial parameters, main model candidate training set, training rounds, and random seeds as the agricultural product risk identification model, but additionally receives evidence root data transferred to the shadow model.

[0121] At the start of each model update cycle, the parameters of the previous effective model version are copied to the agricultural product risk identification model and the shadow model respectively; data that directly participates in the training of the agricultural product risk identification model are updated simultaneously in both models with preset training weights; data with reduced training weights are updated simultaneously in both models according to the corresponding training weight coefficients; data transferred to the shadow model are updated only in the shadow model; data that is paused for training and reviewed are not updated in either model.

[0122] Input the evidence root data one by one according to the training control action sequence. After each evidence root completes the local model update, use the same fixed validation set as in step S330 to calculate the difference in model parameter vectors before and after the update, the change in feature centers of each risk category, and the change in predicted probability vectors. This will yield the actual model parameter increments, actual feature center offsets, and actual decision boundary displacements, all normalized to 0 to 1.

[0123] The actual composite model impact of the r-th evidence root is calculated according to the following formula:

[0124]

[0125] In the formula, Indicates the actual impact of the integrated model; This represents the actual increment of model parameters; Indicates the actual offset of the feature center; This represents the actual displacement of the decision boundary; , and The corresponding weights from step S330 are used.

[0126] The model influence fingerprint is formed by combining the evidence root identifier, member data summary set, model version identifier before and after update, training control action, actual training weight, three actual effects, actual comprehensive model effect, optimized model effect amount, training configuration summary, and knowledge graph version.

[0127] The off-chain training log at least saves the parameters before and after model updates, optimizer status, training batch identifier and order, training epochs, learning rate, batch size, random seed, and model checkpoint storage location; the model impact fingerprint summary and training log storage location index are used for subsequent on-chain evidence storage. The shadow model and the agricultural product risk identification model maintain the same training conditions except for the data transferred to the shadow model, so that the output differences between the two models can be used to evaluate the independent impact of the evidence root.

[0128] S430. Compare the actual integrated model impact of each evidence root with the corresponding optimized model impact amount; if the actual integrated model impact exceeds the optimized model impact amount, or if the deviation between the actual integrated model impact and the estimated integrated model impact in step S330 exceeds the preset deviation threshold, trigger the evidence dependency knowledge graph reverse correction.

[0129] Read the causal lineage vector and off-chain device log saved in step S130, the knowledge graph nodes and directed edges saved in step S210, the model influence fingerprint formed in step S420, and the difference results of the two models, and check whether multiple abnormal evidence roots share network addresses, edge computing nodes, data processing services, calibration standard sources, or unrecorded preprocessing instances.

[0130] Calculate the cosine similarity between the actual model parameter increment vectors of any two abnormal evidence roots; when the cosine similarity reaches the preset directional similarity threshold, and the standard time intervals of the two evidence root member data have an intersection or the time difference does not exceed the preset time window of the corresponding business link, they are listed as candidates for latent common sources.

[0131] After a candidate relationship receives support from at least one of the following: device logs, call logs, calibration records, or network forwarding records, a common data source node, calibration standard source node, or preprocessing service node is added to the evidence dependency knowledge graph. Corresponding directed edges, validity periods, and relationship discovery criteria are also added, generating a new knowledge graph version and graph correction records. The original knowledge graph version is retained. After the update, the affected common key ancestor nodes and evidence root membership relationships are re-determined, and the affected evidence root identifiers are provided to step S230 of the next model update cycle.

[0132] Write the model impact fingerprint set summary, the updated evidence dependency knowledge graph summary, the graph correction record summary, the training control action sequence summary, the version identifiers of the two models, and the off-chain training log storage location index into the blockchain; after the preset confirmation conditions are met, input the model impact fingerprint set and the updated evidence dependency knowledge graph into step S5.

[0133] S5 specifically includes the following sub-steps:

[0134] S510. Obtain the subsequent test results generated after the formation of the evidence lineage record, and determine the invalid evidence root, invalid derived data set and affected model version set based on the subsequent test results.

[0135] Subsequent test results are generated by regulatory spot checks, re-inspections by testing institutions, equipment self-inspections, metrological calibration verifications, gateway security audits, program version audits, or manual on-site verifications. They are used to reconfirm the validity of the original sampling source, test sample, data generation equipment, calibration equipment, clock source, gateway path, or preprocessing procedure. They include at least the test result identifier, the identifier of the tested object, the object version identifier, the test time, the failure start time, the failure end time, the test conclusion, the identifier of the confirming institution, and a digital signature.

[0136] The verification and confirmation agency's digital signature is validated, and subsequent test results are determined to be valid only when any of the following conditions are met: direct confirmation of invalidity by a regulatory agency or an institution with testing qualifications; consistency of conclusions from two independent re-inspections; equipment calibration error exceeding the allowable range; or the security audit certificate object being tampered with during the corresponding time period.

[0137] The detected object identifier, object version identifier, and expiration time period from the valid subsequent detection results are matched with the node attributes in the updated evidence dependency knowledge graph. When a common key ancestor node becomes invalid throughout the entire valid time period, the corresponding evidence root is determined as a global invalidity evidence root; when only part of the time or part of the member data becomes invalid, the corresponding evidence root is determined as a partial invalidity evidence root. An invalidity evidence root is evidence that a common key ancestor node or member data no longer meets the requirements of authenticity, accuracy, or completeness, as confirmed by valid subsequent detection results.

[0138] Starting with the failure evidence root, a forward search is performed along data generation relationships, program processing relationships, and batch splitting, merging, and transfer relationships, retaining only valid directed edges whose time is later than the failure start time and are within the failure's impact range. The search stops at nodes where resampling, recalibration, program replacement, or summary re-verification passes. The searched data constitutes a failure-derived data set, and the impact path corresponding to each failure-derived data item is saved. Then, based on the model impact fingerprint, the model update versions involved in each failure-derived data item are queried, forming a set of affected model versions.

[0139] Taking the calibration device CO2 as an example, when CO2 is inaccurate from June 10 to June 15, only the data generated by referencing the calibration results of CO2 during that period and its derived data are included in the failure derived data set. The data generated after recalibration will not continue to propagate the failure status.

[0140] S520. Based on the set of failed derived data, the set of affected model versions, the model impact fingerprint, and the off-chain training logs, perform impact reversal to generate the agricultural product risk identification model after impact reversal and the impact reversal sequence. Impact reversal is the process of eliminating the historical training influence of failed derived data on the parameters, feature representations, and decision boundaries of the agricultural product risk identification model while preserving the blockchain's historical records.

[0141] Read the off-chain training log saved in step S420 and extract the model parameters, optimizer status, training configuration, training batch order, random seed, and model checkpoint storage location at the beginning and end of each model update cycle. Model checkpoints are datasets used to restore the corresponding model version parameters and training status.

[0142] Based on the set of affected model versions, determine the model version that participated in training earliest with the failed derived data, and determine the most recent model checkpoint before that version that was not affected by the failed derived data as the revocation start checkpoint; restore the model parameters and optimizer state from the revocation start checkpoint, read subsequent training batches according to the original model update order, exclude failed derived data from the training batches to be replayed, skip the training batches that do not contain valid training data after exclusion, and re-execute the model update with the remaining valid training data according to the original training rounds, learning rate, batch size and random seed.

[0143] The model parameters for the j-th reconstructed model version are determined according to the following formula, where j represents the model update version number:

[0144]

[0145] In the formula, This indicates the model parameters that affect the j-th model version after the revocation; This indicates the model parameters of the previous reconstructed model version; This represents the set of model update versions corresponding to the r-th failure evidence root, formed solely from failure-derived data; This represents the valid training data retained after excluding invalid derived data from the j-th original training batch; This indicates a model update operation that is identical to the original training process.

[0146] A sequence of impact reversals is generated, including the invalidation evidence root identifier, the reversal start checkpoint, the skipped model version, the modified training batch, the replayed valid training batch, and the reconstructed model version. The macro-average F1 score, recall rate for each risk category, and output stability of the reconstructed model are verified using a fixed validation set. If all these meet the preset lower limits, the model is identified as the agricultural product risk identification model after impact reversal; otherwise, the reversal is reverted to an earlier model checkpoint and the impact reversal is re-executed.

[0147] When failed derived data is used for training for the first time in model version M05, M04 should be restored and the valid training batches from M06 to M08 should be replayed, instead of simply subtracting the corresponding parameter increments from M05 from M08.

[0148] S530. Using the agricultural product risk identification model after the impact is revoked, the safety risks of the agricultural product batches directly corresponding to the invalid evidence roots, as well as the split batches, merged batches and subsequent circulation batches formed therefrom, are re-identified, and an updated agricultural product safety risk status and impact revocation certificate are generated.

[0149] Based on the impact path saved in step S510 and the parent batch identifier, child batch identifier, merged batch identifier, and parent batch input quality recorded in step S110, the set of batches that need to be re-identified is determined. For each affected batch, the traceability data units that have been verified to be consistent with the summary and are not included in the failure-derived data set, the updated evidence dependency knowledge graph, and the batch relationships are read and input into the agricultural product risk identification model after the impact is withdrawn to obtain the probability of each risk category.

[0150] No. The comprehensive risk value of each batch of agricultural products is calculated according to the following formula, where, Indicates the batch number of agricultural products:

[0151]

[0152] In the formula, Indicates the first The comprehensive risk value of each batch of agricultural products; c represents the total number of risk categories; c represents the risk category number. Indicates the first The probability that a batch of agricultural products belongs to the c-th risk category; This represents the severity weight of the c-th risk category.

[0153] The severity weights of risks are pre-determined and normalized based on regulatory limits, historical recall levels, hazard levels, and historical loss data. The weights and version identifiers are written into the model configuration record. Based on risk thresholds determined and versioned by historically verified data and regulatory limits, the safety risk status of agricultural products is classified into normal, re-inspection, isolation, or recall.

[0154] The risk evidence corresponding to the material source of the split batch is inherited and re-identified in combination with its own valid data; when merging batches, the risk status and input quality ratio of all parent batches are read at the same time. If any parent batch has a high degree of risk and lacks valid re-inspection results, the merged batch shall not be judged as normal.

[0155] The invalidation evidence root identifier, valid subsequent test result summary, invalidation time period, invalidation derived data set summary, affected model version set summary, updated evidence dependency knowledge graph version, revocation start checkpoint identifier, impact revocation sequence summary, model versions before and after revocation, model verification result summary, affected batch identifier set, and updated agricultural product safety risk status summary are combined into an impact revocation proof, which is then digitally signed by the submitting node and written into the blockchain.

[0156] Once the preset confirmation conditions are met, the model version affected by the revocation and the updated agricultural product safety risk status will be set to the current valid version. When multiple invalid evidence roots trigger the revocation at the same time, the earliest affected model version will be executed uniformly to avoid repeated reloads and replays.

[0157] Example 2: Figure 2 As shown, this embodiment provides an agricultural product safety traceability system based on blockchain and the Internet of Things, including:

[0158] The data acquisition and storage module is used to acquire IoT sensing data and detection data from the production, testing, storage, transportation and sales of agricultural products. It extracts the original sampling source, test sample, data generation equipment, calibration equipment, clock source, gateway path, preprocessing program and submission node, generates causal lineage vector, and writes the causal lineage vector and data summary into the blockchain.

[0159] The evidence dependency analysis module is used to construct an evidence dependency knowledge graph based on causal lineage vectors, identify common key ancestor nodes, group data affected by the same common key ancestor node into evidence roots, and calculate the causal independence of each evidence root based on the number of independent paths.

[0160] The model impact optimization module is used to allocate the model impact amount according to the degree of causal independence, optimize the model impact amount using the bat algorithm, and calculate the estimated comprehensive model impact of each evidence root on the agricultural product risk identification model through deletion counterfactual inference, forming the main model candidate training set and the dataset to be isolated;

[0161] The training control correction module is used to determine training control actions based on the impact of the optimized model and the estimated impact of the comprehensive model, update the agricultural product risk identification model and shadow model, generate model impact fingerprints, and reverse correct the evidence dependency knowledge graph when the actual impact of the comprehensive model exceeds the impact of the optimized model.

[0162] The failure impact revocation module is used to determine the failure-derived data and affected model version along the corrected evidence dependency knowledge graph when the evidence root is determined to be invalid in subsequent detection results. Based on the model impact fingerprint, checkpoint reload and effective training batch replay are performed to obtain the agricultural product risk identification model after impact revocation, redetermine the safety risk status of the affected batch, and write the impact revocation proof to the blockchain.

[0163] It should be noted that the various weighting coefficients, penalty term weights, and set parameters in the formulas involved in the above embodiments of the present invention (such as...) , , , The values ​​mentioned (e.g., series A, series A, etc.) are merely preferred illustrative examples. Those skilled in the art can adapt these values ​​to specific agricultural product categories, data scales, and actual computing power constraints during implementation. These parameter optimizations based on conventional experiments do not depart from the scope of this invention.

[0164] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0165] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0166] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for tracing the safety of agricultural products based on blockchain and the Internet of Things, characterized in that, Includes the following steps: S1. Acquire IoT sensing data and detection data from agricultural product production, testing, storage, transportation and sales, extract original sampling sources, test samples, data generation equipment, calibration equipment, clock sources, gateway paths, preprocessing programs and submission nodes, generate causal lineage vectors and associate them with data summaries to write to the blockchain; S2. Construct an evidence dependency knowledge graph based on the causal lineage vector, identify common key ancestor nodes, group data affected by the same common key ancestor node into evidence roots, and calculate causal independence based on the number of independent paths. S3. The impact of the model is allocated according to the degree of causal independence, and the Bat algorithm is used for optimization. The impact of evidence roots on the predicted comprehensive model of agricultural product risk identification model is calculated by deletion counterfactual inference, forming a candidate training set and a dataset to be isolated. S4. Based on the impact of the optimized model and the estimated impact of the comprehensive model, use reinforcement learning to determine the training control actions, update the agricultural product risk identification model and shadow model, and generate the model impact fingerprint; when the actual comprehensive model impact exceeds the limit, reverse-correct the evidence dependency knowledge graph.

2. The method for agricultural product safety traceability based on blockchain and the Internet of Things according to claim 1, characterized in that, Also includes: S5. When subsequent test results determine that the evidence root is invalid, the invalid derived data and the affected model version are determined along the corrected evidence dependency knowledge graph. Based on the model impact fingerprint, checkpoint reloading and effective training batch replay are performed to obtain the agricultural product risk identification model after the impact is revoked. The safety risk status of the affected batch is determined and the impact revocation proof is written into the blockchain.

3. The method for agricultural product safety traceability based on blockchain and the Internet of Things according to claim 1, characterized in that, S1 specifically includes: Data from production, testing, warehousing, transportation and sales processes are acquired, agricultural product batch identifiers are generated, equipment recording times are corrected, and data is aggregated into traceability data unit sets according to batch identifiers, business processes and time windows. From the sampling records and related logs, extract the original sampling source, test sample, data generation device, calibration device, clock source, gateway path, preprocessing program and submission node to generate a causal lineage vector set.

4. The method for agricultural product safety traceability based on blockchain and the Internet of Things according to claim 3, characterized in that, Also includes: The source data unit set and causal lineage vector set are normalized and serialized, and a joint digest is calculated. The joint digest is associated with the batch identifier, business process identifier and original data storage location to form an evidence lineage record. After signing, it is written into the blockchain to obtain the on-chain evidence lineage record set.

5. The method for agricultural product safety traceability based on blockchain and the Internet of Things according to claim 1, characterized in that, S2 specifically includes: Read the on-chain evidence lineage record set and corresponding logs, and construct an evidence dependency knowledge graph with source data units and dependent objects as nodes and collection, detection, calibration, time synchronization, forwarding, processing calls and on-chain submission relationships as directed edges; Ancestor nodes are searched in reverse along the evidence dependency knowledge graph to identify common key ancestor nodes, which are then merged with the affected source data units into evidence roots, and an evidence dependency hypergraph representing multivariate dependencies is constructed. Based on evidence-dependent hypergraph search, independent paths that do not share a common critical ancestor node are calculated, and the maximum number of independent paths, the degree of source difference, the proportion of common critical ancestor nodes, and the proportion of missing nodes are obtained to obtain the set of causal independence for each evidence root.

6. The method for agricultural product safety traceability based on blockchain and the Internet of Things according to claim 1, characterized in that, S3 specifically includes: Read the set of causal independence, divide it into training and validation sets, and generate an initial model impact limit set based on causal independence, risk category distribution contribution, and the importance of business links; Using the initial model influence limit set as the search center, the influence limit of candidate models is encoded as the position of individual bats. The bat algorithm is used to evaluate candidate solutions based on risk identification performance, decision boundary displacement, and limit violation degree to obtain the optimized model influence limit set.

7. The method for agricultural product safety traceability based on blockchain and the Internet of Things according to claim 6, characterized in that, Also includes: For each evidence root, perform deletion-based counterfactual inference, calculate the estimated impact of parameters, feature centers, and decision boundaries, generate the estimated comprehensive model impact, and compare it with the impact of the optimized model to form the main model candidate training set and the dataset to be isolated.

8. The method for agricultural product safety traceability based on blockchain and the Internet of Things according to claim 1, characterized in that, S4 specifically includes: Read the causal independence set, the optimized model influence set, and the evidence root prediction influence set, and combine them with the historical training state to form a state vector. Use reinforcement learning to train the control model and output the training control action sequence. The agricultural product risk identification model and shadow model are updated based on the training control action sequence. The model parameter increment, feature center offset and decision boundary displacement of the evidence roots are calculated to generate the comprehensive model impact and model impact fingerprint. The impact of the comprehensive model is compared with the impact of the optimized model. If the impact exceeds the limit, the implicit common source is identified based on the causal lineage vector, device logs and model impact fingerprint. The evidence dependency knowledge graph is then corrected and the model impact fingerprint and correction results are written into the blockchain.

9. The method for agricultural product safety traceability based on blockchain and the Internet of Things according to claim 2, characterized in that, S5 specifically includes: Obtain subsequent detection results, match object identifiers, version identifiers and failure time periods with the evidence dependency knowledge graph, determine the failure evidence root, search for failure-derived data along effective dependency relationships and query the affected model version; Based on the failure-derived data, the affected model version, the model impact fingerprint, and the off-chain training logs, the initial checkpoint for revocation is determined, the model and optimizer states are restored, and after excluding the failure-derived data, the valid training batches are replayed in the original order to obtain the agricultural product risk identification model after the impact is revoked. The affected batches are identified based on the impact path and batch relationships of the failure-derived data. The safety risk status is redefined using an agricultural product risk identification model, and the impact reversal proof is written into the blockchain.

10. A blockchain- and IoT-based agricultural product safety traceability system, employing the blockchain- and IoT-based agricultural product safety traceability method as described in any one of claims 1 to 9, characterized in that... include: The data acquisition and storage module is used to acquire IoT sensing data and detection data from the production, testing, storage, transportation and sales of agricultural products. The evidence dependency analysis module is used to construct an evidence dependency knowledge graph based on causal lineage vectors, identify common key ancestor nodes, and group data influenced by the same common key ancestor node into evidence roots. The model impact optimization module is used to allocate the model impact amount according to the degree of causal independence, optimize the model impact amount using the bat algorithm, and calculate the estimated comprehensive model impact of each evidence root on the agricultural product risk identification model through deletion counterfactual inference. The training control correction module is used to determine training control actions based on the impact of the optimized model and the estimated impact of the comprehensive model, and to update the agricultural product risk identification model and shadow model using reinforcement learning. The Failure Impact Revocation Module is used to determine the failure-derived data and affected model versions along the corrected evidence dependency knowledge graph when subsequent detection results determine that the evidence root is invalid.