Knowledge graph-driven food safety risk tracing method and system

By constructing a digital twin of food safety and conducting causal mining to generate a dynamic evolutionary knowledge graph, the problem of the lack of deep causal mechanism modeling in existing food safety traceability methods is solved, enabling accurate risk identification and prediction, and improving the level of intelligence in food safety traceability.

CN121787905AInactive Publication Date: 2026-04-03大理白族自治州检验检测院
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing food safety traceability methods lack the ability to model and reason about the complex and deep-seated causal mechanisms behind risk events, making it difficult to proactively discover potential or unknown risks involving multiple factors and cross-process coupling, and limiting the depth of analysis.

Method used

By employing a knowledge graph-driven approach, a digital twin of food safety is constructed, semantic mapping and causal mining are performed, a dynamically evolving knowledge graph is generated, risk identification and causal reasoning are conducted, a target risk transmission chain is generated, and cross-validation and counterfactual simulation are performed to quantify key risk attribution factors, thereby achieving self-verification and optimization of the closed-loop intelligent architecture.

Benefits of technology

It significantly improves the accuracy and foresight of food safety risk tracing, enabling the discovery and quantification of deep-seated potential risks, and realizing the transition from "correlation backtracking" to "causal insight and risk prediction." Furthermore, it continuously enhances the intelligence level of tracing and decision-making through feedback mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121787905A_ABST
    Figure CN121787905A_ABST
Patent Text Reader

Abstract

The invention discloses a food safety risk tracing method and system driven by a knowledge graph, and belongs to the technical field of artificial intelligence and data science, and the method comprises the steps: carrying out the standardized processing of multi-source heterogeneous data, and forming a real-time data flow; constructing a digital twin based on the real-time data stream and outputting a dynamic evolution knowledge graph; performing risk identification and causal reasoning on the knowledge graph to generate a target risk conduction chain; performing anti-fact simulation on the conduction chain to quantify the key risk attribution factor; and executing management and control based on the attribution factor and feeding back a disposal effect to the model. According to the method, a closed-loop intelligent architecture which deeply fuses digital twinning, causal inference and anti-fact analysis is adopted, and deep potential risks formed by multi-factor coupling can be found and quantified by deducing, attributing and verifying the risks in a virtual space, so that the risk analysis accuracy is improved. The accuracy and foresight of food safety risk traceability and the intelligent level of decision making are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and data science, and in particular to a knowledge graph-driven method and system for tracing food safety risks. Background Technology

[0002] Food safety is a critical area related to public health and social stability. The modern food industry is characterized by a long supply chain, numerous links, and heterogeneous data. To achieve effective monitoring of the food distribution process, existing technologies typically focus on building information-based traceability systems. By recording key data from production, processing, and logistics, these systems establish connections between product batches and nodes in the supply chain. This provides fundamental technical support for tracing the source of contamination after a food safety incident.

[0003] However, most existing source tracing methods rely on correlation finding of existing data, with limited analytical depth. They generally lack the ability to model and reason about the complex, deep-seated causal mechanisms behind risk events, making it difficult to proactively discover potential or unknown risks caused by multi-factor, cross-stage coupling. Therefore, how to upgrade from "correlation backtracking" to "causal insight and risk prediction" is the current technical bottleneck facing the field. Summary of the Invention

[0004] To address the aforementioned issues, this invention provides a knowledge graph-driven method and system for tracing food safety risks. It employs a closed-loop intelligent architecture that deeply integrates digital twins, causal inference, and counterfactual analysis. By extrapolating, attributing, and verifying risks in a virtual space, it can discover and quantify deep-seated potential risks resulting from the coupling of multiple factors, significantly improving the accuracy, foresight, and intelligence of food safety risk tracing and decision-making.

[0005] The above objectives can be achieved through the following approach: A knowledge graph-driven food safety risk tracing method includes: acquiring and integrating multi-source heterogeneous data from raw material, production, distribution, and consumption ends; performing unified and standardized processing to form a real-time data stream; constructing a food safety digital twin based on the real-time data stream; using the food safety digital twin to perform semantic mapping and causal mining to output a dynamic evolutionary knowledge graph; performing risk identification and causal reasoning on the dynamic evolutionary knowledge graph to obtain a set of potential risk nodes and causal propagation paths, and performing cross-validation to generate a target risk transmission chain; performing temporal consistency verification on the target risk transmission chain; performing multi-scenario counterfactual simulation on the verified risk transmission chain to quantify and output key risk attribution factors; and based on the key risk attribution factors, implementing control measures and monitoring the disposal effects, and providing feedback to the dynamic evolutionary knowledge graph and the food safety digital twin.

[0006] Optionally, forming a real-time data stream includes: performing named entity recognition and event extraction on unstructured text data in the multi-source heterogeneous data to extract and structure food safety event metadata; performing multi-scale time-series decomposition on sensor time-series data in the multi-source heterogeneous data to separate and quantify trend-seasonal components and residual components; and performing context association and feature embedding on the food safety event metadata, the trend-seasonal components, the residual components, and the structured data in the multi-source heterogeneous data to generate a real-time data stream.

[0007] Optionally, the output dynamic evolutionary knowledge graph includes: mapping the real-time data stream to a computable twin object of the food safety digital twin, and instantiating an initial causal knowledge graph based on the initial association between the food safety digital twins; performing active perturbation and prospective simulation on the computable twin objects to generate virtual process data, and performing causal effect analysis on the virtual process data to generate and verify implicit causal relationships; incrementally updating the topology and relation weights of the initial causal knowledge graph as new knowledge, iteratively evolving and outputting the dynamic evolutionary knowledge graph.

[0008] Optionally, the instantiation of the initial causal knowledge graph includes: semantically aligning the computable twin object with a preset food safety domain ontology to generate an ontology-enhanced twin entity; using the ontology-enhanced twin entity as contextual hints and inputting it into a large language model to generate hypothetical causal triples; and fusing the deterministic causal path of the food safety domain ontology with the hypothetical causal triples to construct and instantiate the initial causal knowledge graph.

[0009] Optionally, generating the target risk transmission chain includes: based on the dynamic evolutionary knowledge graph, performing symbolic reasoning and subgraph anomaly detection processes in parallel to generate a first candidate risk path and a second high-risk node, respectively; performing set operations on the first candidate risk path and the second high-risk node to identify and extract a set of conflict nodes; performing causal attribution verification on the set of conflict nodes, and integrating all verified paths to generate the target risk transmission chain.

[0010] Optionally, the causal attribution verification for the conflict node set includes: performing factual scenario deduction based on the initial conditions of the food safety digital twin for the conflict node set to generate a baseline risk evolution trajectory; performing forced risk injection and forced security constraints on the state of the conflict node set in the food safety digital twin, and performing two independent counterfactual scenario deductions to generate a maximum risk evolution trajectory and a minimum risk evolution trajectory, respectively; and obtaining the causal attribution verification result by quantitatively comparing the differences between the baseline risk evolution trajectory, the maximum risk evolution trajectory, and the minimum risk evolution trajectory.

[0011] Optionally, the quantification of key risk attribution factors includes: extracting multi-dimensional static attributes from the dynamic evolutionary knowledge graph based on the risk transmission chain to obtain static risk weights; performing attenuation accumulation calculations on the static risk weights to quantify and determine the risk contribution of each upstream node to the downstream node, and taking the entity node and relation edge with the largest contribution as key risk attribution factors.

[0012] Optionally, the step of feeding back the dynamic evolutionary knowledge graph and the food safety digital twin includes: using the treatment effect as physical world feedback data, comparing it with the prospective prediction results of the food safety digital twin, quantifying and generating a cognitive-reality bias signal; performing causal attribution backtracking based on the cognitive-reality bias signal, and updating the posterior probability of the dynamic evolutionary knowledge graph; and using the cognitive-reality bias signal to calibrate the internal dynamic model parameters in the food safety digital twin online.

[0013] Optionally, the method further includes: summarizing and pattern mining the key risk attribution factors to generate a structural risk pattern library; associating and matching the structural risk pattern library with the conflict node set to identify and determine high-priority reconstruction targets; and using the high-priority reconstruction targets as core contexts and constraints to reconstruct and instantiate the initial causal knowledge graph.

[0014] Based on the same inventive concept, this invention also provides a knowledge graph-driven food safety risk traceability system. The system includes: a full-link data aggregation and streaming module for acquiring and integrating multi-source heterogeneous data from raw material, production, distribution, and consumption ends, performing unified and standardized processing to form a real-time data stream; a twin-graph collaborative modeling module for constructing a food safety digital twin based on the real-time data stream, and using the food safety digital twin to perform semantic mapping and causal mining, outputting a dynamic evolutionary knowledge graph; a multi-channel risk reasoning and verification module for performing risk identification and causal reasoning on the dynamic evolutionary knowledge graph, obtaining a set of potential risk nodes and causal propagation paths, and performing cross-validation to generate a target risk transmission chain; a counterfactual attribution quantification module for performing temporal consistency verification on the target risk transmission chain, performing multi-scenario counterfactual simulations on the verified risk transmission chain, and quantifying and outputting key risk attribution factors; and a decision control and closed-loop evolution module for performing control and monitoring of the disposal effect based on the key risk attribution factors, and feeding back the dynamic evolutionary knowledge graph and the food safety digital twin.

[0015] Compared with the prior art, the present invention has the following advantages: 1. This invention elevates the paradigm of risk analysis from "static correlation retrospection" of historical data to "dynamic causal deduction" of physical laws and logical relationships in virtual space by constructing a digital twin of food safety and utilizing it for causal mining. This model can discover and verify implicit causal chains that are difficult to detect in historical data due to multi-factor, cross-stage coupling, fundamentally improving the depth and foresight of risk identification; 2. This invention, by introducing a feedback mechanism for the handling effect and a co-evolution mechanism for the model, changes the limitation of the traditional source tracing system's "fixed model," and constructs an intelligent learning closed loop capable of self-verification and capability evolution. The system can transform each "handling experience" into synchronous optimization of the knowledge graph and digital twin, realizing the leap from a "one-time analysis tool" to a "continuous learning partner," ensuring that the accuracy of source tracing and decision-making continuously improves over time.

[0016] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating a knowledge graph-driven food safety risk tracing method according to an embodiment of the present invention.

[0019] Figure 2 This is a structural evolution stacking area diagram of the dynamic evolutionary knowledge graph according to an embodiment of the present invention.

[0020] Figure 3 This is a cross-validation matrix diagram of the dual-channel risk reasoning results in an embodiment of the present invention.

[0021] Figure 4 This is a comparison diagram of multiple counterfactual scenarios in an embodiment of the present invention.

[0022] Figure 5 This is a schematic diagram of the structure of a knowledge graph-driven food safety risk traceability system according to an embodiment of the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] Reference Figure 1 One embodiment of the present invention proposes a knowledge graph-driven food safety risk tracing method, which adopts a closed-loop intelligent architecture that deeply integrates digital twins, causal inference and counterfactual analysis. By extrapolating, attributing and verifying risks in virtual space, it can discover and quantify deep-seated potential risks that are coupled by multiple factors, which significantly improves the accuracy, foresight and intelligence of food safety risk tracing and decision-making.

[0025] The method described in this embodiment specifically includes: Acquire and integrate heterogeneous data from multiple sources, including raw materials, production, distribution, and consumption, and process them in a unified and standardized manner to form a real-time data stream; A food safety digital twin is constructed based on the real-time data stream, and semantic mapping and causal mining are performed using the food safety digital twin to output a dynamic evolutionary knowledge graph. Risk identification and causal reasoning are performed on the dynamic evolutionary knowledge graph to obtain a set of potential risk nodes and causal propagation paths, and cross-validation is performed to generate a target risk transmission chain; Perform time-series consistency verification on the target risk transmission chain, perform multi-scenario counterfactual simulation on the risk transmission chain that passes the verification, and quantify and output key risk attribution factors; Based on the key risk attribution factors, control measures are implemented and the effectiveness of the measures is monitored, and feedback is provided to the dynamic evolutionary knowledge graph and the food safety digital twin.

[0026] By employing a closed-loop intelligent architecture that deeply integrates digital twins, causal inference, and counterfactual analysis, and by extrapolating, attributing, and verifying risks in virtual space, it can discover and quantify deep-seated potential risks that are coupled by multiple factors, significantly improving the accuracy, foresight, and intelligence of food safety risk traceability and decision-making.

[0027] Optionally, forming a real-time data stream includes: Named entity recognition and event extraction are performed on the unstructured text data in the multi-source heterogeneous data to extract and structure food safety event metadata; Specifically, this step aims to automatically extract valuable event information related to food safety from massive amounts of text information without a fixed format. A text processing module handles this task, processing data from various channels, such as government recall notices, industry news reports, or consumer feedback on social media. This module first employs Named Entity Recognition (NER) technology to identify key entities mentioned in the text, such as product names, involved companies, and geographical locations. Subsequently, an event extraction model further analyzes the relationships between these entities to determine the specific event types they are involved in, such as "product recall," "contaminant detection," or "complaint." Through this process, a raw natural language text, such as "The Market Supervision Bureau of City A announced today that batch C of yogurt produced by Company B has been ordered to be removed from shelves due to excessive yeast levels," is automatically and structurally transformed into food safety event metadata containing core elements, providing crucial input for subsequent knowledge graph construction.

[0028] Multi-scale time-series decomposition is performed on the sensor time-series data in the multi-source heterogeneous data to separate and quantify the trend-seasonal component and the residual component. Specifically, this step aims to deeply understand the underlying patterns and anomalous changes from continuous, fluctuating sensor readings. A time-series analysis module processes sensor time-series data from various stages of the supply chain, such as temperature records from cold chain transport vehicles or humidity records from production workshops. This module employs a time-series decomposition algorithm, such as Seasonal-Trend decomposition using Loess (STL), to break down a raw time series into three independent components. The first part is the trend component, which reflects the overall direction of change in the data over a longer time scale, such as a slow, systematic increase in temperature due to equipment aging. The second part is the seasonal component, which captures recurring fluctuations with fixed periods in the data, such as the regular temperature fluctuations caused by the fixed daily defrosting cycle of cold storage. The third part is the residual component, which is the part remaining after removing the trend and seasonal effects from the original series, typically representing random noise or unexpected sudden anomalies. Through this decomposition, it is possible to separate and quantify the inherent long-term patterns and sudden changes from seemingly messy data.

[0029] The food safety event metadata, the trend-seasonal component, the residual component, and the structured data from the multi-source heterogeneous data are context-associated and feature-embedded to generate a real-time data stream.

[0030] Specifically, this step is the final stage in generating a highly condensed data stream that can be used by subsequent models. A feature engineering module deeply integrates the data generated in the preceding steps with the original structured data. This module employs a feature embedding technique, such as Word2Vec or more advanced graph embedding methods, to map all data from different sources and of different types into a unified, high-dimensional vector space. In this space, the complex relationships between different data can be measured through the distance or angle between vectors. Through this process, all the original, scattered data is transformed into a high-dimensional data tensor rich in deep semantics and dynamic temporal features, serving as a real-time data stream.

[0031] Optionally, the output dynamic evolutionary knowledge graph includes: The real-time data stream is mapped to a computable twin object of the food safety digital twin, and an initial causal knowledge graph is instantiated based on the initial association between the food safety digital twins. Specifically, this step aims to transform the input data stream into a dynamic, simulation-enabled digital model. A twin modeling module is responsible for this task. It instantiates each physical or logical entity described in the real-time data stream into one or more computable twin objects within the digital twin environment. Each twin object not only contains its corresponding static attributes and dynamic state data but also encapsulates methods or physical models capable of simulating its behavior and state evolution. For example, a "cold storage" twin object not only has the state attribute of "temperature" but also a heat transfer model that can simulate its internal temperature changes under different external temperatures and door opening / closing frequencies. Subsequently, based on the known, preliminary relationships between these twin objects, an initial causal knowledge graph is instantiated as the symbolic reasoning layer of the digital twin.

[0032] Active perturbation and prospective simulation are performed on the computable twin object to generate virtual process data, and causal effect analysis is performed on the virtual process data to generate and verify implicit causal relationships; Specifically, this step is a core innovation of this invention, aiming to proactively uncover and verify hidden causal relationships in data through "virtual experiments." A causal discovery module is responsible for performing this task. This module selects key parameters of one or more twin objects as intervention variables and actively, algorithmically perturbs them. Subsequently, the module performs two parallel prospective simulations in the digital twin environment: one is a factual simulation based on observational data (ControlGroup), and the other is an intervention simulation with parameter perturbations (Treatment Group). Through multiple simulation runs, the module obtains two sets of virtual process data. Next, the module quantifies the average causal effect (ACE) of the intervention variable on the outcome variable by calculating the difference in expected values ​​of a downstream outcome variable in the two sets of simulation results. An exemplary ACE calculation formula is as follows: , in, This refers to the quantified causal effect, which represents the strength of the verified implicit causal relationship. It is the outcome variable; It is an intervention variable; Represents the variable The value was executed Intervention procedures; Representing variables Maintain its observed values; This represents the mathematical expectation of the simulation results. If the calculated average causal effect is statistically significant and not zero, a new implicit causal relationship is generated and verified.

[0033] The implicit causal relationships are treated as new knowledge and incrementally updated into the topology and relation weights of the initial causal knowledge graph. The graph is then iteratively evolved and output as a dynamically evolving knowledge graph.

[0034] Specifically, this step completes the closed-loop process of knowledge graph "evolution." A graph evolution module receives the implicit causal relationship verified in the previous step and its corresponding average causal effect value. If the causal relationship does not exist in the initial causal knowledge graph, the module creates a new relation edge representing the causal relationship between the corresponding entity nodes, thereby updating the topology of the knowledge graph. If the causal relationship already exists as a hypothesis in the initial graph, the module uses the calculated average causal effect value to update the weight or confidence parameter of the relation edge. Through this incremental, virtual experiment-driven update mechanism, the knowledge graph can continuously iterate and evolve, and its description of causal laws in the real world becomes increasingly accurate and complete. The final output is a dynamically evolving knowledge graph containing the latest knowledge, such as... Figure 2 The diagram illustrates how, as iterations proceed, a "hypothetical causal relationship" gradually transforms into a "verified implicit causal relationship" through verification in a digital twin, while another part is disproven and removed.

[0035] Optionally, the instantiation of the initial causal knowledge graph includes: The computable twin object is semantically aligned with a preset food safety domain ontology to generate an ontology-enhanced twin entity; Specifically, this step aims to endow each digital twin object with rich, scientifically axiom-compliant background knowledge and logical constraints. A semantic alignment module is responsible for this task. This module first loads a pre-built food safety domain ontology. This ontology is a formalized knowledge base that defines the core concepts, attributes, and logical relationships between concepts within the food safety domain. The semantic alignment module analyzes the metadata of the computable twin object and precisely links or "aligns" it with the corresponding concepts in the ontology. For example, a twin object representing "production batch A," if its production process includes "pasteurization," will inherit all axiomatic attributes and constraints related to "pasteurization" from the ontology. Through this step, the originally isolated twin object is transformed into an ontology-enhanced twin entity containing deep domain knowledge.

[0036] The ontology-enhanced twin entity is used as a contextual cue and input into a large language model to generate hypothetical causal triples; Specifically, this step is a core innovation of this invention, aiming to leverage the creative capabilities of artificial intelligence to compensate for the shortcomings of formal knowledge and proactively discover potential, unknown causal relationships. A causal hypothesis generation module performs this task. This module takes one or more ontology-enhanced twin entities generated in the previous step and their associated background knowledge, constructs a structured natural language text as a context prompt, and inputs it into a Large Language Model (LLM). The module poses exploratory questions to the LLM, such as: "Based on the following information, please speculate on potential, yet undocumented, causes that may have shortened the shelf life of this batch of products." The LLM uses its "world knowledge" learned from massive amounts of text data to generate a series of possible, hypothetical causal statements. For example, the LLM might generate: "A prolonged rainy season leads to excessive humidity in the packaging box supplier's warehouse, which may cause slight mold growth in the boxes, thereby affecting the product through contact contamination." These natural language statements are then parsed and converted into a standard "head entity-relationship-tail entity" form, i.e., hypothetical causal triples.

[0037] By integrating the deterministic causal paths of the food safety ontology with the hypothetical causal triples, an initial causal knowledge graph is constructed and instantiated.

[0038] Specifically, this step aims to construct an initial knowledge graph containing both rigorous scientific knowledge and exploratory hypotheses, representing a mixture of beliefs. A graph construction module performs this fusion process. On one hand, it extracts all known, deterministic causal paths from the food safety ontology. On the other hand, it also takes as input the hypothetical causal triples generated by LLM in the previous step. When constructing the final graph, this module assigns different initial confidence weights to the relation edges from these two different sources. Deterministic paths from the ontology are assigned a maximum confidence weight close to 1.0, while hypothetical paths from LLM are assigned a lower initial confidence weight representing uncertainty. Through this process, an initial causal knowledge graph containing both "strong beliefs" and "weak beliefs," which can be subsequently verified and evolved in a digital twin, is constructed and instantiated.

[0039] Optionally, the generation of the target risk transmission chain includes: Based on the dynamic evolutionary knowledge graph, symbolic reasoning and subgraph anomaly detection processes are executed in parallel to generate the first candidate risk path and the second high-risk node, respectively. Specifically, this step aims to comprehensively discover potential risk signals from complex knowledge graphs using two complementary reasoning methods. A risk reasoning engine executes two independent analysis processes in parallel. The first is a symbolic reasoning process, which, based on causal rules defined in the graph, traverses all possible transmission paths through logical reasoning algorithms to generate first candidate risk paths. The second is a subgraph anomaly detection process, which aims to identify nodes that exhibit anomalies in structure or content. The core of this process lies in calculating a joint anomaly score for each node, which integrates structural and semantic anomalies. An exemplary formula for calculating the joint anomaly score is as follows: , in, It is a node The final joint anomaly score; Based on nodes Local subgraph The calculated structural anomaly component, which can be obtained through methods such as the reconstruction error of a graph autoencoder, is used to quantify the rarity of the subgraph in the topological structure. This is the semantic anomaly component, which is used to quantize nodes. Its own feature vector The set of feature vectors of all its neighboring nodes The degree of inconsistency or "repulsion" between nodes can be obtained, for example, by calculating the distance between the node's features and the mean of its neighbors' features; It is a hyperparameter used to balance structural and semantic importance. Nodes with anomaly scores higher than a certain threshold are identified as the second highest-risk nodes.

[0040] Perform set operations on the first candidate risk path and the second high-risk node to identify and extract the set of conflict nodes; Specifically, this step aims to proactively identify "suspicious points" or "blind spots" in the model's cognition by comparing the results of two different inference methods. A result comparison module receives the two outputs generated in the previous step. It performs a symmetric difference operation on all nodes contained in the first candidate risk path and the set of second high-risk nodes. This operation identifies nodes that exist only in one set and not in the other. These nodes are then identified and extracted as conflict node sets. For example, a node is on a logically clear risk path, but its subgraph pattern is not identified as anomalous by the GNN model; therefore, it is a conflict node. Conversely, it is not. These conflict nodes represent "disagreements" between the two inference paradigms and are key objects requiring further in-depth validation.

[0041] Causal attribution verification is performed on the set of conflict nodes, and all verified paths are integrated to generate the target risk transmission chain.

[0042] Specifically, this step is the final decision-making and integration stage, aiming to resolve the "disagreements" identified in the previous step and generate a highly credible final result. A verification and integration module triggers a more refined causal attribution verification process for each conflict node in the conflict node set. This verification process uses multiple counterfactual scenario simulations in the digital twin to ultimately confirm whether the conflict node truly possesses significant risk transmission capabilities. Only those nodes that pass this rigorous verification, and those nodes and their paths that are undisputed in both inference methods, will be ultimately adopted. This module reorganizes and connects all these verified, highly credible nodes and paths, ultimately generating a single, internally consistent target risk transmission chain.

[0043] Optionally, the causal attribution verification for the set of conflict nodes includes: For the set of conflict nodes, a factual scenario is extrapolated based on the initial conditions of the food safety digital twin to generate a baseline risk evolution trajectory; Specifically, this step aims to establish a "baseline world" for subsequent comparisons. A simulation control module performs this operation for each conflict node to be verified in the conflict node set. The module first identifies the corresponding twin object of the conflict node in the food safety digital twin and records its current observational state, driven by real data. Then, within the digital twin environment, the module performs a forward-looking simulation of the node's inherent physical and logical laws without any intervention; this process is called factual scenario extrapolation. Through this extrapolation, the evolution curves of the risk status of other key nodes over time can be obtained, which serve as the baseline risk evolution trajectory for subsequent comparative analysis.

[0044] In the food safety digital twin, forced risk injection and forced security constraints are applied to the state of the conflict node set, and two independent counterfactual scenario simulations are performed to generate the maximum risk evolution trajectory and the minimum risk evolution trajectory, respectively. Specifically, this step aims to construct two extreme "parallel universes" for comparison, maximizing the revelation of the potential causal influence of the node to be verified. The simulation control module executes two independent, parallel counterfactual scenario simulations for the same conflict node. In the first simulation, the module performs a "forced risk injection" intervention on the node's critical risk state, for example, forcing its contamination parameters to a known, clearly defined maximum risk value. The result of this simulation is the maximum risk evolution trajectory. In the second simulation, the module performs a "forced safety constraint" intervention, for example, forcing its risk parameters to an absolutely safe baseline value. The result of this simulation is the minimum risk evolution trajectory. These two simulations together construct the causal influence boundaries of the node under the "worst-case" and "best-case" scenarios.

[0045] The causal attribution verification results are obtained by quantitatively comparing the differences between the baseline risk evolution trajectory, the maximum risk evolution trajectory, and the minimum risk evolution trajectory.

[0046] Specifically, this step is the final quantification and verification of causal effects. A results analysis module compares the three different evolutionary trajectories generated in the previous two steps for the same conflict node. This module calculates the magnitude of the difference between the maximum and minimum risk evolutionary trajectories at a key downstream assessment node. This difference directly quantifies the impact of the change from "absolutely safe" to "absolutely dangerous" state of the conflict node on the downstream, i.e., its true and maximum potential causal transmission capacity. Simultaneously, by analyzing the relative position of the baseline risk evolutionary trajectory between these two extreme trajectories, its actual risk contribution under the current observation state can be determined. If there is a significant difference between the maximum and minimum risk trajectories, the causal attribution verification result for the conflict node is "passed," proving that it is indeed a key risk transmission node.

[0047] Optionally, the quantitative output of key risk attribution factors includes: Based on the risk transmission chain, multi-dimensional static attributes are extracted from the dynamic evolutionary knowledge graph to obtain static risk weights; Specifically, this step aims to assign an initial base score representing the inherent risk level of each link in the risk transmission chain. An attribute extraction module traverses every entity node and relation edge in the target risk transmission chain. For each node or edge, the module extracts multi-dimensional static attributes reflecting its risk characteristics from a dynamically evolving knowledge graph. For example, for a supplier node, its attributes might include its historical quality inspection pass rate, credit rating, and the epidemic risk level of its region. For a "cold chain transportation" relation edge, its attributes might include transportation time, vehicle type, and whether there have been any interruption records. Subsequently, the module uses a predefined risk contribution rule base to map these qualitative or quantitative attributes to a single, standardized static risk weight.

[0048] The static risk weights are decayed and accumulated to quantify and determine the risk contribution of each upstream node to the downstream node, and the entity node and relation edge with the largest contribution are used as key risk attribution factors.

[0049] Specifically, this step aims to quantify the contribution of each link in the final risk event by employing a novel calculation method that simulates the natural decay of risk as it propagates along the chain. A contribution analysis module performs this calculation. This module traces upstream along the risk transmission chain, starting from the point of occurrence of the risk event. For any downstream node in the chain, its accumulated risk contribution is the sum of the risk contributions of all its direct upstream nodes, the static risk weight of the upstream node itself, and a factor that decays with distance. An exemplary decay accumulation algorithm can be expressed by the following formula: , in, It is a downstream node The cumulative risk contribution; It refers to all direct pointers The set of upstream nodes; It is an upstream node Its own static risk weight; It is a decay coefficient between 0 and 1, representing the loss rate of risk at each step of transmission; It is a node arrive The module calculates the quantified contribution of each node to the final risk by iterating along the chain. Ultimately, the entity node or relation edge with the highest contribution value is identified and output as the key risk attribution factor.

[0050] Optionally, the feedback mechanism between the dynamic evolutionary knowledge graph and the food safety digital twin includes: The treatment effect is used as physical world feedback data and compared with the forward-looking prediction results of the food safety digital twin to quantify and generate a cognitive-reality bias signal. Specifically, this step aims to precisely quantify the gap between the model's "cognition" and the "reality" of the physical world. This gap is the fundamental basis for all subsequent learning and optimization. A feedback analysis module is responsible for this task. After a risk control and disposal operation is executed, this module continuously collects real-world data on the effectiveness of the disposal from sensors or manual input channels. For example, the results of subsequent sampling inspections of an isolated batch of products, or changes in the pass rate of downstream products from a rectified production line. This real-world data is the physical world feedback data. Simultaneously, the module retrieves the forward-looking predictions from the food safety digital twin regarding "what will happen in the future if this disposal is carried out" before the disposal was implemented. The module aligns these two sets of data over time and performs a point-by-point quantitative comparison; the difference is then used to generate a cognition-reality bias signal containing rich learning information.

[0051] Causal attribution backtracking is performed based on the aforementioned cognitive-reality bias signal, and posterior probability updates are performed on the aforementioned dynamic evolutionary knowledge graph; and Specifically, this step aims to leverage "cognitive bias" to optimize the "logic layer" of the knowledge graph. A graph update module receives the cognitive-reality bias signal generated in the previous step. This module employs a causal attribution backpropagation algorithm, such as Shapley analysis or gradient backpropagation, to backpropagate the final bias signal upstream along the causal path in the knowledge graph. Through this process, it can identify which one or more causal relationship edges have their weights or confidence levels overestimated or underestimated, leading to the final prediction bias. For example, if it is found that the actual rate of pollution spread far exceeds the prediction, and backpropagation analysis indicates that this is mainly due to underestimating the transmission capacity of the "personnel crossover" causal path, then the module will perform a posterior probability update on the weights or confidence levels of this edge to better align with the newly observed reality.

[0052] The internal dynamics model parameters in the food safety digital twin are calibrated online using the cognitive-reality bias signal.

[0053] Specifically, this step aims to optimize the "physical layer" of the digital twin by leveraging "cognitive bias," which is parallel and complementary to the previous step of optimizing the "logical layer" of the graph. A twin calibration module also receives cognitive-reality bias signals. However, its focus is different; it aims to determine whether the bias is caused by inaccurate parameters in the physical or chemical process model within the twin object. For example, if it finds that the logical transmission path of the knowledge graph is correct, but the predicted growth rate of the total number of colonies in the contaminated product is too slow, the module will determine that the problem lies in the "kinetic model" used to simulate microbial reproduction within the twin object. At this point, the module will use an identification or parameter estimation algorithm, such as least squares or gradient descent, to fine-tune the relevant parameters in the kinetic model online, so that its simulation results better match the new physical world feedback data. Through this dual-track parallel feedback mechanism, both "logical knowledge" and "physical knowledge" are synchronously and purposefully corrected and evolved.

[0054] Optionally, the method further includes: The key risk attribution factors are summarized and patterns are mined to generate a structural risk pattern library. Specifically, this step aims to learn and extract recurring, universal "risk patterns" from independent risk attribution results, thereby transforming experiential knowledge into practical knowledge. A pattern mining module is responsible for this task. This module aggregates historical key risk attribution factors generated for multiple different risk events. It employs an association rule mining algorithm, such as Apriori or FP-Growth, to analyze these historical attribution factors to discover combinations of frequently co-occurring entity nodes or relationship edges. For example, the module might find that the three factors "specific supplier A," "humid summer climate," and "B-type packaging materials" have appeared simultaneously as key attribution factors in over 80% of historical mold risk events. These high-frequency co-occurrence combinations are stored as a structural risk pattern, and all these patterns together constitute a structural risk pattern library.

[0055] The structural risk pattern library is associated and matched with the conflict node set to identify and determine high-priority reconstruction targets; Specifically, this step is a core, non-obvious innovation of this invention, aiming to cross-validate "historical lessons" with "current confusion" to identify potential fundamental flaws in the model. A model diagnostic module performs this association matching. It receives the structural risk pattern library generated in the previous step, as well as the conflict node set generated for the current task. This module checks whether there are any entities or relationships in the current conflict node set that match a pattern in the structural risk pattern library. If a current conflict node happens to match a historically high-frequency risk pattern, then there is reason to highly suspect that the root of its "confusion" lies in the model's biased understanding of this type of risk. This matched node, which simultaneously possesses the dual characteristics of "current uncertainty" and "historically high risk," and its related subgraphs are identified and determined as high-priority reconstruction targets.

[0056] The high-priority reconstruction target is used as the core context and constraint to reconstruct and instantiate the initial causal knowledge graph.

[0057] Specifically, this step involves performing the highest level of "model evolution." It's no longer a simple adjustment of weights, but a surgical reshaping of the knowledge graph's "worldview." A model reconstruction module receives the high-priority reconstruction targets identified in the previous step. This module triggers a re-instantiation process of the initial causal knowledge graph P. However, unlike regular instantiation, this instantiation is "targeted" and "enhanced." When inputting ontology-enhanced twin entities as contextual cues into a large language model, this module inputs the identified high-priority reconstruction targets and their corresponding structural risk patterns as additional, high-weighted constraint information. For example, the cue might become: "Please generate deeper, more specific potential causal hypothesis specifically for the known high-risk pattern 'Supplier A uses Class B packaging materials in humid climates.'" In this way, the AI's "association" and "creation" can be guided to focus on known weaknesses, thereby performing a targeted and fundamental reconstruction of the initial causal knowledge graph.

[0058] Example 1: To verify the feasibility of this invention in practice, it was applied to the global supply chain risk management of a large multinational dairy group. This group's supply chain is complex, involving ranches, processing plants, cold chain logistics, and tens of thousands of retail outlets in multiple countries, facing potential food safety risks arising from minor changes in raw materials, environment, and processes.

[0059] To verify the effectiveness of this invention, the group's supply chain operation data for the entire year of 2024 was selected, and a retrospective analysis and comparative verification were conducted on a real minor product quality complaint incident caused by non-traditional factors. The experimental group used the method and system of this invention; the control group used the group's existing traditional linear traceability system based on ERP and blockchain, which was manually investigated and analyzed by a team of senior quality management experts.

[0060] In this embodiment, the experimental group continuously aggregated and processed multi-source heterogeneous data from the entire supply chain. For example, it not only collected structured data such as production batches and storage temperature and humidity at each stage, but also automatically extracted metadata about a food safety incident, such as "a packaging material supplier being investigated for environmental issues," from industry news using natural language processing technology. Furthermore, it performed time-series decomposition on truck GPS sensor data along a key transportation route, identifying the residual component of "abnormally high afternoon temperatures in summer." All of this was processed and generated into a high-dimensional real-time data stream.

[0061] Subsequently, these data streams were mapped in real time into a food safety digital twin, and an initial causal knowledge graph was instantiated. Notably, during instantiation, the large language model generated a hypothetical causal triple, "Specific packaging materials may release trace amounts of plasticizers at high temperatures," based on contextual cues such as "packaging materials" and "summer high temperatures." This relationship was not defined in the initial domain ontology.

[0062] In August 2024, multiple complaints were received from consumers in Region C regarding a slight "plastic taste" in a batch of premium yogurt, which became an initial risk event. The expert team in the control group traced the product using traditional systems and confirmed that key indicators such as temperature in all production and logistics stages of that batch were "under control," but they were unable to immediately pinpoint the source of the problem.

[0063] A dual-channel risk reasoning approach was employed on the dynamic evolutionary knowledge graph. While the symbolic reasoning process did not identify a clear contamination path, the subgraph anomaly detection process identified "Supplier G's packaging barrels" and "Transport segment L traversing tropical regions" as the second highest-risk nodes. These two nodes subsequently became the core of the "conflict node set." Multiple counterfactual scenario simulations were conducted in a digital twin targeting these two conflict nodes, verifying and confirming that under simulated high-temperature conditions, "Supplier G's packaging barrels" indeed have a significant negative causal effect on the flavor indicators of downstream products.

[0064] Finally, the validated target risk transmission chain was quantitatively attributed, and the key risk attribution factor was identified as a complex causal mechanism: "A specific batch of packaging materials provided by supplier G experienced high temperatures during long-distance summer transportation." Based on this, an intelligent disposal instruction was generated: "Preventively remove all products using this batch of packaging materials and transported through high-temperature areas and send them for testing." Subsequent physical testing confirmed that trace amounts of substances did indeed leach from this batch of packaging materials at high temperatures. This feedback on the disposal effect enhanced the weights of relevant causal paths in the knowledge graph and calibrated the parameters of the chemical substance leaching kinetic model in the digital twin, achieving learning and evolution.

[0065] After analyzing and verifying multiple similar complex risk events throughout the year, this invention has demonstrated significant technical advantages in terms of the depth of risk discovery, the accuracy of source tracing, and the efficiency of decision-making. For specific data, please refer to Tables 1, 2, and 3.

[0066] Table 1. Comparison of Efficiency and Accuracy in Tracing the Sources of Complex Risk Events Test group Average traceability time Root cause positioning accuracy Accuracy of recall of batches with associated risks experimental group 1.5 92.3% 98.1% control group 48.0 38.5% 65.4% Table 2 Comparison of early warning effects for similar potential risks before and after model evolution Model State Risk scenarios Early warning period Early warning confidence level Initial model Packaging materials + high-temperature transportation Unable to provide early warning - Post-evolutionary model Packaging materials + high-temperature transportation 15 85% Initial model Raw material origin + rainy season 3 40% Post-evolutionary model Raw material origin + rainy season 12 91% Table 3. Comparison of the effectiveness of different attribution methods in decision-making. Attribution methods Attribution results example Generated disposal suggestions Recurrence rate of similar risks in the future This invention "Summer High Temperature Transportation Mechanism" Change packaging material suppliers; optimize transportation routes 2.1% conventional methods "SN20240815 batch" Remove and destroy this batch of products. 35.7% Tables 1-3 above record the actual application data of the present invention in handling complex, non-traditional food safety risk events, and demonstrate in detail the system’s superior performance in terms of traceability depth, early warning capability, and decision quality.

[0067] As can be clearly seen in Table 1, for complex risks that go beyond conventional database associations, the method of this invention can reduce the average tracing time from 2 days to 1.5 hours, and the accuracy of root cause location is as high as 92.3%, thanks to its causal inference and digital twin verification capabilities.

[0068] The data in Table 2 verifies the core "learning and evolution" capability of this invention. For initially unknown risks such as "packaging materials + high temperature," after experiencing feedback learning from the first event, the model's warning capability developed from zero to high confidence, reaching as high as 85%. This proves that the closed-loop feedback mechanism of this invention is effective and powerful.

[0069] Table 3 illustrates the significant advantage of this invention in terms of attribution depth. Conventional methods can only attribute the cause to a specific "entity," and their recommendations are merely reactive measures. This invention, however, can attribute the cause to a "mechanism," thereby proposing fundamental, preventative solutions such as "changing suppliers" and "optimizing routes," reducing the recurrence rate of similar risks by more than an order of magnitude.

[0070] Based on the same inventive concept, this invention also provides a knowledge graph-driven food safety risk traceability system, such as... Figure 5 As shown, the system includes: The end-to-end data aggregation and streaming module is used to acquire and integrate multi-source heterogeneous data from the raw material end, production end, distribution end and consumption end, and perform unified and standardized processing to form a real-time data stream; The twin-graph collaborative modeling module is used to construct a food safety digital twin based on the real-time data stream, and to perform semantic mapping and causal mining using the food safety digital twin to output a dynamic evolutionary knowledge graph. The multi-channel risk reasoning and verification module is used to identify risks and reason causally about the dynamic evolution knowledge graph, obtain a set of potential risk nodes and causal propagation paths, and perform cross-validation to generate a target risk transmission chain. The counterfactual attribution quantification module is used to perform time-series consistency verification on the target risk transmission chain, perform multi-scenario counterfactual simulation on the risk transmission chain that passes the verification, and quantify and output key risk attribution factors. The decision control and closed-loop evolution module is used to implement control and monitor the handling effect based on the key risk attribution factors, and to provide feedback on the dynamic evolution knowledge graph and the food safety digital twin.

[0071] It should be noted that the functional division and information interaction between the various modules described above are logical, but in terms of physical implementation, they can be integrated on the same software platform or deployed in a distributed manner. The connections between them represent data flow and control flow, aiming to collaboratively achieve the building energy consumption dynamic optimization goal of this invention. The above descriptions are merely exemplary embodiments of this invention and should not be construed as limiting the scope of protection of this invention.

Claims

1. A knowledge graph-driven method for tracing food safety risks, characterized in that, The method includes: Acquire and integrate heterogeneous data from multiple sources, including raw materials, production, distribution, and consumption, and process them in a unified and standardized manner to form a real-time data stream; A food safety digital twin is constructed based on the real-time data stream, and semantic mapping and causal mining are performed using the food safety digital twin to output a dynamic evolutionary knowledge graph. Risk identification and causal reasoning are performed on the dynamic evolutionary knowledge graph to obtain a set of potential risk nodes and causal propagation paths, and cross-validation is performed to generate a target risk transmission chain; Perform time-series consistency verification on the target risk transmission chain, perform multi-scenario counterfactual simulation on the risk transmission chain that passes the verification, and quantify and output key risk attribution factors; Based on the key risk attribution factors, control measures are implemented and the effectiveness of the measures is monitored, and feedback is provided to the dynamic evolutionary knowledge graph and the food safety digital twin.

2. The knowledge graph-driven food safety risk tracing method according to claim 1, characterized in that, The formation of the real-time data stream includes: Named entity recognition and event extraction are performed on the unstructured text data in the multi-source heterogeneous data to extract and structure food safety event metadata; Multi-scale time-series decomposition is performed on the sensor time-series data in the multi-source heterogeneous data to separate and quantify the trend-seasonal component and the residual component. The food safety event metadata, the trend-seasonal component, the residual component, and the structured data from the multi-source heterogeneous data are context-associated and feature-embedded to generate a real-time data stream.

3. The knowledge graph-driven food safety risk tracing method according to claim 1, characterized in that, The output dynamic evolution knowledge graph includes: The real-time data stream is mapped to a computable twin object of the food safety digital twin, and an initial causal knowledge graph is instantiated based on the initial association between the food safety digital twins. Active perturbation and prospective simulation are performed on the computable twin object to generate virtual process data, and causal effect analysis is performed on the virtual process data to generate and verify implicit causal relationships; The implicit causal relationships are treated as new knowledge and incrementally updated into the topology and relation weights of the initial causal knowledge graph. The graph is then iteratively evolved and output as a dynamically evolving knowledge graph.

4. The knowledge graph-driven food safety risk tracing method according to claim 3, characterized in that, The instantiation of the initial causal knowledge graph includes: The computable twin object is semantically aligned with a preset food safety domain ontology to generate an ontology-enhanced twin entity; The ontology-enhanced twin entity is used as a contextual cue and input into a large language model to generate hypothetical causal triples; By integrating the deterministic causal paths of the food safety ontology with the hypothetical causal triples, an initial causal knowledge graph is constructed and instantiated.

5. The knowledge graph-driven food safety risk tracing method according to claim 4, characterized in that, The generated target risk transmission chain includes: Based on the dynamic evolutionary knowledge graph, symbolic reasoning and subgraph anomaly detection processes are executed in parallel to generate the first candidate risk path and the second high-risk node, respectively. Perform set operations on the first candidate risk path and the second high-risk node to identify and extract the set of conflict nodes; Causal attribution verification is performed on the set of conflict nodes, and all verified paths are integrated to generate the target risk transmission chain.

6. The knowledge graph-driven food safety risk tracing method according to claim 5, characterized in that, The causal attribution verification for the set of conflict nodes includes: For the set of conflict nodes, a factual scenario is extrapolated based on the initial conditions of the food safety digital twin to generate a baseline risk evolution trajectory; In the food safety digital twin, forced risk injection and forced security constraints are applied to the state of the conflict node set, and two independent counterfactual scenario simulations are performed to generate the maximum risk evolution trajectory and the minimum risk evolution trajectory, respectively. The causal attribution verification results are obtained by quantitatively comparing the differences between the baseline risk evolution trajectory, the maximum risk evolution trajectory, and the minimum risk evolution trajectory.

7. The knowledge graph-driven food safety risk tracing method according to claim 6, characterized in that, The key risk attribution factors for the quantitative output include: Based on the risk transmission chain, multi-dimensional static attributes are extracted from the dynamic evolutionary knowledge graph to obtain static risk weights; The static risk weights are decayed and accumulated to quantify and determine the risk contribution of each upstream node to the downstream node, and the entity node and relation edge with the largest contribution are used as key risk attribution factors.

8. The knowledge graph-driven food safety risk tracing method according to claim 1, characterized in that, The dynamic evolutionary knowledge graph and the food safety digital twin mentioned above include: The treatment effect is used as physical world feedback data and compared with the forward-looking prediction results of the food safety digital twin to quantify and generate a cognitive-reality bias signal. Based on the cognitive-reality bias signal, causal attribution backtracking is performed, and the posterior probability of the dynamic evolutionary knowledge graph is updated; and The internal dynamics model parameters in the food safety digital twin are calibrated online using the cognitive-reality bias signal.

9. The knowledge graph-driven food safety risk tracing method according to claim 7, characterized in that, The method further includes: The key risk attribution factors are summarized and patterns are mined to generate a structural risk pattern library. The structural risk pattern library is associated and matched with the conflict node set to identify and determine high-priority reconstruction targets; The high-priority reconstruction target is used as the core context and constraint to reconstruct and instantiate the initial causal knowledge graph.

10. A knowledge graph-driven food safety risk traceability system, applied to a knowledge graph-driven food safety risk traceability method as described in any one of claims 1-9, characterized in that, The system includes: The end-to-end data aggregation and streaming module is used to acquire and integrate multi-source heterogeneous data from the raw material end, production end, distribution end and consumption end, and perform unified and standardized processing to form a real-time data stream; The twin-graph collaborative modeling module is used to construct a food safety digital twin based on the real-time data stream, and to perform semantic mapping and causal mining using the food safety digital twin to output a dynamic evolutionary knowledge graph. The multi-channel risk reasoning and verification module is used to identify risks and reason causally about the dynamic evolution knowledge graph, obtain a set of potential risk nodes and causal propagation paths, and perform cross-validation to generate a target risk transmission chain. The counterfactual attribution quantification module is used to perform time-series consistency verification on the target risk transmission chain, perform multi-scenario counterfactual simulation on the risk transmission chain that passes the verification, and quantify and output key risk attribution factors. The decision control and closed-loop evolution module is used to implement control and monitor the handling effect based on the key risk attribution factors, and to provide feedback on the dynamic evolution knowledge graph and the food safety digital twin.

Citation Information

Cited By

  • Industrial data management method based on big data

    CN122022777A

  • A multi-agent collaborative decision-making system based on digital twinning and knowledge graph

    CN122331610A