A food supply chain sampling method based on reinforcement learning and markov dynamics
Patent Information
- Application Number
- CN202610923501.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-18
AI Technical Summary
[0008]本发明的策略优化方法,用于解决目前食品安全监管中存在的感知维度单一、成本核算粗放、资源错配以及缺乏风险时空动态追踪机制等技术问题;本发明技术方案为:融合大模型先验知识与真实世界物理法规约束,建立风险-规模二维靶向定价模型,利用强化学习引导的多目标优化算法在预算约束内生成最优抽检名单,并通过马尔可夫状态演化推演风险变化,形成策略决策闭环;主要包含以下核心步骤:(1)获取食品供应链多维风险特征矩阵并构建带地理阻断惩罚的时空异构图模型;(2)建立风险-规模二维靶向定价模型,动态核算包含专项试剂费的个体抽检成本;(3)构建强化学习引导的多目标抽检寻优模型,结合高维约束惩罚与贪心修复算子生成帕累托最优策略;(4)建立结合专家知识前置熔断干预的属地化调度机制,输出微观执行工单;(5)引入马尔可夫状态演化动力学对节点特征进行动态更新,并推流至全栈可视化交互大屏
本发明首次将食品生化机理与运筹策略优化进行深度融合,通过构建“风险-规模二维靶向定价模型”,解决了传统“盲抽全项检测”带来的资金浪费,使得系统能够智能决定“对哪些节点进行何种特定项目的抽检”,极大提升了资金利用效能;本发明创新性地引入了强化学习多臂老虎机(RL-MAB)与空间发散代理模型,有效破解了海量节点抽检组合优化过程中的维数灾难,使得多目标策略寻优算法具备了毫秒级的计算响应能力;本发明通过建立包含高维惩罚空间与专家强制前置干预的属地化策略派发机制,使得纯数学输出的抽检名单严格遵守了国家关于信用分级分类监管及“双随机、一公开”底线威慑的要求,输出的微观工单可直接应用于一线监管决策;本发明引入马尔可夫状态转移方程与全栈前端可视化大屏,赋予了系统在时间维度上规划“抽检频次”的能力,直观呈现了高危节点被反复打压与低危节点盲区发酵的全过程,为宏观食品安全风险的治理提供了“数据与策略双轮驱动”的科学决策支持工具。
Smart Images

Figure CN122779477A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to public health governance, large-scale model applications, and operations research optimization methods. Specifically, it relates to a food supply chain sampling method based on reinforcement learning and Markov dynamics, which is particularly suitable for intelligent decision-making scenarios that utilize the prior knowledge of large-scale models to formulate targeted sampling lists, sampling frequencies, and testing items for food safety under complex spatiotemporal environments and multidimensional resource constraints. Background Technology
[0002] Food safety is a major and fundamental issue concerning the national economy and people's livelihood. With the increasing complexity of the modern food supply chain and the frequent cross-regional circulation, food safety risks are showing new characteristics such as high concealment and multiple transmission routes.
[0003] Currently, the formulation of sampling strategies for food safety supervision largely relies on a static blind sampling model based on manual experience, characterized by "two randoms, one public" sampling. This traditional model has the following significant technical drawbacks: (1) Single perception dimension and lack of spatial spillover assessment: Traditional methods often regard each supply chain node as an independent individual, ignoring the cross-contamination and chain reaction of risks and hidden dangers in the logistics topology network and physical space environment.
[0004] (2) The cost accounting is rough, resulting in the misallocation of sampling resources: Traditional sampling plans usually adopt a "one-size-fits-all" cost accounting method, which fails to combine the enterprise size and suspected pollution type (such as the cost difference between routine microbial testing and expensive heavy metal testing) to target resource allocation, resulting in ineffective waste of funds due to blindly implementing "full-item testing".
[0005] (3) The strategy formulation is out of touch with reality and lacks a dynamic tracking mechanism: Existing optimization models often ignore the institutional constraints of the government's "local grid management" and the hard barriers of funding / human resources, resulting in strategies that lack feasibility for implementation. At the same time, traditional static sampling strategies cannot depict the dynamic game law of the natural fermentation and intervention decay of risks and hidden dangers over time, and cannot scientifically guide the allocation of sampling frequency over multiple consecutive days.
[0006] Therefore, there is an urgent need for a new sampling inspection method that can break down the aforementioned physical and legal barriers and integrate the prior knowledge of artificial intelligence risk prediction with the multidimensional constraints of the real world. Summary of the Invention
[0007] Purpose of the invention: To overcome the problems of perception lag, crude cost accounting, lack of targeted assignment and lack of dynamic frequency tracking in the previous food safety sampling inspection, this invention provides a food supply chain sampling inspection method based on reinforcement learning and Markov dynamics.
[0008] The strategy optimization method of the present invention is used to solve the technical problems existing in the current food safety supervision, such as single perception dimension, rough cost accounting, resource misallocation and lack of risk spatiotemporal dynamic tracking mechanism. The technical solution of the present invention is: to integrate the prior knowledge of the large model with the physical regulations of the real world, to establish a risk-scale two-dimensional targeted pricing model, to generate the optimal sampling list within the budget constraint using a multi-objective optimization algorithm guided by reinforcement learning, and to form a closed loop of strategy decision by inferring risk changes through Markov state evolution. The main core steps are as follows: (1) to obtain the multi-dimensional risk feature matrix of the food supply chain and to construct a spatiotemporal heterogeneous graph model with geographical blocking penalty; (2) to establish a risk-scale two-dimensional targeted pricing model and dynamically calculate the individual sampling cost including special reagent fees; (3) to construct a multi-objective sampling optimization model guided by reinforcement learning, and to generate Pareto optimal strategy by combining high-dimensional constraint penalty and greedy repair operator; (4) to establish a localized scheduling mechanism combined with expert knowledge pre-circuit breaker intervention and output micro-execution work orders; (5) to introduce Markov state evolution dynamics to dynamically update node features and push them to the full-stack visualization interactive screen. This invention is applied to the interdisciplinary field of public health governance and operations optimization, which can significantly reduce the cost of ineffective supervision and sampling inspections and the amount of manual work, and realize intelligent prevention and control of food safety risks in complex spatiotemporal environments.
[0009] Technical Solution: This invention provides a food supply chain sampling inspection method based on reinforcement learning and Markov dynamics. The specific method consists of the following steps: Constructing a spatiotemporal heterogeneous graph and sensing the situation: Acquiring attribute data and multi-dimensional risk prediction probabilities of food supply chain nodes, and integrating real logistics allocation relationships with a spatial attenuation network based on geographical Euclidean distance. In distance calculation, a hyperplane equation based on analytical geometry is introduced to define the boundaries between continents and islands. Big-M geographical blocking penalties are applied to cross-boundary nodes to cut off unreasonable cross-regional pollution predictions and subsequent strategy assignments, ultimately aggregating into a network-wide scalar comprehensive risk index. This index serves as the basic input for subsequent decision-making, characterizing the dual propagation characteristics of risk in physical space and the logistics network.
[0010] Risk-scale two-dimensional targeted pricing: Abandoning fixed sampling costs, setting judgment thresholds, and only triggering the high reagent cost of the special test when the risk prediction probability of a node in a certain dimension (such as pesticide and veterinary drug residues) exceeds the limit; at the same time, configuring workload multipliers based on enterprise scale (such as large dairy processing plants with many sampling points) to achieve dynamic and accurate accounting of individual sampling costs, thereby eliminating low-risk redundant testing items within a limited budget.
[0011] The reinforcement learning-guided multi-objective sampling optimization approach aims to minimize the total sampling cost and maximize the interception risk, employing a non-dominated sorting genetic algorithm (NSGA-II). A reinforcement learning multi-armed slot machine (RL-MAB) mechanism is introduced, utilizing Q-Learning temporal difference to update the confidence upper limit and adaptively adjust the mutation exploration rate. In the evaluation phase, spatial dispersion is introduced as a surrogate model to estimate execution overhead. A high-dimensional penalty space is set for global budget, manpower attendance, and coverage of legally mandated core business formats, and a greedy repair operator is used to forcibly remove nodes that exceed the budget and have low cost-effectiveness, outputting a Pareto optimal strategy. This optimization process aims to balance supervisory effectiveness and economic cost, solving the NP-hard problem under massive node combinations.
[0012] Expert pre-intervention and localized strategy generation: Before algorithm optimization, large core enterprises that exceed the limits of critical dimensions are subject to mandatory circuit breaker and pre-deduction of budgets according to national standards, and random blind sampling is conducted at low-risk terminals to maintain long-tail bottom-line deterrence. After merging the AI optimization strategy with the expert list, the sampling targets are allocated to the nearest district-level regulatory agencies according to the geographical matrix with penalties, generating localized micro-execution work orders that include "who to sample, what special items to sample, and how much to spend".
[0013] Markov evolutionary dynamics and front-end visualization: After generating a single strategy, the simulation proceeds over time. For nodes subject to inspection, risk decay calculations are performed according to the corresponding dimensions. For nodes not inspected and in regulatory blind spots, the pollution mechanism is differentiated, with exponential fermentation evolution implemented for microbial contamination and constant-level accumulation for chemical contamination. Subsequently, the updated feature tensors, Pareto front, kernel density contrastive distribution (KDE), and intelligent dispatch tables categorized by region are pushed to a B / S architecture web front-end system, enabling a dynamic large-screen visualization of the inspection strategy and multi-day evolution frequency. Through this step, the system can intuitively demonstrate the risk mitigation effect of regulatory intervention and provide real-time execution feedback.
[0014] Compared with the prior art, the present invention has the following beneficial effects: This invention is the first to deeply integrate food biochemical mechanisms with operations research optimization. By constructing a "risk-scale two-dimensional targeted pricing model," it solves the problem of cost waste caused by traditional "blind sampling of all items" testing, enabling the system to intelligently decide "which nodes to sample and what specific items to test," greatly improving the efficiency of capital utilization. This invention innovatively introduces reinforcement learning multi-armed machine (RL-MAB) and spatial divergent proxy models, effectively overcoming the curse of dimensionality in the optimization process of massive node sampling combinations, giving the multi-objective strategy optimization algorithm millisecond-level computational response capability. This invention establishes a high-level... The localized strategy distribution mechanism, which combines dimensional penalty space and expert-mandated pre-intervention, ensures that the sampling lists generated by pure mathematics strictly adhere to national requirements for credit-based classification and supervision, as well as the "two randoms, one public" bottom-line deterrent. The output micro-level work orders can be directly applied to front-line regulatory decisions. This invention introduces Markov state transition equations and a full-stack front-end visualization dashboard, giving the system the ability to plan "sampling frequency" in the time dimension. It intuitively presents the entire process of high-risk nodes being repeatedly suppressed and low-risk nodes fermenting in blind spots, providing a scientific decision support tool driven by both data and strategy for the governance of macro-level food safety risks. Attached Figure Description
[0015] Figure 1 The diagram below shows the overall control flow and working architecture of the intelligent sampling strategy optimization method of this invention. Figure 2 This is a flowchart illustrating the risk-scale two-dimensional targeted pricing model and dynamic sampling project recommendation of the present invention; Figure 3 This is a schematic diagram of the multi-objective optimization and high-dimensional constraint penalty of the reinforcement learning multi-armed slot machine (RL-MAB) of the present invention; Figure 4 This is a schematic diagram of the expert knowledge intervention and target assignment module and the state evolution dynamics and risk update calculation module of the present invention; Figure 5 The strategy evaluation dashboard output by this invention includes road network scheduling trajectory, Pareto front curve, and reinforcement learning operator evolution trend; Figure 6 This is the interface of the digital intelligent order dispatch system generated by the method of the present invention; Figure 7 This refers to the spatial distribution and heat map of the sampling situation generated by the method of this invention; Figure 8 This is a multi-day spatiotemporal risk variation map generated by the method of this invention. Detailed Implementation
[0016] To enable those skilled in the art to clearly understand the implementation process of the present invention, the accompanying drawings are provided. Figure 1-8The technical solution of the present invention will be further described below with reference to the embodiments. This embodiment uses the complex dairy product supply chain network in Shanghai as the experimental object. By receiving the predicted risk probability output from the upstream multimodal large model, a closed-loop system for optimizing the sampling strategy, consisting of "perception-decision-dispatch-evolution-visualization," is constructed. The overall process is as follows: Figure 1 As shown, the specific implementation steps are as follows: In step S1, during the multidimensional risk perception and geographical blocking stage, the system first acquires the attributes of all network nodes and the upstream predicted 7-dimensional risk tensor. To calculate the overall situation, the system not only constructs a topology propagation matrix based on batch logistics but also introduces a non-Euclidean spatial geographical blocking mechanism. In specific implementation, the system establishes an analytical geometric hyperplane to divide the mainland and islands and uses the Big-M method to reconstruct the actual scheduling effective distance matrix between two points. The engineering calculation formula is as follows: ; In the formula, For nodes With nodes The Euclidean linear distance between them. The penalty constant is extremely large (9999.0 in the example), indicating the function. Internally, it determines whether the coordinates of two nodes lie within the hyperplane equation. The opposite side (the product is less than 0, which means crossing the river) is used to instantly cut off non-compliant cross-regional order dispatch paths at the matrix level.
[0017] Steps S2 and S3, in the targeted pricing calculation and multi-objective strategy optimization stage, such as Figure 2 and Figure 3 As shown. To address the curse of dimensionality caused by optimizing massive node combinations, this embodiment abandons the time-consuming real-physics path planning in the fitness evaluation of the NSGA-II algorithm. Instead, it introduces a surrogate evaluation model based on "spatial dispersion," integrating dynamic targeting cost and multi-dimensional penalty terms. Individual strategy (chromosome) The comprehensive cost objective function The specific implementation formula is as follows: ; In the formula, the first term is based on the 7-dimensional risk characteristics of the nodes. Total cost of dynamically triggered targeted detection reagents; The target detection cost C of a single node i i The specific accounting method is as follows: ; in, Let i be the actual dynamic sampling cost of node i in the current period. The basic attendance sampling fee is K, where K is the total number of risk feature dimensions. For the risk prediction probability of node i in the k-th dimension, The threshold for triggering specific detection, This is an indicator function that takes the value 1 if the condition is true and 0 otherwise. For the specialized testing reagent cost in the k-th dimension, Let i be the enterprise size workload multiplier corresponding to node i.
[0018] The second item is the agent's transportation cost, calculated by the geographic center of the selected node set. Calculate the mean of node divergence and multiply it by the estimated logistics cost rate. ; Third item This is a high-dimensional penalty function used to impose a massive, quadratic penalty on strategies that exceed the total budget limit, manpower limit, or fail to meet the coverage target for core large companies, transforming soft constraints into hard constraints. Meanwhile, in reinforcement learning-guided adaptive search, the multi-armed slot machine (RL-MAB) does not directly use the Q-value to determine the mutation rate, but instead uses the Softmax function to transform the Q-value into a specific action selection probability distribution. The underlying mapping formula of its code is: ; By introducing Numerical stabilization is performed to ensure that the probabilities of each mutation operator (fine-tuning, moderate mutation, global exploration) are equal in the early stage of optimization, and then automatically converge to the high-yield local fine-tuning action in the later stage.
[0019] Building upon this foundation, this invention further introduces a reward feedback mechanism from multi-armed slot machines (RL-MAB) to update the Q-value of each mutation action online. Specifically, a reward function based on Pareto dominance is constructed: if the policy of the mutated offspring policy Pareto dominates its parent policy in the target space, a positive reward value is assigned. =1.0, otherwise assign a negative penalty value. =-0.1. Then, the corresponding action is updated using the temporal difference method. a The Q value is updated using the following formula: ; in, and Let be the expected value of action 'a' selected in the t-th and t+1-th iterations, respectively, and α be the learning rate of reinforcement learning. Let be the reward value obtained in the t-th iteration, so that the optimization process can smoothly evolve from global large-scale exploration to local refined development.
[0020] Step S4, in the stage of expert pre-intervention and localization strategy allocation, such as Figure 4 As shown in the left figure, to bridge the gap between academic models and government logic, an expert intervention database is established at the beginning of each daily simulation. For large, extremely high-risk core enterprises that trigger mandatory national standards, a pre-emptive circuit breaker with a veto is implemented to lock in and pre-deduct emergency response budgets; simultaneously, random blind sampling is conducted at extremely low-risk terminals to maintain a long-tail deterrent. The remaining budget is then optimized using the aforementioned AI algorithm. The final inspection targets are then assigned to the nearest of Shanghai's 16 district-level market supervision and management agencies. Each branch uses a 2-Opt local search algorithm to eliminate route intersections and outputs independent physical navigation work orders.
[0021] Step S5, in the Markov evolution and large-screen visualization stage, such as Figure 4 As shown in the right figure, after the daily strategy is generated, the system performs a heterogeneous matrix update on the risk tensor of the entire network.
[0022] For the nodes that have already been sampled ( The risk characteristics of each dimension are attenuated after the implementation of random sampling intervention, and are uniformly multiplied by the attenuation coefficient β (0<β<1). The calculation formula is as follows: ; For regulatory blind spots that were not sampled ( If so, it is necessary to strictly distinguish the physicochemical properties of pollutants for heterogeneous updating, as follows: In the first Updating risk characteristics across multiple dimensions requires a strict distinction between the physicochemical properties of pollutants: ; In the formula, the dimensions of microorganisms and other organisms possessing biological reproductive characteristics are influenced by exponential growth factors. Driven; while inert dimensions such as heavy metals are only subject to constant-order parameters. The driving force exhibits a small linear accumulation. After completing the above refined tensor update, the new 7-dimensional features are locally stored on disk and fed back to the upstream heterogeneous graph network. Subsequently, the system pushes all decision indicators and evolutionary status to a large visualization screen: as attached. Figure 5 As shown, the Pareto front curve of sampling cost versus interception risk, the road network scheduling trajectory, and the convergence trend of the reinforcement learning operator are illustrated; Appendix Figure 6 A digital intelligent dispatch table is presented, intuitively displaying the total number of inspections, total expenses, and 7-dimensional risk tensor and suggested testing items assigned to each law enforcement agency for the day; (Attached) Figure 7 The heat map visually reflects the spatial risk density distribution and sampling coverage; finally, as attached... Figure 8 As shown, through a multi-day spatiotemporal simulation and comparison from Day 1 to Day 6, the entire process of the game-like evolution of risk reduction under intervention and the fermentation of hidden dangers in regulatory blind spots is presented intuitively.
[0023] The above are merely preferred embodiments of the present invention and do not constitute any limitation on the present invention. Any equivalent substitutions or modifications made by those skilled in the art to the technical solutions and content disclosed in the present invention without departing from the scope of the present invention shall be deemed to have remained within the protection scope of the present invention.
Claims
1. A food supply chain sampling inspection method based on reinforcement learning and Markov dynamics, characterized in that, Includes the following steps: Risk perception steps: Obtain attribute data and multi-dimensional risk feature matrix of food supply chain nodes, construct a spatiotemporal heterogeneous graph model by combining node geographic coordinates and logistics topology, and calculate the comprehensive risk status of all network nodes based on physical geographic blocking mechanisms. Cost accounting steps: Establish a risk-scale two-dimensional targeted pricing model, and dynamically calculate the individual sampling cost of each node based on the risk prediction probability of each dimension in the multi-dimensional risk feature matrix and the enterprise scale parameter; Strategy optimization steps: Construct a reinforcement learning-guided multi-objective sampling optimization model with the optimization objectives of minimizing the total sampling cost and maximizing the interception risk. Combine the preset high-dimensional constraint penalty mechanism and greedy repair operator to generate a Pareto optimal set of sampling strategies. Execution and evolution steps: Establish a localized scheduling mechanism that combines expert knowledge with pre-intervention, extract equilibrium solutions from the Pareto optimal strategy set and generate micro-execution work orders; at the same time, introduce Markov state evolution dynamics, and dynamically update the multidimensional risk characteristics of each node in the spatiotemporal heterogeneous graph in discrete time steps according to the sampling action of the current cycle, and enter the deduction and iteration of the next cycle.
2. The method according to claim 1, characterized in that, The specific implementation method of the physical geographic blocking mechanism is as follows: A strong correlation propagation matrix is constructed by utilizing the real logistics allocation relationships between nodes, and a spatial propagation matrix based on exponential decay is constructed by utilizing the geographical Euclidean distance between nodes. A hyperplane equation based on analytic geometry is introduced to define the boundary between the continent and the island. For node pairs that cross the boundary, a large number penalty constant is applied in the distance calculation to reconstruct the effective distance matrix. The calculation formula is as follows: ; In the formula, For nodes With nodes The Euclidean linear distance between them. A very large penalty constant, As an indicator function, a cross-region blocking penalty is applied by determining whether the coordinates of two nodes are on opposite sides of the hyperplane equation; Finally, the scalar comprehensive risk index after dimensionality reduction is obtained by weighting the maximum value of the node's intrinsic risk and the spillover risk of neighboring nodes.
3. The method according to claim 1, characterized in that, The formula for calculating the sampling cost using the risk-scale two-dimensional targeted pricing model is as follows: ; in, Let i be the actual dynamic sampling cost of node i in the current period. The basic attendance sampling fee is K, where K is the total number of risk feature dimensions. For the risk prediction probability of node i in the k-th dimension, The threshold for triggering specific detection, This is an indicator function that takes the value 1 if the condition is true and 0 otherwise. For the specialized testing reagent cost in the k-th dimension, Let i be the enterprise size workload multiplier corresponding to node i.
4. The method according to claim 1, characterized in that, In the strategy optimization step, the multi-objective sampling optimization model adopts a non-dominated sorting genetic algorithm framework. In the fitness evaluation stage, the spatial dispersion from the selected node to the nearest regulatory agency is used as a proxy model to estimate traffic costs, replacing the full physical path calculation; the spatial dispersion is defined as the average Euclidean distance from the geographic center of the selected node set to each selected node; the high-dimensional constraint penalty mechanism presets penalty functions for the global budget limit, daily manpower attendance limit, single-point overload of regulatory agencies, coverage rate of legal core business formats, and systemic bottom-line coverage rate of large enterprises, and applies high-dimensional penalty values to the over-limit strategy; The specific operation of the greedy repair operator is as follows: after each generation of offspring strategy, the cost-effectiveness index of the selected node is calculated, and the cost-effectiveness index is the ratio of the node's comprehensive risk index to the dynamic sampling cost. If the total cost of the current strategy exceeds the global budget limit, nodes will be forcibly removed in ascending order of the cost-effectiveness index until the total cost meets the budget hard constraint.
5. The method according to claim 1, characterized in that, The reinforcement learning guidance employs a multi-armed slot machine adaptive operator selection strategy, specifically including: Initialize the action space, which contains multiple search strategies with different mutation rates, and initialize the Q-value and selection probability of the corresponding actions; In each step of population iteration, the Q-value is transformed into a Softmax probability distribution based on the confidence upper bound strategy, and the current action is dynamically selected to perform gene mutation; the Softmax probability distribution is numerically stabilized, and its mapping relationship is as follows: ; in, For the first t Next iteration selects action a The probability, For the first t Next iteration action a The expected value; constructing a reward function based on Pareto dominance: if the policy of the mutated offspring policy Pareto dominates its parent policy in the target space, then a positive reward value is assigned. =1.0, otherwise assign a negative penalty value. =-0.1, update the Q value of the corresponding action using the temporal difference method, and the update calculation formula is: ; in, and Let be the expected value of action 'a' selected in the t-th and t+1-th iterations, respectively, and α be the learning rate of reinforcement learning. Let be the reward value obtained in the t-th iteration.
6. The method according to claim 1, characterized in that, The localized scheduling mechanism, which incorporates expert knowledge for pre-intervention, specifically includes: before algorithm optimization, extracting large core source enterprises with critical dimension probability exceeding limits from multi-dimensional risk features, executing mandatory circuit breaker operations, and deducting the corresponding detection costs from the total budget in advance; simultaneously, randomly selecting a fixed number of nodes from extremely low-risk terminal nodes to maintain long-tail bottom-line deterrence; taking the union of the Pareto equilibrium strategy output by AI optimization with the pre-intervention list to generate the final sampling list; assigning nodes to the nearest regulatory agency by calculating the effective distance with geographical blocking penalties from each sampling node to each regional regulatory agency; and, under the constraint of the node's statutory business hours window, using the 2-Opt heuristic algorithm to eliminate cross paths and output the physical execution sequence.
7. The method according to claim 1, characterized in that, The update steps of Markov state evolution dynamics are: dividing the global nodes into a sampled set. Set of blind spots not sampled For pollution dimensions exhibiting biological reproduction characteristics, the formula for calculating the node state evolution is as follows: ; in, and Let be the risk probability of node i in the k-th dimension of the t-th and t+1-th periods, respectively; β be the attenuation coefficient brought about by the implementation of sampling intervention; λ be the basic environmental deterioration rate; γ be the hazard index fermentation rate; and Δt be the time step. For the inert physicochemical pollution dimension, a constant-level linear micro-amplitude accumulation function is applied to the nodes in the unsampled set. After the feature update is completed, it is used as the initial multidimensional risk feature matrix for the next period.
8. A food supply chain sampling strategy optimization system based on reinforcement learning and Markov dynamics, characterized in that, include: The risk perception module is used to acquire attribute data and multi-dimensional risk feature matrices of food supply chain nodes, construct a spatiotemporal heterogeneous graph model by combining node geographic coordinates and logistics topology, and calculate the comprehensive risk status of all network nodes based on physical geographic blocking mechanisms. The cost accounting module is used to establish a risk-scale two-dimensional targeted pricing model, and dynamically calculate the individual sampling cost of each node based on the risk prediction probability of each dimension in the multi-dimensional risk feature matrix and the enterprise scale parameter. The strategy optimization module is used to build a reinforcement learning-guided multi-objective sampling optimization model. With the optimization objectives of minimizing the total sampling cost and maximizing the interception risk, it combines a preset high-dimensional constraint penalty mechanism and a greedy repair operator to generate a Pareto optimal set of sampling strategies. The execution scheduling and evolution module is used to establish a localized scheduling mechanism that combines expert knowledge with pre-intervention, extracting equilibrium solutions from the Pareto optimal strategy set and generating micro-execution work orders. And it is used to introduce Markov state evolution dynamics, and to dynamically update the multidimensional risk characteristics of each node in the spatiotemporal heterogeneous graph in discrete time steps according to the sampling action of the current cycle, and enter the deduction iteration of the next cycle.
9. The system according to claim 8, characterized in that, Also includes: The full-stack visualization and interaction module is used to receive the Pareto front data output by the strategy optimization module, the localized dispatch table output by the execution scheduling and evolution module, and the risk kernel density distribution data pushed by the risk perception module and the evolution module, and to present them dynamically on a large screen through a B / S architecture web front end.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.