Generation method of dynamic supervision optimal strategy and dynamic inspection decision method
By generating dynamic optimal supervision strategies for catering chain enterprises using POMDP and PBVI algorithms, the problems of regulatory lag and resource waste have been solved, and the optimization of inspection resources and improvement of store compliance levels have been achieved.
Patent Information
- Application Number
- CN202511684935.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-02-06
AI Technical Summary
Existing technologies for regulating restaurant chains suffer from problems such as regulatory lag, suboptimal resource allocation, and fragmented decision-making processes, leading to long detection times for violations, waste of resources, and brand damage.
A partially observable Markov decision process (POMDP) model is used, combined with Lagrange relaxation and point base value iteration (PBVI) algorithms, to generate a dynamic optimal regulatory strategy. Through Lagrange multiplier adjustment and Bayesian update, the optimization of inspection resources and the dynamic adjustment of the penalty strategy are realized.
This enabled efficient use of inspection resources, significantly reduced long-term total costs, improved store compliance, and enhanced brand image and operational quality.
Smart Images

Figure CN121481301A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and more specifically, to a method for generating optimal dynamic monitoring strategies and a method for dynamic inspection decision-making. Background Technology
[0002] In the modern restaurant chain industry, especially under the franchise model, one of the core challenges for headquarters is to effectively monitor the operational status of its stores on a large scale and in a standardized manner. The purpose of this monitoring is to ensure that all stores follow standard operating procedures (SOPs), thereby safeguarding brand reputation, food safety, and continued profitability.
[0003] To achieve oversight, existing technologies typically rely on a suite of digital systems, such as point-of-sale (POS) systems, inventory management systems (ERP), and Internet of Things (IoT) devices. These systems continuously generate massive amounts of operational data, providing an objective reflection of store operational status. Oversight primarily employs a passive monitoring mechanism based on fixed thresholds and manual auditing. The specific process is as follows: the headquarters data department periodically (e.g., monthly) compiles data reports from each store's systems, and operational analysts manually review these key data indicators. When a store's indicator consistently exceeds a preset fixed threshold, the system generates a "high-risk" alert, triggering a targeted offline manual inspection. This mechanism is also supplemented by low-frequency random sampling.
[0004] This regulatory approach has the following problems: 1) Significant regulatory lag: This mechanism is reactive. Industry data shows that for chain stores using this mechanism, the average time to discover violations is often several weeks, meaning that significant brand or financial damage may have already occurred by the time the problem is discovered.
[0005] 2) Mismatch and inefficiency of regulatory resources: Due to its reliance on static thresholds and random sampling, this mechanism cannot predict the dynamic evolution of risks. Regulatory results show that more than 60% of inspection resources are invested in stores with "false alarms" or long-term compliance, while truly high-risk stores become regulatory blind spots due to low sampling probability, resulting in inefficient resource allocation.
[0006] 3) Fragmented and non-optimal decision-making process: Inspection decisions and subsequent punishment or intervention decisions are two independent, reactive processes. Existing mechanisms fail to regard inspection behavior itself as a costly technical means aimed at reducing "state uncertainty," nor do they form a closed-loop optimization strategy with the goal of "lowest long-term total cost."
[0007] Therefore, there is an urgent need to develop a new regulatory method to achieve intelligent and precise use of inspection resources and improve the real-time, scientific, and reliable nature of inspections. Summary of the Invention
[0008] The purpose of this invention is to provide a method for generating the optimal dynamic monitoring strategy and a dynamic inspection decision-making method. This method can address the problems of lagging supervision, suboptimal resource allocation, and fragmented decision-making processes in the existing technology of catering chain enterprises. It provides an intelligent decision-making method and system that can comprehensively consider complex factors such as limited resources, unobservable state, and dynamic evolution.
[0009] The embodiments of the present invention are implemented as follows: A method for generating an optimal dynamic monitoring strategy for a restaurant chain enterprise, comprising: S11. Model the regulatory problem of a single store as a partially observable Markov decision process, and define its single-period expected cost function and long-term cumulative discounted total cost objective function; S12. Introduce Lagrange multipliers to internalize the resource constraints of multi-store inspections into the single-cycle expected cost function of a single store, and decompose the multi-store joint decision-making problem into multiple independent single-store decision-making sub-problems. S13. Based on the preset initial values of the Lagrange multipliers, the optimal strategy for each single-store decision sub-problem is solved using a point-based value iteration algorithm, and the total expected number of inspections under the current Lagrange multipliers is calculated. S14. Adjust the Lagrange multipliers and repeat step S13 until the total expected number of inspections matches the preset inspection resource limit to obtain the optimal Lagrange multipliers. S15. Combine the optimal strategies of each store under the optimal Lagrange multiplier to obtain the globally optimal decision strategy.
[0010] Furthermore, in other preferred embodiments of the present invention, in step S11, the single-cycle expected cost function includes inspection cost, expected indirect penalty cost, and expected violation cost.
[0011] Furthermore, in other preferred embodiments of the present invention, in step S11, the single-cycle expected cost function is: , In the formula, C i ( t (for stores) i At any moment t The expected cost per cycle, c I For unit inspection cost, i i ( t (for stores) i At any moment t Inspection decision variables, b i0( t (for stores) i At any moment t The probability of being in a compliant state. c F As a penalty cost coefficient, f i ( t (for stores) i At any moment t The severity of the punishment M This represents the total number of stores with violation status. b ik ( t (for stores) i At any moment t In the first k The probability of a certain violation status. c L ( k (for stores) i In the first k The losses incurred by the unit due to the violation of regulations.
[0012] Furthermore, in other preferred embodiments of the present invention, in step S11, the objective function for total cost is: , In the formula, π Let γ be the time discount factor for long-run costs, and γ∈[0,1]. N Total number of stores E The objective function is the expectation operator; it must satisfy the inspection resource constraints and penalty intensity constraints.
[0013] Furthermore, in other preferred embodiments of the present invention, in step S12, the modified single-cycle expected cost function is: , In the formula, λ For Lagrange multipliers and λ ≥0.
[0014] Furthermore, in other preferred embodiments of the present invention, the globally optimal decision-making strategy is stored as a set. α - Vector set, α - The set of vectors constitutes a piecewise linear convex value function.
[0015] A dynamic inspection decision-making method for restaurant chain enterprises, comprising: S21. At the beginning of each decision cycle, obtain the current belief status of each store; S22. Input the current belief state into the pre-generated global optimal decision-making strategy to generate the optimal action suggestion; S23. Based on the optimal action recommendations, conduct inspections of the selected stores and record the observation results; S24. Update the store's belief status based on the observation results for use in the next decision-making cycle; The globally optimal decision-making strategy is obtained by the method for generating the optimal dynamic monitoring strategy of the aforementioned catering chain enterprises.
[0016] Furthermore, in other preferred embodiments of the present invention, the optimal action recommendation includes inspection decision and penalty intensity.
[0017] Furthermore, in other preferred embodiments of the present invention, in step S23, if the number of selected stores exceeds the inspection resource limit, the selected stores are sorted according to the expected long-term cost reduction of performing inspections compared to not performing inspections, and the stores ranked higher are selected for inspections if the inspection resource limit is met.
[0018] Furthermore, in other preferred embodiments of the present invention, the method for updating the store belief state in step S24 is as follows: For the inspected stores, Bayes' theorem is used to update the prior information into posterior beliefs based on the observation results; then, according to the state transition model, the posterior beliefs of the inspected stores or the prior information of the uninspected stores are extrapolated to the prior information of the next decision cycle.
[0019] The beneficial effects of the embodiments of the present invention are: This invention provides a method for generating the optimal dynamic monitoring strategy for restaurant chains. It utilizes Partially Observable Markov Decision Process (POMDP) modeling, combined with Lagrange relaxation and Point Base Value Iteration (PBVI) algorithms, to optimize inspection resources and dynamically adjust penalty strategies, resulting in the optimal dynamic monitoring strategy. The optimal dynamic monitoring strategy obtained through this method is the best solution that minimizes the long-term cumulative expected cost, providing data support for dynamic monitoring. This invention also provides a dynamic inspection decision-making method. By dynamically assessing the risk of each store, it accurately allocates limited inspection resources to stores most likely to violate regulations or incur the greatest losses. The resulting decisions are scientifically sound and significantly improve the utilization rate of inspection resources, substantially reducing long-term total costs and significantly promoting the overall improvement of store compliance levels. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A flowchart illustrating the method for generating the optimal dynamic monitoring strategy for catering chain enterprises provided in Embodiment 1 of the present invention; Figure 2 This is a flowchart of the dynamic inspection decision-making method provided in Embodiment 2 of the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to represent selected embodiments of the invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] Example 1
[0024] This embodiment provides a method for generating the optimal dynamic monitoring strategy for a catering chain enterprise, the flowchart of which is shown below. Figure 1 As shown, it includes: S11. Model the regulatory problem of a single store as a partially observable Markov decision process, and define its single-period expected cost function and long-term cumulative discounted total cost objective function.
[0025] Partially Observable Markov Decision Processes (POMDPs) are a standard mathematical framework for handling sequential decision-making problems under uncertainty. They transform a non-Markovian observational history problem into a Markovian belief-state decision problem by introducing "belief states" (probability distributions of the true state). This allows headquarters to make rational decisions based on probability and expectation, even without knowing the true compliance status of stores.
[0026] Furthermore, in step S11, the single-cycle expected cost function includes inspection costs, expected indirect penalty costs, and expected violation costs.
[0027] Specifically, the single-cycle expected cost function is: , In the formula, C i ( t (for stores) iAt any moment t The expected cost per cycle, c I For unit inspection cost, i i ( t (for stores) i At any moment t Inspection decision variables, b i0 ( t (for stores) i At any moment t The probability of being in a compliant state. c F As a penalty cost coefficient, f i ( t (for stores) i At any moment t The severity of the punishment M This represents the total number of stores with violation status. b ik ( t (for stores) i At any moment t In the first k The probability of a certain violation status. c L ( k (for stores) i In the first k The losses incurred by the unit due to the violation of regulations.
[0028] In this formula, through c I i i ( t To characterize the inspection cost, i i ( t ) (1- b i0 ( t ))· c F f i ( t ) 2 To characterize the expected indirect penalty cost, This is used to characterize the expected losses from violations.
[0029] Furthermore, the total cost objective function is: , In the formula, π Let γ be the time discount factor for long-run costs, and γ∈[0,1]. NTotal number of stores E This is the expectation operator. The objective function must satisfy the inspection resource constraints. and penalty intensity constraints The ultimate goal of the total cost objective function is to find an optimal strategy. π* .
[0030] S12. Introduce Lagrange multipliers to internalize the resource constraints of multi-store inspections into the single-cycle expected cost function of a single store, thus decomposing the multi-store joint decision-making problem into multiple independent single-store decision-making sub-problems.
[0031] Lagrange relaxation is a classic method in operations research for solving large-scale optimization problems with coupled constraints. Its principle is to apply a penalty term (Lagrange multiplier) to the "hard constraints" (such as the total number of inspections not exceeding K). λ By moving these subproblems into the objective function, the original problem is decomposed into multiple independent subproblems. In other words, the original complex joint decision-making problem with constraints is successfully decomposed into N independently solvable single-store POMDP problems. Here, λ This can be interpreted as the opportunity cost of each additional inspection opportunity. λ When the temperature is high, inspections become "expensive," and the optimal strategy for each store will tend to be to avoid inspections; conversely, when the temperature is low, inspections will be preferred. This can be achieved by adjusting... λ This value can effectively regulate the overall inspection demand and match it with resource supply.
[0032] Furthermore, in step S12, the modified single-cycle expected cost function is: , In the formula, λ For Lagrange multipliers and λ ≥0.
[0033] S13. Based on the preset initial values of the Lagrange multipliers, the optimal strategy for each single-store decision subproblem is solved using a point-based value iteration algorithm, and the total expected number of inspections under the current Lagrange multipliers is calculated.
[0034] Theoretically, directly solving for the value function of POMDP is infeasible on a continuous belief space. The PBVI algorithm operates on the principle of "approximating the whole by focusing on specific points." It proves that the value function of POMDP is a piecewise linear convex function, which can be represented by a set of vectors (α-vectors). PBVI approximates the true value function with an approximate value function by iterating and updating these vectors only on a carefully selected set of representative belief points. This method significantly improves the efficiency of the solution while ensuring solution quality, making it possible to find near-optimal strategies for individual stores.
[0035] Specifically, for a given λ, a point-based value iteration algorithm is used to solve for the optimal strategy for each individual store. π* ( λ The core of PBVI is to iteratively update the value function over a representative set B in the belief space through a value backup process. For each belief point b∈B and each possible action a, its Q-value is updated: , In the formula, It is the probability of observing o when action a is performed. It is based on the observation o to update the belief b. o The value function of .
[0036] Subsequently, the value function of this belief point is updated to the minimum Q value among all actions, i.e.
[0037] This process is iterated repeatedly until the value function converges, ultimately yielding... λ Find the optimal strategy π* ( λ ).
[0038] Furthermore, calculations are performed in the current context. λ Below, the total expected number of inspections I( λ ): .
[0039] S14. Adjust the Lagrange multipliers and repeat step S13 until the total expected number of inspections matches the preset inspection resource limit, thus obtaining the optimal Lagrange multipliers.
[0040] Furthermore, adjustments are made through efficient algorithms such as binary search. λ The value is determined iteratively until an optimal one is found. λ* , so that I( λ* ) ≈ K.
[0041] S15. Combine the optimal strategies of each store under the optimal Lagrange multiplier to obtain the globally optimal decision strategy.
[0042] When the optimal one is found λ* Finally, the system generates the globally optimal strategy. π* This is the combination of optimal strategies for all stores, i.e., { π 1 * ( λ* ), π 2 * ( λ* ), ... πN * ( λ* This strategy is stored in a strategy library using an efficient data structure, allowing online decision-making methods to query and invoke it in real time.
[0043] Furthermore, in the implementation of the computer system, this globally optimal strategy π* It is not a dynamic computation process, but rather stored as a static, queryable data structure. The globally optimal decision strategy is stored as a set... α - A set of vectors, each α - Vectors all define a hyperplane in the belief space, all of these α - The upper envelope of the vectors together constitutes the piecewise linear convex value function V(b) of the entire POMDP problem.
[0044] In summary, the generation method employed in this embodiment utilizes the POMDP, a standard mathematical framework in decision theory for handling sequential decision-making under uncertainty. Its solution process follows the Bellman optimality principle, ensuring the generated strategy... π* Given a model and objective function, it is the optimal solution that minimizes the long-term cumulative expected cost, rather than merely a suboptimal solution based on intuition or experience.
[0045] Advantages of Random and Periodic Inspections: These two strategies are considered "state-independent" in decision theory, completely ignoring the dynamic evolution of "belief states" in each store. b ( t The generation method of this invention is a "state-dependent" method, and its generation strategy... π* It is inevitable to utilize the state of belief b ( t The information contained in the data includes all historical information and risk assessments, and therefore is necessarily better in terms of mathematical expectation than a strategy that discards this information.
[0046] Compared to heuristic methods, which are often short-sighted (such as prioritizing inspections of stores with the most historical violations), the generative method in this embodiment has two major theoretical advantages: first, it is forward-looking, as its value function explicitly considers the expected costs of all future periods; second, it can measure "information value," meaning the strategy may choose to inspect stores with high uncertainty to gain information gain, rather than repeatedly inspecting known problematic stores, thus achieving deeper optimization.
[0047] Meanwhile, theoretically, brute-force solutions to the combined POMDP are computationally infeasible. The generation method of this invention combines the Lagrange relaxation method with the PBVI approximation algorithm to transform a theoretically optimal but computationally infeasible path into a computationally feasible path that closely approximates the optimal solution. This is crucial for achieving large-scale applications.
[0048] Example 2
[0049] This embodiment provides a dynamic inspection decision-making method for restaurant chain enterprises, which utilizes the globally optimal strategy obtained in Embodiment 1. π* To make implementation decisions, refer to the flowchart. Figure 2 As shown, it includes: S21. At the beginning of each decision cycle, obtain the current belief status of each store.
[0050] Specifically, at the start of each decision cycle (e.g., at midnight each day), the system retrieves information from the database for each store within the network. i Current state of belief b i ( t The belief state is a vector, whose components are... b ik ( t (represents the store) i At any moment t In state k The probability of.
[0051] S22. Input the current belief state into the pre-generated global optimal decision-making strategy to generate the optimal action suggestion.
[0052] The system queries the pre-calculated optimal strategy. π* For each store i Enter its current belief state b i ( t This allows us to obtain the optimal action suggestion for the output. a i ( t ) = π i *b i ( t ).
[0053] Furthermore, this optimal action recommendation This clarified the decision-making process for inspecting the store. i i ( t ) and the intensity of punishment f i ( t ).
[0054] S23. Based on the optimal action recommendations, conduct inspections of the selected stores and record the observation results.
[0055] S24. Update the store's belief status based on the observation results for use in the next decision-making cycle.
[0056] Furthermore, in other preferred embodiments of the present invention, the method for updating the store belief state in step S24 is as follows: For the stores being inspected, Bayes' theorem is used based on the observation results. o i ( t Prior information b i ( t Updated to posterior beliefs :
[0057] .
[0058] Subsequently, based on the state transition model, the posterior beliefs of the inspected stores are determined. Or prior information of stores that have not been inspected b i ( t )(at this time, = b i ( t Prior information extrapolated to the next time step (t+1) b i ( t+ 1):
[0059] Updated belief status b i ( t+ 1) It will be stored as input for the next decision-making cycle, thus forming a decision-making loop.
[0060] Furthermore, in step S23, if the selected number of stores (i.e. i i ( t If the number of stores with 1 inspection resource exceeds the inspection resource limit K, then the selected stores are ranked according to the expected long-term cost reduction of performing inspections compared to not performing inspections. Under the condition of meeting the inspection resource limit, the top K stores in the ranking are selected to perform inspections.
[0061] Test case To verify the effectiveness of the dynamic inspection decision-making method provided by this invention, this experimental example compares the dynamic inspection decision-making method (POMDP) provided by this invention with random inspection, periodic inspection, and a heuristic method (i.e., prioritizing inspections of stores with the most violations in history). The long-term total cost, average compliance level of stores, and computational efficiency are compared. The specific experimental settings are as follows: total number of stores N=100; inspection limit per cycle K=10; number of store compliance statuses M=3 (0: compliant, 1: minor violation, 2: serious violation); simulation period T=200.
[0062] The comparison results are shown in Tables 1 and 2.
[0063] Table 1. Comparison of Long-Run Total Costs Strategy Cumulative total cost (unit: RMB 10,000) Cost savings compared to random strategies This invention (POMDP) 185.3 45.8% Heuristic 276.5 19.1% Fixed period (Periodic) 312.8 8.5% Random 342.1 - Table 2. Comparison of Average Compliance Levels of Stores Strategy Percentage of compliant stores (x=0) Percentage of stores with minor violations (x=1) Percentage of stores with serious violations (x=2) This invention (POMDP) 82% 15% 3% Heuristic 65% 25% 10% Fixed period (Periodic) 58% 31% 11% Random 55% 33% 12% As shown in Table 1, the dynamic inspection decision-making method provided by this invention exhibits the best performance in long-term total cost control, saving more than 45% of the cost compared to the traditional random inspection strategy, and also significantly outperforming heuristic and fixed-cycle strategies. This demonstrates that this method can achieve optimal economic benefits through forward-looking planning and resource optimization.
[0064] As shown in Table 2, after adopting the dynamic inspection and decision-making method provided by this invention, the proportion of compliant stores reached as high as 82%, while the proportion of seriously non-compliant stores was controlled to a minimum of 3%, significantly better than other strategies. This indicates that the dynamic inspection and penalty mechanism of this invention can effectively guide and incentivize stores to evolve towards a more compliant state, thereby improving the brand image and operational quality of the entire chain network.
[0065] Furthermore, in order to test the computational efficiency of the dynamic inspection decision-making method provided by this invention in dealing with large-scale problems, this experimental example also tested the offline computing time required to calculate the global optimal strategy under different numbers of stores (N), and the results are shown in Table 3.
[0066] Table 3. Calculation Efficiency Statistics Total number of stores (N) Calculation time (minutes) for the method of this invention Traditional joint decision-making model (estimation) 100 15.2 >12 hours 500 78.5 >24 hours 1000 160.1 Not feasible As shown in Table 3, the dynamic inspection decision-making method provided by this invention, based on its hierarchical framework, allows the computation time to increase approximately linearly with the number of stores, rather than exponentially. Even in a massive scenario with 1000 stores, the strategy can be solved within a few hours, demonstrating its feasibility and superior scalability in practical applications.
[0067] In summary, this invention provides a method for generating an optimal dynamic monitoring strategy for restaurant chains. It utilizes Partially Observable Markov Decision Process (POMDP) modeling, combined with Lagrange relaxation and Point Base Value Iteration (PBVI) algorithms, to optimize inspection resources and dynamically adjust penalty strategies, resulting in an optimal dynamic monitoring strategy. The optimal dynamic monitoring strategy obtained through this method is the best solution that minimizes long-term cumulative expected costs, providing data support for dynamic monitoring. This invention also provides a dynamic inspection decision-making method. By dynamically assessing store risks, it accurately allocates limited inspection resources to stores most likely to violate regulations or incur the greatest losses. The resulting decisions are scientifically sound and significantly improve the utilization rate of inspection resources, thereby reducing long-term total costs and significantly promoting overall improvement in store compliance levels.
[0068] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for generating the optimal dynamic monitoring strategy for a catering chain enterprise, characterized in that, include: S11. Model the regulatory problem of a single store as a partially observable Markov decision process, and define its single-period expected cost function and long-term cumulative discounted total cost objective function; S12. Introduce Lagrange multipliers to internalize the multi-store inspection resource constraints into the single-cycle expected cost function of a single store, and decompose the multi-store joint decision problem into multiple independent single-store decision sub-problems; S13. Based on the preset initial value of the Lagrange multiplier, the optimal strategy for each single-store decision sub-problem is solved using a point-based value iteration algorithm, and the total expected number of inspections under the current Lagrange multiplier is calculated. S14. Adjust the Lagrange multiplier and repeat step S13 until the total expected number of inspections matches the preset inspection resource limit to obtain the optimal Lagrange multiplier. S15. Combine the optimal strategies of each store under the optimal Lagrange multiplier to obtain the globally optimal decision strategy.
2. The generation method according to claim 1, characterized in that, In step S11, the single-cycle expected cost function includes inspection cost, expected indirect penalty cost, and expected violation cost.
3. The generation method according to claim 1, characterized in that, In step S11, the single-cycle expected cost function is: , In the formula, C i ( t (for stores) i At any moment t The expected cost per cycle, c I For unit inspection cost, i i ( t (for stores) i At any moment t Inspection decision variables, b i0 ( t (for stores) i At any moment t The probability of being in a compliant state. c F As a penalty cost coefficient, f i ( t (for stores) i At any moment t The severity of the punishment M This represents the total number of stores with violation status. b ik ( t (for stores) i At any moment t In the first k The probability of a certain violation status. c L ( k (for stores) i In the first k The losses incurred by the unit due to the violation of regulations.
4. The generation method according to claim 3, characterized in that, In step S11, the total cost objective function is: , In the formula, π Let γ be the time discount factor for long-run costs, and γ∈[0,1]. N Total number of stores E The objective function is the expectation operator; it must satisfy the inspection resource constraint and the penalty intensity constraint.
5. The generation method according to claim 4, characterized in that, In step S12, the corrected single-cycle expected cost function is: , In the formula, λ For Lagrange multipliers and λ ≥0.
6. The generation method according to claim 1, characterized in that, The globally optimal decision-making strategy is stored as a set. α - Vector set, the α - The set of vectors constitutes a piecewise linear convex value function.
7. A dynamic inspection decision-making method for catering chain enterprises, characterized in that, include: S21. At the beginning of each decision cycle, obtain the current belief status of each store; S22. Input the current belief state into the pre-generated global optimal decision-making strategy to generate the optimal action suggestion; S23. Based on the optimal action suggestion, conduct inspections of the selected stores and record the observation results; S24. Update the store's belief status based on the observation results for use in the next decision-making cycle; The globally optimal decision-making strategy is obtained by the method for generating the dynamic monitoring optimal strategy for catering chain enterprises as described in any one of claims 1 to 6.
8. The dynamic inspection decision-making method according to claim 7, characterized in that, The recommended optimal action includes inspection decisions and the intensity of penalties.
9. The dynamic inspection decision-making method according to claim 7, characterized in that, In step S23, if the number of selected stores exceeds the inspection resource limit, the selected stores are sorted according to the expected long-term cost reduction of performing inspections compared to not performing inspections. Under the condition of satisfying the inspection resource limit, the stores ranked higher are selected to perform inspections.
10. The dynamic inspection decision-making method according to claim 7, characterized in that, In step S24, the method for updating the store belief status is as follows: For the stores being inspected, Bayes' theorem is used to update prior information into posterior beliefs based on the observation results; Then, based on the state transition model, the posterior beliefs of the inspected stores or the prior information of the uninspected stores are extrapolated to the prior information of the next decision cycle.