Supply chain intelligent scheduling management method and system for fresh perishable products
By constructing a multi-agent game model and Bayesian optimization pricing strategy, and dynamically adjusting order priority and freight settlement, the problem of insufficient supply chain resilience in the fresh food e-commerce origin-based distribution model is solved, achieving efficient and stable supply chain operation and maximizing corporate profits.
Patent Information
- Application Number
- CN202511463403.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-02-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies have failed to effectively simulate the dynamic game between enterprise pricing, freight services, and customer choices in the fresh food e-commerce origin-based distribution model, resulting in insufficient supply chain resilience and making it prone to inventory mismatch, order delays, and additional costs.
A multi-agent game model is constructed to dynamically adjust order priorities, freight settlement, and delivery plans by integrating data from enterprises, customers, and freight forwarders. Bayesian optimization pricing strategies and meta-learning are used to handle objective function drift, thereby achieving intelligent scheduling management.
It enhances the resilience and flexibility of the supply chain, enabling real-time responses to market changes, optimized resource allocation, and improved decision-making accuracy and corporate profits.
Smart Images

Figure CN121563371A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of agricultural economics and technology, specifically to the optimization scheduling of product supply chains and the heterogeneous agent model (HAM), and particularly to intelligent scheduling and management methods and systems for the supply chains of fresh and perishable products. Background Technology
[0002] In the fresh food e-commerce origin-based distribution model, the supply chain revolves around three main entities: enterprises (e-commerce platforms / suppliers), customers (consumers), and freight forwarders (logistics service providers). Its core is to achieve efficient circulation of fresh products through direct delivery from the origin and centralized distribution. Enterprises integrate customer order and freight capacity data to dynamically adjust their harvesting / sorting plans at the origin. The basic distribution logic is: origin pre-cooling → trunk cold chain → regional warehouse distribution → last-mile delivery, relying on the freight forwarder network coverage. Finally, enterprises pay freight charges to the freight forwarders (often including KPI penalty clauses), and customer payments are distributed to various links in the supply chain through the platform.
[0003] Currently, numerous intelligent scheduling and management methods are emerging for the fresh produce e-commerce origin-based distribution model, with increasingly sophisticated algorithmic models. However, the underlying architecture of these various models is not entirely adapted to the aforementioned distribution logic. For example: (1) The literature "Liao Wenjing. Research on Graded Pricing and Logistics Network Optimization of Agricultural Products under the Production Area Distribution Model of Fresh Food E-commerce [D]. Xi'an University of Technology, 2023" discloses an optimization method, which constructs a joint optimization model of agricultural product quality graded pricing and optimized delivery departure time logistics network under the production area distribution model of fresh food e-commerce. However, focusing only on the enterprise's graded pricing and logistics network optimization, the graded pricing model may underestimate the transportation costs of high-value goods (such as cold chain energy consumption), resulting in the enterprise's profits being eroded by the freight cost of freighters. When optimizing the departure time, the elasticity of freight capacity (such as the cost of temporary dispatching during peak season) is not considered, which may cause regional warehouse overload or last-mile delivery delays, ultimately leaving the enterprise to bear the risk of customer compensation.
[0004] (2) The literature “Wang Zhao. Integration Optimization of Multi-Factory Production and Distribution of Perishable Products [D]. Dalian Maritime University, 2023” discloses an optimization method, which includes a two-layer heuristic algorithm of particle swarm optimization and genetic algorithm. The results of solving problems of different scales show that perishable enterprises can not only receive more orders and serve more customers by adopting multi-factory production, but also reduce the integration cost of each factory and improve the economic benefits of the enterprise. However, the multi-factory layout of this scheme may reduce the cost of a single factory, but if it does not match the actual needs of customers (such as changes in the hot spots of community group buying), it will lead to the accumulation of regional warehouse inventory or frequent transfers, increase the secondary delivery cost of freighters, and the enterprise needs to subsidize freighters, squeezing profits.
[0005] (3) The literature "Guo Runzhi. Research and Development of Fresh Food Cold Chain Logistics Vehicle-Inventory Scheduling System [D]. Southeast University, 2022" discloses an optimization method that can effectively solve the optimization problem of fresh food cold chain logistics vehicle-inventory scheduling. The modeling takes into account the energy-saving target of delivery vehicles and the volume constraints of goods stored in each temperature layer, taking into account the characteristics of fresh food cold chain logistics. However, this scheme only pursues the full load rate of vehicles, which may lead to insufficient regional warehouse inventory, forcing enterprises to rush to purchase high-priced goods, or customers to be lost due to shortages. The failure to distinguish the urgency of customer orders (such as the difference in the spoilage rate of fresh food) may lead to the delay of delivery of high-value orders, and enterprises will have to bear additional customer complaint costs.
[0006] (4) The literature "Li Xiaodong. Research on Optimization of Emergency Perishable Material Distribution with Fuzzy Corrosion Rate [D]. Xi'an University of Technology, 2023" discloses an optimization method that considers distance and demand and uses the K-means++ clustering algorithm to cluster regions; then, an improved differential evolution whale optimization algorithm is used to optimize the distribution of perishable materials in different cluster regions; however, pursuing full vehicle load may lead to insufficient regional warehouse inventory (such as during promotional periods), forcing enterprises to rush to purchase high-priced goods, or causing customers to be lost due to shortages. Failure to distinguish the urgency of customer orders (such as differences in the spoilage rate of fresh produce) may lead to delays in the delivery of high-value orders, and enterprises will have to bear additional customer complaint costs.
[0007] In summary, most models fail to consider the triangular game relationship between supply and demand mismatch, ultimately leading to a failure to relay customer demand to the production side (e.g., pre-sale data). This results in a mismatch between regional warehouse inventory and order distribution, requiring freight forwarders to frequently relocate across regions and increasing empty-load rates. In other words, while single-agent optimization improves local efficiency, it disrupts the feedback loop of "enterprise pricing → freight forwarder service → customer choice." A multi-agent game model is needed to incorporate freight settlement, revenue sharing rules, corruption rates, and customer satisfaction into global decision-making to improve supply chain resilience.
[0008] The paper "Liu Xianrun. Research on Disturbance Management of Fresh Agricultural Products Delivery Considering Simultaneous Pickup and Delivery [D]. Southwest Jiaotong University, 2022" discloses an optimization method that measures the disturbance to the delivery system caused by interference events from the perspectives of delivery companies, customers, and freight personnel. Based on the initial delivery model and combined with the system state when the disturbance occurs, it constructs a disturbance management model for fresh agricultural products delivery with simultaneous pickup and delivery as the objective, aiming to minimize the deviation of the generalized cost of the disturbance generated by the system. It also combines simulated annealing and genetic algorithms to design a hybrid algorithm to solve the initial delivery model and the disturbance management model. Although the model framework covers the feedback loop of "company pricing → freight personnel service → customer selection," these three entities only participate in the calculation as corresponding data in the model, and the game relationship between the three entities is not actually considered. The real contradiction lies in: (1) Static decision-making framework; assuming that the decision-making objectives (such as cost, timeliness, and driver fatigue) of enterprises, customers, and freight drivers remain fixed before and after the disturbance occurs, and only one-time optimization is performed through the objective function of "minimizing generalized cost deviation". For example: 1) Enterprises: In disruptive scenarios, they may dynamically adjust order priorities (such as prioritizing the delivery of high-value goods) rather than passively accepting generalized cost calculations.
[0009] 2) Customers: They may actively cancel orders or request compensation due to delivery delays (such as exemption from liability for spoilage of fresh produce), rather than being passively included in the model as "disturbance".
[0010] 3) Freighters: They may refuse some delivery tasks due to high detour costs, or negotiate additional subsidies with companies, rather than completely complying with algorithm scheduling.
[0011] The "optimal solution" output by the model lacks a dynamic negotiation mechanism between the subjects, which may lead to system collapse in actual implementation if one party refuses to cooperate.
[0012] (2) Local disturbance measurement: Only the immediate impact on the delivery route is calculated for a single disturbance event (such as traffic congestion), without considering the cascading effects of the disturbance event on long-term cooperative relationships; for example: 1) Business-customer game: Frequent delivery delays may trigger customer churn, and businesses need to weigh the short-term disruption costs against long-term customer value; 2) Business-freighter game: Demanding that freighters bear excessively high detour costs in the long run may lead to their withdrawal from cooperation or demands for higher freight rates; 3) Customer-freighter game: Customers request expedited delivery due to urgent needs and need to compensate freighters for additional fees, resulting in the transfer of hidden costs; The model cannot predict the erosion of the long-term stability of the supply chain by disruptive events, causing the "optimal solution" to fail after multiple rounds of game theory.
[0013] In summary, the current technical problem to be solved is: how to simulate the dynamic game process between "enterprise pricing → freight service → customer selection" and realize intelligent scheduling and management of products.
[0014] Therefore, this invention proposes a method and system for intelligent scheduling and management of the supply chain of fresh and perishable products. Summary of the Invention
[0015] In view of this, the present invention aims to provide a method and system for intelligent scheduling and management of the supply chain of fresh and perishable products, in order to solve or alleviate the technical problems existing in the prior art, namely: how to simulate the dynamic game process between "enterprise pricing → freight forwarder service → customer selection", including: (1) Subjective decision-making modeling: For businesses: dynamically adjust order priorities based on inventory turnover rate and customer lifetime value (CLV).
[0016] Customer: The order cancellation threshold is dynamically adjusted based on delivery time and product quality.
[0017] Freight drivers: Decide whether to accept detour assignments based on fuel costs and driver fatigue.
[0018] (2) Game rules: Enterprises can temporarily increase freight rates to attract freight drivers to prioritize the delivery of urgent orders.
[0019] (3) Dynamic feedback mechanism: The game result of a single disturbance is used as an input variable to update the parameters of the subsequent delivery plan (regional warehouse inventory threshold).
[0020] Then, based on the above model framework, the invention aims to enhance supply chain resilience in complex environments of "demand-supply mismatch" and achieve rapid and intelligent scheduling and management of fresh and perishable products. The technical solution of this invention is implemented as follows: Firstly, intelligent scheduling and management methods for the supply chain of fresh and perishable products: (I) Overview: This invention integrates data from three main stakeholders—enterprises, customers, and freight forwarders—to construct a model capable of simulating the dynamic game process of "enterprise pricing → freight forwarder service → customer selection." Through offline pre-training and online three-way game simulation, it can dynamically adjust order priorities, freight settlement, and delivery plans to adapt to changes in market demand. Simultaneously, it utilizes Bayesian optimization of pricing strategies (pricing combinations) and meta-learning to handle objective function drift, achieving intelligent pricing and rapid strategy updates. Finally, a dynamic feedback mechanism executes inventory threshold adjustments and route optimization to ensure the efficient and stable operation of the supply chain.
[0021] (II) Technical Solution: 2.1 Step S1, Data Preprocessing and Feature Engineering: Integrate inventory turnover rate from ERP (Electronic Data Interchange) systems. t (i.e., turnover ratio, which is cost of sales / average inventory), customer lifetime value (CLV) (based on the output of an existing LSTM model or CLV model) and preset historical order priority p; The cleaning and delivery timeliness sensor transmits GPS trajectory data G, product quality score s, and order cancellation record o; Integrate driver fatigue monitoring data from onboard sensors and fuel price API data; Crawling competitive pricing c p ; 2.2 Step S2, Offline Pre-training: Input features: Inventory level I, CLV quantiles (CLV) q (Through quantile normalization), historical cancellation rate o'.
[0022] A pre-trained XGBoost model for cancel threshold classification of client agents, where: Tags: Cancel order? Features: Delivery delay minutes, number of historical complaints.
[0023] 2.2.1 Step S200, Input feature vector X: X=[I, CLV q [o′, Delivery delay minutes, Number of historical complaints] Where: I∈R + Inventory level (unit: pieces); o′∈[0,1] is the historical cancellation rate (calculated as: o′=number of historically cancelled orders / total number of orders); delivery delay minutes∈N are the delivery timeout minutes extracted from GPS trajectory data; historical complaint count∈N is the number of complaints made by customers in the past 30 days.
[0024] 2.2.2 Step S201, XGBoost model structure: The XGBoost model is based on Gradient Boosting Trees (GBDT), and its predicted output is the sum of the outputs of each tree. For the binary classification problem in this approach, a logistic function is used to convert the output into a probability. That is, the model output is the probability of canceling the order. : ; Where σ(·) is the Sigmoid activation function; f m (⋅) represents the m-th regression tree (CART); M is the total number of trees (hyperparameter).
[0025] The objective function (binary cross-entropy loss + regularization) is: ; Among them, y i ∈{0,1} is the true label (1 = canceled order); Ω(·) = γ T +(1 / 2)λ'∥w∥ 2 This is the tree complexity penalty term; T is the number of leaf nodes; w is the leaf node weight vector; γ and λ' are regularization coefficients; Each tree f m (X) satisfies: ; Among them: Leaf j It is the feature space partitioning region corresponding to the j-th leaf node; I(⋅) is an indicator function (1 if the condition is met, 0 otherwise).
[0026] The model outputs a feature importance score after training. k : ; Among them, Split j The splitting condition for reaching leaf node j is k, where k is the feature index (corresponding to I, CLV). q (o, delay minutes, number of complaints).
[0027] 2.3 Step S3, Online Three-Party Game Simulation: Suppose a heterogeneous agent model includes three agents: [V1, V2, V3], representing a firm, a freight forwarder, and a customer, respectively. Execute the Vickrey auction process and determine the optimal response function for each agent's strategy based on Nash equilibrium. 2.3.1 Step S300: Execute the Vickrey auction process. S3000, the enterprise publishes an order and sets a reserve price P (based on historical freight distribution): ; Among them, h t ′ represents the historical freight rate distribution, Q a It is the α quantile function of the historical freight rate distribution. For example, when α=0.9, P takes the 90th percentile of the historical freight rate data as the reservation price. k is the risk adjustment coefficient, used to add a safety margin to the expected value (for example, taking k=1.5 means that the reservation price = mean + 1.5 times the standard deviation).
[0028] S3001, Freighter submits sealed quotation B i Including the cost of detours C detour and fatigue compensation C fatigue : ; Among them, C detour (d i ) is the detour cost function for driver i, d i The detour distance (unit: kilometers) is expressed in the form C. detour =γ⋅d i +δ (γ is the cost per kilometer, δ is the fixed cost); C fatigue ( hi ) is the fatigue compensation function, h i The continuous driving time of driver i (in hours) is represented by C. fatigue =η''⋅h i 2(η'' is the fatigue coefficient); λ is the fatigue compensation weighting coefficient, used to balance the relative importance of detour costs and fatigue compensation (λ>1 indicates greater emphasis on fatigue compensation); ϵ i It is a random perturbation term that follows a Gumbel distribution.
[0029] S3002, according to Vickrey's rule, the second-lowest bidder is selected, and the price paid is max(P,C). b ); where C b This is the second lowest bid, i.e., the price to be paid = max(P,C) b ); max(⋅) is the maximum value function, which ensures that the final payment price is not lower than the reservation price P.
[0030] S3003, Solving the optimization problem of a heterogeneous agent model: For a heterogeneous agent model, the optimization problem assumes the existence of continuous infinite agent units V. i The corresponding assets and income are a and y, respectively. Each agent unit V i The utility (i.e., the input) follows the standard expected discounted utility rule, with the expected discounted utility in the future; the steady-state distribution π* of the agent unit's income is calculated using a Hidden Markov Model (HMM): ; Wherein, β is the time discount factor, and the closer β is to 1, the more emphasis is placed on long-term returns; It is a CRRA (constant relative risk aversion) utility function, a it and y it It represents the assets and income of agent unit Vi at the current moment; ρ is the risk aversion coefficient; π∗(θ) i ) is of type θ i The agent's steady-state income distribution is calculated using the forward algorithm of a hidden Markov model. It is the Hidden Markov state transition probability, representing the agent V. i From state s at time t-1 i,t-1 Transition to state s at time t it The probability of A; i (s it ) is agent V i In state s it The set of feasible actions below; D Θ It is a continuous probability distribution over the agent type space, used to describe agent heterogeneity.
[0031] S301, Nash Equilibrium Game: In the later stages of the game (such as every 100 iterations in step S300), calculate the optimal response function of each agent's strategy; including: Business-customer game: Frequent delivery delays trigger customer churn, and businesses will weigh the short-term disruption costs against long-term customer value; The business-freighter game: Demanding that freighters bear excessively high detour costs in the long run will lead to them withdrawing from the cooperation or demanding higher freight rates; Customer-freighter game: When customers request expedited delivery, they may have to compensate the freight company for additional costs, resulting in the transfer of hidden costs. Based on the above game theory strategies, the following conclusions can be drawn: (1) Firm-customer game: ; Where, π e It is the firm's profit function; −θ(d)C loss θ(d) is the loss due to customer churn; θ(d) is the customer churn rate function caused by delivery delay d (the larger d is, the higher θ is); C loss -Cs is the average cost of churning a single customer; -Cs is the short-run disturbance cost (the temporary cost of expedited delivery); δV long It is the discounted value of long-term customer value; δ is the discount factor (0 < δ < 1, reflecting the present value of future earnings); V long It is customer lifetime value (expected returns from long-term cooperation, such as expected returns within one year); π c It is the customer revenue function; β(θ)U service It represents the perceived benefit of service quality; β(θ) is the customer satisfaction coefficient (the lower the θ, the higher the satisfaction); U service It is the basic service utility value; −P switch It is the switching cost (the economic cost of customers switching to competitors); constraint d≤d max This is the delivery delay threshold (exceeding this value increases customer churn rate).
[0032] (2) Enterprise-freighter game: ; Where πf is the firm's profit function; −C detour This refers to detour costs (additional expenses that companies require freight carriers to bear). ρ(ΔF) is the expected benefit of the freight rate adjustment; ρ(ΔF) is the probability function of the freighter's withdrawal caused by the freight rate adjustment amount ΔF; δ t It is a time discount factor (considering long-term cooperation benefits); R k π is the cooperative benefit in period k (such as the profit from improved freight efficiency); r is the discount rate; π is the return on investment in period k. t It is the freight passenger's revenue function; F base This is the basic freight revenue; −μC detour It is the cost burden of detours (μ is the proportion of cost transfer). σ is the exit compensation item; σ is the compensation trigger coefficient (0 < σ ≤ 1). It is an indicator function (triggered when the exit probability exceeds a critical value); K exitThis is the exit compensation amount, triggered by the condition ΔF ≥ ΔF. thr : Freight adjustment threshold (the freighter will be deactivated if the freight rate is exceeded).
[0033] (3) Customer-freighter game: ; Where, π c It is the customer revenue function; γU urgent It represents the utility of expedited service; γ is the probability or response parameter for expedited service requests; U urgent It is the utility value brought by expedited service; It is an additional cost of compensation; P λ It is the penalty coefficient for compensation (λ>0 indicates implicit cost premium); P extra This is the basic additional compensation amount; It is the expedited frequency indication function; π t It is the freight passenger's revenue function; P base It is the basic delivery revenue; η⋅E[γ]C hidden It represents the transfer of implicit costs to benefits; η''' is the transfer efficiency coefficient (0 < η ≤ 1); E[γ] is the expected value of the expedited request probability; C hidden These are unit implicit costs, and implicit cost constraints include: implicit cost ceiling constraint C. hidden ≤τ(γ)⋅Profit margin The transfer ratio function (positively correlated with γ) τ(γ) and freight profit margin margin .
[0034] In the three sets of games mentioned above, the strategy space s e s c (Enterprise / Customer Strategy Set) etc. contain all possible strategy combinations, i.e., pricing combinations.
[0035] Nash equilibrium condition This indicates that agent unit i is in its corresponding equilibrium strategy s i ∗ Below, other agent units maintain s i ∗ At that time, it is impossible to change the strategy unilaterally. i Increase profits.
[0036] 2.4 Step S4, Bayesian optimization of pricing strategy (pricing combination): Objective function: Π=(S×(P−C))−p d Where Π represents corporate profit, and p d It is a freight subsidy (fixed or dynamic value); In the search space: pricing range [P] min ,P maxThe values are discretized into 100 candidate values.
[0037] The price-profit curve is fitted using a Gaussian process (GP), and the next set of pricing experiments is selected by the expected improvement (EI) criterion.
[0038] The plan includes: (1) The objective function is to maximize Π=(S×(P−C))−p d ; Where S is sales volume; P is the pricing decision variable; and C is the unit cost. (2) Gaussian process (GP) fitting: ; Where m(p)=E[Π|P] is the mean function, k(P,P')=Cov(Π(P),Π(P')), and Cov(·) is the covariance function; (3) Set the expected improvement (EI) criterion: EI(P)=E[max(Π(P)-Π)] best ,0)] Among them: Π best The maximum profit value currently observed; E[⋅] is the expectation operation on the GP prediction distribution; (4) Bayesian optimization steps: S400, in [P] min ,P max Randomly sample initial price points within the range; S401, use GP to fit the price-profit data of the sample points; S402, calculate the EI value for each candidate price; S403, select the price point with the highest EI for the next round of experiments; S404, repeat steps S401~S403 until convergence or the budget limit is reached.
[0039] Step 2.5, S5, meta-learning to handle objective function drift: The MAML algorithm is used to quickly update the policy network parameters after drift is detected, based on the experience of historical tasks (balanced mode).
[0040] 2.5.1 Step S500, calculate the reward function: Define the reward function of the enterprise agent as a function of state s, action a, and the next state s': ; Among them, w i R represents the weight of each reward component. basic Based on the reward, R distance For distance bonus, R direction As a directional reward, R environmentAs a reward for environmental characteristics, R task Specific rewards for the task.
[0041] 2.5.2 Step S501, Real-time Profit Weight Monitoring: Set profit weight w profit The monitoring mechanism calculates its increase △w profi : ; in, and $ These are the profit weights for the current time and the previous time, respectively.
[0042] The trigger condition is: when the profit weight increases by more than the threshold θ. drift When this occurs, the objective function drift handling mechanism is triggered.
[0043] 2.5.3 Step S502, MAML algorithm executes objective function drift handling mechanism: For a series of related tasks, starting from the meta-initialization parameter θ', several gradient updates are performed for each task to obtain the task-specific parameter θ. i Then, calculate the loss on the task-specific parameters on the task validation set, and backpropagate to the meta-initialization parameters to update them: ; Where b is the meta-update step size, (·) represents task T i The loss function.
[0044] 2.6 Step S6, execution of the dynamic feedback mechanism: Enterprise inventory threshold IT adjustment: Using a rolling window, calculate the order fulfillment rate η' = total completed orders; if η' < 95%, increase the regional warehouse inventory threshold: S new =S old ×(1+Δ), where Δ is output by the PID controller.
[0045] Call the existing route optimization API, input real-time traffic data and a list of delivery tasks; and feed the optimized route cost back to the driver agent's decision-making model.
[0046] 2.6.1 Step S600, dynamic adjustment of inventory threshold: Order fulfillment rate calculation: The order fulfillment rate η' is calculated using a scrolling window and is defined as the total number of completed orders divided by the total number of orders. If the order fulfillment rate η' is lower than the target threshold, the inventory threshold adjustment mechanism is triggered: the regional warehouse inventory threshold is increased, and the new threshold S new From the old threshold S oldThe adjustment amount Δ is calculated from the output of the PID controller.
[0047] 2.6.2 Step S601, Path Optimization and Cost Feedback: Use Google OR-Tools, input real-time traffic data and a list of delivery tasks, to obtain optimized delivery routes. Based on the optimized delivery routes, calculate the cost C, including distance, time, and fuel consumption.
[0048] 2.6.3 Step S602: Cost feedback to the decision model: The cost C is fed back to the driver agent's decision-making model or administrator to adjust delivery strategies, resource allocation strategies, or update route planning algorithms.
[0049] (III) Mechanisms for solving technical problems: In terms of decision-making modeling, detailed decision-making models were developed for the three main stakeholders: enterprises, customers, and freight forwarders. (1) Enterprises: Dynamically adjust order priorities based on inventory turnover and customer lifetime value (CLV) to balance inventory costs and customer demand.
[0050] (2) Customers: The order cancellation threshold is dynamically adjusted based on delivery time and product quality to ensure a balance between freshness of fresh products and service quality.
[0051] (3) Freighters: Based on fuel costs and driver fatigue, decide whether to accept detours to ensure the economy and safety of transportation.
[0052] In terms of game theory rules, companies are allowed to temporarily increase freight rates in emergency situations to attract freight forwarders to prioritize the delivery of high-value or urgent orders, thereby flexibly responding to changes in market demand. By simulating the dynamic game process between different entities, market behavior can be predicted more accurately, resource allocation can be optimized, and the overall efficiency of the supply chain can be improved.
[0053] Regarding the dynamic feedback mechanism, the outcome of a single disturbance is used as an input variable to update parameters of subsequent delivery plans (such as regional warehouse inventory thresholds) in real time. This dynamic feedback mechanism enables the supply chain to quickly adapt to changes in the external environment, enhancing its resilience and flexibility.
[0054] Secondly, an intelligent scheduling and management system for the supply chain of fresh and perishable products: The system includes a processor and a memory connected to the processor. The memory stores program instructions, which, when executed by the processor, cause the processor to perform the intelligent supply chain scheduling and management method described above. The processor is connected to: (1) Decision modeling module: The deep reinforcement learning (Proximal Policy Optimization, PPO) algorithm is adopted, which can efficiently learn the optimal strategy in the continuous action space. By constructing a heterogeneous agent model, the agent (agent unit) takes actions (order priority ranking, cancellation threshold setting, detour decision, etc.) based on the current state (real-time inventory, customer lifetime value CLV, delivery time, etc.) and continuously optimizes the strategy through the reward mechanism to maximize long-term benefits.
[0055] (2) Three-party game simulation module: Using the Vickrey auction mechanism and Nash equilibrium theory, the game process between suppliers, logistics providers and platforms is simulated.
[0056] A Vickrey auction is a sealed-bid second-price auction that incentivizes participants to submit honest bids; a Nash equilibrium, on the other hand, is a strategy combination where unilaterally changing strategies by either party will not yield a better outcome.
[0057] By analyzing historical game data and market supply and demand, we can predict and achieve equilibrium freight rates and service ratings.
[0058] (3) Bayesian Optimization Pricing Module: The Bayesian Optimization (BO) algorithm is adopted. This algorithm constructs a probabilistic model of the objective function (such as a Gaussian process) to efficiently search for the optimal solution under uncertainty. Combining competitor pricing and demand elasticity analysis, the commodity price is dynamically adjusted to maximize profit or market share.
[0059] (4) Meta-learning Adaptation Module: Utilizing meta-learning algorithms such as MAML (Model-Agnostic Meta-Learning) or Reptile, it can quickly learn and adapt to new tasks on multiple related tasks. By detecting changes in the objective function (sudden changes in market demand), it quickly adjusts policy parameters to adapt to the new market environment.
[0060] Compared with the prior art, the beneficial effects of the present invention are: I. Multi-Agent Game Theory Model: This invention constructs a multi-agent game theory model involving three main entities: enterprises, customers, and freight forwarders. This model can comprehensively simulate the complex interactions within the supply chain. By simulating the dynamic game process between different entities, this invention can more accurately predict market behavior, optimize resource allocation, and improve the overall efficiency of the supply chain.
[0061] II. Dynamic Decision Modeling: Dynamic modeling was implemented for the decisions of enterprises, customers, and freight forwarders, considering multiple factors such as inventory turnover, customer lifetime value, delivery timeliness, product quality, fuel costs, and driver fatigue. This enables the invention to respond to market changes in real time, dynamically adjusting order priorities, freight settlement, and delivery plans, thereby improving the accuracy and flexibility of decision-making.
[0062] III. Bayesian Optimized Pricing Strategy (Pricing Combination): The pricing strategy is optimized using a Bayesian optimization algorithm. A Gaussian process is used to fit the price-profit curve, and the optimal pricing point is selected using the expectation-enhancing criterion. This pricing strategy maximizes corporate profits while reducing the risk of price fluctuations and improving market competitiveness. Simultaneously, a meta-learning algorithm is used to address the objective function drift problem. By monitoring changes in the reward function of the firm's agents, the strategy network parameters are quickly updated. This allows the invention to rapidly adapt to changes in the external environment, maintain the stability and effectiveness of the algorithm, and enhance the resilience and adaptability of the supply chain.
[0063] IV. Enhancing Supply Chain Resilience: Through a multi-agent game model and dynamic feedback mechanism, this invention can respond to market changes in real time, optimize resource allocation, and enhance the resilience and flexibility of the supply chain. Dynamic decision modeling and Bayesian optimization pricing strategies (pricing combinations) enable this invention to more accurately predict market behavior, improving the accuracy and scientific nature of decision-making. By optimizing pricing strategies and resource allocation, this invention can maximize corporate profits and enhance market competitiveness. Attached Figure Description
[0064] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments or technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0065] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a schematic diagram of the classification model architecture for step S2 of the present invention; Figure 3 This is a game-theoretic diagram of the agent model in step S3 of the present invention. Figure 4 This is a schematic diagram of the agent model response function in step S3 of the present invention; Figure 5 This is a schematic diagram of the Bayesian optimization participation game process in step S4 of the present invention. Figure 6 This is a schematic diagram of the meta-parameter update in step S5 of the present invention; Figure 7 This is a schematic diagram of the system composition of the present invention; Figure 8 This is a schematic diagram of the game distribution effect in step S3 of the present invention; Figure 9 This is a schematic diagram of the strategy distribution for step S4 of the present invention. Detailed Implementation
[0066] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below; It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0067] Explanation of relevant terms: (1) Trajectory data G: GPS data transmitted by the delivery time sensor, which records the actual transportation route and location information of the vehicle or goods.
[0068] (2) Product quality score s: A quantitative assessment of the quality of products (especially fresh and perishable products), based on a variety of factors such as sensory inspection and physicochemical indicators, and is a response parameter.
[0069] (3) Order cancellation record: Records historical data of customer order cancellation, including cancellation time, reason and other information.
[0070] (4) Customer Lifetime Value (CLV): The sum of the present value of net profit generated by a customer over its entire life cycle. It is an important indicator for measuring the long-term value of a customer.
[0071] (5) Inventory Level I: The quantity or value of goods in a company’s warehouse at a given moment.
[0072] (6) Historical cancellation rate o': The frequency or proportion of customers canceling orders in the past period.
[0073] (7) Regression Tree (CART): A tree-structured model for classification and regression that builds the model by recursively partitioning the feature space.
[0074] (8) Reserve price P: The lowest acceptable price set by the enterprise.
[0075] (9) Detour cost function Cdetour(d i ): Describes the additional cost incurred due to taking a detour, which is the detour distance d. i The function.
[0076] (10) Firm revenue function π e : Measure the benefits a company receives under a given strategy, including revenue, costs, and customer churn loss.
[0077] (11) Customer revenue function π c : Measures the customer’s benefit under a given strategy, i.e., the benefit remaining after offsetting additional compensation costs.
[0078] (12) Perceived service quality benefit β(θ)U service : The benefits derived from customers' perception of service quality, where β(θ) is the customer satisfaction coefficient, U service It is the utility value of basic services.
[0079] (13) Basic service utility value U service Basic service utility without the influence of additional factors (such as expedited services).
[0080] (14) Firm benefit function π f The total benefit of an enterprise when considering the long-term benefits of cooperation.
[0081] (15) Expected benefits of freight rate adjustment: The benefits that the company expects to gain from adjusting freight rates, including attracting more freight drivers, etc.
[0082] (16) Probability function ρ(Δ) F ): Describes the impact of freight adjustment ΔF on the probability of freighters withdrawing from the freight service.
[0083] (17) Time discount factor δ t : A factor used to discount future returns to the present when calculating long-term returns.
[0084] (18) Freighter revenue function π t The revenue of a freight forwarder under a given strategy includes basic freight revenue, detour cost burden, and exit compensation.
[0085] (19) Urgent frequency indication function: a function that describes the frequency or probability of a customer requesting urgent service or is simply used for response.
[0086] (20) Implicit cost transfer of revenue: The additional revenue obtained by the freight forwarder due to the customer's expedited request is, in principle, borne by the customer; however, when the freight is in force majeure as stipulated in the Civil Code, this part of the revenue is transferred to the enterprise (because the cost is borne by the freight forwarder).
[0087] (21) Expected value of expedited request probability E[γ]: The expected value of the probability of a customer requesting expedited service, or it can be a simple response value.
[0088] (22) Unit implicit cost C hidden The hidden cost incurred per unit of expedited service request.
[0089] (23) Strategy space S e S c (Business / Customer Strategy Set): The set of all possible strategies that a business and a customer might adopt during a game (e.g., when the business's costs increase, the customer's payoff function π). c According to a certain discount reduction, that is, a set of "IF-THEN" rules.
[0090] (24) EI value: Expected improvement value, used in Bayesian optimization to select the next set of pricing experiments to maximize expected returns.
[0091] (25) State s: The state description of the reward function of the firm agent at a certain moment.
[0092] (26) Action a: The action taken by the firm agent when the reward function transitions from the current state to the next state.
[0093] (27) Next state s': The new state that is transitioned from the current state s after taking action a.
[0094] (28) Basic reward R basic Basic reward without additional rewards or penalties.
[0095] (29) Distance reward R distance Additional rewards or penalties based on the transportation distance.
[0096] (30) Directional reward R direction Additional rewards or penalties are given based on the degree of optimization of the transportation direction or route.
[0097] (31) Environmental reward: Additional rewards or penalties given based on external environmental characteristics (such as weather, traffic conditions).
[0098] (32) Task-specific reward R task Additional rewards or penalties given for specific tasks or objectives.
[0099] (33) "Meta-initialization" parameter θ': In meta-learning, the parameter used to initialize the policy network.
[0100] (34) Task-specific parameter θi': For a specific task or environment, the parameter obtained from the "meta-initialization" parameter θ' is updated by gradient.
[0101] (35) Task T i In meta-learning, it is one of a series of related tasks used for training or testing.
[0102] (36) Cost C: The sum of all costs incurred during the delivery process.
[0103] (37) Decision model: A model used to make decisions based on input data (such as order information, inventory level, traffic conditions, etc.); it can be an existing delivery model based on cost C as a consideration factor (independent variable).
[0104] (37) Adjust delivery strategy: Adjust the delivery plan based on cost C according to market changes or customer needs, including changing delivery routes and adjusting delivery time.
[0105] (38) Adjusting resource allocation strategies: Under limited resources, maximize benefits or efficiency through allocation (such as vehicles, personnel, inventory, etc.).
[0106] (39) Update the route planning algorithm: The existing route planning algorithm is mainly based on cost C, and the route planning algorithm is updated according to factors such as real-time traffic data and order changes in order to optimize delivery routes and efficiency; (40) Customer churn rate function θ(d): represents the customer churn rate caused by delivery delay d. This function is an increasing function, that is, the longer the delivery delay, the higher the customer churn rate.
[0107] Example 1: As Figure 1 As shown in the figure, this embodiment discloses a method and system for intelligent scheduling and management of the supply chain of fresh and perishable products; specifically, it includes the following steps S1 to S6.
[0108] In this embodiment, regarding step S1, data preprocessing and feature engineering: integrating the inventory turnover rate s from the ERP system. t (i.e., turnover ratio, which is cost of sales / average inventory), customer lifetime value (CLV) (based on the output of an existing LSTM model or CLV model) and preset historical order priority p; The cleaning and delivery timeliness sensor transmits GPS trajectory data G, product quality score s, and order cancellation record o; Integrating driver fatigue monitoring data from onboard sensors (w) and fuel price API; crawling competitive pricing data (c) p ; Specifically, the GPS trajectory data G transmitted by the delivery timeliness sensors is cleaned, removing noise and outliers. The fusion of driver fatigue monitoring data w helps understand driver fatigue levels and prevent accidents and delivery delays caused by drowsy driving. The integration of fuel price APIs enables companies to obtain real-time fuel price information, providing a basis for cost calculations in scheduling decisions.
[0109] In this embodiment, regarding step S2, offline pre-training: (1) Input features include: inventory level I, CLV quantile CLV q (Through quantile normalization), historical cancellation rate o'.
[0110] A pre-trained XGBoost model for cancel threshold classification of client agents, where: (2) Label: Whether the order has been cancelled; (3) Features: Delivery delay minutes, number of historical complaints.
[0111] Inventory level I directly impacts supply chain responsiveness and costs; high inventory can increase the risk of spoilage, while low inventory can lead to stockouts. CLV quantiles (CLV) q It helps businesses identify high-value customer groups, thereby prioritizing them in scheduling decisions and improving customer satisfaction and loyalty. Historical cancellation rate (o') provides historical patterns of customer order cancellation behavior, predicting future cancellation risk.
[0112] Specifically, in step S200, input the feature vector X: X=[I, CLV q [o′, Delivery delay minutes, Number of historical complaints] Where: I∈R + Inventory level (unit: pieces) reflects the quantity of fresh and perishable products in the current inventory. o′∈[0,1] is the historical cancellation rate (calculated as: o′ = number of historically cancelled orders / total number of orders), representing the frequency with which customers cancel orders. Delivery delay minutes∈N are the number of minutes of delivery delay extracted from GPS trajectory data, reflecting the timeliness of order delivery. Historical complaint count∈N is the number of complaints made by customers in the past 30 days, reflecting the degree of customer dissatisfaction with the service.
[0113] Specifically, in step S201, the XGBoost model structure is as follows: Figure 2As shown: The XGBoost model is based on Gradient Boosting Trees (GBDT), and its predicted output is the sum of the outputs of each tree. Strong learners are built by iteratively training weak learners (decision trees), with each tree attempting to correct the errors of the previous tree. For the binary classification problem in this approach, a logistic function is used to convert the output into a probability. That is, the model output is the probability of canceling the order. : ; Where σ(·) is the Sigmoid activation function, which converts the output of the regression tree into probability values and is suitable for binary classification problems. m (⋅) represents the m-th regression tree (CART); M is the total number of trees (hyperparameter).
[0114] The objective function includes a binary cross-entropy loss (which measures the difference between the predicted probability and the true label) and a regularization term (to prevent the model from overfitting): ; Among them, y i ∈{0,1}: True label (1 = canceled order); Ω(·) = γ T +(1 / 2)λ'∥w∥ 2 This is the tree complexity penalty term; T is the number of leaf nodes; w is the leaf node weight vector; γ and λ' are regularization coefficients; ||·|| 2 It is the square of the Euclidean distance. This objective function reflects the importance of each feature in the model's decision-making, allowing businesses to understand which factors have the greatest impact on order cancellations.
[0115] Furthermore, each tree f m (X) satisfies: ; Among them: Leaf j It is the feature space partitioning region corresponding to the j-th leaf node; I(⋅) is an indicator function (1 if the condition is met, 0 otherwise).
[0116] Furthermore, after model training, the output feature importance score is... k : ; Among them, Split j The splitting condition for reaching leaf node j is k, where k is the feature index (corresponding to I, CLV). q (o, delay minutes, number of complaints).
[0117] Understandably, XGBoost, by integrating multiple weak learners, can capture complex patterns in data. Through gradient boosting and regularization techniques, it continuously optimizes model parameters, enabling the model to gradually approximate the true data distribution and improve the accuracy of cancellation order predictions. Feature importance scores allow companies to understand which factors have the greatest impact on cancellation orders (XGBoost records the contribution of each feature in the decision tree split, thus calculating the feature importance score), providing support for decision-making. In the dynamic game process of "company pricing → freight forwarder services → customer selection," companies can allocate resources more rationally, avoiding unnecessary waste and losses.
[0118] In this embodiment, as Figure 3 As shown, regarding step S3, the online three-way game simulation: Assume the heterogeneous agent model includes three agents: [V1, V2, V3], representing the enterprise, the freight forwarder, and the customer, respectively. Wherein: (1) Agent V1: Responsible for pricing and inventory management, with the goal of maximizing profits while meeting market demand.
[0119] (2) Agent V2: Provides transportation services with the goal of maximizing transportation revenue while meeting the company’s delivery requirements.
[0120] (3) Agent V3: Choose whether to purchase the product, with the goal of maximizing its own utility (such as price, quality, delivery time, etc.).
[0121] Among the three agents, the Vickrey auction process is used to simulate the interaction between agent V1's pricing and agent V2's service offers, as well as agent V3's purchasing decisions based on these offers. By detecting and constraining the optimal response function of each agent's strategy using the Nash equilibrium algorithm, the game can be ensured to reach Nash equilibrium, thereby achieving supply chain stability and efficiency.
[0122] Specifically, in step S300, the Vickrey auction process is executed: S3000, the enterprise publishes an order and sets a reserve price P (based on historical freight distribution): ; Among them, h t ′ represents the historical freight rate distribution, Q a It is the α quantile function of the historical freight rate distribution. For example, when α=0.9, P takes the 90th percentile of the historical freight rate data as the reservation price. k is the risk adjustment coefficient, used to add a safety margin to the expected value (for example, taking k=1.5 means that the reservation price = mean + 1.5 times the standard deviation).
[0123] Understandably, the reserve price, serving as the floor price in an auction, ensures that businesses do not accept services below cost. Using historical data to set the reserve price makes it more reasonable and market-responsive, thereby increasing the safety margin to cope with market volatility and uncertainty.
[0124] S3001, Driver submits sealed quote B i Including detour costs C detour and fatigue compensation C fatigue : ; Among them, C detour (d i ) is the detour cost function for driver i, d i The detour distance (unit: kilometers) is expressed in the form C. detour =γ⋅d i +δ (γ is the cost per kilometer, δ is the fixed cost); C fatigue ( hi ): Fatigue compensation function, h i The continuous driving time of driver i (in hours) is represented by C. fatigue =η''⋅h i 2 (η'' is the fatigue coefficient); λ is the fatigue compensation weighting coefficient, used to balance the relative importance of detour costs and fatigue compensation (λ>1 indicates greater emphasis on fatigue compensation); ϵ i It is a random perturbation term that follows a Gumbel distribution (used to simulate the uncertainty of driver quotes).
[0125] Understandably, this section considers the additional costs incurred by drivers due to detours. It also considers the fatigue costs of continuous driving, quantifying them through a fatigue compensation function. Furthermore, it simulates the uncertainty of driver quotes, making the quotes more realistic.
[0126] S3002, according to Vickrey's rule, the second-lowest bidder is selected, and the price paid is max(P,C). b ); where C b This is the second lowest bid, i.e., the price to be paid = max(P,C) b ); max(⋅) is the maximum value function, which ensures that the final payment price is not lower than the reservation price P.
[0127] It's important to note that the core of the Vickrey auction process is "selecting the second-lowest bidder as the winner and paying the second-lowest bid (or the higher of the reserve prices)," which aims to incentivize bidders to accurately reflect their costs. The requirement that the bid price not be lower than the reserve price ensures that the interests of companies are not harmed and also avoids cutthroat competition.
[0128] S3003, Solving the optimization problem of a heterogeneous agent model: For a heterogeneous agent model, the optimization problem assumes the existence of continuous infinite agent units V. i The corresponding assets and income are a and y, respectively. Each agent unit V i The utility (i.e., the input) follows the standard expected discounted utility rule, and its future expected discounted utility; the steady-state distribution π* of the agent unit's revenue is calculated using a Hidden Markov Model (HMM): ; Where β is the time discount factor (0 < β < 1), which represents the agent's emphasis on future utility. The closer β is to 1, the more emphasis is placed on long-term returns. It is the CRRA (Constant Relative Risk Aversion) utility function, where ρ is the risk aversion coefficient (ρ > 0: risk aversion, ρ = 0: risk neutrality, ρ < 0: risk-seeking); π∗(θ i ) is of type θ i The agent steady-state income distribution is calculated using the forward algorithm of a Hidden Markov Model (HMM); It is the Hidden Markov state transition probability, representing the agent V. i From state s at time t-1 i,t-1 Transition to state s at time t it The probability of A; i (s it ) is agent V i In state s it The following set of feasible actions (accept / reject the order, and adjust the quote); D Θ It is a continuous probability distribution on the agent type space, used to describe agent heterogeneity (the detour cost coefficient γ of different drivers follows a normal distribution).
[0129] Understandably, this section considers the agent's discounted future utility, reflecting the agent's time preference. The Hidden Markov Model (HMM) is used to calculate the steady-state distribution of the agent's income, taking into account the dynamic changes in the agent's state. Constraints ensure that the agent's actions and income conform to their type and state. By considering the heterogeneity and dynamic changes of the agent, decision-making becomes more accurate and efficient. The application of the HMM allows the model to adapt to changes under different environments and conditions. The expected discounted utility rule reflects the agent's time preference through the discount factor β. The HMM calculates the steady-state distribution of the agent's income using a forward algorithm, considering the transition probabilities of the agent's state.
[0130] Specifically, step S301, such as Figure 4 , 8As shown, the Nash equilibrium game involves calculating the optimal response function for each agent's strategy in the later stages of the game (e.g., every 100 iterations in step S300); including: Business-customer game: Frequent delivery delays trigger customer churn, and businesses will weigh the short-term disruption costs against long-term customer value; The business-freighter game: Demanding that freighters bear excessively high detour costs in the long run will lead to them withdrawing from the cooperation or demanding higher freight rates; Customer-freighter game: When customers request expedited delivery, they may have to compensate the freight company for additional costs, resulting in the transfer of hidden costs. In this embodiment, regarding step S4, as follows: Figure 5 , 9 As shown, the Bayesian optimization pricing strategy (pricing combination) aims to enhance supply chain resilience and achieve intelligent scheduling management of fresh and perishable products in a complex environment of "demand-supply mismatch" by simulating the dynamic game process between "firm pricing → freight service → customer choice". This strategy uses a Gaussian process (GP) to fit the price-profit curve and selects the next set of pricing experiments using the expected improvement (EI) criterion to maximize firm profits.
[0131] Specifically, the objective function f: Π=(S×(P−C))−p d Where Π represents corporate profit, and p d It is a freight subsidy (fixed or dynamic value); In the search space: pricing range [P] min ,P max The values are discretized into 100 candidate values.
[0132] The price-profit curve is fitted using a Gaussian process (GP), and the next set of pricing experiments is selected by the expected improvement (EI) criterion.
[0133] The plan includes: (1) The objective function is to maximize Π=(S×(P−C))−p d ; Here, S represents sales volume; P is the pricing decision variable; and C is the unit cost. That is, by adjusting the price P, a company can influence sales volume S, and thus further influence profit Π.
[0134] (2) Gaussian process (GP) fitting: ; Where m(p) = E[Π|P] is the mean function, representing the expected value of profit Π at a given price P. k(P,P') = Cov(Π(P),Π(P')) represents the correlation between profits at different price points. Cov(·) is the covariance function, a nonparametric Bayesian model used to fit the price-profit curve.
[0135] (3) Set the expected improvement (EI) criterion: EI(P)=E[max(Π(P)-Π)] best ,0)] Among them: Π best The currently observed maximum profit value; E[⋅] is the expectation operation of the GP prediction distribution; it provides an effective exploration-utilization equilibrium mechanism, which both explores new price points and optimizes pricing using known information. It can accelerate convergence to the optimal pricing strategy.
[0136] (4) Bayesian optimization steps: S400, in [P] min ,P max Randomly sample initial price points within the range; S401, use GP to fit the price-profit data of the sample points; S402, calculate the EI value for each candidate price; S403, select the price point with the highest EI for the next round of experiments; S404, repeat steps S401~S403 until convergence or the budget limit is reached; then update the policy space based on the price point corresponding to the EI of the last iteration.
[0137] Understandably, this approach utilizes Gaussian processes to fit sampled data, establishes a price-profit model, calculates the expected increase in price for each candidate price, and selects the optimal price point for experimentation, aiming to gradually approach the optimal pricing strategy. It provides a systematic and efficient method for finding the optimal pricing strategy. It can handle high-dimensional, nonlinear, and non-convex optimization problems. Through an iterative optimization process, the price-profit model is progressively updated, approximating the optimal pricing strategy. Each iteration selects the optimal price point for experimentation based on the current model, thereby accelerating convergence.
[0138] In this embodiment, as Figure 6 As shown, regarding step S5, the meta-learning process handles objective function drift: It aims to simulate the dynamic game process between "enterprise pricing → freight forwarder service → customer choice," enhancing supply chain resilience in a complex environment of "demand-supply mismatch" and enabling intelligent scheduling management of fresh and perishable products. This step sets and monitors changes in the reward function of the enterprise's agents. When the profit weight increases beyond a threshold, an objective function drift handling mechanism is triggered. The MAML (Model-Agnostic Meta-Learning) algorithm leverages historical task experience to quickly update the policy network parameters.
[0139] Set and monitor changes in the reward function of the enterprise's agents; when the profit weight increases beyond a threshold θ driftThe event is triggered in real time (corresponding to liability exemption for fresh food spoilage and cost fluctuations in reality); after detecting drift, the MAML algorithm is used to quickly update the policy network parameters based on the experience of historical tasks (balanced mode).
[0140] Specifically, in step S500, the reward function is calculated: Define the reward function of the enterprise agent as a function of state s, action a, and the next state s': ; Among them, w i R represents the weight of each reward component. basic Based on the reward, R distance For distance bonus, R direction As a directional reward, R environment As a reward for environmental characteristics, R task The rewards are specific to the task and are all preset reward factors or response parameters.
[0141] Specifically, in step S501, set the profit weight w. profit The monitoring mechanism calculates its increase △w profi : ; in, and $ These are the profit weights for the current time and the previous time, respectively.
[0142] The trigger condition is: when the profit weight increases by more than the threshold θ. drift When this occurs, the objective function drift handling mechanism is triggered.
[0143] Specifically, in step S502, the MAML algorithm implements a target function drift handling mechanism: The MAML algorithm aims to quickly adapt to new tasks by learning from past tasks. It seeks a set of "meta-initialization" parameters θ' such that, given a small amount of sample data from new tasks, good performance can be achieved with only one or two gradient updates. The solution is: starting from the meta-initialization parameters θ', several gradient updates are performed for each task across a series of related tasks to obtain the task-specific parameters θ. i Then, calculate the loss on the task-specific parameters on the task validation set, and backpropagate to the meta-initialization parameters to update them: ; Where b is the meta-update step size, (·) represents task T i The loss function.
[0144] In the meta-testing phase, given a new target task, the same gradient updates as in the meta-training phase are performed using the current meta-initialization parameters θ' to obtain the task-specific parameters θ''. At this point, the task-specific parameters already possess the ability to quickly adapt to the new task. The MAML algorithm, through the meta-initialization parameters learned in the meta-training phase, can quickly adapt to the new task in the meta-testing phase. By leveraging experience from historical tasks, the MAML algorithm can reduce the sample requirements for new tasks and improve learning efficiency.
[0145] In this embodiment, regarding step S6, the dynamic feedback mechanism is implemented to simulate the dynamic game process between "enterprise pricing → freight service → customer selection," thereby enhancing supply chain resilience in a complex environment of "demand-supply mismatch" and enabling intelligent scheduling management of fresh and perishable products. This step achieves dynamic adjustment and optimization of the supply chain through dynamic adjustment of inventory thresholds and API calls for path optimization.
[0146] Inventory threshold IT adjustment: Using a rolling window (e.g., data from the past 7 days), calculate the order fulfillment rate η' = total completed orders; if η' < 95%, increase the regional warehouse inventory threshold: S new =S old ×(1+Δ), where Δ is output by the PID controller.
[0147] Call the existing route optimization API (Google OR-Tools), input real-time traffic data and a list of delivery tasks; and feed the optimized route cost back to the driver agent's decision-making model.
[0148] Specifically, in step S600, the inventory threshold is dynamically adjusted: the order fulfillment rate η' is calculated using a rolling window (e.g., data from the past 7 days), defined as the total number of completed orders divided by the total number of orders; if the order fulfillment rate η' is lower than the target threshold (e.g., 95%), the inventory threshold adjustment mechanism is triggered: the regional warehouse inventory threshold is increased, and the new threshold S... new From the old threshold S old The adjustment amount Δ from the PID controller output is calculated as follows: ; Wherein, △ is calculated by the PID controller based on the deviation between the order fulfillment rate and the target threshold: e(t) = η target −η(t); ; Where: e(t) is the target threshold η target deviation; K p It is a proportionality coefficient used to adjust the system's immediate response to deviations; K i These are integral coefficients used to eliminate steady-state errors; the integral term is the accumulation of the deviation. Kd These are differential coefficients, used to predict the trend of deviation changes and reduce overshoot; (τ)dτ is the integral of the bias, representing the cumulative effect of historical bias; It is the differential of the deviation, representing the rate of change of the deviation.
[0149] Understandably, step S600 uses the PID controller and the order fulfillment rate as feedback signals to trigger the inventory threshold adjustment mechanism, thereby achieving precise control of inventory levels and reducing inventory backlog and stockouts.
[0150] Specifically, in step S601, route optimization and cost feedback: Google OR-Tools is used, real-time traffic data and a delivery task list are input to obtain optimized delivery routes. Based on the optimized delivery routes, the cost C is calculated, including distance, time, and fuel consumption factors. Distance cost ; Time cost ; fuel consumption cost ; in This represents two adjacent points p on the path. i and p i+1 The distance between them, where n is the number of points on the path; Indicates from point p i Point P i+1 The required time takes into account real-time traffic data; Indicates from point p i Point P i+1 Fuel consumption, taking into account vehicle type; Preferably, d(·), t(·), and f(·) are essentially symbolic functions that, based on the type of input data and a pre-defined parameter lookup table, output pre-defined data. (1) Distance cost function d(p) i ,p i+1 ): A calculation function that uses the Haversine formula based on the geographic coordinates (such as latitude and longitude) between two points, or a lookup table function based on a predefined path network. ; Where r is the Earth's radius, ϕ i and ϕ i+1 It is the latitude of two points, λ i and λ i+1 It is the longitude of the two points.
[0151] (2) Time cost function t(p) i ,p i+1This can be represented as a calculation function based on distance and average speed, or as a dynamic function that takes into account real-time traffic data. ; Among them, v avg This is the average driving speed. t(p) i ,p i+1 You can obtain the estimated travel time between two points by calling the traffic data API, or make a prediction based on historical traffic data and current traffic conditions.
[0152] (3) Fuel consumption cost function f(p) i ,p i+1 This is a calculation function based on distance, vehicle type, and fuel efficiency. ; Among them, FE vehicle Fuel efficiency refers to the fuel efficiency (fuel consumption per unit distance) of a specific vehicle. Fuel efficiency can be further refined into values under different driving speeds, road conditions, and load conditions. In practical applications, fuel efficiency may also need to be dynamically adjusted based on various factors such as vehicle type, driving habits, and road conditions.
[0153] Preferably, for the distance function d(⋅), the parameter lookup table stores the distance between different pairs of geographic coordinates; for the time function t(⋅), it stores the estimated travel time for different route segments under specific traffic conditions; and for the fuel consumption function f(⋅), it stores the fuel efficiency of different vehicle types under different driving conditions. When the function receives input data, it searches for an entry in the parameter lookup table that matches the input data and returns the corresponding output result.
[0154] Specifically, in step S602, cost is fed back to the decision model: the cost C is fed back to the driver agent's decision model or the administrator for adjusting delivery strategies, resource allocation strategies, or updating route planning algorithms.
[0155] It's important to note that the decision-making model is either an existing algorithm- or rule-based automated system, or a decision support system operated by an administrator. If certain delivery strategies are found to be causing excessive costs, the decision-making model will propose adjustments, such as changing delivery routes or adjusting delivery times. The model will adjust resource allocation strategies based on cost data, such as increasing or decreasing the number of delivery vehicles or adjusting vehicle load capacity, to reduce overall costs. By adjusting delivery strategies, resource allocation strategies, and route planning algorithms, delivery costs can be effectively reduced.
[0156] Example 2: This example further provides a specific execution method for the Hidden Markov Algorithm in Example 1: S30030, State transition probability matrix By using the state transition probability matrix, the dynamic changes of the hidden state can be modeled, thereby capturing the time-varying characteristics of the system.
[0157] a ij =P(s t =j∣s t−1 =i) represents the probability that the state at time t is j, given that the state at time t−1 is i.
[0158] S30031, Observation probability matrix
[0159] b jk =P(o t =k∣s t =j) represents the probability of observing value k given that the state is j at time t.
[0160] S30032, initial state distribution π=[π1,π2,...,π] N ], where π i =P(s1=1); By using the initial state distribution, the initial conditions of the system can be modeled, thus providing a more comprehensive description of the system's dynamic behavior.
[0161] S30033, Forward Algorithm: The forward algorithm calculates the probability of the observed sequence through recursion, avoiding the complexity of directly calculating all possible paths.
[0162] S30034, Steady-state distribution calculation (calculated iteratively using a forward algorithm): The steady-state distribution reflects the probability distribution of each hidden state of the system after a long period of operation, and is a manifestation of the system's long-term behavioral pattern. The steady-state distribution of the system can be approximated by using a long-term iterative forward algorithm (or combined with other methods such as backward algorithms).
[0163] In this matrix, the state transition matrix A describes the transition probabilities between hidden states, with each row summing to 1; the observation matrix B describes the probability of observing a specific value in a given state, with each row summing to 1; the forward algorithm recursively calculates the probability of the observation sequence; the steady-state distribution π* is the state distribution obtained through a long-term iterative forward algorithm, reflecting the agent's long-term behavioral pattern; the type parameter θ is a continuous spatial variable used to distinguish the heterogeneity of the agent. N is the total number of hidden states, and M is the total number of observations; s t ∈{1,2,...,N} is the hidden state at time t; O t ∈{1,2,...,M} are the observations at time t; α t (θ) is the forward variable, representing the joint probability of state j up to time t; T is the T-distribution.
[0164] Example 3: This example further provides a method for executing step S301, such as... Figure 4 and 8 As shown.
[0165] In step S301, the concept of Nash equilibrium game is introduced to simulate the dynamic game process between "enterprise pricing → freight service → customer choice". This game process takes place in the later stages of the game (e.g., every 100 iterations) and aims to calculate the optimal response function of each agent's strategy. Through this game process, supply chain resilience can be improved in the complex environment of "demand-supply mismatch", and intelligent scheduling management of fresh and perishable products can be achieved.
[0166] (1) Firm-customer game: Strategy Space: Corporate Strategy Set S e And customer strategy set S c ; Nash equilibrium condition: Under an equilibrium strategy, neither party can increase its payoff by unilaterally changing its strategy; ; Where, π e It is the firm's profit function; −θ(d)C loss θ(d) is the loss due to customer churn; θ(d) is the customer churn rate function caused by delivery delay d (the larger d is, the higher θ is); C loss -Cs is the average cost of churning a single customer; -Cs is the short-run disturbance cost (the temporary cost of expedited delivery); δV long It is the discounted value of long-term customer value; δ is the discount factor (0 < δ < 1, reflecting the present value of future earnings); V long It is customer lifetime value (expected returns from long-term cooperation, such as expected returns within one year); π c It is the customer revenue function; β(θ)U service It represents the perceived benefit of service quality; β(θ) is the customer satisfaction coefficient (the lower the θ, the higher the satisfaction); U service It is the basic service utility value; −P switch It is the switching cost (the economic cost of customers switching to competitors); constraint d≤d max This is the delivery delay threshold (exceeding this value increases customer churn rate).
[0167] The logic is as follows: businesses need to weigh short-term disruption costs against long-term customer value to formulate the optimal pricing strategy. Customers, in turn, choose whether to continue cooperating based on perceived benefits from service quality and conversion rates. The Nash equilibrium condition ensures the stability of both parties' strategies, meaning that neither party can increase revenue by unilaterally changing its strategy.
[0168] Understandably, through this game theory process, businesses can gain a more accurate understanding of customers' sensitivities to service quality and price, thereby developing more reasonable pricing strategies. Customers, in turn, can choose the optimal cooperation strategy based on their own needs and perceived service quality, improving satisfaction and loyalty.
[0169] (2) Enterprise-freighter game: Strategy Space: Corporate Strategy Set S f And freight strategy set S t .
[0170] Nash equilibrium condition: Under an equilibrium strategy, neither party can increase its payoff by unilaterally changing its strategy.
[0171] ; Where πf is the firm's profit function; −C detour This refers to detour costs (additional expenses that companies require freight carriers to bear). ρ(ΔF) is the expected benefit of the freight rate adjustment; ρ(ΔF) is the probability function of the freighter's withdrawal caused by the freight rate adjustment amount ΔF; δ t It is a time discount factor (considering long-term cooperation benefits); R k π is the cooperative benefit in period k (such as the profit from improved freight efficiency); r is the discount rate; π is the return on investment in period k. t It is the freight passenger's revenue function; F base This is the basic freight revenue; −μC detour It is the cost burden of detours (μ is the proportion of cost transfer). σ is the exit compensation item; σ is the compensation trigger coefficient (0 < σ ≤ 1). It is an indicator function (triggered when the exit probability exceeds a critical value); K exit This is the exit compensation amount, triggered by the condition ΔF ≥ ΔF. thr : Freight adjustment threshold (the freighter will be deactivated if the freight rate is exceeded).
[0172] The logic is as follows: companies need to weigh the impact of detour costs and freight rate adjustments on freight forwarders' willingness to cooperate in order to formulate the optimal freight rate strategy. Freight forwarders, in turn, choose whether to continue cooperation based on basic freight rate revenue, detour cost burden, and exit compensation. The Nash equilibrium condition ensures the stability of the strategies of both parties.
[0173] Understandably, through this game theory process, companies can gain a more accurate understanding of freight forwarders' sensitivity to freight rates and detour costs, thereby developing more reasonable freight strategies. Freight forwarders, in turn, can choose the optimal cooperation strategy based on their own cost and benefit situations, increasing their willingness and stability in cooperation.
[0174] (3) Customer-freighter game: Strategy Space: Customer Strategy Set Sc And freight strategy set S t .
[0175] Nash equilibrium condition: Under an equilibrium strategy, neither party can increase its payoff by unilaterally changing its strategy.
[0176] ; Where, π c It is the customer revenue function; γU urgent It represents the utility of expedited service; γ is the probability or response parameter for expedited service requests; U urgent It is the utility value brought by expedited service; It is an additional cost of compensation; P λ It is the penalty coefficient for compensation (λ>0 indicates implicit cost premium); P extra This is the basic additional compensation amount; It is the expedited frequency indication function; π t It is the freight passenger's revenue function; P base It is the basic delivery revenue; η⋅E[γ]C hidden It represents the transfer of implicit costs to benefits; η''' is the transfer efficiency coefficient (0 < η ≤ 1); E[γ] is the expected value of the expedited request probability; C hidden These are unit implicit costs, and implicit cost constraints include: implicit cost ceiling constraint C. hidden ≤τ(γ)⋅Profit margin The transfer ratio function (positively correlated with γ) τ(γ) and freight profit margin margin .
[0177] The logic is as follows: customers need to weigh the utility of expedited service against the additional compensation costs to decide whether to request expedited service. Freight forwarders, on the other hand, choose whether to accept expedited service requests based on basic delivery revenue and the benefits of passing on implicit costs. The Nash equilibrium condition ensures the stability of both parties' strategies.
[0178] Understandably, through the game theory process, customers can more accurately understand the impact of expedited services on freight forwarders' costs and benefits, thus making a more rational choice about whether to request expedited services. Freight forwarders, in turn, can choose whether to accept expedited service requests based on their own cost and benefit situations, improving service efficiency and customer satisfaction. The game theory process iteratively calculates the optimal response function of each agent's strategy, gradually approaching the Nash equilibrium state.
[0179] In the three sets of games mentioned above, S c It is a customer strategy set. c s t It is a set of freight strategy sets. t s f It is a corporate strategy set fIt includes all possible combinations of strategies, i.e., combinations of pricing.
[0180] Nash equilibrium condition This indicates that agent unit i is in its corresponding equilibrium strategy s i ∗ Below, other agent units maintain s i ∗ At that time, it is impossible to change the strategy unilaterally. i Increase profits.
[0181] The parameters that need to be further explained in detail in this embodiment include: (1) Average cost of churning a single customer C loss This represents the average economic loss suffered by a company due to customer churn, including marketing costs and potential profit loss. Based on customer value and purchase frequency, it estimates the direct economic loss (such as marketing costs) and potential profit loss (such as future purchase expectations) resulting from the churn of a single customer. The average cost of churning a single customer, C, is obtained by averaging the cost across all churned customers. loss .
[0182] (2) Short-term disruption cost −Cs: This represents the additional costs incurred due to unforeseen events such as short-term delivery delays or expedited deliveries. Identify cost items related to short-term disruptions, such as expedited delivery fees, temporary storage fees, and additional labor costs. Calculate the specific amount of each cost based on the actual occurrence. Summ up all relevant costs to obtain the short-term disruption cost −Cs.
[0183] (3) Discounted long-term customer value δV long This represents the present value of a customer's lifetime value (Vlong) at the present point in time. It is estimated based on the customer's historical purchase data, purchase frequency, and potential growth rate. long The discount factor δ is determined based on the company's cost of capital and return on investment. The customer lifetime value Vlong is then multiplied by the discount factor δ to obtain the discounted long-term customer value δV. long .
[0184] (4) Discount factor δ: Used to discount future income or costs to the present, reflecting the time value of money. The cost of capital r is determined based on the company's financing cost or rate of return on investment. Discount factor calculation: Use the formula δ=1 / (1+r). t Calculate the discount factor, where t is the time span (in years).
[0185] (5) Customer lifetime value V long Customer lifetime value (Vlong) represents the total expected revenue that a customer will generate for the company over a future period of time.
[0186] (6) Customer revenue function π c This represents the net benefit a customer derives from purchasing goods or services. It calculates the direct benefits a customer receives from the goods or services (such as the value of the goods, the utility of the service, etc.), and deducts the costs incurred by the customer (such as the price of the goods, delivery fees, etc.). Subtracting costs from direct benefits yields the customer's net benefit πc.
[0187] (7) Perceived service quality benefit β(θ)U service This represents the additional benefit derived from customers' perceived service quality. Customer perception of service quality β(θ) is assessed through customer surveys, satisfaction ratings, and other methods. The basic service utility value U is determined based on industry standards and historical data. service The perceived service quality β(θ) is compared with the basic service utility value U. service Multiply by this to obtain the perceived service quality benefit β(θ)U service。
[0188] (8) Customer satisfaction coefficient β(θ): This represents the degree of customer satisfaction with the service quality and is a number between 0 and 1. Customer satisfaction data is collected through customer surveys, online reviews, and other methods.
[0189] (9) Basic service utility value U service This represents the utility a customer derives from a service without any additional service quality enhancements. It references industry standards or competitor service utility values. The base service utility value, Uservice, is estimated based on historical customer data and satisfaction ratings.
[0190] (10) Switching cost - P switch This represents the cost or loss a customer incurs when switching from their current supplier to another. Identify cost items related to customer switching, such as the cost of searching for a new supplier, the cost of establishing a new relationship, and potential lost offers or discounts. Calculate the specific amount for each cost based on actual occurrences or customer survey data. Summate all relevant costs to obtain the switching cost – P. switch .
[0191] (11) Enterprise benefit function π f Net income (π) represents the net profit a company earns under a specific strategy. It calculates the total revenue a company receives after selling goods or providing services. This is after deducting costs associated with the sale of goods or provision of services (such as procurement costs, operating costs, and distribution costs). Subtracting total costs from total revenue gives the company's net income (π). f .
[0192] (12) Detour cost - Cdetour: This represents the additional cost incurred by the freight driver in taking an extra detour to complete the delivery task. Calculate the difference between the actual distance traveled by the freight driver and the shortest distance traveled (detour distance). Estimate the additional cost caused by the detour distance based on the freight driver's unit travel cost (such as fuel costs, vehicle depreciation costs, etc.).
[0193] (13) Expected benefits of freight adjustment: This refers to the additional revenue that the company expects to gain by adjusting its freight strategy. Develop different freight adjustment plans (such as increasing freight, decreasing freight, setting tiered freight rates, etc.). Estimate the expected benefits of each freight adjustment plan based on historical sales data and customer price sensitivity analysis.
[0194] (14) Freighter Exit Probability Function ρ(ΔF) Caused by Freight Adjustment ΔF: The freighter exit probability function ρ(ΔF) caused by freight adjustment ΔF represents the impact of freight adjustment on freighters' willingness to exit cooperation. Collect historical freight adjustment data and freighter exit data. Use logistic regression, Probit model statistics, or other statistical models, with freight adjustment ΔF as the input feature and freighter exit probability as the output target, to train the model. Input the current freight adjustment ΔF into the trained model to obtain the corresponding freighter exit probability ρ(ΔF).
[0195] (15) Time discount factor δ t Discount rate (r) is used to discount future earnings to the present, reflecting the company's emphasis on future returns. The discount rate (r) is determined based on the company's cost of capital, return on investment, or risk appetite. The time span (t) for future earnings is determined (in years). The formula δt = 1 / (1 + r) is used. )t Calculate the time discount factor δ t .
[0196] (16) Freighter revenue function π t This represents the net profit earned by the freight forwarder under a specific strategy. It calculates the total revenue the freight forwarder receives by completing delivery tasks (such as base freight, additional subsidies, etc.). It then deducts costs associated with the delivery tasks (such as fuel costs, vehicle maintenance costs, driver salaries, etc.). Subtracting total costs from total revenue yields the freight forwarder's net profit πt.
[0197] (17) Basic freight revenue F base This represents the fixed income a freight forwarder receives for completing a standard delivery task. The base freight rate is determined based on industry practices, market guidance prices, and customer demand. The base freight rate is then multiplied by the weight, volume, or quantity of the goods in the delivery task to obtain the base freight revenue F. base .
[0198] (18) Detour cost burden −μC detourThis represents the portion of the additional costs incurred by the freight forwarder due to detours that is borne by the company. The cost-sharing ratio μ is determined based on the cooperation agreement or market practice between the company and the freight forwarder. As mentioned above, the detour cost C is calculated. detour Multiplying the detour cost Cdetour by the sharing ratio μ, we obtain the detour cost burden borne by the freight passenger, −μC. detour .
[0199] (19) Expedited service utility γU urgent This represents the additional utility a customer gains from receiving expedited service. Customers' perceived value of expedited service is assessed through methods such as customer surveys and satisfaction ratings. The perceived value of expedited service is then compared to the basic service utility value U. service Combining these, we obtain the expedited service utility γU urgent (Usually represented as γ times U) service ).
[0200] (20) Definition of Expedited Service Request Probability or Response Parameter γ: The expedited service request probability or response parameter γ represents the probability of a customer requesting expedited service or the degree of response of the enterprise to an expedited service request. Collect historical expedited service request data and enterprise response data. Use logistic regression, Probit model statistics, or other statistical models, with customer characteristics (purchase history) as input features and expedited service request probability or response parameter γ as the output target, to train the model. Input the current customer characteristics into the trained model to obtain the corresponding expedited service request probability or response parameter γ.
[0201] (21) The utility value U brought by expedited service urgent The additional utility a customer directly gains from receiving expedited service. This is assessed through customer surveys, satisfaction ratings, and other methods to evaluate customers' perceived direct utility of expedited service. Based on the assessment results and industry standards, the utility value (Uurgent) derived from expedited service is determined.
[0202] (22) Additional Compensation Costs: These represent the extra costs a customer incurs for receiving expedited service (such as expedited delivery fees, additional insurance premiums, etc.). Identify the additional cost items related to expedited service. Calculate the specific amount for each additional cost based on actual occurrences or the company's pricing strategy. Summ up all relevant additional costs to obtain the additional compensation costs.
[0203] (23) Urgent Request Frequency Indicator Function: Used to indicate whether the frequency of customer requests for urgent services exceeds a certain threshold. Based on the company's operational strategy and customer demand analysis, determine a reasonable frequency threshold for urgent service requests. Construct an indicator function that outputs 1 when the frequency of customer requests for urgent services exceeds the threshold, and outputs 0 otherwise.
[0204] (24) Implicit cost transfer benefits η⋅E[γ]C hidden This refers to the additional revenue that freight companies gain by passing on some hidden costs to customers. Identify hidden cost items related to the delivery task (such as vehicle wear and tear, driver fatigue, etc.). Estimate the specific amount of each hidden cost and calculate its expected value E[C]. hidden Based on the cooperation agreement or market practice between the company and the freight forwarder, determine the proportion η of hidden cost transfer. The expected value of hidden costs E[C] is then set. hidden Multiplying this by the transfer ratio η yields the implicit cost transfer benefit η⋅E[γ]C hidden .
[0205] (25) Expected probability of expedited requests E[γ]: Represents the average probability of a customer requesting expedited service. Collect historical expedited service request data. Calculate the average probability of all customers requesting expedited service to obtain the expected probability of expedited requests E[γ].
[0206] (26) Unit implicit cost C hidden This represents the average hidden cost incurred by the freight forwarder for each delivery task completed. As mentioned above, identify and estimate the various hidden costs associated with delivery tasks. Divide the total hidden costs by a benchmark quantity such as the number of delivery tasks or the total distance traveled to obtain the unit hidden cost (Chidden).
[0207] (27) Implicit cost ceiling constraint C hidden ≤τ(γ)⋅Profit margin This indicates that the hidden costs passed on to customers by the freight forwarder cannot exceed a certain percentage of its profit margin. The profit margin is determined based on the freight forwarder's financial statements or industry standards. margin Construct a cost-shifting function τ(γ), which represents the relationship between the probability of expedited service requests γ and the proportion of implicit cost shifting (typically expressed as an increasing function). Then, adjust the profit margin... margin Multiplying by the output value of the transfer ratio function τ(γ), we obtain the upper limit constraint C of the implicit cost. hidden ≤τ(γ)⋅Profit margin .
[0208] (28) Shifting ratio function τ(γ): The relationship between the probability of expedited service requests γ and the implicit cost shifting ratio. Collect historical expedited service request data and corresponding implicit cost shifting ratio data. Use a linear regression model, with the probability of expedited service requests γ as the input feature and the implicit cost shifting ratio as the output target, to train the model. Input the current probability of expedited service requests γ into the trained model to obtain the corresponding implicit cost shifting ratio τ(γ).
[0209] (26) Freighter Profit Marginmargin Profit margin represents the percentage of a freight company's net profit after completing a delivery task, relative to its total revenue. This is calculated by collecting the freight company's financial statements, including total revenue, total costs, and net profit. Dividing the net profit by the total revenue yields the freight company's profit margin. margin .
[0210] Example 4: Figure 7 As shown, this embodiment discloses an intelligent scheduling and management system for the supply chain of fresh and perishable products: The system includes a processor and a memory connected to the processor. The memory stores program instructions, which, when executed by the processor, cause the processor to perform the intelligent supply chain scheduling and management method described above. The processor is connected to: (1) Decision modeling module: Deep reinforcement learning algorithm is adopted to efficiently learn the optimal strategy in continuous action space. By constructing a heterogeneous agent model, the agent (agent unit) takes actions (order priority ranking, cancellation threshold setting, detour decision, etc.) based on the current state (real-time inventory, customer lifetime value CLV, delivery time, etc.) and continuously optimizes the strategy through the reward mechanism to maximize long-term benefits.
[0211] (1.1) Data input: Real-time inventory: The current inventory levels of various goods in the warehouse.
[0212] CLV (Customer Lifetime Value): Long-term customer value forecast, used to assess the importance of an order.
[0213] Delivery time: The estimated delivery time of an order, which affects customer satisfaction and order cancellation rate.
[0214] (1.2) Output target: Order Priority: Orders are sorted according to their importance and urgency.
[0215] Cancellation threshold: Set the conditions under which orders can be automatically cancelled to balance inventory and customer demand.
[0216] Detour decision: During the delivery process, based on real-time traffic conditions and order priorities, decide whether to take a detour to optimize delivery efficiency.
[0217] (2) Three-party game simulation module: Using the Vickrey auction mechanism and Nash equilibrium theory, the game process between suppliers, logistics providers and platforms is simulated.
[0218] A Vickrey auction is a sealed-bid second-price auction that incentivizes participants to submit honest bids; a Nash equilibrium, on the other hand, is a strategy combination where unilaterally changing strategies by either party will not yield a better outcome.
[0219] By analyzing historical game data and market supply and demand, we can predict and achieve equilibrium freight rates and service ratings.
[0220] (2.1) Data input: Historical game theory data: past transaction records and quotes between suppliers, logistics providers, and platforms.
[0221] Market supply and demand: The current supply and demand situation of goods and logistics services in the market.
[0222] (2.2) Output target: Equilibrium freight costs: Logistics costs when the market reaches equilibrium.
[0223] (3) Bayesian Optimization Pricing Module: The Bayesian Optimization (BO) algorithm is adopted. This algorithm constructs a probabilistic model of the objective function (such as a Gaussian process) to efficiently search for the optimal solution under uncertainty. Combining competitor pricing and demand elasticity analysis, the commodity price is dynamically adjusted to maximize profit or market share.
[0224] (3.1) Data input: competitor pricing, and the degree of impact of corresponding product price changes on demand.
[0225] (3.2) The output target is a dynamic pricing strategy that adjusts commodity prices in real time.
[0226] (4) Meta-learning Adaptation Module: Utilizing meta-learning algorithms such as MAML (Model-Agnostic Meta-Learning) or Reptile, it can quickly learn and adapt to new tasks on multiple related tasks. By detecting changes in the objective function (sudden changes in market demand), it quickly adjusts policy parameters to adapt to the new market environment.
[0227] (4.1) The data input is the target function change detection signal.
[0228] (4.2) Output objective is to quickly adjust the policy: When the objective function changes, the policy parameters are quickly adjusted to maintain optimal performance.
[0229] (5) Dynamic Feedback Mechanism Module: Combining the rolling window method and transfer learning techniques, the rolling window method is used to process time series data, updating the model by continuously sliding the window; transfer learning utilizes existing knowledge to accelerate the learning process of new tasks. Based on historical interference events and path optimization results, the inventory threshold and path parameters are dynamically adjusted to improve the robustness and adaptability of the system.
[0230] (5.1) The data input is a historical threshold; (5.2) Output target: Inventory threshold: The inventory level is dynamically adjusted based on historical disturbance events and path optimization results.
[0231] Route parameter update: Update delivery route parameters (such as route, speed, etc.) based on real-time traffic conditions and order changes.
[0232] The system disclosed in this embodiment is summarized in the following table:
[0233] All the above embodiments merely illustrate implementation methods for relevant practical applications of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
[0234] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0235] Furthermore, those skilled in the art will understand that implementing all or part of the processes in all the above-described embodiments can be accomplished by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in this application and embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
Claims
1. A method for intelligent scheduling and management of the supply chain for fresh and perishable products, characterized in that, include: S1, obtain inventory turnover rate s t Customer Lifetime Value (CLV), preset historical order priority (p), GPS trajectory data transmitted from delivery time sensors (G), product quality rating (s), order cancellation records (o), driver fatigue monitoring data from vehicle sensors (w), and fuel price API and competitive pricing (c). p ; S2, XGBoost, a pre-trained cancellation threshold classification model for client agents; S3. Suppose the heterogeneous agent model includes three agents: [V1, V2, V3], representing the enterprise, the freight forwarder, and the customer, respectively; execute the Vickrey auction process and detect the optimal response function of each agent's strategy based on Nash equilibrium; S4, Bayesian optimization pricing strategy (pricing combination), objective function: Π=(S×(P−C))−p d Where Π represents corporate profit, and p d It's a shipping subsidy; within the search space, set a price range [P] min ,P max The candidate values are discretized; the price-profit curve is fitted using a Gaussian process; and the next pricing experiment is selected based on the expectation improvement (EI) criterion. S5, when the firm's profit weight increases beyond the threshold θ drift The network parameters of the policy are updated periodically. S6, adjust the inventory threshold IT.
2. The intelligent supply chain scheduling and management method according to claim 1, characterized in that: In S2, the input features include inventory level I and CLV quantiles (CLV). q And historical cancellation rate; tagged as whether the order was cancelled; features include delivery delay minutes and historical complaint count: X=[I, CLV q [o′, Delivery delay minutes, Number of historical complaints] Where: I∈R + The inventory level is represented by o′∈[0,1], the historical cancellation rate is represented by o′∈[0,1], the delivery delay minutes ∈N are the delivery timeout minutes extracted from GPS trajectory data, the number of historical complaints ∈N are the number of past customer complaints, the XGBoost model prediction output is the sum of the outputs of each tree, and the model output is the probability of order cancellation. : ; Where σ(·) is the Sigmoid activation function; f m (⋅) is the m-th regression tree; M is the total number of trees.
3. The intelligent supply chain scheduling and management method according to claim 1, characterized in that: In S3, the Vickrey auction process includes: S3000: Enterprises publish orders and set reserve prices. ; where h t ′ represents the historical freight rate distribution, Q a It is the alpha quantile function of the historical freight rate distribution; k is the risk adjustment coefficient; S3001, Freighter submits sealed quotation B i Including the cost of detours C detour and fatigue compensation C fatigue : ; Among them, C detour (d i ) is the detour cost function for driver i, d i Indicates the detour distance; C fatigue ( hi ): Fatigue compensation function, h i ϵ represents the continuous driving time of driver i; λ is the fatigue compensation weighting coefficient; i It is a random disturbance term; S3002, according to Vickrey's rule, the second-lowest bidder is selected, and the payment price is max(P,C). b ); where C b This is the second lowest price; max(⋅) is the maximum value function.
4. The intelligent supply chain scheduling and management method according to claim 3, characterized in that: In the Vickrey auction process, the optimization problem of the heterogeneous agent model is: assuming there exists a continuous infinite number of agent units V. i The corresponding assets and income are a and y, respectively; each agent unit V i The utility follows the standard expected discounted utility rule, and its future expected discounted utility is calculated using a hidden Markov algorithm to determine the steady-state distribution π* of the agent unit's revenue. ; Where β is the time discount factor; It is the CRRA utility function, a it and y it It represents the assets and income of agent unit Vi at the current moment; ρ is the risk aversion coefficient; π∗(θ) i ) is of type θ i The steady-state income distribution of agents; It is the Hidden Markov state transition probability, representing the agent V. i From state s at time t-1 i,t-1 Transition to state s at time t it The probability of A; i (s it ) is agent V i In state s it The set of feasible actions below; D Θ It is a continuous probability distribution over the agent type space.
5. The intelligent supply chain scheduling and management method according to claim 3 or 4, characterized in that: The Nash equilibrium game includes: calculating the optimal response function of each agent's strategy: Business-customer game: Delivery delays trigger customer churn, and businesses will weigh the cost of disruption against customer value; Business-freighter game: Demanding that freighters bear the cost of detours may lead to them withdrawing from the cooperation or demanding higher freight rates; Customer-freighter game: When a customer requests expedited delivery, they may have to compensate the freight company for extra costs, resulting in cost shifting.
6. The intelligent supply chain scheduling and management method according to claim 5, characterized in that: In the aforementioned enterprise-customer game: ; Where, π e It is the firm's profit function; −θ(d)C loss θ(d) is the loss term caused by customer churn; θ(d) is the customer churn rate function caused by delivery delay d; C loss -Cs is the average cost of churning a single customer; -Cs is the short-run disturbance cost; δV long It is the discounted value of long-term customers; δ is the discount factor; V long It is customer lifetime value; π c It is the customer revenue function; β(θ)U service It represents the perceived benefit of service quality; β(θ) is the customer satisfaction coefficient; U service It is the basic service utility value; −P switch It is the conversion cost; the constraint d≤d max It is the delivery delay threshold; In the aforementioned enterprise-freighter game: ; Where, π f It is the enterprise benefit function; −C detour It's the cost of taking a detour; ρ(ΔF) is the expected benefit of the freight rate adjustment; ρ(ΔF) is the probability function of the freighter's withdrawal caused by the freight rate adjustment amount ΔF; δ t It is the time discount factor; R k π is the cooperative revenue in period k; r is the discount rate; π is the return in period k. t It is the freight passenger's revenue function; F base This is the basic freight revenue; −μC detour It's about the cost of taking a detour; This is the exit compensation item; σ is the compensation trigger coefficient. It is an indicator function; K exit This is the exit compensation amount, triggered by the condition ΔF ≥ ΔF. thr ; In the aforementioned customer-freighter game: ; Among them, γU urgent It represents the utility of expedited service; γ is the probability or response parameter for expedited service requests; U urgent It is the utility value brought by expedited service; It is an additional cost of compensation; P λ It is the penalty coefficient for compensation; P extra This is the basic additional compensation amount; It is the expedited frequency indication function; π t It is the freight passenger's revenue function; P base It is the basic delivery revenue; η⋅E[γ]C hidden It is the transfer of implicit costs to benefits; η''' is the transfer efficiency coefficient; E[γ] is the expected value of the probability of expedited requests; C hidden These are unit implicit costs, and implicit cost constraints include: implicit cost ceiling constraint C. hidden ≤τ(γ)⋅Profit margin The transfer of the proportional function τ(γ) and the freight forwarder's profit margin margin ; Nash equilibrium condition This indicates that agent unit i is in its corresponding equilibrium strategy s i ∗ Below, other agent units maintain s i ∗ At that time, it is impossible to change the strategy unilaterally. i Increase profits; c It is a customer strategy set. c s t It is a set of freight strategy sets. t s f It is a set of corporate strategies, containing all combinations of strategies.
7. The intelligent supply chain scheduling and management method according to claim 5, characterized in that: In step S4, the steps for optimizing the pricing strategy include: S400, in [P] min ,P max Randomly sample initial price points within the range; S401, use GP to fit the price-profit data of the sample points; S402, calculate the EI value for each candidate price; S403, select the price point with the highest EI for the next round of experiments; S404, repeat S401~S403 until convergence or the budget limit is reached; then update the policy space based on the price point corresponding to the EI of the last iteration.
8. The intelligent supply chain scheduling and management method according to claim 5, characterized in that: The S5 process includes the following steps: S500, Calculate the reward function: Define the reward function of the enterprise agent as a function of state s, action a, and the next state s': ; Among them, w i R represents the weight of each reward component. basic Based on the reward, R distance For distance bonus, R direction As a directional reward, R environment As a reward for environmental characteristics, R task Specific rewards for the task; S501, Set profit weight w profit The monitoring mechanism calculates its increase △w profi : ; in, and $ These are the profit weights for the current time and the previous time, respectively; The trigger condition is: when the profit weight increases by more than the threshold θ. drift When this occurs, the objective function drift handling mechanism is triggered; The S502 MAML algorithm starts from the meta-initialization parameter θ' and performs several gradient updates for each task to obtain the task-specific parameter θ. i '; Calculate the loss on the task-specific parameters on the task validation set, and backpropagate it to the meta-initialization parameters to update them: ; Where b is the meta-update step size, (·) represents task T i The loss function.
9. The intelligent supply chain scheduling and management method according to claim 8, characterized in that: S6 includes the following execution steps: S600, if the calculated order fulfillment rate η' is lower than the target threshold, the inventory threshold adjustment mechanism is triggered: the regional warehouse inventory threshold is increased, and the new threshold S... new From the old threshold S old The adjustment amount Δ from the output of the PID controller is calculated. S601, input real-time traffic data and delivery task list to obtain optimized delivery routes; calculate cost C based on optimized delivery routes, including distance, time and fuel consumption factors; S602 feeds back the cost C to the driver agent's decision-making model or administrator for adjusting delivery strategies, resource allocation strategies, or updating route planning algorithms.
10. A supply chain intelligent scheduling and management system for fresh and perishable products, characterized by: The system includes a processor and a memory connected to the processor. The memory stores program instructions, which, when executed by the processor, cause the processor to perform the intelligent supply chain scheduling and management method as described in any one of claims 1-9.