Supply chain dynamic contract method and system based on artificial intelligence
By adopting a dynamic contract method based on artificial intelligence in the supply chain and using deep reinforcement learning algorithms to adjust the contract terms, the problem that traditional contracts are difficult to adjust flexibly when facing market fluctuations and demand uncertainties is solved, and dynamic interest sharing and long-term cooperation among all parties in the supply chain are realized.
Patent Information
- Application Number
- CN202510233999.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
Traditional supply chain contracts are difficult to adjust flexibly when facing market volatility and demand uncertainty, resulting in imbalance in the interests of all parties in the supply chain, and the existing dynamic contracts have shortcomings in the design of the revenue sharing mechanism.
The dynamic contract method of supply chain based on artificial intelligence is adopted, and the contract terms are automatically adjusted according to historical sales data and market demand functions through deep reinforcement learning algorithms to realize dynamic game and interest sharing between supply and demand parties.
It realizes dynamic interest sharing among all parties in the supply chain under changing conditions, can reflect market demand fluctuations in real time, and adjusts the profit distribution ratio at different time nodes, ensuring long-term cooperation and win-win results, and breaking through the limitations of traditional static contracts.
Smart Images

Figure CN120146911A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of supply chain management, and particularly relates to a method and system for a supply chain dynamic contract based on artificial intelligence. Background Art
[0002] With the increasing complexity of global supply chain management, traditional supply chain contracts, such as revenue-sharing contracts and profit-sharing contracts, often show deficiencies in coping with market fluctuations and demand uncertainties. These contracts usually rely on fixed terms and ratios, making it difficult to flexibly respond to the rapidly changing market environment, resulting in an imbalance of interests among the parties in the supply chain in some cases. Especially when facing demand fluctuations or emergencies, the static nature of traditional contracts cannot be quickly adjusted, affecting the efficiency and stability of the supply chain.
[0003] In modern supply chain management, dynamic contracts have become an important means to optimize the supply-demand relationship and improve supply chain efficiency. Traditional supply chain contracts such as wholesale price contracts, buyback contracts, and revenue-sharing contracts can, to a certain extent, balance the interests of both the supply and demand sides, but they often have limitations in coping with market fluctuations, demand uncertainties, and complex environmental changes. With the improvement of information technology and data analysis capabilities, dynamic contracts based on real-time data have gradually been applied to supply chain management, aiming to adapt to market and environmental changes by flexibly adjusting contract terms. The core of a dynamic supply chain contract lies in quickly adjusting contract terms through real-time monitoring and data analysis, optimizing resource allocation and price setting to achieve the optimal benefits of both the supply and demand sides.
[0004] In this context, although existing dynamic contracts have made some breakthroughs in certain aspects, there is still room for further optimization, especially in the design of the mechanism for sharing the interests of all parties in the supply chain. Traditional supply chain contracts often rely on preset terms and price structures, and this static approach is difficult to cope with the complex and ever-changing market environment and dynamic changes inside and outside the supply chain. Therefore, how to design a revenue-sharing mechanism that can be flexibly adjusted according to market conditions has become a key challenge in current supply chain management.
[0005] Currently, there are mainly two categories of methods to solve this problem:
[0006] 1) Revenue-sharing contract
[0007] In a revenue-sharing contract, all parties in the supply chain (such as suppliers and retailers) share the sales revenue according to an agreed-upon distribution ratio. This contract usually ensures that suppliers and retailers jointly bear risks and share benefits based on sales volume by setting a fixed ratio. Revenue-sharing contracts help to encourage cooperation between the two parties and optimize sales and inventory management.
[0008] 2) Profit-sharing contract
[0009] Profit-sharing contracts allocate profits to all parties in the supply chain, usually based on the total profit of production and sales rather than a single sales revenue. This form of contract can better reflect the contribution of all parties in the supply chain in joint production and operations, and encourage cooperation and reduce conflicts through profit distribution.
[0010] Both current optimization methods have obvious disadvantages:
[0011] 1) Income Sharing Contract
[0012] The core of the revenue sharing contract is to distribute sales revenue through a fixed ratio, which may not be flexible enough to deal with market uncertainties and demand fluctuations in some cases. For example, under the influence of drastic demand fluctuations or market emergencies (such as natural disasters, policy changes, etc.), the fixed revenue distribution ratio may lead to unequal benefits among the parties in the supply chain. When demand is low, the income of suppliers and retailers may be compressed at the same time, and the fixed revenue distribution ratio is difficult to adjust in the short term to adapt to such changes. In addition, revenue sharing contracts usually rely on the transparency and accuracy of sales data, but in actual operations, data asymmetry or errors may lead to unfair profit distribution, which in turn affects the trust and cooperation between the two parties.
[0013] 2) Profit Sharing Contract
[0014] Compared with revenue sharing contracts, profit sharing contracts are more flexible in profit distribution because they take into account the costs and revenues of all parties in the supply chain. However, the implementation of this form of contract is difficult, mainly reflected in the complexity of cost accounting and profit calculation. First, the calculation of profits requires accurate tracking of various costs, including costs of raw materials, production, transportation and other links. In actual operation, the allocation of costs may cause disputes due to different recognition of costs by the parties, resulting in unfair profit distribution. Secondly, since profit sharing contracts rely on the common recognition of profits by both parties, market fluctuations, changes in production efficiency and external factors (such as policy and legal changes) may affect the calculation of profits, thereby affecting the stability and fairness of the contract. If there is a lack of flexible adjustment mechanism, profit sharing contracts are also prone to situations where one party bears too much risk and the other party obtains disproportionate profits.
[0015] In order to solve the above technical problems, it is necessary to develop a method and system for supply chain dynamic contract based on artificial intelligence. Summary of the invention
[0016] The object of the present invention is to provide a method and system for a supply chain dynamic contract based on artificial intelligence to solve the above technical problems, and to achieve dynamic game and interest sharing between the supply and demand sides under changing conditions through an intelligent mechanism. This contract can not only reflect real-time market demand fluctuations, but also adjust the revenue sharing ratio of all parties at different time nodes to ensure long-term cooperation and win-win results among all parties in the supply chain. The design of this contract breaks through the limitations of traditional static contracts and provides a more flexible and sustainable solution.
[0017] To achieve the above object of the invention, the technical solutions adopted by the present invention are as follows:
[0018] A method for a supply chain dynamic contract based on artificial intelligence, including a supplier and a retailer, comprising the following steps:
[0019] S100, system input, the system inputs the information of both the supplier and the retailer, and the information includes basic information and historical sales data;
[0020] S200, setting a transaction contract model, the cooperation between the supplier and the retailer is based on the principle of interest sharing between the two parties, and the revenue sharing ratio is dynamically adjusted according to market conditions. The agent automatically adjusts the contract terms through a deep reinforcement learning algorithm based on historical sales data and market demand functions;
[0021] S300, initializing the Q network and hyperparameters;
[0022] S400, defining the state space, action space, and reward function;
[0023] S500, building a neural network and training the network parameters of Double DQN. By continuously training and updating the network, an optimal decision sequence is finally generated;
[0024] S600, obtaining the supplier profit and retailer profit under the contract conditions through the optimal decision sequence of Double DQN, and making them respectively greater than the maximum values under non-contract conditions to obtain the critical value of the revenue sharing factor.
[0025] Preferably, the basic information includes the unit production cost c of the supplier, time T, the supplier's wholesale price range [w min , w max , the wholesale price step size, the retailer's selling price range [p min , p max , and the selling price step size.
[0026] Preferably, the historical sales data includes the selling price and sales volume to calculate the demand elasticity ε, the benchmark price p 0 , and the benchmark demand d 0; Get the demand d faced by retailers in each period t ;
[0027] The following calculations are involved:
[0028] d t =ab·p t ;
[0029] a=(1+ε)·d 0 ;
[0030]
[0031] Preferably, in S300, during the initialization process of the model, the following steps are performed:
[0032] S301, initialize the training network Q ω and the target network Q ω - , both are deep neural networks;
[0033] S302, initializing the experience replay pool Replay_Buffer to store experience samples of each interaction with the environment;
[0034] S303, initialize hyperparameters: learning rate α, used to control the step size of network parameter update; discount factor γ, used to weigh the impact of immediate rewards and future rewards; batch size N: specifies the number of samples randomly drawn from the experience pool in each training.
[0035] Preferably, the S400 includes:
[0036] State space S: the action sequence of the previous period;
[0037] Action space A: It includes the supplier’s wholesale price and the retailer’s selling price;
[0038] Reward function:
[0039] r t =(p t -c)·d t ;
[0040] At the same time, during time T, the benefits of suppliers and retailers under contractual conditions are greater than those under non-contractual conditions.
[0041] Supplier profit under contractual conditions:
[0042]
[0043] Retailer profit under contractual conditions:
[0044]
[0045] Supplier profit under non - contractual conditions:
[0046]
[0047] Retailer profit under non - contractual conditions:
[0048]
[0049] Obtained respectively by backward induction, the optimal wholesale price w*, the optimal selling price p*, the optimal profit of the supplier The optimal profit of the retailer
[0050]
[0051] where: w t is the wholesale price set by the supplier in each period; c is the unit production cost of the supplier; β t is the revenue sharing factor for dynamic contract formulation, and the supplier will enjoy (1 - β t ) * the retailer's sales volume; p t is the selling price set by the retailer in each period; d t is the demand faced by the retailer in each period and is also the order quantity of the retailer in each period.
[0052] The described construction of the neural network includes constructing a deep neural network Q ω (s,a) for estimating the Q - value corresponding to the current state s and action a;
[0053] The network structure includes:
[0054] Input layer: Receiving the encoded features of the current state s and action a;
[0055] Hidden layer: Using a multi - layer fully - connected network with the ReLU activation function;
[0056] Output layer: Outputting the Q - values of all possible actions, that is, the expected value of choosing each action in the current state;
[0057] Target network Q ω - (s,a) has the same structure as Q ω (s,a), and synchronizes parameters from the training network every C steps;
[0058] The training process of the network parameters for training Double DQN is as follows:
[0059] S521, Initialization: Set the state s 0 =(w min+0.5(w max -w min ), p min +0.5(p max -p min ) as the initial state;
[0060] S522, Selection action: Select action a according to the greedy policy t :
[0061]
[0062] S523, Execute action: Execute action a t , reach the next state s t+1 , calculate the reward function r t ; State transition (s t , a t , r t , s t+1 ) and store it in the replay pool Replay_Buffer;
[0063] S524, Sample extraction and training: Randomly extract N samples (s, a, r, s') from the replay pool. For each sample, calculate the target Q value:
[0064]
[0065] Minimize the loss function:
[0066]
[0067] Update the network parameter ω;
[0068] S525, Update the target network: Copy the parameters of Q ω to Q ω - .
[0069] Preferably, the optimal decision sequence is:
[0070] A optimal = {w, p};
[0071] where w, p are matrices containing t.
[0072] Preferably, in S600, the following formula is involved:
[0073]
[0074] where: w t is the wholesale price set by the supplier in each period; c is the unit production cost of the supplier; β tis the revenue sharing factor for dynamic contract formulation; p t is the selling price set by the retailer in each period; d t is the demand faced by the retailer in each period and also the order quantity of the retailer in each period;
[0075] Finally, output the critical value of β that satisfies the above inequality group and form the key content of the supply chain contract. w for each period within time T t and form the key content of the supply chain contract. w, p t p t β t .
[0076] A system adopting any of the above - mentioned methods includes a basic data module, an algorithm data module, an algorithm training module, an optimal decision - making module, and a stored data module.
[0077] This application has achieved beneficial technical effects:
[0078] The present invention realizes the dynamic game and interest sharing between the supply and demand sides under changing conditions through an intelligent mechanism. This contract can not only reflect the market demand fluctuations in real - time but also adjust the revenue distribution ratio of all parties at different time nodes to ensure the long - term cooperation and win - win situation of all parties in the supply chain. The design of this contract breaks through the limitations of traditional static contracts and provides a more flexible and sustainable solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 The following shows the schematic diagram of the algorithm flow of the present invention;
[0080] Figure 2 The following shows the schematic diagram of the system deployed based on the algorithm of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0081] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will describe the specific embodiments of the present invention with reference to the accompanying drawings. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings and other embodiments can be obtained.
[0082] The following will introduce the technical solutions of the present invention in detail with specific embodiments.
[0083] In this solution, a method and system for a supply chain dynamic contract based on artificial intelligence are proposed.
[0084] The goal of dynamic adjustment is to achieve the dynamic game and interest sharing between the supply and demand sides under changing conditions, enabling both sides to obtain higher revenues compared to the non - contract situation. The following are the embodiments of this method in specific supply chain scenarios:
[0085] 1) System input
[0086] The system first inputs the basic information of both sides of the supply chain, specifically including:
[0087] 1.1) Basic information: The unit production cost c, time T, supplier wholesale price range [w min , w max , wholesale price step, retailer selling price range [p min , p max , selling price step.
[0088] 1.2) Historical sales data: Data such as selling price and sales volume can be used to calculate the demand elasticity ε, benchmark price p 0 , benchmark demand d 0 , and this information is adjusted through the update of historical sales data.
[0089] d t = a - b·p t ;
[0090] a = (1 + ε)·d 0 ;
[0091]
[0092] This information is used to calculate the reward.
[0093] 2) Transaction contract model
[0094] In the present invention, a revenue-sharing contract is adopted as the core mechanism of the transaction contract. The cooperation between the supplier and the retailer is based on the principle of mutual benefit sharing, and the revenue-sharing ratio is dynamically adjusted according to market conditions. The agent automatically adjusts the contract terms through a deep reinforcement learning algorithm based on historical sales data and the market demand function to achieve overall profit maximization. Among them, the historical sales data is the known past real data. The market demand function is a function fitted through historical sales data, that is, d t. .
[0095] 3) Initialize the Q network and hyperparameters
[0096] In the initialization process of the model, the following steps are executed:
[0097] ① Initialize the training network Q ω and the target network Q ω - , both of which are deep neural networks.
[0098] ② Initialize the experience replay pool Replay_Buffer to store the experience samples of each interaction with the environment.
[0099] ③ Initialize the hyperparameters: the learning rate α, which is used to control the step size of network parameter updates; the discount factor γ, which is used to balance the impact of immediate rewards and future rewards; the batch size N: specifying the number of samples randomly drawn from the experience pool in each training, etc.
[0100] 4) Define the state space, action space, and reward function
[0101] ① State space S: The action sequence of the previous period.
[0102] ② Action space A: (The wholesale price of the supplier, the selling price of the retailer).
[0103] ③ Reward function: The total supply chain revenue:
[0104] r t =(p t -c)·d t ;
[0105] At the same time, the revenues of the supplier and the retailer under the contract conditions within time T are greater than those under the non - contract conditions;
[0106] The profit of the supplier under the contract conditions:
[0107]
[0108] The profit of the retailer under the contract conditions:
[0109]
[0110] The profit of the supplier under the non - contract conditions:
[0111]
[0112] The profit of the retailer under the non - contract conditions:
[0113]
[0114] Respectively obtained by backward induction, the optimal wholesale price w*, the optimal selling price p*, the optimal profit of the supplier the optimal profit of the retailer
[0115]
[0116]
[0117] Where: w tis the wholesale price set by the supplier in each period; c is the unit production cost of the supplier; β t is the revenue sharing factor for dynamic contract formulation. The supplier will enjoy (1 - β t ) * the retailer's sales volume; p t is the selling price set by the retailer in each period; d t is the demand faced by the retailer in each period and is also the order quantity of the retailer in each period.
[0118] 5) Build a neural network
[0119] Construct a deep neural network Q ω (s, a) to estimate the Q value corresponding to the current state s and action a. The network structure includes:
[0120] ① Input layer: Receive the encoded features of the current state s and action a.
[0121] ② Hidden layer: Use a multi-layer fully connected network with the activation function ReLU (Rectified Linear Unit).
[0122] ③ Output layer: Output the Q values of all possible actions, that is, the expected value of choosing each action in the current state.
[0123] Target network Q ω - (s, a) has the same structure as Q ω (s, a), and synchronize the parameters from the training network every C steps.
[0124] 6) Train the network parameters of Double DQN
[0125] The training process is as follows:
[0126] ① Initialization: Set the state s 0 =(w min +0.5(w max -w min ), p min +0.5(p max -p min )) as the initial state.
[0127] ② Select an action: Select an action a t :
[0128]
[0129] ③ Execute the action: Execute the action a t , reach the next state s t+1 , and calculate the reward r t . State transition (s t , at , r t , s t+1 ) After that, it is stored in the replay pool Replay_Buffer.
[0130] ④ Sample extraction and training: Randomly extract N samples (s, a, r, s′) from the replay pool. For each sample, calculate the target Q value:
[0131]
[0132] Minimize the loss function:
[0133]
[0134] Update the network parameters ω.
[0135] ⑤ Update the target network: Every C steps, copy the parameters of Q ω to Q ω - .
[0136] 7) Output the optimal price decision
[0137] By continuously training and updating the network, finally generate the optimal decision sequence:
[0138] A optimal = {w, p};
[0139] Among them, w and p are matrices containing t, specifically the vectors corresponding to the supplier's wholesale price and the retailer's selling price respectively.
[0140] 8) Solve the optimal revenue sharing factor
[0141] Based on the optimal decision sequence of Double DQN, obtain the supplier's profit and the retailer's profit under the contract conditions, and make them greater than the maximum value under the non - contract conditions:
[0142]
[0143] Finally, output the β t critical value, and form the key content of the supply chain contract. For each period within time T, w t , p t , β t .
[0144] The dynamic contract system based on artificial intelligence of this technical solution includes a basic data module, an algorithm data module, an algorithm training module, an optimal decision module, and a storage data module.
[0145] First, the user inputs basic data, including the unit production cost of the supplier, time, the wholesale price range of the supplier, the wholesale price step, the retail price range of the retailer, and the retail price step. Then, the initial data of the algorithm is defined. After algorithm training, the optimal price decision and the revenue sharing factor are finally output. The storage module will store the trained algorithm parameters of the user, facilitating the user to directly call and retrain. This system can respond in real time to market demand fluctuations and production capacity changes, and adjust the revenue distribution ratio among all parties in the supply chain at different time points, thereby ensuring long-term cooperation relationships and common interests. Through this contract design, the limitations of traditional static contracts are broken through, providing a more flexible and sustainable solution that can adapt to the ever-changing market environment.
[0146] This system is based on artificial intelligence technology and adopts dynamic contract design, aiming to optimize the pricing and revenue distribution of all parties in the supply chain through intelligent algorithms, and improve cooperation efficiency and shared interests. The workflow of the system is divided into several key modules, and each module is closely integrated with modern technologies to achieve more efficient and flexible supply chain management.
[0147] The specific operation process among these modules is as follows:
[0148] Basic data module: The first step of the system is for the user to input basic data, including the unit production cost of the supplier, production time, the wholesale price range of the supplier, the wholesale price step, the retail price range of the retailer, and the retail price step, etc. These data are input through the front-end interface of the system and directly transmitted to the database system for storage through the data interface to ensure the integrity and consistency of the data.
[0149] Algorithm data module: After the data preparation is completed, the system sets the initial algorithm parameters through data processing and analysis. These parameters are used to define the pricing strategy and the calculation method of the revenue sharing factor. This module processes and analyzes the input data in real time through machine learning algorithms such as regression analysis and optimization algorithms to help the system output a preliminary pricing model.
[0150] Algorithm training module: After the algorithm is defined, the system will conduct training, using the input historical data and market simulation scenarios to gradually adjust the algorithm parameters and optimize the pricing decision and revenue distribution strategy. This training process adopts deep reinforcement learning algorithms, enabling the system to dynamically adjust the model parameters in various market scenarios, thereby improving the accuracy and flexibility of the pricing strategy. The training data and model parameters are stored in the cloud, supporting remote calls and multiple iterative trainings.
[0151] Optimal Decision-making Module: The trained algorithm will output the optimal pricing decision and generate an accurate revenue sharing factor. Through real-time calculation and feedback mechanism, this module ensures that the pricing decision can adapt to uncertain factors such as market demand fluctuations and production capacity changes. The system can automatically update the decision according to real-time data to safeguard the maximum interests of all parties in the supply chain.
[0152] Stored Data Module: The storage module adopts cloud database technology to save the trained algorithm model and optimization parameters. Users can conveniently access and call these data through the system background interface to support the retraining and optimization of the model. At the same time, the system also implements data encryption and backup functions to ensure the security and reliability of the data.
[0153] In this technical solution,
[0154] 1) Dynamic Contract Adjustment Mechanism
[0155] Through the reinforcement learning mechanism of deep Q-learning, the system can automatically adjust contract terms such as wholesale price, order quantity, and revenue sharing factor according to historical sales data and market changes, rather than relying solely on fixed static contracts. Specifically, by learning the demand elasticity and price changes in historical sales data, the system optimizes the interest distribution between suppliers and retailers to form a flexible dynamic contract mechanism. The wholesale price and selling price are output through the algorithm. Since the order quantity is a function of the selling price, that is, the market demand function, the automatic adjustment of the order quantity is realized. By introducing demand elasticity, the relevant reward function is affected, and the optimal decision matrices of relevant w and p are obtained, thus affecting the subsequent obtained β. t 。
[0156] 2) Intelligent Optimization of Revenue Sharing Mechanism
[0157] In the traditional revenue sharing model, the interests of suppliers and retailers are often distributed based on a fixed ratio. However, a fixed ratio cannot adapt to the dynamic changes of the market. In the present invention, the intelligent agent dynamically adjusts the revenue sharing factor βt, enabling the interests of retailers and suppliers to be more fairly adjusted in real time according to market fluctuations. This dynamic revenue sharing mechanism not only enhances the mutual trust between the two cooperative parties but also improves the overall adaptability of the supply chain.
[0158] Among them, the contract terms involve not only the revenue sharing factor but also the supplier's wholesale price, the retailer's purchase quantity, and the selling price in the contract terms.
[0159] Through the optimal decision sequence of Double DQN, the supplier's profit and the retailer's profit under the contract conditions are obtained, making them greater than the maximum values under the non-contract conditions, and finally outputting the critical values that satisfy the above inequality group to form the key content of the supply chain contract.
[0160] This technical solution proposes a supply chain dynamic contract method based on artificial intelligence. By using deep reinforcement learning technology, it analyzes market data and historical sales information in real time, and dynamically adjusts the revenue distribution mechanism in the supply chain through intelligent algorithms. Through this method, all parties in the supply chain can flexibly adjust the contract terms under different market conditions to achieve benefit sharing and risk sharing. In the existing technology, suppliers sell products to retailers at a relatively low wholesale price in exchange for a certain percentage of the revenue from the retailers' sales revenue. When the market demand is low and the retailers' sales volume is not high, due to the relatively low wholesale price, the cost pressure on the retailers is relatively small. Although the suppliers' revenue will also decrease, by obtaining a part of the retailers' revenue, they can also make up for the loss of selling products at a low price to a certain extent. On the contrary, when the market demand is strong and the retailers obtain high sales revenue, the suppliers can also obtain more profits by sharing the revenue, and both parties can benefit from the high demand. In this way, the risk brought by market demand fluctuations is shared between suppliers and retailers, avoiding the situation where one party bears all the risks alone.
[0161] The core technical solution of this technical solution enables all parties in the supply chain to optimize the cooperation terms according to real-time market changes through the deep Q-learning algorithm, thereby improving the adaptability and overall profit of the supply chain. Compared with traditional static contracts, this method can enhance the flexibility and response speed of the supply chain and provide guarantee for long-term stable cooperation. This invention breaks through the limitations of traditional contract mechanisms and provides a more efficient and sustainable supply chain management solution with broad application prospects.
[0162] This technical solution proposes a supply chain revenue sharing contract based on dynamic adjustment, aiming to achieve dynamic game and benefit sharing between supply and demand parties under changing conditions through an intelligent mechanism. This contract can not only reflect market demand fluctuations in real time, where the demand function is fitted through historical data, and the historical data can reflect demand fluctuations. It can also adjust the revenue distribution ratio of all parties at different time nodes to ensure long-term cooperation and win-win situation among all parties in the supply chain. The design of this contract breaks through the limitations of traditional static contracts and provides a more flexible and sustainable solution.
[0163] The method protected by this invention aims to solve the above problems. It realizes dynamic game and benefit sharing between supply and demand parties under changing conditions through the deep reinforcement learning method. This contract can not only reflect market demand fluctuations and production capacity changes in real time, but also adjust the revenue distribution ratio of all parties at different time nodes to ensure long-term cooperation and win-win situation among all parties in the supply chain. The design of this contract breaks through the limitations of traditional static contracts and provides a more flexible and sustainable solution.
[0164] The dynamic contract of the model specifically adjusts the profit distribution of the supplier and the retailer by jointly considering the wholesale price, the order quantity of the retailer in each period, and the revenue sharing factor.
[0165] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0166] The above-described embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent should be subject to the appended claims.
[0167] The above has elaborated in detail the embodiments of a method and system for a supply chain dynamic contract based on artificial intelligence provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the core idea of the present invention. It should be noted that for those of ordinary skill in the technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A method for dynamic supply chain contract based on artificial intelligence, including suppliers and retailers, characterized in that: The steps include: S100, system input, the system inputs information of both suppliers and retailers, the information including basic information and historical sales data; S200, sets a transaction contract model, where the cooperation between suppliers and retailers is based on the principle of shared benefits between both parties, and the profit sharing ratio is dynamically adjusted as market conditions change. The intelligent agent automatically adjusts the contract terms based on historical sales data and market demand functions through a deep reinforcement learning algorithm; S300, initialize Q network and hyperparameters; S400, define state space, action space and reward function; S500, builds a neural network and trains the network parameters of Double DQN. By continuously training and updating the network, the optimal decision sequence is finally generated. S600, through the optimal decision sequence of Double DQN, the supplier profit and retailer profit under the contractual conditions are obtained, and they are made greater than the maximum values under the non-contractual conditions, and the critical value of the profit sharing factor is obtained.
2. The method according to claim 1, characterized in that The basic information includes the supplier's unit production cost c, time T, supplier wholesale price range [w min ,w max ], wholesale price step, retailer selling price range [p min ,p max ], sales price step.
3. The method according to claim 1, characterized in that The historical sales data includes sales price and sales volume, so as to calculate demand elasticity ε, base price p0, and base demand d0; and obtain the demand d faced by the retailer in each period. t ; The following calculations are involved: d t =a-b·p t ; a=(1+ε)·d0; 4. The method according to claim 1, characterized in that: In S300, during the initialization process of the model, the following steps are performed: S301, initialize the training network Q ω and the target network Q ω - , both are deep neural networks; S302, initializing the experience replay pool Replay_Buffer to store experience samples of each interaction with the environment; S303, initialize hyperparameters: learning rate α, used to control the step size of network parameter update; Discount factor γ, used to weigh the impact of immediate rewards and future rewards; batch size N: specifies the number of samples randomly drawn from the experience pool in each training.
5. The method according to claim 3, characterized in that: The S400 includes: State space S: the action sequence of the previous period; Action space A: It includes the supplier’s wholesale price and the retailer’s selling price; Reward function: r t =(p t -c)·d t ; At the same time, during time T, the benefits of suppliers and retailers under contractual conditions are greater than those under non-contractual conditions.
6. The method according to claim 5, characterized in that Supplier profit under contractual conditions: Retailer profit under contractual conditions: Supplier profit under non-contractual conditions: Retailer profit under non-contractual conditions: Through the reverse induction method, we can obtain the optimal wholesale price w*, optimal sales price p*, and optimal profit of the supplier under non-contractual conditions. Optimal profit for retailers Where: w t is the wholesale price set by the supplier each period; c is the supplier's unit production cost; β t is the revenue sharing factor of dynamic contract formulation, and the supplier will enjoy (1-β t )*Retailer sales; p t is the selling price set by the retailer each period; t It is the demand faced by retailers in each period, and it is also the order quantity of retailers in each period.
7. The method according to claim 3, characterized in that The construction of the neural network includes constructing a deep neural network Q ω (s,a), used to estimate the Q value corresponding to the current state s and action a; The network structure includes: Input layer: receives the encoded features of the current state s and action a; Hidden layer: A multi-layer fully connected network is used, and the activation function uses ReLU; Output layer: Outputs the Q value of all possible actions, that is, the expected value of selecting each action in the current state; Target Network Q ω - (s,a) has the same ω (s,a) Same structure, synchronizing parameters from the training network every C steps; The training process of the network parameters for training Double DQN is as follows: S521, initialization: set state s0=(w min +0.5(w max -w min ), p min +0.5(p max -p min )) as the initial state; S522, select action: select action a according to the greedy strategy t : S523, Execute action: Execute action a t , reach the next state s t+1 , calculate the reward function r t ;State transfer(s t ,a t ,r t ,s t+1 ), and then store it in the replay pool Replay_Buffer; S524, sample extraction training: randomly extract N samples (s, a, r, s′) from the playback pool, and calculate the target Q value for each sample: Minimize the loss function: Update network parameters ω; S525, update the target network: every C steps, Q ω The parameters are copied to Q ω - .
8. The method according to claim 7, characterized in that The optimal decision sequence is: A optimal ={w,p}; Among them, w and p are matrices containing t.
9. The method according to claim 8, characterized in that In the S600, the following formula is involved: Where: w t is the wholesale price set by the supplier each period; c is the supplier's unit production cost; β t is the revenue sharing factor for dynamic contract formulation; p t is the selling price set by the retailer each period; t It is the demand faced by the retailer in each period, and it is also the order quantity of the retailer in each period; The final output satisfies the above inequality group β t The critical value of w in each period within time T forms the key content of the supply chain contract. t 、p t , β t .
10. A system using the method according to any one of claims 1 to 9, characterized in that: It includes basic data module, algorithm data module, algorithm training module, optimal decision module and storage data module.