Decision-making method for combining inventory control and dynamic pricing based on competitive environment
By calculating benchmark prices and demand, combined with deep reinforcement learning algorithms, real-time adjustment of inventory and pricing, the joint optimization problem of inventory control and dynamic pricing in competitive markets is solved, and higher supply chain efficiency and profitability are achieved.
Patent Information
- Application Number
- CN202510135950.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-06-06
AI Technical Summary
It is difficult for the existing technology to effectively jointly optimize inventory control and dynamic pricing in competitive markets, especially when facing market fluctuations and competitive environments, traditional methods are difficult to respond to changes in real time.
By collecting historical sales data, calculating the benchmark prices and benchmark demands of products, combining deep reinforcement learning algorithms, inventory levels and price strategies are adjusted in real time to cope with changes in market demand and inventory conditions.
It has achieved dynamic adjustment of pricing strategies and inventory levels in competitive markets, optimized overall supply chain efficiency, avoided inventory surplus or shortages, and improved the adaptability and profitability of enterprises in complex markets.
Smart Images

Figure CN120106879A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of joint decision-making, and in particular to a decision-making method for joint inventory control and dynamic pricing based on a competition environment. Background Art
[0002] With the continuous development of the global economy and the changes in market demand, enterprises are facing more and more complex supply chain management issues in their operations, especially in terms of inventory control and pricing strategies. Inventory management and pricing decisions are two key links in modern business operations. They are intertwined and jointly affect the profitability, market share and operational efficiency of enterprises.
[0003] In order to improve supply chain efficiency and reduce inventory costs, many companies try to combine inventory management with pricing strategies and adopt dynamic pricing and joint inventory control methods. Dynamic pricing refers to the real-time adjustment of product prices based on demand, inventory, market competition and other factors during the sales process, while joint inventory control refers to the coordinated optimization of inventory levels and pricing strategies during inventory management and pricing decision-making to maximize overall profits and remain flexible in a competitive environment.
[0004] Although existing studies have proposed some theoretical models for inventory control and pricing, most of these models focus on a single optimization objective and ignore the mutual influence between inventory management and pricing strategies and their synergy in competitive markets. For example, traditional inventory management models such as the economic order quantity (EOQ) model assume that pricing is fixed, but in the actual market, different pricing and price fluctuations for the same product will directly affect consumers' purchasing decisions, and thus affect inventory levels. With the intensification of market competition, companies face more complex challenges in inventory management and pricing decisions, especially in a competitive environment. Traditional inventory control and pricing methods are difficult to effectively respond to market fluctuations and competitive environments. In other words, in a competitive environment, other prior pricing strategies and changes in the same product will affect market demand, which may lead to excess or shortages of inventory, thereby affecting the overall supply chain efficiency. Traditional methods cannot respond to changes in market demand and the competitive environment in real time. Therefore, how to jointly optimize inventory control and dynamic pricing in a competitive market is an urgent problem to be solved.
[0005] The prior art provides a joint strategy optimization method based on reinforcement learning, which realizes dynamic inventory control and pricing decisions through an adaptive learning mechanism. This method regards the optimization of inventory levels and pricing strategies as a joint reinforcement learning problem, and makes optimal decisions in a dynamic and complex market environment by continuously interacting with the environment through reinforcement learning agents. In this method, reinforcement learning models (such as Q-learning, SARSA) take variables such as inventory status, market demand, and price as inputs to generate corresponding pricing strategies and inventory adjustment strategies. Through continuous trial and error and strategy updates, the agent can learn how to flexibly adjust pricing and inventory levels according to market changes to optimize overall profits.
[0006] Although the joint policy optimization method based on reinforcement learning provides a flexible decision-making mechanism in a dynamic environment, it still has certain limitations in practical applications. First, this method usually uses a fixed reward function when modeling the environment. If the market environment changes drastically, the model may not be able to adapt quickly. Secondly, the training process of reinforcement learning usually takes a long time to converge, and it is easy to be challenged in the balance between exploration and utilization, which may lead to slow convergence or fall into the local optimal solution. In addition, the expansion of the scale of state and action space will significantly increase the computational complexity, resulting in low training efficiency, especially in the absence of deep learning support, the expressive power of the algorithm is limited. Finally, the reinforcement learning method has a weak ability to respond to sudden external changes. For instantaneous demand fluctuations or changes in market competition, it may not be able to quickly adjust the strategy to maintain the optimal decision. Summary of the invention
[0007] In view of the problems existing in the prior art, the present invention provides a decision-making method for joint inventory control and dynamic pricing based on a competitive environment, which can perform joint inventory control and dynamic pricing in a competitive environment to maximize the expected total sales profit.
[0008] The technical solution of the present invention is achieved in this way:
[0009] A decision method for joint inventory control and dynamic pricing based on a competitive environment includes the following steps:
[0010] S1. Collect basic information about a target product from n objects; select two of the objects according to the requirements, and record them as the first object and the second object;
[0011] The basic information includes unit ordering cost c, unit holding cost h and historical sales data; the historical sales data is an ordered data sequence within a preset time period, including the sales price of each of the objects and the sales volume of the first object;
[0012] S2. Obtain the maximum inventory level and dynamic pricing range of each object according to the historical sales data; calculate the self-demand elasticity ε of the target product according to the historical sales data 1 and the cross-price elasticity ε 2 , and calculating a base price for the object and a base demand for the first object;
[0013] S3. Build training network Q ω and the target network Q ω - ; Construct an experience replay pool; Initialize hyperparameters; Initialize state space S and action space A; In the state space, the state is represented as
[0014] s t ={(I t,1 ,I t,2 ,…,I t,i ,…,I t,n ),(sp t,1 ,sp t,2 ,…,sp t,i ,…,sp t,n )}, in the action space, the action is represented by a t ={(q t,1 ,q t,2 ,…,q t,i ,…,q t,n ),(bp t,1 ,bp t,2 ,…,bp t,i ,…,bp t,n )};
[0015] Where t represents the period index, I t,i 、sp t,i ,q t,i and bp t,i They represent the initial inventory level, sales price, order quantity, and pricing of the i-th object in the t-th period respectively;
[0016] In deep reinforcement learning, the state space is a set that contains all possible states in the environment. Each state is a unique description of the environment. The action space is a set that includes all actions that can be performed in a specific state. The new state obtained by any state after selecting any action must be a state that exists in the state space.
[0017] s t That is, the state of the tth cycle, a t Represents the action selected in the tth cycle.
[0018] S4, is the current state st Select an action t , and perform the action a t , get the next state s t+1 ; According to the elasticity of own demand ε 1 , the cross price elasticity ε 2 , the benchmark price and the benchmark demand, calculate the demand for each of the objects; using the state s t and the action a t , combined with the demand, the reward value r is calculated t ; Construction state transfer (s t ,a t ,r t ,s t+1 ), and stored in the experience replay pool;
[0019] Specifically, use the greedy strategy to select action a t ,Right now:
[0020]
[0021] Among them, ε represents the preset exploration rate. The above greedy strategy formula obtains a probability value; each action in the action space corresponds to a probability value.
[0022] Calculating the reward value based on demand reflects the consideration of competitive environment factors, that is, how to choose the best strategy for an object. The demand and pricing of other objects are also important influencing factors.
[0023] S5. Select several samples from the experience replay pool, where the samples are recorded as (s, a, r, s'); for each of the samples, use the training network Q ω and the target network Q ω - Calculate the training Q value and the target Q value respectively; update the training network Q based on the training Q value and the target Q value ω ;
[0024] S6, repeat steps S4 and S5 to train the network Q ω training; regularly training the training network Q ω Synchronize to the target network Q ω - ;
[0025] S7, using the target network Q ω - , generate the optimal strategy a for the first object optimal , that is: a optimal =(q optimal ,bpoptimal ), where q optimal is the optimal order quantity, bp optimal For optimal pricing.
[0026] The present invention inputs all historical sales data sequences within a certain period of time T, calculates and obtains the benchmark price and benchmark demand of the commodity based on the historical sales data, effectively extracts the relationship between market demand and price, and provides key input for subsequent inventory control and pricing optimization.
[0027] By introducing behavioral models of different objects and combining deep reinforcement learning algorithms, inventory levels and pricing strategies can be adjusted in real time to cope with pricing changes of other objects and market demand fluctuations. By simulating market games in a competitive environment, the present invention can optimize supply chain efficiency while ensuring its own profitability, avoid market share losses due to pricing imbalances or improper inventory, and ultimately help companies achieve higher adaptability and profitability in complex competitive markets. Repeating steps S1 to S7 at a fixed cycle and dynamically updating the collected data can assist companies in making better decisions. Through accurate demand elasticity estimation, the present invention can dynamically adjust pricing strategies and improve the accuracy of pricing decisions. By adjusting the selection of the first object, the optimal strategies can be output for different objects.
[0028] As a further optimization of the above solution, the base price and the base demand are calculated as follows:
[0029]
[0030] Among them, p base,i represents the base price of the i-th object; d base represents the benchmark demand; T represents the preset time period; p ht,i represents the sales price of the i-th object at time ht in the historical sales data; d ht,1 Represents the sales volume of the first object at time ht in the historical sales data.
[0031] As a further optimization of the above scheme, the reward value r t The calculation is:
[0032]
[0033] Among them, d t,i represents the demand for the i-th object in the t-th period; one object corresponds to one r t,i ; min() means selecting the minimum value.
[0034] The reward value is the expected sales profit.
[0035] As a further optimization of the above scheme, the demand is calculated as follows:
[0036]
[0037] Among them, p base,1 represents the base price of the first object; base,2 represents the base price of the second object; t,i Represents the demand for the i-th object in the t-th cycle.
[0038] As a further optimization of the above solution, the maximum inventory level of the i-th object is I i , the dynamic pricing interval is [p i,min ,p i,max ];
[0039] In step S3, the starting state is also included, which is recorded as
[0040] s 0 ={(I 0,1 ,I 0,2 ,…,I 0,i ,…,I 0,n ),(sp 0,1 ,sp 0,2 ,…,sp 0,i ,…,sp 0,n )};
[0041] Initialization obtains the starting state, namely:
[0042]
[0043] Among them, BASE is the preset multiple.
[0044] As a further optimization of the above solution, the next state s t+1 Recorded as
[0045] s t+1 ={(I t+1,1 ,I t+1 ,2,…,I t+1 ,i,…,I t+1,n ),(sp t+1,1 ,sp t+1 ,2,…,sp t+1 ,i,…,sp t+1,n )};
[0046] The calculation process is:
[0047]
[0048] As a further optimization of the above solution, for each sample, the training Q value is calculated as:
[0049] Q ω,α,β (s,a)=V ω,α (s)+A ω,β (s,a);
[0050] Among them, Q ω,α,β (s,a) is the training Q value of the sample, V ω,α (·) is the state value function, A ω,β (·) is the advantage function.
[0051] α and β are the parameters of the state value function and the advantage function respectively, and ω is the shared parameter of the state value function and the advantage function. The architecture of the training network and the target network is the same, both including input layer, hidden layer and output layer. The input layer is used to receive the encoding features of the current state s and action a; the hidden layer uses a multi-layer fully connected network to extract state features, and the activation function uses ReLU (rectified linear unit) to enhance nonlinear fitting capabilities; the output layer is divided into a state value branch and an advantage branch, and these two parts are combined to calculate the training Q value.
[0052] As a further optimization of the above solution, for each sample, the training Q value is calculated as:
[0053]
[0054] Among them, Q ω,α,β (s,a) is the training Q value of the sample, V ω,α (·) is the state value function, A ω,β (·) is the advantage function; |A| is the size of the action space; ∑ a′ A ω,β (s,a′) means that all actions in the action space are combined with the state s respectively, and the advantage function value is calculated and then accumulated.
[0055] The term is used to decentralize the advantage function and make the training more stable.
[0056] As a further optimization of the above scheme, in step S3, the initialized hyperparameters include γ; in step S5, the target Q value of the sample is calculated, that is:
[0057] y i =r i +γ·max a′ Q ω -(s′,a′);
[0058] Among them, r iis the i-th component of the reward value r of the sample; (y 1 ,y 2 ,…,y i ,…,y n ) is the target Q value of the sample; max a′ Q ω -(s′,a′) means that given a state s′, find an action a′ that makes Q ω The value of -(s′,a′) is the largest.
[0059] The hyperparameter γ is used to weigh the impact of immediate rewards and future rewards.
[0060] As a further optimization of the above scheme, for each of the samples, the training network Q is updated by minimizing the loss function ω The network parameters ω are:
[0061] Compared with the prior art, the present invention achieves the following beneficial effects:
[0062] (1) The present invention calculates and obtains the benchmark price and benchmark demand of commodities based on historical sales data, effectively extracts the relationship between market demand and price, and provides key input for subsequent inventory control and pricing optimization. Through accurate demand elasticity estimation, the present invention can dynamically adjust pricing strategies, improve the accuracy of pricing decisions, and thus optimize the overall supply chain efficiency and profit maximization goals.
[0063] (2) In traditional models, inventory and pricing decisions are usually considered independently, ignoring the dynamic changes in market competition. The present invention introduces behavioral models of different objects and combines deep reinforcement learning algorithms to adjust inventory levels and pricing strategies in real time to cope with pricing changes of other objects and market demand fluctuations. By simulating market games in a competitive environment, the present invention can optimize supply chain efficiency while ensuring its own profitability and avoid market share losses due to pricing imbalances or improper inventory. This method can help companies achieve higher adaptability and profitability in complex competitive markets.
[0064] (3) The present invention combines a dynamic pricing mechanism to adjust product prices and inventory levels in real time to adapt to changes in market demand and inventory status. By adopting a deep reinforcement learning algorithm, the system can flexibly adjust pricing strategies and optimize inventory management based on historical data and real-time feedback, thereby reducing the risk of excess or shortage inventory. This method not only improves the accuracy of pricing, but also enhances the flexibility of inventory management, ultimately improving overall profits and the adaptability of the supply chain. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1It is a flowchart of a decision method for joint inventory control and dynamic pricing based on a competitive environment provided by an embodiment of the present invention;
[0066] Figure 2 It is a schematic diagram of the structure of a deep neural network provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0067] In order to make the purpose, technical solution and advantages of the present invention more clear, the technical solution in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0068] like Figure 1 As shown, this embodiment provides a decision method for joint inventory control and dynamic pricing based on a competitive environment, including the following steps:
[0069] S1. Collect basic information about a target product from n objects; select two objects from them according to the requirements, and record them as the first object and the second object;
[0070] The basic information includes unit ordering cost c, unit holding cost h and historical sales data; the historical sales data is an ordered data sequence combination within a preset time period, including the sales price of each object and the sales volume of the first object;
[0071] S2. Obtain the maximum inventory level and dynamic pricing range of each object based on historical sales data; in this embodiment, the maximum inventory level of the i-th object is recorded as I i , the dynamic pricing interval is recorded as [p i,min ,p i,max ]; Calculate the target product's own demand elasticity ε based on historical sales data 1 and the cross-price elasticity ε 2 , and calculating a base price for the object and a base demand for the first object.
[0072] In this embodiment, the base price and base demand are calculated as follows:
[0073]
[0074] Among them, p base,i represents the base price of the i-th object; d base represents the baseline demand; T represents the preset time period; p ht,i represents the sales price of the ith object at time ht in the historical sales data; d ht,1represents the sales volume of the first object at time ht in the historical sales data. In this embodiment, the base price p of the first object is calculated. base,1 and the base price p of the second object base,2 .
[0075] S3. Build training network Q ω and the target network Q ω - ; Construct an experience replay pool. Initialize hyperparameters, including γ; also include learning rate α, which is used to control the step size of network parameter update. Initialize state space S and action space A; in state space, the state is represented as
[0076] s t ={(I t,1 ,I t,2 ,…,I t,i ,…,I t,n ),(sp t,1 ,sp t,2 ,…,sp t,i ,…,sp t,n )}, in the action space, the action is represented by a t ={(q t,1 ,q t,2 ,…,q t,i ,…,q t,n ),(bp t,1 ,bp t,2 ,…,bp t,i ,…,bp t,n )};
[0077] Where t represents the period index, I t,i 、sp t,i ,q t,i and bp t,i They represent the beginning inventory level, sales price, order quantity, and pricing of the i-th object in the t-th period, respectively.
[0078] In this embodiment, the initialization starting state is recorded as
[0079] s 0 ={(I 0,1 ,I 0,2 ,…,I 0,i ,…,I 0,n ),(sp 0,1 ,sp 0,2 ,…,sp 0,i ,…,sp 0,n )};
[0080]
[0081] Wherein, BASE is a preset multiple. In this embodiment, BASE=0.5, that is,
[0082] s 0 =
[0083] {(0.5I 1 ,0.5I 2 ,…,0.5I i ,…,0.5I n ),(p 1,min +0.5·(p 1,max -p 1,min ),(p 2,min +0.5·(p 2,max -p 2,min ),…,(p i,min +0.5·(p i,max -p i,min ),…,(p n,min +0.5
[0084] (p n,max -p n,min ))}.
[0085] In deep reinforcement learning, the state space is a set that contains all possible states in the environment. Each state is a unique description of the environment. The action space is a set that includes all actions that can be performed in a specific state. The new state obtained by any state after selecting any action must be a state that exists in the state space.
[0086] s t That is, the state of the tth cycle, a t Represents the action selected in the tth cycle.
[0087] S4, is the current state s t Select an action t Specifically, we use the greedy strategy to select action a t ,Right now:
[0088]
[0089] Among them, ε represents the preset exploration rate. The above greedy strategy formula obtains a probability value; each action in the action space corresponds to a probability value.
[0090] Calculating the reward value based on demand reflects the consideration of competitive environment factors, that is, how to choose the best strategy for an object. The demand and pricing of other objects are also important influencing factors.
[0091] Execute action a t , get the next state st+1 ; According to own demand elasticity ε 1 , cross price elasticity ε 2 , base price and base demand, calculate the demand for each object; use state s t and action a t , combined with the demand, the reward value r is calculated t .
[0092] Specifically, the demand is calculated as:
[0093]
[0094] Among them, d t,i Represents the demand for the i-th object in the t-th cycle.
[0095] Use the reward function to calculate the reward value r t ,Right now:
[0096]
[0097] Among them, d t,i represents the demand for the i-th object in the t-th period; one object corresponds to one r t,i ;min() means to select the minimum value. The reward value is the expected sales profit.
[0098] Specifically, the next state s t+1 Recorded as
[0099] s t+1 ={(I t+1,1 ,I t+1,2 ,…,I t+1,i ,…,I t+1,n ),(sp t+1,1 ,sp t+1,2 ,…,sp t+1,i ,…,sp t+1,n )};
[0100] The calculation process is:
[0101]
[0102] Construction state transfer (s t ,a t ,r t ,s t+1 ) and stored in the experience replay pool.
[0103] S5. Select several samples from the experience replay pool, and record them as (s, a, r, s'); for each sample, use the training network Q ω and the target network Q ω- Calculate the training Q value and target Q value respectively.
[0104] In this embodiment, for each sample, the training Q value is calculated as:
[0105]
[0106] Among them, Q ω,α,β (s,a) is the training Q value of the sample, V ω,α (·) is the state value function, A ω,β (·) is the advantage function; |A| is the size of the action space; ∑ a′ A ω,β (s,a′) means that all actions in the action space are combined with the state s respectively, and the advantage function value is calculated and then accumulated.
[0107] α and β are the parameters of the state value function and advantage function respectively, and ω is the shared parameter of the state value function and advantage function. The architecture of the training network and the target network is the same, both are deep neural networks, such as Figure 2 As shown in the figure, they all include input layer, hidden layer and output layer. The input layer is used to receive the encoding features of the current state s and action a; the hidden layer uses a multi-layer fully connected network to extract state features, and the activation function uses ReLU (rectified linear unit) to enhance nonlinear fitting capabilities; the output layer is divided into a state value branch and an advantage branch, and these two parts are combined to calculate the training Q value. The term is used to decentralize the advantage function and make the training more stable.
[0108] In this embodiment, the calculation of the target Q value of the sample is:
[0109] y i =r i +γ·max a 'Q ω -(s′,a′);
[0110] Among them, r i is the i-th component of the sample reward value r; (y 1 ,y 2 ,…,y i ,…,y n ) is the target Q value of the sample; max a 'Q ω -(s′,a′) means that given a state s′, find an action a′ that makes Q ω The value of -(s′,a′) is the largest. The hyperparameter γ is used to weigh the impact of immediate rewards and future rewards.
[0111] Combine the training Q value and the target Q value to update the training network Q ω, that is, updating the training network Q by minimizing the loss function ω The network parameters ω are:
[0112] S6, repeat steps S4 and S5 to train the network Q ω until the model converges; regularly train the network Q ω Synchronize to target network Q ω - ; For example, if the training network is updated C times (10 times), it is synchronized to the target network.
[0113] S7. Utilize the target network Q ω - , generate the optimal strategy a for the first object optimal , that is: a optimal =(q optimal ,bp optimal ), where q optimal is the optimal order quantity, bp optimal For optimal pricing.
[0114] Repeating steps S1 to S7 according to a fixed cycle, such as one month as a model training cycle, and dynamically updating the collected data can assist enterprises in making better decisions. Adjusting and changing the selection of the first object according to needs can generate the optimal strategy for all objects.
[0115] According to the disclosure and teaching of the above description, those skilled in the art to which the present invention belongs may also make changes and modifications to the above embodiments. Therefore, the present invention is not limited to the specific embodiments disclosed and described above, and some modifications and changes to the present invention should also fall within the scope of protection of the claims of the present invention. In addition, although some specific terms are used in this specification, these terms are only for the convenience of description and do not constitute any limitation to the present invention.
Claims
1. A decision-making method for joint inventory control and dynamic pricing based on a competitive environment, characterized in that: The following steps are involved: S1. Collect basic information about a target product from n objects; select two of the objects according to the requirements, and record them as the first object and the second object; The basic information includes unit ordering cost c, unit holding cost h and historical sales data; the historical sales data is an ordered data sequence within a preset time period, including the sales price of each of the objects and the sales volume of the first object; S2, obtaining the maximum inventory level and dynamic pricing range of each object according to the historical sales data; calculating the own demand elasticity ε1 and cross-price elasticity ε2 of the target product according to the historical sales data, and calculating the benchmark price for the object and the benchmark demand for the first object; S3. Build training network Q ω and the target network Q ω - ; Construct an experience replay pool; Initialize hyperparameters; Initialize state space S and action space A; In the state space, the state is represented by s t ={(I t,1 ,I t,2 ,…,I t,i ,…,I t,n ),(sp t,1 ,sp t,2 ,…,sp t,i ,…,sp t,n )}, in the action space, the action is represented by a t ={(q t,1 ,q t,2 ,…,q t,i ,…,q t,n ),(bp t,1 ,bp t,2 ,…,bp t,i ,…,bp t,n )}; Where t represents the period index, I t,i 、sp t,i ,q t,i and bp t,i They represent the initial inventory level, sales price, order quantity, and pricing of the i-th object in the t-th period respectively; S4, is the current state s t Select an action t , and perform the action a t , get the next state s t+1 ; Calculate the demand quantity of each object according to the own demand elasticity ε1, the cross price elasticity ε2, the benchmark price and the benchmark demand; Utilize the state s t and the action a t , combined with the demand, the reward value r is calculated t ; Construction state transfer (s t ,a t ,r t ,s t+1 ), and stored in the experience replay pool; S5. Select several samples from the experience replay pool, where the samples are recorded as (s, a, r, s'); for each of the samples, use the training network Q ω and the target network Q ω - Calculate the training Q value and the target Q value respectively; update the training network Q based on the training Q value and the target Q value ω ; S6, repeat steps S4 and S5 to train the network Q ω training; regularly training the training network Q ω Synchronize to the target network Q ω - ; S7, using the target network Q ω - , generate the optimal strategy a for the first object optimal ,Right now: a optimal =(q optimal ,bp optimal ), where q optimal is the optimal order quantity, bp optimal For optimal pricing.
2. The decision method for joint inventory control and dynamic pricing based on a competitive environment according to claim 1, characterized in that: The base price and base demand are calculated as follows: Among them, p base,i represents the base price of the i-th object; d base represents the benchmark demand; T represents the preset time period; p ht,i represents the sales price of the i-th object at time ht in the historical sales data; d ht,1 Represents the sales volume of the first object at time ht in the historical sales data.
3. The decision method for joint inventory control and dynamic pricing based on a competitive environment according to claim 2, characterized in that: The reward value r t The calculation is: Among them, d t,i represents the demand for the i-th object in the t-th period; one object corresponds to one r t,i ; min() means selecting the minimum value.
4. The decision method for joint inventory control and dynamic pricing based on a competitive environment according to claim 2, characterized in that: The demand is calculated as: Among them, p base,1 represents the base price of the first object; base,2 represents the base price of the second object; t,i Represents the demand for the i-th object in the t-th cycle.
5. The decision-making method for joint inventory control and dynamic pricing based on a competitive environment according to claim 1, characterized in that: The maximum inventory level of the i-th object is I i , the dynamic pricing interval is [p i,min ,p i,max ]; In step S3, the starting state is also included, which is recorded as s0={(I 0,1 ,I 0,2 ,…,I 0,i ,…,I 0,n ),(sp 0,1 ,sp 0,2 ,…,sp 0,i ,…,sp 0,n )}; Initialization obtains the starting state, namely: Among them, BASE is the preset multiple.
6. A decision method for joint inventory control and dynamic pricing based on a competitive environment according to claim 3 or 4, characterized in that: Next state s t+1 Recorded as s t+1 ={(I t+1,1 ,I t+1 ,2,…,I t+1 ,i,…,I t+1,n ),(sp t+1,1 ,sp t+1 ,2,…,sp t+1 ,i,…,sp t+1,n )}; The calculation process is:
7. The decision method for joint inventory control and dynamic pricing based on a competitive environment according to claim 1, characterized in that: For each of the samples, the training Q value is calculated as: Q ω,α,β (s,a)=V ω,α (s)+A ω,β (s,a); Among them, Q ω,α,β (s,a) is the training Q value of the sample, V ω,α (·) is the state value function, A ω,β (·) is the advantage function.
8. The decision method for joint inventory control and dynamic pricing based on a competitive environment according to claim 1, characterized in that: For each of the samples, the training Q value is calculated as: Among them, Q ω,α,β (s,a) is the training Q value of the sample, V ω,α (·) is the state value function, A ω,β (·) is the advantage function; |A| is the size of the action space; ∑ a′ A ω,β (s,a′) means that all actions in the action space are combined with the state s respectively, and the advantage function value is calculated and then accumulated.
9. A decision method for joint inventory control and dynamic pricing based on a competitive environment according to claim 7 or 8, characterized in that: In step S3, the initialized hyperparameters include γ; in step S5, the target Q value of the sample is calculated, that is: y i =r i +γ·max a′ Q ω -(s′,a′); Among them, r i is the i-th component of the reward value r of the sample; (y1,y2,…,y i ,…,y n ) is the target Q value of the sample; max a′ Q ω -(s′,a′) means that given a state s′, find an action a′ that makes Q ω The value of -(s′,a′) is the largest.
10. A decision-making method for joint inventory control and dynamic pricing based on a competitive environment according to claim 9, characterized in that: For each of the samples, update the training network Q by minimizing the loss function ω The network parameters ω are: