Reinforcement learning dynamic pricing method based on game feedback

By introducing a data freshness assessment mechanism and a Stackelberg game framework into the online data market, and combining multi-armed slot machines and reinforcement learning models, the pricing strategy is dynamically adjusted to solve the pricing lag problem under changing market conditions, and to achieve real-time response and stability optimization of the market strategy.

CN121685018APending Publication Date: 2026-03-17NORTHEASTERN UNIV CHINA

Patent Information

Application Number
CN202511837452.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing pricing methods in the online data market are unable to dynamically respond to sudden changes in the market environment and network conditions, and lack adaptability to the timeliness of real-time data, resulting in decision-making delays and bottlenecks in strategy optimization.

Method used

By introducing a data freshness assessment mechanism and combining Stackelberg game theory and multi-armed slot machine exploration—utilizing the balance principle—a reinforcement learning-based dynamic pricing method is constructed by dynamically adjusting data weights and pricing strategies through a reinforcement learning model, ensuring high real-time performance and strong relevance of the data.

Benefits of technology

It achieves real-time response and stability of pricing strategies in the online data market, avoids decision lag caused by outdated data, optimizes returns and strategy stability in non-stationary environments, and adapts to market competition with multiple participants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121685018A_ABST
    Figure CN121685018A_ABST
Patent Text Reader

Abstract

The invention provides a reinforcement learning dynamic pricing method based on game feedback, and relates to the technical field of online data market dynamic pricing. Through combination of leader-follower strategy interaction of the Stackelberg game and the exploration-utilization balance principle of the dobby machine, the problems of price adjustment lag, large income fluctuation and insufficient strategy stability in a non-stable environment are solved. By introducing a data freshness evaluation mechanism, data weight is dynamically adjusted by integrating data timeliness attenuation and market correlation analysis, it is ensured that data input into a model always has high real-time performance and strong correlation, and the problem of decision lag caused by static data input in a comparison scheme is avoided; the exploration-utilization balance principle and the reinforcement learning model of the dobby machine are fused, the exploration rate is dynamically adjusted, a historical optimal strategy can be fully utilized in a non-stationary environment, a potential better scheme can be explored, and the static optimization bottleneck that a game model is limited to a preset objective function in a comparison scheme is broken through.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dynamic pricing technology in online data markets, and in particular to a reinforcement learning-based dynamic pricing method based on game feedback. Background Technology

[0002] As the online data market expands, user demands become more diverse, and competitors accelerate their strategy iterations, traditional static pricing models are no longer adequate. There is an urgent need for dynamic pricing mechanisms that integrate multiple technologies to achieve rapid responses to real-time market feedback, while also ensuring revenue stability and strategy adaptability. This technology has become one of the core directions for optimizing online business operations.

[0003] Traditional online data market dynamic pricing technologies primarily rely on two implementation paths. One is a static rule-based pricing method. This method pre-sets fixed price adjustment thresholds (e.g., reducing prices by a fixed percentage when inventory falls below a certain level, or increasing prices by a fixed percentage when user traffic increases by more than a certain percentage), and combines this with historical sales data to formulate standardized pricing rules. It doesn't require real-time analysis of the interactions between multiple market participants; price adjustments are triggered simply through conditional judgments. A typical application scenario is the early promotional pricing of e-commerce platforms. The other is a pricing method based on a single machine learning model. This method often uses regression analysis, traditional neural networks, and other models. It uses static data such as historical revenue and user purchase records as input to train a pricing prediction model. The model's output determines the price range. For example, some retail platforms use price prediction models trained based on historical sales data from the same period. However, these models can only infer pricing based on historical patterns and cannot dynamically capture the impact of sudden market changes (such as sudden competition or shifts in user demand) and changes in data timeliness on pricing.

[0004] Chinese patent CN202011273448.2 discloses a power control and interference pricing method and apparatus for heterogeneous renewable energy networks. The method includes: constructing a system model of the heterogeneous renewable energy network; constructing state models of the Main Power Controller (MBS) and Subsequent Power Controller (SBS) in the heterogeneous network; constructing cost objective functions for the MBS and SBS; wherein the MBS objective is to find the optimal interference pricing strategy, and the SBS objective is to control its own transmission power according to the interference price provided by the MBS to minimize costs; constructing a dynamic game model with the MBS as the leader and the SBS as the follower, and solving the Nash equilibrium solutions under open-loop and feedback modes of the dynamic game model to obtain the optimal power allocation and optimal interference pricing strategy. This invention can optimize power resource allocation and effectively manage power resources to reduce interference in heterogeneous networks.

[0005] The above-mentioned inventions focus on power control and interference pricing in heterogeneous renewable energy networks. They rely on the open-loop and feedback Nash equilibrium solution of Stackelberg dynamic game. However, the strategy optimization of the game model is highly dependent on the preset objective function and fixed parameters (such as weight constants and discount factors), lacking adaptability to real-time dynamic data and unable to dynamically respond to sudden changes in market environment or network status. The above solutions do not establish a data timeliness assessment mechanism and make decisions based only on static inputs such as fixed channel gain and energy storage status, ignoring the impact of data freshness on pricing and power control strategies, which may lead to suboptimal decisions driven by outdated data. Summary of the Invention

[0006] The technical problem this invention aims to solve is to address the shortcomings of the existing technologies by providing a reinforcement learning-based dynamic pricing method based on game feedback. This method introduces a data freshness assessment mechanism and dynamically adjusts data weights by integrating data timeliness decay and market relevance analysis. This ensures that the data input to the model always possesses high real-time performance and strong relevance, avoiding the decision-making lag problem caused by static data input in comparative schemes. Secondly, it integrates the exploration-utilization principle of multi-armed slot machines with reinforcement learning models. By dynamically adjusting the exploration rate, it can fully utilize historical optimal strategies and explore potential better solutions in non-stationary environments, overcoming the bottleneck of static optimization limited to a preset objective function in comparative schemes.

[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0008] A reinforcement learning-based dynamic pricing method based on game-theoretic feedback, suitable for online data market environments with time dynamics and data timeliness, includes the following steps:

[0009] Step 1: Real-time market data collection and freshness assessment. Through data interfaces deployed in an online data marketplace, multi-source dynamic data related to market conditions are collected in real time, including user behavior data, competitor pricing data, supply and demand data, and market environment data. Simultaneously, a timeliness assessment module is set up for each type of data collected. Using timestamps and data update frequency monitoring, a data freshness index is calculated. This index reflects the relevance weight of the data to the current market state. If the data timeliness is below a preset threshold, the weight of the corresponding data in the subsequent pricing model is reduced, or redundant data is removed to ensure the real-time nature and effectiveness of the model input.

[0010] Step 2: Construct a Stackelberg game framework and define the roles of market participants. In the online data market, based on Stackelberg game theory, the platform or core pricing entity is defined as the "Leader," and other participating merchants, service providers, or competitors are defined as "Followers." The Leader actively guides market equilibrium through dynamic pricing strategies, while followers adjust their own strategies based on the Leader's pricing decisions. Specifically, the Leader sets an initial price decision space based on real-time collected market data and freshness index, and transmits it to the followers as the initial signal for the game. Followers, based on the received signal and combined with their own costs, inventory, and local market data, generate a strategy response to the Leader, thus forming the basis for the strategic interaction of the dynamic game.

[0011] Step 3: Within the Stackelberg game framework, the leader adopts the "exploration-exploitation problem" approach to model the price adjustment strategy as a dynamic equilibrium problem. Based on feedback from market data, the leader balances the exploration of new price strategies with the exploitation of known high-yield strategies, thereby optimizing the price adjustment decision.

[0012] Step 4: Based on game feedback and real-time price adjustments using the reinforcement learning model, the leader combines the exploration results of the multi-armed slot machine in Step 3 with the game framework in Step 2 to construct a reinforcement learning (RL) model and perform training and optimization. The reinforcement learning model specifically includes:

[0013] Define the state space: Integrate real-time market data and data freshness index to form a dimension of Multidimensional state vector ,in The number of market data indicators and the number of data source types are both determined by the quantity of market data indicators. Each dimension corresponds to a standardized market characteristic or freshness index, ensuring that the model can perceive the correlation between market status and data quality in real time.

[0014] Defining the action space: Based on high-yield pricing strategies explored in multi-armed slot machines, the top 20% of historically profitable pricing schemes are selected as candidate actions. Each action... It includes two core parameters: price adjustment range. and pricing model identifier ;

[0015] Define the reward function: Taking both profit maximization and strategy stability as dual objectives, the reward function comprehensively considers both short-term returns and long-term strategy stability; reward function The calculation formula is:

[0016] ;

[0017] in, These are weighting coefficients, based on data freshness. Dynamic adjustment The higher The larger the value, the more emphasis is placed on short-term gains; The lower The smaller the value, the more emphasis is placed on strategy stability; For the first Actual revenue during the pricing cycle, in yuan; For the first Price adjustment range of wheels This represents the largest price adjustment in history, expressed in percent.

[0018] The model is trained using deep reinforcement learning algorithms to learn the optimal pricing strategy in non-stationary environments. The model updates its strategy network by continuously receiving market feedback and dynamically adjusts its learning rate and exploration rate based on data freshness to cope with sudden market changes.

[0019] In each pricing cycle, the model generates the optimal price decision based on the current state and pushes it to market participants in real time through the API interface, while recording decision and feedback data for subsequent model iterations;

[0020] Step 5: Dynamic verification and updating of strategy stability and return optimization. After the price decision generated in Step 4 is executed, the leader continuously monitors market feedback data and return performance; specifically including:

[0021] Step 5.1: Establish a strategy stability assessment module: Use statistical methods to determine whether the current price strategy deviates from the historical stable range; In order to strengthen the quantitative basis of the strategy stability assessment module, a sliding window weighted variance calculation model based on data freshness is introduced to accurately capture the volatility characteristics of the price series.

[0022] Step 5.2: Establish a revenue optimization evaluation module: Compare the revenue difference between the current strategy and the historical strategy. If the revenue decline exceeds a preset threshold, a strategy update mechanism is triggered. In order to establish a revenue optimization evaluation module, a dynamic revenue benchmark model based on data freshness is introduced to accurately compare the performance difference between the current strategy and the historical strategy.

[0023] Step 5.3: Feed the stability and reward evaluation results back to the reinforcement learning model as new training samples to update the state-action mapping relationship; In order to deeply integrate the stability and reward evaluation results into the update mechanism of the reinforcement learning model, design a policy optimization objective function based on evaluation feedback to dynamically correct the state-action value mapping relationship.

[0024] Step 5.4: Based on the updated model parameters, recalculate the price decision, generate the final price decision, and adjust the exploration rate in conjunction with data freshness;

[0025] Step 5.5: Perform A / B testing on the updated strategy and the original strategy to ensure that the new strategy can improve returns and maintain user confidence in price changes in a non-stationary environment.

[0026] Furthermore, the specific method of step 1 is as follows:

[0027] First, define the timeliness decay function for each data source. ,in It is the difference between the last updated time of the data and the current time, in hours. Its purpose is to reflect the impact of time on data timeliness; in order to make the timeliness decay non-linearly over time, the following formula is used to represent the exponential decay of data timeliness:

[0028] ;

[0029] in, It is the decay factor, which controls the rate of decay; a larger one... This means that the timeliness of data decays faster; This indicates the timeliness of data updates; the closer the value is to 1, the fresher the data, and the closer the value is to 0, the more outdated the data.

[0030] Next, for each data source, a freshness index F is introduced. i To comprehensively evaluate the timeliness and relevance of the data to the current market situation; for each data source i, its freshness index F i The calculation formula is as follows:

[0031] ;

[0032] in, This represents the time-sensitivity decay value of data i. It is the weight coefficient of data source i; It is the market relevance index of data source i, which reflects the strength of the relationship between the data source and current market decisions; The freshness index, derived from historical data analysis, represents the contribution of the data source to market decisions, ranging from 0 to 1. A higher value indicates a more relevant and timely data source, and vice versa;

[0033] Calculate the freshness index for each data source. After determining the value, adjust the weight of each data source in the pricing model; if the freshness index of some data sources... If the value is lower than a preset threshold, the weight of the corresponding data source in the model will be reduced; in extreme cases, i.e., the freshness index F of the data source is reduced... i If the data is far below the preset threshold and has almost no effective reference value for current market pricing decisions, the corresponding data source will be completely removed.

[0034] Furthermore, the specific method for step 2 is as follows:

[0035] The leader sets the initial price decision space based on real-time collected market data and freshness index; the leader's initial price decision is represented as a price strategy space. Each price point represents a potential pricing decision; based on known market conditions and freshness index. The leader determines the initial price range and passes it on to the followers through game theory;

[0036] In the game, followers adjust their pricing behavior according to the leader's pricing strategy. The follower's price adjustment decisions depend on two key factors: the leader's pricing decisions and their own costs, inventory levels, and local market data. A follower reaction function is introduced. The reaction function represents how followers, given a leader's pricing, adjust their pricing optimally based on their own circumstances.

[0037] ;

[0038] in, This indicates the leader's pricing decisions. It is the pricing decision of the follower. It is the production cost of the follower. It is the inventory of the followers; It is the follower's payoff function, reflecting the follower's profit at a given price and cost; in the reaction function... In the middle, followers attempt to adjust To maximize its profits.

[0039] Furthermore, the specific method for step 3 is as follows:

[0040] The leader views potential pricing strategies as the "arms" of a multi-armed slot machine, each arm corresponding to a possible pricing scheme. The choice of these pricing strategies directly affects market response, thus requiring dynamic adjustments to maximize total revenue. To evaluate the effectiveness of each pricing strategy, the leader uses a market data freshness index. and historical earnings data The expected revenue for each arm is calculated using the UCB (Upper Confidence Bound) algorithm, where the revenue estimate for each arm is updated using the following formula:

[0041] ;

[0042] ;

[0043] in, For the arm In time Average return at that time For the arm In the Feedback benefits from rounds of trials For the arm The number of times it was selected For the arm The upper confidence interval reflects the maximum potential gain that the arm may obtain in future rounds;

[0044] After calculating the UCB value for each arm, the leader adjusts the exploration probability based on these values, i.e., how to choose between known high-yield strategies and new strategies. When market data is fresh and the market state is stable, the leader reduces the exploration probability, relying more on known high-yield strategies. Conversely, when the market state changes abruptly or the data is not fresh, the leader increases the exploration probability, increasing attempts at new pricing strategies. This adjustment is achieved through dynamically changing exploration rates. To achieve this, specifically expressed as:

[0045] ;

[0046] in, and To adjust the parameters, As a freshness index for market data, For the threshold; when When the exploration rate is close to or greater than the threshold, Lower values ​​mean more reliance on known strategies; when When the exploration rate is less than the threshold, The higher level encourages further exploration.

[0047] Finally, the leader uses the multi-armed slot machine's reward feedback mechanism to update the profit estimates of each arm in real time and adjust the priority of subsequent price decisions; reward feedback Key metrics derived from user purchasing behavior, conversion rates, and inventory depletion rates are presented in the following formats:

[0048] ;

[0049] in, Indicates the frequency of purchase behavior. Indicates conversion rate. Indicates the rate at which inventory is consumed; , , These correspond to the weighting coefficients of the three key indicators—purchase frequency, conversion rate, and inventory consumption rate—in the reward feedback calculation, and are used to adjust the contribution of different indicators to the reward feedback Rᵢ(t).

[0050] Furthermore, in step 4, the model is trained using a deep reinforcement learning algorithm as follows:

[0051] enter:

[0052] State space: This includes real-time collected market data, specifically the current supply-demand ratio, competitor pricing strategies, inventory status, user profile characteristics, and data freshness index; the state space is represented by a multi-dimensional vector. Each of them This represents a specific market factor;

[0053] Action Space: In step 3, the leader explores high-yield pricing strategies through MAB, including price adjustment range and pricing patterns; the action space is represented as... ,in This indicates a pricing scheme;

[0054] Reward Function: The reward function comprehensively considers both short-term returns and long-term strategy stability; in each round of pricing decisions, the reward function is defined as:

[0055] ;

[0056] in, It's an immediate benefit. It is a long-term effect; and Corresponding to short-term returns R short-term and long-term effects R long-term In the reward function r t The weighting coefficients in the calculation are used to adjust the contribution ratio of short-term returns and long-term strategy stability to the rewards;

[0057] Output:

[0058] Optimal pricing strategy: Obtain the optimal pricing strategy under a given market environment through reinforcement learning model training. ;

[0059] Policy Network: The trained policy network Able to adjust according to real-time market conditions Make the optimal pricing decision;

[0060] The specific training process is as follows:

[0061] Step 4.1: Initialize the policy network and value network ,in and These are the parameters for the policy network and the value network; initializing the experience replay pool. , marks the beginning of the empty set;

[0062] Step 4.2: Determine the initial state This refers to the initial state of the market environment, including the current supply and demand relationship, competitor pricing, inventory status, user profile characteristics, and data freshness index; at this point, the timestamps of all market data are marked, and the data freshness index is initially calculated.

[0063] Step 4.3: Select Action: The agent selects an action based on the current state. Select a pricing action from the action space. In the initial stage, an exploration strategy is used to explore and randomly select actions; as training progresses, the probability of utilizing existing strategies is gradually increased.

[0064] Step 4.4: Market Feedback: After implementing the pricing decision, the environment provides feedback based on the agent's actions, returning immediate rewards. and the new market situation Reward function Calculated based on user behavior and changes in competitors' strategies;

[0065] Step 4.5: Update the experience replay pool: Store the current experience in the experience replay pool. In the middle; samples are randomly selected from the playback pool through batch sampling. To train the model;

[0066] Step 4.6: Strategy Update:

[0067] Update the policy network using Q-learning or policy gradient methods;

[0068] If Q-learning is used, the value function is updated using the following formula. :

[0069] ;

[0070] in, It's the learning rate. It is a discount factor. Indicates selecting the next state. The Q value maximized at time;

[0071] If the policy gradient method is used, the policy network parameters are updated using the following formula. :

[0072] ;

[0073] in, It's the learning rate. Indicates the expected cumulative reward; This represents taking the partial derivative with respect to the policy network parameters θ;

[0074] Step 4.7: Adjust the exploration rate: based on the data freshness index And the stability of the market environment, dynamically adjust the exploration rate When market data is new and stable, the exploration rate is reduced to utilize known strategies; when market conditions change abruptly, the exploration rate is increased to explore new pricing strategies.

[0075] Step 4.8: Evaluation and Adjustment: Validate the effectiveness of the pricing strategy by evaluating its performance. If the strategy does not meet expectations, adjust the model's learning rate, exploration rate, discount factor, and decay factor in the data timeliness decay function to continue optimizing the training process.

[0076] Step 4.9: Iterative Update: Repeat the above steps to iterate the training until the model converges; the policy network during the training process will gradually tend towards the optimal pricing strategy.

[0077] Step 4.10: Final Output of the Optimal Pricing Strategy: After training, use the optimized policy network. The final pricing decision is generated; at this point, the model makes the optimal pricing decision based on real-time market data to maximize profits and maintain the stability of the market strategy.

[0078] Further, step 5.1 specifically includes:

[0079] Define the sliding window size as W, corresponding to the most recent W pricing cycles. For each time point t, the price sequence within the window is: ;in, This represents the price decision made by the leader in round i.

[0080] To incorporate the impact of data timeliness, the data freshness index F from step 1 is used. i As a weight, the weighted average price within the window is calculated first:

[0081] ;

[0082] Then, the sliding window weighted variance is derived:

[0083] ;

[0084] This formula is passed through F i The contribution of prices in each period is dynamically adjusted to ensure that high-freshness data dominates volatility assessment, thereby accurately reflecting the degree of price dispersion under real-time market conditions; if the V(t) value is high, it indicates that the price strategy fluctuates sharply in the recent window and may deviate from the stable range.

[0085] To further quantify the volatility threshold, an adaptive threshold function is designed:

[0086] ;

[0087] in, , represents the historical long-term variance mean, and L is the historical window size. Sensitivity adjustment parameter; adaptive threshold function pass It captures the lowest data freshness within a window, and automatically increases the threshold tolerance when data timeliness decreases to avoid misjudgments caused by outdated data; the stability judgment condition is... If true, it indicates that the current strategy has deviated from the historical stable range, triggering the subsequent strategy update mechanism.

[0088] Furthermore, step 5.2 specifically includes:

[0089] Define the historical window size as L, corresponding to the past L pricing cycles. For each cycle time t, the historical return benchmark is... Calculated as a freshness-weighted average:

[0090] ;

[0091] in, It is the first The data freshness index of the round, It is the first The actual benefit of the wheel; this formula is obtained through The contribution of revenue in each period is dynamically adjusted to ensure that the benchmark more accurately reflects the revenue level during periods of high real-time data, and to avoid distortion of the benchmark by outdated data.

[0092] Next, define the profit difference ratio. To quantify the relative performance of the current strategy:

[0093] ;

[0094] Earnings Difference Ratio Directly measure current returns relative to historical return benchmark The degree of deviation, with negative values ​​indicating a decrease in returns;

[0095] To enhance the robustness of the triggering mechanism, a cumulative deterioration metric is introduced:

[0096] ;

[0097] in, For a fixed threshold, when And three consecutive rounds When this occurs, the strategy update mechanism is triggered; cumulative deterioration indicators... By measuring the extent to which the accumulated negative deviation exceeds a threshold, the severity of persistent revenue deterioration is captured, ensuring that updates are only initiated when there is a significant and continuous decline in performance.

[0098] Furthermore, step 5.3 specifically includes:

[0099] Define the comprehensive strategy health index:

[0100] ;

[0101] in, It is the sliding window weighted variance described in step 5.1. It is the profit difference ratio defined in step 5.2. and For balancing weighting coefficients;

[0102] Comprehensive Strategy Health Index By combining price volatility and the degree of return decline, the overall health status of a quantitative strategy is determined. Its value ranges from 0 to 1, with a lower value indicating a more serious problem with the strategy.

[0103] Then, an enhanced advantage function is constructed:

[0104] ;

[0105] in, It is the standard dominance function. This function represents the policy gradient; it sets the policy health metric. As a modulating factor for advantage estimation, when health is low, it enhances the contribution of the policy gradient term, forcing the model to update the state-action mapping more aggressively to correct bias.

[0106] Finally, this mechanism is embedded into the objective function of the proximal policy optimization algorithm (PPO algorithm) to form the optimization objective:

[0107] ;

[0108] in, The probability ratio is used to dynamically adjust the policy update magnitude through the health index, ensuring that the model prioritizes adjusting the mapping weights of problematic actions when stability or returns deteriorate, thereby achieving efficient transmission of feedback information to policy parameters.

[0109] Furthermore, step 5.4 specifically includes:

[0110] When recalculating pricing decisions, a comprehensive data freshness index is first constructed to quantify the overall data timeliness; a weighted average freshness is defined as follows:

[0111] ;

[0112] in, It is the importance weight of data source i, and the weight in step 1. echo, This is the current data freshness index, where N is the total number of data sources. The weighted average freshness index reflects the overall real-time level of multi-source data through weighted aggregation; the higher the value, the more reliable the input data.

[0113] To further capture the internal variation of data freshness, we define freshness dispersion:

[0114] ;

[0115] This value measures the degree to which the freshness of each data source deviates from the mean. High dispersion indicates inconsistency in the data state, which may introduce uncertainty in decision-making.

[0116] Based on the above indicators, the exploration rate is adjusted as follows:

[0117] ;

[0118] in, The base exploration rate is derived from the initial settings in step 3; and As a positive adjustment parameter, it controls the intensity of the influence of the freshness mean and dispersion; when At higher levels, the exponential term negatively dominates, thus reducing the exploration rate and prioritizing known high-yield strategies; when When the index is high, the positive drive of the index increases the exploration rate and helps to mitigate the risk of market changes caused by data inconsistency.

[0119] Subsequently, using For action selection: explore new policies by randomly and uniformly sampling from the action space with this probability; otherwise, select from the updated policy network. Select the optimal action This generates the final price decision.

[0120] Furthermore, step 5.5 specifically includes:

[0121] To scientifically verify the comprehensive effectiveness of the updated strategy in a non-stationary market environment, a dynamic A / B testing framework based on Bayesian decision theory is constructed; a comprehensive performance index for the strategy is defined:

[0122] ;

[0123] in, and These represent the average returns of the new strategy and the original strategy within the test window, respectively. The standard deviation of price volatility for the new strategy. The historical stable range serves as the volatility benchmark, derived from the sliding window variance analysis in step 5.1; This is a trade-off coefficient between returns and stability; the comprehensive performance index of this strategy quantifies both the magnitude of return improvement and volatility changes, with a positive value indicating that the new strategy maintains or improves stability while increasing returns.

[0124] To further test the significance of the results, a weighted likelihood ratio statistic was designed:

[0125] ;

[0126] in, and These correspond to the old and new strategies in the state. The following generates revenue The probability density, The weighted average freshness index defined in step 5.4; this weighted likelihood ratio statistic weights the likelihood ratio at each time point by the data freshness, ensuring that high-real-time data has a higher weight in decision-making and effectively dealing with distribution drift in non-stationary environments; if simultaneously satisfying and ,in and If a strict threshold is set, the new strategy is deemed to have passed verification and can be officially deployed.

[0127] The beneficial effects of adopting the above technical solution are as follows: The reinforcement learning dynamic pricing method based on game feedback provided by this invention, through the construction of a data freshness evaluation mechanism, achieves accurate screening and dynamic weight adjustment of multi-source market data, ensuring that the data input into the pricing model always has high real-time performance and strong correlation, avoiding the price adjustment lag problem caused by outdated data, and making pricing decisions more in line with current market dynamics. Utilizing the Stackelberg game framework, the roles of market players are clearly defined, and a dynamic strategy interaction mechanism is constructed. The platform or core pricing entity is defined as the "leader," and other merchants and service providers are defined as "followers," allowing the leader to actively guide market equilibrium while verifying the effectiveness of its own strategy through follower feedback, forming a two-way interactive pricing optimization model, which is more adaptable to market competition environments with multiple participants compared to existing technologies. By integrating the exploration of multi-armed slot machines (MAB) – utilizing the balance principle and reinforcement learning models – a balance between intelligent decision optimization and stability is achieved, allowing the model to continuously learn the optimal pricing strategy in non-stationary environments while avoiding frequent price fluctuations, achieving the dual goals of revenue and stability. Attached Figure Description

[0128] Figure 1 A dynamic pricing market model diagram provided for embodiments of the present invention;

[0129] Figure 2 A flowchart illustrating the overall dynamic pricing process provided in this embodiment of the invention;

[0130] Figure 3 A flowchart illustrating the data freshness assessment logic provided in this embodiment of the invention;

[0131] Figure 4 This is a flowchart of the reinforcement learning model training process provided in an embodiment of the present invention;

[0132] Figure 5 A closed-loop flowchart for strategy verification and update provided in embodiments of the present invention. Detailed Implementation

[0133] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0134] A reinforcement learning-based dynamic pricing method based on game feedback, applicable to, for example... Figure 1The online data market environment shown here, characterized by time dynamism and data timeliness, aims to address the technical challenges of lagging price adjustments, large payout fluctuations, and insufficient strategy stability in non-stationary environments by combining the leader-follower strategy interaction of Stackelberg games with the exploration of Multi-Armed Bandits (MABs) using equilibrium principles. This will enable adaptive price responses to real-time market feedback and data freshness, maintaining optimal payouts and strategy stability under dynamically changing market conditions. Figure 2 As shown, the method of this embodiment is specifically described below.

[0135] Step 1: Real-time Market Data Collection and Freshness Assessment. Through a data interface deployed in an online data marketplace, multi-source dynamic data related to the market status is collected in real time, including user behavior data, competitor pricing data, supply and demand data, and market environment data. Simultaneously, a timeliness assessment module is set up for each type of data collected. Using timestamps and data update frequency monitoring, a data freshness index is calculated. This index reflects the relevance weight of the data to the current market status. If the data's timeliness is below a preset threshold (e.g., more than 24 hours without updating), its weight in the subsequent pricing model is reduced, or redundant data is removed to ensure the real-time nature and effectiveness of the model input. Figure 3 As shown in the figure, this is a flowchart of the data freshness assessment logic provided in this embodiment, which clearly presents the complete process of data timeliness decay calculation, freshness index determination and data weight adjustment.

[0136] First, define the timeliness decay function for each data source. ,in It is the difference between the last updated time of the data and the current time, in hours. Its function is to reflect the impact of time on data timeliness. To ensure that timeliness decays non-linearly over time, the following formula is used to represent the exponential decay of data timeliness:

[0137] ;

[0138] in, It is the decay factor, which controls the rate of decay. A larger value indicates a lower decay rate. This means that the data's timeliness decays more rapidly. This function's output... This indicates the timeliness of data updates; the closer the value is to 1, the fresher the data, and the closer the value is to 0, the outdated the data.

[0139] Next, a freshness index is introduced for each data source. To comprehensively evaluate the timeliness of the data and its relevance to the current market situation. For each data source Its freshness index The calculation formula is as follows:

[0140] ;

[0141] in, Representing data The time-dependent decay value, It is a data source The weighting coefficients, It is a data source The market relevance index reflects the strength of the relationship between the data source and current market decisions. This can be obtained through historical data analysis and represents the contribution of the data source to market decisions, ranging from 0 to 1. Freshness Index A higher value indicates a more relevant and timely data source, and vice versa.

[0142] Calculate the value of each data source After setting the values, adjust the weights of each data source in the pricing model. If some data sources... If the value is lower than a preset threshold (e.g., 0.5), the weight of the corresponding data source in the model will be reduced, and in extreme cases, the freshness index F of the data source may even be reduced. i If the data is significantly below a preset threshold and has virtually no effective reference value for current market pricing decisions, the corresponding data source will be completely removed. This ensures that the data input into the pricing model is always of high quality and its real-time nature is guaranteed. This dynamic adjustment mechanism effectively prevents price decisions from being influenced by outdated data, especially in a rapidly changing market environment.

[0143] This process not only improves the responsiveness of the pricing model, but also ensures that price adjustments can reflect the latest market changes in a timely manner through timeliness assessment.

[0144] Step 2: Construct a Stackelberg game framework and define the roles of market participants. In the online data market, based on Stackelberg game theory, the platform or core pricing entity is defined as the "Leader," and other participating merchants, service providers, or competitors are defined as "Followers." The Leader actively guides market equilibrium through dynamic pricing strategies, while followers adjust their own strategies based on the Leader's pricing decisions. Specifically, the Leader sets an initial price decision space based on real-time collected market data and freshness index, and transmits this as the initial signal to the followers. Followers, based on the received signal and combined with their own costs, inventory, and local market data, generate a strategic response to the Leader, thus forming the basis for the strategic interaction of the dynamic game. Figure 1The dynamic pricing market model diagram shown can intuitively illustrate the strategic interaction between leaders and followers in the online data market.

[0145] Specifically, leaders set an initial pricing decision space based on real-time market data and freshness index. This decision space considers various factors, such as market supply and demand, competitor pricing, and user behavior. Therefore, the leader's initial pricing decision is represented as a pricing strategy space. Each price point represents a potential pricing decision. This is based on known market conditions and freshness index. The leader determines the initial price range and passes it on to the followers through game theory.

[0146] In the game, followers adjust their pricing behavior according to the leader's pricing strategy. Specifically, the follower's price adjustment decision depends on two key factors: first, the follower considers the leader's pricing decision; second, the follower also reacts based on its own costs, inventory levels, and local market data. To quantify this strategic response, a follower reaction function is introduced. This represents how followers, given a leader's pricing, adjust their pricing optimally based on their own circumstances. The reaction function is expressed as:

[0147] ;

[0148] in, This indicates the leader's pricing decisions. It is the pricing decision of the follower. It is the production cost of the follower. It is the inventory of the followers. This is the follower's payoff function, reflecting the follower's profit at a given price and cost. In this function, the follower attempts to adjust... To maximize their profits. Follower profits are influenced by a variety of factors, including the leader's price, the follower's own costs, inventory levels, and demand elasticity.

[0149] The purpose of the above follower reaction function is to clarify how followers respond to the leader's pricing decisions in the market, thus forming a complete game theory model. In this model, the leader's goal is to set prices... This guides the market toward an optimal equilibrium, while followers adjust their prices accordingly. To maximize profits, while taking into account the leader's strategies and the dynamic changes in the market.

[0150] Step 3: Within the Stackelberg game framework, the leader employs an "exploration-exploitation problem" approach, modeling the price adjustment strategy as a dynamic equilibrium problem. Based on market data feedback, the leader balances the exploration of new price strategies with the exploitation of known high-yield strategies, thereby optimizing the price adjustment decision. Specific steps include:

[0151] S3.1: Treat potential price adjustment strategies (such as different price points, discount coefficients) as arms in MAB, with each arm corresponding to a possible pricing scheme;

[0152] S3.2: Calculate the expected returns for each arm based on the freshness index of current market data and historical revenue data;

[0153] S3.3: Dynamically allocate the exploration probability (ExplorationRate), reduce the exploration probability when the data is fresh and the market is stable, and increase the exploration probability when the data is not fresh or the market changes abruptly, in order to explore new price strategy combinations;

[0154] S3.4: Update the revenue estimates of each arm in real time through the MAB reward feedback mechanism and adjust the priority of subsequent price decisions.

[0155] The above steps, through the exploratory capabilities of MAB, prevent leaders from getting trapped in local optima due to over-reliance on historical data. At the same time, through follower feedback in the game theory framework, the effectiveness of the exploratory strategy is verified and subsequent decisions are optimized.

[0156] In practice, leaders view potential pricing strategies (such as different price points and discount coefficients) as "arms" in a multi-armed slot machine, with each arm corresponding to a possible pricing scheme. The choice of these pricing strategies directly impacts market response, thus requiring dynamic adjustments to maximize total revenue. To evaluate the effectiveness of each pricing strategy, leaders use a market data freshness index... and historical earnings data To calculate the expected revenue for each arm, the UCB (Upper Confidence Bound) algorithm is used, where the revenue estimate for each arm is updated using the following formula:

[0157] ;

[0158] ;

[0159] in, For the arm In time Average return at that time For the arm In the Feedback benefits from rounds of trials For the arm The number of times it was selected For the arm The upper confidence interval reflects the maximum potential gain that the arm may obtain in future rounds. The core of the UCB algorithm is to maximize the upper limit of future gains by balancing exploration and exploitation.

[0160] After calculating the UCB value for each arm, the leader adjusts the exploration probability based on these values, i.e., how to choose between known high-yield strategies and new strategies. This depends on the freshness of the market data. When market conditions are high and stable, leaders tend to reduce exploration probability and rely more on known high-yield strategies. Conversely, when market conditions change abruptly or data freshness is low, leaders increase exploration probability, experimenting with new pricing strategies. This adjustment can be achieved through dynamically changing exploration rates. To achieve this, specifically expressed as:

[0161] ;

[0162] in, and To adjust the parameters, As a freshness index for market data, The threshold value is used. When the probability of exploration is close to or greater than the threshold Lower values ​​mean more reliance on known strategies; when When the probability of exploration is low A higher exploration rate encourages more exploration. By dynamically allocating exploration rates in this way, leaders can respond flexibly to market changes and avoid strategy lag caused by over-reliance on historical data.

[0163] Finally, the leader continuously updates the revenue estimates for each arm and adjusts the priority of subsequent pricing decisions through the MAB's reward feedback mechanism. Key metrics derived from user purchasing behavior, conversion rates, and inventory depletion rates are presented in the following formats:

[0164] ;

[0165] in, Indicates the frequency of purchase behavior. Indicates conversion rate. This indicates the rate at which inventory is consumed. , , These correspond to the weighting coefficients of the three key indicators—purchase frequency, conversion rate, and inventory consumption rate—in the reward feedback calculation, and are used to adjust the contribution of different indicators to the reward feedback Rᵢ(t).

[0166] Step 4: Based on game feedback and real-time price adjustments using the reinforcement learning model, the leader combines the exploration results of MAB in Step 3 with the game framework of Step 2 to construct a reinforcement learning (RL) model and perform training and optimization. The reinforcement learning model specifically includes:

[0167] Define the State Space: Integrate real-time market data (such as current supply and demand ratio, competitor price distribution, and user profile characteristics) with data freshness index to form a multi-dimensional state vector;

[0168] Define the Action Space: High-yield pricing strategies explored in MAB are used as candidate actions, which include price adjustment range, pricing models (such as fixed discounts, dynamic premiums), etc.

[0169] Define a reward function: taking both profit maximization and strategy stability as dual objectives, the reward function comprehensively considers short-term returns and long-term strategy stability.

[0170] The model is trained using deep reinforcement learning algorithms (such as DQN and PPO) to learn the optimal pricing strategy in non-stationary environments. The model updates its policy network by continuously receiving market feedback (such as user purchasing behavior and competitor strategy adjustments) and dynamically adjusts the learning rate and exploration rate based on data freshness to cope with sudden market changes.

[0171] Within each pricing cycle, the model generates the optimal price decision based on the current state and pushes it to market participants (such as merchants and users) in real time via API interface, while recording decision and feedback data for subsequent model iterations.

[0172] When building a reinforcement learning model, leaders begin with the precise definition of the state space. This space needs to comprehensively integrate two types of core information: one is real-time market data, covering indicators that directly reflect market dynamics, such as the current supply-demand ratio, competitor price distribution, and user profile characteristics; the other is the freshness index of each data source calculated in step 1. After normalizing these information, they are concatenated to form a dimension of Multidimensional state vector ,in The number of market data indicators and the number of data source types are jointly determined. Each dimension corresponds to a standardized market characteristic or freshness index, ensuring the model can perceive the correlation between market status and data quality in real time. The action space is defined strictly based on the high-yield pricing strategies obtained from the MAB exploration in step 3, selecting the top 20% of historically profitable pricing schemes as candidate actions. Each action... It includes two core parameters: price adjustment range. (Values ​​range from -15% to 15%, with a negative sign indicating a price decrease and a positive sign indicating a price increase) and pricing model identifier. The size of the action space is dynamically adjusted based on the number of MAB exploration rounds. The candidate action set is updated every 10 exploration rounds to ensure that the action space always includes high-potential pricing strategies under the current market conditions. The reward function design must simultaneously satisfy the dual objectives of maximizing returns and ensuring strategy stability. Therefore, a reward function is constructed. The calculation formula is:

[0173] ;

[0174] in Weighting coefficients (based on data freshness) Dynamic adjustment The higher The larger the value, the more emphasis is placed on short-term gains; The lower The smaller the value, the more emphasis is placed on strategy stability. For the first Actual revenue per pricing cycle (in yuan). For the first Price adjustment range of wheels This represents the historical maximum price adjustment (in percentage). The core purpose of this formula is to balance short-term gains with strategy continuity in reward calculation. When the price adjustment fluctuation is small, the second term of the formula takes a value close to 1, even if short-term gains are low. Even with a slight decrease, a high reward value can be maintained, preventing the model from experiencing frequent price fluctuations due to excessive pursuit of short-term gains. When data freshness is high and market conditions are stable, increasing γ allows the model to focus more on... To fully capture profit opportunities.

[0175] During the model training phase, the PPO algorithm is used. This algorithm addresses the step size issue in policy updates by constructing a clipped objective function, adapting to the needs of non-stationary market environments. During training, the model receives market feedback data after each pricing cycle, including transaction volume generated by user purchasing behavior and new pricing after competitors' strategy adjustments. This feedback data is then integrated with the state vector. ,action ,award Together, they form the training samples and are stored in the experience replay pool, while also being processed according to the freshness index in step one. Dynamically adjust the learning rate Adjust the formula to ,in The base learning rate (with a value of 0.001) is used when... When reduced (data timeliness decreases). Synchronous reduction slows down model updates and prevents policy shifts caused by outdated data; when When the data is elevated (data real-time performance is enhanced). To improve and accelerate the model's learning of new market information, the exploration rate adjustment still follows the method in step 3. The calculation method ensures that the exploration behavior is consistent with the quality of market data and the state of game interaction. The number of model training iterations is set to a complete parameter update every 50 rounds of pricing cycle. After each update, the strategy performance is evaluated through the validation set. If the performance improvement exceeds 5%, the current model parameters are saved; otherwise, the previous optimal parameters are backtracked. Through this training mechanism, the model can continuously learn the optimal price strategy in a non-stationary environment, providing a reliable model foundation for the dynamic verification and updating of subsequent strategy stability and profit optimization.

[0176] like Figure 4 As shown, the specific training process for the model in step 4 is as follows:

[0177] enter:

[0178] State Space: This includes real-time collected market data, specifically the current supply-demand ratio, competitor pricing strategies, inventory status, user profile characteristics, and data freshness index. The state space is represented by a multi-dimensional vector. Each of them It represents a specific market factor.

[0179] Action Space: In step 3, the leader explores high-yield pricing strategies through MAB, including price adjustment ranges and pricing patterns. The action space is represented as... ,in This indicates a pricing scheme.

[0180] Reward Function: The reward function comprehensively considers both short-term returns and long-term strategy stability. In each round of pricing decisions, the reward function is defined as: ,in These are immediate benefits (such as user purchase conversion rate, inventory consumption rate, etc.). These are long-term effects (such as brand value, market share growth, etc.). and Corresponding to short-term returns R short-term and long-term effects R long-term In the reward function r t The weighting coefficients used in the calculation are used to adjust the contribution ratio of short-term gains and long-term policy stability to the reward. This reward function shares the same core objective as the reward function defined in the reinforcement learning model, both aiming to balance short-term gains and long-term policy stability. This reward function is a concrete quantitative expression of this dual objective within the model training process of the defined reward function.

[0181] Output:

[0182] Optimal pricing strategy: Obtain the optimal pricing strategy under a given market environment through reinforcement learning model training. .

[0183] Policy Network: The trained policy network Able to adjust according to real-time market conditions Make the optimal pricing decision.

[0184] The specific algorithm flow is as follows:

[0185] Step 4.1: Initialize the policy network and value network ,in and These are the parameters for the policy and the value network, respectively. Initialize the experience replay pool. , which marks the beginning of an empty set.

[0186] Step 4.2: Determine the initial state This refers to the initial state of the market environment, including current supply and demand, competitor pricing, inventory levels, user profile characteristics, and data freshness index. At this point, the timestamps of all market data are marked, and the data freshness index is initially calculated.

[0187] Step 4.3: Select Action: The agent selects an action based on the current state. Select a pricing action from the action space. In the initial stage, an exploration strategy is used to explore and randomly select actions; as training progresses, the probability of utilizing existing strategies is gradually increased.

[0188] Step 4.4: Market Feedback: After implementing the pricing decision, the environment provides feedback based on the agent's actions, returning immediate rewards. and the new market situation Reward function Calculated based on user behavior (purchases, browsing, conversion rates, etc.) and changes in competitor strategies.

[0189] Step 4.5: Update the experience replay pool: Store the current experience in the experience replay pool. In the middle, samples are randomly selected from the playback pool using a batch sampling method. To train the model.

[0190] Step 4.6: Strategy Update:

[0191] Update the policy network using Q-learning or policy gradient.

[0192] If Q-learning is used, the value function is updated using the following formula. :

[0193] ;

[0194] in, It's the learning rate. It is a discount factor. Indicates selecting the next state. The Q value maximized at that time.

[0195] If using PolicyGradient, update the policy network parameters using the following formula. :

[0196] ;

[0197] in, It's the learning rate. This represents the expected cumulative reward. This indicates taking the partial derivative with respect to the policy network parameter θ.

[0198] Step 4.7: Adjust the exploration rate: based on the data freshness index And the stability of the market environment, dynamically adjust the exploration rate When market data is new and stable, the exploration rate is reduced to utilize known strategies; when market conditions change abruptly, the exploration rate is increased to explore new pricing strategies.

[0199] Step 4.8: Evaluation and Adjustment: Validate the effectiveness of the pricing strategy by evaluating its performance (e.g., through metrics such as actual market purchase conversion rate and sales volume). If the strategy does not meet expectations, adjust parameters such as the model's learning rate, exploration rate, discount factor, and decay factor in the data timeliness decay function to continue optimizing the training process.

[0200] Step 4.9: Iterative Update: Repeat the above steps to iterate the training until the model converges. During the training process, the policy network will gradually converge to the optimal pricing strategy.

[0201] Step 4.10: Final Output of the Optimal Pricing Strategy: After training, use the optimized policy network. The final pricing decision is then generated. At this point, the model can make the optimal pricing decision based on real-time market data, maximizing profits and maintaining the stability of the market strategy.

[0202] Through the steps described above, reinforcement learning models can learn and optimize pricing strategies through continuous interaction with the market environment. Especially within a game theory framework, leaders can continuously adjust their strategies and optimize price decisions based on feedback from competitors, thereby achieving more refined market pricing and equilibrium.

[0203] Step 5: Dynamic verification and updating of strategy stability and return optimization. After the price decision generated in Step 4 is executed, the leader continuously monitors market feedback data and return performance. For example... Figure 5 As shown, it specifically includes:

[0204] Step 5.1: Establish a strategy stability assessment module: Use statistical methods (such as sliding window ANOVA and price volatility threshold detection) to determine whether the current price strategy deviates from the historical stable range.

[0205] To strengthen the quantitative foundation of the strategy stability assessment module, a sliding window weighted variance calculation model based on data freshness is introduced to accurately capture the volatility characteristics of price series. The sliding window size is defined as... , corresponding to the most recent Pricing cycle, for each point in time The price sequence within the window is as follows: ,in Indicates the first The price decisions of wheel leaders.

[0206] To incorporate the impact of data timeliness, the data freshness index from step 1 is used. As a weight, the weighted average price within the window is calculated first:

[0207] ;

[0208] Then, the weighted variance is derived:

[0209] ;

[0210] The formula is passed The contribution of prices in each period is dynamically adjusted to ensure that high-freshness data dominates volatility assessment, thereby accurately reflecting the degree of price dispersion under real-time market conditions. If A high value indicates that the price strategy has fluctuated wildly in the recent window and may have deviated from the stable range.

[0211] To further quantify the volatility threshold, an adaptive threshold function is designed:

[0212] ;

[0213] in This represents the average of the long-term historical variance. For the size of the history window, This is the sensitivity adjustment parameter. The threshold function is achieved through... It captures the lowest data freshness within a window. When data timeliness decreases, it automatically increases the threshold tolerance to avoid misjudgments caused by outdated data. The stability judgment condition is... If true, it indicates that the current strategy has deviated from the historical stable range, triggering a subsequent strategy update mechanism. These formulas, by integrating data freshness and historical volatility patterns, enhance the robustness and adaptability of stability assessment, providing a reliable input basis for the return optimization module.

[0214] Step 5.2: Establish a profit optimization evaluation module: compare the profit difference between the current strategy and the historical strategy. If the profit decline exceeds the preset threshold (e.g., the profit is 10% lower than the average for three consecutive rounds), the strategy update mechanism will be triggered.

[0215] To establish a revenue optimization evaluation module, a dynamic revenue benchmark model based on data freshness is introduced to accurately compare the performance differences between the current strategy and historical strategies. The historical window size is defined as follows: Corresponding to the past Pricing cycle, for each round of time Historical return benchmark Calculated as a freshness-weighted average:

[0216] ;

[0217] in It is the first The data freshness index for this round originates from step 1. It is the first The actual reward of the round comes from the output of the reward function in step 4. This formula is derived through... The contribution of revenue in each period is dynamically adjusted to ensure that the benchmark more accurately reflects the revenue level during periods of high real-time data, thus avoiding distortion of the benchmark by outdated data.

[0218] Next, define the profit difference ratio. To quantify the relative performance of the current strategy:

[0219] ;

[0220] This difference in returns Directly measure current returns relative to historical benchmark The degree of deviation, with negative values ​​indicating a decrease in returns.

[0221] To enhance the robustness of the triggering mechanism, a cumulative deterioration metric is introduced:

[0222] ;

[0223] in For a fixed threshold such as 0.1, when And three consecutive rounds When the cumulative negative bias exceeds a threshold, a policy update mechanism is triggered. This metric captures the severity of persistent performance degradation, ensuring that updates are only initiated when there is a significant and continuous decline in performance, thereby reducing misjudgments and improving response efficiency. This quantitative evaluation module complements the policy stability evaluation, providing reliable input for model updates and driving iterative optimization of reinforcement learning policies.

[0224] Step 5.3: Feed the stability and reward evaluation results back to the reinforcement learning model as new training samples to update the state-action mapping relationship.

[0225] To deeply integrate stability and return evaluation results into the update mechanism of the reinforcement learning model, a policy optimization objective function based on evaluation feedback is designed to dynamically correct the state-action value mapping relationship. A comprehensive policy health index is defined as follows:

[0226] ;

[0227] in It is the sliding window weighted variance described in step 5.1. It is the profit difference ratio defined in step 5.2. and These are the balancing weighting coefficients.

[0228] This indicator quantifies the overall health of a strategy by integrating price volatility and the degree of return decline; its value range is [value range missing]. arrive The lower the value, the more serious the strategy problem.

[0229] Then, an enhanced advantage function is constructed:

[0230] ;

[0231] in It is the standard dominance function. This is the policy gradient. This function measures the policy health. As a modulating factor for advantage estimation, it enhances the contribution of the policy gradient term when health is low, forcing the model to update the state-action mapping more aggressively to correct bias.

[0232] Finally, this mechanism is embedded into the objective function of the Proximal Policy Optimization (PPO) algorithm to form the optimization objective:

[0233] ;

[0234] in This represents the probability ratio. The objective is to dynamically adjust the policy update magnitude using a health index, ensuring that the model prioritizes adjusting the mapping weights of problem actions when stability or returns deteriorate, thereby achieving efficient transmission of feedback information to policy parameters. This design enables the reinforcement learning model to autonomously optimize its decision logic based on evaluation results, laying an adaptive foundation for subsequent adjustments to the exploration rate based on data freshness.

[0235] Step 5.4: Recalculate the price decision based on the updated model parameters, and adjust the exploration rate in combination with data freshness (e.g., increase the exploration ratio when the data update frequency decreases, in order to recapture market changes).

[0236] After updating the model parameters, when recalculating price decisions, a comprehensive data freshness index is first constructed to quantify the overall data timeliness, and a weighted average freshness is defined:

[0237] ;

[0238] in It is a data source Importance weights, compared with those in step 1 echo; This is the current data freshness index. This represents the total number of data sources. This metric reflects the overall real-time performance level of multi-source data through weighted aggregation; a higher value indicates more reliable input data.

[0239] To further capture the internal variation of data freshness, we define freshness dispersion:

[0240] ;

[0241] This value measures the degree to which the freshness of each data source deviates from the mean. High dispersion indicates inconsistency in data status, which may introduce uncertainty into decision-making. Based on the above indicators, the exploration rate is adjusted as follows:

[0242] ;

[0243] in The base exploration rate is derived from the initial settings in step 3. and The positive adjustment parameter controls the strength of the influence of freshness mean and dispersion. When When the exponential term is high, it negatively dominates, thus reducing the exploration rate and prioritizing known high-yield strategies; when A higher exponential rate positively drives the exploration rate, mitigating the risk of market changes caused by data inconsistency. This dynamic adjustment mechanism ensures that the exploration ratio is automatically increased when the data update frequency decreases or the quality fluctuates, thus re-capturing market dynamics and avoiding strategy lag.

[0244] Subsequently, using For action selection: explore new policies by randomly and uniformly sampling from the action space with this probability; otherwise, select from the updated policy network. Select the optimal action This process generates the final price decision. The data-driven exploration rate optimization enhances the strategy's adaptability, providing a foundation for subsequent A / B testing and validation.

[0245] Step 5.5: Perform A / B testing on the updated strategy and the original strategy to ensure that the new strategy can improve returns and maintain user confidence in price changes in a non-stationary environment.

[0246] To scientifically verify the comprehensive effectiveness of the updated strategy in a non-stationary market environment, a dynamic A / B testing framework based on Bayesian decision theory is constructed. A comprehensive performance index for the strategy is defined as follows:

[0247] ;

[0248] in and These represent the average returns of the new strategy and the original strategy within the test window, respectively. The standard deviation of price volatility for the new strategy. The historical stable range serves as the volatility benchmark, derived from the sliding window variance analysis in step 5.1. This is a trade-off coefficient between returns and stability. This indicator quantifies both the magnitude of the return improvement and the change in volatility; a positive value indicates that the new strategy maintains or improves stability while increasing returns.

[0249] To further test the significance of the results, a weighted likelihood ratio statistic was designed:

[0250] ;

[0251] in and These correspond to the old and new strategies in the state. The following generates revenue The probability density, This refers to the weighted average freshness index defined in step 5.4. This statistic weights the likelihood ratio at each time point based on data freshness, ensuring that highly real-time data carries greater weight in decision-making and effectively addressing distribution drift in non-stationary environments. If simultaneously satisfying... and ,in and If a strict threshold is set, the new strategy is deemed validated and officially deployed. This dual-verification mechanism ensures that strategy optimization not only pursues revenue growth but also maintains users' long-term trust in the pricing system, ultimately forming a closed-loop system iteration process.

[0252] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the present invention.

Claims

1. A reinforcement learning dynamic pricing method based on game feedback, characterized in that: The method is suitable for online data market environment with time dynamics and data timeliness, comprising the following steps: Step 1: Real-time market data collection and freshness evaluation, through the deployment of data interface in online data market, real-time collection of multi-source dynamic data related to market state, including user behavior data, competitor pricing data, supply and demand relationship data and market environment data, at the same time, setting timeliness evaluation module for each type of collected data, calculating the freshness index of data through time stamp marking and data update frequency monitoring, the freshness index reflects the correlation weight of data and current market state; if the data timeliness is lower than the preset threshold, the weight of the corresponding data in the subsequent pricing model is reduced or the redundant data is removed, to ensure the real-time and effectiveness of model input; Step 2: Constructing Stackelberg game framework and defining market subject role, in the online data market, based on Stackelberg game theory, the platform or core pricing subject is defined as "leader", and other participants in pricing, service providers or competitors are defined as "followers"; the leader initiatively guides market equilibrium through dynamic pricing strategy, while the followers adjust their own strategy based on the pricing decision of the leader; specifically, the leader sets the initial price decision space according to the real-time collected market data and freshness index, and transmits it to the followers as the initial signal of the game; the followers generate strategy response to the leader according to the received signal, combined with their own cost, inventory and local market data, to form the basis of strategy interaction in dynamic game; Step 3: In the Stackelberg game framework, the leader uses "exploration-exploitation problem" method to model the price adjustment strategy as a dynamic balance problem, and balances the exploration of new price strategy and the exploitation of known high-yield strategy according to the feedback of market data, so as to optimize the price adjustment decision; Step 4: Real-time price adjustment based on game feedback and reinforcement learning model, the leader combines the exploration results of multi-armed bandit in step 3 with the game framework in step 2 to construct and optimize the reinforcement learning model; the reinforcement learning model specifically includes: Define state space: integrate real-time market data and data freshness index to form a multi-dimensional state vector with dimensionality where determined by the number of market data indicators and the number of data source types, Each dimension corresponds to a standardized market feature or freshness index, ensuring that the model can perceive the correlation between market state and data quality in real time.​ Definition of action space: Based on the high-yield price strategy explored in the multi-arm bandit, the top 20% of historical yield pricing schemes are selected as candidate actions, and each action includes two core parameters, namely the price adjustment range and the pricing mode identifier ; The definition of the reward function: maximize the revenue and the stability of the strategy as a double goal, the reward function takes into account the short-term revenue and long-term strategy stability; the calculation formula of the reward function is: ​ ; wherein, is a weight coefficient, according to the data freshness Dynamic adjustment, The higher The greater, focusing on short-term returns; The lower The smaller, focusing on strategy stability; is the actual return of the first round pricing cycle, in yuan; is the price adjustment range of the first round, is the historical maximum price adjustment range, in %; Train the model through deep reinforcement learning algorithm to learn the optimal price strategy in non-stationary environment; the model updates the strategy network by continuously receiving market feedback, and dynamically adjusts the learning rate and exploration rate according to the data freshness, to cope with market mutations; In each pricing cycle, the model generates the optimal price decision according to the current state, and pushes it to market participants in real time through API interface, while recording the decision and feedback data for subsequent model iteration; Step 5: Dynamic verification and update of strategy stability and revenue optimization, after the price decision generated in step 4 is executed, the leader continuously monitors market feedback data and revenue performance; specifically including: Step 5.1: Establish strategy stability evaluation module: determine whether the current price strategy deviates from the historical stable interval through statistical method; in order to strengthen the quantitative basis of strategy stability evaluation module, introduce sliding window weighted variance calculation model based on data freshness, to accurately capture the volatility characteristics of price sequence; Step 5.2: Establishing a revenue optimization evaluation module: comparing the revenue difference between the current strategy and the historical strategy, if the revenue decline exceeds the preset threshold, triggering the strategy update mechanism; in order to establish the revenue optimization evaluation module, a dynamic revenue benchmark model based on data freshness is introduced to accurately compare the performance difference between the current strategy and the historical strategy; Step 5.3: Feedback the stability and revenue evaluation results to the reinforcement learning model as new training samples to update the state-action mapping relationship; in order to deeply integrate the stability and revenue evaluation results into the update mechanism of the reinforcement learning model, a strategy optimization objective function based on evaluation feedback is designed to dynamically correct the state-action value mapping relationship; Step 5.4: Recalculate the price decision based on the updated model parameters to generate the final price decision and adjust the exploration rate combined with data freshness; Step 5.5: A / B test the updated strategy with the original strategy to ensure that the new strategy can not only improve revenue but also maintain users' trust in price changes in non-stationary environments. 2.The game feedback based reinforcement learning dynamic pricing method of claim 1, wherein: The specific method of step 1 is: Define the timeliness decay function of each data source wherein is the difference between the time of the last update of the data and the current time, in hours, The role of is to reflect the influence of time on the timeliness of the data; in order to make the timeliness present a nonlinear decay as time goes by, the following formula is used to express the exponential decay of the timeliness of the data: ; wherein, is a decay factor, controlling the speed of decay; a greater means the data timeliness decays faster; represents the timeliness of data update, the value closer to 1 means the data is fresher, the value closer to 0 means the data is outdated. Next, for each data source, introduce a freshness index F i to synthetically evaluate the timeliness of the data and its relevance to the current market state; for each data source i, the calculation formula of its freshness index F i is as follows: ; wherein, represents the timeliness decay value of data i, is the weight coefficient of data source i; is the market correlation index of data source i, which reflects the strength of the relationship between the data source and the current market decision; is obtained by analyzing historical data, indicating the contribution of the data source to market decision, ranging from 0 to 1; freshness index The higher the value of the freshness index, the more relevant and timely the data source is, and vice versa; The freshness index of each data source is calculated After the value of the freshness index F of each data source is calculated, the weight of each data source in the pricing model is adjusted; if the value of the freshness index F of some data sources is lower than a preset threshold, the weight of the corresponding data source in the model is reduced; in an extreme case, i.e., the freshness index F i of a data source is much lower than the preset threshold, and the data has almost no effective reference value for the current market pricing decision, the corresponding data source is completely excluded. 3.The game feedback based reinforcement learning dynamic pricing method of claim 2, wherein: The specific method of step 2 is: The leader sets an initial price decision space according to real-time collected market data and freshness index; the initial price decision of the leader is expressed as a price strategy space , wherein each price point represents a potential pricing decision; according to the known market state and freshness index , the leader determines the initial price space and transmits it to the follower through game. In the game process, the follower adjusts its pricing behavior according to the pricing strategy of the leader; the price adjustment decision of the follower depends on two key factors: the price decision of the leader, and its own cost, inventory level and local market data; a follower reaction function is introduced , which represents how the follower optimally adjusts its pricing based on its own situation given the leader's pricing, and the reaction function is represented as: ; where, Pleaderis the price decision of the leader, Pfolloweris the pricing decision of the follower, Cfolloweris the production cost of the follower, Ifolloweris the inventory level of the follower; Rfolloweris the revenue function of the follower, reflecting the profit of the follower at a given price and cost; in the reaction function , the follower tries to maximize its profit by adjusting . 4.The method of claim 3, wherein: The specific method of step 3 is: The leader considers potential pricing strategies as "arms" in a multi-armed bandit, each arm corresponding to a possible pricing scheme, the selection of which directly affects market response and thus needs to be maximized through dynamic adjustment; in order to evaluate the effect of each pricing strategy, the leader calculates the expected return of each arm according to the freshness index of market data and historical return data ; the calculation of expected return is carried out through the UCB algorithm, in which the return estimate of each arm is updated by the following formula: ; ; wherein, is the arm the average reward at time , t, is the arm the feedback reward in the th trial, is the arm the number of times it was selected, is the arm the upper confidence interval, reflecting the maximum reward the arm can obtain in future rounds; After calculating the UCB value of each arm, the leader adjusts the exploration probability according to these values, that is, how to choose between known high-yield strategies and new strategies; when the freshness of market data is high and the market state is stable, the leader will reduce the exploration probability and rely more on known high-yield strategies for use; when the market state mutates or the freshness of data is low, the leader will increase the exploration probability and increase the attempt of new pricing strategies; this adjustment is realized through the dynamic change of the exploration rate , which is specifically expressed as: ; wherein, and is a tuning parameter, is a freshness index for market data, is a threshold value; when the proximity or is greater than the threshold value, the rate of exploration is lower, meaning more reliance on known strategies; when is less than the threshold value, the rate of exploration is higher, encouraging more exploration; Finally, the leader updates the reward estimates of each arm in real time through the reward feedback mechanism of the multi-armed bandit, and adjusts the priority of the subsequent price decision; reward feedback Key indicators from user purchase behavior, conversion rate, inventory consumption speed, and specific forms are: ; wherein, represents the frequency of purchase behavior, represents the conversion rate, represents the speed of inventory consumption; , , respectively correspond to the weight coefficients of the three key indicators of the frequency of purchase behavior, the conversion rate, and the speed of inventory consumption in the reward feedback calculation, for adjusting the contribution degree of different indicators to the reward feedback Rᵢ(t). 5.The game feedback based reinforcement learning dynamic pricing method of claim 4, wherein: In step 4, the model is trained by deep reinforcement learning algorithm as follows: Input: State space: includes real-time collected market data, including current supply-demand ratio, competitor pricing strategy, inventory status, user profile features, and data freshness index; the state space is represented by a multi-dimensional vector, where each represents a specific market factor; Action space: In step 3, the high-reward price strategy explored by the leader through the MAB, including the price adjustment range, pricing mode; the action space is represented as wherein represents a pricing scheme; Reward function: The reward function considers both short-term revenue and long-term strategy stability; in each round of pricing decision, the reward function is defined as: ; wherein, is the immediate reward, is the long-term effect; and correspond to the short-term reward R short-term and the long-term effect R long-term respectively, in the reward function r t is a weight coefficient in the calculation of the reward function, used to adjust the contribution ratio of the short-term reward and the long-term policy stability in the reward. Output: Optimal pricing strategy: through reinforcement learning model training, get the optimal pricing strategy under the given market environment ; Policy network: trained policy network be able to make optimal pricing decisions in real-time based on real-time market conditions make optimal pricing decisions; The specific training process is as follows: Step 4.1: Initialize policy network and value network where and are parameters of the policy and value networks, respectively; Initialize experience replay buffer to an empty set; Step 4.2: Determine Initial State i.e. the initial state of the market environment, including the current supply-demand relationship, competitor pricing, inventory status, user profile characteristics, and data freshness index; at this time, the timestamps of all market data are marked, and the data freshness index is preliminarily calculated; Step 4.3: Select action: agent selects a pricing action based on current state from the action space ; in initial phase, use exploration strategy to explore, randomly select action; as training proceeds, gradually increase probability of using existing strategy Step 4.4: Market Feedback: The environment provides feedback according to the agent's actions, returning an immediate reward after executing a pricing decision and new market state ; reward function computed according to user behavior and competitor strategy changes; Step 4.5: Update experience replay buffer: store current experience into experience replay buffer Train the model by batch sampling from the replay buffer to train the model; Step 4.6: Strategy update: Update the strategy network using Q-learning or policy gradient method; If Q-learning is used, the value function is updated by the following equation : ; wherein, is a learning rate, is a discount factor, denotes the Q-value that is maximized when selecting the next state . If using the policy gradient method, the policy network parameters are updated by the following equation : ; wherein, is a learning rate, represents a desired cumulative reward; represents a partial derivative with respect to the policy network parameters θ. Step 4.7: Adjust exploration rate: based on data freshness index and market environment stability, dynamically adjust exploration rate When market data is fresh and stable, reduce exploration rate to exploit known strategies; when market state is abrupt, increase exploration rate to explore new pricing strategies; Step 4.8: Evaluation and adjustment: Evaluate the performance of the current strategy to verify the effect of the pricing strategy; if the strategy effect does not meet the expectation, adjust the learning rate, exploration rate, discount factor, and decay factor in the data time effectiveness decay function to continue optimizing the training process; Step 4.9: Iterative update: Repeat the above steps to continuously iterate and train until the model converges; the strategy network in the training process will gradually tend to the optimal pricing strategy; Step 4.10: Final output optimal pricing strategy: After training is complete, the optimized strategy network is used Generate final pricing decisions; at this point, the model makes optimal pricing decisions based on real-time market data, maximizing returns and maintaining the stability of the market strategy. 6.The method of claim 5, wherein: The step 5.1 specifically includes: The size of the sliding window is defined as W, which corresponds to the latest W pricing periods. For each time point t, the price sequence within the window is: ; where, denotes the price decision of the i-th leader. To incorporate the impact of data timeliness, use the data freshness index F from Step 1 i As weights, first compute the weighted average price within the window: ; Then derive the sliding window weighted variance: ; The formula passes F i The contribution degree of each period price is dynamically adjusted to ensure that high freshness data dominates the volatility evaluation, thereby accurately reflecting the price dispersion degree under real-time market conditions. If the value of V(t) is high, it indicates that the price strategy has fluctuated violently in the recent window and may deviate from the stable interval. To further quantify the volatility threshold, an adaptive threshold function is designed: ; wherein, , represents the historical long-term variance average value, L is the historical window size, is a sensitivity adjustment parameter; an adaptive threshold function by capturing the lowest data freshness within the window, when the data timeliness decreases, automatically raising the threshold tolerance, avoiding false judgments caused by outdated data; the stability determination condition is , if it is true, it indicates that the current strategy deviates from the historical stable interval, triggering the subsequent strategy update mechanism. 7.The game feedback based reinforcement learning dynamic pricing method of claim 6, wherein: The step 5.2 specifically includes: Define the history window size as L, corresponding to the past L rounds of pricing period, for each round time t, the historical revenue benchmark is calculated as the freshness weighted average: ; in, It is the first The data freshness index of the round, It is the first The actual benefit of the wheel; this formula is obtained through The contribution of revenue in each period is dynamically adjusted to ensure that the benchmark more accurately reflects the revenue level during periods of high real-time data, and to avoid distortion of the benchmark by outdated data. Next, define the earnings differential ratio to quantify the relative performance of the current strategy: ; Earnings surprise ratio Direct measure of current earnings Relative to historical earnings benchmark Degree of deviation, negative values indicate earnings decline; To enhance the robustness of the triggering mechanism, a cumulative deterioration index is introduced: ; wherein, is a fixed threshold, when and three consecutive rounds have the strategy update mechanism triggered; the cumulative deterioration indicator captures the severity of the sustained deterioration in returns by accumulating the magnitude of negative deviations beyond the threshold, ensuring that the update is initiated only when there is a significant and continuous decline in performance. 8.The method of claim 7, wherein: The step 5.3 specifically includes: Define a comprehensive strategy health index: ; wherein, is the sliding window weighted variance as described in step 5.1, is the difference in returns ratio defined in step 5.2, and is the balancing weight coefficient; Comprehensive strategy health indicator By fusing price volatility and the degree of drawdown, the overall health status of a quantitative strategy is quantified, with a value range of 0 to 1, and the lower the value, the more serious the strategy problem. Then build an enhanced advantage function: ; where, is the standard advantage function, is the policy gradient; this function will be used to monitor the health of the policy as a modulation factor for the advantage estimate, when the health is low, it enhances the contribution of the policy gradient term, forcing the model to more aggressively update the state-action mapping to correct the bias; Finally, embed this mechanism into the objective function of the proximal policy optimization algorithm (PPO algorithm) to form the optimization objective: ; wherein, is the probability ratio; the target updates the amplitude of the health index dynamic adjustment strategy, ensuring that the model prioritizes adjusting the mapping weight of the problem action when stability or returns deteriorate, thereby achieving efficient transmission of feedback information to the strategy parameters. 9.The game feedback based reinforcement learning dynamic pricing method of claim 8, wherein: The step 5.4 specifically includes: When recalculating the price decision, first build a comprehensive data freshness index to quantify the overall data timeliness; define the weighted average freshness: ; wherein, is the importance weight of data source i, and in step 1, is the current data freshness index, N is the total number of data sources; the weighted average freshness index reflects the overall real-time level of multi-source data through weighted aggregation, and the higher the value, the more reliable the input data; To further capture the internal variation of data freshness, define the freshness dispersion: ; This value measures the deviation of each data source freshness from the mean, and high dispersion indicates inconsistent data state, which may introduce decision uncertainty; Based on the above index, adjust the exploration rate as: ; wherein, is the base exploration rate, derived from the initial setting in Step 3; and is the positive regulation parameter, controlling the influence strength of the freshness mean and dispersion; when is higher, the exponential term is negatively dominant, thus reducing the exploration rate and prioritizing known high-reward strategies; when is higher, the exponential term is positively driven, thus increasing the exploration rate and addressing the market change risk brought by data inconsistency; Subsequently, using For action selection: explore new policies by sampling uniformly at random from the action space with this probability, otherwise select the optimal action from the updated policy network Generate final price decision.​ 10.The method of claim 9, wherein: The step 5.5 specifically includes: To scientifically verify the comprehensive performance of the updated strategy in non-stationary market environments, a dynamic A / B testing framework based on Bayesian decision theory is constructed; define a comprehensive strategy performance index: ; wherein, and respectively represent the average returns of the new strategy and the original strategy within the test window, is the price volatility standard deviation of the new strategy, is the volatility benchmark of the historical stable interval, derived from the sliding window variance analysis in step 5.1; is the trade-off coefficient between returns and stability; this strategy performance composite indicator quantifies both the return improvement magnitude and the volatility change, a positive value indicates that the new strategy has improved returns while maintaining or improving stability. To further test the significance of the results, a weighted likelihood ratio statistic was designed: ; wherein, with corresponding to the probability density of the yield of the new and old policies in the state , respectively, is the weighted average freshness indicator defined in step 5.4; the weighted likelihood ratio statistic weights the likelihood ratios of each time point by the data freshness, ensuring that high real-time data occupies a higher weight in decision-making, effectively dealing with distribution drift in non-stationary environments; if the following conditions are met simultaneously and , wherein and are strict thresholds, it is determined that the new policy passes the verification and is formally deployed.​

Citation Information

Patent Citations

  • Power control and interference pricing methods and devices for heterogeneous renewable energy networks

    CN112533275B

Cited By

  • Power implicit energy storage user identification method based on game analysis

    CN122364785A