Method and system for inhibiting spontaneous monopoly of pricing algorithm, computer storage medium and computer equipment

By combining multi-subject game model and reinforcement learning algorithm, dynamically calculate the penalty coefficient and real-time monitoring mechanism, the monopoly tendency of pricing algorithms is suppressed, and the problem of spontaneous formation of monopoly by pricing algorithms in the market environment is solved, and fair competition in the market and reasonable allocation of resources are achieved.

CN120106880APending Publication Date: 2025-06-06HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510137373.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In a market environment where multi-subject games are played, pricing algorithms based on reinforcement learning may spontaneously form monopoly pricing behavior, resulting in reduced market competitiveness, reduced consumer choice and inefficient resource allocation.

Method used

By combining multi-subject game model and reinforcement learning algorithm, the penalty coefficient is calculated dynamically, the punishment mechanism based on competitive prices and monopoly prices, as well as real-time monitoring and adjustment strategy mechanisms, the monopoly tendency of pricing algorithms is suppressed.

Benefits of technology

Effectively curb the formation of monopoly behavior in the process of independent learning, promote fair competition in the market and reasonable allocation of resources, and protect the interests of consumers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106880A_ABST
    Figure CN120106880A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for inhibiting spontaneous monopoly of a pricing algorithm, a computer storage medium and computer equipment, relates to the technical field of artificial intelligence, and aims to control pricing algorithm behaviors by introducing a dynamic adjustment mechanism, inhibit spontaneous monopoly of the algorithm in market competition and improve the pricing efficiency. And a partial competition pricing strategy is adopted in the algorithm. The algorithm combines a multi-subject game model and a reinforcement learning algorithm, and can reduce dead weight loss and prevent the price from being excessively concentrated in a monopoly range by calculating market demands and price conditions and utilizing a penalty coefficient to adjust an award mechanism of the algorithm. A multi-dimensional game model is specifically designed, wherein a plurality of participants compete by selecting different pricing strategies. The punishment coefficient in the Q value updating function is dynamically adjusted, the dead weight loss area of the Halberg triangle is introduced as a punishment mechanism, and the behavior of monopoly pricing is punished, so that the competitive behavior in the market is promoted. Meanwhile, by calculating consumer residue, dead weight loss and collusion degree indexes in real time, whether the algorithm tends to monopolize the market or not can be effectively monitored, and a reward strategy is adjusted to optimize market behaviors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and specifically relates to a method and system for suppressing spontaneous monopoly of pricing algorithms, a computer storage medium, and a computer device. Background Art

[0002] With the rapid development of information technology and artificial intelligence, especially the widespread application of reinforcement learning algorithms, more and more companies and research institutions have begun to adopt intelligent pricing strategies to maximize profits and increase market share. These intelligent agents can make efficient decisions in complex market environments by continuously learning and optimizing pricing strategies. However, in a multi-agent game market environment, these pricing algorithms based on reinforcement learning may gradually tend to form monopolistic pricing behavior during the autonomous learning process.

[0003] Monopolistic pricing behavior refers to the phenomenon that a few or a single market participant obtains excess profits by controlling prices and reducing market competition. This behavior not only weakens the competitiveness of the market and reduces consumers' choice and welfare, but may also lead to inefficient allocation of market resources and hinder the healthy development of the market. Traditional pricing mechanisms and market supervision methods have certain limitations when dealing with monopoly phenomena spontaneously formed by intelligent algorithms, and it is difficult to effectively intervene and adjust the pricing behavior of algorithms in real time.

[0004] In the prior art, although some studies have attempted to constrain the behavior of reinforcement learning algorithms by introducing penalty mechanisms or restrictions, these methods often lack dynamic adjustment capabilities and are difficult to adapt to real-time changes in the market environment. In addition, many methods fail to fully consider the complex interactive relationships in multi-agent games, resulting in limited effectiveness in practical applications. Therefore, there is an urgent need for a method and system that can dynamically suppress the spontaneous formation of monopoly phenomena in pricing algorithms, so as to promote fair competition in the market and the rational allocation of resources, and protect the interests of consumers.

[0005] To solve the above problems, the present application provides a method and system, a computer storage medium, and a computer device for suppressing the spontaneous monopoly of pricing algorithms. Summary of the invention

[0006] The purpose of the present invention is to provide a method and system, a computer storage medium, and a computer device for suppressing the spontaneous monopoly of pricing algorithms to solve the above-mentioned technical problems. The method combines a multi-agent game model with a reinforcement learning algorithm, and effectively suppresses the monopoly tendency of pricing algorithms by dynamically calculating penalty coefficients, a penalty mechanism based on competitive prices and monopoly prices, and a real-time monitoring and adjustment strategy mechanism.

[0007] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is as follows:

[0008] A method for suppressing the spontaneous formation of monopoly by a pricing algorithm based on reinforcement learning, comprising the following steps:

[0009] Construct an AI agent pricing algorithm model based on reinforcement learning, which simulates the pricing decision-making behavior of multiple participants in the market through multi-agent games. Specifically, define the pricing strategy space of multiple participants, set the market demand function and cost parameters, and use the reinforcement learning algorithm to enable each participant to optimize its pricing strategy in the process of continuous learning and game.

[0010] Dynamically calculate the penalty coefficient, which is dynamically calculated based on the ratio of deadweight loss in the market to collusion deadweight loss, so as to adjust the reward coefficient in reinforcement learning and implement the punishment for monopoly pricing behavior. Specifically, it includes the following steps:

[0011] Step 1: Calculate the deadweight loss (dwl) caused by algorithmic pricing behavior in the market i ), that is, by comparing the demand differences and price differences generated by competitive prices and algorithmic pricing, the social welfare losses caused by the pricing behavior of participants are calculated. The calculation formula is:

[0012] dwl i =0.5×(d competitive -d i )×(p i -p competitive );

[0013] Among them, p competitive is the optimal price in a perfectly competitive market, d competitive is the algorithm demand in a perfectly competitive market, p i is the price determined by the algorithm under the current pricing behavior, d i The algorithmic demands faced by the algorithm under current pricing behavior.

[0014] Step 2: Calculate the deadweight loss (dwl) of complete collusion cartel ), that is, by calculating the demand difference and price difference between the competitive price and the monopoly price, the welfare loss caused by collusive pricing is determined. The calculation formula is:

[0015] dwl cartel =0.5×(d competitive -d cartel )×(p cartel -p competitive );

[0016] Among them, p competitive For in full

[0017] The optimal price in a competitive market, d competitive is the algorithm demand in a perfectly competitive market, pcartel is the optimal price in a monopoly market, d cartel For the demand in a completely monopolistic market.

[0018] Step 3: Based on the above calculation results, the dead weight loss ratio (penalty_index) is obtained, the formula is:

[0019]

[0020] This ratio is used to adjust the reward mechanism in reinforcement learning to dynamically suppress the monopoly phenomenon that spontaneously arises from the algorithm.

[0021] Step 4: When the algorithmic pricing behavior is close to monopoly, the ratio penalty_index will guide the reinforcement learning model to increase the penalty for the behavior, forcing the pricing behavior to be more in line with competitive market rules, thereby avoiding monopoly in the market.

[0022] The penalty mechanism based on competitive prices and monopoly prices quantifies the degree to which algorithmic pricing deviates from the competitive state by comparing competitive prices with monopoly prices, and applies it to the adjustment of the reward mechanism. Specifically, it includes the following steps:

[0023] Step 1: Calculate the competitive price p competitive and the monopoly price p cartel , where p competitive is the optimal price in a perfectly competitive market, p cartel is the optimal price in a completely monopolistic market.

[0024] Step 2: Calculate the deadweight loss (dwl) based on the competitive price and the monopoly price i ) and collusion dead weight loss (dwl cartel ), the specific formula is as follows:

[0025] dwl i =0.5×(d competitive -d i )×(p i -p competitive );

[0026] dwl cartel =0.5×(d competitive -d cartel )×(p cartel -p competitive );

[0027] Step 3: Based on the deadweight loss calculated above, calculate the deadweight loss ratio penalty_index, which is used to quantify the degree to which the algorithm deviates from competitive pricing in the market.

[0028] Step 4: During the reinforcement learning update process, adjust the reward coefficient adjusted_π by multiplying the penalty_index ratio by the manually set control factor game.control_factor i , the specific formula is:

[0029] adjusted_π i =π i ×(1-game.control_factor×penalty_index[n]);

[0030] Among them, π i It is the actual profit obtained by the algorithm and the current pricing behavior; n is defined as the market entity number in the current market. For example, if the market entity number is 1, then its corresponding value is penalty_index[1].

[0031] Through the above steps, it is possible to punish behaviors that deviate from competitive pricing and reduce the risk of algorithms spontaneously forming a monopoly.

[0032] The penalty coefficient penalty_index is introduced into the reward update process of the reinforcement learning algorithm, and the actual reward of the current pricing strategy is dynamically adjusted according to the size of the penalty coefficient penalty_index, forming a dynamic penalty mechanism, so that when the strategy tends to monopoly, it will be punished more strongly, and when the strategy tends to competition, the punishment will be relatively weakened; the method of introducing penalty_index is:

[0033] adjusted_π i =π i ×(1-game.control_factor×penalty_index[n]),

[0034] The conventional pricing algorithm only aims at profit, and its reward is exactly equal to profit π i However, this behavior pattern will cause the algorithm to remain in a collusion state. In this technical solution, the profit is reduced based on the degree of collusion, i.e., penalty_index, to obtain a new adjusted_π i The mechanism of dynamically adjusting the control factor, which controls the weight of the penalty coefficient penalty_index in the algorithm adjustment through the manually set control factor game.control_factor, specifically includes the following steps:

[0035] Step 1: Set a manually set control factor game.control_factor, which is used to adjust the weight of the penalty coefficient penalty_index in the reward update of the reinforcement learning algorithm.

[0036] Step 2: Each time the Q value of the reinforcement learning algorithm is updated, the penalty coefficient penalty_index is adjusted using the control factor game.control_factor. Specifically, the penalty intensity of each participant is adjusted by calculating the adjustment_factor. The specific formula is:

[0037] adjustment_factor=1-game.control_factor×penalty_index[n];

[0038] Step 3: The control factor game.control_factor is manually set by the user during the initialization phase and remains unchanged throughout the training process to ensure that the penalty mechanism consistently suppresses the algorithm's tendency to spontaneously form a monopoly during the learning process.

[0039] Real-time monitoring and adjustment strategy mechanism, which monitors pricing behavior in real time at each time step and dynamically adjusts the reward function based on the monitoring results to suppress potential monopolistic behavior. Specifically, it includes the following steps:

[0040] Step 1: At each time step t, monitor the price p generated by the algorithmic strategy in real time i Competitive price p competitive and the monopoly price p cartel The deviation between the two is calculated and the dead weight loss dwl is updated. i and conspiracy to die heavy losses dwl cartel , to assess the impact of the algorithm’s current strategy on the market.

[0041] Step 2: According to the real-time calculated deadweight loss ratio penalty_index and cumulative collusion index cumulative_collusion_index, adjust the algorithm’s reward mechanism. The specific formula is:

[0042]

[0043] Among them, π i is the actual profit obtained by the algorithm and the current pricing behavior, π competitive is the profit of the agent algorithm in a fully competitive environment, π cartel is the profit of the agent algorithm in a completely monopolistic environment.

[0044] By calculating the cumulative degree of collusion, cumulative_collusion_index, the overall collusion tendency of the algorithm behavior is evaluated. The specific formula is:

[0045]

[0046] Among them, mean_t_collusion_index is the arithmetic mean of the degree of collusion of a single agent algorithm in all periods, game.n is the number of participants, and the specific formula is:

[0047]

[0048] Where T is the total number of cycles of the agent algorithm action.

[0049] Step 3: In each training round, adjust the reward coefficient adjusted_π based on the penalty_index and the manually set control factor game.control_factor i , the specific formula is:

[0050] adjusted_π i =π i ×(1-game.control_factor×penalty_index[n]);

[0051] Through this adjustment, the algorithm is guided to adopt more competitive pricing strategies and reduce potential monopolistic behavior.

[0052] Step 4: During the algorithm training process, the real-time feedback and adjustment mechanism ensures that the reward mechanism is dynamically adjusted by combining the manually set control factor game.control_factor with the penalty coefficient penalty_index calculated in real time to avoid the algorithm from forming pricing behaviors that are unfavorable to consumers due to excessive exploration.

[0053] Step 5: Through real-time monitoring and adjustment, ensure that the pricing strategy at each time step complies with market competition rules, dynamically respond to market changes, continuously suppress the tendency of algorithms to form monopolies, and safeguard fair competition in the market and the interests of consumers.

[0054] A competitive equilibrium ensuring mechanism with multiple rounds of iterative updates, which achieves market competition equilibrium through multiple rounds of game iterations and Q-value updates of reinforcement learning algorithms. Specifically, it includes the following steps:

[0055] Step 1: During the training process of the reinforcement learning algorithm, multiple rounds of game iterations are performed, and each participant in each round of the game selects a pricing strategy based on the current Q value.

[0056] Step 2: After each round of game, calculate the profit π of each participant according to the pricing strategy i , and adjust the reward function through the penalty coefficient penalty_index to ensure that the algorithm gradually optimizes the pricing strategy.

[0057] Step 3: Through continuous Q value updates, the pricing strategies of each participant gradually tend to the market competition equilibrium state, ensuring that the pricing behavior of all participants remains within a reasonable competitive range.

[0058] A pricing control algorithm applicable to price-sensitive markets, which ensures that pricing strategies remain competitive in a market environment with high price sensitivity and effectively prevents price manipulation and monopoly. Specifically, the following steps are included:

[0059] Step 1: Define the price sensitivity parameters for the applicable market, using parameters b and b 1 Reflects the market's sensitivity to price changes. Parameter b represents the direct impact of price on demand. 1 Represents the indirect effect of price difference on demand.

[0060] Step 2: In the multi-agent game model, the impact of price sensitivity on market demand is considered, and the market demand is dynamically adjusted through the function demand(p) to obtain the demand d of each agent algorithm. The market demand function calculates the market demand of each participant under different pricing according to the pricing strategy and price sensitivity parameters of the participants. The specific formula is:

[0061] d=ab×p+b 1 ×p 1 ;

[0062] Among them, parameter a is the maximum demand scale in the linear demand function, parameter b represents the direct impact of price on demand, b1 represents the indirect impact of price difference on demand, p is the algorithm's own pricing, and p 1 Pricing another algorithm in a market environment.

[0063] Step 3: By introducing the penalty coefficient penalty_index, price manipulation behavior is penalized in the reinforcement learning algorithm to reduce the motivation of participants to adopt excessively high pricing strategies. Specifically, the penalty coefficient affects the reward function, prompting the algorithm to optimize the pricing strategy to avoid high-price manipulation:

[0064] Step 4: During the algorithm training process, monitor the impact of price changes on market demand and profits in real time to ensure that the algorithm does not deviate from market competition rules while optimizing profits. The monitoring mechanism includes calculating market demand changes, profit fluctuations and their correlation with pricing strategies.

[0065] Step 5: By dynamically adjusting the reward mechanism and penalty coefficient, the algorithm is encouraged to maintain competitive pricing in a market with high price sensitivity, prevent price manipulation and monopoly, and ensure the rational allocation of market resources and the maximization of consumer interests. The adjustment mechanism dynamically optimizes the pricing strategy based on the penalty coefficient penalty_index calculated in real time and market feedback data.

[0066] The system is suitable for pricing decision-making environments in highly price-sensitive markets and highly competitive markets to achieve real-time adaptive suppression of pricing strategies with collusion or monopoly tendencies without the need to trigger penalty adjustments through preset thresholds.

[0067] The market fairness guarantee mechanism of collusion index and penalty coefficient calculates the cumulative collusion degree index and dynamically adjusts the penalty coefficient to ensure market fairness. Specifically, it includes the following steps:

[0068] Step 1: Calculate the collusion index, the specific formula is:

[0069]

[0070] The degree of collusion among participants is evaluated by collusion_index, π i is the actual profit obtained by the algorithm and the current pricing behavior, π competitive is the profit of the agent algorithm in a fully competitive environment, π cartel is the profit of the agent algorithm in a completely monopolistic environment.

[0071] Step 2: At each time step t, calculate the sum of the collusion degrees of each period before, sum_t_collusion_index. The specific formula is:

[0072]

[0073] And calculate the average collusion degree mean_t_collusion_index of a single agent algorithm. The specific formula is:

[0074]

[0075] Where t is the current cycle number of the agent algorithm action.

[0076] The cumulative collusion index is further obtained. The specific formula is:

[0077]

[0078] Among them, game.n is the number of participants.

[0079] Step 3: According to the cumulative collusion index cumulative_collusion_index, adjust the penalty coefficient penalty_index through the formula:

[0080] adjusted_π i =π i ×(1-game.control_factor×penalty_index[n]);

[0081] Ensure effective punishment for collusion, prevent excessive concentration of market prices, and safeguard market diversity and fairness.

[0082] In the process of using the cumulative collusion index, at the end of each training session, the size of game.control_factor is controlled according to the size of the cumulative collusion index cumulative_collusion_index, so that

[0083] adjusted_π i =π i ×(1-game.control_factor×penalty_index[n]); the degree of influence of penalty index penalty_index; the specific implementation method is that the cumulative collusion degree index of the current period

[0084] The higher the cumulative_collusion_index is, the greater the game.control_factor will be in the next period, which will make the penalty_index in the next period punish collusion to a greater extent.

[0085] System architecture, the system includes a data processing module, a computing module, a reinforcement learning module, a control module, a monitoring module and a user interface, which is used to implement a method for suppressing the spontaneous formation of monopoly phenomena by suppressing pricing algorithms based on reinforcement learning. The specific components and their functions are as follows:

[0086] The data processing module is used to receive and process market data, including competitive prices, monopoly prices, algorithmic pricing, market demand, consumer surplus, etc. The data processing module ensures the accuracy and consistency of input data through data preprocessing, cleaning and standardization, and provides a high-quality data foundation for subsequent calculations and algorithm training.

[0087] The calculation module is used to calculate the deadweight loss and collusion deadweight loss, and dynamically adjust the penalty coefficient penalty_index according to the calculation results. Specific functions include:

[0088] Step 1: Based on the received market data, calculate the market demand d faced by the current agent algorithm pricing i , and market demand in a competitive market competitive and the market demand in a monopoly market cartel ;

[0089] Step 2: Calculate the deadweight loss dwl caused by the current pricing behavior i , the specific formula is:

[0090] dwl i =0.5×(d competitive -d i )×(p i -p competitive );

[0091] Step 3: Calculate the deadweight loss dwl under full collusion pricing cartel The specific formula is:

[0092] dwl cartel =0.5×(d competitive -d cartel )×(p cartel -p competitive );

[0093] Step 4: Calculate the penalty_index, the ratio of the deadweight loss, and use it to adjust the reward coefficient in the reinforcement learning algorithm to achieve dynamic suppression of monopoly pricing behavior. The specific formula is:

[0094]

[0095] The reinforcement learning module is used to perform pricing strategy optimization based on the reinforcement learning algorithm. The module includes the learning and decision-making units of the intelligent agent, and its specific functions include:

[0096] Step 1: Calculate the algorithm’s profit π i , the specific formula is:

[0097] π i =d i ×(p i -c);

[0098] Among them, p i is the price determined by the algorithm under the current pricing behavior, d i is the algorithm demand faced by the algorithm under the current pricing behavior, and c is the unit cost of the goods sold.

[0099] Step 2: Based on the calculated profit π iAnd the dynamically adjusted reward coefficient adjusted_πi is used to update the Q-value function in the reinforcement learning algorithm to optimize the pricing decisions of participants and promote competitive pricing behavior in the market.

[0100] The control module is used to set and manage the manually set control factor game.control_factor to accurately adjust the weight of the penalty coefficient penalty_index in the reward mechanism. Specifically, it includes the following steps:

[0101] Step 1: The user manually sets the control factor game.control_factor through the user interface and adjusts its weight in the algorithm reward mechanism to control the severity of the penalty.

[0102] Step 2: The control module monitors the game.control_factor set by the user, and applies the factor to the adjustment of the penalty coefficient penalty_index during the algorithm operation to ensure that the penalty intensity meets the preset standard.

[0103] The monitoring module is used to monitor market pricing behavior and the learning process of the algorithm in real time to ensure that the pricing strategy complies with market competition rules. Specific functions include:

[0104] Step 1: At each time step t, collect and analyze the algorithm’s pricing behavior data in real time, and calculate the current market competition index and collusion degree index.

[0105] Step 2: Based on the monitoring data, evaluate the competition status and potential monopoly risks in the market and generate corresponding feedback information.

[0106] Step 3: The monitoring module passes the feedback information to the calculation module and the control module so as to adjust the penalty coefficient and reward mechanism in real time.

[0107] User interface, used for user interaction and system configuration. Specific functions include:

[0108] Step 1: Provide an interface for users to manually set the control factor game.control_factor, allowing users to adjust the penalty intensity according to market conditions.

[0109] Step 2: Provide a real-time monitoring data display interface to show key indicators such as market competition status, deadweight loss, degree of collusion, and penalty coefficient penalty_index, to assist users in making decisions and adjustments.

[0110] This application has achieved beneficial technical effects:

[0111] The present invention prevents the pricing algorithm from forming monopoly behavior in the autonomous learning process by dynamically adjusting the penalty coefficient and real-time monitoring mechanism, promotes fair competition in the market, and protects the interests of consumers; by introducing a dynamic adjustment mechanism to control the behavior of the pricing algorithm, the algorithm is inhibited from spontaneously forming a monopoly phenomenon in market competition, thereby prompting the algorithm to adopt a competitive pricing strategy. The algorithm combines a multi-agent game model with a reinforcement learning algorithm, and can reduce deadweight losses and avoid excessive concentration of prices within a monopoly range by calculating market demand and price conditions, and adjusting the algorithm's reward mechanism using a penalty coefficient.

[0112] Specifically, the present invention designs a multi-dimensional game model in which multiple participants compete by choosing different pricing strategies. By dynamically adjusting the penalty coefficient in the Q-value update function and introducing the deadweight loss area of ​​the Harberger triangle as a penalty mechanism, the behavior of monopolistic pricing is punished, thereby promoting competitive behavior in the market. At the same time, by calculating consumer surplus, deadweight loss and collusion degree indicators in real time, it is possible to effectively monitor whether the algorithm tends to monopolize the market and adjust the reward strategy to optimize market behavior.

[0113] The advantage of the present invention is that it can respond to market changes in real time and avoid the formation of monopoly in the game through a reasonable penalty mechanism, thus ensuring the fairness and competitiveness of the market. This method is applicable to a variety of pricing issues, especially in highly competitive and price-sensitive markets, which can effectively reduce monopoly and improve the overall efficiency of the market. BRIEF DESCRIPTION OF THE DRAWINGS

[0114] Figure 1 The figure shows the overall architecture diagram of the system that suppresses the spontaneous formation of monopoly by the pricing algorithm based on reinforcement learning, which shows the structural relationship between the modules of the system.

[0115] Figure 2 It is a flow chart of the method for suppressing the spontaneous formation of monopoly by a pricing algorithm of the present invention, which corresponds to the process from S10 to S60.

[0116] Figure 3 It is a schematic diagram of the interaction process between the data processing module and the calculation module of the present invention, which involves the calculation of dead weight loss by the calculation module after data processing, which can correspond to S40 dead weight loss calculation and S90 price calculation.

[0117] Figure 4 It is a schematic diagram of the interaction process between the reinforcement learning module and the control module and the monitoring module of the present invention, which involves the reinforcement learning S20, S30, S60 and the adjustment-related steps of the control modules S50, S130 and the monitoring modules S160, S170, S180, S190, S200.

[0118] Figure 5 It is a flow chart of the real-time monitoring and adjustment strategy mechanism of the present invention, which corresponds to S160 real-time monitoring, S170 calculation of collusion degree, S180 cumulative collusion index, S190 dynamic adjustment of reward mechanism, and S200 feedback of adjustment results.

[0119] Figure 6 It is a flow chart of the competitive equilibrium ensuring mechanism of multiple rounds of iterative updates of the present invention, which involves multiple rounds of game S210, profit and reward adjustment S220, continuous update of Q value S230, real-time monitoring and dynamic adjustment S240, and final convergence S250.

[0120] Figure 7 Flow chart of the artificial intelligence algorithm used in the present invention. DETAILED DESCRIPTION

[0121] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0122] A reinforcement learning-based method to suppress the spontaneous formation of monopoly by pricing algorithms. Based on the construction of an AI agent pricing model based on reinforcement learning, the pricing decision-making behaviors of multiple participants in the market are simulated through multi-agent game, and the market pricing strategy is dynamically adjusted to suppress the spontaneous formation of monopoly by the algorithm during the autonomous learning process.

[0123] Reference Figure 1 The overall architecture of the system for suppressing the spontaneous formation of monopoly phenomenon by the pricing algorithm based on reinforcement learning provided in this embodiment includes six parts: data processing module, calculation module, reinforcement learning module, control module, monitoring module and user interface. The functions of each module and their interactive relationship are described in detail as follows.

[0124] The data processing module is used to receive and process market-related data, including competitive prices, monopoly prices, algorithmic pricing, market demand, consumer surplus, etc. Specifically, the data processing module ensures the accuracy and consistency of input data through data preprocessing, cleaning and standardization, providing a high-quality data foundation for subsequent calculations and algorithm training.

[0125] The calculation module is responsible for calculating the deadweight loss and collusion deadweight loss, and dynamically adjusting the penalty coefficient penalty_index according to the calculation results. The specific steps include:

[0126] Calculate the market demand d under the current algorithm pricing i , and market demand in a competitive market competitive and the market demand in a monopoly marketcartel ;

[0127] Calculate the deadweight loss dwl caused by current pricing behavior i , the specific formula is:

[0128] dwl i =0.5×(d competitive -d i )×(p i -p competitive );

[0129] Calculate the current deadweight loss dwl under complete collusion pricing cartel , the specific formula is:

[0130] dwl cartel =0.5×(d competitive -d cartel )×(p cartel -p competitive );

[0131] Calculate the penalty_index of the deadweight loss ratio and use it to adjust the reward coefficient in the reinforcement learning algorithm to achieve dynamic suppression of monopoly pricing behavior. The specific formula is:

[0132]

[0133] The reinforcement learning module is used to perform pricing strategy optimization based on reinforcement learning algorithms. Specific functions include:

[0134] Calculate the algorithm's profit π i , the specific formula is:

[0135] π i =d i ×(p i -c);

[0136] Among them, p i is the price determined by the algorithm under the current pricing behavior, d i is the algorithm demand faced by the algorithm under the current pricing behavior, and c is the unit cost of the goods sold.

[0137] According to the calculated profit π i and the dynamically adjusted reward coefficient adjusted_π i , update the Q-value function in the reinforcement learning algorithm to optimize the pricing decisions of participants and promote competitive pricing behavior in the market.

[0138] The control module is used to set and manage the manually set control factor game.control_factor to accurately adjust the weight of the penalty coefficient penalty_index in the reward mechanism. The specific steps include:

[0139] The user manually sets the control factor game.control_factor through the user interface to adjust its weight in the algorithm reward mechanism to control the intensity of the penalty;

[0140] The control module monitors the game.control_factor set by the user and applies it to the adjustment of the penalty coefficient penalty_index during the algorithm operation to ensure that the penalty intensity meets the preset standards.

[0141] The monitoring module is used to monitor market pricing behavior and the learning process of the algorithm in real time to ensure that the pricing strategy complies with market competition rules. Specific functions include:

[0142] At each time step t, the pricing behavior data of the algorithm is collected and analyzed in real time to calculate the current market competition index and collusion degree index;

[0143] Based on monitoring data, evaluate the competition status and potential monopoly risks in the market and generate corresponding feedback information;

[0144] The monitoring module passes feedback information to the calculation module and the control module so as to adjust the penalty coefficient and reward mechanism in real time.

[0145] The user interface is used for user interaction and system configuration. Its specific functions include:

[0146] Provides an interface for users to manually set the control factor game.control_factor, allowing users to adjust the penalty intensity according to market conditions;

[0147] Provides a real-time monitoring data display interface to show key indicators such as market competition status, deadweight loss, degree of collusion, and penalty coefficient penalty_index, to assist users in making decisions and adjustments.

[0148] Reference Figure 2 The method for suppressing the spontaneous formation of monopoly phenomenon by the pricing algorithm of this embodiment includes the following main steps: model construction, strategy selection, profit calculation, deadweight loss calculation, penalty coefficient adjustment, Q value update and convergence check. The specific implementation of each step is as follows.

[0149] Step S10, model construction: Construct an AI agent pricing algorithm model based on reinforcement learning, define the pricing strategy space of multiple participants, set the market demand function and cost parameters. Through multi-agent game simulation, the pricing decision-making behavior of multiple participants in the market is ensured to ensure that the model can accurately reflect the actual market environment.

[0150] Step S20, strategy selection: At each time step t, the AI ​​agent selects a pricing strategy a through the exploration and exploitation mechanism according to the current state s and time step t. The exploration probability gradually decreases with the increase of time steps to balance the relationship between exploring new strategies and exploiting learned strategies.

[0151] Step S30, profit calculation: Calculate the profit π of each participant according to the selected pricing strategy a i Profit calculations take into account the impact of market demand and pricing strategies, ensuring that profits reflect the actual impact of pricing actions on the market.

[0152] Step S40, dead weight loss calculation: calculate the dead weight loss dwl under the current pricing behavior i and deadweight loss under perfect collusion pricing cartel . The impact of pricing behavior on social welfare is quantified by comparing demand differences and price differences between competitive pricing and algorithmic pricing and monopoly pricing.

[0153] Step S50, penalty coefficient adjustment: dynamically adjust the reward coefficient adjusted_π in the reinforcement learning algorithm according to the deadweight loss ratio penalty_index i By adjusting the reward mechanism and penalizing behaviors that deviate from competitive pricing, the risk of algorithms spontaneously forming a monopoly can be reduced.

[0154] Through this continuous adjustment method, the degree of deviation of the pricing strategy is proportional to the severity of the penalty, without setting a fixed threshold. The more the pricing strategy deviates from the competitive price level, the higher the penalty coefficient penalty_index, and the greater the reward reduction. Conversely, the reward reduction is relatively small.

[0155] Step S60, Q value update and convergence check: according to the adjusted reward coefficient adjusted_π i , update the Q value function and optimize the pricing strategy of the participants. At the same time, check whether the algorithm has reached the convergence conditions, such as the strategy no longer changes during the stable period, or the maximum number of iterations is reached, to ensure that the algorithm can run stably and achieve the expected goals.

[0156] Reference Figure 3 , the interaction process between the data processing module and the computing module in the system of this embodiment is as follows:

[0157] After receiving the market data, the data processing module performs preprocessing and standardization;

[0158] The preprocessed data is passed to the calculation module;

[0159] The calculation module calculates the dead weight loss and the collusion dead weight loss based on the received data;

[0160] The calculation results are fed back to the reinforcement learning module to dynamically adjust the reward mechanism.

[0161] Reference Figure 4 In this embodiment, the interaction process between the reinforcement learning module, the control module and the monitoring module is as follows:

[0162] The reinforcement learning module selects a pricing strategy based on the current Q value and calculates the profit;

[0163] The control module adjusts the penalty coefficient according to the control factor set by the user;

[0164] The monitoring module monitors market pricing behavior in real time and evaluates the state of competition;

[0165] The monitoring results are fed back to the reinforcement learning module and the control module to dynamically adjust the algorithm behavior.

[0166] Reference Figure 5 The detailed process of the real-time monitoring and adjustment strategy mechanism in this embodiment includes:

[0167] Real-time monitoring of pricing behavior: At each time step t, monitor the deviation between the current pricing strategy and the competitive price and monopoly price;

[0168] Calculate key indicators: calculate the deadweight loss ratio penalty_index and the collusion degree index collusion_index;

[0169] Dynamically adjust the reward mechanism: Based on the penalty coefficient penalty_index and the control factor, adjust the reward coefficient to guide the algorithm to adopt a more competitive pricing strategy;

[0170] Feedback adjustment results: Feedback the adjustment results to the reinforcement learning module to ensure that the algorithm continues to optimize the pricing strategy and avoid the formation of a monopoly.

[0171] Reference Figure 6 The detailed process of the competitive balance ensuring mechanism of multiple rounds of iterative updates in this embodiment includes:

[0172] Multi-round game iteration: During the training process of the reinforcement learning algorithm, multiple rounds of games are conducted, and each participant in each round of the game chooses a pricing strategy based on the current Q value;

[0173] Profit calculation and reward adjustment: After each round of game, the profit of each participant is calculated, and the reward function is adjusted according to the penalty coefficient penalty_index. Through the penalty mechanism, the algorithm is ensured to gradually optimize the pricing strategy to avoid the formation of monopoly pricing behavior.

[0174] Q value update and strategy optimization: Through continuous Q value update, the pricing strategies of all participants gradually tend to the market competition equilibrium state. Ensure that the pricing behavior of all participants remains within a reasonable competition range and does not deviate from the market competition rules.

[0175] Real-time monitoring and dynamic adjustment: In each round of iteration, the changes in pricing behavior in the market are monitored in real time, and the reward mechanism is adjusted dynamically based on the monitoring results. This prevents the algorithm from forming monopoly pricing during the training process and ensures the competitiveness and fairness of the market pricing strategy.

[0176] Final convergence: Through multiple rounds of iterations and dynamic adjustments, a stable pricing strategy is eventually achieved to ensure that the market maintains a healthy pricing level while ensuring fair competition and avoiding monopolistic behavior caused by over-optimization of the algorithm.

[0177] The suppression method uses multiple rounds of iterative training in a multi-agent game model, so that each participant can eventually form a relatively stable pricing strategy combination, avoiding long-term concentration of prices at a monopoly level, thereby protecting consumer interests and ensuring reasonable allocation of market resources.

[0178] This inhibition method combines the multi-agent game model with the Q-learning algorithm. By dynamically adjusting the penalty coefficient and real-time monitoring mechanism, it can effectively inhibit the pricing algorithm from forming monopolistic behavior during the autonomous learning process, promote market competitive pricing strategies, and protect consumer interests and market fairness.

[0179] By introducing a dynamic penalty coefficient based on the Harberger triangle and combining it with a real-time monitoring and adjustment mechanism, the algorithm can be effectively restrained from forming monopolistic behavior in a multi-agent game environment, ensuring the competitiveness and fairness of market pricing strategies.

[0180] Reference Figures 1 to 6 The system described in this embodiment realizes the function of suppressing the spontaneous formation of monopoly phenomenon by suppressing pricing algorithm based on reinforcement learning through close cooperation of various modules. The data processing module provides high-quality data foundation, the calculation module dynamically adjusts the penalty coefficient, the reinforcement learning module optimizes the pricing strategy, the control module and the monitoring module ensure the stable operation of the system, and the user interface provides convenient system configuration and monitoring means.

[0181] In this embodiment, the implementation steps of the system initialization and parameter setting part are as follows.

[0182] Step S70, system initialization: When the system starts, the user sets the initial parameters through the user interface, including the number of participants n, product differentiation parameter alpha, exploration parameter beta, discount factor delta, control factor game.control_factor, marginal cost c, market demand parameter a, price sensitivity parameters b and b1, and grid dimension k, etc.

[0183] Step S80, initializing the state and action space: Initialize the state space and action space of the system according to the set parameters. The state space sdim is defined as the combination of pricing strategies, and the action space A is defined as the set of pricing strategies that each participant can choose. All possible pricing actions are generated through the preset price range and step size to ensure that the algorithm can cover all possibilities of market pricing.

[0184] Step S90, calculate competition and monopoly prices: use the calculation module to solve the market competition equilibrium price p based on the defined demand function and marginal cost competitive and the monopoly equilibrium price p cartel . Through numerical solution methods, such as fsolve, the accuracy of price calculation is ensured, providing a basis for subsequent deadweight loss calculation.

[0185] Step S100, initializing the profit matrix and Q value function: Initialize the profit matrix PI and Q value function Q of each participant according to the calculated competition and monopoly prices. The profit matrix PI records the profit value of each participant under different pricing strategy combinations, and the Q value function Q is used to guide the participants to choose the optimal pricing strategy during the learning process.

[0186] In this embodiment, the steps for implementing the reinforcement learning and dynamic adjustment mechanism are as follows.

[0187] Step S110, strategy selection and action execution: At each time step t, the reinforcement learning module selects pricing strategy a through the exploration and utilization mechanism according to the current state s and time step t. The exploration probability gradually decreases with the increase of time steps to balance the relationship between exploring new strategies and utilizing learned strategies. The selected pricing strategy a is then executed, affecting market demand and profit calculation.

[0188] Step S120, profit and deadweight loss calculation: After executing pricing strategy a, calculate the profit π of each participant i , and calculate the corresponding dead weight loss dwl according to the current pricing strategy i and conspiracy to die heavy losses dwl cartel The calculation of deadweight loss reflects the impact of pricing strategies on social welfare, and collusion deadweight loss is used to quantify the degree of collusion in pricing strategies.

[0189] Step S130, penalty coefficient penalty_index adjustment:

[0190] Calculate the dead weight loss ratio, the specific formula is:

[0191]

[0192] Dynamically adjust the reward coefficient adjusted_π i By adjusting the reward mechanism and penalizing behaviors that deviate from competitive pricing, the risk of algorithms spontaneously forming a monopoly can be reduced.

[0193] Calculate the adjustment factor adjustment_factor to ensure that the penalty is proportional to the degree of pricing deviation. The specific formula is:

[0194] adjustment_factor=1-game.control_factor×penalty_index[n];

[0195] Step S140, Q value function update: according to the adjusted reward coefficient adjusted_π i , update the Q-value function in the reinforcement learning algorithm. Through the reinforcement learning algorithm, the pricing strategy of the participants is gradually optimized to maximize the cumulative reward.

[0196] Step S150, convergence check and iteration termination: After each round of iteration, check whether the algorithm has reached the convergence condition. If the pricing strategy remains stable for multiple consecutive time steps, or reaches the preset maximum number of iterations, the algorithm converges and stops iterating. Otherwise, continue to the next round of iteration to ensure that the algorithm can find a stable pricing strategy combination.

[0197] In this embodiment, the implementation steps of the real-time monitoring and feedback adjustment part are as follows.

[0198] Step S160: Real-time monitoring of pricing behavior: At each time step t, monitor the current pricing strategy and the competitive price p competitive and the monopoly price p cartel By calculating and updating the dead weight loss dwl i and conspiracy to die heavy losses dwl cartel , evaluate the impact of the algorithm's current strategy on the market. The specific formula is:

[0199] dwl i =0.5×(d competitive -d i )×(p i -p competitive );

[0200] dwl cartel =0.5×(dcompetitive -d cartel )×(p cartel -p competitive );

[0201] Step S170: Calculate the collusion index collusion_index, the specific formula is:

[0202]

[0203] Based on current profit π i , competitive profit π competitive and monopoly profit π cartel , calculate the collusion degree index collusion_index, which is used to quantify the degree of collusion between participants and ensure that pricing behavior does not deviate from market competition rules. The specific formula is:

[0204]

[0205] Step S180, cumulative collusion index cumulative_collusion_index, the specific formula is:

[0206]

[0207] The cumulative collusion index cumulative_collusion_index is specifically the geometric mean of the cumulative average collusion index of all algorithm entities in the market, which is used to comprehensively evaluate the collusion tendency of all algorithms in the entire decision-making cycle; if the average collusion index of algorithm entity 1 is 0.2, and the average collusion index of algorithm collusion entity 2 is 0.8, then the cumulative collusion index of the two is 0.4;

[0208] Among them, at each time step t, the average degree of collusion mean_t_collusion_index is calculated to comprehensively evaluate the collusion tendency of a single algorithm in the entire decision-making cycle. The specific formula is:

[0209]

[0210] Step S190, dynamically adjust the reward mechanism: adjust the reward mechanism according to the cumulative collusion index cumulative_collusion_index. Calculate adjusted_π by combining the penalty coefficient penalty_index with the control factor game.control_factor i , ensuring effective punishment for collusive behavior and promoting the formation of competitive pricing strategies. The specific formula is:

[0211] adjusted_πi =π i ×(1-game.control_factor×penalty_index[n]);

[0212] Specifically, the reward mechanism is adjusted by combining the cumulative collusion index cumulative_collusion_index with the penalty coefficient penalty_index; at the end of each training session, the size of game.control_factor is controlled according to the size of the cumulative collusion index cumulative_collusion_index, so that

[0213] adjusted_π i =π i ×(1-game.control_factor×penalty_index[n]) affects the degree of effect of penalty index penalty_index; its specific implementation method is that the current cumulative collusion degree index

[0214] The higher the cumulative_collusion_index is, the greater the game.control_factor will be in the next period, which will make the penalty_index in the next period punish collusion to a greater extent, thus influencing and adjusting the reward mechanism through indirect combination.

[0215] Step S200: Feedback the adjustment result: adjust the reward coefficient adjusted_π i Feedback is given to the reinforcement learning module to guide participants to optimize pricing strategies in the next iteration. Through real-time feedback and dynamic adjustments, the algorithm is ensured to continuously suppress monopolistic behavior during the training process and maintain fair competition in the market.

[0216] In this embodiment, the steps for implementing the multi-round iteration and competitive balance ensuring part are as follows.

[0217] Step S210, multiple rounds of game iteration: During the training process of the reinforcement learning algorithm, multiple rounds of game iteration are performed. In each round of the game, each participant selects a pricing strategy based on the current Q value and executes the strategy to influence market demand and profit.

[0218] Step S220, profit calculation and reward adjustment: After each round of game, calculate the profit π of each participant i , and adjust the reward function according to the penalty coefficient penalty_index. Through the penalty mechanism, the algorithm is ensured to gradually optimize the pricing strategy and avoid the formation of monopoly pricing behavior.

[0219] Step S230, continuous Q value update: Through continuous Q value update, the pricing strategy of each participant gradually tends to the market competition equilibrium state, ensuring that the pricing behavior of all participants remains within a reasonable competition range and does not deviate from the market competition rules.

[0220] Step S240, real-time monitoring and dynamic adjustment: In each round of iteration, real-time monitoring of changes in pricing behavior in the market is performed, and the reward mechanism is dynamically adjusted according to the monitoring results. This prevents the algorithm from forming monopoly pricing during the training process and ensures the competitiveness and fairness of the market pricing strategy.

[0221] Through continuous iteration and Q-value updating, the pricing strategy of the AI ​​agent gradually tends to be consistent with the market competition equilibrium state, thereby suppressing the spontaneously formed monopoly pricing behavior in real time during the autonomous learning process of reinforcement learning;

[0222] The Q value represents the expected future reward of the agent taking a certain action in a certain state during the reinforcement learning process. The reinforcement learning algorithm helps the agent learn the optimal strategy by continuously updating these Q values, so as to make better decisions in uncertain environments.

[0223] Step S250, final convergence and stabilization: Through multiple rounds of iterations and dynamic adjustments, a stable pricing strategy combination is finally achieved, ensuring that the market maintains a healthy pricing level under the premise of fair competition and avoiding monopoly behavior caused by over-optimization of the algorithm.

[0224] In this embodiment, the implementation steps of the pricing control part applicable to the price-sensitive market are as follows.

[0225] Step S260, define price sensitive parameters: define price sensitive parameters b and b1 according to the price sensitivity of the applicable market. Parameter b reflects the direct impact of price on demand, and parameter b1 reflects the indirect impact of price difference on demand.

[0226] Step S270, dynamically adjust the market demand function: In the multi-agent game model, consider the impact of price sensitivity on market demand. Dynamically adjust the market demand through the function demand(p) to obtain the demand d of each agent algorithm. The market demand function calculates the market demand of each participant under different pricing according to the pricing strategy and price sensitivity parameters of the participants. The specific formula is:

[0227] d=ab×p+b 1 ×p 1 ;

[0228] Step S280, price manipulation behavior penalty: by introducing the penalty coefficient penalty_index, price manipulation behavior is punished in the reinforcement learning algorithm, reducing the motivation of participants to adopt excessively high pricing strategies and promoting the formation of competitive pricing strategies.

[0229] Step S290, real-time monitoring and profit impact analysis: During the algorithm training process, the impact of price changes on market demand and profits is monitored in real time. By calculating changes in market demand, profit fluctuations and their correlation with pricing strategies, the algorithm is ensured to optimize profits while not deviating from market competition rules.

[0230] Step S300, dynamically optimize pricing strategy: By dynamically adjusting the reward mechanism and penalty coefficient, the algorithm is encouraged to maintain competitive pricing in the market with high price sensitivity. This prevents price manipulation and monopoly, ensures the rational allocation of market resources and maximizes consumer benefits. In the market with high price sensitivity, based on the real-time feedback of dwl i and dwl cartel Changes in the penalty coefficient penalty_index are more sensitive, allowing the AI ​​agent to more quickly reduce the probability of the emergence of monopoly-oriented pricing strategies in markets with higher price elasticity.

[0231] In this embodiment, the steps for implementing part of the market fairness protection mechanism of the collusion index and the penalty coefficient are as follows.

[0232] Step S310: Calculate the collusion index to evaluate the degree of collusion among participants. This index is used to quantify the collusion tendency of pricing strategies and ensure that pricing behavior does not damage market fairness. The specific formula is:

[0233]

[0234] Step S320: Calculate the cumulative collusion index cumulative_collusion_index through the formula to comprehensively evaluate the collusion tendency of the overall market.

[0235] Step S330: observe the overall performance of the agent algorithm according to the cumulative collusion index cumulative_collusion_index. The specific formula is:

[0236]

[0237] Step S340, real-time adjustment of reward mechanism: In the Q learning update process, the cumulative collusion index is combined with the penalty coefficient to adjust the reward mechanism in real time. Ensure that the pricing decision of each participant does not lead to excessive concentration of market prices and maintain the diversity and fairness of market pricing behavior. Use the cumulative collusion degree, and control the size of game.control_factor according to the size of the cumulative collusion degree index cumulative_collusion_index at the end of each training period, so as to

[0238] adjusted_π i =π i ×(1-game.control_factor×penalty_index[n]) affects the degree of effect of the penalty index penalty_index; specifically, the higher the cumulative collusion index of the current period, the larger the game.control_factor of the next period, which makes the penalty_index of the next period more likely to affect profits, so as to adjust the reward mechanism.

[0239] Through the above specific implementation methods, the present invention provides a method and system for suppressing the spontaneous formation of monopoly by pricing algorithms based on reinforcement learning. The system effectively suppresses the formation of monopoly behavior by algorithms in the process of autonomous learning through the collaborative work of multiple modules, combined with dynamic adjustment mechanisms and real-time monitoring methods, promotes fair competition in the market and rational allocation of resources, and protects the interests of consumers.

[0240] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0241] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0242] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0243] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0244] The present invention has been disclosed above with preferred embodiments, but they are not intended to limit the present invention. Any technical solutions obtained by adopting equivalent replacement or equivalent transformation solutions fall within the protection scope of the present invention.

Claims

1. A method for suppressing spontaneous monopoly of pricing algorithms, characterized in that: The method comprises: Obtain market-related data; Based on the data, the dead weight loss dwl is calculated and obtained. i With complete collusion dead heavy loss dwl cartel , and based on this, the deadweight loss ratio used to quantify the degree of monopoly deviation is calculated and defined as the penalty coefficient. The penalty coefficient penalty_index is introduced into the reward update process of the reinforcement learning algorithm, and the actual reward of the current pricing strategy is dynamically adjusted according to the size of the penalty coefficient penalty_index, so as to form a dynamic penalty mechanism, so that when the strategy tends to be monopolistic, it will be punished more strongly, and when the strategy tends to be more competitive, the punishment will be relatively weakened; Through continuous iteration and Q-value updating, the pricing strategy of the AI ​​agent gradually tends to be in line with the market competition equilibrium state.

2. The method according to claim 1, characterized in that: The dynamic penalty mechanism includes: i and dwl cartel Calculate the penalty coefficient penalty_index; Use the control factor game.control_factor to weight the penalty coefficient penalty_index according to adjusted_π i =π i ×(1-game.control_factor×penalty_index[n]) formula for current profit π i Adjust the reward and get the dynamically adjusted reward coefficient adjusted_π i ;.

3. The method according to claim 2, characterized in that The reinforcement learning algorithm recalculates and stores the Q value of each participant in each round of price strategy update: The original profit π obtained by the participant based on state s and action a i Adjust to generate adjusted_π i ; Adjusted_π is updated by the Q value formula i Taking the Q function into account, the AI ​​agent tends to choose a price strategy that is less punishing and more competitive in subsequent strategy selection; Continuously iterate to achieve dynamic optimization and stabilization of pricing strategies.

4. The method according to claim 1 or 2, characterized in that: The dead weight loss dwl i With complete collusion dead heavy loss dwl cartel , calculated by the following method: Calculate the deadweight loss caused by algorithmic pricing behavior in the market i , that is, by comparing the demand differences and price differences generated by competitive prices and algorithmic pricing, the social welfare losses caused by the pricing behavior of participants are calculated. The calculation formula is: stupid i =0.5×(d competitive -d i )×(p i - p competitive ); Among them, p competitive is the optimal price in a perfectly competitive market, d competitive is the algorithm demand in a perfectly competitive market, p i is the price determined by the algorithm under the current pricing behavior, d i The algorithmic demands faced by the algorithm under current pricing behavior; The complete collusion dead weight loss dwl cartel , calculated by the following method: Calculate the deadweight loss dwl for complete collusion cartel , that is, by calculating the demand difference and price difference between the competitive price and the monopoly price, the welfare loss caused by collusive pricing is determined. The calculation formula is: stupid cartel =0.5×(d competitive -d cartel )×(p cartel - p competitive ); Among them, p competitive is the optimal price in a perfectly competitive market, d competitive is the algorithm demand in a perfectly competitive market, p cartel is the optimal price in a monopoly market, d cartel For the demand in a completely monopolistic market.

5. The method according to claim 1, characterized in that The suppression method also includes calculating the collusion degree index collusion_index to evaluate the degree of collusion among participants. This index is used to quantify the collusion tendency of pricing strategies and ensure that pricing behavior does not damage market fairness. The specific formula is: where π i is the actual profit obtained by the algorithm and the current pricing behavior, π competitive is the profit of the agent algorithm in a fully competitive environment, π cartel The profit of the agent algorithm in a completely monopolistic environment; At each time step t, the sum of the collusion degrees of each period before is calculated, sum_t_collusion_index. The specific formula is: And calculate the average collusion degree mean_t_collusion_index of a single agent algorithm. The specific formula is: Where t is the current cycle number of the agent algorithm action; Calculate the cumulative collusion index through the formula cumulative_collusion_index; According to the cumulative collusion index, cumulative_collusion_index, the overall performance of the agent algorithm is observed. The specific formula is: Among them, game.n is the number of participants; Adjust the reward mechanism in real time. During the Q learning update process, the cumulative collusion degree index is combined with the penalty coefficient to adjust the reward mechanism in real time.

6. The method according to claim 1, characterized in that The data includes price sensitivity parameters, number of participants, cost parameters and historical pricing information.

7. A system for suppressing spontaneous monopoly of pricing algorithms, characterized in that: include A data processing module for receiving and preprocessing market data, including historical pricing, market demand parameters, price sensitivity, and marginal costs; Calculation module, used to calculate competition and monopoly prices, as well as deadweight loss dwl based on data i With complete collusion dead heavy loss dwl cartel , and then get the penalty coefficient penalty_index; Reinforcement learning module, used to execute Q learning algorithm, according to the penalty coefficient penalty_index profit π i Perform dynamic penalty adjustments so that the Q value update reflects the degree of pricing deviation; The control module is used to set and maintain game.control_factor, so that the penalty intensity can be continuously adjusted linearly according to the user-preset control factor; Monitoring module, used to observe price changes in real time and send dwl i 、dwl cartel The calculation results of the penalty coefficient penalty_index are fed back to the reinforcement learning module and the control module to realize a threshold-free and continuous penalty adjustment process.

8. The system according to claim 7, characterized in that The reinforcement learning module calculates the profit π i and the dynamically adjusted reward coefficient adjusted_π i , update the Q-value function in the reinforcement learning algorithm.

9. A computer storage medium, characterized in that: The computer storage medium stores a computer program, which implements the steps of the method according to any one of claims 1 to 6 when the computer program is executed by a processor.

10. A computer device, characterized in that: Comprising the computer storage medium of claim 9.