Game theory-based e-commerce live broadcast supply chain income optimization calculation method and system
By using a game theory-based mathematical model and reinforcement learning to train an intelligent agent, the dynamic optimization problem of supply chain management in e-commerce live streaming was solved, optimizing the decisions of streamers and manufacturers, reducing customer churn, and improving market reputation and profits.
Patent Information
- Application Number
- CN202510943509.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-11-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional supply chain management systems struggle to accurately balance the interests of livestreamers, manufacturers, and consumers. In particular, the problems of discrepancies between goods and descriptions and exaggerated quality issues in e-commerce livestreaming, which lead to customer churn and negative impacts on the market ecosystem, are difficult to resolve effectively.
By establishing a game theory-based mathematical model, the interaction between product quality and livestreamer promotion strategies is quantified, a supply chain model is constructed, and reinforcement learning is used to train livestreamer and manufacturer agents to optimize decision support.
It enables dynamic supply chain management for products of different qualities and promoted by livestreamers, optimizing the profits of manufacturers and livestreamers, reducing customer churn, and enhancing market reputation and competitiveness.
Smart Images

Figure CN120892902A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of live broadcast supply chain management, in particular to an e-commerce live broadcast supply chain revenue optimization calculation method based on game theory, and further relates to a system adopting the e-commerce live broadcast supply chain revenue optimization calculation method based on game theory. BACKGROUND
[0002] As a new sales model, although live broadcast sales improves the efficiency of commodity sales, it also exposes some problems in the development process, such as mismatching, exaggerating quality and using inferior goods as good ones, etc. When the information spread by the anchor and the actual quality of the product have a large deviation, the purchase decision of the consumer may be misled, at this time, the quality problem not only weakens the trust of the consumer to the product, but also may lead to the disappointment of the consumer to the anchor, which will cause losses to both the anchor and the manufacturer. When the manufacturer pursues higher profits, its strategy often depends on the valuation of the product by the consumer, and such valuation is largely influenced by the promotion of the anchor. In the long run, the loss of consumers due to quality problems may cause a chain reaction to the entire market ecology, affecting the reputation and income of the anchor, and also weakening the market competitiveness of the manufacturer. In addition, the change of consumer surplus is closely related to the information dissemination strategy of the anchor. When the anchor improves the promotion, the willingness to pay of the consumer may rise; however, if the product quality fails to match the promotion level, the actual satisfaction of the consumer will decrease, leading to the fall of the willingness to pay.
[0003] Therefore, this dynamic game relationship involves the interaction of the anchor, the manufacturer and the consumer, so that the traditional supply chain management system is difficult to accurately and balance the interests of all parties, and cannot realize dynamic e-commerce live broadcast supply chain management for various product qualities and different anchor promotions. SUMMARY
[0004] The technical problem to be solved by the present application is to provide an e-commerce live broadcast supply chain revenue optimization calculation method based on game theory, aiming to quantify the interactive influence of product quality and anchor promotion when different strategies are adopted by establishing a mathematical model, and to train anchor agents and manufacturer agents through reinforcement learning, so as to provide a technical basis for supply chain management for various product qualities and different anchor promotions, and to provide optimization decision support for anchors and manufacturers and other members.
[0005] To this end, the present application provides an e-commerce live broadcast supply chain revenue optimization calculation method based on game theory, comprising the following steps: Step S1, obtaining historical transaction data through the API interface of the e-commerce platform, wherein the historical transaction data includes product actual quality parameters, anchor behavior data and consumer evaluation data; Step S2: Construct a supply chain model based on the relationships between manufacturers, livestreamers, and consumers. Cluster the quality of products produced by manufacturers into high-quality or low-quality products. When a product is low-quality, the strategies in the supply chain model include strategy LN (not revealing the true low quality), strategy LL (revealing the true low quality), and strategy LH (exaggerating product quality). When a product is high-quality, the strategies in the supply chain model include strategy HN (not revealing the true high quality) and strategy HH (revealing the true high quality). Set adjustment factors. Used to indicate changes in fan base; among which... For the actual quality of the product, Let be the consumer's expectation of product quality, and be the sensitivity factor. Step S3: Obtain the strategy through the numerical computation library. The equilibrium price under and strive for balanced streamer ,in, This indicates the manufacturer's production quality strategy. This indicates the anchor's quality information promotion strategy; Step S4, based on backward induction and the equilibrium price... and strive for balanced streamer Acquisition Strategy Manufacturer equilibrium profit Balance profits with streamers ; Step S5: Based on the output data from steps S3 and S4, train the anchor agent and the manufacturer agent through reinforcement learning. Step S6: For low-quality and high-quality products, output regional maps of product quality and entertainment time based on the profit when the anchor adopts different strategies.
[0006] A further improvement of the present invention is that, in step S1, the streamer behavior data includes promotional quality, entertainment interaction duration, and historical fan churn rate; the consumer evaluation data includes willingness to pay and return rate. When constructing the supply chain model in step S2, the formula is used. The probability that a broadcaster delivers accurate information is expressed by the formula. This represents the probability that the broadcaster conveys inaccurate information, where, and These respectively indicate the status of the information conveyed by the broadcaster regarding high-quality and low-quality products; and These respectively represent the actual quality of high-quality products and low-quality products; This indicates the information conveyed by the broadcaster. , representing the ability of the anchor to convey the product quality information, t represents the entertainment time, representing the time used by the anchor to introduce the product; In the numerical calculation library, the relationship table corresponding to the equilibrium price of each strategy , the equilibrium anchor effort , the manufacturer equilibrium profit and the anchor equilibrium profit is stored in advance; when the current strategy is selected, the analytical solution corresponding to the strategy is obtained by calling and querying the relationship table.
[0007] The further improvement of the present application is that when the strategy uses low-quality products and the strategy LN of not revealing the true low quality is adopted, in the step S3, the equilibrium price of the strategy LN is calculated by the formula , and the equilibrium anchor effort of the strategy LN is calculated by the formula , wherein, represents the probability of the manufacturer producing high-quality products, represents the product quality, represents the commission rate, and t represents the entertainment time; in the step S4, the manufacturer equilibrium profit of the strategy LN is calculated by the formula , and the anchor equilibrium profit of the strategy LN is calculated by the formula .
[0008] The further improvement of the present application is that when the strategy uses low-quality products and the strategy LL of revealing the true low quality is adopted, in the step S3, the equilibrium price of the strategy LL is calculated by the formula , and the equilibrium anchor effort of the strategy LL is calculated by the formula ; in the step S4, the manufacturer equilibrium profit of the strategy LL is calculated by the formula , and the anchor equilibrium profit of the strategy LL is calculated by the formula ; wherein, represents the probability of the manufacturer producing high-quality products, t represents the entertainment time, represents the ability of the anchor to convey the product quality information, represents the product quality, represents the commission rate; , and These are intermediate calculation parameters. , , .
[0009] A further improvement of the present invention is that, when the strategy When a strategy LH that uses low-quality products and exaggerates product quality is employed, in step S3, the formula is used... Calculate the equilibrium price of strategy LH And through the formula Equilibrium of Calculation Strategy LH In step S4, the formula is used. Calculate the manufacturer's equilibrium profit for strategy LH And through the formula Calculate the anchor equilibrium profit of strategy LH ; in, Let t represent the probability that a manufacturer produces a high-quality product, and let t represent the entertainment time. This indicates the anchor's ability to convey product quality information. Indicates product quality. Indicates the commission rate; , and These are intermediate calculation parameters. , , .
[0010] A further improvement of the present invention is that, when the strategy When using high-quality products and not disclosing the true high-quality strategy HN, in step S3, the formula is used... Calculate the equilibrium price of strategy HN And through the formula Equilibrium broadcaster effort of calculation strategy HN In step S4, the formula is used. Calculate the manufacturer's equilibrium profit for strategy HN And through the formula Calculate the anchor equilibrium profit of strategy HN ; in, This indicates the probability that a manufacturer will produce high-quality products. Indicates product quality. represents the commission rate, and t represents the entertainment time; , and are intermediate calculation parameters, respectively, , , .
[0011] A further improvement of the present application is that when the strategy is a high-quality product, and the strategy HH of revealing the real high-quality strategy is adopted, in the step S3, the equilibrium price of the strategy HN is calculated by the formula , and the equilibrium anchor effort of the strategy HN is calculated by the formula ; in the step S4, the manufacturer equilibrium profit of the strategy HN is calculated by the formula , and the anchor equilibrium profit of the strategy HN is calculated by the formula ; ; wherein, represents the probability of the manufacturer producing high-quality products, represents the product quality, represents the commission rate, and t represents the entertainment time, represents the ability of the anchor to deliver product quality information; , , and are intermediate calculation parameters, respectively, , , , .
[0012] A further improvement of the present application is that in the step S5, the anchor agent includes a first state space, a first action space and a first reward function, the first state space includes product quality, information delivery accuracy, entertainment time, last period demand and fan change amount; the first action space includes the strategy LN, the strategy LL, the strategy LH, the strategy HN and the strategy HH; the manufacturer agent includes a second state space, a second action space and a second reward function, the second state space includes product quality, high-quality probability, entertainment time and anchor strategy, and the second action space includes high-quality products and low-quality products; In the training process of the step S5, the anchor agent and the manufacturer agent are trained by the DQN algorithm of reinforcement learning, and the training process includes the following sub-steps: Step S501, using the output data of step S3 and step S4 as the basis data for training, input into the DQN network of the anchor agent and the manufacturer agent; Step S502, the anchor agent selects a strategy from the first action space according to the current state, and the manufacturer agent selects a corresponding high-quality product or low-quality product from the second action space, and then calls the numerical calculation library for calculation to generate a new state and a reward; Step S503, updating the DQN network according to the new state and the reward; Step S504, through iterative training, the strategy of the anchor agent and the manufacturer agent is constantly optimized, and gradually converges to the optimal strategy.
[0013] Further improvement of the application is that the step S6 includes the following sub-steps: Step S601, for low-quality products, the profit boundary analytical solution of each strategy is calculated according to the profit of the anchor when adopting different strategies; Step S602, according to the profit boundary analytical solution of the low-quality products in different strategies calculated in step S601, a first area graph of product quality and entertainment time is drawn; Step S603, for high-quality products, the profit boundary analytical solution of each strategy is calculated according to the profit of the anchor when adopting different strategies; Step S604, according to the profit boundary analytical solution of the high-quality products in different strategies calculated in step S603, a second area graph of product quality and entertainment time is drawn; Step S605, comparing and integrating the profit boundary analytical solution of the low-quality products in different strategies with the profit boundary analytical solution of the high-quality products in different strategies, generating the final comprehensive strategy profit analytical solution and area graph.
[0014] The application also provides an e-commerce live broadcast supply chain revenue optimization calculation system based on game theory, which adopts the e-commerce live broadcast supply chain revenue optimization calculation method based on game theory as described above, and includes: A historical transaction data acquisition module acquires historical transaction data through an API interface of an e-commerce platform; A supply chain model construction module constructs a supply chain model through the association among manufacturers, anchors and consumers; A price acquisition module acquires the equilibrium price and the equilibrium anchor effort under the strategy through a numerical calculation library. ; The profit-generating module, based on backward induction and the equilibrium price, and strive for balanced streamer Acquisition Strategy Manufacturer equilibrium profit Balance profits with streamers ; The agent training module trains the broadcaster agent and the manufacturer agent through reinforcement learning based on the output data of the price acquisition module and the profit acquisition module. The region map drawing module outputs region maps of product quality and entertainment time for both low-quality and high-quality products, based on the profit generated when the anchor adopts different strategies.
[0015] Compared with existing technologies, the advantages of this invention are as follows: First, historical transaction data is obtained through the API interface of e-commerce platforms; then, a supply chain model is constructed through the relationship between manufacturers, livestreamers, and consumers; and finally, strategies are obtained through a numerical calculation library. The equilibrium price under and strive for balanced streamer And by using backward induction, the manufacturer's equilibrium profit is obtained. Balance profits with streamers Finally, by training the anchor agent and manufacturer agent through reinforcement learning, and outputting regional maps of product quality and entertainment time based on the profits when the anchor adopts different strategies, this invention can quantify the interaction between product quality and different anchor promotion strategies by establishing a mathematical model, and provide a technical foundation for supply chain management to cope with various product qualities and different anchor promotions through reinforcement learning training of the anchor agent and manufacturer agent, and provide optimization decision support for anchors, manufacturers and other members. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the workflow of one embodiment of the present invention. Detailed Implementation
[0017] In the description of this invention, the term "several" means one or more; the term "multiple" means two or more; the terms "greater than," "less than," and "exceeding" are all understood to exclude the stated number; and the terms "above," "below," and "within" are all understood to include the stated number. The terms "first," "second," etc., are understood to be used only to distinguish identical or similar technical feature names, and should not be construed as implying / indicating the relative importance of the technical features, the number of technical features, or the sequential relationship between the technical features.
[0018] The preferred embodiments of the present application will be further described in details below with reference to the accompanying drawings.
[0019] As Figure 1 shown, the embodiment provides a game theory-based e-commerce live broadcast supply chain revenue optimization calculation method, including the following steps: Step S1, obtaining historical transaction data through the API (application programming interface) of the e-commerce platform, wherein the historical transaction data includes product actual quality parameters, anchor behavior data and consumer evaluation data; Step S2, constructing a supply chain model through the association among manufacturers, anchors and consumers, clustering the product quality produced by the manufacturer as high-quality products or low-quality products, when the product is a low-quality product, the strategies in the supply chain model include a strategy LN of not revealing the true low quality, a strategy LL of revealing the true low quality, and a strategy LH of exaggerating the product quality; when the product is a high-quality product, the strategies in the supply chain model include a strategy HN of not revealing the true high quality and a strategy HH of revealing the true high quality; setting an adjustment factor for representing the change of fans; wherein, is the actual quality of the product, is the expectation of the consumer for the product quality, is a sensitivity factor; Step S3, obtaining the equilibrium price and the equilibrium anchor effort under the strategy through a numerical calculation library, wherein, represents the production quality strategy of the manufacturer, represents the quality information promotion strategy of the anchor; Step S4, obtaining the manufacturer equilibrium profit and the anchor equilibrium profit under the strategy according to the backward induction method and based on the equilibrium price and the equilibrium anchor effort ; Step S5, training the anchor agent and the manufacturer agent through reinforcement learning according to the output data of Step S3 and Step S4; Step S6, outputting a region graph of product quality and entertainment time according to the profits of the anchor when adopting different strategies for low-quality products and high-quality products, respectively.
[0020] The equilibrium price in the embodiment refers to the optimal price based on the game theory under the strategy ; and the equilibrium anchor effort refers to the optimal effort of the anchor under the strategy The optimal anchor effort parameter based on game theory is used as the optimal effort level invested in the equilibrium state; the manufacturer equilibrium profit refers to the strategy The optimal profit of the manufacturer based on game theory; the anchor equilibrium profit refers to the strategy The optimal profit of the anchor based on game theory.
[0021] In this embodiment, first, in step S1, the historical transaction data is obtained through the API interface of the e-commerce platform; then, the supply chain model is constructed through the association between the manufacturer, the anchor and the consumer, and when constructing the supply chain model, the product quality produced by the manufacturer is clustered into high-quality products or low-quality products, and five different strategies are proposed accordingly; then, the equilibrium price and the equilibrium anchor effort under the strategy are obtained through the numerical calculation library, focusing on how the anchor calculates the optimal information transmission mode and allocates entertainment time when facing different quality products in the e-commerce live broadcast scene, and analyzing its impact on various aspects of the supply chain, including manufacturer production and consumer welfare, etc., and then, in step S3, a supply chain model based on game theory is established; and in step S4, according to the backward induction method, the equilibrium price and the equilibrium anchor effort are obtained to obtain the manufacturer equilibrium profit and the anchor equilibrium profit ; finally, the anchor agent and the manufacturer agent are trained through reinforcement learning, and the region graph of product quality and entertainment time is output according to the profit of the anchor when adopting different strategies.
[0022] For the decision-making interaction mechanism of the manufacturer and the anchor, in the live e-commerce supply chain, the product quality decision (high / low quality) of the manufacturer and the information transmission mode (such as real description or exaggerated promotion) of the anchor exist complex interaction. The supply chain model constructed in this embodiment refers to constructing a stackelberg game model, analyzing how the manufacturer adjusts the production decision based on the anchor behavior, and how the anchor selects the information transmission mode according to the principle of maximizing the profit. By solving the equilibrium solution, through mathematical modeling and iterative optimization, the mutual influence law of the decisions of both parties is revealed, and the profit changes under different decision combinations are quantified.
[0023] For the dynamic influence of consumer feedback on the anchor, the consumer will adjust the purchase behavior according to the difference between the actual product quality and the anchor description. If the anchor transmits information that deviates greatly from the true quality of the product, it may lead to a decline in consumer satisfaction, and then cause user loss. In the modeling process of the embodiment, a consumer behavior response model is established to quantify the direct loss of the anchor's short-term revenue caused by user loss, and analyze the potential impact on the anchor's long-term reputation and market share. Further, combined with dynamic game simulation, the long-term revenue changes of the anchor under different information transmission modes are evaluated.
[0024] Therefore, the embodiment can establish a mathematical model based on game theory, thereby quantifying the interactive influence of product quality and anchor promotion strategies, and train anchor agents and manufacturer agents through reinforcement learning, providing a technical basis for supply chain management in response to various product qualities and different anchor promotions, and providing optimization decision support for anchors and manufacturers and other members.
[0025] In step S1 of the embodiment, the anchor behavior data includes promotion quality, entertainment interaction time, and historical fan loss rate; and the consumer evaluation data includes willingness to pay and return rate. After obtaining the historical transaction data, the missing values and abnormal values are preferably processed by a data cleaning tool (Python Pandas library), and then normalized and stored in a preset database.
[0026] In the embodiment, the quality of low-quality products is uniformly represented as L; represents the quality of the product, and the quality of high-quality products can be represented as the quality of the product . represents the probability that the manufacturer produces high-quality products, that is, the probability that the consumer considers the product quality to be high; represents the probability that the manufacturer produces low-quality products, that is, the probability that the consumer considers the product quality to be low. represents the commission rate; t represents the entertainment time, including the interaction between the anchor and the fans, the time for entertainment or showing one's own talent; represents the time used by the anchor to introduce the product.
[0027] When constructing the supply chain model in step S2, the anchor transmits accurate information through the formula represents the probability that the anchor transmits accurate information, through the formula represents the probability that the anchor transmits inaccurate information, wherein and respectively represent the information state of the anchor transmitting information about high-quality products and low-quality products; and respectively represent the actual quality of high-quality products and low-quality products; represents the information transmitted by the anchor, , represents the ability of the anchor to convey product quality information. Specifically, p : represents the probability that the anchor conveys accurate high-quality information in the case of high-quality products; p : represents the probability that the anchor conveys accurate low-quality information in the case of low-quality products; p : represents the probability that the anchor conveys inaccurate low-quality information in the case of high-quality products; p : represents the probability that the anchor conveys inaccurate high-quality information in the case of low-quality products.
[0028] The accuracy of the anchor's delivery of product information is positively related to the anchor's own ability and negatively related to the proportion of entertainment time allocated. Among them, the higher the ability (such as rich professional knowledge and clear expression), the higher the information accuracy; if the proportion of entertainment time increases, it will squeeze out the product introduction time (denoted as ), and the information accuracy will decrease, for example, excessive entertainment may lead to insufficient understanding of consumers on the key characteristics of the product.
[0029] From this, we can deduce . The quality probability of Bayesian update depends on the information state and the information state : ; ; ; .
[0030] The strategy choices of manufacturers and anchors are largely determined by the consumer's pre-expected value of the anchor's service or information minus the actual value after the fact. Therefore, this embodiment takes into account that the anchor's expression of product information during live streaming will have a significant impact on the consumer's stickiness. The anchor's false description of product quality may lead to a decline in consumer trust, thereby triggering fan loss. Conversely, when the anchor chooses not to disclose product quality and the consumer ultimately receives a high-quality product, this unexpected positive experience may enhance the consumer's favorability and trust, further promoting fan growth and stickiness. However, if the product itself is of low quality and the anchor chooses not to disclose it, it may also lead to fan loss / user loss, which is less impactful than exaggerating product quality, and the sensitivity factor controls this situation. The sensitivity factor which is a pre-set sensitivity parameter to control the impact of the difference between actual quality and expected value on demand. Correspondingly, the embodiment also represents the change of fans by normalizing the difference between expected value and actual quality, and sets an adjustment factor .
[0031] When the manufacturer produces low-quality products and the anchor does not disclose to the consumers, , the consumer updates the posterior expectation . The embodiment normalizes the difference between the consumer's expectation and the actual quality of the received product. The anchor's concealment of quality has a smaller impact on fan change than false exaggeration of product quality, so the sensitivity factor .
[0032] The fan change rate is represented by , where the fan change rate under the LN strategy is: . The consumer utility is represented by .
[0033] Let , the consumer purchase rate is , and the demand is .
[0034] The manufacturer's profit is represented as: ; the optimal solution of the manufacturer's profit , i.e. the manufacturer's equilibrium profit , will be introduced later. The anchor's profit is represented as: , and the optimal solution of the anchor's profit , i.e. the anchor's equilibrium profit , will be introduced later.
[0035] When the manufacturer produces low-quality products and the anchor discloses to the consumers, the consumer updates the posterior expectation , .
[0036] Let , we get , At this time, the fans who receive the expected product do not leave, but the anchor's information transmission accuracy needs to be considered: ; .
[0037] When the manufacturer produces low-quality products but the anchor discloses to the consumers that the quality is high, the consumers may leave after receiving the product due to the gap between the anchor's raised product quality and the actual quality, and the impact is greater than when the anchor does not disclose the quality information, so the sensitivity factor , .
[0038] ; ; .
[0039] Let , ; ; .
[0040] When the manufacturer produces high-quality products and the anchor does not disclose to the consumer, the consumer's updated posterior expectation is the same as the prior expectation, , when the consumer gets a high-quality product higher than his own expectation, it may make the anchor popular, . ; ; .
[0041] When the manufacturer produces high-quality products and the anchor discloses to the consumer, the consumer's posterior expectation and LH strategy are the same, ; ; ; .
[0042] Take strategy LN as an example, solve it by reverse induction, and get the anchor's profit Take the derivative of and set it equal to zero, get : .
[0043] Substitute into the manufacturer's profit formula , take the derivative of and set it equal to zero, get the optimal price : .
[0044] Substitute the optimal price into to get the optimal effort level , .
[0045] Substitute and into the manufacturer's and anchor's profit formulas to get the manufacturer's equilibrium profit and the anchor's equilibrium profit .
[0046] After getting the profit expression of each strategy, set = = , that is, the boundary value of each strategy, expressed as quality Regarding the expression of time allocation , draw a region diagram, and the corresponding region is the strategy that maximizes the profit.
[0047] Analysis shows that the anchor's entertainment time and the manufacturer's quality differentiation jointly determine the quality selection. Shorter entertainment time and greater quality differentiation will prompt the manufacturer to produce high-quality products. Analysis of user churn rate shows that when the quality difference is small, entertainment content can reduce fan churn rate, and when the quality difference is large, real quality information needs to be disclosed. The anchor's professional ability can maintain the loyalty of fans better than exaggerated propaganda. The impact of different strategies on consumer surplus shows that high-quality products need to increase entertainment to increase surplus, while low-quality products need to strengthen professional information to avoid expected gap. From the perspective of comprehensive strategy combination, manufacturers and anchors should focus on disclosing real information rather than relying on short-term false advertising. The LL strategy combination (low-quality products, disclose real low quality) produces the highest profit. Low-quality products can also maximize consumer willingness to buy through self-designed mathematical models with accurate information disclosure.
[0048] In the following, the technical solutions of the present application will be further described through specific preferred embodiments.
[0049] Preferably, when constructing the supply chain model, the numerical calculation library is used to pre-store the relationship tables of the equilibrium price , the equilibrium anchor effort , the manufacturer's equilibrium profit , and the anchor's equilibrium profit corresponding to each strategy . The relationship tables are shown in Table 1-1 and Table 1-2 below. After selecting the current strategy , the corresponding analytical solution is obtained by calling and querying the relationship table of the strategy . That is, the relationship tables of different strategies can also be pre-stored, queried and called in the constructed supply chain model.
[0050] Table 1-1 Relationship table of different strategies when low-quality products
[0051] Table 1-2 Relationship table of different strategies when high-quality products
[0052] When the strategy adopts low-quality products, whether or not to disclose real low quality, whether or not to exaggerate product quality, the manufacturer's production quality strategy refers to the low-quality parameter corresponding to the low-quality product Ll . Thus, the equilibrium price can be represented as the equilibrium price of Table 1-1. The equilibrium anchor effort can be represented as the equilibrium anchor effort of Table 1-1. The manufacturer equilibrium profit can be represented as the manufacturer equilibrium profit of Table 1-1.
[0053] More specifically, in the present embodiment, when the strategy is a low-quality product, and the strategy LN of not revealing the true low quality is adopted, in the step S3, the equilibrium price of the strategy LN is calculated by the formula , and the equilibrium anchor effort of the strategy LN is calculated by the formula ; in the step S4, the manufacturer equilibrium profit of the strategy LN is calculated by the formula , and the anchor equilibrium profit of the strategy LN is calculated by the formula .
[0054] In the present embodiment, when the strategy is a low-quality product, and the strategy LL of revealing the true low quality is adopted, in the step S3, the equilibrium price of the strategy LL is calculated by the formula , and the equilibrium anchor effort of the strategy LL is calculated by the formula ; in the step S4, the manufacturer equilibrium profit of the strategy LL is calculated by the formula , and the anchor equilibrium profit of the strategy LL is calculated by the formula .
[0055] In the present embodiment, when the strategy is a low-quality product, and the strategy LH of exaggerating the product quality is adopted, in the step S3, the equilibrium price of the strategy LH is calculated by the formula , and the equilibrium anchor effort of the strategy LH is calculated by the formula ; in the step S4, the manufacturer equilibrium profit of the strategy LH is calculated by the formula , and the anchor equilibrium profit of the strategy LH is calculated by the formula .
[0056] in, , , , , and These are intermediate calculation parameters, which are custom parameters used for intermediate calculations. Their functions are twofold: firstly, they can simplify the above formulas; secondly, they can be used in different calculation formulas to minimize unnecessary repetitive calculations and improve dynamic response speed. , , , , , .
[0057] When the strategy When using high-quality products, regardless of whether the manufacturer discloses the true high quality, it reflects their production quality strategy. These all refer to the high-quality parameters corresponding to high-quality product H. h Therefore, the equilibrium price This can be represented as Table 1-2. Balanced efforts from livestreamers This can be represented as Table 1-2. Manufacturer's equilibrium profit This can be represented as Table 1-2. Balanced Profits for Livestreamers This can be represented as Table 1-2. .
[0058] More specifically, in this embodiment, when the strategy When using high-quality products and not disclosing the true high-quality strategy HN, in step S3, the formula is used... Calculate the equilibrium price of strategy HN And through the formula Equilibrium broadcaster effort of calculation strategy HN In step S4, the formula is used. Calculate the manufacturer's equilibrium profit for strategy HN And through the formula Calculate the anchor equilibrium profit of strategy HN .
[0059] In this embodiment, when the strategy The high-quality product is adopted, and the real high-quality strategy HH is disclosed. In step S3, the equilibrium price of the strategy HN is calculated by formula , The equilibrium anchor effort of the strategy HN is calculated by formula , In step S4, the manufacturer equilibrium profit of the strategy HN is calculated by formula , The anchor equilibrium profit of the strategy HN is calculated by formula , .
[0060] , , , , , and are intermediate calculation parameters, that is, custom parameters for intermediate calculation. On the one hand, they can be used to simplify the above formulas. On the other hand, the same intermediate calculation parameter can be shared in different calculation formulas to minimize unnecessary repeated calculation and improve dynamic response speed. , , , , , .
[0061] In step S5 of the embodiment, the anchor agent includes a first state space, a first action space and a first reward function. The first state space includes product quality, transmission information accuracy, entertainment time, last period demand and fan change amount. The first action space includes strategy LN, strategy LL, strategy LH, strategy HN and strategy HH. The first reward function preferably adopts the calculation formula of anchor equilibrium profit . The manufacturer agent includes a second state space, a second action space and a second reward function. The second state space includes product quality, high-quality probability, entertainment time and anchor strategy. The second action space includes high-quality product and low-quality product. The second reward function preferably adopts the calculation formula of manufacturer equilibrium profit .
[0062] In the training process of step S5, the anchor agent and the manufacturer agent are trained by the DQN algorithm of reinforcement learning. The training process includes the following sub-steps: Step S501, using the output data of step S3 and step S4 as the basis data for training, input into the DQN network of the anchor agent and the manufacturer agent; Step S502, the anchor agent selects a strategy from the first action space according to the current state, and the manufacturer agent selects a corresponding high-quality product or low-quality product from the second action space, and then calls the numerical calculation library for calculation to generate a new state and a reward; Step S503, updating the DQN network according to the new state and the reward; Step S504, through iterative training, the strategy of the anchor agent and the manufacturer agent is constantly optimized, and gradually converges to the optimal strategy.
[0063] The step S6 of the embodiment includes the following sub-steps: Step S601, for low-quality products, the profit boundary analytical solution of each strategy is calculated according to the profit of the anchor adopting different strategies; these analytical solutions will clearly define the profit level that each strategy can achieve under the combination of low-quality products and entertainment time; Step S602, according to the profit boundary analytical solution of low-quality products under different strategies calculated in step S601, a first area graph of product quality and entertainment time is drawn; the first area graph can be in the form of a graph, or in the form of Table 2 below, which directly shows the profit distribution of different strategies under the low-quality product scenario, helping to identify the applicable range of each strategy; Step S603, for high-quality products, the profit boundary analytical solution of each strategy is calculated according to the profit of the anchor adopting different strategies; these analytical solutions will clearly define the profit level that each strategy can achieve under the combination of high-quality products and entertainment time; Step S604, according to the profit boundary analytical solution of high-quality products under different strategies calculated in step S603, a second area graph of product quality and entertainment time is drawn; this second area graph can also be in the form of a graph, or in the form of Table 2 below, which directly shows the profit distribution of different strategies under the high-quality product scenario; Step S605, compare and integrate the profit boundary analytical solution of low-quality products under different strategies with the profit boundary analytical solution of high-quality products under different strategies to generate the final comprehensive strategy profit analytical solution and area graph. Through comparison and analysis, determine which strategy can achieve the optimal profit under different product quality and entertainment time combinations, to finally generate the comprehensive strategy profit analytical solution, and generate the total area graph which integrates the first area graph and the second area graph, this graph integrates the strategy profit information under the low-quality product and high-quality product scenarios, providing comprehensive decision support for supply chain members.
[0064] This embodiment also provides a game theory-based e-commerce live streaming supply chain revenue optimization calculation system, which adopts the game theory-based e-commerce live streaming supply chain revenue optimization calculation method described above, and includes: The historical transaction data acquisition module obtains historical transaction data through the e-commerce platform's API interface; The supply chain model building module constructs a supply chain model by exploring the relationships between manufacturers, livestreamers, and consumers. The price acquisition module obtains strategies through a numerical calculation library. The equilibrium price under and strive for balanced streamer ; The profit-generating module, based on backward induction and the equilibrium price, and strive for balanced streamer Acquisition Strategy Manufacturer equilibrium profit Balance profits with streamers ; The agent training module trains the broadcaster agent and the manufacturer agent through reinforcement learning based on the output data of the price acquisition module and the profit acquisition module. The region map drawing module outputs region maps of product quality and entertainment time for both low-quality and high-quality products, based on the profit generated when the anchor adopts different strategies.
[0065] Below, Table 2 shows some of the profit results. Essentially, step S4 calculates the equilibrium profit for various strategies, allowing us to determine the optimal solution by calculating the profit corresponding to different entertainment times based on different strategies for low-quality or high-quality products. Table 2 below shows some examples of the calculated results.
[0066] Profit results in Table 2 entertainment time commission rate high quality probability quality manufacturer profit (LH) manufacturer profit (LN) manufacturer profit (LL) anchor profit (LH) anchor profit (LN) anchor profit (LL) manufacturer profit (HH) manufacturer profit (HN) anchor profit (HH) anchor profit (HN) 0.299 0.856 0.398 2.374 0.048 0.056 0.050 0.281 0.323 0.287 0.281 0.061 0.401 0.355 0.654 0.734 0.038 1.378 0.097 0.098 0.098 0.207 0.209 0.212 0.207 0.101 0.208 0.214 0.844 0.078 0.894 3.369 0.331 0.535 0.778 0.028 0.045 0.065 0.028 0.732 0.048 0.061 0.068 0.972 0.025 3.654 0.009 0.008 0.007 0.305 0.259 0.245 0.305 0.010 0.410 0.338 0.678 0.901 0.322 3.512 0.038 0.052 0.069 0.295 0.410 0.552 0.295 0.066 0.369 0.499 0.247 0.642 0.912 4.872 0.148 0.242 0.361 0.265 0.432 0.643 0.265 0.410 0.765 0.731 0.994 0.607 0.612 1.011 0.245 0.244 0.244 0.096 0.100 0.103 0.096 0.244 0.096 0.100 0.519 0.227 0.408 4.260 0.273 0.387 0.474 0.079 0.112 0.137 0.079 0.503 0.131 0.146 0.800 0.738 0.977 2.316 0.104 0.153 0.190 0.271 0.387 0.467 0.271 0.188 0.447 0.463 0.688 0.718 0.473 3.811 0.105 0.158 0.230 0.246 0.372 0.544 0.246 0.206 0.354 0.475 0.738 0.098 0.021 2.819 0.240 0.247 0.262 0.025 0.026 0.028 0.025 0.303 0.025 0.032 0.350 0.853 0.666 3.576 0.056 0.081 0.089 0.324 0.463 0.509 0.324 0.106 0.657 0.602 0.902 0.234 0.726 2.604 0.250 0.388 0.526 0.071 0.113 0.154 0.071 0.459 0.085 0.133 0.982 0.111 0.284 4.484 0.251 0.413 0.987 0.030 0.050 0.122 0.030 0.541 0.030 0.065 0.134 0.009 0.343 1.735 0.302 0.306 0.261 0.003 0.003 0.002 0.003 0.318 0.004 0.003 0.795 0.805 0.567 4.201 0.078 0.123 0.208 0.284 0.468 0.790 0.284 0.175 0.407 0.641 0.607 0.328 0.992 3.698 0.268 0.425 0.640 0.130 0.205 0.307 0.130 0.639 0.305 0.306 0.642 0.856 0.336 2.521 0.053 0.067 0.078 0.270 0.342 0.404 0.270 0.077 0.319 0.382 0.219 0.952 0.814 1.215 0.013 0.014 0.014 0.264 0.283 0.272 0.264 0.015 0.292 0.285 0.001 0.222 0.478 1.497 0.233 0.238 0.195 0.066 0.068 0.055 0.066 0.243 0.083 0.069 0.253 0.868 0.927 4.447 0.054 0.087 0.128 0.354 0.568 0.830 0.354 0.141 0.951 0.917 0.558 0.858 0.537 1.899 0.050 0.061 0.066 0.271 0.331 0.356 0.271 0.066 0.334 0.351 0.573 0.662 0.903 4.621 0.139 0.229 0.387 0.270 0.441 0.738 0.270 0.381 0.707 0.726 0.107 0.093 0.090 3.307 0.304 0.272 0.233 0.031 0.028 0.024 0.031 0.333 0.047 0.034 0.828 0.170 0.089 3.115 0.238 0.271 0.374 0.046 0.053 0.074 0.046 0.335 0.047 0.064 0.625 0.436 0.708 3.518 0.210 0.321 0.451 0.158 0.241 0.338 0.158 0.424 0.281 0.317 0.110 0.271 0.546 3.308 0.276 0.351 0.237 0.102 0.130 0.088 0.102 0.427 0.209 0.159 The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.
Claims
1. A method for optimizing the revenue calculation of e-commerce live streaming supply chain based on game theory, characterized in that, Includes the following steps: Step S1: Obtain historical transaction data through the API interface of the e-commerce platform. The historical transaction data includes actual product quality parameters, live streamer behavior data, and consumer review data. Step S2: Construct a supply chain model based on the relationships between manufacturers, livestreamers, and consumers. Cluster the quality of products produced by manufacturers into high-quality or low-quality products. When a product is low-quality, the strategies in the supply chain model include strategy LN (not revealing the true low quality), strategy LL (revealing the true low quality), and strategy LH (exaggerating product quality). When a product is high-quality, the strategies in the supply chain model include strategy HN (not revealing the true high quality) and strategy HH (revealing the true high quality). Set adjustment factors. Used to indicate changes in fan base; among which... For the actual quality of the product, To meet consumers' expectations for product quality, Sensitivity factor; Step S3: Obtain the strategy through the numerical computation library. The equilibrium price under and strive for balanced streamer ,in, This indicates the manufacturer's production quality strategy. This indicates the anchor's quality information promotion strategy; Step S4, based on backward induction and the equilibrium price... and strive for balanced streamer Acquisition Strategy Manufacturer equilibrium profit Balance profits with streamers ; Step S5: Based on the output data from steps S3 and S4, train the anchor agent and the manufacturer agent through reinforcement learning. Step S6: For low-quality and high-quality products, output regional maps of product quality and entertainment time based on the profit when the anchor adopts different strategies.
2. The method for optimizing e-commerce live streaming supply chain revenue based on game theory as described in claim 1, characterized in that, In step S1, the streamer behavior data includes promotion quality, entertainment interaction duration, and historical fan churn rate; the consumer evaluation data includes willingness to pay and return rate. When constructing the supply chain model in step S2, the formula is used. The probability that a broadcaster delivers accurate information is expressed by the formula. This represents the probability that the broadcaster conveys inaccurate information, where, and These respectively indicate the status of the information conveyed by the broadcaster regarding high-quality and low-quality products; and These respectively represent the actual quality of high-quality products and low-quality products; This indicates the information conveyed by the broadcaster. , This indicates the host's ability to convey product quality information; t represents entertainment time. This indicates the time the host spends introducing the product; The numerical computation library pre-stores various strategies. Corresponding equilibrium price Balanced efforts from streamers Manufacturer's Equilibrium Profit and the balanced profits of streamers Relationship table; when selecting the current strategy Then, by calling and querying the relationship table, the strategy can be obtained. The corresponding analytical solution.
3. The method for optimizing e-commerce live streaming supply chain revenue based on game theory according to claim 1 or 2, characterized in that, When the strategy When a low-quality product is used and the strategy LN that does not disclose the true low quality is employed, in step S3, the formula is used... Calculate the equilibrium price of strategy LN And through the formula Equilibrium broadcaster effort of computational strategy LN ,in, This indicates the probability that a manufacturer will produce high-quality products. Indicates product quality. The commission rate is represented by t, and the entertainment time is represented by t; in step S4, the formula is used... Calculate the manufacturer's equilibrium profit for strategy LN And through the formula Calculate the anchor equilibrium profit of strategy LN .
4. The method for optimizing e-commerce live streaming supply chain revenue based on game theory according to claim 1 or 2, characterized in that, When the strategy When a low-quality product is used, and the strategy LL reveals the true low quality, in step S3, the formula is used... Equilibrium price of strategy LL And through the formula Equilibrium of LL Computational Strategy In step S4, the formula is used. Calculate the manufacturer's equilibrium profit using strategy LL And through the formula Calculate the streamer equilibrium profit of strategy LL ; in, Let t represent the probability that a manufacturer produces a high-quality product, and let t represent the entertainment time. This indicates the anchor's ability to convey product quality information. Indicates product quality. Indicates the commission rate; , and These are intermediate calculation parameters. , , .
5. The method for optimizing the revenue calculation of e-commerce live streaming supply chain based on game theory according to claim 1 or 2, characterized in that, When the strategy When a strategy LH that uses low-quality products and exaggerates product quality is employed, in step S3, the formula is used... Calculate the equilibrium price of strategy LH And through the formula Equilibrium of Calculation Strategy LH In step S4, the formula is used. Calculate the manufacturer's equilibrium profit for strategy LH And through the formula Calculate the anchor equilibrium profit of strategy LH ; in, Let t represent the probability that a manufacturer produces a high-quality product, and let t represent the entertainment time. This indicates the anchor's ability to convey product quality information. Indicates product quality. Indicates the commission rate; , and These are intermediate calculation parameters. , , 。 6. The method for optimizing the revenue calculation of e-commerce live streaming supply chain based on game theory according to claim 1 or 2, characterized in that, When the strategy When using high-quality products and not disclosing the true high-quality strategy HN, in step S3, the formula is used... Calculate the equilibrium price of strategy HN And through the formula Equilibrium broadcaster effort of calculation strategy HN ; In step S4, the formula is used. Calculate the manufacturer's equilibrium profit for strategy HN And through the formula Calculate the anchor equilibrium profit of strategy HN ; in, This indicates the probability that a manufacturer will produce high-quality products. Indicates product quality. This represents the commission rate, and t represents the entertainment time. , and These are intermediate calculation parameters. , , 。 7. The method for optimizing the revenue calculation of e-commerce live streaming supply chain based on game theory according to claim 1 or 2, characterized in that, When the strategy When using high-quality products and revealing a truly high-quality strategy HH, in step S3, the formula is used... Calculate the equilibrium price of strategy HN And through the formula Equilibrium broadcaster effort of calculation strategy HN ; In step S4, the formula is used. Calculate the manufacturer's equilibrium profit for strategy HN And through the formula Calculate the anchor equilibrium profit of strategy HN ; in, This indicates the probability that a manufacturer will produce high-quality products. Indicates product quality. This represents the commission rate, and t represents the entertainment time. This indicates the anchor's ability to convey product quality information; , , and These are intermediate calculation parameters. , , , 。 8. The method for optimizing the revenue calculation of e-commerce live streaming supply chain based on game theory according to claim 1 or 2, characterized in that, In step S5, the anchor agent includes a first state space, a first action space, and a first reward function. The first state space includes product quality, information transmission accuracy, entertainment time, previous demand, and fan change. The first action space includes policy LN, policy LL, policy LH, policy HN, and policy HH. The manufacturer agent includes a second state space, a second action space, and a second reward function. The second state space includes product quality, high quality probability, entertainment time, and anchor strategy. The second action space includes high-quality products and low-quality products. During the training process in step S5, the broadcaster agent and the manufacturer agent are trained using the DQN algorithm of reinforcement learning. The training process includes the following sub-steps: Step S501: Use the output data from steps S3 and S4 as the basic training data and input it into the DQN network of the anchor agent and the manufacturer agent. In step S502, the anchor agent selects a strategy from the first action space based on the current state, and the manufacturer agent selects the corresponding high-quality product or low-quality product from the second action space. Then, the numerical calculation library is called to perform calculations to generate a new state and reward. Step S503: Update the DQN network according to the new state and reward; Step S504: Through iterative training, continuously optimize the strategies of the anchor agent and the manufacturer agent, so that they gradually converge to the optimal strategy.
9. The method for optimizing the revenue calculation of e-commerce live streaming supply chain based on game theory according to claim 1 or 2, characterized in that, Step S6 includes the following sub-steps: Step S601: For low-quality products, calculate the analytical solution of the profit boundary for each strategy based on the profit when the anchor adopts different strategies. Step S602: Based on the analytical solution of the profit boundary of low-quality products under different strategies calculated in step S601, draw the first region map of product quality and entertainment time. Step S603: For high-quality products, calculate the analytical solution of the profit boundary for each strategy based on the profit when the anchor adopts different strategies. Step S604: Based on the analytical solution of the profit boundary of high-quality products under different strategies calculated in step S603, draw a second region diagram of product quality and entertainment time. Step S605: Compare and integrate the analytical solutions of the profit boundaries of low-quality products under different strategies with the analytical solutions of the profit boundaries of high-quality products under different strategies to generate the final comprehensive strategy profit analytical solution and region map.
10. A game theory-based e-commerce live streaming supply chain revenue optimization calculation system, characterized in that, The method for optimizing the revenue of e-commerce live streaming supply chain, based on game theory as described in any one of claims 1 to 9, is adopted and includes: The historical transaction data acquisition module obtains historical transaction data through the e-commerce platform's API interface; The supply chain model building module constructs a supply chain model by exploring the relationships between manufacturers, livestreamers, and consumers. The price acquisition module obtains strategies through a numerical calculation library. The equilibrium price under and strive for balanced streamer ; The profit-generating module, based on backward induction and the equilibrium price, and strive for balanced streamer Acquisition Strategy Manufacturer equilibrium profit Balance profits with streamers ; The agent training module trains the broadcaster agent and the manufacturer agent through reinforcement learning based on the output data of the price acquisition module and the profit acquisition module. The region map drawing module outputs region maps of product quality and entertainment time for both low-quality and high-quality products, based on the profit generated when the anchor adopts different strategies.