E-commerce full-link income collaborative decision-making agent based on security reinforcement learning
By constructing an intelligent agent for collaborative decision-making across the entire e-commerce value chain, and utilizing the meta-strategy decision-making center and the agent's collaborative decision-making instructions, the coordination problem between product pricing and advertising bidding systems was solved, thereby achieving an improvement in globally optimal business revenue.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANMIAO TECHNOLOGY (HANGZHOU) CO LTD
- Filing Date
- 2026-03-25
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, product pricing and advertising bidding systems lack a unified optimization framework and coordination mechanism, leading to strategy conflicts and resource inefficiencies, making it difficult to achieve globally optimal business returns.
Construct a collaborative decision-making intelligent agent for the entire e-commerce revenue chain based on secure reinforcement learning, including a meta-policy decision-making hub, a product pricing intelligent agent, and an advertising bidding intelligent agent. Coordinate the decisions between the two through collaborative decision-making instructions, and combine target weight coefficients and virtual budget allocation to achieve dynamic resource allocation and iterative optimization under security constraints.
It effectively alleviates strategic conflicts, achieves synchronization and coordination between product pricing and advertising bidding, improves overall business revenue, and ensures efficient allocation and business reliability of marketing resources from a holistic perspective.
Smart Images

Figure CN121903684A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular to an intelligent agent for collaborative decision-making on the entire revenue chain of e-commerce based on security reinforcement learning. Background Technology
[0002] In e-commerce operations, merchants typically manage two key but significantly different types of business decisions: one is a low-frequency strategic decision, exemplified by product pricing, whose core objective is to achieve long-term profit goals through price optimization; the other is a high-frequency tactical decision, exemplified by real-time advertising bidding, whose core objective is to compete for limited traffic to achieve immediate conversions and sales growth. Both types of decisions jointly influence final revenue, but they differ inherently in their decision frequency, core objectives, and risk characteristics.
[0003] Existing technologies typically employ independent systems to handle these two types of decisions. Pricing systems often build price-volume models based on historical transaction data, adjusting prices periodically with profit maximization as the guiding principle. Bidding systems, on the other hand, rely on real-time auction feedback, making high-frequency adjustments based on clicks or conversions within budget constraints. Due to the lack of a unified optimization framework and coordination mechanism, this fragmented execution model is prone to strategy conflicts: for example, when an advertising system pays high costs for traffic acquisition to boost volume, the pricing system may not be in an optimal state matching the traffic acquisition strategy, resulting in diluted returns on marketing investment and making it difficult to achieve overall operational efficiency at the global optimum. Summary of the Invention
[0004] To address the above issues, this application provides an intelligent agent for collaborative decision-making on the entire e-commerce revenue chain based on security reinforcement learning.
[0005] In a first aspect, embodiments of this application provide an e-commerce end-to-end revenue collaborative decision-making intelligent agent based on secure reinforcement learning, comprising: a meta-strategy decision-making center, a product pricing intelligent agent, and an advertising bidding intelligent agent; the meta-strategy decision-making center is configured to generate collaborative decision-making instructions based on a global market state, the collaborative decision-making instructions being used to coordinate the decisions between the product pricing intelligent agent and the advertising bidding intelligent agent; wherein, the global market state includes at least the lifecycle stage of the target product and the decision execution results fed back by the product pricing intelligent agent and the advertising bidding intelligent agent respectively; the product pricing intelligent agent is configured to execute a corresponding pricing decision according to the collaborative decision-making instructions and feed back the execution result of the pricing decision to the meta-strategy decision-making center; the advertising bidding intelligent agent is configured to execute a corresponding bidding decision according to the collaborative decision-making instructions and feed back the execution result of the bidding decision to the meta-strategy decision-making center.
[0006] In one possible implementation of the first aspect above, the collaborative decision instruction includes a target weight coefficient and a virtual budget allocation; the target weight coefficient indicates that, at the life cycle stage of the target product, the pricing of the target product focuses on profit or sales volume, and indicates that the advertising bid for the target product focuses on traffic or return on investment; the virtual budget allocation indicates that, under global budget constraints, a virtual resource quota is dynamically allocated for advertising consumption and product discounts.
[0007] In one possible implementation of the first aspect above, the product pricing agent and the advertising bidding agent are configured to make decisions under security constraints, and both are iteratively optimized using a framework that combines secure online exploration with conservative offline training. The execution results of the pricing decision include product sales and profit data, and the execution results of the bidding decision include advertising consumption and conversion data.
[0008] In one possible implementation of the first aspect above, the product pricing agent has a first price-volume model. The product pricing agent is configured to execute corresponding pricing decisions under security constraints according to the collaborative decision-making instructions, including: the product pricing agent performs a security exploration within a preset dynamic pricing security domain based on the first price-volume model, determines the pricing of the target product, and stores the product sales and profit data generated from sales based on the pricing of the target product in the experience replay buffer in the e-commerce full-link revenue collaborative decision-making agent; The sales and profit data include data generated from both unsold and sold items of the target product.
[0009] In one possible implementation of the first aspect above, the advertising bidding agent has a second price-volume model, the advertising bidding agent includes a budget allocator and a real-time bidder; the advertising bidding agent is configured to perform corresponding pricing decisions under security constraints according to the collaborative decision instruction, including: According to the collaborative decision-making instructions, the advertising bidding agent executes a high-level strategy through the budget allocator: determining the budget for each advertising campaign based on the global state of the target product's advertising; and executes a low-level strategy through the real-time bidder: determining the keyword bid adjustment amount and activation probability in the target product's advertising based on the budget of each advertising campaign. During the safe exploration period, the advertising bidding agent determines the advertising bid for the target product based on the budget of each advertising campaign, the keyword bid adjustment amount and activation probability in the target product's advertisement, and the second price-volume model, within a preset dynamic bidding safety domain; or, If the period is not in the safe exploration phase, the advertising bidding agent executes a safety baseline strategy to determine the advertising bid for the target product; The advertising bidding agent stores the advertising consumption and conversion data generated from the advertising bids for the target product in the experience replay buffer of the e-commerce end-to-end revenue collaborative decision-making agent.
[0010] In one possible implementation of the first aspect above, the commodity pricing agent is configured to perform conservative offline training based on the commodity sales and profit data stored in the experience replay buffer of the e-commerce full-link revenue collaborative decision-making agent within a historical time period, so as to update the first price-volume model, and feed the update result back to the meta-strategy decision center, which then updates the global market state.
[0011] In one possible implementation of the first aspect above, the advertising bidding agent is configured to perform conservative offline training based on the advertising consumption and conversion data stored in the experience replay buffer of the e-commerce full-link revenue collaborative decision-making agent within a historical time period, so as to update the second price-volume model, and feed the update result back to the meta-strategy decision center, which then updates the global market state.
[0012] In one possible implementation of the first aspect above, the global market state further includes the target product's status label, historical price and volume data, real-time advertising metrics, and competitor prices; The meta-strategy decision-making center is configured to generate collaborative decision-making instructions based on the global market state, including: The meta-strategy decision center determines the target route decision for the target product based on the life cycle stage of the target product as determined by the global market status. The meta-strategy decision-making center, based on the target routing decision, performs product-level value assessment, channel-level efficiency assessment, and joint optimization allocation based on opportunity cost, and generates the collaborative decision-making instruction.
[0013] In one possible implementation of the first aspect above, the lifecycle stages include a new product cold start period, a growth period, a maturity period, and a clearance / promotion period; within different lifecycle stages, the target product has different routing decisions; each routing decision has a corresponding target weight coefficient and the virtual budget allocation.
[0014] In one possible implementation of the first aspect above, the meta-strategy decision center is configured to perform the commodity-level value assessment, including: the meta-strategy decision center determining the commodity value of the target commodity based on the marginal value benefit of the target commodity; wherein the commodity value of the target commodity represents the importance of the investment, the higher the commodity value, the more worthwhile the investment, and vice versa; the marginal value benefit is proportional to the investment budget of the target commodity.
[0015] In one possible implementation of the first aspect above, the meta-strategy decision center is configured to perform the channel-level performance evaluation, including: the meta-strategy decision center determining the return on investment of the target product in the advertising channel and the discount channel, respectively, based on the advertising channel performance and the discount channel performance of the target product; wherein, the advertising channel performance of the target product is determined by the incremental profit and consumption generated by the advertising of the target product within a specified time period obtained by the meta-strategy decision center through the advertising bidding agent; the discount channel performance is determined by the expected incremental profit and discount budget of the product obtained by the meta-strategy decision center through the product pricing agent.
[0016] In one possible implementation of the first aspect above, the meta-strategy decision center is configured to perform joint optimization allocation based on opportunity cost, including: the meta-strategy decision center determining the target weight coefficient and the virtual budget allocation based on the product value of the target product, the advertising channel effectiveness and the discount channel effectiveness of the target product. Attached Figure Description
[0017] Figure 1 According to some embodiments of this application, a schematic diagram of the structure of an e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning is shown. Figure 2 According to some embodiments of this application, a flowchart of a collaborative decision-making process for the entire revenue chain of e-commerce based on security reinforcement learning is shown.
[0018] Figure 3 According to some embodiments of this application, a flowchart of a collaborative decision-making method is shown.
[0019] Figure 4 According to some embodiments of this application, a schematic diagram of the structure of an electronic device is shown. Detailed Implementation
[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] In current e-commerce operations, product pricing and advertising bidding, as two core decisions affecting e-commerce revenue, are typically managed and optimized by separate systems. The pricing system focuses on periodically adjusting prices based on historical transaction data and price-volume models, while the bidding system strives to compete for traffic in a real-time bidding environment to achieve immediate conversions.
[0022] In some embodiments, due to inherent differences in decision-making frequency, core objectives, and operational logic between these two types of systems, and the lack of a unified coordination mechanism, strategic conflicts and resource inefficiencies often arise in actual operation. For example, when the bidding system uses high prices to attract traffic in order to boost sales, the pricing system may not be in a profit-optimal state that matches it, causing the high marketing investment to fail to be converted into corresponding overall profits, ultimately resulting in a dilemma where the overall business benefits fall into a local optimum but a global suboptimal state.
[0023] Based on this, this application constructs an e-commerce end-to-end revenue collaborative decision-making intelligent agent, consisting of a meta-strategy decision-making center, a product pricing intelligent agent, and an advertising bidding intelligent agent. The meta-strategy decision-making center is responsible for analyzing the overall market state (including product lifecycle stages, feedback results from each intelligent agent, etc.) and generating collaborative decision-making instructions. These instructions do not directly specify concrete prices or bids, but rather dynamically coordinate the optimization tendencies and resource usage of the two executing intelligent agents by issuing target weight coefficients and virtual budget allocations. Upon receiving the collaborative decision-making instructions, the product pricing intelligent agent and the advertising bidding intelligent agent make decisions in the pricing and bidding domains respectively, and feed back the execution results to the center, thus forming a complete closed loop of perception-decision-execution-feedback.
[0024] Therefore, through the unified scheduling of the meta-strategy decision-making center, the two major decision-making processes of product pricing and advertising bidding are deeply coupled at the levels of goal setting and resource allocation, ensuring that the actions of both serve the globally optimal goal calculated by the meta-strategy decision-making center. This not only effectively alleviates strategic conflicts but also allows marketing resources that might otherwise be wasted to be dynamically and efficiently allocated from a global perspective.
[0025] See Figure 1 and Figure 2 , Figure 1 This is a schematic diagram of the structure of an e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning, provided in an embodiment of this application. Figure 2This is a flowchart of a collaborative decision-making process for the entire revenue chain of e-commerce based on security reinforcement learning, provided in an embodiment of this application.
[0026] This application provides a collaborative decision-making agent for e-commerce end-to-end revenue based on security reinforcement learning. The collaborative decision-making agent for e-commerce end-to-end revenue includes: a meta-strategy decision center, a product pricing agent, and an advertising bidding agent.
[0027] The meta-strategy decision-making hub is configured to generate collaborative decision-making instructions based on the global market state. These instructions coordinate the decisions between the product pricing agent and the advertising bidding agent. The global market state includes at least the target product's lifecycle stage and the decision execution results reported by the product pricing agent and the advertising bidding agent, respectively.
[0028] The product pricing agent is configured to execute corresponding pricing decisions based on collaborative decision-making instructions and feed back the execution results of the pricing decisions to the meta-strategy decision-making center.
[0029] The advertising bidding agent is configured to execute corresponding bidding decisions based on collaborative decision-making instructions and feed back the execution results of the bidding decisions to the meta-strategy decision center.
[0030] It is understood that, in the embodiments of this application, the e-commerce end-to-end revenue collaborative decision-making intelligent agent can be an electronic device containing a collaborative decision-making system or a coordinated decision-making system. The meta-strategy decision-making center does not directly output specific product prices or advertising bids. Instead, the meta-strategy decision-making center continuously receives and integrates information from multiple sources. This information collectively constitutes its understanding of the current business environment, i.e., the global market state.
[0031] The global market status can be considered a dynamic and comprehensive data view. It's not merely a list of historical transaction data for commodities, but also includes dimensions that can influence long-term returns. A key dimension is the product's lifecycle stage, indicating whether it's in a new product phase requiring rapid market penetration, a growth phase requiring balanced growth, a mature phase requiring stable profits, or a clearance / promotion phase requiring inventory reduction. Furthermore, the global market status integrates decision-making execution results from the product pricing agent and the advertising bidding agent. These feedback results are the direct basis for the meta-strategy decision-making center to evaluate the effectiveness of instructions and perceive real-time market responses. Based on this deeply integrated global perspective, the meta-strategy decision-making center generates strategy instructions—collaborative decision-making instructions.
[0032] The essence of collaborative decision-making instructions can be a set of high-order policy parameters. These parameters convey to the product pricing agent and the advertising bidding agent, for example, the intention of "what goals should be prioritized at the current stage and how many resources can be used," rather than specific operational instructions.
[0033] A product pricing agent can be a system or module focused on making product sales price decisions. The pricing agent receives collaborative decision-making instructions from the meta-strategy decision-making center and combines these instructions with the product's own state (such as cost and inventory). This allows it to transform collaborative decision-making instructions from the meta-strategy decision-making center, such as those emphasizing sales volume at this stage, into an executable, specific product price.
[0034] An ad bidding agent can be a system or module focused on bidding decisions in the ad auction market. It receives the same collaborative decision-making instructions and combines these instructions with real-time ad performance (such as click-through rate and budget). This translates collaborative decision-making instructions from the meta-strategy decision center—for example, prioritizing high-quality traffic in the current phase—into a series of specific operations on the ad platform, such as bidding and budget allocation.
[0035] After executing their decisions, both the product pricing agent and the advertising bidding agent will feed back the resulting business outcomes (such as sales volume and profit after pricing, and advertising consumption and conversion after bidding) to the meta-strategy decision-making center. This forms a complete closed-loop decision-making process: the meta-strategy decision-making center issues instructions → the agents execute instructions → the results are fed back to the center.
[0036] To address this, the aforementioned architecture establishes a unified command system for the e-commerce end-to-end revenue collaborative decision-making agent based on security reinforcement learning. The meta-policy decision-making center ensures that the two major business actions—product pricing and advertising bidding—remain synchronized and coordinated in terms of strategic goals and resource budgets, thereby enabling two potentially conflicting systems to work together to improve overall business revenue.
[0037] In some embodiments, collaborative decision-making instructions may include target weighting coefficients and virtual budget allocations.
[0038] Among them, the target weight coefficient can indicate whether the pricing of the target product focuses on profit or sales volume at the stage of the target product's life cycle, and whether the advertising bid for the target product focuses on traffic or return on investment.
[0039] Virtual budget allocation can indicate the dynamic allocation of virtual resources for advertising spending and product discounts under global budget constraints.
[0040] In some embodiments, the collaborative decision-making instructions are designed to be implemented through two core control dimensions: target weight coefficients and virtual budget allocation. This design translates high-level strategic intent into quantifiable and executable operational parameters, thereby establishing a precise collaborative language between the product pricing agent and the advertising bidding agent.
[0041] In some embodiments, the target weight coefficient can be a multi-dimensional vector or a set of parameters, whose core function is to dynamically decompose and convey the meta-strategy decision-making center's understanding of the global objective. The target weight coefficient is not fixed but can be adjusted according to the life cycle stage of the target product.
[0042] As an example, when the e-commerce end-to-end revenue collaborative decision-making agent determines that a target product is in its new product launch phase, the meta-strategy decision center can generate a set of target weight coefficients that highly favor maximizing product sales and acquiring advertising traffic. This set of coefficients conveys to the product pricing agent that, at the current stage, the product pricing agent, within its existing decision model, prioritizes increasing sales volume over pursuing profit per unit. Simultaneously, this set of coefficients conveys to the advertising bidding agent that, at the current stage, the advertising bidding agent's task is to acquire sufficient market exposure and clicks, and the stringent requirements for short-term ROI can be appropriately relaxed. In other words, this involves setting a low data richness score, high advertising weight, low discount weight, and a relaxed safety domain.
[0043] The data richness score is an indicator that measures whether the available data for a product is sufficient to support the model in making reliable decisions. The higher the score, the more abundant the data, and the more confidently the model can rely on the data to make decisions; the lower the score, the sparser the data, and the more conservative the decisions need to be, relying more on rules, prior knowledge, or exploration strategies.
[0044] As another example, during the maturity stage, the target weight coefficients will tilt towards profit maintenance and ROI optimization. That is, a high data richness score, a high-price-volume model weight, a low exploration weight, and a tightened safety domain are assigned. In this way, the target weight coefficients ensure that two agents in different business stages have a consistent understanding of what is most important at the current stage of the product's lifecycle.
[0045] In some embodiments, virtual budget allocation can refer to a secondary dynamic allocation performed by a meta-strategy decision-making center within a pre-defined, periodic global marketing budget framework. This allocation is virtual because it does not directly disburse cash, but rather sets authorized caps and inclinations for resource consumption within the scope of action of two agents.
[0046] Specifically, virtual budget allocation comprises at least two parts: one is the advertising expenditure explicitly allocated to the advertising bidding agent, and the other is the promotional resource allocation implicitly allocated to the product pricing agent, manifested through price discounts. By adjusting the ratio of these two parts, the meta-strategy decision center essentially adjusts whether more resources are allocated to external advertising for new customer acquisition or to internal price reductions and promotions to directly stimulate sales. For example, for a new product that needs to quickly build brand awareness, the meta-strategy decision center can allocate a higher proportion of the virtual budget to advertising channels. Conversely, for a clearance item with high inventory pressure, the meta-strategy decision center can allocate more of the virtual budget to discount channels.
[0047] It is understandable that virtual budget allocation and target weighting coefficients work together; aggressive growth targets are usually accompanied by more ample advertising budgets, while profit protection targets often correspond to tighter discount budgets.
[0048] In this way, by combining the target weight coefficient with the virtual budget allocation, the meta-strategy decision-making center can simultaneously control the decision-making direction and action boundaries of the product pricing agent and the advertising bidding agent in a clear, measurable and flexible manner, fundamentally realizing the alignment and synchronization of goals and resources between the two heterogeneous decision-making systems.
[0049] In some embodiments, the product pricing agent and the advertising bidding agent are configured to make decisions under security constraints, and both employ an iterative optimization framework combining secure online exploration with conservative offline training. The results of the pricing decision may include product sales and profit data, while the results of the bidding decision may include advertising consumption and conversion data.
[0050] In some embodiments, to ensure the commercial reliability of automated decision-making and achieve continuous optimization, the product pricing agent and the advertising bidding agent can be co-architected within an iterative framework that integrates active learning and risk management. The core features of this framework are decision execution under security constraints and an iterative paradigm that combines secure online exploration with conservative offline training.
[0051] Among these, safety constraints are not mere limitations, but a dynamic protective system that permeates the entire decision-making and execution process. For a product pricing agent, safety constraints can manifest as hard boundaries for price fluctuations (such as not allowing prices to fall below cost), protection thresholds for profit margins, or credible price ranges defined based on historical market acceptance. For an advertising bidding agent, safety constraints are typically associated with strict budget consumption rates, upper limits on the Advertising Cost of Sales (ACOS), and maximum bid limits set based on historical channel performance. These constraints collectively constitute the protective barriers of the agent's action space, ensuring that any automatically generated decisions do not cross the pre-defined business risk threshold.
[0052] Safe online exploration is a key mechanism within the aforementioned safety constraints, where the product pricing agent and the advertising bidding agent proactively try new strategies to seek better solutions. Safe online exploration differs from blind random experimentation; it is a controlled trial-and-error process guided by theory or experience. For example, within a safe price range, the product pricing agent can conduct more intensive trials near a high-potential-profit price point based on current expectations of the demand curve. Similarly, the advertising bidding agent can slightly increase bids for certain high-conversion keywords below the ACOS safety line to compete for better ad placements. This process allows the e-commerce end-to-end revenue collaborative decision-making agent to collect data on the effects of unknown actions in a real business environment, making it a necessary step to break local optima and adapt to market changes.
[0053] Conservative offline training can be the core of the e-commerce end-to-end revenue collaborative decision-making agent's self-evolution using experience accumulated through online exploration. Conservative offline training specifically refers to a cautious and robust algorithm built into the model update phase. When training the model using historical data (including successful and non-immediately effective exploration records) in the experience replay buffer of the e-commerce end-to-end revenue collaborative decision-making agent, the conservative offline training algorithm deliberately reduces overly optimistic value estimates for high-reward actions that are rare or occur only once in the data.
[0054] In this embodiment, conservative offline training systematically "penalizes" the model's tendency to overestimate out-of-distribution actions by introducing a specific regularization term into the loss function. This makes the trained policy more inclined to choose actions that have been repeatedly verified as safe and effective by historical data, or to maintain necessary skepticism towards high-risk, high-reward actions. This effectively prevents the agent from becoming overly aggressive due to accidental success and falling into a decision-making trap that is detrimental in the long run.
[0055] The sales and profit data included in the pricing decision execution results, and the advertising consumption and conversion data included in the bidding decision execution results, serve as training samples for this iterative framework. The pricing decision execution results represent the true impact of price changes on market demand and the final profit or loss, while the bidding decision execution results reflect the efficiency of advertising resource investment in a competitive environment in real time. These two types of data are fed back to the meta-strategy decision center and their respective training processes, used to refresh the global market state and drive conservative updates to the agent's strategy, respectively.
[0056] By deeply integrating security constraints, secure online exploration, and conservative offline training, the e-commerce end-to-end revenue collaborative decision-making agent achieves a balance: bold exploration under the protection of rigid security constraints, while remaining robust through data-driven learning. This enables the product pricing agent and the advertising bidding agent to continuously adapt and optimize while mitigating significant business risks, ultimately driving the entire security reinforcement learning-based e-commerce end-to-end revenue collaborative decision-making agent towards a higher-return equilibrium state.
[0057] In some embodiments, the product pricing agent may have a first price-quantity model, and the product pricing agent is configured to execute corresponding pricing decisions under security constraints according to collaborative decision-making instructions, including: The product pricing agent conducts a safe exploration within a preset dynamic pricing safety domain based on the first price-volume model, determines the price of the target product, and stores the product sales and profit data generated from sales based on the target product's price in the experience playback buffer of the e-commerce full-link revenue collaborative decision-making agent.
[0058] The sales and profit data include data generated separately for unsold and sold target products.
[0059] In some embodiments, the core decision-making logic of the commodity pricing agent is embodied in an interactive learning process that integrates prior knowledge and proactive trial and error. This process is based on the first price-quantity model as its cognitive foundation, uses the dynamic pricing safety domain as its action boundary, collects feedback in the market through safe exploration, and systematically stores all results in an experience replay buffer, laying a data foundation for subsequent in-depth optimization.
[0060] In the embodiments of this application, the first price-volume model can provide robust prior guidance when data is sparse or the market environment changes. For example, for a new product, the first price-volume model can be initialized as a coarse estimation model calibrated based on the category average price, cost structure, and initial small-scale test data. For products with existing data, the first price-volume model can be a continuously updated predictor that reflects recent price elasticity. The core function of the first price-volume model is to provide the product pricing agent with a preliminary estimate of "what the expected sales volume and profit might be if a certain price is set" before each pricing decision, thereby guiding the exploration direction and reducing completely blind searches.
[0061] Setting a dynamic pricing safety domain is a key practice for ensuring business security. The dynamic pricing safety domain is not a fixed percentage fluctuation range, but a dynamically calculated value that is linked to the product's status and external instructions. The "dynamic" nature of the dynamic pricing safety domain is reflected in at least two aspects: First, the benchmark center point of the dynamic pricing safety domain is dynamically selected according to the business strategy. For example, during the new product phase, the average price of competitors can be used as the benchmark center point, while during the mature phase, the benchmark center point is the product's own historical best gross profit price. Second, the radius of the dynamic pricing safety domain is adjusted by the collaborative decision-making instructions issued by the meta-strategy decision-making center. For example, when the collaborative decision-making instructions indicate aggressive volume expansion, the dynamic pricing safety domain can be significantly widened towards price reduction, while when the collaborative decision-making instructions indicate profit protection, the dynamic pricing safety domain will narrow overall, and exploration on the high-price side may be strictly limited. Therefore, the dynamic pricing safety domain constitutes a hard boundary that all exploratory actions of the product pricing agent must adhere to.
[0062] Safe exploration within the aforementioned boundaries constitutes a guided sampling experiment. The product pricing agent does not uniformly try every price point within the dynamic pricing safety domain. Instead, it combines predictions from the first price-volume model (e.g., predicting high-profit points) with a degree of randomness to generate an exploratory price. Whether this price ultimately leads to a successful sale or a failed transaction, the exploration is fully recorded. Therefore, the "unsold" state is also considered important market feedback information. A failed attempt to sell at a high price is just as valuable as a successful sale at a low price in building an understanding of the complete market demand boundary.
[0063] All sales and profit data generated during the exploration process, including specific sales volume, sales revenue, and profit at the time of a transaction, as well as zero sales records, corresponding prices, and inventory costs for unsold transactions, are structured and stored in the experience replay buffer of the e-commerce end-to-end revenue collaborative decision-making intelligent agent. This ensures the integrity and authenticity of the training data for the first price-volume model, recording not only successful experiences but also negative market reactions to certain prices. This allows subsequent conservative offline training to learn on a more comprehensive data distribution containing both positive and negative samples, preventing the strategy from becoming overly optimistic or aggressive, and truly understanding and respecting the boundaries of the market.
[0064] In some embodiments, the commodity pricing agent is configured to perform conservative offline training based on the commodity sales and profit data stored in the experience replay buffer of the e-commerce full-link revenue collaborative decision-making agent within a historical time period, in order to update the first price-volume model, and feed the update results back to the meta-policy decision center, which then updates the global market state.
[0065] Specifically, the product pricing agent can periodically (e.g., at midnight each day) sample a batch of historical data from the experience replay buffer. These data samples include the current state (e.g., inventory, competitor prices), the action taken (i.e., the price explored), the immediate reward obtained (e.g., the profit generated at that price, or zero or negative cost if the product is not sold), and the subsequent state.
[0066] For the "unsold" samples with very low or negative immediate rewards, the standard training algorithm may simply learn to avoid these prices without processing, but it cannot distinguish whether this is because the price itself is infeasible or simply because of bad luck during exploration.
[0067] To address this, in this embodiment, conservative offline training introduces a targeted data processing and loss calculation method. Instead of discarding these "unsold" samples, it utilizes a separate, lightweight sales prediction model pre-trained only on successful transaction data to conservatively estimate the potential sales volume of these samples. For example, for an exploratory high-priced unsold sample, the sales prediction model provides a sales prediction range, and the training algorithm intentionally uses the lower bound of this range (i.e., a pessimistic prediction) to re-estimate a conservative reward. In subsequent model parameter updates, the algorithm uses a specific loss term (such as a data imputation loss term) to encourage the first price-volume model of the product pricing agent to move its long-term value prediction for this high-price action closer to this lower conservative reward. In this way, the e-commerce end-to-end revenue collaborative decision-making agent learns to be wary of high-price areas that have not been fully validated by the market, reducing the tendency to overly optimistically and frequently attempt these high-risk prices in future online explorations. The updated first price-volume model will be more sensitive to price risk assessment, and this update result will be fed back to the meta-policy decision center, giving it a more accurate global understanding of the market price boundaries of the product.
[0068] See Figure 2 In specific application scenarios, for a product pricing intelligence agent, the first step is to determine the authenticity of the data and whether there are actual sales figures. If there are actual sales figures, the e-commerce end-to-end revenue collaborative decision-making intelligence agent possesses reliable data that can be directly used to calculate profits and update the model. If there are no actual sales figures, a data filling mechanism is initiated, which is crucial for handling products in the "cold start period" or with "sparse data."
[0069] When there are no actual sales figures, the e-commerce end-to-end revenue collaborative decision-making intelligent agent cannot blindly set prices; the product pricing intelligent agent calls L... Impute A "virtual sales volume" is estimated using a fill mechanism to calculate expected profits. Furthermore, the product pricing agent uses a trained sales prediction model to predict a sales volume value y based on the current state (such as price, traffic, and competitive environment). pred For security reasons, the product pricing agent does not directly use the sales value y. pred Instead, it uses a conservative estimate that is adjusted downwards. The product pricing agent assumes that sales volume is at least y. pred -k* So much, and based on this, calculate the expected profit y. pred -k* .in, This represents the uncertainty (standard deviation) of the model's prediction, and k is a safety factor (e.g., k=2, corresponding to a lower bound of approximately 95% confidence interval). This risk-averse strategy can reduce the risk of overpricing and inventory buildup due to overestimating sales volume.
[0070] Understandably, regardless of whether sales figures are real or fabricated, the product pricing agent will perform the following steps: Execute a pricing action: Set a price based on a pricing strategy (possibly an exploratory strategy or a safety baseline). Observe actual sales / profit: In the next cycle, obtain real sales and profit data. Feedback to the meta-strategy decision center: Update the product status score; for example, if sales remain zero, the product may enter a slow-moving state, and the score will decrease. Update sales model predictions: Retrain the sales prediction model with new real data to make it more accurate. Update data accumulation.
[0071] Afterwards, the product pricing AI will generate a complete transfer sample {s}. t ,a t ,r t ,s t+1}. Among them, s t Indicates the current state (price, inventory, market sentiment, etc.), a t Indicates the price action taken, r t Indicates the reward (profit) received, s t+1 This indicates the new state transitioned to. This sample is stored in the experience replay buffer for subsequent conservative offline training.
[0072] Following this, while in the exploration phase, the pricing agent can perform safe online exploration (trying new prices within a safe price range). If not in the exploration phase, the pricing agent can execute a safe baseline strategy (using known, superior pricing). Upon reaching the offline training period, the pricing agent can trigger offline training. If not, the pricing agent can continue collecting data online, returning to the beginning of the process, and repeating the cycle continuously.
[0073] In some embodiments, the advertising bidding agent may have a second price-volume model, and the advertising bidding agent may include a budget allocator and a real-time bidder. The advertising bidding agent may be configured to perform corresponding pricing decisions under security constraints based on collaborative decision-making instructions, including: Based on collaborative decision-making instructions, the advertising bidding agent can execute high-level strategies through the budget allocator: determining the budget for each advertising campaign based on the global state of the target product's advertising; and execute low-level strategies through the real-time bidder: determining the keyword bid adjustment amount and activation probability in the target product's advertising based on the budget of each advertising campaign.
[0074] During the safe exploration period, the advertising bidding agent can determine the advertising bid for the target product based on the budget of each advertising campaign, the adjustment amount and activation probability of keyword bids in the target product's advertisement, and the second price-volume model within the preset dynamic bidding safety domain.
[0075] Alternatively, if not in a safe exploration period, the advertising bidding agent can execute a safe baseline strategy to determine the advertising bid for the target product.
[0076] The advertising bidding agent stores the advertising consumption and conversion data generated by the advertising bids for the target products into the experience replay buffer in the e-commerce full-link revenue collaborative decision-making agent.
[0077] In some embodiments, the internal structure of the advertising bidding agent is designed as a hierarchical decision-making system to address the coexistence of macro-level resource planning and micro-level real-time game theory in advertising bidding scenarios. This hierarchical decision-making system can consist of a budget allocator, a real-time bidder, and a second price-volume model as the core of strategy evaluation, and achieves a dynamic balance between robust operation and active learning by introducing a safe exploration period.
[0078] In this embodiment, the budget allocator and the real-time bidder undertake tasks at different time scales and decision levels. As a high-level strategy module, the budget allocator has a relatively long decision cycle (e.g., in hours or days). The core responsibility of the budget allocator is to perform strategic resource allocation across advertising campaigns, receive virtual budget allocations from the meta-strategy decision center, and comprehensively consider the historical performance of each advertising campaign (such as ROI trends), current stage goals (such as new user acquisition or user activation), and competitive landscape to decompose the total budget into budget packages for different activities, products, or channels.
[0079] The real-time bidding engine, acting as the underlying execution module, operates within high-frequency decision-making cycles (e.g., every minute or every 15 minutes). Working within the budget constraints set by the budget allocator, the real-time bidding engine focuses on making optimal, instantaneous decisions in the ever-changing bidding market. This includes dynamically adjusting specific bids (i.e., keyword bid adjustments) based on real-time keyword quality scores, competition levels, and conversion rate estimates. Simultaneously, the real-time bidding engine also manages the activation probability of keywords and negative keyword lists; for example, it may attempt to activate new, potentially high-value keywords with a lower probability, or pause high-cost but low-conversion keywords, thereby optimizing the traffic structure.
[0080] The second price-volume model can be a strategy-value prediction model. It learns the long-term value (such as total conversions, clicks under ACOS constraints) that different bidding strategies and on / off actions can bring under given advertising campaign status (such as consumption progress, click rate) and budget constraints, providing a valuable reference for the real-time bidder's decision-making.
[0081] In some embodiments, the e-commerce end-to-end revenue collaborative decision-making agent does not engage in aggressive exploration at all times, but rather confines it to specific, monitored time windows. During the safe exploration period, the advertising bidding agent can, within the dynamic bidding safety domain (whose range may be dynamically adjusted based on ACOS risk and budget consumption rate), and guided by the second price-volume model, attempt exploratory actions such as slightly increasing bids for core keywords or enabling a batch of new keywords.
[0082] During non-exploration periods, the e-commerce end-to-end revenue collaborative decision-making agent can strictly execute a long-term validated and stable safety baseline strategy, such as using a rule-based bidding formula, to ensure the stability of core performance indicators. This alternating exploration and utilization mechanism allows the e-commerce end-to-end revenue collaborative decision-making agent to both accumulate new knowledge and ensure the reliability of daily operations. All advertising consumption and conversion data generated from exploration and baseline execution, regardless of immediate results, are stored in the experience replay buffer, providing training samples for the evolution of the second-price-volume model.
[0083] In some embodiments, the advertising bidding agent can be configured to perform conservative offline training based on the advertising consumption and conversion data stored in the experience replay buffer of the e-commerce full-link revenue collaborative decision-making agent within a historical time period, in order to update the second price-volume model, and feed the update results back to the meta-strategy decision center, which then updates the global market state.
[0084] Similarly, conservative offline training was also applied to the iteration of the ad bidding agent, but it was adapted to the specific business characteristics. The training data for the ad bidding agent also came from the experience replay buffer, which recorded detailed ad consumption and conversion data. Unlike product pricing scenarios, ad interactions typically provide immediate feedback such as impressions and clicks.
[0085] However, in this embodiment, to reduce the over-reliance of the second price-volume model on aggressive strategies that occasionally yield high returns but carry significant long-term risks (e.g., drastically increasing bids across the board during a certain period), a loss function is designed for the conservative offline training of the second price-volume model of the advertising bidding agent. Specifically, the training algorithm includes a conservative regularization term, which mathematically produces two effects: first, it tends to lower the expected value estimate for all possible actions; second, it relatively raises the value estimate for actions that frequently occur in historical data and bring stable returns. In short, the second price-volume model assigns a lower value score to high-risk bidding actions and budget allocation schemes that are not sufficiently present in the data, especially those that may lead to a surge in ACOS or excessive budget depletion. For example, even if there is a record in the historical data showing that increasing the bid for a certain keyword by 200% in a certain exploration resulted in a large number of clicks, the conservative offline training will not easily consider this a high-value action unless similar strategies have been successfully verified multiple times at different times and under different conditions.
[0086] Through conservative offline training, the strategies learned by the advertising bidding agent will naturally lean towards robustness, tending to combine proven effective bid adjustments, budget allocations, and keyword management methods, while remaining cautious in unknown areas. After the second price-volume model is updated, its new behavioral characteristics (such as more stable ACOS control and stricter adherence to budget constraints) will be reported as feedback signals to the meta-strategy decision center. Based on this, the meta-strategy decision center can update its understanding of the risk appetite and effectiveness boundaries of advertising channels in the global market state, thereby adjusting the target weight coefficients or virtual budget allocations sent to the advertising bidding agent in the next cycle, forming a closed loop of global strategy optimization based on safe evolutionary feedback.
[0087] Continue reading Figure 2 In some specific application scenarios, for advertising bidding agents, the goal is to collect new data while ensuring safety (not exceeding the budget and not causing performance failure).
[0088] During the security exploration period, the advertising bidding agent executes the following security online exploration strategy steps: First, continuous action sampling. In the dynamic security domain... Internal sampling of a continuous action a continuous .in, These are the reference actions given by the security baseline policy. and The first is the offset of the upper and lower bounds, used to limit actions within a safe range. The second is discrete action sampling. Flipping a switch with a low probability (e.g., 5%) represents a discrete action selection. This means that most of the time, the e-commerce end-to-end revenue collaborative decision-making agent uses continuous actions, but occasionally tries a discrete action that may have high risk or high reward. The third is sampling distribution calculation. Define a probability distribution. πe Used for sampling in a continuous action space: The first term is the Gaussian term. Used to encourage actions to be closer to safe actions. (To ensure security), the second term: value function term This is used to encourage the selection of high-value actions. It is a temperature parameter used to balance exploration and utilization.
[0089] During non-exploration periods, the advertising bidding agent can directly execute the safety baseline policy. s (t), no exploration is performed. After the action is executed, the environment returns to the new state s. t+1 and reward r t+1 The reward is usually a comprehensive indicator (such as profit considering ACOS constraints). Then the sample {s} t ,a t ,s t+1 ,r t+1 Stored in the experience replay buffer for subsequent conservative offline training.
[0090] Furthermore, for conservative offline training, the goal is to utilize the collected data to update the policy network and the second valence model offline, thereby improving overall performance. Specifically, the first step is to determine whether the offline training period has been reached. If it has, the offline training process begins; otherwise, online data collection continues.
[0091] Understandably, the experience replay buffer only contains data collected from historical policies (which may be safe but mediocre). For actions that never occur or occur very rarely (i.e., "out-of-distribution" actions), Conservative Q-Learning tends to give blindly optimistic overestimations (overestimation). To address this, during the offline training process, a batch of historical data (states, actions, rewards, new states) is sampled from the experience replay buffer. The total loss L is calculated. total Using timing difference error loss L TD The goal is to teach an intelligent agent how to predict the value of actions based on data to improve long-term advertising effectiveness (reduce ACOS), using a loss function L. CQL The forced advertising bidding agent respects historical experience and remains vigilant against unknown actions. Then, the parameters in the second valence model are updated via gradient descent, resulting in the updated conservative Q-network. It can more accurately and conservatively assess the value of actions. The updated policy network outputs new, better, and safer action distributions. It outputs new safety baselines and safety domains, transforming the learned policies into safe action benchmarks and fluctuation ranges that can be directly executed in the online phase.
[0092] In some embodiments, the global market status may also include the target product's status label, historical price and volume data, real-time advertising metrics, and competitor prices.
[0093] The meta-strategy decision-making center can be configured to generate collaborative decision-making instructions based on the global market state. This includes: determining the target product's lifecycle stage based on the global market state, and making a target routing decision for the target product. Based on the target routing decision, the meta-strategy decision-making center performs product-level value assessment, channel-level efficiency assessment, and joint optimization allocation based on opportunity cost, generating collaborative decision-making instructions.
[0094] In some embodiments, the process by which the meta-strategy decision-making center generates collaborative decision-making instructions is a complete cognitive-decision chain, from multi-source information fusion to refined strategy solving. As the global market state becomes more three-dimensional due to the integration of richer dimensions (such as real-time advertising metrics and competitor price matrices), the decision-making logic of the meta-strategy decision-making center also deepens, evolving from a simple lifecycle-based mapping to a systematic calculation that includes three stages: routing, evaluation, and optimization.
[0095] First, determining the target routing decision is a classification and retrieval process. The meta-strategy decision center retrieves the corresponding basic strategy configuration from a predefined or dynamically learned strategy mapping table based on the product's lifecycle stage, promotional objectives, and other status tags. This target routing decision can be viewed as a preset template, initially indicating whether resources should be biased towards growth or profit at the current stage, and whether marketing should focus on advertising or discounts, etc.
[0096] However, relying solely on static templates is insufficient to cope with dynamic markets. Therefore, the meta-strategy decision-making center initiates a real-time quantitative evaluation process based on this.
[0097] Specifically, commodity-level valuation aims to indicate which commodities are more worthy of investment among all commodities. The meta-strategy decision center estimates the marginal global benefit gain from adding a unit of budget to each commodity through, for example, an internal value network or predictive model, thereby ranking all commodities to be managed by value.
[0098] Channel-level performance evaluation aims to indicate which marketing method is currently more efficient for this product. The meta-strategy decision center calculates the ROI of advertising channels and discount channels in near real-time. The former is based on the actual consumption and conversion data fed back by the advertising bidding agent, while the latter relies on the price reduction promotion effect prediction given by the product pricing agent based on the price-volume model.
[0099] Ultimately, the joint optimization allocation based on opportunity cost is the meta-strategy decision-making center's process of combining the outputs of the first two steps—the relative value of each product, the current effectiveness of each channel—along with global budget constraints, product inventory, and strategic importance, into a constrained optimization. The goal is to maximize expected global returns given resources, specifically the target weight coefficient and virtual budget amount that should be allocated to each product across advertising and discount channels. Essentially, this process involves making optimal asset allocation decisions among multiple investment targets (products) and two investment tools (channels), ensuring that every budget flows to the product-channel combination with the highest expected marginal return.
[0100] In some embodiments, the product lifecycle stages include a new product launch period, a growth period, a maturity period, and a clearance / promotion period. Within each of these lifecycle stages, the target product has different routing decisions; each routing decision has a corresponding target weight coefficient and virtual budget allocation.
[0101] In some embodiments, the new product launch period, growth period, maturity period, and clearance / promotion period each correspond to distinct business objectives, market challenges, and resource constraints, thus requiring highly differentiated collaborative decision-making instructions.
[0102] This differentiation is not achieved through manual configuration, but rather through automated management of pre-defined, lifecycle-stage-bound target routing decisions. Each routing decision is a data structure containing default parameter biases and strategic logic. For example, the routing decision for the cold start period of a new product might be set up as follows: it heavily relies on advertising channels for market exploration and data collection, thus assigning the advertising bidding agent a very high weight for traffic acquisition and a relatively generous initial budget, while instructing the product pricing agent to adopt a conservative price exploration strategy to avoid prematurely damaging the price image. However, the routing decision for the clearance / promotion period completely shifts its logic: the core objective is to quickly reduce inventory, thus assigning the product pricing agent a very high weight for sales volume and a clear discount budget, while the task of the advertising bidding agent simultaneously changes to accurately drive traffic in conjunction with price reductions, and its target weight coefficient might be adjusted to "conversion rate" rather than simply "click rate".
[0103] When the meta-strategy decision-making center determines that a product has entered a certain lifecycle stage, it automatically activates the corresponding target routing decision. This decision not only directly provides initial suggested values or strong inclinations for target weight coefficients and virtual budget allocation, but more importantly, it defines the weights and key considerations for various parameters in subsequent product-level value assessments and channel-level performance assessments at that stage. This enables the entire e-commerce end-to-end revenue collaborative decision-making intelligence agent to automatically switch its operational focus and strategy based on the product's stage, achieving adaptive synchronization between operational strategies and the product lifecycle.
[0104] In some embodiments, the meta-strategy decision center can be configured to perform commodity-level value assessment, including: the meta-strategy decision center determining the commodity value of the target commodity based on the marginal value benefit of the target commodity.
[0105] Among them, the commodity value of the target commodity represents the importance of the investment. The higher the commodity value, the more worthwhile the investment is, and vice versa. The marginal value return is directly proportional to the investment budget of the target commodity.
[0106] In some embodiments, commodity-level value assessment is a key computational step in the meta-strategy decision-making center to achieve precise resource allocation and avoid equal distribution. Its core is calculating the marginal value benefit of each target commodity and deriving the commodity value accordingly.
[0107] Marginal value is a dynamic and forward-looking indicator. It measures how much incremental revenue the entire product portfolio is expected to generate if an additional unit of marketing budget (whether for advertising or discounts) is added to a product under current market conditions and system strategies. This estimation is typically performed through a value network within the meta-strategy decision-making center, which is trained to simulate the cascading effects of budget allocation changes on long-term outcomes. For example, for a strategic product with significant inventory backlog and nearing the end of the season, increasing the promotional budget may not only boost sales of the product itself but also free up warehousing costs and recoup funds for new product launches, potentially resulting in a high marginal value. Conversely, for a star product with stable sales and high profit margins, the incremental revenue from a large additional budget may be limited, resulting in a lower marginal value.
[0108] The value of a product, calculated based on its marginal value benefit, is a standardized score or ranking used for cross-product comparisons. A higher product value means that, at the current moment, marketing investment in that product is expected to contribute more to the overall value, making it a more worthwhile investment. This assessment is crucial because it ensures that, with a limited total budget, the meta-strategy decision-making center can prioritize resources to products that generate the greatest global marginal return, thereby improving the overall efficiency of budget utilization. This assessment is dynamic and is updated in real time based on changes in factors such as inventory, competitor actions, and sales progress.
[0109] In some embodiments, the meta-strategy decision center is configured to perform channel-level performance evaluation, including: the meta-strategy decision center determining the return on investment of the target product in the advertising channel and the discount channel, respectively, based on the advertising channel performance and the discount channel performance of the target product.
[0110] The advertising channel effectiveness of the target product is determined by the incremental profit and cost generated by advertising within a specified time period, obtained by the advertising bidding agent through the meta-strategy decision-making center. The discount channel effectiveness is determined by the expected incremental profit and discount budget of the product, obtained by the product pricing agent through the meta-strategy decision-making center.
[0111] In some embodiments, channel-level performance evaluation is the core basis for investment selection by the meta-strategy decision-making center. It conducts independent efficiency audits of advertising channels and discount channels, calculates real-time return on investment, and determines which marketing method is more cost-effective for a specific product in the current market environment.
[0112] To evaluate the effectiveness of advertising channels, the meta-strategy decision center obtains the incremental profit (i.e., advertising-related sales after deducting product costs) and the corresponding total advertising expenditure for a product within a recent time window (e.g., the past 24 hours). The ratio (or a smoothed ratio) of these two figures constitutes a direct measure of advertising channel effectiveness. This metric reflects in real time how efficient the purchase of traffic is under the current competitive landscape, platform algorithms, and user attention. If this value continues to decline, it means that advertising costs are rising or conversion rates are decreasing, and the attractiveness of the advertising channel is diminishing.
[0113] When assessing the effectiveness of discount channels, the meta-strategy decision center, through the product pricing agent, performs the following actions when considering whether to promote sales through price reductions: Assuming a price reduction equivalent to a certain amount (discount budget), the agent predicts, based on its first price-volume model, how much additional sales and incremental profit will be generated. The ratio of the expected incremental profit calculated based on this prediction to the assumed discount budget is the estimated value of the discount channel effectiveness. This reflects consumers' price sensitivity to the product. This value may be high during clearance sales or in response to competition; however, during periods of strong brand premium, indiscriminate price reductions may be ineffective, and this value will be lower.
[0114] By comparing these two performance metrics in parallel, the meta-strategy decision center can clearly determine whether, for this product, it is more cost-effective to invest money in advertising platforms to acquire new customers, or to directly offer discounts to consumers to stimulate conversion of existing traffic. This provides crucial quantitative input for subsequent joint optimization.
[0115] In some embodiments, the meta-policy decision center can be configured to perform joint optimization allocation based on opportunity cost, including: The meta-strategy decision-making center determines the target weight coefficient and virtual budget allocation based on the target product's value, advertising channel effectiveness, and discount channel effectiveness.
[0116] Opportunity cost-based joint optimization allocation is the final convergence step in the meta-strategy decision-making central decision chain. It integrates the results of product-level value assessment and channel-level performance assessment, solving for the optimal collaborative decision-making instruction within a unified mathematical framework.
[0117] In some embodiments, the core idea of this process is opportunity cost. Each unit of budget, if allocated to advertising for product A, cannot simultaneously be allocated to discounting for product B. Therefore, the decision must compare the expected returns of all possible product-channel combinations. The meta-strategy decision center treats product value as a priority weight relative to the product itself and channel effectiveness as an efficiency multiplier for using that channel on a particular product. Joint optimization is the process of finding the optimal solution to an objective function (e.g., maximizing the weighted sum of effectiveness for all products) while satisfying global total budget constraints, individual product inventory constraints, and other constraints.
[0118] In this optimization problem, the decision variables are the target weight coefficients and the virtual budget allocation to be output. For example, the optimization result might show that most of the budget should be allocated to the advertising channel for product X because product X has high value and its current advertising effectiveness is excellent; at the same time, a small portion of the budget should be allocated to the discount channel for product Y because product Y, although of moderate value, has exceptionally high predictive effectiveness for discount promotions (possibly due to seasonal demand); and for product Z, no additional budget should be allocated for the time being because the current effectiveness of both its channels is below the investment threshold. The target weight coefficients obtained at the final solution will reflect this resource allocation tendency, for example, giving the advertising bidding agent for product X a higher weight for traffic acquisition, and giving the product pricing agent for product Y a higher weight for sales volume. Through this rigorous optimization calculation, the collaborative decision-making instruction transforms from empirical strategy selection to a data-driven, globally optimal, and precise output.
[0119] By implementing the above solution, this e-commerce end-to-end revenue collaborative decision-making intelligent agent can significantly improve overall operational revenue. Specifically, it achieves dynamic collaboration between daily product pricing and real-time advertising bidding, thereby continuously driving overall business revenue towards a better state in complex market environments and overcoming the revenue bottleneck caused by fragmented decision-making in related technologies.
[0120] See Figure 3 , Figure 3 This is a flowchart illustrating a collaborative decision-making method provided in an embodiment of this application.
[0121] This application embodiment also provides a collaborative decision-making method for an e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning. The e-commerce end-to-end revenue collaborative decision-making intelligent agent includes: a meta-policy decision center, a product pricing intelligent agent, and an advertising bidding intelligent agent; the method includes steps S101-S103: S101: The meta-strategy decision-making center generates collaborative decision-making instructions based on the global market state. These instructions coordinate the decisions between the product pricing agent and the advertising bidding agent. The global market state includes at least the target product's lifecycle stage and the decision execution results reported by the product pricing agent and the advertising bidding agent, respectively. S102: The commodity pricing agent executes the corresponding pricing decision according to the collaborative decision instruction and feeds back the execution result of the pricing decision to the meta-strategy decision center.
[0122] S103: The advertising bidding agent executes the corresponding bidding decision according to the collaborative decision instruction and feeds back the execution result of the bidding decision to the meta-strategy decision center.
[0123] This application also provides an electronic device, which includes the e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning as described in any of the above claims, for implementing the collaborative decision-making method of the e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning as described in any of the above claims.
[0124] This application also provides a computer-readable medium storing instructions that, when executed on a server, cause the server to perform the collaborative decision-making method for an e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning mentioned in any of the above embodiments of this application.
[0125] This application also provides a computer program product, including: a computer program / instruction, which, when executed by a processor, implements the collaborative decision-making method of the e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning mentioned in any of the above embodiments of this application.
[0126] See Figure 4 , Figure 4 This is a structural block diagram of an electronic device provided in an embodiment of this application.
[0127] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method provided in the above embodiments.
[0128] The electronic device may include: a memory 110, a processor 120, and a communication interface 130. The memory 110, the processor 120, and the communication interface 130 are connected through internal connection paths.
[0129] The memory 110 is used to store computer programs, which in some implementations may include code for implementing the methods of the embodiments of this application.
[0130] The processor 120 executes the computer program stored in the memory 110 to control the communication interface 130 to receive input data and information, and output operation results and other data. In some implementations, when the solutions of the embodiments of this application are implemented by software or firmware, the computer program used to implement the solutions of the embodiments of this application can be stored in the processor 120 and executed by the processor 120.
[0131] The memory 110 may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM). It should be noted that the memory 110 described herein is intended to include, but is not limited to, any memory of these and other suitable types. As an example, the memory 110 includes random access memory (RAM), cache memory, and read-only memory (ROM). The memory 110 stores a computer program that can be executed by processor 120, causing processor 120 to implement the steps of any of the methods described above.
[0132] The processor 120 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor 120 can be any conventional processor.
[0133] In implementation, each step of the above method can be completed by the integrated logic circuitry of the hardware in the processor 120 or by instructions in software form. The method disclosed in the embodiments of this application can be directly implemented by the hardware processor, or by a combination of hardware and software modules in the processor 120. The software modules can be located in mature storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in the memory 110, and the processor 120 reads the information in the memory 110 and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, detailed descriptions are not provided here.
[0134] In some implementations, in addition to the hardware units described above, electronic devices may also include software modules, such as operating systems, basic input / output systems (BIOS), and application software.
[0135] An operating system is used to manage one or more hardware and software resources of an electronic device; it is the kernel and foundation of the electronic device. The operating system handles fundamental tasks such as managing and configuring memory, determining the priority of system resource allocation and demand, controlling input and output devices, operating networks, and managing file systems. To facilitate user operation, most operating systems provide a user interface for interaction with the system.
[0136] The BIOS is used to perform hardware initialization during the power-on boot phase and to provide runtime services for the operating system and applications. In some implementations, the BIOS can also monitor and display processor temperature and execute temperature protection strategies.
[0137] Application software, also known as an application program, can be understood as software written for a specific user application purpose, and is one of the main categories of computer software. For example, application software can be a program used to achieve purposes such as power control and temperature management.
[0138] It is understood that the specific examples in this application are only intended to help those skilled in the art better understand the implementation of this application, and are not intended to limit the scope of protection of this application.
[0139] It is understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of this application.
[0140] It is understood that the various implementation methods described in this application can be implemented individually or in combination, and this application does not limit them.
[0141] Unless otherwise stated, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The term "one or more" as used in this application includes any and all combinations of one or more of the associated listed items. The singular forms "a," "the," and "the" as used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0142] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0143] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes and beneficial effects of the embodiments described above can be referred to the corresponding processes and beneficial effects in other embodiments, and will not be repeated here.
[0144] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0145] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the technical solution in this application, depending on actual needs.
[0146] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0147] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, essentially, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0148] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A collaborative decision-making agent for the entire e-commerce revenue chain based on security reinforcement learning, characterized in that, The e-commerce full-link revenue collaborative decision-making intelligent agent includes: a meta-strategy decision-making center, a product pricing intelligent agent, and an advertising bidding intelligent agent; The meta-strategy decision-making center is configured to generate collaborative decision-making instructions based on the global market state. These instructions are used to coordinate the decisions between the product pricing agent and the advertising bidding agent. The global market state includes at least the lifecycle stage of the target product and the decision execution results fed back by the product pricing agent and the advertising bidding agent, respectively. The commodity pricing intelligent agent is configured to execute corresponding pricing decisions according to the collaborative decision instructions and feed back the execution results of the pricing decisions to the meta-strategy decision center; The advertising bidding agent is configured to execute corresponding bidding decisions according to the collaborative decision instructions and feed back the execution results of the bidding decisions to the meta-strategy decision center.
2. The e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning according to claim 1, characterized in that, The collaborative decision-making instructions include target weight coefficients and virtual budget allocation; The target weighting coefficient indicates that, at the life cycle stage of the target product, the pricing of the target product focuses on profit or sales volume, and the advertising bid for the target product focuses on traffic or return on investment. The virtual budget allocation instruction dynamically allocates virtual resource quotas for advertising consumption and product discounts under global budget constraints.
3. The e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning according to claim 2, characterized in that, The product pricing agent and the advertising bidding agent are configured to make decisions under security constraints, and both adopt a framework that combines secure online exploration with conservative offline training for iterative optimization. The execution results of the pricing decision include product sales and profit data, and the execution results of the bidding decision include advertising consumption and conversion data.
4. The e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning according to claim 3, characterized in that, The product pricing agent has a first price-quantity model, and is configured to execute corresponding pricing decisions under security constraints according to the collaborative decision-making instructions, including: The product pricing intelligence agent performs a security exploration within a preset dynamic pricing security domain based on the first price-volume model, determines the pricing of the target product, and stores the product sales and profit data generated from sales based on the pricing of the target product in the experience replay buffer of the e-commerce full-link revenue collaborative decision-making intelligence agent. The sales and profit data include data generated from both unsold and sold items of the target product.
5. The e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning according to claim 3, characterized in that, The advertising bidding agent has a second price-volume model, and the advertising bidding agent includes a budget allocator and a real-time bidder; the advertising bidding agent is configured to execute corresponding pricing decisions under security constraints according to the collaborative decision-making instructions, including: According to the collaborative decision-making instructions, the advertising bidding agent executes a high-level strategy through the budget allocator: determining the budget for each advertising campaign based on the global state of the target product's advertising; and executes a low-level strategy through the real-time bidder: determining the keyword bid adjustment amount and activation probability in the target product's advertising based on the budget of each advertising campaign. During the safe exploration period, the advertising bidding agent determines the advertising bid for the target product based on the budget of each advertising campaign, the keyword bid adjustment amount and activation probability in the target product's advertisement, and the second price-volume model, within a preset dynamic bidding safety domain; or, If the period is not in the safe exploration phase, the advertising bidding agent executes a safety baseline strategy to determine the advertising bid for the target product; The advertising bidding agent stores the advertising consumption and conversion data generated from the advertising bids for the target product in the experience replay buffer of the e-commerce end-to-end revenue collaborative decision-making agent.
6. The e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning according to claim 4, characterized in that, The product pricing agent is configured to perform conservative offline training based on the product sales and profit data stored in the experience replay buffer of the e-commerce full-link revenue collaborative decision-making agent within a historical time period, in order to update the first price-volume model, and feed the update result back to the meta-strategy decision center, which then updates the global market state.
7. The e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning according to claim 5, characterized in that, The advertising bidding agent is configured to perform conservative offline training based on the advertising consumption and conversion data stored in the experience replay buffer of the e-commerce full-link revenue collaborative decision-making agent within the historical time period, so as to update the second price-volume model, and feed the update result back to the meta-strategy decision center, which then updates the global market state.
8. The e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning according to claim 1, characterized in that, The global market status also includes the target product's status label, historical price and volume data, real-time advertising metrics, and competitor prices; The meta-strategy decision-making center is configured to generate collaborative decision-making instructions based on the global market state, including: The meta-strategy decision center determines the target route decision for the target product based on the life cycle stage of the target product as determined by the global market status. The meta-strategy decision-making center, based on the target routing decision, performs product-level value assessment, channel-level efficiency assessment, and joint optimization allocation based on opportunity cost, and generates the collaborative decision-making instruction.
9. The e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning according to claim 8, characterized in that, The product lifecycle stages include the new product launch period, growth period, maturity period, and clearance / promotion period; within different product lifecycle stages, the target product has different routing decisions; each routing decision has a corresponding target weight coefficient and virtual budget allocation.
10. The e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning according to claim 8, characterized in that, The meta-strategy decision-making center is configured to perform the commodity-level value assessment, including: The meta-strategy decision-making center determines the commodity value of the target commodity based on its marginal value benefit. The value of the target commodity represents the importance of the investment; the higher the commodity value, the more worthwhile the investment, and vice versa. The marginal value return is proportional to the investment budget for the target commodity.
11. The e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning according to claim 10, characterized in that, The meta-strategy decision-making center is configured to perform the channel-level performance evaluation, including: The meta-strategy decision center determines the return on investment of the target product in the advertising channel and the discount channel based on the effectiveness of the advertising channel and the discount channel, respectively. The advertising channel effectiveness of the target product is determined by the incremental profit and consumption generated by the advertising of the target product within a specified time period, obtained by the meta-strategy decision center through the advertising bidding agent. The effectiveness of the discount channel is determined by the expected incremental profit and discount budget of the product obtained by the product pricing agent through the meta-strategy decision center.
12. The e-commerce end-to-end revenue collaborative decision-making intelligent agent based on security reinforcement learning according to claim 11, characterized in that, The meta-policy decision center is configured to perform joint optimization allocation based on opportunity cost, including: The meta-strategy decision-making center determines the target weight coefficient and the virtual budget allocation based on the target product's value, advertising channel effectiveness, and discount channel effectiveness.