E-commerce operation strategy optimization method based on multi-AI agent cooperation

By employing a multi-AI agent collaborative e-commerce operation strategy optimization method, which utilizes deep reinforcement learning and real-time data analysis, the problem of low operational efficiency in traditional e-commerce operation methods is solved. This enables the generation of precise and automated operation strategies, thereby improving market responsiveness and marketing effectiveness.

CN121921041APending Publication Date: 2026-04-24SHENZHEN MAGIC FISH FASHION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN MAGIC FISH FASHION TECHNOLOGY CO LTD
Filing Date
2025-11-20
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Traditional e-commerce operation methods lack a dynamic adjustment mechanism based on real-time data, resulting in low operational efficiency, an inability to respond promptly to market changes, and a lack of refined and personalized user behavior analysis, making it difficult to maximize marketing effectiveness.

Method used

The e-commerce operation strategy optimization method based on multi-AI agent collaboration classifies operation scenarios by acquiring historical user behavior data and product sales peak data, constructs state vectors and decision models, uses deep reinforcement learning for real-time strategy generation and iterative optimization, and combines extreme scenario validation sets to evaluate the model's adaptability and timeliness.

Benefits of technology

It enables precise and automated e-commerce operation strategies, allowing for rapid response to demands and fluctuations in different scenarios, improving operational efficiency, enhancing system robustness and responsiveness, increasing user conversion rates and marketing effectiveness, and reducing operating costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921041A_ABST
    Figure CN121921041A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of e-commerce marketing, in particular to an e-commerce operation strategy optimization method based on multi-AI agent collaboration, which comprises the following steps: acquiring historical user behavior data and commodity sales peak data of a target e-commerce platform, performing operation scene classification on the target e-commerce platform based on the historical user behavior data and the commodity sales peak data to obtain at least one e-commerce operation sub-scene; and obtaining real-time operation monitoring data of at least one e-commerce operation sub-scene, and performing feature extraction on the real-time operation monitoring data to obtain a state vector corresponding to the e-commerce operation sub-scene. According to the method, the user behavior data, the commodity sales data and the like of the e-commerce platform are acquired in real time, a customized operation strategy can be accurately generated for each e-commerce operation sub-scene based on historical data and real-time monitoring data, the strategies can accurately cope with requirements and fluctuations in different scenes, and the operation efficiency is maximized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of e-commerce marketing technology, specifically to a method for optimizing e-commerce operation strategies based on multi-AI agent collaboration. Background Technology

[0002] With the acceleration of globalization, the market has gradually become highly competitive. In the past, companies may have relied on large-scale offline stores and traditional sales channels to occupy market share. However, this model has many limitations, such as high costs, inefficient inventory management, and limited customer reach.

[0003] Currently, traditional e-commerce operation methods mostly rely on preset rules and fixed operational strategies, lacking a dynamic adjustment mechanism based on real-time data. Therefore, when market demand changes or unexpected events occur, traditional methods often cannot react in time, resulting in low operational efficiency and even missed opportunities. Moreover, they usually require a lot of human intervention and judgment, especially when facing complex e-commerce environments and volatile market conditions, where human decision-making and adjustments are difficult to be accurate and efficient.

[0004] Furthermore, traditional methods may lack sufficient refinement and personalization in user behavior analysis and marketing strategy formulation, making it difficult to maximize marketing effectiveness. They also often fail to effectively integrate data from multiple channels and dimensions, resulting in an inability to comprehensively analyze market dynamics and user behavior. Summary of the Invention

[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for optimizing e-commerce operation strategies based on multi-AI agent collaboration, comprising: Obtain historical user behavior data and peak product sales data of the target e-commerce platform, and classify the target e-commerce platform into operational scenarios based on the historical user behavior data and peak product sales data to obtain at least one e-commerce operation sub-scenario. Obtain real-time operation monitoring data for at least one of the e-commerce operation sub-scenarios, extract features from the real-time operation monitoring data, and obtain the state vector corresponding to the e-commerce operation sub-scenarios. The pre-constructed basic decision model is trained based on at least one of the state vectors to obtain the specific decision model corresponding to the e-commerce operation sub-scenario. At least one of the special decision-making models is evaluated based on a pre-built extreme scenario validation set to obtain a comprehensive evaluation result for the corresponding special decision-making model. Based on at least one of the comprehensive evaluation results, the parameters of at least one of the specific decision-making models are fused to obtain the global optimal parameters. Based on the global optimal parameters, the at least one of the specific decision-making models are iteratively updated to obtain the optimized decision-making model corresponding to the e-commerce operation sub-scenario. Obtain at least one real-time state vector of the e-commerce operation sub-scenario, input the at least one real-time state vector into the corresponding optimization decision model to generate a strategy, and obtain a precise operation strategy for at least one e-commerce operation sub-scenario.

[0006] Preferably, the target e-commerce platform is classified into operational scenarios based on the historical user behavior data and the peak sales data of the goods, resulting in at least one e-commerce operation sub-scenario, including: Obtain historical user behavior data from the target e-commerce platform, and perform statistical analysis on the historical user behavior data to obtain a user behavior concentration index; Obtain peak sales data of the target e-commerce platform, and perform trend fitting on the peak sales data to obtain the sales fluctuation coefficient; Based on the user behavior concentration index and combined with the category association constraints of the target e-commerce platform, an initial scenario segmentation result is obtained by performing a preliminary scenario segmentation of the target e-commerce platform. If the sales fluctuation coefficient exceeds the preset fluctuation range, the scenario boundary of the initial scenario division result is redefined to obtain at least one e-commerce operation sub-scenario.

[0007] Preferably, the pre-constructed basic decision model is trained based on at least one of the aforementioned state vectors to obtain a specialized decision model corresponding to the e-commerce operation sub-scenario, including: Deploy scenario-based intelligent agents for each e-commerce operation sub-scenario. Each scenario-based intelligent agent includes an online decision-making network and a target decision-making network. Define the state space and action space of the scene intelligent agent in different e-commerce operation sub-scenarios, and establish a special decision model for each scene intelligent agent.

[0008] Preferably, the state space and action space of the scene intelligent agent are defined in different e-commerce operation sub-scenarios, including: In response to the e-commerce operation sub-scenario being a major promotion preparation scenario, a state space for the major promotion preparation period is constructed based on user accumulation data, product inventory data, marketing budget data, and historical major promotion conversion data for each e-commerce operation sub-scenario. All control actions taken by the scenario intelligent agent during the major promotion preparation period are defined, including user reach frequency adjustment actions, product inventory adjustment actions, and marketing budget allocation actions, forming a major promotion preparation period action space. In response to the e-commerce operation sub-scenario being a major promotional event scenario, a state space for the major promotional event is constructed based on real-time user access data, real-time product inventory data, real-time marketing conversion data, and user demand fluctuation data within the e-commerce operation sub-scenario under the jurisdiction of each scenario's intelligent agent. The user reach strategy adjustment amount, product inventory transfer amount, and marketing placement ratio adjustment amount are defined as the action space for the major promotional event. In response to the e-commerce operation sub-scenario being a real-time operation scenario, a real-time state space is constructed based on the real-time order data, real-time inventory data, and real-time user behavior data of the e-commerce system during the real-time operation phase. The real-time action space is defined as the emergency transfer volume of goods, the temporary marketing replenishment volume, and the user priority reach volume.

[0009] Preferably, a specific decision-making model is established for each scenario agent, including: Initialize the online decision-making network, the target decision-making network, and the operational experience replay pool; The online decision network randomly selects an operational strategy from the action space and calculates the operational feedback after the selected operational strategy is executed in each operational unit. The operational feedback includes the sales growth rate, cost saving ratio, and user conversion improvement rate. Based on the operational status and state transition probability corresponding to the operational strategy, and combined with operational feedback, the next operational status is predicted; The operational strategies, operational feedback, operational status, and the next operational status are combined into operational experience and stored in the operational experience replay pool. The online decision-making network and the target decision-making network are trained based on the operational experience to update the parameters of the decision-making model; The online decision network randomly selects an operational strategy from the initial operational strategy set again, calculates operational experience, and stores it in the operational experience replay pool. The online decision network and the target decision network are trained based on the operational experience, and the parameters of the special decision model are updated. The above steps are repeated until the preset convergence threshold is met for a preset number of consecutive times.

[0010] Preferably, the online decision network and the target decision network are trained based on the operational experience to update the parameters of the decision model, including: The online decision network randomly extracts operational experience from the operational experience replay pool, calculates the strategy benefit value of the current operational state based on the operational experience, and the target decision network calculates the maximum strategy benefit value of the next operational state based on the operational experience. Calculate the target strategy return value based on the strategy return value under the current operating state and the maximum strategy return value under the next operating state; The loss value is calculated based on the target policy return value and the policy return value predicted by the online decision network, and the network parameters of the decision model are updated based on the gradient descent method.

[0011] Preferably, at least one of the specific decision-making models is evaluated based on a pre-built extreme scenario validation set to obtain a comprehensive evaluation result for the corresponding specific decision-making model, including: Obtain a validation set of extreme scenarios for sudden emergencies in e-commerce; Perform state vector adaptability testing on the extreme scenario validation set to obtain adaptability evaluation results for at least one of the specific decision-making models; Based on the extreme scenario validation set, the timeliness of the policy response of at least one of the special decision-making models is evaluated to obtain the timeliness evaluation results of the corresponding special decision-making models; The at least one adaptability assessment result and at least one timeliness assessment result are weighted and fused to obtain a comprehensive assessment result corresponding to the specific decision-making model.

[0012] Preferably, based on at least one of the comprehensive evaluation results, parameter fusion is performed on at least one of the specialized decision-making models to obtain globally optimal parameters. Then, based on the globally optimal parameters, at least one of the specialized decision-making models is iteratively updated to obtain an optimized decision-making model corresponding to the e-commerce operation sub-scenario, including: A normalization function is used to convert at least one of the comprehensive evaluation results into corresponding parameter fusion weights; Based on at least one of the parameter fusion weights, the core parameters of at least one of the special decision-making models are weighted and integrated to obtain the globally optimal parameters; Based on the global optimal parameters, parameter migration and fusion are performed on at least one of the specific decision models to obtain the optimized decision model corresponding to the e-commerce operation sub-scenario.

[0013] Preferably, at least one of the real-time state vectors is input into the corresponding optimization decision model to generate a strategy, thereby obtaining a precise operation strategy for at least one of the e-commerce operation sub-scenarios, including: At least one of the real-time state vectors is input into the corresponding optimization decision model to generate a strategy, thereby obtaining an initial operation strategy for at least one of the e-commerce operation sub-scenarios; Obtain real-time inventory dynamic data and user interaction feedback data for at least one of the e-commerce operation sub-scenarios, and dynamically modify at least one of the initial operation strategies based on at least one of the real-time inventory dynamic data and user interaction feedback data to obtain the intermediate operation strategy corresponding to the e-commerce operation sub-scenarios. Based on at least one of the aforementioned intermediate operation strategies, and combined with the marketing resource quotas of the e-commerce platform, a precise operation strategy corresponding to the aforementioned e-commerce operation sub-scenario is generated.

[0014] Preferably, based on at least one of the aforementioned real-time inventory dynamic data and user interaction feedback data, at least one of the aforementioned initial operational strategies is dynamically modified to obtain an intermediate operational strategy corresponding to the aforementioned e-commerce operational sub-scenario, including: Obtain real-time inventory dynamic data and user interaction feedback data for at least one of the aforementioned e-commerce operation sub-scenarios; The feasibility of at least one of the initial operation strategies is verified based on at least one of the real-time inventory dynamic data and user interaction feedback data. The parameters of the infeasible initial operation strategies are reconstructed to obtain the intermediate operation strategies corresponding to the e-commerce operation sub-scenario.

[0015] Compared with the prior art, the beneficial effects of the present invention are: (1) By acquiring user behavior data and product sales data of e-commerce platforms in real time, this invention can accurately generate customized operation strategies for each e-commerce operation sub-scenario based on historical data and real-time monitoring data. These strategies can accurately respond to the needs and fluctuations in different scenarios, maximize operational efficiency, and by applying deep reinforcement learning to e-commerce operations, the method can generate dynamic adjustment strategies based on real-time state vectors, so that the operation model of each e-commerce operation sub-scenario can continuously optimize itself and gradually approach the global optimum. Through continuous parameter fusion and iterative updates, it can ensure that the model can quickly adapt to market changes, reduce human intervention, and improve the level of automation. (2) In response to emergencies and extreme scenarios, the method can evaluate the adaptability and timeliness of the model through the extreme scenario validation set, ensuring that the operation strategy can respond quickly in emergency or uncertain environments, thereby enhancing the robustness and adaptability of the system. By analyzing data such as user behavior concentration and sales fluctuations, the method can gain a deep understanding of user needs and behavior patterns, thereby providing accurate marketing strategies for promotional preparation and real-time operation. This refined analysis can improve user conversion rate and marketing effect, while reducing operating costs. Attached Figure Description

[0016] Figure 1 This is a schematic flowchart of the overall method in one embodiment of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Example 1, please refer to Figure 1This invention provides a technical solution: an e-commerce operation strategy optimization method based on multi-AI agent collaboration, comprising: S1. Obtain historical user behavior data and peak sales data of the target e-commerce platform. Based on the historical user behavior data and peak sales data, classify the target e-commerce platform into operational scenarios to obtain at least one e-commerce operation sub-scenario. S2. Obtain real-time operation monitoring data for at least one e-commerce operation sub-scenario, extract features from the real-time operation monitoring data, and obtain the state vector of the corresponding e-commerce operation sub-scenario. S3. Train the pre-constructed basic decision model based on at least one state vector to obtain the special decision model for the corresponding e-commerce operation sub-scenario; S4. Evaluate at least one specific decision-making model based on a pre-built extreme scenario validation set to obtain a comprehensive evaluation result for the corresponding specific decision-making model; S5. Based on at least one comprehensive evaluation result, perform parameter fusion on at least one specific decision model to obtain the global optimal parameters, and iteratively update at least one specific decision model according to the global optimal parameters to obtain the optimized decision model for the corresponding e-commerce operation sub-scenario. S6. Obtain the real-time state vector of at least one e-commerce operation sub-scenario, input the at least one real-time state vector into the corresponding optimization decision model to generate a strategy, and obtain the precise operation strategy of at least one e-commerce operation sub-scenario.

[0019] It's important to note that collecting historical data from the target e-commerce platform primarily includes: user behavior data (such as user browsing, clicking, purchasing, and saving); and peak sales data (such as sales data for certain products during specific holidays or promotional events). This data allows us to identify different scenarios during platform operations, such as normal sales, promotional activities, and seasonal changes. These scenarios can be categorized based on different conditions (such as user groups, time periods, and sales peaks) to determine different operational sub-scenarios. For example, when analyzing an electronics e-commerce platform, historical data might show a surge in smartphone sales for certain brands during the "Double 11" shopping festival, while sales remain stable at other times. Therefore, the platform's operational scenarios can be divided into two sub-scenarios: "regular operational scenarios" and "promotional peak operational scenarios." By monitoring real-time platform operational data, such as user traffic, real-time sales, and inventory status, features are extracted to construct a state vector reflecting the platform's current state. This state vector represents the current operational scenario and reflects key indicators during the operation, such as sales revenue and user conversion rate. For example, during peak promotional periods, the following real-time monitoring data might be acquired: Number of users: the number of users browsing products; Sales data: sales revenue within the current time period; Inventory status: whether any products are out of stock; Page views: the number of clicks on the product details page. This data will be converted into a state vector, for example: [5000 (number of users), 20000 (sales revenue), 50 (remaining inventory), 30000 (clicks)], which is the state vector for this scenario. The extracted state vectors are used to train the decision-making model. Each sub-scenario will have a dedicated decision-making model, which will adopt different operational strategies based on different states, gradually optimizing the effectiveness of the strategies. The training process mainly uses reward and penalty mechanisms to allow the model to continuously learn and optimize in various operational environments. For example, assuming it is for the "Double 11" promotion, the model uses real-time state vectors (such as sales volume, number of users, and inventory) to decide how to allocate the marketing budget, adjust product pricing, and recommend products. The model will try different strategies, such as increasing advertising investment and pushing coupons, and adjust the strategies based on the results (such as improving conversion rates). The model is tested using a pre-built extreme scenario validation set to verify its performance under abnormal conditions. Extreme scenarios may include a sudden surge in sales or system failures. Through evaluation, the model is ensured to make effective decisions under various conditions. For example, the model will be tested to see if it can adjust its strategy in time when there is a shortage of inventory or a surge in website traffic. For example, on "Double 11" day, if some popular products are running out of stock, can the model quickly adjust its recommendation strategy or reduce losses by adding inventory warnings? By evaluating the results of various specialized models, including the performance of decision-making models in normal and extreme scenarios, parameters are fused to find the globally optimal parameters. Then, the model is iteratively updated to optimize its decision-making process. For example, if different advertising and pricing strategies are tested during the "Double 11" promotion and it is found that certain strategies perform better under specific circumstances, the model parameters will be updated based on this feedback to ensure that it can make better decisions in similar promotional activities in the future. In the final stage, an optimized decision-making model is used to generate precise operational strategies based on real-time state vectors. These strategies will include decisions on pricing, promotions, and recommendations to maximize the platform's operational efficiency. For example, assuming the optimized model predicts sales trends for the next hour on "Double 11" based on real-time state vectors (such as user numbers, sales revenue, and inventory), it will automatically adjust product recommendations, advertising budgets, and inventory management to ensure maximum operational efficiency for the platform. For instance, during periods of high demand, the model may push best-selling products and increase the advertising budget for those products to ensure continued sales growth.

[0020] In an optional embodiment, the target e-commerce platform is classified into operational scenarios based on historical user behavior data and peak product sales data to obtain at least one e-commerce operation sub-scenario, including: Obtain historical user behavior data from the target e-commerce platform, and conduct statistical analysis on the historical user behavior data to obtain the user behavior concentration index; Obtain peak sales data of products from the target e-commerce platform, and perform trend fitting on the peak sales data to obtain the sales fluctuation coefficient; Based on the user behavior concentration index and combined with the category association constraints of the target e-commerce platform, the initial scenario segmentation results are obtained by performing a preliminary scenario segmentation of the target e-commerce platform. If the sales fluctuation coefficient exceeds the preset fluctuation range, the scenario boundary of the initial scenario division result is redefined to obtain at least one e-commerce operation sub-scenario.

[0021] It's important to note that historical user behavior data from e-commerce platforms includes various user behaviors such as browsing, clicking, searching, purchasing, and saving. Statistical analysis of these behaviors yields an indicator reflecting user behavior concentration—the user behavior concentration index. This index measures the distribution of user behavior on the platform; higher concentration indicates that most users are focused on a few product categories or behaviors, while lower concentration indicates more dispersed user behavior. For example, suppose an e-commerce platform sells a wide variety of products, including mobile phones, home appliances, clothing, and cosmetics. Statistical analysis of user behavior data yields a user behavior concentration index. If the result shows that user behavior is mainly concentrated on mobile phones and home appliances, accounting for 80% of total pageviews, while other products (clothing, cosmetics, etc.) account for only 20%, this indicates a high concentration of user behavior, with most users focusing on a few product categories. The platform can use this information to make initial adjustments to its operational strategies. Peak sales data refers to the peak sales volume of certain products within a specific time period, such as during promotional events or holidays, when sales of some products surge rapidly. By fitting the trend of these sales data, a sales volatility coefficient can be obtained, which represents the amplitude of fluctuations in product sales. For example, suppose that during the "Double 11" shopping festival on an e-commerce platform, the sales volume of a certain mobile phone reached 10 times its normal level within 24 hours. By fitting the trend of the peak sales data of this product, the sales volatility coefficient is found to be 2.5, meaning that the sales fluctuation amplitude is 2.5 times that under normal circumstances. This indicates that the sales volume of this product fluctuates greatly during the promotion period, so its inventory and supply situation need to be closely monitored. After obtaining the user behavior concentration index and sales volatility coefficient, and combining them with category association constraints, a preliminary scenario division of the e-commerce platform is made. Category association constraints refer to the correlation between the sales trends of certain products, such as mobile phones and accessories, or clothing and shoes, which often appear together when users make purchases. The user behavior concentration index is used to determine the degree of concentration of user behavior. Category association constraints help to understand the relationship between different products. Combining these two factors, the e-commerce platform can be divided into different operational sub-scenarios. For example, one scenario might be a "regular sales scenario," while another might be a "promotional peak scenario." For example, assuming the platform's user behavior concentration index is high, it means that most users are only interested in a few types of products. Combining category association data, the platform might be divided into two preliminary scenarios: Regular sales scenario: Users mainly purchase mobile phones, headphones, accessories, etc.; Promotional peak scenario: During promotional festivals like "Double 11," sales are concentrated in specific product categories, such as smart home products and clothing. Based on the initial scenario segmentation, if the sales volatility coefficient of certain products exceeds the preset fluctuation range, the boundaries of these scenarios need to be redefined. This means re-segmenting the scenarios according to the actual sales volatility of the products. The aim is to enable the model to better adapt to actual market fluctuations. For example, suppose that during the "Double 11" promotion, the sales volatility coefficient of certain products (such as smart home products) is found to be much higher than normal, exceeding the preset fluctuation range. Specifically, the sales volatility coefficient of this product is 5, while the preset fluctuation range is 2 to 3. Therefore, it may be necessary to redefine the initially segmented scenarios, potentially dividing the original "regular sales scenario" into multiple sub-scenarios, including: high-volatility sales scenario: for products with large sales fluctuations, special inventory management, pricing, and promotional strategies are adopted; low-volatility sales scenario: for products with small sales fluctuations, conventional sales strategies are adopted. Through this adjustment, the platform can better cope with the sales fluctuations of different types of products, ensuring maximum operational efficiency.

[0022] In an optional embodiment, a pre-built basic decision model is trained based on at least one state vector to obtain a specialized decision model for the corresponding e-commerce operation sub-scenario, including: Deploy scenario-based intelligent agents for each e-commerce operation sub-scenario. Each scenario-based intelligent agent includes an online decision-making network and a target decision-making network. Define the state space and action space of the scene intelligent agent in different e-commerce operation sub-scenarios, and establish a special decision model for each scene intelligent agent.

[0023] It's important to note that in different operational sub-scenarios of an e-commerce platform, a scenario-based intelligent agent needs to be deployed for each sub-scenarios. This agent is implemented through a decision-making model, primarily consisting of two parts: an online decision-making network (ADIN) and a goal-oriented decision-making network. The ADIN is responsible for making decisions at every moment, selecting an action based on the current state and feeding it back to the environment. These decisions are typically made in a real-time operational environment. The goal-oriented decision-making network provides the target value (i.e., the future goal). It's a long-term policy network, usually used to estimate the long-term returns of the optimal policy, helping the agent optimize its decision path. For example, suppose the e-commerce platform divides the "Double 11" promotional period into two... Sub-scenarios: Regular sales scenarios, such as non-promotional periods, and promotional sales scenarios, such as during "Double 11". For each scenario, a scenario-specific intelligent agent will be deployed: For regular sales scenarios, the online decision-making network will decide whether to adjust pricing or run ads based on real-time information such as current inventory levels, user demand, and pageviews; For promotional sales scenarios, the online decision-making network will formulate promotional strategies based on factors such as product sales fluctuations, user click-through rates, and traffic, such as limited-time discounts and buy-one-get-one-free promotions; The target decision-making network will provide the target direction for decision-making based on historical data and long-term goals (such as total sales and user satisfaction) and guide the intelligent agent to continuously optimize strategies; Each scenario-based intelligent agent has two key components: state space and action space. State space represents the agent's current environmental state. In the e-commerce platform example, state space might include current user behavior data, product inventory, order status, traffic data, and pricing information. Action space contains the actions or behaviors the agent can choose. Based on the current state, the agent will decide what action to take according to its strategy. Action space can include pricing adjustments, advertising, changes in promotional strategies, and inventory adjustments. For example, for a scenario-based intelligent agent in a typical sales scenario: State space includes current product inventory, real-time user pageviews, current product price, and the last promotion... The system considers factors such as the effectiveness of the system and user search behavior; the action space includes, for example, adjusting product prices, optimizing ad display, increasing or decreasing ad budgets, and issuing coupons; for scenario-based intelligent agents in promotional sales scenarios (such as during "Double 11"), the state space includes product inventory levels, traffic conditions, peak sales traffic data, user interaction behavior, current discount levels, and the effectiveness of promotional activities; the action space includes, for example, increasing product discounts, setting up limited-time flash sales, periodically pushing coupons, and adjusting inventory allocation; by defining suitable state and action spaces for each e-commerce operation sub-scenario, reinforcement learning models can find the optimal operational decisions through continuous trial and error and learning. After defining the state space and action space, a decision model will be used to train the agent specifically for each scenario. "Specific" here refers to each sub-scenario having a dedicated training process and model optimization path. The core idea of ​​reinforcement learning is to continuously adjust the agent's behavior through rewards and penalties to maximize returns. In e-commerce platforms, returns typically refer to goals such as profit, user satisfaction, and conversion rates. For example, for a specific decision model for regular sales scenarios, the goal is to teach the agent how to optimize pricing and advertising strategies based on user behavior and inventory levels to increase sales and profits. After each decision (such as adjusting prices or launching promotions), the platform will reward or penalize the agent based on actual sales results (such as increased or decreased sales), thus continuously adjusting the strategy. However, for a specific decision model for promotional sales scenarios (such as during "Double 11"), the agent will face more uncertainties, such as the complexity of promotional activities, traffic fluctuations, and competitor strategies. The model will adjust activity details through trial-and-error decisions and feedback, such as setting optimal discounts and selecting the best advertising timing, to achieve higher sales and user engagement.

[0024] In an optional embodiment, the state space and action space of the scene intelligence agent under different e-commerce operation sub-scenarios are defined, including: Responding to the e-commerce operation sub-scenario as the promotion preparation scenario, based on the user accumulation data, product inventory data, marketing budget data, and historical promotion conversion data of each e-commerce operation sub-scenario, a state space for the promotion preparation period is constructed; all control actions taken by the scenario intelligent agent during the promotion preparation period are defined, including user reach frequency adjustment actions, product inventory quantity adjustment actions, and marketing budget allocation actions, forming the promotion preparation period action space. In response to the e-commerce operation sub-scenario being a major promotional event, the state space for the major promotional event is constructed based on the real-time user access data, real-time product inventory data, real-time marketing conversion data, and user demand fluctuation data within the e-commerce operation sub-scenario under the jurisdiction of each scenario's intelligent agent. The action space for the major promotional event is defined as the user reach strategy adjustment amount, product inventory transfer amount, and marketing placement ratio adjustment amount. In response to the e-commerce operation sub-scenario being a real-time operation scenario, a real-time state space is constructed based on the real-time order data, real-time inventory data, and real-time user behavior data of the e-commerce system during the real-time operation phase. The real-time action space is defined as the amount of emergency transfer of goods, the amount of temporary marketing replenishment, and the amount of priority user outreach.

[0025] It's important to note that during the preparation period for a major promotional event, e-commerce platforms need to prepare and plan based on several key factors. Specifically, the state space includes the following important data: User engagement data refers to the user participation, registration, and browsing data accumulated by a platform during pre-sale marketing activities aimed at attracting users, such as advertising and pre-sales. This data reflects users' interest in and preparation for the upcoming promotional event.

[0026] Product inventory data: Based on past sales data and projected demand, the platform prepares and reserves product inventory. For example, it shows which products have been prepared in what quantities to ensure sufficient supply during major promotional periods.

[0027] Marketing budget data: refers to the budget that e-commerce platforms invest during the preparation period for major promotions in order to increase exposure and attract users, including advertising expenses, promotional activity budgets, etc.

[0028] Historical promotional conversion data: This includes conversion rates, sales figures, and user activity data from similar past promotional periods. This data can help the platform predict potential performance during future promotional periods.

[0029] Action space: During the preparation phase, the scene-based intelligent agent needs to make adjustments based on the current state. Specific adjustment actions include: User reach frequency adjustment: This refers to adjusting the frequency of reaching potential customers through different channels (such as SMS, email, APP push, etc.) before a major promotion to ensure the widespread dissemination of marketing information.

[0030] Product inventory adjustment: Adjust product inventory based on anticipated demand and market trends during the promotion to ensure that products can meet expected sales demand.

[0031] Marketing Budget Allocation: Based on the current preparation phase, decide how to allocate the marketing budget across different channels and activities. For example, should advertising investment be increased for certain products, or promotional efforts be intensified for certain product categories? Specifically, consider this scenario: Assume an e-commerce platform plans to sell smartphones during the "Double 11" shopping festival. The platform's statistics show that approximately 20% of users have already added items to their shopping carts during the pre-sale phase. The platform anticipates needing to stock 5,000 smartphones, but based on data from the past two years, peak sales typically occur around 8 PM. Therefore, additional stock is needed for peak periods. Merchandise; The marketing budget has been allocated 3 million yuan, with 70% for social media advertising and 30% for search engine advertising; Action plan: User reach frequency adjustment: The agent decided to increase the frequency of SMS and APP push notifications from once a week to three times a week at the beginning of the preparation period; Merchandise inventory adjustment: Based on historical data, the agent decided to increase the inventory of smartphones from 5,000 units to 6,000 units to cope with demand fluctuations; Marketing budget allocation: The agent decided to increase the budget for social media advertising by 10% and reduce the budget for search advertising by 5% to improve the reach of potential users; During peak sales periods (such as "Double 11"), the intelligent agent needs to make adjustments and decisions based on real-time data. Specific state spaces include: Real-time user access data: tracking user visits, page views, and dwell time during the promotion; this data reflects user engagement and interest in specific products; Real-time product inventory data: the platform's real-time inventory status for each product during the promotion; inventory data directly affects product availability and order processing; Real-time marketing conversion data: conversion rates generated through different channels (such as advertising, promotional activities, coupons, etc.) during the promotion, including click-through rates and purchase rates; User demand fluctuation data: referring to changes in user demand during the promotion, such as a sudden surge in demand due to certain product categories becoming popular; Action space: During peak sales periods, the scenario-based intelligent agent mainly takes the following control actions: User reach strategy adjustment: the agent adjusts user reach strategies based on real-time data, such as increasing advertising pushes for certain products or adjusting reach frequency; Product inventory transfer: based on inventory status and user demand, the agent decides to transfer certain best-selling products from low-demand areas to high-demand areas or accelerate replenishment; Marketing Spending Adjustment: The agent adjusts the spending ratio of different marketing channels based on real-time marketing performance data, such as increasing the budget for a certain ad slot or reducing the spending of ineffective ads. Specific examples include: State Space: On the day of a major promotion, page views on smartphones increased by 300% compared to usual, user demand surged, but inventory was insufficient; smartphone ad conversion rates were low, with only 5% of clicks converting into actual purchases, and users were more inclined to choose other brands after comparing prices; due to the great success of certain promotional activities (such as limited-time discounts), the inventory of some products was almost sold out. Action Space: User Reach Strategy Adjustment: The agent decides to increase the frequency of push notifications about smartphones to twice per hour and send personalized reminders to specific user groups (such as users who added items to their cart but did not purchase); Inventory Transfer: Based on real-time inventory data, the agent decides to transfer the remaining smartphone inventory from low-demand areas (such as second-tier cities) to first-tier cities to ensure that inventory can meet high demand; Marketing Spending Adjustment: Based on real-time marketing performance, the agent decides to increase the budget for search ads by 20% while reducing the budget for ineffective ads. State Space: In real-time operation scenarios, the e-commerce platform's intelligent agent needs to make decisions based on real-time operational data at every moment. The state space includes the following data: Real-time order data: including real-time data during order generation, payment, and delivery; this data reflects the actual situation of user purchases; Real-time inventory data: the current inventory level of goods, helping the intelligent agent understand whether emergency replenishment or allocation is needed; Real-time user behavior data: user behavior data on the platform, including browsing products, adding to cart, clicking to purchase, etc., helping the intelligent agent assess user interests and purchasing trends; Action Space: In real-time operation scenarios, the scenario's intelligent agent needs to take the following regulatory actions based on real-time data: Emergency allocation of goods: based on inventory data, the intelligent agent decides whether goods need to be allocated from other warehouses or regions to ensure supply chain stability; Temporary marketing replenishment: when it is found that the sales of certain products are lower than expected, the intelligent agent may decide to temporarily increase... Marketing budget: Boost sales through specific ads or coupons; Prioritized user outreach: In some cases, the agent can decide to prioritize outreach to specific user groups based on user behavior and preferences, such as sending limited-time offers to potential high-value customers; Specific examples: State space: Smartphone orders increase significantly starting at 8 AM, but only 100 units remain in stock, expected to sell out within 2 hours; User behavior data indicates strong interest in smartphone accessories (such as phone cases), but only 50 units are in stock; Action space: Emergency product allocation: The agent decides to allocate 200 smartphones from a nearby warehouse to meet high demand; Temporary marketing replenishment: To address the demand for phone cases, the agent decides to increase the marketing budget by 100,000 yuan to promote bundled sales of phone cases; Prioritized user outreach: The agent decides to send personalized promotional messages to users who have previously purchased phones, reminding them to buy accessories.

[0032] In an optional embodiment, a specialized decision-making model is established for each scenario agent, including: Initialize the online decision-making network, the target decision-making network, and the operational experience replay pool; The online decision network randomly selects an operational strategy from the action space and calculates the operational feedback after the selected operational strategy is executed in each operational unit. The operational feedback includes the sales growth rate, cost saving ratio, and user conversion improvement rate. Based on the operational status and state transition probability corresponding to the operational strategy, and combined with operational feedback, predict the next operational status; Operational strategies, operational feedback, operational status, and the next operational status are combined into operational experience and stored in the operational experience replay pool. The online decision-making network and the target decision-making network are trained based on operational experience to update the parameters of the decision-making model; The online decision network randomly selects an operational strategy from the initial operational strategy set again, calculates operational experience, and stores it in the operational experience replay pool. The online decision network and the target decision network are trained based on the operational experience, and the parameters of the special decision model are updated. The above steps are repeated until the preset convergence threshold is met for a preset number of consecutive times.

[0033] It's important to note that the online decision network is a deep neural network model responsible for selecting the optimal operational strategy based on the current operational status, such as product recommendation and advertising strategies. It continuously learns and updates to make better decisions in actual operations. The target decision network is similar to the online decision network, but it's a "lagging" network, typically used to stabilize the training process. The parameters of the target decision network are periodically synchronized from the online decision network to reduce fluctuations during training and make learning smoother. The operational experience replay pool is a database storing past operational experiences, including past "experiences"—states, actions, feedback, and the next state. By replaying these experiences, the model can learn more efficiently, avoiding excessive reliance on the immediate feedback of the current strategy. For example, suppose you're operating an e-commerce platform; the online decision network decides whether to promote a certain product, and the target decision network acts as its "backup," periodically updating and synchronizing the latest strategies from the online decision network. The experience replay pool stores all data from past promotional activities, such as sales volume before and after the promotion, user conversion rates, and costs. Action space is the set of all possible operations the model can choose; in an e-commerce platform example, actions might include selecting promotional items, choosing discount levels, and choosing ad display methods; Operational feedback: the feedback on the effect of implementing a strategy, usually quantified by some indicators; such as sales growth rate, cost savings rate, and user conversion rate improvement; for example: suppose the online decision network randomly selects "offer a 10% discount on a certain mobile phone product" as the strategy; then, the following feedback is calculated: Sales growth rate: 10%, sales increased by 10 after the promotion compared to before the promotion; Cost savings rate: 5%, transportation costs were reduced by optimizing logistics; User conversion rate improvement: 8%, more potential customers were attracted by promotional advertising; Operational status: refers to the current... The set of all variables in the environment, such as product inventory, user activity, and competitors' pricing strategies; state transition probability: the probability of how the current state will transition to the next state after executing a certain operational strategy; for example, the probability of increased sales after a promotion, or the probability of user registration after advertising; prediction of the next operational state: based on the current strategy and feedback, predicting the possible operational results after executing the strategy, such as predicting changes in sales volume and inventory after a promotion; for example: the operational strategy chosen by the online decision network is "to offer a 10% discount on a certain mobile phone"; based on previous experience and the current state, such as inventory and user activity, the possible state transitions after executing this strategy are predicted to be: a 10% increase in sales, a 5% decrease in inventory, and an 8% increase in user activity; Store the feedback and status obtained in the previous step as an "experience" record; this experience data will be used for subsequent model training. For example, store the following data in the experience replay pool: Current operational status: 5,000 units of a certain mobile phone product are in stock, user activity is moderate, and competitors are conducting similar promotions; Operational strategy: Offer a 10% discount on this product; Operational feedback: Sales growth of 10%, cost savings of 5%, and user conversion rate improvement of 8%; Next state prediction: Sales increase of 10%, inventory reduction of 5%, and user activity improvement of 8%; By leveraging operational experience from the replay pool, deep reinforcement learning algorithms are used to adjust and optimize the parameters of the online decision network and the target decision network, enabling the networks to more accurately predict and select the optimal strategy in the future. For example, the online decision network updates its weights based on the experience data stored in the replay pool, learning how to better predict the sales growth and cost savings brought about by promotional strategies; the target decision network is updated synchronously on a regular basis to avoid overfitting. The convergence threshold is the stopping condition for model training. It usually refers to the point at which the model's performance reaches a certain stable state after multiple training iterations and no longer shows significant improvement. For example, the accuracy of the model's strategy selection or the growth rate of revenue reaches the expected target after several updates. For instance, after 100 strategy selections and updates, the model's sales forecast accuracy has stabilized and reached the set performance standard, which means that the model has converged and can begin actual operation.

[0034] In an optional embodiment, the online decision network and the target decision network are trained based on operational experience to update the parameters of the decision model, including: The online decision network randomly selects operational experience from the operational experience replay pool and calculates the strategy reward value of the current operational state based on the operational experience. The target decision network calculates the maximum strategy reward value of the next operational state based on the operational experience. Calculate the target strategy return value based on the strategy return value under the current operating state and the maximum strategy return value under the next operating state; The loss value is calculated based on the target policy return value and the policy return value predicted by the online decision network, and the network parameters of the decision model are updated based on the gradient descent method.

[0035] It's important to note that the online decision-making network is a neural network model that makes decisions based on the current operational state. In deep reinforcement learning, the training process uses past experiences to improve the decision-making process. The operational experience replay pool is a database storing historical experiences, including information such as (state, action, feedback, next state). The online decision-making network randomly selects one experience from this pool as the basis for current training. For example, suppose we operate an e-commerce platform. The replay pool stores different promotional strategies and their corresponding sales performance. For instance, the experience of a certain promotional activity includes: Current state: 1000 units of a certain product in stock, medium user activity, competitors have promotions; Action performed: Offer a 20% discount on the product; Feedback received: Sales increased by 30% after the promotion; Next state: Stock decreased to 700 units, user activity increased to high. The online decision-making network randomly selects such experiences from the replay pool. For each extracted experience, the online decision network needs to calculate the payoff value of executing a certain strategy under the current operating state. This payoff value can be calculated in various ways, such as expected payoff based on historical data or short-term payoff predicted by a model. For example, assuming the scenario of the e-commerce platform mentioned above, the online decision network will calculate the payoff value of the strategy based on the current state and the promotional strategy. If a 20% discount is applied, the expected payoff is a 30% increase in sales volume. However, due to the reduced revenue from the discount and advertising costs, the actual net payoff might be 20%. Therefore, the payoff value of the strategy under the current operating state is 20%. The target decision network is a lagging neural network used to avoid instability caused by rapid updates in the online decision network during training. Based on the next state, the target decision network calculates the maximum policy reward that can be obtained in that state. Typically, it estimates the rewards of all possible actions in the next state and selects the maximum value. For example, in the e-commerce platform example above, the next state is that inventory has decreased to 700 units and user activity has increased significantly. The target decision network evaluates possible strategies based on this new state. For example, it might evaluate the following possible actions: continue the promotion, which is expected to bring an additional 20% sales increase; adjust advertising, which is expected to bring a 15% sales increase. The target decision network calculates that in this new state, the maximum policy reward is 20%, i.e., the maximum reward from continuing the promotion. The target policy payoff is calculated based on the current policy and the maximum policy payoff in the next state. The target policy payoff is typically updated using the Bellman equation. Specifically, it is calculated by combining the payoff of the current policy with the maximum payoff in the next state. The loss value is calculated by comparing the difference between the target policy payoff and the policy payoff predicted by the online decision network. This loss value guides model updates. Typically, the loss function is mean squared error or another suitable metric. By calculating the loss value, the model uses gradient descent to update its network parameters. Gradient descent is an optimization algorithm that minimizes the loss value by adjusting the weights of the neural network. The purpose of the update is to enable the online decision network to better predict the policy's payoff in the next operation, thereby improving decision quality. For example, in an e-commerce platform scenario, with a loss value of 5.29, the decision model uses gradient descent to adjust the weights in the neural network, reducing future prediction errors and enabling the online decision network to more accurately assess the benefits of the promotional strategy in the next selection.

[0036] In an optional embodiment, at least one specialized decision-making model is evaluated based on a pre-built extreme scenario validation set to obtain a comprehensive evaluation result for the corresponding specialized decision-making model, including: Obtain a validation set of extreme scenarios for sudden emergencies in e-commerce; Perform state vector adaptability testing on the extreme scenario validation set to obtain adaptability evaluation results for at least one specific decision model; Based on the extreme scenario validation set, the timeliness of the policy response of at least one special decision-making model is evaluated to obtain the timeliness evaluation results of the corresponding special decision-making model. The results of at least one adaptability assessment and at least one timeliness assessment are weighted and fused to obtain the comprehensive assessment result of the corresponding special decision-making model.

[0037] It's important to note that the extreme scenario validation set is a specially designed test set containing extreme or limit situations, used to examine the performance of the decision-making model in dealing with special or unexpected circumstances. In e-commerce platforms, these unexpected scenarios might include traffic surges, sudden inventory changes, or unexpected large-scale promotions by competitors. For example, assuming you're running an online store, the extreme scenario validation set for e-commerce unexpected scenarios might include the following scenarios: Traffic surge: A sudden influx of users visiting the store, with traffic increasing fivefold compared to normal; Inventory shortage: The inventory of a popular product suddenly runs out; Competitor surprise attack: A competitor launches a promotional campaign with significant discounts, leading to a sudden drop in sales. These scenarios constitute the "extreme scenario validation set," used to verify the model's decision-making ability under these abnormal conditions. State vector adaptability testing verifies whether a model can adapt to state vectors (i.e., current environmental information) in extreme scenarios and make reasonable decisions. A state vector is a feature vector describing the current operational state, containing information such as inventory, sales, user activity, and price. Adaptability testing assesses whether the model can quickly understand and correctly respond to these "extreme" states. For example, when an e-commerce platform suddenly experiences a surge in traffic, the state vector might include: number of users: a significant increase; order volume: a rapid increase; server response time: an increase. Whether the model can understand this sudden scenario, adapt to the state change, and provide reasonable promotional strategies or inventory adjustment measures is the content of adaptability assessment. Strategy response timeliness assessment evaluates the model's reaction speed under extreme scenarios. In actual operations, rapid response is crucial when facing unexpected events. Timeliness assessment focuses on the time required for the model to take new actions (such as adjusting prices or optimizing advertising) from receiving a new situation (e.g., a surge in traffic or inventory shortage). For example, suppose an e-commerce platform encounters a competitor launching a major promotion. The model needs to assess this unexpected situation within a short time and adjust prices or promotional strategies accordingly. If the model calculates a new pricing strategy and updates the page within one minute, its timeliness is good; if the model needs 10 minutes to update the strategy, its timeliness is poor. Therefore, timeliness assessment checks the model's reaction speed to ensure that decisions can be made quickly in a highly competitive market environment. When evaluating a specific decision-making model, simply considering adaptability and timeliness is insufficient; these two evaluation results need to be combined to obtain a comprehensive assessment. This can be achieved through weighted fusion, which involves weighting the importance of adaptability and timeliness according to the actual situation to arrive at a comprehensive score. For example, suppose we are evaluating a decision-making model for an e-commerce platform. In a certain e-commerce scenario, timeliness (response speed) is more important than adaptability (the ability to correctly understand extreme states) because the market changes rapidly, and timely responses can bring more revenue. In this case, the timeliness evaluation result can be given a higher weight (e.g., 0.7), while the adaptability evaluation result can be given a lower weight (e.g., 0.3).

[0038] In an optional embodiment, at least one specialized decision-making model is fused based on at least one comprehensive evaluation result to obtain globally optimal parameters. The at least one specialized decision-making model is then iteratively updated according to the globally optimal parameters to obtain an optimized decision-making model for the corresponding e-commerce operation sub-scenario, including: A normalization function is used to transform at least one comprehensive evaluation result into corresponding parameter fusion weights; The core parameters of at least one specialized decision-making model are weighted and integrated based on at least one parameter fusion weight to obtain the globally optimal parameters. Based on the globally optimal parameters, parameter transfer and fusion are performed on at least one specialized decision-making model to obtain the optimized decision-making model for the corresponding e-commerce operation sub-scenario.

[0039] It's important to note that the comprehensive evaluation results need to be transformed into weights that can be used for optimization. The normalization function converts the comprehensive evaluation results into standardized weight values, making it easier to weight the parameters of different models. For example, suppose there are two evaluation results: a suitability evaluation result of 80 points and a timeliness evaluation result of 90 points. After comprehensive evaluation, the resulting comprehensive evaluation result is 87 points (as calculated earlier). This score needs to be converted into an appropriate weight value. For example, a normalization function (such as minimum-maximum normalization) can be used to map the comprehensive evaluation result to a weight range (e.g., [0,1]). If the normalized score is 0.87, then this means that the evaluation result has an importance of 87% in the final model parameter fusion. Once the comprehensive evaluation results are converted into weights, these weights can be used to weight and integrate the core parameters in the decision model. Each decision model may have multiple parameters, such as learning rate, discount factor, and exploration rate. By weighting and integrating these parameters, a globally optimal parameter set can be obtained, which performs well across multiple evaluation dimensions. For example, suppose a decision model has the following core parameters: learning rate: initial value 0.01, discount factor: initial value 0.95, exploration rate: initial value 0.1. If the normalized comprehensive evaluation result is 0.87, and this evaluation result is timely... With higher weights (e.g., 0.7) and lower fitness weights (e.g., 0.3), the values ​​of these parameters can be adjusted based on these weights: Learning rate: Due to the importance of timeliness, adjusting the learning rate may enable the model to converge faster when facing unexpected events; it can be adjusted to 0.015 to increase the model's reaction speed; Discount factor: Considering the importance of fitness assessment, the discount factor is adjusted to 0.92 to balance the weights of current and long-term rewards, ensuring the model can better adapt to unexpected scenarios; Exploration rate: Due to the importance of timeliness, the exploration rate may be reduced to 0.05 to reduce excessive exploration and ensure the model quickly executes the optimal action; Through these weighted adjustments, a new set of globally optimal parameters is finally obtained. The parameter transfer fusion process involves applying previously obtained globally optimal parameters to different specialized decision-making models. Through parameter transfer, the performance of these models is optimized in different e-commerce operation sub-scenarios. Transfer learning helps the model apply the experience gained in one sub-scenario to another similar sub-scenario, thereby accelerating training and improving the model's generalization ability. For example, suppose two specialized decision-making models have been trained: an inventory management model (optimizing inventory levels and product replenishment) and a pricing strategy model (optimizing product pricing strategies to increase sales). After the weighted integration and acquisition of globally optimal parameters, a set of parameters is obtained, such as a learning rate of 0.015 and a discount factor of 0.92. These parameters have already performed well in the inventory management scenario. These optimal parameters are applied to the pricing strategy model. Through parameter transfer fusion, this pricing strategy model can leverage the successful experience of the inventory management model and be optimized in new pricing strategy sub-scenarios. Specific steps include: parameter transfer: applying globally optimal learning rate, discount factor, and other parameters to the pricing strategy model so that the model can adapt to market changes more quickly in pricing decisions; fusion optimization: combining the training data and experience of the two models, using transfer learning techniques to optimize the pricing strategy and improve the model's ability to cope with different price adjustment strategies on e-commerce platforms; ultimately, an optimized decision model will be obtained, capable of making better decisions in different sub-scenarios of e-commerce operations (such as inventory management and pricing strategies).

[0040] In an optional embodiment, at least one real-time state vector is input into a corresponding optimization decision model to generate a strategy, thereby obtaining a precise operation strategy for at least one e-commerce operation sub-scenario, including: Input at least one real-time state vector into the corresponding optimization decision model to generate a strategy, thereby obtaining an initial operation strategy for at least one e-commerce operation sub-scenario; Obtain real-time inventory dynamic data and user interaction feedback data for at least one e-commerce operation sub-scenario, and dynamically modify at least one initial operation strategy based on at least one real-time inventory dynamic data and user interaction feedback data to obtain the intermediate operation strategy for the corresponding e-commerce operation sub-scenario. Based on at least one intermediate operation strategy, and combined with the marketing resource quota of the e-commerce platform, a precise operation strategy is generated for the corresponding e-commerce operation sub-scenario.

[0041] It's important to note that the real-time state vector is input into the optimized Deep Reinforcement Learning (DRL) model. This state vector may contain the current state of the e-commerce platform, such as inventory, sales data, and user behavior data. Based on this input, the decision-making model generates a preliminary operational strategy—the optimal decision or action plan for the current state. For example, consider the sub-scenario of managing product pricing and inventory on an e-commerce platform. In this scenario, the real-time state vector could include: the current inventory quantity of the product (e.g., 100 units); the current sales rate of the product (e.g., 20 units sold in the past hour); the current price of the product (e.g., 10 yuan / unit); and user interaction data for the product, such as 50 users viewing the product in the past hour. After inputting this data into the decision-making model, the model generates an initial operational strategy based on historical experience and optimized parameters. For example, the model might suggest: price adjustment: raising the price of the product to 12 yuan to increase profits; replenishment: based on sales rate prediction, suggesting replenishing 50 units to meet upcoming demand. After obtaining the initial operational strategy, it is necessary to adjust the strategy based on real-time data feedback. This real-time data includes inventory dynamics data and user interaction feedback data. Inventory dynamics data can include changes in product inventory, while user interaction feedback data includes user browsing, clicking, and purchasing behavior data. This real-time data helps to dynamically revise the initial strategy, that is, to continuously adjust decisions based on new information during strategy execution. For example, suppose the initial strategy is to raise the price of a certain product to 12 yuan and suggest replenishing 50 units. Next, the platform obtains the following feedback data in real time: Inventory dynamics data: 5 units of this product have been sold in the past two hours. 0 items remaining, 50 items in stock; User interaction feedback data: Although 50 users viewed the product, the purchase rate is low, possibly due to a decrease in purchase intent caused by the price increase; Based on this data, the model will dynamically adjust the initial strategy; For example: Price adjustment: Considering that the price increase has led to a decrease in the purchase rate, the model may suggest adjusting the price back to 10 yuan to improve the purchase conversion rate; Inventory replenishment: Since there are only 50 items left in stock and the sales speed is relatively fast, the model may suggest increasing the replenishment quantity from 50 to 100 items; This revised strategy is called the intermediate operation strategy, which is more accurate than the initial strategy and can reflect the actual operating status of the current e-commerce platform; Based on the intermediate strategy, and combined with the e-commerce platform's marketing resource quota, a final precise operational strategy is generated. Marketing resource quotas may include advertising budgets, coupon distribution limits, and the number of recommended placements. These resources will help the platform rationally allocate resources among different operational strategies to achieve optimal results. For example, continuing with the above example, after deriving the intermediate strategy, the e-commerce platform also needs to consider marketing resources. For instance, the platform has an advertising budget that can be used to promote products. Based on the intermediate operational strategy, the model combines the following resource quotas: Advertising budget: 5000 yuan (can be used for advertising promotion); Coupon quota: 1 coupon issued. 00 coupons (each coupon can reduce the price by 10 yuan); combined with intermediate operational strategies, such as adjusting the price to 10 yuan and suggesting replenishing 100 units, the model will generate a precise operational strategy, such as: advertising: allocate an advertising budget of 5,000 yuan to promote the product, increase exposure and click-through rate; coupon distribution: distribute 100 coupons to potential high-conversion users to attract users to buy through promotion; in this way, the decision model integrates real-time data feedback, marketing resources, and strategy optimization to finally obtain a precise operational strategy, which can more effectively improve the operational effect of the e-commerce platform in this sub-scenario.

[0042] In one optional embodiment, at least one initial operational strategy is dynamically revised based on at least one real-time inventory dynamic data and user interaction feedback data to obtain an intermediate operational strategy corresponding to the e-commerce operation sub-scenario, including: Obtain real-time inventory dynamic data and user interaction feedback data for at least one e-commerce operation sub-scenario; The feasibility of at least one initial operation strategy is verified based on at least one real-time inventory dynamic data and user interaction feedback data. The parameters of the infeasible initial operation strategy are reconstructed to obtain the intermediate operation strategy for the corresponding e-commerce operation sub-scenario.

[0043] It's important to note that the system acquires real-time inventory dynamic data and user interaction feedback data: Inventory dynamic data describes the current inventory status of products, such as the quantity of goods in stock and the rate of inventory change; User interaction feedback data describes user interactions with products, including the quantity of products viewed, the quantity added to the shopping cart, the quantity purchased, and the duration of user interaction. When the initially generated operational strategy is executed, the system uses real-time data to verify the feasibility of the strategy. If the conditions or results of the strategy execution do not match expectations, the strategy is considered infeasible and requires adjustment. Infeasible strategies include, for example, initial strategies that suggest excessive replenishment or overly high price adjustments, leading to excess inventory or low purchase conversion rates, which require verification and correction. For infeasible initial operational strategies, parameter reconstruction is performed: if the strategy does not meet actual operational conditions, the system will reconstruct the parameters in the strategy, adjusting key factors such as price, inventory, and promotional activities to better suit actual needs, resulting in an intermediate operational strategy that is more targeted and operational than the initial strategy. For example: Suppose you are managing an e-commerce platform and responsible for the operation and management of a best-selling product. An initial strategy has been generated, and based on certain assumptions, the following initial operational strategies have been formulated: Product pricing: The product price is 100 yuan; Inventory replenishment: The current inventory is 500 units, and the initial strategy recommends replenishing 500 units to ensure expected demand is met; Promotional activities: A discount activity is planned, with a discount rate of 10%. After implementing the initial strategy, the system will monitor the following data in real time: Inventory dynamic data: The current inventory is 500 units; Sales speed is slow, with 20 units sold in the past 24 hours; User interaction feedback data: In the past 24 hours, 1000 people viewed the product; Only 50 people added the product to their shopping cart, and 10 people ultimately purchased it; The purchase rate is low, especially after the discount activity started, the number of items added to the shopping cart increased, but the number of final purchases did not increase significantly. Based on real-time data, the system verifies the feasibility of the initial operational strategy: the current product price is 100 yuan, while the discount only reduces it by 10 yuan; user feedback data reveals that although user browsing volume has increased, the conversion rate (i.e., the conversion rate from adding to the shopping cart to final purchase) is low due to the high price; the initial pricing strategy has problems, the product price may be too high, and users have not generated sufficient purchasing desire; the product inventory is still sufficient, but due to the slow sales speed, replenishing 500 units seems to be unrealistic; based on current sales data, replenishing 500 units may lead to overstocking; the replenishment quantity is too large, and the actual demand has not met expectations, making the replenishment strategy infeasible; during the discount activity, the number of items added to the shopping cart increased, but the purchase conversion rate did not improve significantly; this may mean that the discount is not enough to stimulate purchasing decisions, or users expect a larger discount; the discount activity is not attractive enough, and the promotional efforts may need to be further strengthened; After feasibility verification, issues were found in the initial strategy, necessitating parameter restructuring: Pricing adjustment: Considering the excessively high price, the model recommends reducing the product price from 100 yuan to 85 yuan to improve the purchase conversion rate; price reductions can attract more price-sensitive users, thereby increasing the conversion rate; Inventory adjustment: Considering the slow sales speed, inventory replenishment is recommended to be reduced from 500 units to 100 units; by analyzing sales data and inventory turnover speed, unnecessary inventory backlog can be reduced, lowering inventory costs; Promotional activity adjustment: To improve the purchase conversion rate, the model recommends increasing the discount from 10% to 20% and adding a limited-time promotion period, such as setting it to be valid for 72 hours, to stimulate users' sense of urgency to purchase; The adjusted intermediate operation strategy is: Price: 85 yuan; Inventory: Replenishment of 100 units; Promotional activity: 20% discount, lasting for 3 days; This intermediate operation strategy is more in line with the current market situation than the initial strategy, better able to respond to changes in inventory and user purchasing behavior, ultimately improving the product's sales conversion rate.

[0044] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited thereto. Various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention.

Claims

1. A method for optimizing e-commerce operation strategies based on multi-AI agent collaboration, characterized in that, include: Obtain historical user behavior data and peak product sales data of the target e-commerce platform, and classify the target e-commerce platform into operational scenarios based on the historical user behavior data and peak product sales data to obtain at least one e-commerce operation sub-scenario. Obtain real-time operation monitoring data for at least one of the e-commerce operation sub-scenarios, extract features from the real-time operation monitoring data, and obtain the state vector corresponding to the e-commerce operation sub-scenarios. The pre-constructed basic decision model is trained based on at least one of the state vectors to obtain the specific decision model corresponding to the e-commerce operation sub-scenario. At least one of the special decision-making models is evaluated based on a pre-built extreme scenario validation set to obtain a comprehensive evaluation result for the corresponding special decision-making model. Based on at least one of the comprehensive evaluation results, the parameters of at least one of the specific decision-making models are fused to obtain the global optimal parameters. Based on the global optimal parameters, the at least one of the specific decision-making models are iteratively updated to obtain the optimized decision-making model corresponding to the e-commerce operation sub-scenario. Obtain at least one real-time state vector of the e-commerce operation sub-scenario, input the at least one real-time state vector into the corresponding optimization decision model to generate a strategy, and obtain a precise operation strategy for at least one e-commerce operation sub-scenario.

2. The e-commerce operation strategy optimization method based on multi-AI agent collaboration as described in claim 1, characterized in that, Based on the historical user behavior data and the peak sales data of the products, the target e-commerce platform is classified into operational scenarios to obtain at least one e-commerce operation sub-scenario, including: Obtain historical user behavior data from the target e-commerce platform, and perform statistical analysis on the historical user behavior data to obtain a user behavior concentration index; Obtain peak sales data of the target e-commerce platform, and perform trend fitting on the peak sales data to obtain the sales fluctuation coefficient; Based on the user behavior concentration index and combined with the category association constraints of the target e-commerce platform, an initial scenario segmentation result is obtained by performing a preliminary scenario segmentation of the target e-commerce platform. If the sales fluctuation coefficient exceeds the preset fluctuation range, the scenario boundary of the initial scenario division result is redefined to obtain at least one e-commerce operation sub-scenario.

3. The e-commerce operation strategy optimization method based on multi-AI agent collaboration according to claim 2, characterized in that, Based on at least one of the aforementioned state vectors, a pre-constructed basic decision model is trained to obtain a specialized decision model corresponding to the aforementioned e-commerce operation sub-scenario, including: Deploy scenario-based intelligent agents for each e-commerce operation sub-scenario. Each scenario-based intelligent agent includes an online decision-making network and a target decision-making network. Define the state space and action space of the scene intelligent agent in different e-commerce operation sub-scenarios, and establish a special decision model for each scene intelligent agent.

4. The e-commerce operation strategy optimization method based on multi-AI agent collaboration according to claim 3, characterized in that, Define the state space and action space of the scene-based intelligent agent in different e-commerce operation sub-scenarios, including: In response to the e-commerce operation sub-scenario being a major promotion preparation scenario, a state space for the major promotion preparation period is constructed based on user accumulation data, product inventory data, marketing budget data, and historical major promotion conversion data for each e-commerce operation sub-scenario. All control actions taken by the scenario intelligent agent during the major promotion preparation period are defined, including user reach frequency adjustment actions, product inventory adjustment actions, and marketing budget allocation actions, forming a major promotion preparation period action space. In response to the e-commerce operation sub-scenario being a major promotional event scenario, a state space for the major promotional event is constructed based on real-time user access data, real-time product inventory data, real-time marketing conversion data, and user demand fluctuation data within the e-commerce operation sub-scenario under the jurisdiction of each scenario's intelligent agent. The user reach strategy adjustment amount, product inventory transfer amount, and marketing placement ratio adjustment amount are defined as the action space for the major promotional event. In response to the e-commerce operation sub-scenario being a real-time operation scenario, a real-time state space is constructed based on the real-time order data, real-time inventory data, and real-time user behavior data of the e-commerce system during the real-time operation phase. The real-time action space is defined as the emergency transfer volume of goods, the temporary marketing replenishment volume, and the user priority reach volume.

5. The e-commerce operation strategy optimization method based on multi-AI agent collaboration according to claim 4, characterized in that, Establish a specialized decision-making model for each scenario's intelligent agent, including: Initialize the online decision-making network, the target decision-making network, and the operational experience replay pool; The online decision network randomly selects an operational strategy from the action space and calculates the operational feedback after the selected operational strategy is executed in each operational unit. The operational feedback includes the sales growth rate, cost saving ratio, and user conversion improvement rate. Based on the operational status and state transition probability corresponding to the operational strategy, and combined with operational feedback, the next operational status is predicted; The operational strategies, operational feedback, operational status, and the next operational status are combined into operational experience and stored in the operational experience replay pool. The online decision-making network and the target decision-making network are trained based on the operational experience to update the parameters of the decision-making model; The online decision network randomly selects an operational strategy from the initial operational strategy set again, calculates operational experience, and stores it in the operational experience replay pool. The online decision network and the target decision network are trained based on the operational experience, and the parameters of the special decision model are updated. The above steps are repeated until the preset convergence threshold is met for a preset number of consecutive times.

6. The e-commerce operation strategy optimization method based on multi-AI agent collaboration according to claim 5, characterized in that, The online decision-making network and the target decision-making network are trained based on the operational experience to update the parameters of the decision-making model, including: The online decision network randomly extracts operational experience from the operational experience replay pool, calculates the strategy benefit value of the current operational state based on the operational experience, and the target decision network calculates the maximum strategy benefit value of the next operational state based on the operational experience. Calculate the target strategy return value based on the strategy return value under the current operating state and the maximum strategy return value under the next operating state; The loss value is calculated based on the target policy return value and the policy return value predicted by the online decision network, and the network parameters of the decision model are updated based on the gradient descent method.

7. The e-commerce operation strategy optimization method based on multi-AI agent collaboration according to claim 6, characterized in that, At least one of the aforementioned specialized decision-making models is evaluated based on a pre-built extreme scenario validation set to obtain a comprehensive evaluation result for the corresponding specialized decision-making model, including: Obtain a validation set of extreme scenarios for sudden emergencies in e-commerce; Perform state vector adaptability testing on the extreme scenario validation set to obtain adaptability evaluation results for at least one of the specific decision-making models; Based on the extreme scenario validation set, the timeliness of the policy response of at least one of the special decision-making models is evaluated to obtain the timeliness evaluation results of the corresponding special decision-making models; The at least one adaptability assessment result and at least one timeliness assessment result are weighted and fused to obtain a comprehensive assessment result corresponding to the specific decision-making model.

8. The e-commerce operation strategy optimization method based on multi-AI agent collaboration according to claim 7, characterized in that, Based on at least one of the comprehensive evaluation results, at least one of the specific decision-making models is fused to obtain globally optimal parameters. Then, based on the globally optimal parameters, at least one of the specific decision-making models is iteratively updated to obtain an optimized decision-making model corresponding to the e-commerce operation sub-scenario, including: A normalization function is used to convert at least one of the comprehensive evaluation results into corresponding parameter fusion weights; Based on at least one of the parameter fusion weights, the core parameters of at least one of the special decision-making models are weighted and integrated to obtain the globally optimal parameters; Based on the global optimal parameters, parameter migration and fusion are performed on at least one of the specific decision models to obtain the optimized decision model corresponding to the e-commerce operation sub-scenario.

9. The e-commerce operation strategy optimization method based on multi-AI agent collaboration according to claim 8, characterized in that, Inputting at least one of the real-time state vectors into the corresponding optimization decision model to generate a strategy, thereby obtaining a precise operation strategy for at least one of the e-commerce operation sub-scenarios, including: At least one of the real-time state vectors is input into the corresponding optimization decision model to generate a strategy, thereby obtaining an initial operation strategy for at least one of the e-commerce operation sub-scenarios; Obtain real-time inventory dynamic data and user interaction feedback data for at least one of the e-commerce operation sub-scenarios, and dynamically modify at least one of the initial operation strategies based on at least one of the real-time inventory dynamic data and user interaction feedback data to obtain the intermediate operation strategy corresponding to the e-commerce operation sub-scenarios. Based on at least one of the aforementioned intermediate operation strategies, and combined with the marketing resource quotas of the e-commerce platform, a precise operation strategy corresponding to the aforementioned e-commerce operation sub-scenario is generated.

10. The e-commerce operation strategy optimization method based on multi-AI agent collaboration according to claim 9, characterized in that, Based on at least one of the aforementioned real-time inventory dynamic data and user interaction feedback data, at least one of the aforementioned initial operational strategies is dynamically revised to obtain intermediate operational strategies corresponding to the aforementioned e-commerce operational sub-scenario, including: Obtain real-time inventory dynamic data and user interaction feedback data for at least one of the aforementioned e-commerce operation sub-scenarios; The feasibility of at least one of the initial operation strategies is verified based on at least one of the real-time inventory dynamic data and user interaction feedback data. The parameters of the infeasible initial operation strategies are reconstructed to obtain the intermediate operation strategies corresponding to the e-commerce operation sub-scenario.