Cross-border e-commerce big data intelligent processing system based on reinforcement learning

By applying a big data intelligent processing system based on reinforcement learning in the cross-border e-commerce field, the problem of difficulty in quickly extracting valuable information in traditional processing methods is solved, real-time monitoring and strategic adjustment of market dynamic changes are achieved, and the company's market competitiveness and profitability are improved.

CN119991177AInactive Publication Date: 2025-05-13QUANZHOU YUNMOXING TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510261237.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The volume of data transactions in the cross-border e-commerce field has shown explosive growth. It is difficult for traditional processing methods to extract valuable information from massive data quickly and accurately, and it is impossible to adjust marketing strategies in a timely manner to deal with dynamic markets.

Method used

A cross-border e-commerce big data intelligent processing system based on reinforcement learning is adopted, including real-time monitoring module, strategy formulation module and decision-making adjustment module. The real-time monitoring module collects data through network crawling technology, formulates a strategy module to build an agent-environment interaction model using singular value decomposition algorithm and deep Q network algorithm, and the decision adjustment module optimizes the strategy through a policy gradient algorithm.

Benefits of technology

Real-time monitoring and strategy adjustment of dynamic changes in the cross-border e-commerce market has been achieved, the accuracy of pricing, inventory management and advertising has been improved, and the profitability and market competitiveness of the company has been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991177A_ABST
    Figure CN119991177A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cross-border e-commerce, in particular to a cross-border e-commerce big data intelligent processing system based on reinforcement learning. The system comprises a real-time monitoring module, a strategy making module and a decision adjustment module. According to the invention, the real-time monitoring module collects data in increasing changes in the field of cross-border e-commerce, monitors and records the dynamic changes of the market in real time, utilizes the strategy making module to extract the features of the recorded data, and constructs the intelligent agent-environment interaction model. The strategy is evaluated and optimized according to the influence of the optimal strategy selected by the strategy making module on rewards and the long-term stability of the strategy by utilizing a decision adjustment module, the adjustment strategy is constrained in the aspects of actual resource conditions, market and competition and laws and regulations, and resources are reasonably distributed and utilized; and the maximum benefits can be brought into play by limited resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cross-border e-commerce, and in particular to a cross-border e-commerce big data intelligent processing system based on reinforcement learning. Background Art

[0002] Reinforcement learning is the process of interaction between an agent and the environment. The agent learns the optimal behavior strategy by taking a series of actions in the environment with the goal of maximizing the cumulative reward. In this process, the environment will feedback a reward signal and a new state to the agent based on the actions taken by the agent. The agent adjusts its behavior strategy based on this information, thereby continuously optimizing its performance in the environment.

[0003] At present, in the field of cross-border e-commerce, the volume of data transactions has exploded, covering multi-dimensional data such as market trends, consumer preferences, supplier information, logistics conditions, etc. Traditional processing methods are difficult to extract valuable information from massive data quickly and accurately to cope with the dynamic international market. In order to be able to monitor the real-time data and dynamics of the cross-border e-commerce market, timely adjust various cross-border e-commerce decisions, improve the accuracy of pricing, inventory management and advertising, enhance the profitability and market competitiveness of enterprises, and help enterprises seize the initiative in the rapidly changing cross-border e-commerce market, we propose a cross-border e-commerce big data intelligent processing system based on reinforcement learning. Summary of the invention

[0004] The purpose of the present invention is to solve the problem that the data transaction volume in the cross-border e-commerce field is growing explosively. Traditional processing methods are difficult to extract valuable information from massive data quickly and accurately, and cannot make timely adjustments to the marketing strategies of cross-border e-commerce according to changes in data. In order to be able to monitor the real-time data and dynamics of the cross-border e-commerce market, timely adjust various decisions of cross-border e-commerce, and help enterprises seize the initiative in the rapidly changing cross-border e-commerce market.

[0005] To achieve the above-mentioned purpose, the present invention provides a cross-border e-commerce big data intelligent processing system based on reinforcement learning, including a real-time monitoring module, a strategy formulation module and a decision adjustment module;

[0006] The real-time monitoring module collects data from major global e-commerce platforms, social media, logistics websites, and customs data platforms through web crawler technology, and monitors and records the prices, sales, inventory, and market dynamics of goods in the cross-border e-commerce market in real time, and transmits the recorded data to the strategy formulation module;

[0007] The strategy formulation module extracts the key features of price, sales volume and inventory of commodities from the data monitored and recorded by the real-time monitoring module, performs feature dimension reduction on the data using the singular value decomposition algorithm, defines state variables with the information variables of the key features, defines action sets to improve pricing, inventory management and advertising, establishes an agent-environment interaction model, and combines reward functions of profit maximization, market share growth and customer satisfaction improvement into a comprehensive reward function, and determines the optimal strategy selected by the agent in the cross-border e-commerce environment with the value of the comprehensive reward function;

[0008] The decision adjustment module evaluates and optimizes the optimal strategy selected by the strategy formulation module based on its impact on the reward and the long-term stability of the strategy, uses the policy gradient algorithm to analyze the contribution of each strategy to the reward function, adjusts the optimal strategy, and constrains the adjusted strategy based on actual resource conditions, market and competition, and regulations and compliance.

[0009] As a further improvement of the technical solution, the strategy formulation module includes a model building unit and a reward feedback unit;

[0010] The model building unit adopts a deep Q network algorithm as a basic algorithm to build an intelligent agent-environment interaction model, based on a Q learning algorithm, and uses a deep neural network to approximate a Q value function;

[0011] The reward feedback unit determines a comprehensive reward function by weighted summing the reward functions of profit maximization, market share growth and customer satisfaction improvement.

[0012] As a further improvement of the present technical solution, the model building unit adopts an experience replay mechanism to store the experience of the interaction between the intelligent agent and the environment in an experience pool, randomly extracts experience for training, updates the neural network, and breaks the correlation between the data.

[0013] As a further improvement of the technical solution, the model building unit calculates the target Q value using the Bellman equation through the target network technology, reduces the deviation of the Q value estimation, and determines the optimal Q value function.

[0014] As a further improvement of the technical solution, the formula for determining the optimal Q value function by the model building unit is:

[0015] ;

[0016] in, is the current state, For the current action, For the next state, For the next action, For the status Take action The best value, To perform actions Rewards, is the discount factor, For the next state Maximum value.

[0017] As a further improvement of the technical solution, the formula for determining the comprehensive reward function by the reward feedback unit is:

[0018] ;

[0019] in, is the comprehensive reward function, is the reward function for maximizing profit, is the market share growth reward function, Improve the reward function for customer satisfaction, for The weight of for The weight of for The weight of .

[0020] As a further improvement of the technical solution, the decision adjustment module includes an evaluation strategy unit and an adjustment constraint unit;

[0021] The evaluation strategy unit uses a policy gradient algorithm to analyze the contribution of each strategy to the reward function. In multiple evaluation cycles, the parameters and dynamic characteristics of the environment are kept unchanged, and the strategies are run under the same conditions to evaluate the stability of the strategies.

[0022] The adjustment constraint unit utilizes resource demand forecasting, market environment monitoring mechanism and compliance review mechanism to constrain the adjustment strategy in terms of actual resource conditions, market and competition, and regulations and compliance.

[0023] As a further improvement of the present technical solution, the evaluation strategy unit uses a local linear approximation method to perform a sensitivity analysis on the input features of the strategy, and uses the sensitivity value to determine the contribution of each strategy to the reward function.

[0024] As a further improvement of the technical solution, the adjustment constraint unit arranges the resource demand data in chronological order using a time series analysis method to form a time series, and uses a moving average method to predict future resource demand trends.

[0025] As a further improvement of the present technical solution, the real-time monitoring module stores the data of the real-time monitoring records by combining a distributed file system with a NoSQL database.

[0026] Compared with the prior art, the present invention has the following beneficial effects:

[0027] 1. This cross-border e-commerce big data intelligent processing system based on reinforcement learning collects data that is growing in the cross-border e-commerce field through a real-time monitoring module, and monitors and records the dynamic changes of the market in real time. It uses a strategy-making module to extract features from the recorded data and builds an agent-environment interaction model. The agent takes actions based on the current data status to adjust pricing, inventory management, and advertising. The environment feedbacks rewards and new data status based on the agent's actions. Through continuous iterative training, the agent learns the optimal strategy and adjusts various cross-border e-commerce decisions in a timely manner.

[0028] 2. Use the decision adjustment module to evaluate and optimize the optimal strategy selected by the strategy formulation module based on its impact on rewards and its long-term stability. Use the policy gradient algorithm to analyze the contribution of each strategy to the reward function, adjust the optimal strategy, and constrain the adjustment strategy based on actual resource conditions, market and competition, and regulations and compliance. Rationally allocate and utilize resources to maximize the benefits of limited resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0030] Figure 2 It is a schematic diagram of the overall details of the present invention.

[0031] The meaning of each number in the figure is:

[0032] 100. Real-time monitoring module; 200. Strategy formulation module; 210. Model building unit; 220. Reward feedback unit; 300. Decision adjustment module; 310. Evaluation strategy unit; 320. Constraint adjustment unit. DETAILED DESCRIPTION

[0033] The following will be combined with the accompanying drawings in the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0034] At present, the volume of data transactions in the cross-border e-commerce field is growing explosively. Traditional processing methods make it difficult to quickly and accurately extract valuable information from massive data, and cannot make timely adjustments to cross-border e-commerce marketing strategies based on data changes. In order to be able to monitor real-time data and dynamics of the cross-border e-commerce market, timely adjust various cross-border e-commerce decisions, and help companies seize the initiative in the rapidly changing cross-border e-commerce market.

[0035] Therefore, the present invention proposes to collect data showing growth changes in the field of cross-border e-commerce through a real-time monitoring module, and to monitor and record the dynamic changes of the market in real time, to extract features from the recorded data using a strategy-making module, and to construct an agent-environment interaction model, to evaluate and optimize the strategy based on the impact of the optimal strategy selected by the strategy-making module on the reward and the long-term stability of the strategy using a decision-making adjustment module, to constrain the adjustment strategy based on actual resource conditions, market and competition, and regulations and compliance, to rationally allocate and utilize resources and maximize the benefits of limited resources.

[0036] The details are as follows:

[0037] See also Figure 1 As shown, the present invention provides a cross-border e-commerce big data intelligent processing system based on reinforcement learning, including a real-time monitoring module 100, a strategy formulation module 200 and a decision adjustment module 300;

[0038] The real-time monitoring module 100 collects data from data sources of major global e-commerce platforms, social media, logistics websites, and customs data platforms through web crawler technology, and monitors and records the prices, sales volume, inventory, and market dynamics of commodities in the cross-border e-commerce market in real time, and transmits the recorded data to the strategy formulation module 200;

[0039] The strategy formulation module 200 extracts the key features of the price, sales volume and inventory of the goods from the data monitored and recorded by the real-time monitoring module 100, uses the singular value decomposition algorithm to perform feature dimension reduction, defines the state variables with the information variables of the key features, defines the action set to improve pricing, inventory management and advertising, establishes the intelligent agent-environment interaction model, combines the reward functions of profit maximization, market share growth and customer satisfaction improvement into a comprehensive reward function, and determines the optimal strategy selected by the intelligent agent in the cross-border e-commerce environment with the value of the comprehensive reward function;

[0040] The decision adjustment module 300 evaluates and optimizes the optimal strategy selected by the strategy formulation module 200 based on its impact on the reward and the long-term stability of the strategy, uses the policy gradient algorithm to analyze the contribution of each strategy to the reward function, adjusts the optimal strategy, and constrains the adjusted strategy based on actual resource conditions, market and competition, and regulations and compliance.

[0041] The real-time monitoring module 100 collects data from major global e-commerce platforms, social media, logistics websites, customs data platforms and other data sources through web crawlers, API interfaces, etc., and uses the distributed data collection framework Scrapy cluster to ensure the efficiency and stability of data collection, and performs preliminary cleaning on the collected data to remove duplicate, erroneous and invalid data.

[0042] like Figure 2 As shown, the strategy formulation module 200 includes a model building unit 210 and a reward feedback unit 220;

[0043] The model building unit 210 uses a deep Q network algorithm as a basic algorithm to build an intelligent agent-environment interaction model, based on a Q learning algorithm, and uses a deep neural network to approximate a Q value function;

[0044] The reward feedback unit 220 determines a comprehensive reward function by weighted summing the reward functions of profit maximization, market share growth and customer satisfaction improvement;

[0045] Q-learning is a model-free reinforcement learning algorithm used to find the optimal strategy in a given environment. It is based on the Q-value function, where Q(s,a) represents the expected cumulative reward that can be obtained by taking action a in state s. At each step, the agent selects an action based on the current state, and then updates the Q-value based on the reward feedback from the environment and the next state. The goal is to make the Q-value function converge to the optimal value, thereby finding the optimal strategy.

[0046] When the state space and action space are very large or continuous, traditional Q learning methods have difficulty effectively representing and calculating the Q value function. The Deep Q Network algorithm (DQN algorithm) uses the powerful function approximation ability of deep neural networks, takes the state as the input of the neural network, and outputs the Q value corresponding to each action. By training the neural network, it can accurately estimate the Q value of each action under different states, thereby achieving approximation of the Q value function.

[0047] In order to better improve the utilization efficiency of data and the stability of the algorithm, the model building unit 210 adopts an experience replay mechanism to store the experience of the agent's interaction with the environment in an experience pool, randomly extract experience for training, update the neural network, and break the correlation between data;

[0048] In the process of interacting with the environment, the agent stores the experience (state, action, reward, next state) of each step in the experience replay buffer. During training, a batch of experience data is randomly sampled from the buffer to update the neural network, which can break the correlation between data, reduce fluctuations during training, and improve the convergence speed and stability of the algorithm.

[0049] In order to improve the convergence speed and stability of the algorithm, the model building unit 210 calculates the target Q value by using the Bellman equation through the target network technology, reduces the deviation of the Q value estimation, and determines the optimal Q value function;

[0050] The target network has the same structure as the main network used to estimate the Q value, but its parameters are updated relatively slowly. During the training process, the main network is used to generate the current Q value estimate, while the target network is used to provide the target Q value, that is, the target value used to calculate the loss function. By regularly copying the parameters of the main network to the target network, the target Q value can be made relatively stable, avoiding the main network from excessively chasing the changing target during the update process, thereby improving the convergence of the algorithm.

[0051] In order to better determine the optimal Q value function, the formula for determining the optimal Q value function by the model building unit 210 is:

[0052] ;

[0053] in, is the current state, For the current action, For the next state, For the next action, For the status Take action The best value, To perform actions Rewards, is the discount factor, For the next state Maximum value.

[0054] Since the target network's parameter updates are relatively slow, the target Q value it calculates is relatively stable. During the training of the main network, using a relatively stable target Q value as the learning target can prevent the Q value estimation of the main network from being affected by the Q value fluctuations in the current state, thereby reducing the deviation of the Q value estimation. If the main network itself is used directly to calculate the target Q value, since the parameters of the main network are constantly updated during the training process, the calculated target Q value will also change continuously, which will cause the learning process of the main network to be unstable and prone to large deviations;

[0055] The target network decouples the learning process of Q-value estimation from the target calculation process. The main network focuses on adjusting parameters based on current empirical data to better estimate the Q-value, while the target network provides a relatively fixed target, so that the learning of the main network has a clear and stable direction. This decoupling method helps to improve the convergence and stability of the algorithm and further reduce the deviation of Q-value estimation.

[0056] The process of the deep Q network algorithm is:

[0057] S1. Initialize the parameters of the deep neural network (main network and target network), set the size of the experience replay buffer, initialize the environment, and set hyperparameters such as learning rate and discount factor;

[0058] S2. The agent uses the main network to select an action and execute it based on the current state, observes the reward and next state fed back by the environment, and stores the experience of this step (state, action, reward, next state) in the experience replay buffer;

[0059] S3. When the data in the experience replay buffer reaches a certain amount, a batch of experience data is randomly sampled from the buffer. For each sample, the Q value estimate of the action to be performed in the current state is calculated based on the main network, the maximum Q value of the next state is calculated based on the target network, and the target Q value is calculated in combination with the reward. The error between the estimated value and the target Q value is calculated using a loss function such as mean square error, and the parameters of the main network are updated through the back propagation algorithm to reduce the loss;

[0060] S4, at a certain number of steps or time intervals, copy the parameters of the main network to the target network and update the parameters of the target network;

[0061] S5. Repeat steps S2-S4 until the preset number of training steps is reached or other termination conditions are met.

[0062] In order to better determine the reward function value, the reward feedback unit 220 determines the formula of the comprehensive reward function as follows:

[0063] ;

[0064] in, is the comprehensive reward function, is the reward function for maximizing profit, is the market share growth reward function, Improve the reward function for customer satisfaction, for The weight of for The weight of for The weight of .

[0065] For example, the reward function design based on profit maximization: Assume that the current commodity price is , sales volume is The purchase cost is , the logistics cost is , the advertising cost is , then the reward function can be designed as:

[0066] ;

[0067] When market demand is high, the same price may lead to higher sales; Action Adjusting prices, advertising, etc. will directly change the values ​​of these variables. Environmental feedback will reflect the actual impact of these actions on profits;

[0068] The state includes information about market competition, competitor prices and sales, which will affect the company's sales and total market sales. Actions such as lowering prices and launching new marketing strategies will change the company's sales and thus affect market share. Environmental feedback will reflect the changes in market share caused by these actions;

[0069] Status includes information such as product quality, logistics speed, and after-sales service, which will affect customer satisfaction. Actions such as improving product quality, optimizing logistics solutions, and strengthening after-sales service will improve these indicators and thus increase customer satisfaction. Environmental feedback will be reflected in the form of actual return rates and customer evaluation scores.

[0070] The decision adjustment module 300 includes an evaluation strategy unit 310 and an adjustment constraint unit 320;

[0071] The evaluation strategy unit 310 uses a policy gradient algorithm to analyze the contribution of each strategy to the reward function. In multiple evaluation cycles, the parameters and dynamic characteristics of the environment are kept unchanged, and the strategies are run under the same conditions to evaluate the stability of the strategies.

[0072] The adjustment constraint unit 320 uses resource demand forecasting, market environment monitoring mechanisms, and compliance review mechanisms to constrain the adjustment strategy in terms of actual resource conditions, market and competition, and regulations and compliance;

[0073] The policy gradient indicates how a small change in the policy parameters will affect the reward function. By observing the size and direction of the policy gradient, we can understand whether a policy is changing in the direction of increasing or decreasing the reward in the current state, and the magnitude of the change. If the policy gradient is large, it means that the adjustment of the policy has a more significant impact on the reward function; otherwise, the impact is small.

[0074] Observe the gradient size of each parameter in the policy gradient. The larger the absolute value of the gradient, the greater the impact of the adjustment of the policy by the parameter on the reward function. For example, if the gradient of a parameter is large, it means that a small adjustment to the parameter may cause a large change in the policy, which will have a significant impact on the reward function. By sorting the gradients, you can find the parameters that contribute most to the reward function and then analyze the corresponding policy characteristics.

[0075] In order to better determine the contribution value of the strategy, the evaluation strategy unit 310 uses a local linear approximation method to perform a sensitivity analysis on the input features of the strategy, and uses the sensitivity value to determine the contribution of each strategy to the reward function;

[0076] Perform sensitivity analysis on the input features of the strategy, that is, change the value of a certain feature and observe the changes in the policy gradient and reward function. If changing the value of a certain feature causes a large change in the policy gradient and reward function, it means that the feature contributes more to the decision-making and reward function of the strategy. For example, in an autonomous driving scenario, analyze the impact of features such as vehicle speed and distance to obstacles ahead on the policy gradient and reward;

[0077] For the policy function, if it can be expressed in mathematical expressions, calculate the partial derivative of the output result for each input feature. The magnitude of the partial derivative reflects the local sensitivity of the input feature to the output result near the current point. The larger the absolute value of the partial derivative, the greater the impact of the input feature on the output result in the local range;

[0078] Near a certain input feature value, a linear function is used to approximate the policy function. Through linear approximation, we can get the approximate relationship between a small change in the input feature and the change in the output result, thereby evaluating the sensitivity of the input feature.

[0079] In order to better add constraints when adjusting the strategy, the constraint adjustment unit 320 arranges the resource demand data in chronological order using the time series analysis method to form a time series, and uses the moving average method to predict the demand trend of future resources;

[0080] Collect historical resource demand data and arrange them in chronological order to form a time series. Then use methods such as moving average, exponential smoothing, and ARIMA model to analyze and model the time series to predict future resource demand. For example, based on the raw material procurement volume data for each quarter in the past few years, the company uses the moving average method to calculate the raw material demand trend for the next few quarters;

[0081] It is applicable to situations where resource demands have certain time patterns and trends, such as the demand for raw materials by manufacturing companies and the demand for manpower by service companies, which may show a stable growth or fluctuating trend within a certain period of time.

[0082] In order to better record and store the collected data, the real-time monitoring module 100 uses a distributed file system combined with a NoSQL database to store the real-time monitoring recorded data;

[0083] Distributed file systems are used to store massive amounts of unstructured and semi-structured data to ensure data reliability and scalability; NoSQL databases are used to store structured data to meet the needs of fast query and read and write. At the same time, data warehouse tools (such as Hive) are used to store and manage data in layers to facilitate subsequent data processing and analysis.

[0084] In summary, the working principle of this solution is as follows:

[0085] The cross-border e-commerce big data intelligent processing system based on reinforcement learning collects data with increasing changes in the cross-border e-commerce field through the real-time monitoring module 100, and monitors and records the dynamic changes of the market in real time. The strategy-making module 200 is used to extract features from the recorded data and build an intelligent agent-environment interaction model. The intelligent agent takes actions according to the current data status to adjust pricing, inventory management and advertising. The environment feedbacks rewards and new data status according to the actions of the intelligent agent. Through continuous iterative training, the intelligent agent learns the optimal strategy and adjusts various decisions of the cross-border e-commerce in a timely manner. The decision adjustment module 300 is used to evaluate and optimize the optimal strategy selected by the strategy-making module 200 according to its impact on rewards and the long-term stability of the strategy. The policy gradient algorithm is used to analyze the contribution of each strategy to the reward function, adjust the optimal strategy, and constrain the adjustment strategy based on actual resource conditions, market and competition, and regulations and compliance. Resources are reasonably allocated and utilized to maximize the benefits of limited resources.

[0086] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions are only preferred examples of the present invention and are not intended to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A cross-border e-commerce big data intelligent processing system based on reinforcement learning, characterized by: It includes a real-time monitoring module (100), a strategy formulation module (200) and a decision adjustment module (300); The real-time monitoring module (100) collects data from data sources of major global e-commerce platforms, social media, logistics websites, and customs data platforms through web crawler technology, and monitors and records the prices, sales volume, inventory, and market dynamics of commodities in the cross-border e-commerce market in real time, and transmits the recorded data to the strategy formulation module (200); The strategy formulation module (200) extracts the key features of the price, sales volume and inventory of the commodity from the data monitored and recorded by the real-time monitoring module (100), performs feature dimension reduction on the data using a singular value decomposition algorithm, defines state variables using information variables of the key features, defines action sets to improve pricing, inventory management and advertising, establishes an intelligent agent-environment interaction model, combines reward functions of profit maximization, market share growth and customer satisfaction improvement into a comprehensive reward function, and determines the optimal strategy selected by the intelligent agent in the cross-border e-commerce environment using the value of the comprehensive reward function; The decision adjustment module (300) evaluates and optimizes the optimal strategy selected by the strategy formulation module (200) based on its impact on the reward and the long-term stability of the strategy, uses a policy gradient algorithm to analyze the contribution of each strategy to the reward function, adjusts the optimal strategy, and constrains the adjusted strategy based on actual resource conditions, market and competition, and regulations and compliance.

2. The cross-border e-commerce big data intelligent processing system based on reinforcement learning according to claim 1 is characterized by: The strategy formulation module (200) includes a model building unit (210) and a reward feedback unit (220); The model building unit (210) uses a deep Q network algorithm as a basic algorithm to build an intelligent agent-environment interaction model, based on a Q learning algorithm, and uses a deep neural network to approximate a Q value function; The reward feedback unit (220) determines a comprehensive reward function by weighted summing the reward functions of profit maximization, market share growth and customer satisfaction improvement.

3. The cross-border e-commerce big data intelligent processing system based on reinforcement learning according to claim 2 is characterized by: The model building unit (210) adopts an experience replay mechanism to store the experience of the interaction between the intelligent agent and the environment in an experience pool, randomly extracts experience for training, updates the neural network, and breaks the correlation between data.

4. The cross-border e-commerce big data intelligent processing system based on reinforcement learning according to claim 2 is characterized by: The model building unit (210) calculates the target Q value using the Bellman equation through the target network technology, reduces the deviation of the Q value estimation, and determines the optimal Q value function.

5. The cross-border e-commerce big data intelligent processing system based on reinforcement learning according to claim 4 is characterized by: The formula for determining the optimal Q value function by the model building unit (210) is: ; in, is the current state, For the current action, For the next state, For the next action, For the status Take action The best value, To perform actions Rewards, is the discount factor, For the next state Maximum value.

6. The cross-border e-commerce big data intelligent processing system based on reinforcement learning according to claim 2 is characterized by: The reward feedback unit (220) determines the formula of the comprehensive reward function as follows: ; in, is the comprehensive reward function, is the reward function for maximizing profit, is the market share growth reward function, Improve the reward function for customer satisfaction, for The weight of for The weight of for The weight of .

7. The cross-border e-commerce big data intelligent processing system based on reinforcement learning according to claim 1 is characterized by: The decision adjustment module (300) includes an evaluation strategy unit (310) and an adjustment constraint unit (320); The strategy evaluation unit (310) uses a strategy gradient algorithm to analyze the contribution of each strategy to the reward function, and in multiple evaluation cycles, keeps the parameters and dynamic characteristics of the environment unchanged, allows the strategies to run under the same conditions, and evaluates the stability of the strategies; The adjustment constraint unit (320) utilizes resource demand forecasting, market environment monitoring mechanism and compliance review mechanism to constrain the adjustment strategy in terms of actual resource conditions, market and competition, and regulations and compliance.

8. The cross-border e-commerce big data intelligent processing system based on reinforcement learning according to claim 7 is characterized by: The strategy evaluation unit (310) uses a local linear approximation method to perform a sensitivity analysis on the input features of the strategy, and uses the sensitivity value to determine the contribution of each strategy to the reward function.

9. The cross-border e-commerce big data intelligent processing system based on reinforcement learning according to claim 7 is characterized by: The constraint adjustment unit (320) arranges the resource demand data in chronological order using a time series analysis method to form a time series, and uses a moving average method to predict future resource demand trends.

10. The cross-border e-commerce big data intelligent processing system based on reinforcement learning according to claim 1 is characterized by: The real-time monitoring module (100) uses a distributed file system combined with a NoSQL database to store data recorded in real-time monitoring.

Citation Information

Cited By

  • Cross-border intelligent advertisement putting optimization method and system based on reinforcement learning

    CN121329506A

  • Cross-border intelligent advertisement putting optimization method and system based on reinforcement learning

    CN121329506B