A dynamic marketing strategy optimization system combined with reinforcement learning
By combining multi-source data processing and reinforcement learning, a causal perturbation decomposition and reward feedback mechanism is constructed to optimize marketing strategies. This solves the problems of noise perturbation and unstable strategy training in existing technologies, and achieves accurate analysis and dynamic adaptation of marketing strategies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GAODEZHONGCAI TECH CO LTD
- Filing Date
- 2026-04-24
- Publication Date
- 2026-07-24
AI Technical Summary
Existing reinforcement learning marketing optimization solutions fail to effectively eliminate noise disturbances and lack a reasonable feedback mechanism for the final conversion revenue to preceding marketing actions, affecting the stability of strategy training and the model's adaptability to dynamic scenarios.
By combining multi-source data collection and preprocessing, causal perturbation decomposition, net contribution calculation, purification state representation, cross-cycle behavior graph construction, reward feedback, and reinforcement learning training, a closed-loop iterative update system is constructed to optimize marketing strategies through purification marketing state representation and feedback reward values.
It improves the accuracy of marketing effectiveness analysis, enhances the model's adaptability to dynamic scenarios, improves the stability of strategy training and the rationality of long-term conversion benefits, and realizes the output of personalized optimal marketing strategies.
Smart Images

Figure CN122453431A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of dynamic marketing strategy optimization, and more specifically, to a dynamic marketing strategy optimization system that incorporates reinforcement learning. Background Technology
[0002] With the development of digital marketing, existing technologies typically rely on user behavior records, channel reach records, historical preference data, and conversion result data to identify, classify, and configure strategies for marketing targets. However, in dynamic marketing scenarios, user responses are often simultaneously influenced by marketing actions, external events, user cycle fluctuations, and multi-channel collaborative outreach. This makes it difficult for existing technologies to accurately distinguish the true impact of targeted marketing actions from response changes caused by non-target factors. Especially in multi-cycle, long-link marketing processes, different marketing actions have temporal transmission and path coupling relationships. If effectiveness is evaluated solely based on click-through rates, conversion rates, or stage-specific gains, it can easily lead to marketing attribution bias.
[0003] Existing reinforcement learning-based marketing optimization solutions typically use raw behavioral features directly as state inputs and immediate conversion results as reward signals. This fails to effectively eliminate noise disturbances and lacks a reasonable feedback mechanism for the final conversion benefits to preceding marketing actions, thus affecting the stability of strategy training, convergence efficiency, and the model's adaptability to dynamic scenarios. Therefore, there is an urgent need for a dynamic marketing strategy optimization system that incorporates reinforcement learning to improve the accuracy of marketing effect attribution, the authenticity of state representation, and the ability to update strategies. Summary of the Invention
[0004] The purpose of this invention is to provide a dynamic marketing strategy optimization system that combines reinforcement learning. This system addresses the problems of existing reinforcement learning marketing optimization schemes, which typically use original behavioral features directly as state inputs and instant conversion results directly as reward signals. These schemes fail to effectively eliminate noise disturbances and lack a reasonable feedback mechanism for the final conversion benefits to preceding marketing actions. Consequently, they affect the stability of strategy training, convergence efficiency, and the model's adaptability to dynamic scenarios, and thus cannot meet the usage requirements.
[0005] This invention achieves the above objectives through the following technical solution: a dynamic marketing strategy optimization system combining reinforcement learning, the system comprising: Multi-source data acquisition and preprocessing unit, causal perturbation decomposition unit, net contribution calculation unit, purification state characterization unit, cross-cycle behavior graph construction unit, reward feedback calculation unit, reinforcement learning training and strategy decision-making unit, closed-loop iterative update unit; The multi-source data acquisition and preprocessing unit is used to collect marketing action data, user behavior sequence data, channel reach data, external event data, user historical preference data and conversion result data, and output a unified time series dataset. The causal disturbance decomposition unit is used to decompose changes in user response into components directly contributed by marketing actions, external event disturbances, user cycle fluctuations, and channel coupling effects. The net contribution calculation unit is used to generate a counterfactual reference response trajectory and calculate the net contribution value of the target marketing action; The purification status characterization unit is used to construct a purification marketing status characterization based on net contribution value, channel coupling residual, user status migration trend and intensity of behavior change. The cross-cycle behavior graph construction unit is used to construct a cross-cycle marketing behavior link graph that includes marketing action nodes, user behavior nodes, and conversion result nodes. The reward feedback calculation unit is used to generate feedback reward values based on time correlation strength, behavioral link integrity, channel coordination degree and actual marketing contribution strength. The reinforcement learning training and strategy decision-making unit is used to output target marketing strategies based on the purified marketing state representation and the feedback reward value. The closed-loop iterative update unit is used to update the parameters of the causal perturbation decomposition unit, the cross-cycle behavior graph construction unit, and the reinforcement learning training and policy decision unit based on user behavior feedback data.
[0006] Furthermore, the multi-source data acquisition and preprocessing unit includes: Timing alignment subunit, anomaly correction subunit, standardization subunit, and data partitioning subunit; The time-series alignment subunit is used to perform a unified time-granularity mapping on data from different sources according to a preset marketing observation period; The anomaly correction subunit is used to identify and correct missing values, mutated values, and duplicate values; The standardization subunit is used to normalize the original data based on the sample statistics of each dimension, and to use the threshold instead of the standard deviation in the calculation when the standard deviation is lower than the threshold. The data partitioning subunit is used to divide the preprocessed data into a training set and a validation set.
[0007] Furthermore, the causal perturbation decomposition unit includes: Component modeling subunit, decomposition loss calculation subunit, verification loss monitoring subunit, and parameter update subunit; The component modeling subunit is used to establish the correspondence between the total change in user response and the direct contribution component of marketing actions, the external event disturbance component, the user cycle fluctuation component, and the channel coupling influence component. The decomposition loss calculation subunit is used to calculate the decomposition loss between the model output and the observed response; The validation loss monitoring subunit is used to generate the model convergence state based on the validation set results; The parameter update subunit is used to iterate the model parameters based on the model convergence state.
[0008] Furthermore, the time-series alignment subunit is also used to organize marketing action data, user behavior sequence data, channel reach data, external event data, user historical preference data, and conversion result data into a multi-dimensional dynamic marketing time-series matrix based on a unified dimension index; The multidimensional dynamic marketing time series matrix is output to the causal perturbation decomposition unit to ensure consistency of different data sources on the time axis and feature axis.
[0009] Furthermore, the net contribution calculation unit includes: The system includes a counterfactual trajectory generation subunit, a difference calculation subunit, a contribution correction subunit, and a significance determination subunit. The counterfactual trajectory generation subunit is used to generate a reference response trajectory under the condition that the target marketing action has not been implemented; The difference calculation subunit is used to calculate the difference between the actual response trajectory and the reference response trajectory; The contribution correction subunit is used to correct the difference by combining the external event disturbance component, the user period fluctuation component, and the channel coupling influence component. The significance determination subunit is used to compare the corrected difference with the preset net contribution threshold and output the effectiveness determination result of the target marketing action.
[0010] Furthermore, the parameter update subunit includes: Learning rate control subunit, batch training subunit, optimal parameter saving subunit, and early stopping subunit; The learning rate control subunit is used to control the iteration step size of the decomposition model; The batch training subunit is used to perform forward propagation and backward propagation according to a preset batch size; The optimal parameter storage subunit is used to save the current model parameters when verifying loss reduction; The early stop subunit is used to terminate the current training process if the verification loss does not decrease for a preset number of consecutive rounds.
[0011] Furthermore, the purification status characterization unit includes: Feature extraction subunit, weight determination subunit, fusion encoding subunit, and dimensionality compression subunit; The feature extraction subunit is used to extract net contribution value, channel coupling residual, user state migration trend and intensity of behavior change; The weight determination subunit is used to determine the corresponding weights based on the explanatory power of each feature for the transformation response. The fusion coding subunit is used to encode the weighted multidimensional features into a fusion state vector; The dimensional compression subunit is used to convert the fused state vector into a low-dimensional clean marketing state representation suitable for input to a reinforcement learning model.
[0012] Furthermore, the reward feedback calculation unit includes: The sub-units are: time-related calculation, link integrity calculation, channel collaboration calculation, real contribution mapping, and revenue distribution. The time correlation calculation subunit is used to generate the time correlation strength based on the time interval between the marketing action and the conversion result; The link integrity calculation subunit is used to generate behavioral link integrity based on the ratio between the actual number of nodes in the link and the theoretical number of nodes. The channel collaboration calculation subunit is used to generate the channel collaboration level based on the ratio between the number of times the multiple channels jointly reach each other and the total number of times the channels reach each other. The true contribution mapping subunit is used to generate the true marketing contribution intensity based on the net contribution value; The revenue distribution subunit is used to generate feedback reward values for each marketing action based on the time correlation strength, behavioral link integrity, channel synergy degree, and actual marketing contribution strength.
[0013] Furthermore, the reinforcement learning training and policy decision-making unit adopts a deep deterministic policy gradient model, which includes an Actor network, a Critic network, a target Actor network, a target Critic network, an experience replay storage structure, and a policy evaluation subunit. The Actor network is used to output marketing actions based on the clean marketing status representation; The Critic network is used to output action value based on the marketing action and the clean marketing status representation. The experience replay storage structure is used to cache states, actions, rewards, and state transition samples; The target Actor network and the target Critic network are used to maintain training stability according to the soft update rule; The strategy evaluation subunit is used to determine whether to save the strategy model based on the average cumulative reward during the verification phase.
[0014] Furthermore, the closed-loop iterative update unit includes: The system includes a feedback receiving subunit, an incremental training triggering subunit, a state reconstruction subunit, a reward revaluation subunit, and a policy fine-tuning subunit. The feedback receiving subunit is used to receive a new round of user behavior feedback data generated after the target marketing strategy is implemented; The incremental training triggering subunit is used to trigger the multi-source data acquisition and preprocessing unit and the causal perturbation decomposition unit to perform incremental updates. The state reconstruction subunit is used to reconstruct the clean marketing state representation based on the updated net contribution value; The reward revaluation subunit is used to revalue and transmit the reward value based on the updated cross-cycle marketing behavior link graph. The strategy fine-tuning subunit is used to fine-tune the reinforcement learning strategy model based on the reconstructed clean marketing state representation and the re-evaluated feedback reward value, and output the updated target marketing strategy.
[0015] The beneficial effects of this invention are as follows: 1. By performing causal perturbation decomposition on changes in user response, it is possible to distinguish between the direct contribution of marketing actions and the impact of external event perturbations, user cycle fluctuations, and channel coupling, thereby improving the accuracy of marketing effectiveness analysis.
[0016] 2. By constructing a counterfactual reference response trajectory and calculating the net contribution value, it is possible to isolate response fluctuations caused by non-targeted marketing actions and improve the quantitative accuracy of the true contribution of a single marketing action.
[0017] 3. By constructing a clean marketing state representation, the net contribution value, channel coupling residual, user state migration trend and behavioral change intensity are fused and compressed, which can reduce the interference of noise features on reinforcement learning training and improve the effectiveness of state representation.
[0018] 4. By constructing a cross-cycle marketing behavior chain map and combining the strength of time correlation, the completeness of the behavior chain, the degree of channel synergy, and the intensity of actual marketing contribution for reward feedback, the rationality of long-term conversion revenue attribution can be improved, and the adaptability of strategy training to long-term goals can be enhanced.
[0019] 5. By setting up a closed-loop iterative update mechanism, the new round of user behavior feedback is fed back to the causal perturbation decomposition, state representation, reward calculation and reinforcement learning training process, which can enhance the model's adaptability to dynamic marketing scenario changes and improve the stability and real-time effectiveness of strategy output. Attached Figure Description
[0020] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart illustrating the overall system workflow of the present invention. Figure 2This is a flowchart of the causal perturbation decomposition unit of the present invention; Figure 3 This is a flowchart of the closed-loop iterative update process of the present invention. Detailed Implementation
[0021] The present application will now be described in further detail with reference to the accompanying drawings. It should be noted that the following specific embodiments are only used to further illustrate the present application and should not be construed as limiting the scope of protection of the present application. Those skilled in the art can make some non-essential improvements and adjustments to the present application based on the above application content.
[0022] Example 1: Please see Figure 1-3 This invention provides a technical solution: a dynamic marketing strategy optimization system combining reinforcement learning, the system comprising: Multi-source data acquisition and preprocessing unit, causal perturbation decomposition unit, net contribution calculation unit, purification state characterization unit, cross-cycle behavior graph construction unit, reward feedback calculation unit, reinforcement learning training and strategy decision-making unit, closed-loop iterative update unit; The multi-source data acquisition and preprocessing unit is used to collect multi-source heterogeneous data in dynamic marketing scenarios. The multi-source heterogeneous data includes at least marketing action data, user behavior sequence data, channel reach data, external event data, user historical preference data, and conversion result data. The data includes: Marketing Action Data, which records various actions taken by the company during the marketing process, such as sending promotional text messages, pushing advertisements, and holding offline events; User Behavior Sequence Data, which is an orderly record of a series of user behaviors in the marketing scenario, such as the time sequence and specific content of browsing product pages, adding items to the shopping cart, and clicking on advertisements; Channel Reach Data, which reflects the data related to the company's contact with users through different channels, such as social media, email, and text messages, including the time, method, and frequency of contact; External Event Data, which is data on external factors related to marketing activities but not directly controlled by the company, such as holidays, competitor activities, and changes in industry policies; User Historical Preference Data, which is information on user preferences for different products, services, or marketing methods summarized based on past user behavior and feedback; and Conversion Result Data, which records whether users ultimately achieve the company's expected goals, such as data on conversion behaviors such as purchasing goods and registering as members. The causal disturbance decomposition unit is used to build a causal disturbance decomposition model of marketing response based on multi-source heterogeneous data, and decompose changes in user response into components directly contributed by marketing actions, external event disturbances, user cycle fluctuations, and channel coupling effects. Among them, the Marketing Response Causal Disturbance Decomposition Model is a model used to analyze the causes of changes in user response. This model can decompose changes in user response into the following components: the direct contribution component of marketing actions, which is the part of user response change directly caused by the marketing actions taken by the company; the external event disturbance component, which is the part of user response affected by external events, such as holidays, competitor activities, etc.; the user cycle fluctuation component, which is the part of user response change caused by the periodic changes in user behavior and needs over time; and the channel coupling influence component, which is the part of user response change caused by the interaction and influence between different marketing channels. The net contribution calculation unit is used to generate a counterfactual reference response trajectory under the condition that the target marketing action has not been implemented, based on the marketing response causal perturbation decomposition model, and compare the actual response trajectory with the counterfactual reference response trajectory to obtain the net contribution value of the target marketing action. Among them, the counterfactual reference response trajectory is the expected change trajectory of user response assuming no target marketing action was implemented, used for comparative analysis with the actual response trajectory; the net contribution value is the actual impact of the target marketing action on user response, calculated by the difference between the actual response trajectory and the counterfactual reference response trajectory, reflecting the true effect of the marketing action after excluding interference from other factors. The purification status characterization unit is used to construct a purification marketing status characterization that represents the actual marketing effect based on net contribution value, channel coupling residual, user status migration trend and intensity of behavior change. Among them, channel coupling residual, in channel coupling impact analysis, is the difference between the actual channel coupling effect and the model prediction effect, used to measure the complexity and uncertainty of the channel coupling impact; user state migration trend, is the trend of user state changes at different marketing stages or points in time, such as the conversion process of users from potential customers to interested customers and then to paying customers; behavioral change intensity, is the degree of drastic change in user behavior under the influence of different periods or different marketing actions, reflecting the user's sensitivity to marketing activities; and purified marketing state representation, is a state representation constructed after comprehensively considering multiple factors, which can accurately reflect the real marketing effect and provide more accurate state input for reinforcement learning models; The cross-cycle behavior graph construction unit is used to construct a cross-cycle marketing behavior link graph consisting of marketing action nodes, user behavior nodes, and conversion result nodes based on users' touch records, interaction records, and conversion records in multiple marketing cycles. The marketing cycle is a complete period of time during which a company conducts a marketing campaign, typically including multiple marketing phases and activities. Reach records are relevant records of contact between the company and users across different marketing cycles, such as reach time and channels. Interaction records are records of user interactions with the company's marketing content or marketing personnel, such as comments, likes, and inquiries. Conversion records are records of users achieving the company's expected goals across different marketing cycles, such as purchasing goods or registering as members. The cross-cycle marketing behavior chain graph is a graph composed of marketing action nodes, user behavior nodes, and conversion result nodes, showing the complete path from user contact with marketing actions to generating behavior and ultimately conversion across different marketing cycles. The reward feedback calculation unit is used to backpropagate the conversion revenue to the corresponding preceding marketing action based on the cross-cycle marketing behavior link map, combined with the time correlation strength between marketing actions and final conversion, the completeness of the behavior link, the degree of channel synergy, and the intensity of actual marketing contribution, so as to obtain the feedback reward value of each marketing action. Among these factors, the temporal correlation strength refers to the closeness between the marketing action and the final conversion in time; the shorter the time interval, the higher the correlation strength is likely to be. The completeness of the behavioral path refers to the completeness of the user's behavioral path from initial contact with the marketing action to the final conversion; a complete behavioral path helps to more accurately evaluate the effectiveness of the marketing action. The degree of channel synergy refers to the degree to which different marketing channels cooperate and work together during the user conversion process; good channel synergy can improve marketing effectiveness. The intensity of actual marketing contribution refers to the actual contribution of each marketing action to the final user conversion; this is a comprehensive evaluation considering multiple factors. The backpropagation reward value is the reward value of each marketing action calculated based on the backpropagation of conversion revenue, used in the reinforcement learning model to evaluate the merits of the marketing actions. The reinforcement learning training and strategy decision-making unit is used to take the purified marketing state representation as the state input of the reinforcement learning model, take the feedback reward value as the reward input of the reinforcement learning model, and use the reinforcement learning model to output the target marketing strategy for the corresponding user or user group. Among them, reinforcement learning model is a machine learning model that learns the optimal strategy by interacting with the environment. In marketing strategy optimization, the environment is the marketing scenario, and the agent is the enterprise. By continuously trying different marketing actions and adjusting the strategy according to the feedback reward value, the marketing goal is maximized. State input is the information received by the reinforcement learning model about the current state of the marketing environment, which is the purification marketing state representation. Reward input is the feedback signal obtained by the reinforcement learning model based on the effect of the marketing action, which is the feedback reward value. Target marketing strategy is the optimal combination of marketing actions and execution plan formulated based on the output of the reinforcement learning model for a specific user or user group. The closed-loop iterative update unit is used to execute the target marketing strategy and re-input the new round of user behavior feedback generated after execution into the marketing response causal perturbation decomposition model, the cross-cycle marketing behavior link graph and the reinforcement learning model to complete the dynamic update of strategy parameters and form a closed-loop marketing optimization process. Among these, the new round of user behavior feedback refers to the new behavioral data generated by users after the execution of the target marketing strategy, including reach records, interaction records, conversion records, etc.; the dynamic updating of strategy parameters involves adjusting the parameters in the marketing response causal perturbation decomposition model, the cross-cycle marketing behavior link graph, and the reinforcement learning model based on the new user behavior feedback, so that the model can better adapt to the ever-changing marketing environment and user behavior, and achieve continuous optimization of the marketing strategy; the closed-loop marketing optimization process involves continuously executing marketing strategies, collecting user feedback, and updating model parameters to form a cyclical and continuously improving marketing optimization process, in order to improve marketing effectiveness and return on investment.
[0023] It should be noted that, when using this system, multi-source data collection and preprocessing can comprehensively acquire marketing-related information, providing rich material for subsequent analysis; causal perturbation decomposition can accurately analyze the reasons for changes in user response, clarify the influence of each factor, and provide a scientific basis for strategy formulation; net contribution calculation can eliminate interference and accurately measure the effectiveness of marketing actions; purification state representation can truly reflect the marketing effect, making model input more accurate; cross-cycle behavior mapping can present a complete picture of user behavior and marketing conversion path; reward feedback calculation can reasonably allocate conversion revenue and accurately evaluate the value of each marketing action; reinforcement learning training and strategy decision-making can output personalized and optimal marketing strategies; and closed-loop iterative updates can dynamically adjust parameters based on new feedback, enabling the system to continuously optimize, adapt to the ever-changing marketing environment, improve marketing efficiency and effectiveness, and achieve optimal allocation of marketing resources and maximize return on investment.
[0024] In one embodiment, collecting multi-source heterogeneous data in a dynamic marketing scenario includes: Obtain multi-dimensional dynamic marketing time-series data:
[0025] in, Indicates data dimensions, Indicates the length of the marketing observation period; Standardization is performed on multi-source heterogeneous data to eliminate differences in dimensions and numerical fluctuations between different dimensions, thereby improving the stability and decomposition accuracy of subsequent model training. The standardization formula is as follows:
[0026] in, Indicates the first On the data dimension The raw data at each time step; Indicates the first The mean of all samples in each dimension; Indicates the first The standard deviation of all samples in each dimension; when At that time, directly ordered To avoid numerical anomalies caused by dividing by extremely small values, the threshold value is used. Pick ; The standardized data is divided into training and validation sets in an 8:2 ratio for subsequent model training and performance evaluation.
[0027] This design allows multi-dimensional time-series data to comprehensively record marketing scenario information, providing a rich data foundation for subsequent analysis. Standardization ensures that data from different dimensions are within similar numerical ranges, avoiding model training bias caused by different units of measurement, thus improving training stability and decomposition accuracy. Setting thresholds prevents abnormal division by extremely small values, ensuring computational reliability. Proportionally dividing the training and validation sets allows for model training and performance evaluation, enabling timely identification and optimization of model issues, ensuring the effective application of the model in real-world scenarios, and providing reliable data support for the formulation of precision marketing strategies.
[0028] In one embodiment, the training steps of the marketing response causal perturbation decomposition model include: Initialize the decomposition model parameters and set the learning rate. Batch size Maximum number of iterations ; Input the training set data into the model, perform forward propagation to output each response component, and calculate the decomposition loss:
[0029] The Adam optimizer is used to backpropagate and update the parameters. The validation loss is calculated on the validation set every 10 rounds. Early stopping is triggered when the validation loss does not decrease for 15 consecutive rounds, and the optimal model is saved. The decomposition formula is:
[0030] in, express Constant changes in total user response time This indicates that marketing activities directly contribute to the score. This represents the disturbance component caused by external events. Represents the user's periodic fluctuation components. This indicates the influence of channel coupling. The weights of each component are determined by minimizing the causal inference loss, and the optimal decomposition result with the minimum loss is taken.
[0031] This design involves initializing parameters, inputting them into the training set, calculating the loss through forward propagation, updating the model using the Adam optimizer through backpropagation, setting an early stopping mechanism to save the optimal model, and providing a good starting point for model training by reasonably initializing parameters. The parameters are continuously adjusted through forward and backpropagation, allowing the model to gradually fit the data. Batch processing and setting the maximum number of iterations improve training efficiency. The early stopping mechanism prevents overfitting by stopping training when the validation loss does not decrease, saving the model with the minimum validation loss, ensuring the model's generalization ability, accurately decomposing changes in user response, and providing a basis for subsequent precise quantitative marketing actions.
[0032] In one embodiment, the net contribution value of the target marketing action is obtained. This is calculated by subtracting the difference between the actual response and the counterfactual response, thus eliminating response fluctuations caused by non-marketing actions and accurately quantifying the true contribution of a single marketing action. The calculation formula is as follows:
[0033] in, This indicates the net contribution value of the target marketing activities. express Actual response trajectory value at any given moment express Counterfactual reference response trajectory values at all times; Set net contribution threshold ,when At that time, it was determined that the current marketing efforts had no significant effect, among which Take 1 / 10 of the standard deviation of the historical response fluctuation.
[0034] This design calculates the net contribution value by comparing the actual and counterfactual responses, and sets a threshold to determine the effect. It isolates response fluctuations caused by non-marketing actions, accurately quantifies the true contribution of a single marketing action, and allows marketers to clearly understand the effect of each action. Setting a net contribution threshold can effectively filter out marketing actions with no significant effect, avoiding the waste of resources on ineffective actions. Using 1 / 10 of the standard deviation of historical response fluctuations as the threshold has a certain degree of scientific validity and rationality, making the judgment criteria consistent with the actual situation of the marketing scenario, which helps to optimize the allocation of marketing resources and improve marketing efficiency.
[0035] In one embodiment, a clean marketing status representation is constructed, and a weighted fusion method is used to integrate multi-dimensional effective features into a low-dimensional compact representation, filtering out noise and redundant information. The representation fusion formula is as follows:
[0036] in, express Constantly purify the marketing status indicators. Net contribution value For channel coupling residuals, For user state migration trends, Intensity of behavioral change; Weight Determined by iterative training of information gain rate, satisfying ,and initial value set up; The state representation dimension is uniformly compressed to 64 dimensions for use as input to reinforcement learning models.
[0037] This design employs a weighted fusion approach to construct a purified marketing state representation. By determining weights and compressing dimensions, weighted fusion integrates effective features from multiple dimensions, filtering out noise and redundant information, making the state representation more accurately reflect the actual marketing effect. Weights are determined through iterative training using information gain rate, ensuring that each feature contributes reasonably to the state representation. The initial weight values and their magnitudes take into account the importance of each feature, compressing the state representation dimension to 64 dimensions, reducing data complexity and computational load, while retaining key information to facilitate input to reinforcement learning models, thereby improving model training efficiency and decision accuracy.
[0038] In one embodiment, the feedback reward value for each marketing action is obtained. Based on the behavioral link graph, the final conversion revenue is allocated level by level according to the multi-dimensional correlation strength to achieve accurate attribution of long-term conversion revenue. The backpropagation calculation formula is as follows: in, Indicates the first The return reward value for each marketing action, Indicates the first The revenue value of each conversion result; The time correlation strength is determined by the exponential decay function. calculate, The attenuation coefficient is... The number of days between actions and conversions; The behavior link completeness is determined by the ratio of the actual number of nodes in the link to the theoretical number of complete nodes. The value represents the level of channel synergy, calculated as the number of times the message was jointly delivered across multiple channels divided by the total number of deliveries. The intensity of contribution to real marketing equals the net contribution value. Normalization results; All coefficients are normalized to The interval, with the reward value normalized to... .
[0039] This design, based on the behavioral link graph, allocates conversion revenue according to the strength of multi-dimensional associations and calculates the feedback reward value. It can achieve accurate attribution of long-term conversion revenue. It considers multiple dimensions such as the strength of time association, the completeness of the behavioral link, the degree of channel synergy, and the strength of actual marketing contribution. This makes the feedback reward value more reasonable and accurate in reflecting the value of each marketing action. The calculation of time association strength through an exponential decay function conforms to the law that the impact of time on conversion gradually weakens in reality. The normalization of each coefficient ensures that the reward value is within a reasonable range, which makes it easier for the reinforcement learning model to use reward signals to optimize the strategy and improve the long-term effectiveness and profitability of the marketing strategy.
[0040] In one embodiment, the reinforcement learning model employs a Deep Deterministic Policy Gradient (DDPG) structure, and the specific training steps are as follows: The network structure is set with 64 dimensions for state input and 16 dimensions for action output. The Actor network has 3 fully connected layers with 128→64→32 nodes; the Critic network has 3 fully connected layers with 256→128→64 nodes. Training parameter settings, learning rate , Discount factor Target network update rate Experience replay pool size Batch size ; The training process will purify the marketing status representation. Input Actor network output marketing actions ,Will and Input the Critic network to evaluate the value of actions; return reward values. To monitor the signal, update the Critic network; update the Actor network using policy gradients; and softly update the target network parameters. The objective function for strategy optimization is: in, Let the policy objective function be... For Actor model parameters, For policy networks, for Rewards are transmitted back in real time.
[0041] This design, employing the DDPG structure, sets the network structure and training parameters and clarifies the training process. The DDPG structure is suitable for handling marketing strategy optimization problems in continuous action spaces. A reasonable network structure setting, such as the number of fully connected layer nodes in the Actor and Critic networks, can balance model complexity and performance, effectively learning the mapping relationship between states and actions. Appropriate training parameters, such as learning rate and discount factor, ensure the stability and convergence of model training. The clear training process, through steps such as inputting states and outputting actions, evaluating the value of actions, and updating the network, enables the model to continuously optimize the strategy to maximize the policy objective function, output a better marketing strategy, and improve marketing effectiveness and return on investment.
[0042] In one embodiment, the reinforcement learning model sets the model evaluation and saving rules: The average cumulative reward is calculated on the validation set after every 50 training episodes. ,when Improvement Save the current optimal model; Training stops when there is no performance improvement after 100 consecutive episodes, and the optimal target marketing strategy is output.
[0043] This design calculates the average cumulative reward after a certain number of training episodes. Based on the reward improvement, the model is saved or training is stopped. By periodically calculating the average cumulative reward, the model's performance on the validation set can be understood in a timely manner. When the reward improvement exceeds a certain percentage, the model is saved to ensure that the saved model version has better performance for subsequent applications. Training is stopped when there is no performance improvement for a certain number of consecutive episodes to avoid ineffective training and wasting resources. At the same time, the optimal target marketing strategy is output, making the model training process efficient and targeted. It can provide reliable strategic support for marketing decisions in a timely manner, improving the responsiveness and effectiveness of marketing activities.
[0044] In one embodiment, the specific process of closed-loop iterative update is as follows: The newly collected user behavior feedback data is fed back to the multi-source data collection and preprocessing module to complete the standardization; Input the causal perturbation decomposition model to complete incremental training and update the decomposition parameters; Recalculate net contribution value, clean up marketing status indicators, and return reward value; The reinforcement learning model is fine-tuned for 5–10 rounds to update the policy parameters. The closed-loop iteration cycle is set to 1-7 days, and dynamically adjusted according to the business scenario.
[0045] This design, by processing new data feedback, sequentially updates each model and parameter, setting a closed-loop iteration cycle. The feedback processing of newly collected user behavior data enables the system to obtain the latest marketing information in a timely manner, ensuring data timeliness. The causal perturbation decomposition model is updated sequentially, correlation values are calculated, and the reinforcement learning model is fine-tuned, forming a closed-loop iteration. This allows each model to cooperate and continuously optimize. The closed-loop iteration cycle is dynamically adjusted according to business scenarios, ensuring the system adapts to market changes while avoiding excessively frequent updates that waste resources. This ensures the system maintains a strong ability to optimize marketing strategies, continuously improving marketing effectiveness and competitiveness.
[0046] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0047] The above embodiments provide a detailed description of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A dynamic marketing strategy optimization system incorporating reinforcement learning, characterized in that, The system includes: Multi-source data acquisition and preprocessing unit, causal perturbation decomposition unit, net contribution calculation unit, purification state characterization unit, cross-cycle behavior graph construction unit, reward feedback calculation unit, reinforcement learning training and strategy decision-making unit, closed-loop iterative update unit; The multi-source data acquisition and preprocessing unit is used to collect marketing action data, user behavior sequence data, channel reach data, external event data, user historical preference data and conversion result data, and output a unified time series dataset. The causal disturbance decomposition unit is used to decompose changes in user response into components directly contributed by marketing actions, external event disturbances, user cycle fluctuations, and channel coupling effects. The net contribution calculation unit is used to generate a counterfactual reference response trajectory and calculate the net contribution value of the target marketing action; The purification status characterization unit is used to construct a purification marketing status characterization based on net contribution value, channel coupling residual, user status migration trend and intensity of behavior change. The cross-cycle behavior graph construction unit is used to construct a cross-cycle marketing behavior link graph that includes marketing action nodes, user behavior nodes, and conversion result nodes. The reward feedback calculation unit is used to generate feedback reward values based on time correlation strength, behavioral link integrity, channel coordination degree and actual marketing contribution strength. The reinforcement learning training and strategy decision-making unit is used to output target marketing strategies based on the purified marketing state representation and the feedback reward value. The closed-loop iterative update unit is used to update the parameters of the causal perturbation decomposition unit, the cross-cycle behavior graph construction unit, and the reinforcement learning training and policy decision unit based on user behavior feedback data.
2. The dynamic marketing strategy optimization system combining reinforcement learning according to claim 1, characterized in that, The multi-source data acquisition and preprocessing unit includes: Timing alignment subunit, anomaly correction subunit, standardization subunit, and data partitioning subunit; The time-series alignment subunit is used to perform a unified time-granularity mapping on data from different sources according to a preset marketing observation period; The anomaly correction subunit is used to identify and correct missing values, mutated values, and duplicate values; The standardization subunit is used to normalize the original data based on the sample statistics of each dimension, and to use the threshold instead of the standard deviation in the calculation when the standard deviation is lower than the threshold. The data partitioning subunit is used to divide the preprocessed data into a training set and a validation set.
3. The dynamic marketing strategy optimization system combining reinforcement learning according to claim 1, characterized in that, The causal perturbation decomposition unit includes: Component modeling subunit, decomposition loss calculation subunit, verification loss monitoring subunit, and parameter update subunit; The component modeling subunit is used to establish the correspondence between the total change in user response and the direct contribution component of marketing actions, the external event disturbance component, the user cycle fluctuation component, and the channel coupling influence component. The decomposition loss calculation subunit is used to calculate the decomposition loss between the model output and the observed response; The validation loss monitoring subunit is used to generate the model convergence state based on the validation set results; The parameter update subunit is used to iterate the model parameters based on the model convergence state.
4. The dynamic marketing strategy optimization system combining reinforcement learning according to claim 2, characterized in that: The time-series alignment subunit is also used to organize marketing action data, user behavior sequence data, channel reach data, external event data, user historical preference data, and conversion result data into a multi-dimensional dynamic marketing time-series matrix based on a unified dimension index; The multidimensional dynamic marketing time series matrix is output to the causal perturbation decomposition unit to ensure consistency of different data sources on the time axis and feature axis.
5. The dynamic marketing strategy optimization system combining reinforcement learning according to claim 1, characterized in that, The net contribution calculation unit includes: The system includes a counterfactual trajectory generation subunit, a difference calculation subunit, a contribution correction subunit, and a significance determination subunit. The counterfactual trajectory generation subunit is used to generate a reference response trajectory under the condition that the target marketing action has not been implemented; The difference calculation subunit is used to calculate the difference between the actual response trajectory and the reference response trajectory; The contribution correction subunit is used to correct the difference by combining the external event disturbance component, the user period fluctuation component, and the channel coupling influence component. The significance determination subunit is used to compare the corrected difference with the preset net contribution threshold and output the effectiveness determination result of the target marketing action.
6. The dynamic marketing strategy optimization system combining reinforcement learning according to claim 3, characterized in that, The parameter update subunit includes: Learning rate control subunit, batch training subunit, optimal parameter saving subunit, and early stopping subunit; The learning rate control subunit is used to control the iteration step size of the decomposition model; The batch training subunit is used to perform forward propagation and backward propagation according to a preset batch size; The optimal parameter storage subunit is used to save the current model parameters when verifying loss reduction; The early stop subunit is used to terminate the current training process if the verification loss does not decrease for a preset number of consecutive rounds.
7. The dynamic marketing strategy optimization system combining reinforcement learning according to claim 1, characterized in that, The purification status characterization unit includes: Feature extraction subunit, weight determination subunit, fusion encoding subunit, and dimensionality compression subunit; The feature extraction subunit is used to extract net contribution value, channel coupling residual, user state migration trend and intensity of behavior change; The weight determination subunit is used to determine the corresponding weights based on the explanatory power of each feature for the transformation response. The fusion coding subunit is used to encode the weighted multidimensional features into a fusion state vector; The dimensional compression subunit is used to convert the fused state vector into a low-dimensional clean marketing state representation suitable for input to a reinforcement learning model.
8. The dynamic marketing strategy optimization system combining reinforcement learning according to claim 5, characterized in that, The reward feedback calculation unit includes: The sub-units are: time-related calculation, link integrity calculation, channel collaboration calculation, real contribution mapping, and revenue distribution. The time correlation calculation subunit is used to generate the time correlation strength based on the time interval between the marketing action and the conversion result; The link integrity calculation subunit is used to generate behavioral link integrity based on the ratio between the actual number of nodes in the link and the theoretical number of nodes. The channel collaboration calculation subunit is used to generate the channel collaboration level based on the ratio between the number of times the multiple channels jointly reach each other and the total number of times the channels reach each other. The true contribution mapping subunit is used to generate the true marketing contribution intensity based on the net contribution value; The revenue distribution subunit is used to generate feedback reward values for each marketing action based on the time correlation strength, behavioral link integrity, channel synergy degree, and actual marketing contribution strength.
9. The dynamic marketing strategy optimization system combining reinforcement learning according to claim 7, characterized in that: The reinforcement learning training and policy decision-making unit adopts a deep deterministic policy gradient model, which includes an Actor network, a Critic network, a target Actor network, a target Critic network, an experience replay storage structure, and a policy evaluation subunit. The Actor network is used to output marketing actions based on the clean marketing status representation; The Critic network is used to output action value based on the marketing action and the clean marketing status representation. The experience replay storage structure is used to cache states, actions, rewards, and state transition samples; The target Actor network and the target Critic network are used to maintain training stability according to the soft update rule; The strategy evaluation subunit is used to determine whether to save the strategy model based on the average cumulative reward during the verification phase.
10. The dynamic marketing strategy optimization system combining reinforcement learning according to claim 9, characterized in that, The closed-loop iterative update unit includes: The system includes a feedback receiving subunit, an incremental training triggering subunit, a state reconstruction subunit, a reward revaluation subunit, and a policy fine-tuning subunit. The feedback receiving subunit is used to receive a new round of user behavior feedback data generated after the target marketing strategy is implemented; The incremental training triggering subunit is used to trigger the multi-source data acquisition and preprocessing unit and the causal perturbation decomposition unit to perform incremental updates. The state reconstruction subunit is used to reconstruct the clean marketing state representation based on the updated net contribution value; The reward revaluation subunit is used to revalue and transmit the reward value based on the updated cross-cycle marketing behavior link graph. The strategy fine-tuning subunit is used to fine-tune the reinforcement learning strategy model based on the reconstructed clean marketing state representation and the re-evaluated feedback reward value, and output the updated target marketing strategy.