Advertisement optimization method and device, electronic equipment and storage medium
By acquiring user behavior data, extracting local time-series and long-term dependency features, constructing user profiles, and optimizing advertising optimization models, the problem of traditional advertising placement being unable to be optimized in real time is solved, resulting in more efficient advertising placement and better return on investment.
Patent Information
- Application Number
- CN202511372181.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2026-02-17
AI Technical Summary
Traditional advertising relies on manual operation and fixed rules, making it difficult to optimize in real time based on dynamically changing user behavior, resulting in low efficiency and wasted advertising resources.
By acquiring user behavior data, extracting local temporal features and long-term dependency features, constructing user profile features, and using advertising optimization models to generate delivery strategies, the model is optimized by combining reinforcement learning and supervised learning, forming a closed-loop adaptive optimization process.
It improves the automation level of ad delivery, reduces human intervention, enhances the matching degree between ads and users, and improves the overall effectiveness and return on investment of ad delivery.
Smart Images

Figure CN121544324A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of advertisement optimization, and in particular to an advertisement optimization method and device, an electronic device, and a storage medium. BACKGROUND
[0002] Currently, marketing and advertisement delivery face many challenges. Traditional advertisement delivery relies on manual operation and fixed rules, which is difficult to optimize in real time according to dynamic changes in user behavior.
[0003] However, this approach is not only inefficient, but also prone to waste of advertising resources and unstable marketing effects. With the popularity of the Internet and mobile Internet, user behavior data, interests, purchase intentions, and other information are increasingly complex, and traditional advertisement optimization methods have been unable to adapt to this change. SUMMARY
[0004] The present application provides an advertisement optimization method, device, electronic device, and storage medium to solve the defects of traditional advertisement delivery relying on manual operation and fixed rules, which is difficult to optimize in real time according to dynamic changes in user behavior.
[0005] The present application provides an advertisement optimization method, comprising the following steps: Obtain user behavior data, extract local time sequence features and long-term dependence features of the user behavior data respectively, and determine user portrait features based on the local time sequence features and the long-term dependence features; Input the user portrait features into an advertisement optimization model to obtain an advertisement delivery strategy output by the advertisement optimization model; Execute the advertisement delivery strategy to obtain an execution effect feedback of the advertisement delivery strategy; In response to the execution effect feedback, optimize the advertisement optimization model, and input the user portrait features into the optimized advertisement optimization model to obtain an updated advertisement delivery strategy output by the optimized advertisement optimization model.
[0006] According to the advertisement optimization method provided by the present application, the training step of the advertisement optimization model comprises: Obtain an initial model to be trained, and determine a current state vector based on user portrait data; the current state vector is used to represent the current advertisement delivery environment; Determine an advertisement delivery action corresponding to the current state vector based on the initial model, and execute the advertisement delivery action to generate a next state vector; determine a cumulative total reward corresponding to the advertising action based on the execution effect corresponding to the benchmark advertising delivery strategy, the execution effect corresponding to the current state vector, and the execution effect corresponding to the next state vector, and train the initial model according to the cumulative total reward to obtain the advertising optimization model; The execution effect corresponding to the benchmark advertising delivery strategy is determined based on a benchmark state vector. The benchmark state vector is used to represent an initial advertising delivery environment without applying any automatic optimization strategy.
[0007] According to the advertising optimization method provided by the application, the cumulative total reward corresponding to the advertising action is determined based on the execution effect corresponding to the benchmark advertising delivery strategy, the execution effect corresponding to the current state vector, and the execution effect corresponding to the next state vector, and the initial model is trained according to the cumulative total reward to obtain the advertising optimization model. The immediate reward corresponding to the advertising action at each decision time step is determined based on the execution effect corresponding to the benchmark advertising delivery strategy, the execution effect corresponding to the current state vector, and the execution effect corresponding to the next state vector. The cumulative total reward corresponding to the advertising action is determined based on the immediate reward corresponding to the advertising action at each decision time step.
[0008] According to the advertising optimization method provided by the application, the cumulative total reward corresponding to the advertising action is determined based on the immediate reward corresponding to the advertising action at each decision time step, and the initial model is trained according to the cumulative total reward to obtain the advertising optimization model. The cumulative total reward corresponding to the advertising action is determined based on the immediate reward corresponding to the advertising action at the current decision time step, the immediate reward corresponding to a preset number of time steps after the current decision time step, and the discount factor corresponding to the preset number of time steps. The discount factor is used to measure the relative importance of the immediate reward corresponding to the preset number of time steps.
[0009] According to the advertising optimization method provided by the application, the advertising action corresponding to the current state vector is determined based on the initial model, and the initial model comprises: Obtain a candidate advertising vector set representing a plurality of candidate advertisements. For each candidate advertisement in the plurality of candidate advertisements, the initial model performs the following steps to generate a recommended ranking score thereof: Based on the current state vector and the advertising vector of the candidate advertisement, a cross-feature matrix is constructed; the cross-feature matrix is used to represent the fine-grained interaction information between the current state vector and the advertising vector. input the cross feature matrix into a convolutional neural network layer to extract and enhance key interaction modes in the cross feature matrix, to obtain an enhanced cross vector; input the enhanced cross vector into a multilayer perceptron to predict and output a recommended ranking score value of the candidate advertisement; sort the plurality of candidate advertisements according to the recommended ranking score values of the plurality of candidate advertisements, and select a target advertisement with the highest recommended ranking score value; take an advertisement launching action corresponding to the target advertisement as an advertisement launching action corresponding to the current state vector.
[0010] According to the advertisement optimization method provided by the application, the local time sequence feature and the long-term dependence feature of the user behavior data are extracted respectively, and the user portrait feature is determined based on the local time sequence feature and the long-term dependence feature, which comprises: Based on the user portrait model, the local time sequence feature and the long-term dependence feature of the user behavior data are extracted respectively, and the user portrait feature is determined based on the local time sequence feature and the long-term dependence feature; The training step of the user portrait model comprises: Obtain sample behavior data; input the sample behavior data into an initial user portrait model to obtain a predicted user portrait feature output by the initial user portrait model; Based on the predicted user portrait feature, a predicted advertisement launching strategy is determined, and a predicted execution effect feedback of the predicted advertisement launching strategy is determined; According to the difference between the predicted execution effect feedback and the label execution effect feedback of the sample behavior data, a target loss is determined, and the initial user portrait model is trained based on the target loss to obtain the user portrait model.
[0011] According to the advertisement optimization method provided by the application, the advertisement launching strategy comprises advertisement launching content, launching time and launching platform.
[0012] The application also provides an advertisement optimization device, comprising the following units: An acquisition unit is configured to acquire user behavior data, extract local time sequence features and long-term dependence features of the user behavior data respectively, and determine user portrait features based on the local time sequence features and the long-term dependence features; An input unit is configured to input the user portrait features into an advertisement optimization model to obtain an advertisement launching strategy output by the advertisement optimization model; An execution unit is configured to execute the advertisement launching strategy to obtain an execution effect feedback of the advertisement launching strategy; An adjustment unit is used to respond to the execution effect feedback, optimize the advertising optimization model, and input the user profile features into the optimized advertising optimization model to obtain the updated advertising delivery strategy output by the optimized advertising optimization model.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the advertising optimization method as described above.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the advertising optimization method as described above.
[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the advertising optimization method as described above.
[0016] The present invention provides an advertising optimization method, apparatus, electronic device, and storage medium that acquires user behavior data, extracts local temporal features and long-term dependency features from the user behavior data, and determines user profile features based on these features. The user profile features are then input into an advertising optimization model to obtain an advertising delivery strategy output by the model. The advertising delivery strategy is executed, and feedback on its execution effect is obtained. In response to the feedback, the advertising optimization model is optimized, and the user profile features are input into the optimized model to obtain an updated advertising delivery strategy output by the optimized model. This method not only improves the automation level of advertising delivery and reduces manual intervention, but also continuously improves the matching degree between advertisements and users based on real-time execution effect feedback, thereby effectively improving the overall effect and return on investment of advertising delivery. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 This is one of the flowcharts illustrating the advertising optimization method provided by the present invention.
[0019] Figure 2 This is the second flowchart of the advertising optimization method provided by the present invention.
[0020] Figure 3This is a schematic diagram of the advertising optimization device provided by the present invention.
[0021] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0023] Figure 1 This is one of the flowcharts illustrating the advertising optimization method provided by the present invention, such as... Figure 1 As shown, the method includes steps 110, 120, 130 and 140.
[0024] Step 110: Obtain user behavior data, extract local temporal features and long-term dependency features from the user behavior data, and determine user profile features based on the local temporal features and the long-term dependency features; Step 120: Input the user profile features into the advertising optimization model to obtain the advertising delivery strategy output by the advertising optimization model; Step 130: Execute the advertising delivery strategy and obtain feedback on the execution effect of the advertising delivery strategy; Step 140: In response to the execution effect feedback, optimize the advertising optimization model and input the user profile features into the optimized advertising optimization model to obtain the updated advertising delivery strategy output by the optimized advertising optimization model.
[0025] Specifically, firstly, user behavior data can be acquired. User behavior data refers to various types of data that reflect a user's interactions with a product, service, or platform over a period of time. User behavior data may include a user's web browsing history, product click records, search queries, application usage records, social media interactions, video viewing history, purchase and add-to-cart behavior, and geolocation information. Social media interactions include, for example, likes, comments, and shares; however, this embodiment of the invention does not specifically limit these interactions.
[0026] User behavior data can be collected in real time or in batches from first-party data sources (such as the company's own website or APP) or authorized third-party data platforms through various technologies such as integrated Application Programming Interfaces (APIs), Software Development Kits (SDKs), and server log monitoring. The collected raw user behavior data usually needs to undergo preprocessing such as data cleaning, deduplication, format standardization, and missing value imputation to ensure data quality.
[0027] Then, local temporal features and long-term dependency features of user behavior data can be extracted separately, and user profile features can be determined based on these features. Local temporal features refer to those that characterize a user's behavioral patterns and immediate intentions within a short time window. These features are crucial for capturing users' sudden, fleeting interests. For example, if a user continuously searches and browses "hiking shoes" and "trekking poles" within one hour, it indicates a strong current intention for outdoor activities.
[0028] Here, various methods can be used to extract local temporal features, such as sliding window method, convolutional neural network (CNN) and attention mechanism, etc., and the embodiments of the present invention do not specifically limit them.
[0029] Here, the sliding window method refers to statistically analyzing or encoding user behavior sequences within a fixed-size time window. Convolutional neural networks utilize the characteristics of their local receptive fields to perform convolution operations on user behavior sequences corresponding to user behavior data in order to capture short-term behavioral patterns.
[0030] Here, long-term dependency features refer to features that characterize stable preferences, consumption habits, and interest evolution trends formed by users over a relatively long period. These features help construct a basic interest profile of the user. For example, a user who has purchased maternity and baby products multiple times in the past six months indicates their status as a "mother" and a long-term, stable demand for maternity and baby products. Extracting long-term dependency features typically requires a model capable of handling long-sequence data, such as Recurrent Neural Networks (RNNs) and their variants, Transformer models, etc. This embodiment of the invention does not specifically limit this approach.
[0031] Here, recurrent neural networks and their variants, such as Long Short-Term Memory (LSTM) networks or Gated Recurrent Units (GRUs), are designed to capture long-term dependencies in sequential data.
[0032] Here, user profile features refer to the comprehensive feature vector generated by effectively fusing extracted local temporal features and long-term dependency features, which can comprehensively and dynamically describe the current user state.
[0033] Here, the user profile features can be determined by concatenating the local time-series features and long-term dependency features, or by weighting the local time-series features and long-term dependency features before concatenation, etc. The embodiments of the present invention do not specifically limit this.
[0034] Then, user profile features can be input into the advertising optimization model to obtain the advertising delivery strategy output by the model. The advertising optimization model is a core decision-making model whose function is to output an optimal advertising delivery decision based on the input user profile features. This model can be built based on various machine learning paradigms, such as supervised learning models (e.g., logistic regression, gradient boosting trees), deep learning models, or reinforcement learning models, etc., and this embodiment of the invention does not specifically limit this. In a preferred embodiment, the advertising optimization model can be a Deep Reinforcement Learning (DRL) model, which models the advertising delivery process as a Markov Decision Process (MDP), learning a delivery strategy that maximizes long-term cumulative rewards through interaction with the environment (i.e., users) and trial and error.
[0035] Here, the advertising placement strategy refers to a set of decision parameters used to guide specific advertising placement behaviors. This strategy may include specific advertising content, target audience profile, placement duration, time period (e.g., weekday evenings, weekend afternoons), platform or channel (e.g., a social media app, a news website), advertising bid, daily or total budget allocation, etc., though this embodiment of the invention does not impose specific limitations on these aspects. The advertising placement strategy output by the advertising optimization model can be a specific advertising ID or structured data containing multiple decision parameters.
[0036] After obtaining the advertising strategy, it can be executed to obtain feedback on the strategy's effectiveness. The system will execute advertising according to the strategy output by the advertising optimization model through the interface with the advertising platform. For example, it can display specified advertising content to target users matching user profile characteristics at specified times and on specified platforms.
[0037] Performance feedback refers to the quantitative evaluation and data collection of advertising effectiveness after the campaign has been launched. Performance feedback may include click-through rate (CTR), conversion rate (CVR), return on investment (ROI), interaction data, and subsequent behavioral data, etc., but this embodiment of the invention does not specifically limit these metrics.
[0038] Here, click-through rate (CTR) refers to the ratio of the number of times an ad is clicked to the number of times it is displayed. Conversion rate refers to the ratio of the number of users who complete a preset goal to the number of users who click on the ad, where the preset goal can be purchase, registration, download, etc. Return on investment (ROI) is the ratio of the revenue generated by advertising to the cost of advertising. Interaction data includes, for example, the playback duration of video ads and the number of likes and comments on social ads. Subsequent behavior data refers to the time users spend on the landing page after clicking on the ad, bounce rate, etc., but this embodiment of the invention does not specifically limit these.
[0039] Finally, in response to the feedback on execution performance, the advertising optimization model is optimized, and user profile features are input into the optimized model to obtain the updated advertising delivery strategy output by the optimized model. This is a closed-loop learning and adaptive optimization process. The system uses the collected execution performance feedback as a supervision signal or reward signal to update the parameters of the advertising optimization model.
[0040] Here, optimizing the advertising optimization model in response to performance feedback specifically refers to adjusting the model's internal parameters (such as the weights and biases of the neural network). If the advertising optimization model is a reinforcement learning model, performance feedback (such as CTR, CVR) can be transformed into reward signals to update the model's policy or value network, enabling it to make decisions that yield higher rewards in the future. If the advertising optimization model is a supervised learning model, the feedback can serve as labels for new training samples, allowing for model fine-tuning using algorithms such as gradient descent.
[0041] This process forms a closed loop of "deployment-feedback-learning-re-deployment". By continuously inputting the latest user profile characteristics into the continuously optimized model, the system can generate updated advertising strategies that are more up-to-date and effective, thereby dynamically adapting to changes in the market environment and user interests.
[0042] The method provided in this invention acquires user behavior data, extracts local temporal features and long-term dependency features from the user behavior data, and determines user profile features based on these features. The user profile features are then input into an advertising optimization model to obtain an advertising delivery strategy output by the model. The advertising delivery strategy is executed, and feedback on its execution effect is obtained. In response to the feedback, the advertising optimization model is optimized, and the user profile features are input into the optimized model to obtain an updated advertising delivery strategy output by the optimized model. This method not only improves the automation level of advertising delivery and reduces manual intervention, but also continuously improves the matching degree between advertisements and users based on real-time execution effect feedback, thereby effectively improving the overall effect and return on investment of advertising delivery.
[0043] Based on the above embodiments, the training steps of the advertising optimization model include: Step 210: Obtain the initial model to be trained, and determine the current state vector based on user profile data; the current state vector is used to represent the current advertising environment. Step 220: Based on the initial model, determine the advertising action corresponding to the current state vector, and execute the advertising action to generate the next state vector; Step 230: Based on the execution effect of the benchmark advertising strategy, the execution effect of the current state vector, and the execution effect of the next state vector, determine the cumulative total reward corresponding to the advertising action, and train the initial model according to the cumulative total reward to obtain the advertising optimization model; The execution effect of the benchmark advertising strategy is determined based on the benchmark state vector. The baseline state vector is used to characterize the initial advertising environment when no automatic optimization strategy is applied.
[0044] Specifically, firstly, an initial model to be trained can be obtained, and based on user profile data, the current state vector can be determined, where the current state vector is used to represent the current advertising environment.
[0045] The initial model to be trained can be a neural network model with randomly initialized parameters, or it can be a model pre-trained on a relevant task to accelerate convergence. In the context of reinforcement learning, this model is usually referred to as an agent.
[0046] In advertising scenarios, the current state vector is used to comprehensively describe the environment in which the agent makes decisions. The generated user profile features can be used as the core part of the current state vector. In addition, the state vector can also include other contextual information, such as the current time, user device type, network environment, and currently available advertising budget, to provide a richer environmental description. This embodiment of the invention does not specifically limit this aspect.
[0047] Then, based on the initial model, the advertising delivery action corresponding to the current state vector can be determined and executed to generate the next state vector. Here, the advertising delivery action refers to the operation that the agent can perform in the current state. In the advertising delivery scenario, an action can specifically refer to "selecting one advertisement from multiple candidate advertisements for delivery". The initial model (agent) will determine the advertising delivery action based on the current state vector S. t It outputs an action At according to its internal strategy. In the early stages of training, in order to encourage exploration, strategies such as ε-greedy are usually adopted, that is, randomly selecting an action with a certain probability and selecting the action that the model considers optimal with a higher probability.
[0048] After executing the ad delivery action At, the environment will change and generate the next state vector S. t+1 The next state vector reflects changes in user behavior or the progression of time.
[0049] Finally, based on the execution effect of the benchmark advertising strategy, the execution effect of the current state vector, and the execution effect of the next state vector, the cumulative total reward corresponding to the advertising action can be determined. Based on the cumulative total reward, the initial model can be trained to obtain the advertising optimization model.
[0050] Here, the performance of the benchmark ad delivery strategy is determined based on the benchmark state vector; the benchmark state vector is used to characterize the initial ad delivery environment when no automatic optimization strategy is applied.
[0051] The baseline state vector can be understood as a representative or initial advertising environment state before applying the method of this embodiment of the invention. The baseline advertising strategy is a fixed, non-learning, simple strategy. For example, it could be a random placement strategy, a strategy that places ads based on the highest historical average CTR, or a default strategy that does not take any optimization actions. The performance corresponding to the baseline advertising strategy is a relatively fixed reference value, for example, a historical average CTR of 0.5%.
[0052] Understandably, by introducing a benchmark, reward calculation is no longer based on absolute performance, but rather on the improvement of the current policy relative to the benchmark policy. For example, the reward can be defined as the CTR of the current action minus the CTR of the benchmark policy. The advantage of this approach is that it effectively reduces the variance of the reward signal, ensuring that the reward value fluctuates around a meaningful center, thus making the model training process more stable and efficient.
[0053] The cumulative total reward refers to the sum of discounts on all immediate rewards from the current time step until a future time period. This is the optimization objective of reinforcement learning; the model's goal is to learn a policy that maximizes the expected value of this cumulative total reward.
[0054] Based on the cumulative total reward, the initial model is trained using reinforcement learning algorithms such as Proximal Policy Optimization (PPO) or Deep Q-Learning to update its parameters. Specifically, the cumulative total reward is used to evaluate the quality of the actions performed. For PPO, it is used to calculate the advantage function and update the policy network accordingly; for Q-Learning, it is used to construct a temporal difference objective and update the Q-value network accordingly. Through multiple rounds of iterative training in a "state-action-reward-new state" cycle, the initial model gradually learns how to select actions that yield the highest long-term rewards in different states.
[0055] The method provided in this invention significantly improves the stability and convergence speed of the reinforcement learning training process by introducing a benchmark policy to calculate relative rewards. This method allows the initial model to focus on learning a policy better than the benchmark, avoiding training difficulties caused by excessive fluctuations in the absolute value of the reward signal, and ultimately training a more robust and higher-performing advertising optimization model.
[0056] Based on the above embodiments, step 230, which involves determining the cumulative total reward corresponding to the advertising action based on the execution effect corresponding to the baseline advertising strategy, the execution effect corresponding to the current state vector, and the execution effect corresponding to the next state vector, includes: Step 231: Based on the execution effect of the benchmark advertising strategy, the execution effect of the current state vector, and the execution effect of the next state vector, determine the immediate reward corresponding to the advertising action at each decision time step. Step 232: Determine the cumulative total reward corresponding to the advertising placement action based on the immediate reward corresponding to the advertising placement action at each decision time step.
[0057] Specifically, firstly, the immediate reward corresponding to the advertising action can be determined based on the execution effect of the benchmark advertising strategy, the execution effect of the current state vector, and the execution effect of the next state vector.
[0058] Here, the immediate reward is the single-step reward signal returned by the environment after the agent performs action At at time step t. This embodiment defines a more refined method for calculating the immediate reward. In a specific example, the immediate reward... It can consist of the following parts: = (Effect of the current action - Effect of the baseline strategy) + γ * V(S) t+1 ) - V(S t ) The effect of the current action refers to the observed performance result after performing action At in state St, such as CTR and CVR. The effect of the baseline strategy refers to the performance result corresponding to the baseline ad placement strategy, which is a constant or slowly changing value used as a reference.
[0059] V(S t ) is the value function in the current state S t The estimated value under state S represents the value obtained from state S. t Begin by following the expected future cumulative reward achievable with the current strategy; V(S t+1 ) is the value function in the next state S t+1 The estimated value is given below; γ is the discount factor.
[0060] By subtracting the baseline effect from the current action effect, the reward signal directly reflects the "gain" of the current action relative to the baseline, making the reward signal more instructive. Simultaneously, the change in the value function γ * V(S) is introduced. t+1 )-V(S t This helps reduce the variance of rewards.
[0061] Furthermore, based on the immediate rewards corresponding to the ad placement actions at each decision time step, the cumulative total reward corresponding to the ad placement actions is determined. This is done after calculating the immediate reward for each time step. Then, the cumulative total reward can be obtained by summing or weighted summing multiple instantaneous rewards over a time series.
[0062] The method provided in this invention offers a high-quality, low-variance training signal for reinforcement learning models by defining an immediate reward calculation method that includes baseline comparison and value function estimation. This immediate reward directly reflects the marginal contribution of each decision action, enabling the advertising optimization model to learn the merits of each action more accurately, thereby accelerating the formation of effective advertising delivery strategies.
[0063] Based on the above embodiments, step 232 includes: Step 2321: Based on the immediate reward corresponding to the advertising action at the current decision time step, the immediate reward corresponding to a preset number of time steps after the current decision time step, and the discount factor corresponding to the preset number of time steps, determine the cumulative total reward corresponding to the advertising action. The discount factor is used to measure the relative importance of the instantaneous rewards corresponding to the preset number of time steps.
[0064] Specifically, the cumulative total reward for an advertising action can be determined based on the immediate reward corresponding to the advertising action at the current decision time step, the immediate rewards corresponding to a preset number of time steps after the current decision time step, and the discount factor corresponding to the preset number of time steps. The discount factor is used to measure the relative importance of the immediate rewards corresponding to the preset number of time steps.
[0065] In this embodiment of the invention, the definition is in the first... t Total cumulative reward at each time step as follows: in, Indicates the first t Instant rewards earned at each time step Indicates the discount factor. It is the set maximum number of time steps, i.e. This indicates a preset number of time steps. express The state vector at each time step State value, Represents the step from the current time t The total cumulative reward actually obtained from the start to the maximum set time step.
[0066] The discount factor is a hyperparameter between 0 and 1 (e.g., 0.99). Its function is to measure the relative importance of future rewards compared to current rewards. A discount factor close to 1 means the ad optimization model is more forward-looking and will give equal importance to long-term rewards; a discount factor close to 0 means the ad optimization model is more "short-sighted" and will focus more on immediate rewards. By adjusting the discount factor, the short-term gains and long-term goals of the ad optimization model can be balanced.
[0067] The method provided in this invention calculates the cumulative total reward by introducing a discount factor, enabling the advertising optimization model to balance immediate effects and long-term impacts during training. This helps the model avoid adopting "short-sighted" strategies that may bring high short-term returns but could harm long-term user experience or brand value, thereby learning more sustainable advertising strategies that deliver stable long-term returns.
[0068] Based on the above embodiments, step 220 includes: Step 221: Obtain the candidate ad vector set representing multiple candidate ads; Step 222: For each of the plurality of candidate ads, the initial model performs the following steps to generate its recommendation ranking score: Based on the current state vector and the advertisement vector of the candidate advertisement, a cross-feature matrix is constructed; the cross-feature matrix is used to characterize the fine-grained interaction information between the current state vector and the advertisement vector. The cross feature matrix is input into a convolutional neural network layer to extract and enhance key interaction patterns in the cross feature matrix, thereby obtaining an enhanced cross vector; The enhanced cross vector is input into a multilayer perceptron to predict and output the recommendation ranking score of the candidate advertisement; Step 223: Sort the multiple candidate advertisements according to their recommendation ranking scores, and select the target advertisement with the highest recommendation ranking score. Step 224: Take the advertising delivery action corresponding to the target advertisement as the advertising delivery action corresponding to the current state vector.
[0069] Specifically, firstly, a candidate ad vector set representing multiple candidate ads is obtained. The candidate ad vector set refers to the collection of all available ads under the current delivery opportunity, where each ad is represented as an ad vector. Ad vectors can include ad content features, attribute features, and delivery requirement features. Content features can be keywords, image style, etc.; attribute features can be the industry, advertiser ID, etc.; and delivery requirement features can be bidding and target audience characteristics, etc. This embodiment of the invention does not impose specific limitations on these aspects.
[0070] Then, for each of the multiple candidate ads, the initial model performs the following steps to generate its recommendation ranking score: a) Construct a cross-feature matrix based on the current state vector and the candidate ad vectors. The cross-feature matrix is designed to capture all fine-grained interaction information between user features and ad features. Traditionally, the user vector and ad vector are simply concatenated and input into an MLP (Multi-Layer Perceptron), but this results in the loss of significant second-order interaction information. In this embodiment, the cross-feature matrix can be constructed using an outer product operation. Assuming the current state vector has dimension Du and the ad vector has dimension Da, a Du × Da cross-feature matrix is generated, where each element (i, j) represents the interaction value between the i-th user feature and the j-th ad feature.
[0071] (b) The cross-feature matrix is input into a convolutional neural network layer to extract and enhance key interaction patterns within it, resulting in an enhanced cross vector. Since the cross-feature matrix has a two-dimensional structure similar to an image, CNNs can be used to effectively extract local interaction patterns. The CNN's convolutional kernels slide across the cross-feature matrix, automatically detecting and learning which combinations of user features and advertising features are meaningful or strongly correlated, such as the interaction between a user's "sports enthusiast" feature and an advertisement's "new running shoes" feature. After convolution and pooling operations, an enhanced cross vector is output, which is a compact representation of all key interaction patterns.
[0072] c) The enhanced cross vectors are input into a multilayer perceptron to predict and output the recommendation ranking score of the candidate ads. The recommendation ranking score can be understood as the Q-value or probability of the action predicted by the model, which intuitively represents the expected return of showing this ad to the user.
[0073] Then, based on the recommendation ranking scores of multiple candidate ads, the multiple candidate ads are ranked, and the target ad with the highest recommendation ranking score is selected.
[0074] Finally, the ad delivery action corresponding to the target ad is used as the ad delivery action corresponding to the current state vector.
[0075] The method provided in this invention constructs a unique network structure of "cross-feature matrix + convolutional neural network layer + multilayer perceptron," enabling the initial model to automatically learn and extract high-order, non-linear interaction features between users and advertisements from the original features. Compared to simple feature concatenation, this approach can gain a deeper understanding of the fine-grained matching relationship between user preferences and advertisement attributes, thereby making more accurate advertisement recommendations and rankings, and significantly improving the hit rate and effectiveness of advertisement delivery.
[0076] Based on the above embodiments, step 110, which involves extracting the local temporal features and long-term dependency features of the user behavior data, and determining the user profile features based on the local temporal features and the long-term dependency features, includes: Step 111: Based on the user profile model, extract the local temporal features and long-term dependency features of the user behavior data respectively, and determine the user profile features based on the local temporal features and long-term dependency features; The training steps of the user profile model include: Obtain sample behavior data; The sample behavior data is input into the initial user profile model to obtain the predicted user profile features output by the initial user profile model. Based on the predicted user profile features, a predicted advertising delivery strategy is determined, and the predicted execution effect feedback of the predicted advertising delivery strategy is determined. Based on the difference between the predicted execution effect feedback and the labeled execution effect feedback of the sample behavior data, a target loss is determined, and the initial user profile model is trained based on the target loss to obtain the user profile model.
[0077] Specifically, based on the user profile model, local temporal features and long-term dependency features of user behavior data can be extracted respectively, and user profile features can be determined based on the local temporal features and long-term dependency features.
[0078] Here, the training steps for the user profile model include: a) Obtain sample behavior data. Sample behavior data refers to a large amount of labeled historical user behavior data. Each sample typically contains a user's behavior sequence and the labeled performance feedback generated by that user after receiving a particular advertisement.
[0079] b) Input the sample behavior data into the initial user profile model to obtain the predicted user profile features output by the initial user profile model.
[0080] c) Based on the predicted user profile characteristics, determine the predicted advertising delivery strategy and determine the feedback on the predicted execution effect of the predicted advertising delivery strategy.
[0081] d) Based on the difference between the predicted execution effect feedback and the labeled execution effect feedback of the sample behavior data, determine the target loss, and train the initial user profile model based on the target loss to obtain the user profile model.
[0082] Here, the target loss, which is the difference between the predicted value and the true value, can be quantified by mean squared error (MSE) or cross-entropy loss.
[0083] Repeat the above training steps until the initial user profile model converges, and finally obtain the trained user profile model.
[0084] Based on any of the above embodiments Figure 2 This is the second flowchart illustrating the advertising optimization method provided by the present invention, as follows: Figure 2 As shown, the method includes: First, the data acquisition module captures user behavior and market data in real time, including multi-source information such as web browsing, search history, purchasing behavior, and social media interactions, laying the foundation for subsequent analysis. Next, the data preprocessing module cleans, deduplicates, standardizes, and performs feature engineering on the raw data, transforming messy logs into well-structured feature vectors readable by the model, ensuring data consistency and high quality.
[0085] Subsequently, the user behavior analysis module utilizes deep learning models (such as LSTM or Transformer) to perform in-depth analysis of the cleaned time-series data, completing the construction of user profiles and outputting interest prediction and demand analysis results to accurately depict the potential purchase intentions of each user. These profiles and prediction results are input into the advertising optimization module, which uses deep reinforcement learning algorithms (Deep Q-Learning or PPO) as its core to model the advertising process as a Markov decision process, generating advertising strategies in real time. This covers key variables such as ad content, target audience, timing of placement, and budget allocation, striving to maximize click-through rate, conversion rate, and return on investment in a dynamic market.
[0086] Once the strategy is generated, the system immediately executes the campaign and monitors it in real time through the performance evaluation module. It continuously collects key metrics such as click-through rate (CTR) and conversion rate, and simultaneously runs A / B testing to compare different creatives, audience segments, and time slots, generating objective performance evaluation conclusions. The feedback samples generated from the evaluation are sent to the closed-loop learning and adaptive optimization module. Here, reinforcement learning and supervised learning work together to update model parameters online using the latest feedback data, adjusting the strategy to make the next round of campaigns more precise.
[0087] Each iteration of the entire process is accompanied by the automatic output of report generation and integration modules: the system summarizes structured information such as advertising performance and optimization suggestions into standard reports, and seamlessly integrates them with the enterprise's OA, DMS (Document Management System) and other business systems, making it convenient for the marketing team to read and make decisions. Through the interconnected links of "collection—preprocessing—analysis—optimization—evaluation—closed-loop learning—reporting", the platform realizes data-driven adaptive marketing, continuously using fresh data to reset, deconstruct, and reconstruct the campaign logic, ultimately driving a continuous increase in advertising performance.
[0088] The advertising optimization device provided by the present invention is described below. The advertising optimization device described below and the advertising optimization method described above can be referred to in correspondence.
[0089] Based on any of the above embodiments, the present invention provides an advertising optimization device. Figure 3 This is a schematic diagram of the advertising optimization device provided by the present invention, as shown below. Figure 3 As shown, the device includes: The acquisition unit 310 is used to acquire user behavior data, extract local temporal features and long-term dependency features of the user behavior data respectively, and determine user profile features based on the local temporal features and the long-term dependency features. Input unit 320 is used to input the user profile features into the advertising optimization model to obtain the advertising delivery strategy output by the advertising optimization model; The execution unit 330 is used to execute the advertising delivery strategy and obtain feedback on the execution effect of the advertising delivery strategy. The adjustment unit 340 is used to optimize the advertising optimization model in response to the execution effect feedback, and input the user profile features into the optimized advertising optimization model to obtain the updated advertising delivery strategy output by the optimized advertising optimization model.
[0090] The apparatus provided in this invention acquires user behavior data, extracts local temporal features and long-term dependency features from the user behavior data, and determines user profile features based on these features. The user profile features are then input into an advertising optimization model to obtain an advertising delivery strategy output by the model. The advertising delivery strategy is executed, and feedback on its execution effect is obtained. In response to the feedback, the advertising optimization model is optimized, and the user profile features are input into the optimized model to obtain an updated advertising delivery strategy output by the optimized model. This method not only improves the automation level of advertising delivery and reduces manual intervention, but also continuously improves the matching degree between advertisements and users based on real-time execution effect feedback, thereby effectively improving the overall effect and return on investment of advertising delivery.
[0091] Based on any of the above embodiments, a training unit is further included, wherein the training unit specifically includes: The model acquisition unit is used to acquire the initial model to be trained and determine the current state vector based on user profile data; the current state vector is used to represent the current advertising environment. The generation unit is used to determine the advertising delivery action corresponding to the current state vector based on the initial model, and execute the advertising delivery action to generate the next state vector; The training subunit is used to determine the cumulative total reward corresponding to the advertising action based on the execution effect corresponding to the benchmark advertising strategy, the execution effect corresponding to the current state vector, and the execution effect corresponding to the next state vector, and to train the initial model based on the cumulative total reward to obtain the advertising optimization model. The execution effect of the benchmark advertising strategy is determined based on the benchmark state vector. The baseline state vector is used to characterize the initial advertising environment when no automatic optimization strategy is applied.
[0092] Based on any of the above embodiments, the training subunit specifically includes: The instant reward unit is used to determine the instant reward corresponding to the advertising action at each decision time step based on the execution effect corresponding to the baseline advertising strategy, the execution effect corresponding to the current state vector, and the execution effect corresponding to the next state vector. A cumulative total reward unit is determined to determine the cumulative total reward corresponding to the advertising action based on the immediate reward corresponding to the advertising action at each decision time step.
[0093] Based on any of the above embodiments, the step of determining the cumulative total reward unit is specifically used for: Based on the immediate reward corresponding to the advertising action at the current decision time step, the immediate rewards corresponding to a preset number of time steps after the current decision time step, and the discount factors corresponding to the preset number of time steps, the cumulative total reward corresponding to the advertising action is determined. The discount factor is used to measure the relative importance of the instantaneous rewards corresponding to the preset number of time steps.
[0094] Based on any of the above embodiments, the generation unit is specifically used for: Obtain the candidate ad vector set representing multiple candidate ads; For each of the plurality of candidate ads, the initial model performs the following steps to generate its recommendation ranking score: Based on the current state vector and the advertisement vector of the candidate advertisement, a cross-feature matrix is constructed; the cross-feature matrix is used to characterize the fine-grained interaction information between the current state vector and the advertisement vector. The cross feature matrix is input into a convolutional neural network layer to extract and enhance key interaction patterns in the cross feature matrix, thereby obtaining an enhanced cross vector; The enhanced cross vector is input into a multilayer perceptron to predict and output the recommendation ranking score of the candidate advertisement; Based on the recommendation ranking scores of the multiple candidate advertisements, the multiple candidate advertisements are ranked, and the target advertisement with the highest recommendation ranking score is selected. The ad delivery action corresponding to the target ad is taken as the ad delivery action corresponding to the current state vector.
[0095] Based on any of the above embodiments, the acquisition unit 310 is specifically used for: Based on the user profile model, local temporal features and long-term dependency features of the user behavior data are extracted respectively, and the user profile features are determined based on the local temporal features and long-term dependency features. It also includes a user profile model training unit, which is specifically used for: Obtain sample behavior data; The sample behavior data is input into the initial user profile model to obtain the predicted user profile features output by the initial user profile model. Based on the predicted user profile features, a predicted advertising delivery strategy is determined, and the predicted execution effect feedback of the predicted advertising delivery strategy is determined. Based on the difference between the predicted execution effect feedback and the labeled execution effect feedback of the sample behavior data, a target loss is determined, and the initial user profile model is trained based on the target loss to obtain the user profile model.
[0096] Based on any of the above embodiments, the advertising delivery strategy includes advertising content, delivery duration, and delivery platform.
[0097] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 4As shown, the electronic device may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute an advertising optimization method. This method includes: acquiring user behavior data; extracting local temporal features and long-term dependency features from the user behavior data; determining user profile features based on the local temporal features and the long-term dependency features; inputting the user profile features into an advertising optimization model to obtain an advertising delivery strategy output by the advertising optimization model; executing the advertising delivery strategy to obtain feedback on the execution effect of the advertising delivery strategy; and optimizing the advertising optimization model in response to the execution effect feedback, and inputting the user profile features into the optimized advertising optimization model to obtain an updated advertising delivery strategy output by the optimized advertising optimization model.
[0098] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0099] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the advertising optimization method provided by the above methods. The method includes: acquiring user behavior data; extracting local temporal features and long-term dependency features from the user behavior data; determining user profile features based on the local temporal features and the long-term dependency features; inputting the user profile features into an advertising optimization model to obtain an advertising delivery strategy output by the advertising optimization model; executing the advertising delivery strategy to obtain feedback on the execution effect of the advertising delivery strategy; and optimizing the advertising optimization model in response to the execution effect feedback, and inputting the user profile features into the optimized advertising optimization model to obtain an updated advertising delivery strategy output by the optimized advertising optimization model.
[0100] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the advertising optimization method provided by the above methods. The method includes: acquiring user behavior data; extracting local temporal features and long-term dependency features from the user behavior data; determining user profile features based on the local temporal features and the long-term dependency features; inputting the user profile features into an advertising optimization model to obtain an advertising delivery strategy output by the advertising optimization model; executing the advertising delivery strategy to obtain feedback on the execution effect of the advertising delivery strategy; and optimizing the advertising optimization model in response to the execution effect feedback, and inputting the user profile features into the optimized advertising optimization model to obtain an updated advertising delivery strategy output by the optimized advertising optimization model.
[0101] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0102] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An advertising optimization method, characterized in that, include: Acquire user behavior data, extract local temporal features and long-term dependency features from the user behavior data, and determine user profile features based on the local temporal features and the long-term dependency features; The user profile features are input into the advertising optimization model to obtain the advertising delivery strategy output by the advertising optimization model; Execute the advertising delivery strategy and obtain feedback on the execution effect of the advertising delivery strategy; In response to the execution effect feedback, the advertising optimization model is optimized, and the user profile features are input into the optimized advertising optimization model to obtain the updated advertising delivery strategy output by the optimized advertising optimization model.
2. The advertising optimization method according to claim 1, characterized in that, The training steps of the advertising optimization model include: Obtain the initial model to be trained, and determine the current state vector based on user profile data; the current state vector is used to represent the current advertising environment. Based on the initial model, the advertising action corresponding to the current state vector is determined and executed to generate the next state vector; Based on the execution effect of the benchmark advertising strategy, the execution effect of the current state vector, and the execution effect of the next state vector, the cumulative total reward corresponding to the advertising action is determined, and the initial model is trained according to the cumulative total reward to obtain the advertising optimization model. The execution effect of the benchmark advertising strategy is determined based on the benchmark state vector. The baseline state vector is used to characterize the initial advertising environment when no automatic optimization strategy is applied.
3. The advertising optimization method according to claim 2, characterized in that, The process of determining the cumulative total reward for the advertising action based on the execution effect of the baseline advertising strategy, the execution effect of the current state vector, and the execution effect of the next state vector includes: Based on the execution effect of the benchmark advertising strategy, the execution effect of the current state vector, and the execution effect of the next state vector, the immediate reward corresponding to the advertising action at each decision time step is determined. Based on the immediate reward corresponding to the advertising action at each decision time step, the cumulative total reward corresponding to the advertising action is determined.
4. The advertising optimization method according to claim 3, characterized in that, The determination of the cumulative total reward corresponding to the advertising campaign based on the immediate reward corresponding to the advertising campaign at each decision time step includes: Based on the immediate reward corresponding to the advertising action at the current decision time step, the immediate rewards corresponding to a preset number of time steps after the current decision time step, and the discount factors corresponding to the preset number of time steps, the cumulative total reward corresponding to the advertising action is determined. The discount factor is used to measure the relative importance of the instantaneous rewards corresponding to the preset number of time steps.
5. The advertising optimization method according to claim 2, characterized in that, The step of determining the advertising delivery action corresponding to the current state vector based on the initial model includes: Obtain the candidate ad vector set representing multiple candidate ads; For each of the plurality of candidate ads, the initial model performs the following steps to generate its recommendation ranking score: Based on the current state vector and the advertisement vector of the candidate advertisement, a cross-feature matrix is constructed; the cross-feature matrix is used to characterize the fine-grained interaction information between the current state vector and the advertisement vector. The cross feature matrix is input into a convolutional neural network layer to extract and enhance key interaction patterns in the cross feature matrix, thereby obtaining an enhanced cross vector; The enhanced cross vector is input into a multilayer perceptron to predict and output the recommendation ranking score of the candidate advertisement; Based on the recommendation ranking scores of the multiple candidate advertisements, the multiple candidate advertisements are ranked, and the target advertisement with the highest recommendation ranking score is selected. The ad delivery action corresponding to the target ad is taken as the ad delivery action corresponding to the current state vector.
6. The advertising optimization method according to any one of claims 1 to 5, characterized in that, The step of extracting local temporal features and long-term dependency features from the user behavior data, and determining user profile features based on the local temporal features and the long-term dependency features, includes: Based on the user profile model, local temporal features and long-term dependency features of the user behavior data are extracted respectively, and the user profile features are determined based on the local temporal features and long-term dependency features. The training steps of the user profile model include: Obtain sample behavior data; The sample behavior data is input into the initial user profile model to obtain the predicted user profile features output by the initial user profile model. Based on the predicted user profile features, a predicted advertising delivery strategy is determined, and the predicted execution effect feedback of the predicted advertising delivery strategy is determined. Based on the difference between the predicted execution effect feedback and the labeled execution effect feedback of the sample behavior data, a target loss is determined, and the initial user profile model is trained based on the target loss to obtain the user profile model.
7. The advertising optimization method according to any one of claims 1 to 5, characterized in that, The advertising strategy includes the advertising content, duration, and platform.
8. An advertising optimization device, characterized in that, include: The acquisition unit is used to acquire user behavior data, extract local temporal features and long-term dependency features of the user behavior data respectively, and determine user profile features based on the local temporal features and the long-term dependency features; The input unit is used to input the user profile features into the advertising optimization model to obtain the advertising delivery strategy output by the advertising optimization model; An execution unit is used to execute the advertising delivery strategy and obtain feedback on the execution effect of the advertising delivery strategy; An adjustment unit is used to respond to the execution effect feedback, optimize the advertising optimization model, and input the user profile features into the optimized advertising optimization model to obtain the updated advertising delivery strategy output by the optimized advertising optimization model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the advertising optimization method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the advertising optimization method as described in any one of claims 1 to 7.