Method for optimizing advertisement putting effect by combining user group and pre-putting

By detecting the drift of user feature distribution in real time and dynamically dividing user groups, combining causal model and multi-agent reinforcement learning framework, a multi-objective optimization strategy is generated, which solves the problems of strategy lag and resource mismatch in existing advertising delivery technologies, and achieves continuous optimization of advertising resource utilization and conversion rates.

CN120181925AActive Publication Date: 2025-06-20BEIJING HONGTU XINDA TECH CO LTD

Patent Information

Application Number
CN202510251023.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

Existing advertising delivery technologies are difficult to capture the real-time distribution of user feature drift, resulting in the failure of static user grouping models, traditional counterfactual gain estimation ignores the interaction effect between groups, and multi-objective optimization strategies lack joint modeling of dynamic grouping structure and causal gain, resulting in policy lag and resource mismatch.

Method used

By obtaining user behavior data flow in real time, using Wasserstein distance to detect feature distribution drift, dynamically divide user groups, and generate user group tags and user-group membership matrix. Dual machine learning is used to construct a causal model, calculate the potential gain of the user group under different advertising strategies, and generate a counterfactual gain matrix. The user-group membership degree matrix and counterfactual gain matrix are input into the multi-agent reinforcement learning framework, and a multi-objective optimization strategy is dynamically generated, and the user group offset degree and counterfactual prediction error are evaluated through Mahayana distance, triggering group model reconstruction and causal model calibration.

Benefits of technology

Real-time response to the advertising delivery system is realized, the lag and causal inference deviation of the static grouping model are solved, the utilization and conversion rate of advertising resources are dynamically optimized, user fatigue is reduced, and strategy iteration cycle is shortened.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181925A_ABST
    Figure CN120181925A_ABST
Patent Text Reader

Abstract

The invention discloses an advertisement putting effect optimization method combining user groups and pre-putting, and relates to the technical field of advertisement calculation, and the method comprises the steps: obtaining a user behavior data stream in real time, detecting the user feature distribution drift through a Wasserstein distance, triggering an adversarial generative network to dynamically divide the user groups when the drift exceeds a drift judgment threshold, and obtaining a user behavior data stream; outputting a user group label and a user-group membership matrix; and collecting real-time conversion data after execution of the multi-objective optimization strategy, evaluating a user group offset degree and an anti-fact prediction error through a mahalanobis distance, and triggering grouping model reconstruction and causal model calibration. Through closed-loop optimization of dynamic grouping, anti-fact gain calculation, multi-target strategy cooperation and feedback calibration, the advertisement resource utilization rate and conversion rate are continuously improved, the user fatigue is effectively reduced, and the strategy iteration period is shortened to the minute level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computational advertising, and in particular to an optimization method for combining user groups with the advertising delivery effect of pre-delivery. Background Art

[0002] Current advertising delivery technologies have gradually shifted from extensive delivery to precision marketing based on user segmentation. Existing methods mostly adopt static clustering and grouping combined with historical conversion rate prediction models, use collaborative filtering or matrix factorization to mine user interests, and achieve dynamic optimization through the multi-armed bandit algorithm. In the field of causal inference, the double machine learning framework has been introduced into advertising effect evaluation to improve the robustness of treatment effect estimation by decoupling confounding variables. At the same time, the policy optimization system based on reinforcement learning begins to integrate context information to achieve joint decision-making of bidding and frequency control.

[0003] However, the existing technologies have significant defects: static user grouping models are difficult to capture real-time feature distribution drift, resulting in the invalidation of historical grouping rules (such as the problem of clustering center offset caused by user interest migration); traditional counterfactual gain estimation ignores the potential interaction effects between groups and the non-linear superposition characteristics of policy combinations, resulting in aggregation errors of individual treatment effects; multi-objective optimization strategies lack joint modeling of dynamic grouping structures and causal gain confidence intervals, and are prone to action space dimension explosion and policy oscillation. The above problems lead to policy lag and resource mismatch in the advertising delivery system when facing highly dynamic user behaviors. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides an optimization method for combining user groups with the advertising delivery effect of pre-delivery to solve the problems of policy lag caused by dynamic user feature distribution drift and resource mismatch caused by counterfactual gain estimation deviation.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides an optimization method for combining user groups with the advertising delivery effect of pre-delivery, which includes: obtaining a user behavior data stream in real time, detecting user feature distribution drift through the Wasserstein distance, and triggering an adversarial generation network to dynamically divide user groups when the drift determination threshold is exceeded, and outputting user group labels and a user-group membership matrix;

[0008] Constructing a causal model by using double machine learning, calculating the potential gains of each user group under different advertising strategies, and generating a counterfactual gain matrix;

[0009] The user-group membership matrix and counterfactual gain matrix are input into the multi-agent reinforcement learning framework to dynamically generate a multi-objective optimization strategy including bid coefficients, frequency control, and creative selection;

[0010] Collect real-time conversion data after the execution of the multi-objective optimization strategy, evaluate the degree of user group deviation and counterfactual prediction error through Mahalanobis distance, and trigger the reconstruction of the clustering model and the calibration of the causal model.

[0011] As a preferred solution of the optimization method of combining user groups and pre-delivered advertising delivery effects of the present invention, wherein: the output of user group labels and user-group membership matrix, the specific steps are as follows:

[0012] The sliding window intercepts the user behavior stream, encodes the behavior sequence through the time-aware Transformer and normalizes it, and outputs the current / baseline feature matrix;

[0013] Calculate the entropy regularized Wasserstein distance between the current distribution and the benchmark distribution, and dynamically adjust the drift judgment threshold by combining the exponentially weighted moving average;

[0014] If drift is detected, the adversarial generative network is activated to generate virtual features, and the generation process is constrained by the cluster center alignment loss to output the original membership matrix;

[0015] The original membership matrix is ​​Top-3 sparsely processed, adjacent clusters with center distance less than K are merged, and compressed and stored in the distributed database.

[0016] As a preferred solution of the optimization method of combining user groups and pre-delivered advertisement delivery effects of the present invention, wherein: the drift determination threshold refers to a critical value for determining whether the current distribution has changed significantly;

[0017] The dynamically adjusting drift determination threshold comprises:

[0018] Set the initial threshold based on the mean and standard deviation of the historical distribution distance;

[0019] The real-time drift determination threshold is calculated by combining the current distance and the threshold at the previous moment through exponentially weighted moving average.

[0020] As a preferred solution of the optimization method of combining user groups and pre-delivered advertising delivery effects of the present invention, the specific steps of generating the counterfactual gain matrix are as follows:

[0021] The user behavior features are weightedly concatenated with the sparse clustering matrix, and the enhanced features are generated by combining the group statistics interaction terms;

[0022] The individual treatment effects were calculated by orthogonalizing the residuals, weighted aggregation was performed, and a time decay factor was introduced to correct the potential gain of the group;

[0023] Construct a three-dimensional counterfactual gain tensor across strategies and groups, and apply KL divergence constraints to merge the differences in similar group strategies;

[0024] Use BCa Bootstrap sampling to evaluate significance and output the gain matrix with confidence interval corrected.

[0025] As a preferred embodiment of the optimization method for combining user groups with the pre-launch advertising delivery effect of the present invention, wherein: the weighted splicing of user behavior characteristics and the sparse dynamic clustering matrix, and the generation of enhanced features by combining group statistic interaction terms are specifically carried out as follows.

[0026] Standardize the user behavior characteristics and splice them with the cluster membership degree according to the preset weight;

[0027] Calculate the group mean and activity decay value, generate behavior statistical indicators and interact with user characteristics item by item;

[0028] After merging, perform dimensionality reduction through PCA and add Gaussian noise to form enhanced features.

[0029] As a preferred embodiment of the optimization method for combining user groups with the pre-launch advertising delivery effect of the present invention, wherein: the dynamic generation of a multi-objective optimization strategy including bid coefficient, frequency control, and creative selection is specifically carried out as follows.

[0030] Splice the user-group membership matrix and the counterfactual gain matrix into a joint state vector, and superimpose real-time context features;

[0031] Define the bid coefficient as a continuous action, the frequency control as a three-level discrete action, and the creative selection as a Top10 material discrete action;

[0032] Coordinate the agent to dynamically couple the sub-strategy weights through LSTM, and use the FQI algorithm to update the strategy in parallel;

[0033] Use the lower bound of the confidence interval of the counterfactual gain matrix as a hard constraint, and generate a strategy candidate set through the Lagrange multiplier method;

[0034] Adopt the VCG mechanism to screen the Pareto optimal strategy.

[0035] As a preferred embodiment of the optimization method for combining user groups with the pre-launch advertising delivery effect of the present invention, wherein: the generation of the strategy candidate set is specifically carried out as follows.

[0036] Map the lower bound of the confidence interval of BCa Bootstrap sampling to the hard boundary of the bid action;

[0037] Construct an optimization model with constraints and solve it through the Lagrange multiplier method;

[0038] Generate a bidding action within the feasible region and perform a Cartesian product combination with frequency control and creative strategies.

[0039] As a preferred solution of the optimization method for combining user groups and pre-launch advertising delivery effects according to the present invention, wherein: the real-time conversion data after the execution of the multi-objective optimization strategy is collected, and the degree of user group deviation and counterfactual prediction error are evaluated through Mahalanobis distance, triggering the reconstruction of the clustering model and the calibration of the causal model. The specific steps are as follows.

[0040] Align the real-time conversion data with the clustering matrix and calculate the Mahalanobis distance offset of the user group distribution.

[0041] Compare the actual conversion gain with the predicted value to calculate the global error index.

[0042] If the offset exceeds the offset determination threshold and the error exceeds the tolerance range, immediately trigger the reconstruction of the clustering model and the calibration of the causal model.

[0043] If a single index exceeds the limit, trigger after continuously monitoring multiple time windows.

[0044] In a second aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the optimization method for combining user groups and pre-launch advertising delivery effects as described in the first aspect of the present invention is implemented.

[0045] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, wherein: when the computer program is executed by the processor, any step of the optimization method for combining user groups and pre-launch advertising delivery effects as described in the first aspect of the present invention is implemented.

[0046] The beneficial effects of the present invention are as follows: The present invention encodes the user behavior sequence through time-decaying Transformer, dynamically detects the feature distribution drift by combining the entropy-regularized Wasserstein distance, and uses the generative adversarial network to generate a sparse user-group membership matrix to solve the lag of traditional static clustering; Subsequently, enhanced features are constructed by combining the group statistic interaction terms, and the interference of confounding variables is stripped through double machine learning and orthogonalized residual calculation to generate a counterfactual gain matrix with confidence interval correction, eliminating causal inference bias; Furthermore, the sparse structure of the user-group membership matrix is used to isolate the training of multi-agent parameters, and the bidding, frequency control, and creative selection strategies are coupled through the attention mechanism, and the VCG mechanism is combined to screen the Pareto optimal solution; Based on the Mahalanobis distance to quantify the correlation between user group deviation and prediction error, dynamically trigger model reconstruction and calibration, form a "detection - decision - feedback" closed loop, realize the continuous optimization of advertising resource utilization rate and conversion rate, effectively reduce user fatigue, and shorten the strategy iteration cycle to the minute level. Brief Description of the Drawings

[0047] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0048] Figure 1 It is a flowchart of an optimization method for combining user groups with the advertising placement effect of pre-placement in Embodiment 1.

[0049] Figure 2 It is a flowchart of generating a strategy candidate set in Embodiment 1. Detailed Description of the Embodiments

[0050] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention in conjunction with the drawings of the specification.

[0051] Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0052] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments.

[0053] Embodiment 1, referring to Figure 1 and Figure 2 , is the first embodiment of the present invention. This embodiment provides an optimization method for combining user groups with the advertising placement effect of pre-placement, including the following steps:

[0054] S1: Real-time obtain the user behavior data stream, detect the drift of the user feature distribution through the Wasserstein distance, trigger the adversarial generative network to dynamically divide the user groups when the drift determination threshold is exceeded, and output the user group labels and the user-group membership matrix.

[0055] Specifically, it includes the following steps:

[0056] S1.1: The sliding window intercepts the user behavior stream, encodes the behavior sequence through a time-aware Transformer model, reduces the dimension to a fixed dimension by PCA and standardizes it, and outputs the current feature matrix and the reference feature matrix.

[0057] Among them, the sliding window is a method for dynamically intercepting data streams, which obtains the user behavior sequence by a fixed time interval or the number of events. Its advantage is that it can capture the changing trend of user behavior in real time and provide a dynamic data basis for subsequent feature extraction.

[0058] S1.2: Calculate the entropy-regularized Wasserstein distance between the current distribution and the reference distribution, dynamically adjust the drift judgment threshold by exponential weighted moving average, and determine whether the distribution drifts.

[0059] Among them, the drift judgment threshold refers to the critical value for determining whether the current distribution has changed significantly.

[0060] Specifically, the steps for dynamically adjusting the drift judgment threshold are as follows.

[0061] Regard the current feature matrix and the reference feature matrix as the current probability distribution and the reference probability distribution respectively.

[0062] It should be noted that the "current probability distribution" reflects the latest distribution state of real-time user behavior characteristics and is used to capture the dynamic changes of user behavior; the "reference probability distribution" represents the stable distribution state of historical user behavior characteristics and is used as a reference benchmark for comparison to detect whether the current distribution has a significant drift.

[0063] Use the entropy-regularized Wasserstein distance formula to calculate the distance between the current probability distribution and the reference probability distribution.

[0064] Preferably, the entropy-regularized Wasserstein distance can robustly detect the drift of the user feature distribution, and the dynamic threshold adjustment adapts to the change of the data distribution by exponential weighted moving average to avoid misjudgment.

[0065] Within the historical time period, the entropy-regularized Wasserstein distance calculated based on historical data is used as the historical distribution distance.

[0066] It should be noted that the historical data includes the user's behavior data, attribute data, and context data within the historical time period. The specific content includes the user behavior type, behavior time, user age, gender, region, time, location, and scene where the behavior occurs, etc. Through feature extraction, data cleaning, and standardization processing, a feature matrix for calculating the entropy-regularized Wasserstein distance is generated, thereby obtaining the historical distribution distance. This step has important application value in tasks such as user behavior analysis and distribution drift detection.

[0067] Set an initial threshold based on the mean and standard deviation of the historical distribution distance.

[0068] Among them, historical data refers to user behavior data within a certain past time period, which is used to generate a historical feature matrix and calculate the historical distribution distance, for initializing the threshold and providing a reference for dynamically adjusting the threshold.

[0069] Example illustration: Assume that the historical time period is the past 30 days. Calculate an entropy-regularized Wasserstein distance every day to obtain 30 historical distribution distances D1, D2, …, D 30 . Calculate the mean μ history = 1.2, and the standard deviation σ history = 0.3. The adjustment coefficient α = 2. Then the initial threshold T0 is: T0 = 1.2 + 2 × 0.3 = 1.8.

[0070] Use the exponential weighted moving average (EWMA) to calculate the real-time drift determination threshold based on the distance between the prior probability distribution and the reference probability distribution and the drift determination critical value at the previous moment.

[0071] Among them, the "drift determination critical value" refers to the statistical boundary value calculated by a dynamic algorithm, which is used to determine whether the user behavior distribution has a significant shift. Its technical essence is equivalent to the dynamic drift detection threshold.

[0072] S1.3: If drift is detected, activate the adversarial generation network.

[0073] Preferably, activating the adversarial generation network (GAN) can, when detecting the drift of the user behavior distribution, synthesize high-quality virtual features through the generator to alleviate the data sparsity problem. At the same time, the discriminator outputs the user-group membership matrix to ensure the accuracy of the clustering result; the generation process is constrained by the clustering center alignment loss to make the distribution of the generated virtual features consistent with that of the real features, further improving the robustness and generalization ability of the clustering model, so as to dynamically adapt to the changes in user behavior and provide a more accurate user clustering basis for subsequent multi-objective optimization strategies.

[0074] S1.4: The generator synthesizes virtual features, the discriminator outputs the user-group membership, and the generation process is constrained by the clustering center alignment loss to output the original membership matrix.

[0075] Specifically, the generator receives user behavior features as input, synthesizes virtual features through a multi-layer neural network, and the discriminator receives the virtual features and real user features output by the generator, distinguishes virtual features from real features through adversarial learning, and outputs a user-group membership matrix at the same time; during the generation process, a clustering center alignment loss is adopted to calculate the distance between the generated features and the clustering centers of the real features, and the generator is optimized through backpropagation to ensure that the distribution of the generated features is consistent with that of the real features; finally, the original membership matrix is output, which contains the membership of each user to all user groups and is used for subsequent sparsification and compression processing.

[0076] S1.5: Perform Top-3 sparsification on the original membership matrix, merge adjacent clusters with a center distance less than K, compress it using floating-point precision and store it in a distributed database, output the final user-group labels and membership matrix, and store them in the distributed database.

[0077] Among them, "Top-3 sparsification" is a data sparsification technology for the user-group membership matrix. Its core idea is to retain the membership of each user to the top three user groups, and set the membership of the remaining user groups to zero or ignore them. This significantly reduces the complexity of data storage and calculation, while retaining the most important user grouping information.

[0078] It should be noted that "K" is a preset clustering center distance threshold, which is used to determine whether two clusters are close enough to be merged. Specifically, when the center distance between two clusters is less than K, it is considered that they belong to the same category, and they are merged into one cluster to reduce the redundancy of user groups and improve the simplicity and practicality of the grouping results.

[0079] The specific steps to obtain the "preset clustering center distance threshold" are as follows: First, run a clustering algorithm (such as K-means or hierarchical clustering) based on historical data or sample data, and calculate the distance distribution between different clustering centers; then, according to business requirements and data characteristics, select a suitable distance threshold as the K value, such as selecting the median or a specific percentile of the distance distribution; finally, evaluate the rationality of the K value through cross-validation or actual application effects, and adjust it if necessary to optimize the clustering results.

[0080] S2: Construct a causal model using dual machine learning, calculate the potential gains of each user group under different advertising strategies, and generate a counterfactual gain matrix.

[0081] Specifically, it includes the following steps:

[0082] S2.1: Read the user-group membership matrix from the distributed database, decompress it and reconstruct it into a sparsified dynamic grouping matrix.

[0083] Among them, during the decompression process, the ZFP floating-point precision compression algorithm is used for decompression, and the decompressed data is restored to the original structure of the sparse dynamic clustering matrix to ensure that subsequent calculations can be directly used.

[0084] S2.2: Weightedly splice the user behavior characteristics with the sparse dynamic clustering matrix, generate enhanced features by combining the group statistic interaction terms, respectively fit the propensity score model and the potential outcome prediction model, and calculate the individual treatment effect through orthogonalized residuals.

[0085] Specifically, based on the user unique identifier, extract user behavior characteristics (such as click-through rate, conversion rate, and activity decay value) from the user behavior data stream, and obtain the membership degree of the user's group from the sparse dynamic clustering matrix;

[0086] It should be noted that only the Top-3 user group membership degrees are retained in the sparse dynamic clustering matrix.

[0087] Perform standardization processing on the user behavior characteristics (such as Z-score standardization);

[0088] Match the standardized user behavior characteristics with the sparse dynamic clustering matrix according to the user unique identifier to ensure a one-to-one correspondence between the characteristics and the membership degrees;

[0089] Splice the user behavior characteristics and the membership degrees of the user's group according to a preset weight ratio to generate a preliminary joint feature matrix;

[0090] Specifically, splice the user behavior characteristics (weight 0.4) and the membership degrees of the user's group (weight 0.6) by column.

[0091] According to the historical user behavior data of the user's group, calculate the group mean and the group activity decay value to generate behavior statistical indicators;

[0092] Specifically, first, based on the historical user behavior data of the user's group, extract the behavior characteristics of each user within the historical time period, such as behavior frequency, behavior type distribution, etc.; then calculate the mean value of all user behavior characteristics within each group to obtain the group mean; then calculate the activity decay value of each user behavior based on the time decay function, and combine the activity decay values of the users within the group to calculate the group activity decay value; finally, perform a weighted combination of the group mean and the group activity decay value to generate behavior statistical indicators for describing the behavior characteristics and active status of the group.

[0093] Example: Suppose a group contains 100 users. Calculate the average click frequency of these users in the past 30 days as the group average. At the same time, calculate the decay value of each user's click activity based on the exponential decay function and take the average as the group activity decay value. Finally, generate a behavior statistical indicator by weighting the two to describe the click behavior and activity level of the group.

[0094] Interact the user behavior characteristics with the behavior statistical indicators of the corresponding group item by item according to the feature dimension to generate group enhanced interaction terms;

[0095] Specifically, extract the user behavior characteristics and the behavior statistical indicators of the corresponding group, and perform dot product or weighted summation operations item by item according to the feature dimension to generate group enhanced interaction terms. For example, multiply the click frequency of the user by the average click frequency of the group, or perform weighted summation of the purchase amount of the user and the decay value of the purchase amount of the group. Finally, generate enhanced interaction terms that reflect the interaction relationship between user behavior and group characteristics.

[0096] Example: Suppose a user's click frequency is 5 times, and the average click frequency of the group to which the user belongs is 10 times. Then multiply the two to generate a group enhanced interaction term of 50. If the user's purchase amount is 200 yuan and the decay value of the group's purchase amount is 150 yuan, then perform weighted summation of the two to generate a group enhanced interaction term of 350, which is used to describe the interaction relationship between user behavior and group characteristics.

[0097] Merge the preliminary joint feature matrix and the group enhanced interaction terms by column, perform PCA dimensionality reduction (retaining 95% variance) and add Gaussian noise (standard deviation 0.01) to form enhanced features.

[0098] It should also be noted that the model objective of the propensity score model: estimate the conditional probability that a user accepts a specific advertising strategy under given context conditions, and the expression is as follows:

[0099] v(X) = P(T = t|C);

[0100] Among them, v(C) represents the propensity score, and the value range is [0,1]. T is the intervention variable (such as whether to apply a high-price strategy), usually a binary variable T ∈ {0,1}, and X is the input feature, including enhanced features such as user behavior characteristics, cluster membership, and group statistical interaction terms;

[0101] The input layer architecture of the propensity score model:

[0102] Hierarchy Dimension Description User behavior characteristics 128 Click-through rate / conversion rate, etc. after standardization Cluster membership degree 3 Top-3 membership degree vectors after sparsification Group statistical interaction 64 Hadamard product of user characteristics and group mean Noise layer 0.01σ Gaussian noise injection

[0103] Furthermore, the model objective of the potential outcome prediction model: predict the potential outcomes of users under different advertising strategies, and the expression is as follows:

[0104]

[0105] Among them, represents the orthogonalized potential outcome under policy a, and f a (X) represents the linear prediction function for policy a, is the set of policies (such as bid coefficient bins), represents for the set of policies any policy a in;

[0106] Input layer architecture:

[0107]

[0108]

[0109] S2.3: Combine the user-group membership matrix to weighted aggregate the individual treatment effect (ITE), and use a time decay factor to correct the group potential gain.

[0110] Specifically, use a time decay factor to correct the group potential gain, and the expression is:

[0111]

[0112] Among them, G k,a is the original potential gain of group k under advertising policy a, representing the average conversion rate improvement of this group under advertising policy a, k represents the group number, α represents the timeliness adjustment coefficient, and the value range is [0, 0.5], which is determined by grid search (the optimal value for the customer scenario is 0.3);

[0113] Furthermore, Recency k is the activity decay factor of group k, and the expression is:

[0114]

[0115] Among them, Δt i represents the time interval (unit: days) of the most recent behavior of user u i , N k represents the number of users in group k, i is the index number of users in group k, and the value range is i ∈ {1, 2,..., N k}}, c k represents the set of users in group k, that is, all users belonging to group k, expresses the behavior freshness decay value of user u i ;

[0116] Example illustration: Assume that group k has 3 users, and the time intervals of their most recent behaviors are: user u1: Δt1 = 1 day; user u2: Δt2 = 1 day; user u3: Δt3 = 1 day.

[0117] Then Recency k The calculation process is as follows:

[0118] Calculate the decay value of the freshness of each user's behavior: User u1: e -0.1×1 = e -0.1 ≈0.9048; User u2: e -0.1×3 = e -0.3 ≈0.7408; User u3: e -0.1×5 = e -0.5 ≈0.6065.

[0119] Sum and take the average:

[0120] Preferably, compared with the traditional static gain aggregation method, this step effectively solves the prediction deviation problem caused by the insufficient timeliness of historical data by dynamically correcting the group potential gain using the time decay factor. Based on the exponential decay function of the time interval of the user's recent behavior, it can quantify the freshness of the user's behavior. For example, the weight of the user's behavior 3 days ago is retained relatively high, while the weight 30 days ago is significantly reduced, thus effectively suppressing the interference of stale data and improving the real-time performance. At the same time, by adjusting the coefficient to control the decay intensity and combining the confidence interval calibration, the gain correction effect of the long-tail user group is more stable. In the scenario where the user's behavior changes rapidly, the prediction accuracy is significantly improved, and the adaptability of the model to the dynamic environment is significantly enhanced.

[0121] S2.4: Construct a three-dimensional counterfactual gain tensor across strategies and groups, and apply KL divergence constraints to merge the differences of similar group strategies.

[0122] It should be noted that after calculating the group gain, a three-dimensional counterfactual gain tensor across strategies and groups is further constructed to quantify the substitution relationship between strategies and support the joint optimization of multiple strategies.

[0123] Specifically, the expression is:

[0124]

[0125] Among them, represents the gain difference between group k under strategies a and a ′ in the three-dimensional counterfactual gain matrix. a and a ′ represent two different strategies in the advertising strategy pair, used to compare the gain differences of group k under strategies a and a ′ . k represents the group number, G k,a is the original gain of group k under strategy a, representing the average conversion rate improvement of this group. τ is the temperature parameter (τ = 0.1), used to control the discrimination degree of strategy differences, |G k,a - Gk,a′ | represents the absolute difference in gain between strategy a and a ′ for retaining magnitude information;

[0126] It should be noted that by converting the strategy differences into probability distributions, joint optimization across strategies is supported. For example, if represents the probability that strategy a is better than a ′ is 80%.

[0127] Furthermore, by using KL-divergence constraints to merge the strategy differences of similar groups, the balance between the granularity of clustering and the stability of strategies is ensured.

[0128] Specifically, calculate the KL-divergence between two groups k and k ′ , and the expression is:

[0129] D KL (M k ||M k′ );

[0130] where D KL represents the Kullback-Leibler divergence, i.e., the KL-divergence, M k represents the user-group membership distribution vector of group k, and M k′ represents the user-group membership distribution vector of group k ′ , and k ′ represents the group number of another group compared with group k;

[0131] If D KL < 0.1, then the strategy differences of the two groups are forced to be merged to ensure the balance between the granularity of clustering and the stability of strategies. The formula is as follows:

[0132] Basis for parameter values: α = 0.3: Customer A / B testing shows that this value increases the long-tail user coverage rate by 52% and has the smallest prediction error. τ = 0.1: Experiments show that this temperature parameter can effectively distinguish the head strategy differences (the discrimination of the top-3 strategies is increased by 40%). KL threshold 0.1: When D KL < 0.1, the similarity of the user behavior distributions of the two groups exceeds 95% (KS test result).

[0133] Preferably, compared with the isolated analysis of the conventional two-dimensional gain matrix, this step constructs a three-dimensional counterfactual gain tensor across strategies - groups, quantifies the differences of different strategies among groups, and uses normalization to convert the absolute gain differences into probability distributions, so as to more intuitively reflect the relative advantages among strategies. Further, the KL divergence constraint is introduced to automatically merge similar groups. In the advertising budget allocation scenario, it not only reduces the number of segmented groups, but also maintains a high strategy discrimination degree, significantly reduces the computational complexity and improves the strategy stability, and fundamentally avoids the risk of strategy oscillation caused by over-segmentation of groups.

[0134] S2.5: Use BCa Bootstrap sampling to evaluate statistical significance, output the counterfactual gain matrix with the confidence interval corrected, and store it in the distributed database.

[0135] Among them, "BCa Bootstrap sampling" is an improved self-sampling method. By adjusting the confidence interval through bias correction and acceleration factors, it can more accurately evaluate the distribution and significance of statistics.

[0136] Specifically, based on the counterfactual gain matrix, BCa Bootstrap sampling is carried out. Multiple sample matrices are generated through multiple resamplings, the statistics (such as mean, variance, etc.) of each sample matrix are calculated, the distribution of the statistics is constructed, the bias correction factor and acceleration factor are calculated according to the distribution, the upper and lower limits of the confidence interval are adjusted, the counterfactual gain matrix is corrected based on the adjusted confidence interval, the counterfactual gain matrix with the confidence interval corrected is output, and it is stored in the distributed database for subsequent generation of multi-objective optimization strategies.

[0137] S3: Input the user-group membership matrix and the counterfactual gain matrix into the multi-agent reinforcement learning framework to dynamically generate multi-objective optimization strategies including bid coefficients, frequency control, and creative selection.

[0138] Specifically, it includes the following steps:

[0139] S3.1: Concatenate the user-group membership matrix and the counterfactual gain matrix into a joint state vector, and superimpose real-time context features to form a multi-dimensional state space.

[0140] Specifically, expand the user-group membership matrix in vector form according to the user dimension, concatenate it with the counterfactual gain matrix according to the same user dimension to generate a joint state vector, extract real-time context features (such as time, location, scenario, etc.), encode the real-time context features into numerical vectors, and perform feature superposition with the joint state vector to form a multi-dimensional state space for input into the subsequent multi-agent reinforcement learning framework to support the dynamic generation of multi-objective optimization strategies.

[0141] Preferably, "forming a multi-dimensional state space" can unify and integrate user clustering information, counterfactual gains, and real-time context features, providing comprehensive input information for the multi-agent reinforcement learning framework, enhancing the decision-making model in the multi-agent reinforcement learning framework's perception ability of the dynamic changes in user behavior and the context environment, thereby generating more accurate multi-objective optimization strategies including bid coefficients, frequency control, and creative selection, and improving the advertising delivery effect and user experience.

[0142] S3.2: Define the bid coefficient as a continuous action, the frequency control as a three-level discrete action, and the creative selection as a discrete action of the top 10 materials, and dynamically couple the relationships among the three through an attention mechanism.

[0143] Among them, "Top10" refers to the top 10 materials selected from the candidate creative material library, and these materials are usually sorted based on indicators such as user interest, click-through rate, and conversion rate, and are used to optimize the creative selection strategy to ensure that the advertising materials delivered are the most attractive and effective.

[0144] S3.3: Use the joint state vector to calculate the click-through rate (CTR), user lifetime value (LTV), and fatigue penalty term in real time, and calibrate the reward weights through the group potential gains in the counterfactual gain matrix to strictly align the reward signal with the causal inference result.

[0145] S3.4: Deploy a long short-term memory neural network (LSTM) to coordinate the agents, and according to the temporal dependence relationship of the user-group membership matrix, coordinate the policy weights of each sub-agent to achieve multi-objective collaborative optimization.

[0146] Specifically, input the temporal data of the user-group membership matrix into the long short-term memory neural network (LSTM), train the LSTM model to learn the dynamic change rules of the user group, and output the weight coefficients of each sub-agent (such as bid, frequency, and creative agents); capture the temporal dependence relationship of the user-group membership through the hidden state of the LSTM, dynamically adjust the policy weights of the sub-agents, and generate a coordinated policy combination by combining the current state and temporal information for the execution of subsequent multi-objective optimization strategies.

[0147] Preferably, coordinating the policy weights of each sub-agent can dynamically balance the priorities of multi-objective policies (such as bid, frequency, and creativity) according to the temporal changes of the user group, avoid policy conflicts, and improve the synergy of the policy combination; through the LSTM's modeling of temporal dependence, ensure that the policy weights are consistent with the user behavior trend, enhance the dynamic adaptability of advertising delivery, and thus optimize the real-time conversion effect and user experience.

[0148] S3.5: Use the coordinated policy weights of each sub-agent to guide the FQI algorithm for parallel policy update. Isolate the training processes of different group parameters through the sparse structure of the user-group membership matrix to obtain the original policy action set.

[0149] It should be noted that the specific steps for parallel policy update are as follows: Based on the coordinated policy weights of each sub-agent, distribute the FQI algorithm to multiple computing nodes. Each node independently calculates the policy update of a specific sub-agent, uses the distributed computing framework to synchronize the intermediate results between nodes, and aggregates the updated policy weights to achieve parallel policy update.

[0150] Example illustration: Suppose there are three sub-agents responsible for bid coefficient, frequency control, and creative selection respectively. Assign the policy update tasks of each sub-agent to independent computing nodes. Node 1 loads the user data and model parameters related to the bid coefficient, Node 2 loads the user data and model parameters related to frequency control, and Node 3 loads the user data and model parameters related to creative selection. Each node independently calculates the policy update based on the FQI algorithm, synchronizes the intermediate results through the distributed computing framework (such as Spark or Ray), aggregates the updated policy weights, and completes the parallel policy update.

[0151] S3.6: Use the lower bound of the confidence interval of the counterfactual gain matrix as a hard boundary, and reconstruct the feasible region of the bid action through the Lagrange multiplier method to generate a set of policy candidates.

[0152] Specifically, through the BCa Bootstrap method, extract the lower bound estimate of the confidence interval of the counterfactual gain matrix and map it to a hard constraint on the bid action.

[0153] Specifically, based on the counterfactual gain matrix, perform BCa Bootstrap sampling to generate multiple sample matrices, calculate the statistics (such as the mean) of each sample matrix, construct the distribution of the statistics, calculate the bias correction factor and acceleration factor according to the distribution, adjust the upper and lower limits of the confidence interval, extract the lower bound estimate of the confidence interval, and map it to a hard constraint on the bid action to limit the bid range and ensure the robustness and reliability of the bid strategy.

[0154] Preferably, "mapping to a hard constraint on the bid action" can directly convert the lower bound estimate of the confidence interval of the counterfactual gain matrix into the upper or lower limit of the bid range, ensuring that the bid strategy is executed within a robust range and avoiding waste of advertising budget or insufficient exposure caused by too high or too low bids. This hard constraint provides a reliable basis for subsequent multi-objective optimization strategies, improves the effect and budget utilization rate of advertising placement, and enhances the interpretability and controllability of the strategy.

[0155] Define the basic revenue function based on the business objective and construct a constrained optimization model in combination with the hard constraint.

[0156] Specifically, define the basic revenue function based on business goals, and the steps are as follows: Determine the key variables of the revenue function according to business goals (such as maximizing the ad click-through rate or conversion rate), such as the number of clicks, the number of conversions, or ad revenue; Quantify the key variables into mathematical expressions, such as the number of clicks multiplied by the click unit price, or the number of conversions multiplied by the conversion unit price; Introduce a weight coefficient to balance the priorities of different business goals, for example, the click-through rate weight is 0.6, and the conversion rate weight is 0.4; Finally, generate the basic revenue function for the construction of the subsequent optimization model.

[0157] Construct an optimization model with constraints by combining hard constraints, and the steps are as follows: Map the lower bound estimate value of the confidence interval of the counterfactual gain matrix to the hard constraints of the bidding action, such as the upper and lower limits of the bid; Convert the hard constraints into mathematical inequalities, such as the bid range constraint is:

[0158] b min ≤b≤b max ;

[0159] where b represents the variable of the bidding action, that is, the bid value set during ad placement, b min represents the lower limit of the bidding action, that is, the lowest allowable bid value, b max represents the upper limit of the bidding action, that is, the highest allowable bid value.

[0160] Example: Assume b min = 0.5 yuan, b max = 5 yuan, then the value range of the bidding action b is [0.5, 5] yuan, ensuring that the bid is neither lower than 0.5 yuan (to avoid insufficient exposure) nor higher than 5 yuan (to avoid budget waste).

[0161] Furthermore, integrate the basic revenue function with the hard constraints to construct an optimization problem with constraints, such as maximizing the revenue function under the bid constraints; Solve the optimization problem through the Lagrange multiplier method or the penalty function method, and output the bid strategy that meets the business goals and constraint conditions for subsequent ad placement execution.

[0162] Solve the optimization model through the Lagrange multiplier method to obtain the optimal solution that satisfies the constraints;

[0163] Execute action generation (including sampling and enumeration) within the feasible region defined by the hard constraints, and perform a Cartesian product combination with the frequency control and creative selection strategies to generate a set of strategy candidates.

[0164] Specifically, generate a set of bidding actions by sampling or enumeration within the bid feasible region defined by the hard constraints, such as within the bid range [b min ,b maxUniformly sample 100 bid values within; perform a Cartesian product combination of the bid action set with the frequency control strategy (such as high, medium, and low gears) and the creative selection strategy (such as the top 10 materials) to generate a set of candidate strategies. For example, 100 bid values × 3 gears of frequency control × 10 creative materials = 3,000 strategy combinations, which are used for the evaluation and selection of subsequent multi-objective optimization strategies.

[0165] Preferably, generating a set of candidate strategies provides a comprehensive basis of strategy combinations for the evaluation and selection of subsequent multi-objective optimization strategies, ensuring the coordination of bid, frequency control, and creative selection strategies; through the Cartesian product combination, all possible strategy combinations are covered, enhancing the flexibility and accuracy of strategy optimization, thereby maximizing the advertising placement effect within the hard constraints while meeting the business requirements of frequency control and creative selection.

[0166] S3.7: Use BCa Bootstrap (Bias-Corrected Accelerated Bootstrap) to verify the set of candidate strategies, and screen the Pareto-optimal strategies through the Vickrey-Clarke-Groves (incentive-compatible auction mechanism design framework) mechanism as the multi-objective optimization strategies.

[0167] It should be noted that the execution results of the Pareto-optimal strategies will be written into the user behavior log in real time through the Kafka stream, triggering an incremental update of the user-group membership matrix (update period ≤ 5 minutes).

[0168] S4: Collect the real-time conversion data after the execution of the multi-objective optimization strategies, evaluate the degree of user group deviation and the counterfactual prediction error through the Mahalanobis distance, and trigger the reconstruction of the clustering model and the calibration of the causal model.

[0169] Specifically, it includes the following steps:

[0170] S4.1: Capture user conversion events through the real-time log stream of the advertising placement engine, associate the sparse dynamic clustering matrix and the counterfactual gain matrix according to the user ID, align the time window and standardize the features, and output the aligned real-time conversion data set.

[0171] Among them, the real-time log of the advertising placement engine is the core component of the advertising system, responsible for making real-time decisions on advertising display strategies (such as bidding, frequency control, and creative selection).

[0172] S4.2: Load the benchmark feature matrix and the current feature matrix, calculate the covariance matrix and the mean vector, and quantify the deviation of the user group distribution through the Mahalanobis distance.

[0173] S4.3: Input the historical Mahalanobis distance sequence into the exponentially weighted moving average model, dynamically generate the deviation determination threshold based on the historical mean and volatility, and determine whether the current deviation of the user group distribution exceeds the stable range.

[0174] Among them, the historical mean refers to the average value of the entropy-regularized Wasserstein distance within a historical time period. The volatility refers to the standard deviation of the historical distribution distance, which is used to quantify the dispersion degree of historical data.

[0175] S4.4: Extract the actual conversion gain of the user group from the real-time conversion dataset, compare it with the predicted value in the counterfactual gain matrix, calculate the group-level error, and weighted aggregate it into a global prediction error metric.

[0176] S4.5: If the distribution offset of the user group exceeds the offset determination threshold and the global prediction error metric exceeds the tolerance range (0.15), immediately trigger the reconstruction of the clustering model and the calibration of the causal model. If a single metric exceeds the limit (such as only the global prediction error exceeds 0.15), it will be triggered after continuously monitoring multiple time windows.

[0177] It should be noted that the tolerance range (0.15) is set based on the distribution of the global prediction error metric in historical data. Calculate the mean and standard deviation of the error, and determine the tolerance range in combination with business requirements; set the error tolerance threshold by analyzing the impact of the error on business objectives (such as ad click-through rate, conversion rate); for example, if the historical error mean is 0.10 and the standard deviation is 0.05, select the mean plus one standard deviation (0.15) as the tolerance range; verify the impact of the threshold of 0.15 on the model performance and business metrics through experiments to ensure that the trigger conditions for model reconstruction and calibration are reasonable and effective, and finally determine 0.15 as the tolerance range.

[0178] Among them, the reconstruction of the clustering model refers to the retraining of the generative adversarial network and the update of the user-group membership matrix; the calibration of the causal model refers to the iterative correction of the propensity score and the potential outcome prediction model.

[0179] It should be noted that the clustering model is the core component for dynamically dividing user groups. Based on the user behavior data stream, it intercepts the behavior sequence through a sliding window, encodes features using a time-decaying Transformer, and uses a generative adversarial network (GAN) to generate the user-group membership matrix. The core output of this clustering model is a sparsified dynamic clustering matrix, which is used to describe the membership relationship between users and groups (such as the probability that user A belongs to group 1 is 0.8 and the probability of group 2 is 0.2).

[0180] The causal model is the core component for counterfactual gain calculation. Based on the user clustering results, it constructs a propensity score model and a potential outcome prediction model through double machine learning (DML) to estimate the causal effects of different advertising strategies on each user group (such as strategy A increases the conversion rate of group 1 by 15%). The core output is the counterfactual gain matrix, which is used to guide the generation of multi-objective optimization strategies.

[0181] It should also be noted that the reconstruction of the clustering model means that when the distribution shift of the user group is detected (Mahalanobis distance exceeds the threshold), the adversarial generation network needs to be retrained to generate a new clustering matrix to overwrite the original result. The reconstruction process directly depends on the GAN generation, sparsification, and clustering merging rules to ensure that the clustering structure is dynamically updated with the data distribution.

[0182] Causal model calibration means that when the counterfactual prediction error exceeds the tolerance range (0.15) of the global prediction error metric, it is necessary to regenerate enhanced features (weighted splicing and interaction terms) based on the latest clustering matrix, iteratively update the propensity score model (Focal Loss), and correct the potential outcome prediction model (SHAP constraint). The calibration process directly invokes the gain calculation process and feature enhancement rules to ensure that causal inference is consistent with real-time data.

[0183] S4.6: Retrain the adversarial generation network based on the current feature matrix, generate a new sparsified dynamic clustering matrix using the clustering center alignment loss, and overwrite the original clustering result through the sparsification and merging rules.

[0184] Specifically, retrain the adversarial generation network based on the current feature matrix to generate new virtual features. Constrain the distribution consistency between the virtual features and the real features through the clustering center alignment loss. Use the K-means clustering algorithm to cluster the new feature matrix to generate a new sparsified dynamic clustering matrix. Merge similar clusters according to the clustering center distance threshold to overwrite the original clustering result, and output the updated user-group membership matrix for the optimization and execution of subsequent advertising placement strategies.

[0185] This embodiment also provides a computer device, which is applicable to the situation of optimizing the advertising placement effect by combining user groups and pre-placed advertisements, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method for optimizing the advertising placement effect by combining user groups and pre-placed advertisements as proposed in the above embodiment.

[0186] The computer device may be a terminal, which includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, carrier networks, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, trackball, or touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0187] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the optimization method for realizing the combination of user groups and the advertising delivery effect proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disc.

[0188] In summary, the present invention encodes the user behavior sequence through a time decay Transformer, dynamically detects the feature distribution drift by combining the entropy-regularized Wasserstein distance, and uses a generative adversarial network to generate a sparse user-group membership matrix to solve the lag of traditional static grouping. Subsequently, an enhanced feature is constructed by combining the group statistic interaction term, and the interference of confounding variables is stripped through double machine learning and orthogonalized residual calculation to generate a counterfactual gain matrix with confidence interval correction, eliminating causal inference bias. Furthermore, the sparse structure of the user-group membership matrix is used to isolate the multi-agent parameter training, and the bid, frequency control, and creative selection strategies are coupled through an attention mechanism. The VCG mechanism is combined to screen the Pareto optimal solution. Based on the Mahalanobis distance, the correlation between user group deviation and prediction error is quantified, and the model reconstruction and calibration are dynamically triggered to form a "detection-decision-feedback" closed loop, realizing the continuous optimization of the advertising resource utilization rate and conversion rate, effectively reducing the user fatigue degree, and shortening the strategy iteration cycle to the minute level.

[0189] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A method for optimizing the effect of advertising delivery by combining user groups and pre-delivery, characterized in that: include, Obtain user behavior data streams in real time, detect user feature distribution drift through Wasserstein distance, trigger the adversarial generative network to dynamically divide user groups when the drift determination threshold is exceeded, and output user group labels and user-group membership matrix; Use dual machine learning to build a causal model, calculate the potential gains of each user group under different advertising strategies, and generate a counterfactual gain matrix; The user-group membership matrix and counterfactual gain matrix are input into the multi-agent reinforcement learning framework to dynamically generate a multi-objective optimization strategy including bid coefficients, frequency control, and creative selection; Collect real-time conversion data after the execution of the multi-objective optimization strategy, evaluate the degree of user group deviation and counterfactual prediction error through Mahalanobis distance, and trigger the reconstruction of the clustering model and the calibration of the causal model.

2. The method for optimizing the effect of advertising delivery by combining user groups and pre-delivery as claimed in claim 1, characterized in that: The specific steps of outputting user group labels and user-group membership matrix are as follows: The sliding window intercepts the user behavior stream, encodes the behavior sequence through the time-aware Transformer and normalizes it, and outputs the current / baseline feature matrix; Calculate the entropy regularized Wasserstein distance between the current distribution and the benchmark distribution, and dynamically adjust the drift judgment threshold by combining the exponentially weighted moving average; If drift is detected, the adversarial generative network is activated to generate virtual features, and the generation process is constrained by the cluster center alignment loss to output the original membership matrix; The original membership matrix is ​​Top-3 sparsely processed, adjacent clusters with center distance less than K are merged, and compressed and stored in the distributed database.

3. The method for optimizing the effect of advertising delivery by combining user groups and pre-delivery as claimed in claim 2, characterized in that: The drift determination threshold refers to a critical value for determining whether the current distribution has changed significantly; The dynamically adjusting drift determination threshold comprises: Set the initial threshold based on the mean and standard deviation of the historical distribution distance; The real-time drift determination threshold is calculated by combining the current distance and the threshold at the previous moment through exponentially weighted moving average.

4. The method for optimizing the effect of advertising delivery by combining user groups and pre-delivery as claimed in claim 3, characterized in that: The specific steps of generating the counterfactual gain matrix are as follows: The user behavior features are weightedly concatenated with the sparse clustering matrix, and the enhanced features are generated by combining the group statistics interaction terms; The individual treatment effects were calculated by orthogonalizing the residuals, weighted aggregation was performed, and a time decay factor was introduced to correct the potential gain of the group; Construct a three-dimensional counterfactual gain tensor across strategies and groups, and apply KL divergence constraints to merge strategy differences between similar groups; BCa Bootstrap sampling was used to evaluate significance, and the confidence interval-corrected gain matrix was output.

5. The method for optimizing the effect of advertising delivery by combining user groups and pre-delivery as claimed in claim 4, characterized in that: The user behavior features are weightedly concatenated with the sparse dynamic clustering matrix, and the enhanced features are generated by combining the group statistics interaction items. The specific steps are as follows: Standardize user behavior characteristics and combine them with group membership according to preset weights; Calculate the group mean and activity decay value, generate behavioral statistical indicators and interact with user characteristics item by item; After merging, PCA is used to reduce the dimension and Gaussian noise is added to form enhanced features.

6. The method for optimizing the effect of advertising delivery by combining user groups and pre-delivery as claimed in claim 5, characterized in that: The dynamic generation of a multi-objective optimization strategy including bid coefficients, frequency control and creative selection has the following specific steps: The user-group membership matrix and the counterfactual gain matrix are concatenated into a joint state vector, and real-time context features are superimposed; Define the bid coefficient as a continuous action, frequency control as three discrete actions, and creative selection as a Top 10 material discrete action; Through LSTM, the agent dynamically couples the sub-strategy weights and updates the strategies in parallel using the FQI algorithm. Taking the lower bound of the confidence interval of the counterfactual gain matrix as a hard constraint, the strategy candidate set is generated by the Lagrange multiplier method; The VCG mechanism is used to screen the Pareto optimal strategy.

7. The method for optimizing the effect of advertising delivery by combining user groups and pre-delivery as claimed in claim 6, characterized in that: The specific steps of generating a strategy candidate set are as follows: Map the lower bound of the confidence interval of the BCa Bootstrap sampling to the hard boundary of the bidding action; Construct a constrained optimization model and solve it using the Lagrange multiplier method; Generate bidding actions within the feasible domain and perform Cartesian product combination with frequency control and creative strategies.

8. The method for optimizing the effect of advertising delivery by combining user groups and pre-delivery as claimed in claim 7, characterized in that: The real-time conversion data after the execution of the multi-objective optimization strategy is collected, the user group deviation degree and the counterfactual prediction error are evaluated by Mahalanobis distance, and the clustering model reconstruction and causal model calibration are triggered. The specific steps are as follows: Align the real-time conversion data with the clustering matrix and calculate the Mahalanobis distance offset of the user group distribution; Compare the actual conversion gain with the predicted value to calculate the global error index; If the offset exceeds the offset determination threshold and the error exceeds the tolerance range, the clustering model reconstruction and causal model calibration are immediately triggered; If a single indicator exceeds the limit, it will be triggered after continuous monitoring of multiple time windows.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for optimizing the effect of advertising delivery combining user groups and pre-delivery are implemented as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the optimization method combining user groups and pre-delivered advertisement delivery effects described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Micro-grid flexible supply multi-period energy storage value evaluation method

    CN118917736A

  • Private domain traffic scheduling and content distribution method and system based on group behavior analysis

    CN119377498A

  • Data resource integration system based on cloud computing

    CN119475073A

  • Model-agnostic multi-factor metric drift attribution

    US20240420009A1

  • Systems, methods, kits, and apparatuses for generative artificial intelligence, graphical neural networks, transformer models, and converging technology stacks in value chain networks

    WO2024226801A2

Cited By

  • Non-invasive driver driving state sensing system

    CN121549825A

  • Advertisement putting effect monitoring and optimizing method and system

    CN121616363A