Marketing flow dynamic regulation and control method based on real-time data feedback
By combining real-time data feedback and dynamic adjustment methods with budget consumption stage perception and value competition preference function, the problem of collaborative modeling of budget consumption and traffic value in real-time bidding advertising system is solved, achieving the effects of budget smoothing and high-value traffic capture.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-03-10
AI Technical Summary
Existing real-time bidding advertising systems struggle to dynamically adjust the balance between budget consumption and traffic value, resulting in uneven budget spending and insufficient ability to capture high-value traffic.
By leveraging real-time data feedback, a budget consumption phase awareness module and a value competition preference function are employed to dynamically adjust bidding strategies. Combined with pulse-style bidding management of traffic clusters, this optimizes the processing of strategy preferences and traffic characteristics during the budget consumption phase.
It achieves smooth budget consumption and real-time capture of high-value traffic, improving budget utilization efficiency and campaign effectiveness, and enhancing responsiveness to changes in the market environment.
Smart Images

Figure CN121639288A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of Internet online advertising, and particularly relates to a marketing traffic dynamic regulation method based on real-time data feedback. BACKGROUND
[0002] In online marketing scenarios such as real-time bidding advertising, advertisers usually expect that a fixed budget can be effectively consumed within a limited delivery period and converted into as many high-quality exposures and conversions as possible after setting the budget; this demand makes the advertising delivery system face the trade-off between budget consumption speed and traffic acquisition quality: the system needs to dynamically decide the bidding level for each exposure request in limited bidding opportunities; if the consumption speed is excessively pursued, the value of the traffic may be generally acquired at a high cost in the early delivery period, and the subsequent budget space may be squeezed; if the traffic quality is excessively emphasized, the high-value traffic may be missed due to conservative bidding, and the budget consumption may be slow, which makes it difficult to complete the exposure or conversion target; with the continuous change of traffic value and bidding competition intensity in real time, this contradiction is particularly prominent.
[0003] In order to cope with the above challenges, some schemes propose a dynamic bidding strategy based on deep reinforcement learning; such methods usually take the budget state, traffic features and bidding environment information as state input, learn the bidding decision strategy through the continuous interaction of the agent and the real-time bidding environment, and aim to balance the smoothness of budget consumption and the optimization of overall value income in the entire delivery period; compared with the fixed bidding rule, this scheme can adaptively adjust the bidding strength according to the real-time feedback, and to a certain extent, improve the budget utilization efficiency and delivery effect.
[0004] However, the existing reinforcement learning bidding strategy still has limitations in dealing with the above budget-quality contradiction; on the one hand, many strategies tend to use a relatively unified optimization target or a simple stage-by-stage preset in the entire budget consumption period, and fail to carefully depict the differentiated delivery focus corresponding to different time periods and different remaining budget levels in the budget consumption process; on the other hand, the system lacks explicit modeling of the dynamic correlation between the budget consumption stage and the real-time traffic value-competition situation, and it is difficult to adjust the strategy in time and accurately according to the budget progress and environmental changes.
[0005] In order to make up for the above shortcomings, the existing scheme introduces a budget smoothing control algorithm, or divides the daily delivery time into several fixed periods and sets sub-targets for each period; the budget smoothing control can prevent the budget from being consumed too early to a certain extent, but its constraints may weaken the system's real-time capture ability when discovering high-value traffic clusters; the time period control relies on the pre-set time division and target configuration, and the division is usually static, which is difficult to respond to the nonlinear budget consumption process and value opportunity distribution caused by market competition changes in a single advertising activity.
[0006] Therefore, the prior art has not fully solved the problem of collaborative modeling between the phased dynamic characteristics of budget consumption process and the real-time game of traffic value, and it is necessary to provide a marketing traffic dynamic regulation method capable of adaptively depicting budget consumption stages and linking bidding strategies based on real-time data feedback. SUMMARY
[0007] In view of the above existing problems, the present application is proposed.
[0008] The present application provides a marketing traffic dynamic regulation method based on real-time data feedback to solve the problem that the existing RTB mostly adopts unified or static bidding and budget control, which is difficult to balance the real-time capture of budget smooth consumption and high-value traffic.
[0009] To solve the above technical problems, the present application provides the following technical solutions:
[0010] The present application provides a marketing traffic dynamic regulation method based on real-time data feedback, applied to a real-time bidding (RTB) advertising system, comprising:
[0011] Step S1, obtaining budget consumption state data for a target advertising activity and real-time bidding environment data, wherein the budget consumption state data at least includes initial budget, remaining budget, consumed budget proportion, and time progress of delivery, and the real-time bidding environment data at least includes traffic characteristics of a current bidding request and historical bidding records related to the traffic characteristics;
[0012] Step S2, inputting the budget consumption state data into a budget consumption stage perception module to determine the budget consumption stage of the current advertising activity, wherein the budget consumption stage at least includes early consumption stage, middle consumption stage, and late consumption stage;
[0013] Step S3, according to the budget consumption stage, selecting a target value competition preference function bound to the current budget consumption stage from a plurality of preset value competition preference functions, wherein the value competition preference function is used to represent the trade-off preference between the estimated traffic value index and the real-time competition intensity index;
[0014] Step S4, based on the real-time bidding environment data, calculating the estimated traffic value index and the real-time competition intensity index of the current bidding request;
[0015] Step S5, inputting the estimated traffic value index and the real-time competition intensity index into the target value competition preference function to obtain the strategy bidding value corresponding to the current bidding request;
[0016] Step S6, based on the strategy bidding value, initiating a bidding request in the RTB advertising system.
[0017] As a preferred scheme of the marketing flow dynamic regulation method based on real-time data feedback, the value competition preference function has different configurations in different budget consumption stages.
[0018] In the early consumption stage, high competition intensity is preferentially inhibited to control bidding cost, in the middle consumption stage, the estimated flow value index and the real-time competition intensity index are balanced, and in the late consumption stage, the weight of the estimated flow value index is increased to preferentially obtain high value flow.
[0019] As a preferred scheme of the marketing flow dynamic regulation method based on real-time data feedback, the budget consumption stage perception module divides and switches the budget consumption stages through the relationship between the budget consumption ratio and the time progress ratio, at least including: calculating a deviation index according to the consumed budget proportion and the delivery time progress, and comparing the deviation index with a plurality of stage thresholds, to output one of the early consumption stage, the middle consumption stage or the late consumption stage.
[0020] As a preferred scheme of the marketing flow dynamic regulation method based on real-time data feedback, in step S4, the estimated flow value index is obtained by calculating a trained estimation model based on historical conversion data, the input of the estimation model includes at least one of user features, media position features and advertisement material features, and the output is a numerical index representing the expected revenue per exposure.
[0021] As a preferred scheme of the marketing flow dynamic regulation method based on real-time data feedback, in step S4, the real-time competition intensity index is obtained by constructing and dynamically maintaining a value competition heat map, the value competition heat map is used to depict the corresponding relationship between historical conversion value and recent winning price level under different flow feature combinations, and the corresponding real-time competition intensity index is obtained from the value competition heat map according to the flow features of the current bidding request, wherein the mapping relationship and interpolation method of the value competition heat map are determined in advance when constructing the value competition heat map.
[0022] As a preferred scheme of the marketing flow dynamic regulation method based on real-time data feedback, the dynamic maintenance of the value competition heat map is event-driven, and the triggering events include:
[0023] When a preset update time period is reached, or the fluctuation amplitude of the winning price level under a specific flow feature combination exceeds a stability threshold, the historical conversion value and the recent winning price level of the corresponding region are updated.
[0024] As a preferred scheme of the marketing traffic dynamic regulation method based on real-time data feedback, the method further comprises a pulse bidding management step of the traffic cluster, comprising:
[0025] In step S71, the continuously-arriving bidding requests are divided into one or more traffic clusters according to the similarity of traffic features.
[0026] In step S72, a first preset time window after the traffic cluster is identified is marked as a pulse bidding period of the traffic cluster, and in the pulse bidding period, the bidding requests belonging to the traffic cluster are bid based on the strategy bid value calculated according to the value competition preference function bound to the current budget consumption stage.
[0027] In step S73, the input-output ratio feedback index of the traffic cluster is counted in the pulse bidding period, and whether to continue to perform bidding on the subsequent bidding requests of the traffic cluster is judged according to the input-output ratio feedback index after the pulse bidding period ends.
[0028] As a preferred scheme of the marketing traffic dynamic regulation method based on real-time data feedback, in step S71, the continuously-arriving bidding requests are divided into the traffic clusters according to the similarity of traffic features, comprising:
[0029] The traffic features of each bidding request are mapped into a feature vector in an embedding space, the distance between the feature vectors of adjacent bidding requests is calculated, and the distance is compared with a dynamic distance threshold, so as to determine whether the bidding request is included in an existing traffic cluster or a new traffic cluster is started.
[0030] As a preferred scheme of the marketing traffic dynamic regulation method based on real-time data feedback, the input-output ratio feedback index is the ratio of the conversion value accumulated in the pulse bidding period to the cumulative bid cost of the traffic cluster, and when the input-output ratio feedback index is lower than a preset threshold, the bidding on the subsequent bidding requests of the traffic cluster is terminated, and the internal weight parameter in the value competition preference function bound to the current budget consumption stage is dynamically fine-tuned based on the input-output ratio feedback indexes of a plurality of traffic clusters.
[0031] As a preferred scheme of the marketing traffic dynamic regulation method based on real-time data feedback, in step S5, on the basis of calculating the strategy bid value according to the target value competition preference function, an exploration strategy is executed with a preset probability, and the exploration strategy comprises introducing a random disturbance on the basis of the strategy bid value or sampling an exploratory bid value from a preset distribution.
[0032] The application has the beneficial effects that: the application introduces a cooperative mechanism of budget consumption phase perception and value competition preference function binding in the real-time bidding advertising system, so that an explicit and dynamic linkage relationship is established between the budget progress and the traffic value and the competition intensity: on the one hand, the budget consumption phase perception module no longer simply relies on a fixed time slice, but comprehensively considers the consumed budget proportion and the time progress to form a stable judgment of fast running and slow running, so that different bidding preferences are automatically selected in the early, middle and late stages of consumption, thereby avoiding the problem of excessive money burning in the early stage or difficulty in spending the budget in the late stage caused by the use of a single strategy in the whole period in the traditional scheme; on the other hand, by constructing and maintaining a value competition heat map, the application encodes the historical conversion value and the recent winning price in the feature space at the same time, converts the abstract traffic quality and competition intensity into continuous indexes that can be queried and interpolated, and supports the real-time response of bidding decision to market environment changes. In addition, the application introduces pulse bidding management of traffic clusters at the local level, forms a short-term exploration window for similar requests, uses the input-output ratio feedback as the standard for whether to continue to invest, can concentrate the budget and amplify exposure when a high-value traffic cluster is found, and can stop loss and recover the budget in time for low-efficiency clusters, thereby improving the allocation efficiency of the budget among different traffic groups; at the same time, the internal weights of the value competition preference function in each stage are dynamically adjusted based on the feedback of multiple traffic clusters, so that the strategy can continuously adapt and evolve with the bidding process, avoiding the invalidation of fixed weights under environmental drift; in combination with the exploratory bidding mechanism based on the strategy value, the application retains the ability to discover new traffic combinations while ensuring overall revenue, and alleviates the local optimal dilemma of repeatedly turning around in the vicinity of existing high-quality traffic.
[0033] In summary, compared with the existing scheme relying on a single optimization target or static budget smoothing control, the application has comprehensive improvements in the controllability of the budget consumption curve, the high-value traffic capture ability and the robustness to market fluctuations. BRIEF DESCRIPTION OF DRAWINGS
[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and should not be regarded as limiting the scope of the present application.
[0035] Figure 1 The flowchart of the marketing traffic dynamic regulation method based on real-time data feedback in the embodiments. DETAILED DESCRIPTION
[0036] In order to make the purpose, technical solutions and advantages of the present application more clear, the following will further describe the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and should not be regarded as limiting the present application.
[0037] All terms used herein, including technical and scientific terms, have the meanings as commonly understood by one of ordinary skill in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning that is consistent with the context of the specification, and should not be interpreted in an idealized or overly formal way.
[0038] For example, the terms "first", "second", and the like as used herein are used only to distinguish one object from another, and do not necessarily indicate a particular order or sequence, nor are they used to indicate or imply relative importance of the objects so qualified.
[0039] The present application provides a marketing traffic dynamic regulation method based on real-time data feedback, applied to a real-time bidding (RTB) advertising system, which combines Figure 1 As shown in the figure, comprising:
[0040] Step S1, obtaining budget consumption state data and real-time bidding environment data for a target advertising campaign, wherein the budget consumption state data at least includes initial budget, remaining budget, consumed budget proportion, and time progress, and the real-time bidding environment data at least includes traffic characteristics of a current bidding request and historical bidding records related to the traffic characteristics;
[0041] Step S2, inputting the budget consumption state data into a budget consumption stage perception module to determine the budget consumption stage of the current advertising campaign, wherein the budget consumption stage at least includes the initial consumption stage, the middle consumption stage, and the final consumption stage;
[0042] Step S3, according to the budget consumption stage, selecting a target value competition preference function bound to the current budget consumption stage from a plurality of preset value competition preference functions, wherein the value competition preference function is used to represent the trade-off preference between the estimated traffic value index and the real-time competition intensity index;
[0043] Step S4, based on the real-time bidding environment data, calculating the estimated traffic value index and the real-time competition intensity index of the current bidding request;
[0044] Step S5, inputting the estimated traffic value index and the real-time competition intensity index into the target value competition preference function to obtain a strategy output value corresponding to the current bidding request;
[0045] Step S6, initiating a bidding request in the RTB advertising system based on the strategy output value;
[0046] In one embodiment, the value competition preference function has different configurations in different budget consumption stages:
[0047] In the early consumption stage, the high competition intensity is preferentially inhibited to control the bidding cost, in the middle consumption stage, the estimated traffic value index and the real-time competition intensity index are balanced, and in the late consumption stage, the weight of the estimated traffic value index is increased to preferentially obtain high-value traffic.
[0048] In one embodiment, the budget consumption stage perception module divides and switches the budget consumption stage through the relationship between the budget consumption ratio and the time progress ratio, at least including: calculating a deviation index according to the consumed budget proportion and the delivery time progress, and comparing the deviation index with a plurality of stage thresholds, to output one of the early consumption stage, the middle consumption stage or the late consumption stage.
[0049] The budget consumption stage division and switching judgment mode is as follows:
[0050] Step S21, for each time Calculate the consumed budget proportion, the instant deviation, the smooth deviation and the fluctuation intensity:
[0051] , , ,
[0052] , ,
[0053] Among them, is the consumed budget proportion at the time, is the initial budget amount, is the remaining budget amount at the time, is the instant deviation at the time, is the delivery time progress at the time, is the smooth deviation at the time, is the smooth deviation at the time, is the exponential weighted variance of the deviation at the time, is the exponential weighted variance at the time, is the fluctuation intensity of the deviation at the time, is the deviation smoothing coefficient, is the variance smoothing coefficient, is the numerical stability term.
[0054] Step S22, first calculate the budget runway correction amount, and then generate the entering / leaving threshold values (including hysteresis) of the three stages:
[0055] ,
[0056] ,
[0057] ,
[0058] wherein, is the runway correction amount at the moment, is the upper / lower bound truncation function, is the runway correction coefficient, and are the lower / upper correction limits, is the deviation threshold for the initial entry consumption period, is the deviation threshold for the initial exit consumption period, is the deviation threshold for the final entry consumption period, is the deviation threshold for the final exit consumption period, is the fluctuation-to-threshold ratio coefficient, is the threshold baseline, is the hysteresis width, with superscript , denote the entry and exit threshold types, respectively;
[0059] Step S23, set the minimum dwell time switching enablement and update the phase label with a hysteresis segmented rule:
[0060]
[0061]
[0062] wherein, is the moment switching enablement indication, is the indicator function, is the last phase switching moment, is the minimum dwell time window, is the budget consumption phase label at the moment, is the previous moment phase label, denotes the initial consumption period, denotes the middle consumption period, denotes the final consumption period;
[0063] Initialize the acceptable , ; when a budget or time progress jump is detected, skip the current update and inherit ; wherein, is the initial smooth deviation degree, is the initial immediate deviation degree, For initial exponential weighted variance;
[0064] The offline grid selection can be based on historical logs, The system latency and traffic fluctuation can be combined to set a robust interval;
[0065] Specifically, the above embodiment takes the difference between the consumed budget proportion and the time progress as the deviation, and adopts exponential smoothing and variance recursion to obtain robust trend and fluctuation, which are used to depict the current progress deviation and its uncertainty; On this basis, the stage boundary is expanded from a fixed threshold to a dynamic threshold that is self-adaptive to the fluctuation and budget runway, and the hysteresis width of entering / leaving is introduced to reduce the back-and-forth switching near the critical point; Then the enabling rule with minimum residence time is given, forming a finite state machine type stage update logic, which takes into account the response speed and stability; The binding relationship of this link and the subsequent value competition preference function naturally connects: the stage is updated as soon as possible to drive internal weight selection or fine-tuning, realize the segmented regulation of bidding strategy under different budget progress and market fluctuation, and improve the controllability of cost and the ability to capture high-value traffic of the bid;
[0066] In this embodiment, the budget consumption state data is automatically aggregated by the billing and monitoring log of the advertisement delivery system according to a fixed scheduling period. At the end of each scheduling period, the consumed budget proportion and time progress are calculated based on the initial budget, the latest remaining budget, and the delivery start and end time of the current activity, so that the deviation and related statistics are always updated with the same time granularity. The scheduling period can be defined according to natural time or according to bidding request batches. By default, it is updated in rounds of 5 to 10 seconds, which can be shortened to the order of seconds in scenarios with particularly dense traffic, and can be relaxed to the order of 30 seconds in long tail activities or at night when traffic is less, to ensure that the stage label can respond to budget deviation in time and will not be frequently jittered due to fluctuation noise. Further, the deviation smoothing coefficient used in exponential smoothing can be preferentially taken as a value close to one, for example 0.09-0.99, to emphasize the change in budget consumption rhythm in the last few minutes; the smoothing coefficient in variance recursion can be slightly smaller than the deviation smoothing coefficient, for example 0.8-0.95, to more quickly raise the fluctuation intensity estimate when the deviation suddenly increases; the numerical stability term can be set to a positive number in the order of one ten-thousandth to one thousandth, to avoid division by zero or underflow when the sample is extremely small or the deviation is close to zero. Specifically, the deviation threshold and its baseline and hysteresis width for entering and exiting different stages can be obtained by offline playback of the budget consumption curve of historical healthy activities: select a batch of advertisement activities with good matching of budget consumption and time progress, and calculate the standard deviation and quantile range of the deviation distribution, set the entry threshold slightly higher than the normal fluctuation level, and lower the exit threshold by a fixed width to form a threshold pair with obvious hysteresis, so as to avoid frequent switching of stage labels when approaching the boundary. Optionally, the budget deviation correction coefficient can be set separately according to different budget size segments, for example, smaller correction intensity is given to activities with larger budget amount, to prevent short-term deviation from causing frequent stage adjustment, while activities with smaller budget and shorter period can have higher correction intensity to return to the target consumption track more quickly. Similarly, the minimum residence time window can be measured by natural time or by the number of bidding requests, and is usually set to tens of seconds to a few minutes, or a few hundred to a few thousand requests, to ensure that the stage is maintained for at least one complete feedback period before allowing entry into the next switching judgment. Optionally, when detecting scenarios such as budget reset, activity suspension and recovery, or cross-natural-day segmentation that cause discontinuous jumps in budget or time progress, the system can use the estimated value of the current deviation and its variance from the last scheduling period, and skip the stage update and switching judgment, while marking the time with a reset flag to avoid misleading the stage recognition due to single abnormal writing; if billing or remaining budget data is missing in some scheduling periods, the stage label can be temporarily frozen, and the deviation trajectory can be recalculated through playback after data recovery, to ensure that the budget consumption stage perception link still has realizability and robustness in abnormal scenarios.
[0067] In one embodiment, in step S4, the estimated traffic value index is calculated by an estimation model trained based on historical conversion data. The input of the estimation model includes at least one user feature, media position feature, and advertising material feature, and the output is a numerical index representing the expected revenue per unit exposure.
[0068] In one embodiment, in step S4, the real-time competition intensity index is obtained by constructing and dynamically maintaining a value competition heatmap. The value competition heatmap is used to depict the correspondence between historical conversion value and recent winning price level under different traffic feature combinations. The corresponding real-time competition intensity index is obtained from the value competition heatmap based on the traffic features of the current bidding request. The mapping relationship and interpolation method of the value competition heatmap are predetermined when constructing the value competition heatmap.
[0069] The mapping relationship and interpolation method of the value competition heatmap are as follows:
[0070] Step S41: Embed the traffic features of the current bidding request into a unified interval and then map them onto a discrete grid to form a deterministic index of feature combination → grid points:
[0071] ,
[0072] in, express Moment The dimensional flow feature vector has the following components: , Indicates the first Monotonic feature transformation of dimension, This indicates the first estimate based on historical samples. Marginal cumulative distribution function, Indicates will Mapped to the Grid point index of a 3D grid, Indicates the first The number of bins in the dimension, Indicates the first Wei Zai The first The upper boundary of each box (satisfies) ), Indicates the time index. Indicates a dimension index;
[0073] Define the center coordinates for subsequent interpolation:
[0074] ,
[0075] ,
[0076] in, Indicates the first Veth center position of the box, center coordinate vector of the multi-dimensional grid point, subscript denotes the multi-dimensional grid point index;
[0077] Step S42, for each grid point, accumulates the historical conversion value and the recent winning price level, and forms a pair of indicators:
[0078] , , , ,
[0079] , , ,
[0080] wherein, denotes the time decay weight of the sample , denotes the current time, denotes the occurrence time of the sample , denotes the time half-life (decay) parameter, denotes the sample set mapped to the grid point , denotes the conversion value of the sample , denotes the winning price level of the sample , and and respectively denote the weighted cumulative value and price sum on the grid point , denotes the weighted sample amount, and respectively denote the value and price estimates of the grid point , denotes the binary correspondence relationship vector, and superscript is used to distinguish the value and the price; Step S43, when querying , the barycentric interpolation is performed on the vertex set of the hypercube where it is located to obtain a continuous estimate: , ,
[0081]
[0082] , ,
[0083] ,
[0084] , ,
[0085] in, express In the Normalized position within the box. express The set of vertices of the hypercube. Represents vertices linear weights, Indicates an indicator function, and They represent continuous estimates of value and price obtained through linear interpolation, respectively, with subscripts... Represents the vertex grid index, superscript Indicates linear interpolation;
[0086] Step S44: To mitigate the fluctuations caused by sparse grid points, kernel regression smoothing at the grid center is introduced and fused with linear interpolation according to confidence level.
[0087] , ,
[0088] , ,
[0089] , , ,
[0090] in, Represents the bandwidth matrix Gaussian kernel function, The kernel distance vector, It is a diagonal bandwidth matrix. For the first Core bandwidth, Indicates the confidence weight of the grid points. For confidence temperature parameters, Indicates embedded coordinates, These represent the value and price estimates obtained through kernel regression smoothing, respectively. express The grid point, Represents the fusion coefficient. This represents the minimum weighted sample size threshold using linear interpolation. These represent continuous estimates of the merged value and price, respectively, with superscript indicating the final value. Indicates nuclear regression, superscript Indicates confidence;
[0091] Mesh Boundary It can be set according to equal frequency or equal width. When affected by the discreteness of category features, the discrete combinations can be directly encoded into a one-dimensional index and then Cartesian productd with the remaining continuous dimensions to maintain a unified mapping process.
[0092] Specifically, the above implementation completes the value competition heatmap from data to function: by transforming the marginal distribution, features of each dimension are mapped to a unified interval, facilitating the setting of a regularized multidimensional grid. After establishing a deterministic index of feature combination-grid points, historical conversion value and recent winning price are accumulated for each grid point with time decay weights to obtain queryable pairwise statistics. To improve compatibility in sparse and dense regions, two types of continuous methods are introduced: one is centroid interpolation based on the vertices of the hypercube, which has the advantages of locality and low computational cost; the other is kernel regression based on the grid center, which uses the bandwidth matrix to adaptively smooth each dimension to suppress outliers and local holes. The two are adaptively switched and weighted through a fusion coefficient driven by the sample size. When the sample is sufficient, linear interpolation is relied upon more to maintain structural sensitivity, while kernel regression is switched to obtain robust estimation when the sample is sparse. This mapping and interpolation framework provides a continuous, differentiable, and controllable basic representation for the subsequent acquisition of real-time competition intensity and strategy trade-offs.
[0093] Furthermore, historical conversion value and winning price samples can be uniformly provided by the advertising platform's exposure, bidding receipts, and conversion attribution logs. These samples are bucketed according to traffic characteristics and written into the value competition heatmap storage structure in offline batch processing or near real-time streaming computing. This allows the online query stage to obtain statistical results simply by indexing the corresponding grid point or its neighborhood based on the characteristics of the current bidding request, without needing to access the original detailed logs. In actual deployment, several features highly correlated with traffic value and competition intensity can be prioritized as heatmap dimensions, such as media placement type, ad placement level, region, terminal device type, main interest categories, and coarse-grained time periods, keeping the total number of dimensions between 5 and 20 to achieve a balance between expressive power and storage / computing costs. The number of bins for each dimension can be set between 8 and 64 based on historical traffic distribution. Dimensions with more concentrated traffic use fewer bins, while dimensions with a long tail of distribution use more granular bins to avoid a large number of samples falling into a very small number of grid points. Specifically, the half-life parameter in the time decay weight can be set between half an hour and twenty-four hours according to the business rhythm and the requirement for freshness. A shorter half-life emphasizes recent samples and is suitable for activities with rapidly changing competitive environments, while a longer half-life is smoother and suitable for stable, long-term activities. The bandwidth of kernel regression can be set to a fraction of the empirical standard deviation of each feature in the standardized space, so that the kernel function has obvious weights near similar features, while the weights naturally decay on combinations of features with large differences. The confidence temperature parameter and the sample size fusion threshold can be selected by selecting typical intervals through offline simulation experiments of estimation errors under different sample sizes. For example, when the weighted sample size of a single grid point reaches tens to hundreds, it is considered that the sample is sufficient, and the weight of linear interpolation is gradually increased within this range. When the sample size is lower than this range, more reliance is placed on kernel smoothing results. Optionally, to reduce the computational overhead of kernel regression, in actual calculations, kernel weights can be accumulated only on grid points whose distance from the query point is less than a few times the bandwidth, while grid points that are farther away and have extremely small weights can be directly ignored to form an approximately sparse kernel. If a grid point and its neighborhood have almost no effective samples in the historical period, the system can degenerate into using a coarser-grained feature combination or the global average at the activity level as the default conversion value and winning price estimate, and automatically switch back to heatmap-based estimation after sufficient log samples have been accumulated. Similarly, when a feature value falls within a range not covered during the training phase, the feature value can be truncated or projected to the nearest boundary box to ensure that the index is always within the multidimensional grid. If incremental samples are missing due to log collection or attribution link anomalies within a certain time period, the system can temporarily maintain the heatmap estimation and mark that time period as expired. After the preset update time period arrives, the corresponding area can be re-estimated using the supplemented logs to ensure the feasibility of the value competition heatmap in engineering and the controllability of data quality.
[0094] In one embodiment, the dynamic maintenance of the value competition heatmap is event-driven, and the triggering events include:
[0095] When the preset update time period is reached, or when the fluctuation of the winning price level under a specific combination of traffic characteristics exceeds the stability threshold, the historical conversion value and recent winning price level of the corresponding region will be updated.
[0096] In one embodiment, the method further includes a pulsed bidding management step for traffic clusters, including:
[0097] Step S71: Divide consecutively arriving bidding requests into one or more traffic clusters based on the similarity of traffic characteristics;
[0098] Step S72: Mark the first preset time window after each traffic cluster is identified as the pulse bidding period of the traffic cluster. During the pulse bidding period, bids are made on all bidding requests belonging to the traffic cluster based on the strategy bid value calculated by the value competition preference function bound to the current budget consumption stage.
[0099] Step S73: During the pulse bidding period, the input-output ratio feedback index of the traffic cluster is calculated, and after the pulse bidding period ends, it is determined whether to continue bidding on subsequent bidding requests for the traffic cluster based on the input-output ratio feedback index.
[0100] For example, traffic clusters can be maintained online in chronological order during the ad request access chain. Each new bidding request generates a corresponding embedding vector from the existing feature encoding component in the system. This vector is then compared to the embedding vectors of the most recent requests preceding it. Based on whether the distance is below a dynamic distance threshold and whether the arrival time interval meets the constraint of continuous arrival, a decision is made whether to classify the request into the current active cluster or to create a new traffic cluster. A pulse bidding period can be initiated immediately when a traffic cluster is first identified. Its duration is defined by a preset time window or a preset upper limit on the number of requests. For example, a pulse period is typically defined as 3-10 minutes or 200 to 1000 bidding requests. On high-traffic media, priority can be given to controlling the number of requests, while on low-traffic media, priority can be given to controlling the duration, thus ensuring that each cluster has sufficient samples within the pulse period for evaluating the return on investment. In terms of numerical settings, the dynamic distance threshold can be determined by offline statistical analysis of the distance distribution between homogeneous and heterogeneous requests in the embedding space. A level near a high similarity quantile is preferred, ensuring that only requests with truly similar features are clustered together. The return on investment (ROI) during the pulse bidding period can be accumulated along the cluster dimension based on near real-time billing and conversion attribution links. Costs are the actual charges corresponding to the requests in that cluster, and revenues are the conversion value successfully attributable to these requests. This is refreshed on a rolling basis at the minute or request batch level to provide sufficiently timely feedback while ensuring acceptable conversion feedback latency. Furthermore, the ROI threshold for determining whether to retain a traffic cluster can be set slightly higher than the global average or the target return level of the campaign, referencing the average delivery efficiency of similar advertising campaigns in historical campaigns. This prioritizes retaining clusters with higher overall efficiency and promptly terminates significantly inefficient clusters. Optionally, when the actual number of bidding requests accumulated by a cluster during the pulse bidding period is far below the preset lower limit, or when not enough attributable conversions have been generated during the entire pulse period, the observation window for that cluster can be extended, or the weight of its return on investment (ROI) estimation can be reduced during decision-making to avoid misjudgment due to insufficient samples. In scenarios where there is a long delay in conversion return, the system can also make a preliminary judgment based on some of the returned conversions and supplement and correct the ROI of the cluster in subsequent cycles. Similarly, when a billing log or attribution log is detected to be temporarily unavailable, the traffic cluster within the corresponding time period can be marked as having incomplete data, and the final decision on whether to continue bidding or terminate bidding can be postponed until the data is recovered and complete, thereby improving the feasibility and stability of pulse bidding management under abnormal operating conditions.
[0101] In one embodiment, step S71, dividing consecutively arriving bidding requests into traffic clusters based on the similarity of traffic characteristics, includes:
[0102] The traffic characteristics of each bidding request are mapped to feature vectors in the embedding space. The distance between the feature vectors of adjacent bidding requests is calculated and compared with a dynamic distance threshold to determine whether to classify the bidding request into an existing traffic cluster or to open a new traffic cluster.
[0103] In one embodiment, the input-output ratio feedback index is the ratio of the cumulative conversion value obtained by the traffic cluster to the cumulative bidding cost during the pulse bidding period. When the input-output ratio feedback index is lower than a preset threshold, the bidding for subsequent bidding requests of the traffic cluster is terminated, and the internal weight parameters in the value competition preference function bound to the current budget consumption stage are dynamically fine-tuned based on the input-output ratio feedback indices of multiple traffic clusters.
[0104] In one embodiment, in step S5, based on the calculated value of the strategy according to the target value competition preference function, an exploration strategy is executed with a preset probability. The exploration strategy includes: introducing random perturbation or sampling exploratory value from a preset distribution based on the value of the strategy, so as to improve the adaptability of the strategy to unknown traffic feature combinations.
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0106] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are meant to be within the scope of this application and form different embodiments. For example, all the embodiments above can be used in any combination. The information disclosed in this background section is intended only to enhance the understanding of the general background of this application and should not be construed as an admission or in any way implying that such information constitutes prior art known to those skilled in the art.
Claims
1. A marketing traffic dynamic regulation method based on real-time data feedback, applied to a real-time bidding (RTB) advertising system, characterized in that, The method comprises the following steps: Step S1, obtaining budget consumption state data and real-time bidding environment data for a target advertising campaign, wherein the budget consumption state data at least includes an initial budget, a remaining budget, a consumed budget proportion, and a delivery time progress, and the real-time bidding environment data at least includes traffic characteristics of a current bidding request and historical bidding records related to the traffic characteristics; Step S2, inputting the budget consumption state data into a budget consumption phase perception module to determine a budget consumption phase in which the current advertising campaign is located, wherein the budget consumption phase at least includes an early consumption phase, a middle consumption phase, and a late consumption phase; Step S3, selecting a target value competition preference function bound to the current budget consumption phase from a plurality of preset value competition preference functions according to the budget consumption phase, wherein the value competition preference function is used to represent a trade-off preference between an estimated traffic value index and a real-time competition intensity index; Step S4, calculating the estimated traffic value index and the real-time competition intensity index of the current bidding request based on the real-time bidding environment data; Step S5, inputting the estimated traffic value index and the real-time competition intensity index into the target value competition preference function to obtain a strategy output value corresponding to the current bidding request; Step S6, initiating a bidding request in the RTB advertising system based on the strategy output value.
2. The marketing flow dynamic regulation method based on real-time data feedback according to claim 1, characterized in that, The value competition preference function has different configurations in different budget consumption phases: In the early consumption phase, high competition intensity is preferentially suppressed to control bidding cost, in the middle consumption phase, the estimated traffic value index and the real-time competition intensity index are balanced, and in the late consumption phase, the weight of the estimated traffic value index is increased to preferentially obtain high-value traffic.
3. The marketing flow dynamic regulation method based on real-time data feedback according to claim 1, characterized in that, The budget consumption phase perception module divides and switches the budget consumption phase through the relationship between the budget consumption proportion and the time progress proportion, and at least includes: calculating a deviation index according to the consumed budget proportion and the delivery time progress, and comparing the deviation index with a plurality of phase thresholds to output one of the early consumption phase, the middle consumption phase, or the late consumption phase.
4. The marketing flow dynamic regulation method based on real-time data feedback according to claim 1, characterized in that, In step S4, the estimated traffic value index is obtained by calculating a prediction model trained based on historical conversion data, wherein the input of the prediction model includes at least one of user characteristics, media position characteristics, and advertising material characteristics, and the output is a numerical index representing expected revenue per exposure.
5. The marketing flow dynamic regulation method based on real-time data feedback according to claim 1, characterized in that, In step S4, the real-time competition intensity index is obtained by constructing and dynamically maintaining a value competition heat map, wherein the value competition heat map is used to depict the corresponding relationship between historical conversion value and recent winning price level under different traffic characteristic combinations, and the corresponding real-time competition intensity index is obtained from the value competition heat map according to the traffic characteristics of the current bidding request, wherein the mapping relationship and interpolation method of the value competition heat map are determined in advance when constructing the value competition heat map.
6. The marketing traffic dynamic regulation method based on real-time data feedback according to claim 5, characterized in that, The dynamic maintenance of the value competition heat map is event-driven, and the triggering events include: When a preset update time period is reached, or when the fluctuation amplitude of the winning price level under a specific traffic characteristic combination exceeds a stability threshold, the historical conversion value and the recent winning price level of the corresponding region are updated.
7. The marketing flow dynamic regulation method based on real-time data feedback according to claim 1, characterized in that, The method further comprises a pulsing bidding management step of the traffic cluster, comprising: Step S71, dividing the continuously arrived bidding requests into one or more traffic clusters according to the similarity of traffic features; Step S72, marking the first preset time window after the traffic cluster is identified as a pulsing bidding period of the traffic cluster, and in the pulsing bidding period, bidding of the bidding requests belonging to the traffic cluster is performed based on the strategy bid value calculated according to the value competition preference function bound with the current budget consumption stage; Step S73, in the pulsing bidding period, the input-output ratio feedback index of the traffic cluster is counted, and whether to continue to perform bidding on the subsequent bidding requests of the traffic cluster is judged according to the input-output ratio feedback index after the pulsing bidding period ends.
8. The marketing traffic dynamic regulation method based on real-time data feedback according to claim 7, characterized in that, In step S71, the continuously arrived bidding requests are divided into the traffic cluster according to the similarity of traffic features, comprising: The traffic features of each bidding request are mapped into a feature vector in an embedding space, the distance between the feature vectors of adjacent bidding requests is calculated, and the distance is compared with a dynamic distance threshold, which is used to determine whether the bidding request is included in an existing traffic cluster or a new traffic cluster is started.
9. The marketing flow dynamic regulation method based on real-time data feedback according to claim 7, characterized in that, The input-output ratio feedback index is the ratio of the conversion value accumulated in the pulsing bidding period to the cumulative bid cost of the traffic cluster, and when the input-output ratio feedback index is lower than a preset threshold, the bidding on the subsequent bidding requests of the traffic cluster is terminated, and the internal weight parameter in the value competition preference function bound with the current budget consumption stage is dynamically fine-tuned based on the input-output ratio feedback indexes of multiple traffic clusters.
10. The marketing flow dynamic regulation method based on real-time data feedback according to claim 1, characterized in that, In step S5, on the basis of calculating the strategy bid value according to the target value competition preference function, an exploration strategy is executed with a preset probability, and the exploration strategy comprises introducing a random disturbance on the basis of the strategy bid value or sampling an exploratory bid value from a preset distribution.