A DMP data integrated intelligent management method based on dynamic regulation

By integrating real-time data management and dynamic control, the problems of model lag and rule rigidity in DMP under highly dynamic business scenarios have been solved, enabling real-time updates of user tags and efficient allocation of resources, thereby improving the accuracy of ad delivery and the stability of the system.

CN122115038APending Publication Date: 2026-05-29BEIJING GREY INNOVATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING GREY INNOVATION TECHNOLOGY CO LTD
Filing Date
2026-02-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing DMPs suffer from lagging model updates, rigid rule configurations, and insufficient data integration in highly dynamic business scenarios, leading to delayed user tag updates, inaccurate targeting, and wasted resources.

Method used

By ingesting multi-source user behavior data streams in real time, an integrated user behavior data view is formed. The advertising budget is dynamically adjusted using a contextual multi-armed gambling machine algorithm. Combined with real-time feedback signals and online learning models, incremental updates and priority allocation of audience segmentation are achieved, solving the real-time and adaptive capabilities of the data management platform.

Benefits of technology

It significantly improves response speed, targeting accuracy, and system stability in high-concurrency scenarios, reduces latency and deviation caused by inconsistent labeling and cross-system fragmentation, and enhances the real-time performance of user profiles and the robustness of resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115038A_ABST
    Figure CN122115038A_ABST
Patent Text Reader

Abstract

The application discloses a kind of DMP data integration intelligent management method based on dynamic regulation, it is related to internet advertisement and big data processing technical field, the application is by constructing integrated user behavior data view in data management platform and with streaming mode continuously update, make from browsing, interaction and device etc. Multi-source data under unified standard is aggregated, significantly reduce the delay and deviation caused by label inconsistent and cross-system fragmentation;On the audience side, with dynamic clustering incremental maintenance subdivision boundary, combine context multi-arm bandit machine to candidate subdivision Real-time evaluation and selection, both use the effective experience of historical sedimentation, also retain necessary exploration to quickly capture emerging interest groups during promotion;On the learning and decision side, introduce online incremental update and robust feedback processing, can the instant and lag information of click and conversion is continuously assimilated, shorten the path from behavior change to strategy adjustment, reduce the lag effect of periodic retraining.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of internet advertising and big data processing technology, and in particular to an integrated intelligent management method for DMP data based on dynamic control. Background Technology

[0002] With the rapid development of e-commerce, large-scale promotional events (such as annual shopping festivals and themed promotional days) have become an important way for platforms to acquire traffic and conversions. During such peak events, users' browsing, searching, and purchasing behaviors fluctuate dramatically in a short period of time, and their interests and interaction patterns change frequently. To improve the effectiveness of advertising and marketing campaigns, existing technologies widely use data management platforms (DMPs) to aggregate and process multi-source user data, and achieve personalized advertising and marketing strategy optimization through audience targeting.

[0003] Most existing Data Management Platforms (DMPs) integrate multiple data sources such as browsing logs, device information, and transaction records. They utilize machine learning models to segment and tag users, and combine them with stream processing systems to ingest, clean, and aggregate data. Some solutions also introduce prediction models and rule engines based on offline training, and periodically adjust audience profiles and targeting strategies through batch processing or near real-time update mechanisms. At the same time, they integrate stream computing frameworks to reduce data processing latency, completing some audience segmentation within seconds or minutes, thereby alleviating the real-time pressure in high-concurrency environments to a certain extent.

[0004] While the above-mentioned technical solutions have achieved some success in personalized targeting, the following problems still exist in highly dynamic business scenarios such as promotional activities:

[0005] Model update lag: Existing machine learning models rely on periodic offline retraining, and the feature construction and model deployment process is long. When user behavior patterns change drastically in a short period of time, the model cannot reflect the latest interests and preferences in time, resulting in a lag in user tag updates, which in turn leads to inaccurate audience targeting.

[0006] Rigid rule configuration: Rule engines are mostly based on pre-set static thresholds and fixed logic. When new interest categories, new traffic sources or new interaction paths frequently appear during promotions, traditional rules are difficult to cover in a timely manner, requiring a lot of manual intervention and repeated parameter tuning, which can easily lead to unreasonable budget allocation and waste of marketing resources.

[0007] Insufficient data integration: During the multi-source data access process, user identifiers from different channels, devices, and account systems conflict and are inconsistent. The limited real-time entity parsing capability makes it difficult to merge the profiles of the same user in different systems in a timely manner, resulting in audience overlap or omissions. At the same time, the data links between offline data warehouses, real-time streaming engines, and advertising systems are relatively fragmented. The lack of a unified integrated management mechanism for tag definitions, update rhythm, and management strategies limits the system's responsiveness and accuracy in highly dynamic scenarios.

[0008] Therefore, how to achieve integrated management and intelligent processing of multi-source data while ensuring high concurrency processing capabilities, and improve the real-time performance and adaptability of user profiling and audience targeting strategies, has become an urgent problem for those skilled in the art. Summary of the Invention

[0009] In view of the aforementioned existing problems, the present invention is proposed.

[0010] This invention provides a DMP data integration intelligent management method based on dynamic control to solve the problems of sudden behavioral changes during promotional peaks, as well as the problems of lagging retraining, rigid rules, and unstable real-time targeting caused by ID conflicts in existing DMPs.

[0011] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0012] This invention provides a DMP data integrated intelligent management method based on dynamic control, comprising:

[0013] Step S1: Real-time capture of multi-source user behavior data streams, which include at least user browsing history, interaction events, and device information; preprocessing of the multi-source user behavior data streams and association by user identifier to form an integrated user behavior data view.

[0014] Step S2: Based on the integrated user behavior data view, the multiple audience segments are defined as multiple audience arms;

[0015] Step S3: Obtain real-time context features based on the current ad request to be delivered, input the real-time context features into the context multi-armed gambling machine algorithm, and calculate the expected return for each audience arm;

[0016] Step S4: Based on the expected return of each audience arm, a priority allocation mechanism is used to dynamically adjust the allocation of the advertising budget among multiple audience arms;

[0017] Step S5: Output the audience targeting strategy corresponding to the advertising budget allocation result to drive real-time ad delivery.

[0018] As a preferred embodiment of the DMP data integrated intelligent management method based on dynamic control described in this invention, the real-time context features include at least one of the following: user real-time browsing sequence, current timestamp, geographic location information, and terminal device type.

[0019] As a preferred embodiment of the DMP data integration intelligent management method based on dynamic control described in this invention, wherein: the context multi-armed gambling machine algorithm updates the expected return of each audience arm according to real-time feedback signals, and the real-time feedback signals include at least one of click events or conversion events;

[0020] The step of updating the expected returns of each audience arm based on real-time feedback signals includes:

[0021] Step S301: Establish an online attribution channel for exposure-click-conversion; the system will monitor the data in real time. Select Audience Arm The click and conversion results of the request are then collected to form a real-time feedback signal that can be used for subsequent updates. Unselected audience arms only undergo time decay processing of historical information.

[0022] Step S302: Maintain the posterior state of click-through rate and single conversion value estimate for each audience arm, which are used as two components of expected return, and refresh online when new feedback arrives;

[0023] Step S303: Treat the click as a Bernoulli event and introduce a time discount, applying it to the selected audience arm at time... Incremental update of posterior parameters:

[0024] ,

[0025] in, Indicates the audience arm At any moment Success count parameter, Indicates the audience arm At any moment The failure count parameter, and For the parameters of the previous time step, This represents the time discount factor for click-through rate updates. Indicates time Should audience arm be selected? The indicated quantity, This indicates whether a click occurred during the request. Indicates the audience arm index, Indicates a discrete-time index;

[0026] Step S304: When a conversion event occurs, update the single conversion value estimate of the audience arm in an online smooth manner. For samples that only have exposure or clicks but no conversion is observed, retain the existing estimate and perform a slight time decay. For conversions with lag attribution, update the corresponding historical moment when the event arrives.

[0027] Step S305: Multiplicatively combine the posterior mean of the click-through rate with the value of a single conversion to obtain the expected return for real-time decision-making.

[0028] ,

[0029] in, Indicates the audience arm At any moment Expected returns This represents the mean click-through rate based on the Beta posterior distribution obtained in step S303. Indicates the audience arm At any moment Estimate the value of a single conversion. Indicates the audience arm index, This represents a discrete-time index.

[0030] As a preferred embodiment of the DMP data integration intelligent management method based on dynamic control described in this invention, the context multi-armed gambling machine algorithm adopts an exploration and exploitation balance mechanism, constructs a probability distribution based on the expected returns of each audience arm, and selects the target audience arm according to the probability distribution.

[0031] As a preferred embodiment of the DMP data integration intelligent management method based on dynamic control described in this invention, the method further includes: updating the parameters of the context multi-armed gambling machine algorithm in real time through an online learning model, wherein the online learning model uses an incremental learning method to iteratively adjust the parameters.

[0032] As a preferred embodiment of the DMP data integration intelligent management method based on dynamic control described in this invention, the real-time ingestion of multi-source user behavior data streams is achieved through a distributed stream processing platform, which is used for distributed access, buffering, and computation of the multi-source user behavior data streams.

[0033] As a preferred embodiment of the DMP data integration intelligent management method based on dynamic control described in this invention, multiple audience segments are generated based on a dynamic clustering algorithm. The dynamic clustering algorithm adopts a density-based clustering method and incrementally updates the audience segments according to changes in user behavior characteristics.

[0034] As a preferred embodiment of the DMP data integration intelligent management method based on dynamic control described in this invention, the method includes: associating multi-source user behavior data streams by user identifier, including an integrated entity parsing algorithm, constructing a graph structure for user identifiers from different data sources, and merging nodes belonging to the same entity based on graph similarity calculation.

[0035] The steps for merging nodes belonging to the same entity based on graph similarity calculation include:

[0036] Step S308: Using the user identifier as a node, generate weighted undirected edges based on evidence of the same session, same device fingerprint, same account login, and same geographic-temporal co-occurrence, and extract the attribute set and adjacency relationship for each node;

[0037] Step S309: Measure the attribute overlap based on the attribute set. When numerical stabilization is required for the empty set, a small stabilization term is introduced, and the similarity is written into the node pair.

[0038] ,

[0039] in, Represents a node With nodes Attribute similarity, Represents a node The set of attributes, Represents a node The set of attributes Represents the cardinality of a set. Represents the numerically stable term. Indicates the node index. Indicates the node index;

[0040] Step S310: Perform time decay accumulation on co-occurrence events on the same screen, in the same session, or in a short time window to obtain the co-occurrence intensity updated over time;

[0041] ,

[0042] in, Represents a node With nodes At any moment Temporal co-occurrence similarity, This represents the set of timestamps indicating the co-occurrence of the two. , Each represents its own set of related timestamps. Represents a timestamp element. Represents a discrete-time index. Indicates the time decay coefficient. This is the symbol for an exponential function;

[0043] Step S311: Use adjacency similarity to measure structural proximity, and adopt normalized common neighbors;

[0044] ,

[0045] in, Represents a node With nodes Topological similarity, Represents a node The neighborhood group, Represents a node The set of neighbors;

[0046] Step S312: The three components of attribute, temporal co-occurrence and topological structure are fused into graph similarity by convex combination and updated over time.

[0047] ,

[0048] in, Represents a node With nodes At any moment Fusion similarity, , , These represent the fusion weights of the three components, satisfying... and ;

[0049] Step S313: Node merging is performed when the fusion similarity reaches the threshold and there is no hard conflict. Consistency is determined by the alignment of key fields.

[0050] ,

[0051] ,

[0052] in, Indicates at time Merge nodes With nodes The judgment quantity, Indicates an indicator function, Indicates the merger threshold. Indicates a consistency indicator. This represents the set of key field indexes that are included in hard constraints. Indicates the node at the 1st The values ​​for each key field This indicates that the field is missing. Indicates logical OR;

[0053] Step S314, when When nodes are merged into the same entity set, the disjoint-set data structure or connected component update transitive closure is used, and the attribute set and adjacency set of the entity representative node are merged and deduplicated.

[0054] As a preferred embodiment of the DMP data integration intelligent management method based on dynamic control described in this invention, the priority allocation mechanism includes:

[0055] The priority of each audience arm is determined based on its expected returns and preset business metrics; while meeting the overall budget constraints, a base budget is allocated to high-priority audience arms, and a budget share for exploration is reserved for low-priority audience arms.

[0056] As a preferred embodiment of the DMP data integration intelligent management method based on dynamic control described in this invention, the method is executed collaboratively by a data access module, a streaming processing module, an audience management module, a real-time decision-making module, and a strategy output module in a data management platform, wherein the real-time decision-making module is used to execute the contextual multi-armed gambling machine algorithm and generate the audience targeting strategy.

[0057] The beneficial effects of this invention are as follows: By constructing an integrated user behavior data view within a data management platform and continuously updating it in a streaming manner, this invention aggregates multi-source data from browsing, interaction, and devices under a unified standard, significantly reducing the latency and bias caused by inconsistent labeling and cross-system fragmentation. On the audience side, dynamic clustering incrementally maintains segmentation boundaries, and combined with contextual multi-armed gambling machines, candidate segments are evaluated and selected in real time, utilizing both historically accumulated effective experience and retaining necessary exploration to quickly capture new interest groups emerging during promotional periods. On the learning and decision-making side, online incremental updates and robust feedback processing are introduced, enabling continuous assimilation of immediate and lagging information on clicks and conversions, shortening the path from behavioral changes to strategy adjustments, and mitigating the lagging impact of periodic retraining. On the resource orchestration side, a budget priority allocation mechanism is used to roll back resources under overall constraints, balancing the conversion efficiency of the main audience and the exploration efficiency of the long-tail audience, improving the robustness of budget usage and the stability of revenue in high-concurrency scenarios. On the identifier fusion side, attribute, time co-occurrence, and topological three-way evidence are used to complete graph entity parsing, reducing audience overlap and omissions, and improving the consistency of the targeting loop. Therefore, this invention achieves end-to-end adaptive capabilities from data aggregation and audience modeling to real-time decision-making and budget control without relying on cumbersome manual parameter tuning, significantly improving response speed, targeting accuracy and system stability during peak promotional periods. Attached Figure Description

[0058] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation on the scope of this application.

[0059] Figure 1 This is a flowchart illustrating the DMP data integration intelligent management method based on dynamic control in the embodiment. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0061] All terms used in this application (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0062] For example, the terms “first” and “second” used in this application are only used to distinguish and describe similar objects, to differentiate the first object from another object, and are not used to describe a specific order or sequence, nor should they be interpreted as indicating or implying relative importance.

[0063] This application proposes an integrated intelligent management method for DMP data based on dynamic control, combining... Figure 1 As shown, the method includes:

[0064] Step S1: Ingest multi-source user behavior data streams in real time. The multi-source user behavior data streams include at least user browsing history, interaction events, and device information. Preprocess the multi-source user behavior data streams and associate them according to user identifiers to form an integrated user behavior data view.

[0065] In this embodiment, multi-source user behavior data streams refer to event streams originating from in-site tracking points, server logs, and third-party feedback. Preprocessing includes deduplication, time zone unification, and abnormal field repair. Association by user identifier uses login account, device fingerprint, and persistent identifier as primary keys to form a continuous behavior view of the same user. The default end-to-end latency from event arrival to availability is no more than 3 seconds, and out-of-order arrival is tolerated for no more than 120 seconds. The default session time window for cross-identifiers is 30 minutes, and the default weak association window for cross-source identifiers is 7 days. These criteria are set based on experience with concurrency and log jitter during peak business periods. The user identifiers required for association are obtained through server authentication or client identifier collection. Session division starts from the most recent active time. Optionally, when entity parsing is not yet complete, a partial view is temporarily formed by aggregating known identifiers, and historical data is supplemented after entity merging is completed. If necessary, when fields are missing or location permissions are restricted, city-level coarse-grained location and general device type are used as default values ​​to ensure that subsequent steps can be executed continuously.

[0066] Step S2: Based on the integrated user behavior data view, define multiple audience segments as multiple audience arms;

[0067] Step S3: Obtain real-time context features based on the current ad request to be delivered, input the real-time context features into the context multi-armed gambling machine algorithm, and calculate the expected return for each audience arm;

[0068] Step S4: Based on the expected return of each audience arm, a priority allocation mechanism is used to dynamically adjust the allocation of the advertising budget among multiple audience arms;

[0069] Step S5: Output the audience targeting strategy corresponding to the advertising budget allocation results to drive real-time ad delivery;

[0070] In one embodiment, the real-time context features include at least one of the following: the user's real-time browsing sequence, the current timestamp, geolocation information, and the terminal device type;

[0071] Specifically, the real-time browsing sequence refers to the most recent one or more page or product interaction records sorted by time. The timestamp uses millisecond-level sampling in a unified time zone. Geographic information comes from network addresses or city-level tags returned by system positioning. The terminal device type is mapped to a limited set such as mobile terminals, personal computing devices, and TV terminals based on user agents or system identifiers. The default sequence length is the most recent 20 records or records within the last 15 minutes. When the upper limit is exceeded, a first-in-first-out (FIFO) elimination method is used. When the geographic information accuracy is insufficient, the city level is used as the standard. When the device type cannot be determined, it is classified as "other". To ensure consistency, the feature timeout threshold is 10 minutes by default. If the feature is not updated after the threshold, it is considered outdated and a feature refresh is triggered. Optionally, the browsing sequence can be truncated by page level or category tag to reduce noise. When positioning is rejected or the network address is unavailable, the location feature is set to empty and participates in the calculation with default weight in subsequent processing.

[0072] In one embodiment, the contextual multi-armed slot machine algorithm updates the expected reward for each audience arm based on real-time feedback signals, which include at least one of click events or conversion events.

[0073] The steps for updating the expected returns of each audience arm based on real-time feedback signals include:

[0074] Step S301: Establish an online attribution channel for exposure-click-conversion; the system will monitor the data in real time. Select Audience Arm The click and conversion results of the request are then collected to form a real-time feedback signal that can be used for subsequent updates. Unselected audience arms only undergo time decay processing of historical information.

[0075] Step S302: Maintain the posterior state of click-through rate and single conversion value estimate for each audience arm, which are used as two components of expected return, and refresh online when new feedback arrives;

[0076] Step S303: Treat the click as a Bernoulli event and introduce a time discount, applying it to the selected audience arm at time... Incremental update of posterior parameters:

[0077] ,

[0078] in, Indicates the audience arm At any moment Success count parameter, Indicates the audience arm At any moment The failure count parameter, and For the parameters of the previous time step, This represents the time discount factor for click-through rate updates. Indicates time Should audience arm be selected? The indicated quantity, This indicates whether a click occurred during the request. Indicates the audience arm index, Indicates a discrete-time index;

[0079] For example, click feedback from selected audience arms is deduplicated on the server side and written to the online channel. Success and failure counts are recorded as two cumulative quantities of the click-through rate posterior state. The time discount coefficient is used to control the influence intensity of old samples and achieve smooth decay with a value close to 1. The default time discount is 0.98, and the adjustable range is 0.95 to 1.00, which is set according to the stability and latency characteristics of the peak click distribution. The selection indicator is directly given by the real-time decision result, and the click indicator is determined by the deduplicated click event and is verified twice with request identifier and time window. Optionally, to avoid amplification caused by sudden high-frequency repeated reporting, only one incremental update is allowed per second for the same request. If necessary, when the log delay exceeds 1 hour or repeated reporting occurs, only the decay update is performed or the duplicate is discarded to ensure monotonous status and traceability.

[0080] Step S304: When a conversion event occurs, update the single conversion value estimate of the audience arm in an online smooth manner. For samples that only have exposure or clicks but no conversion is observed, retain the existing estimate and perform a slight time decay. For conversions with lag attribution, update the corresponding historical moment when the event arrives to reduce bias.

[0081] Step S305: Multiplicatively combine the posterior mean of the click-through rate with the value of a single conversion to obtain the expected return for real-time decision-making.

[0082] ,

[0083] in, Indicates the audience arm At any moment Expected returns This represents the mean click-through rate based on the Beta posterior distribution obtained in step S303. Indicates the audience arm At any moment Estimate the value of a single conversion. Indicates the audience arm index, Indicates a discrete-time index;

[0084] Similarly, the expected return is derived by multiplying the posterior mean of the click-through rate (CTR) with the value per conversion (VPC). The VPC is measured in a uniform currency and aligned with tax and discount rules. The posterior mean of the CTR is truncated with upper and lower limits to avoid the influence of extreme values. The default lower limit of the CTR is 0.01% and the upper limit is 30%. The VPC is initialized with a robust average of the past seven days and continuously smoothed during the online phase. To facilitate coordination with budget allocation, the expected return is normalized within the same advertising campaign, and the sliding average of the entire audience arm is scaled proportionally by default. Optionally, if the VPC is temporarily unavailable, the median of the same campaign is used and replenished after a valid conversion is received. When abnormally high values ​​or suspected fraud signals appear in a short period of time, the contribution of the batch of samples is truncated to the upper quantile and an audit mark is recorded.

[0085] Step S306, when At that time, the parameters are only based on Attenuation is used to reflect information aging, and new audience arms employ gentle priors to gain early exploration opportunities;

[0086] Step S307: Truncate or robustly weight extreme feedback to reduce the impact of occasional anomalies on subsequent budget allocation and audience targeting strategies.

[0087] Specifically, this paper uses a discounted Bayesian approach to characterize the random event of a click, enabling the model to maintain a rapid response capability even when the distribution changes over time. By selecting an indicator, incremental learning is limited to the selected audience arm, while the unselected audience arm is maintained by decay to avoid noise propagation. The expected return adopts a multiplicative structure of the posterior mean of the click-through rate and the value of a single conversion, which facilitates direct connection with budget priority allocation and audience targeting strategies. At the same time, it retains engineering-level delay compensation and anomaly suppression to improve robustness.

[0088] Optionally, the default wide window for delayed attribution is 7 days. Recent online updates prioritize feedback from the last 24 hours, and the state at historical moments is corrected by backfilling when delayed events occur. Abnormal feedback is robustly weighted, and samples exceeding the upper quantile of the historical distribution are downweighted, with the upper quantile set to 99.5% by default. During the cold start phase, new audience arms receive minimum exploration guarantees to ensure early estimates are available. The guarantee ratio is enabled in the first update cycle and then taken over according to the regular strategy. If necessary, when no feedback is received for several consecutive update cycles, incremental learning of the audience arm is paused and the recent state is retained to avoid noise-driven oscillations.

[0089] In one embodiment, the contextual multi-armed gambling machine algorithm employs an exploration-exploitation balance mechanism, constructs a probability distribution based on the expected returns of each audience arm, and selects the target audience arm according to the probability distribution.

[0090] Furthermore, target selection can be performed by random sampling based on the expected return ratio to reflect exploration, or by selecting based on confidence limits to improve utilization efficiency. Both methods are based on the same input features and expected returns and keep the output as the audience arm index. The default exploration intensity is between 5% and 15%, which is adjusted according to the traffic volume and return variance. When the expected returns of multiple arms are similar, a stable random seed is used to scatter them to prevent load concentration. If necessary, when there is only a single available audience arm or other arms are in a circuit breaker state, it degenerates into a greedy selection to ensure service availability.

[0091] In one embodiment, the algorithm further includes updating the parameters of the contextual multi-armed gambling machine algorithm in real time through an online learning model, wherein the online learning model uses an incremental learning approach to iteratively adjust the parameters.

[0092] In one embodiment, the real-time ingestion of multi-source user behavior data streams is achieved through a distributed stream processing platform, which is used to perform distributed access, buffering, and computation of multi-source user behavior data streams.

[0093] In this embodiment, event processing is based on event time and a waiting level is set for out-of-order arrivals. Arrival and processing are decoupled to reduce the impact of jitter. The default allowed delay time is 120 seconds, the buffer refresh cycle is 1 second, and the status is stored in partitions according to user identifier and advertising plan. The status lifespan is 7 days by default to support recovery. Optionally, when cluster clock drift or event time unavailable is detected, processing time advancement is temporarily switched and consistent alignment is achieved after the clock is restored. If necessary, when computing resources are insufficient or back pressure is continuous, low-priority traffic is sampled at a fixed ratio to ensure the main path delay.

[0094] In one embodiment, multiple audience segments are generated based on a dynamic clustering algorithm, which employs a density-based clustering method and incrementally updates the audience segments according to changes in user behavior characteristics.

[0095] Specifically, the features of dynamic clustering can be derived from normalized statistics such as access frequency, category preference, and price range preference over a recent period. Clustering uses density threshold and minimum sample size as core parameters and supports addition and deletion. The default incremental update cycle is 5 minutes. The density threshold is determined by grid search within an empirical range after normalization based on distance units. The minimum sample size is 50 by default to suppress noise. Optionally, when the amount of data in a short period is insufficient, the subdivision results of the previous cycle can be temporarily maintained, and incremental updates can be resumed after the minimum sample size is reached. If necessary, when frequent oscillations of the subdivision boundary are detected in a very short period of time, the update can be temporarily frozen for no more than 2 minutes to stabilize the downstream strategy.

[0096] In one embodiment, associating multi-source user behavior data streams by user identifier includes integrating an entity parsing algorithm, constructing a graph structure for user identifiers from different data sources, and merging nodes belonging to the same entity based on graph similarity calculation to resolve identifier conflicts.

[0097] For example, attribute evidence includes stable fields such as account, device, and network; temporal co-occurrence evidence comes from shared activities within the same screen or short time window; and topological evidence measures structural proximity by common neighbor relationships. The three types of evidence generate similarity components under a unified mode and are fused on the same scale. By default, attribute and temporal co-occurrence components each account for about 40% of the weight, and topological components account for about 20%. The merging threshold is 0.8 by default and is jointly calibrated based on the accuracy and false merging rate of the historical annotation set. Optionally, when there are conflicts in key fields, merging is not performed even if the fusion score exceeds the threshold. If necessary, node pairs with sparse evidence or unreliable sources are marked as pending verification and re-evaluated after subsequent evidence arrives or the expiration date expires.

[0098] The steps for merging nodes belonging to the same entity based on graph similarity calculation include:

[0099] Step S308: Using the user identifier as a node, generate weighted undirected edges based on evidence such as same session, same device fingerprint, same account login, and same geographic-time co-occurrence. Extract attribute sets (such as account, device, network, geographic-time segment) and adjacency relationships for each node for subsequent similarity calculation and merging determination.

[0100] Step S309: Measure the overlap of attributes based on the attribute set. When numerical stabilization is required for the empty set, a small stabilization term is introduced. The similarity is written into the node pair for fusion.

[0101] ,

[0102] in, Represents a node With nodes Attribute similarity, Represents a node The set of attributes Represents a node The set of attributes Represents the cardinality of a set. Represents the numerically stable term. Indicates the node index. Indicates the node index;

[0103] Step S310: Perform time decay accumulation on co-occurrence events on the same screen, in the same session, or in a short time window to obtain the co-occurrence intensity updated over time;

[0104] ,

[0105] in, Represents a node With nodes At any moment Temporal co-occurrence similarity, This represents the set of timestamps indicating the co-occurrence of the two. , Each represents its own set of related timestamps. Represents a timestamp element. Represents a discrete-time index. Indicates the time decay coefficient. This is the symbol for an exponential function;

[0106] Step S311: Measure structural proximity using adjacency similarity, employing the normalized common neighbor (Salton index).

[0107] ,

[0108] in, Represents a node With nodes Topological similarity, Represents a node The neighborhood group, Represents a node The set of neighbors;

[0109] Step S312: The three components of attribute, temporal co-occurrence and topological structure are fused into graph similarity by convex combination and updated over time.

[0110] ,

[0111] in, Represents a node With nodes At any moment Fusion similarity, , , These represent the fusion weights of the three components, satisfying... and ;

[0112] Step S313: Node merging is performed when the fusion similarity reaches the threshold and there is no hard conflict. Consistency is determined by the alignment of key fields.

[0113] ,

[0114] ,

[0115] in, Indicates at time Merge nodes With nodes The judgment quantity, Indicates an indicator function, Indicates the merger threshold. Indicates a consistency indicator. This represents the set of key field indexes (such as strong identifier fields) that are included in hard constraints. Indicates the node at the 1st The values ​​for each key field This indicates that the field is missing. Indicates logical OR;

[0116] Step S314, when When nodes are merged into the same entity set, the transitive closure is updated using disjoint-set data structure or connected component analysis, and the attribute set and adjacency set of the entity representative node are merged and deduplicated so that the next round of similarity calculation can be performed at the entity level.

[0117] Similarly, to maintain the traceability and maintainability of entity sets, source evidence and threshold versions are recorded during merging. The attribute set and adjacency set representing nodes are updated using union and deduplication methods while maintaining priority coverage of key fields. The default shelf life of edges is 30 days, and expired evidence is removed from the calculation to avoid the accumulation of historical noise. Optionally, when there is multi-source batch import or large-scale correction, merging is performed in batches to control peak resource usage. If necessary, when serious conflicts or erroneous merging clues are found, further merging of the entity is suspended and handed over to the offline verification process.

[0118] Specifically, this approach aims at entity parsing, constructing similarity through three chains of evidence: attribute overlap, temporal co-occurrence, and graph structure. A single score is formed using convex combinations, and consistency constraints on key fields are then used to determine whether to merge. The attribute part covers stable elements such as accounts, devices, and networks; the temporal part strengthens recent evidence through a decay mechanism; and the structural part uses normalized indicators of common neighbors to characterize proximity in the graph. The merged score has an interpretable weight decomposition, facilitating adjustments and canary deployments according to business requirements. Setting thresholds and consistency constraints in the merging process can avoid erroneous merging under strong conflicts.

[0119] In one embodiment, the priority allocation mechanism includes:

[0120] The priority of each audience arm is determined based on its expected returns and preset business metrics; while meeting the overall budget constraints, a base budget is allocated to high-priority audience arms, and a budget share for exploration is reserved for low-priority audience arms.

[0121] Optionally, the budget allocation is recalculated on a rolling basis within a fixed period and the actual consumption of the previous period is corrected. The default recalculation period is 1 minute, which can be shortened to 10 seconds when traffic surges. Under the total budget constraint, the main audience arm is guaranteed a basic level of security, and the exploration budget is distributed proportionally among all audience arms and adaptively adjusted according to changes in returns. To prevent oscillations caused by short-term fluctuations, an upper limit on the change range is introduced, with a default single adjustment not exceeding 20% ​​of the budget of the previous period. If necessary, when there is no effective feedback for several consecutive periods, the current allocation is maintained and the exploration ratio is reduced to control risks.

[0122] In one embodiment, the method is executed collaboratively by a data access module, a streaming processing module, an audience management module, a real-time decision-making module, and a strategy output module in a data management platform, wherein the real-time decision-making module is used to execute the contextual multi-armed gambling machine algorithm and generate an audience targeting strategy;

[0123] Furthermore, the modules exchange structured messages through lightweight interfaces, including essential fields such as timestamps, request identifiers, and policy versions. By default, the target latency from decision to downstream deployment does not exceed 100 milliseconds. Policy outputs include target audience arm, budget guidance, and expiration date markers. Optionally, when real-time decision-making is unavailable, it degenerates to the most recently valid policy and performs consistency reconciliation after recovery. If necessary, to ensure cross-module consistency, a monotonically increasing version number and idempotent writing strategy are adopted for key fields to avoid duplicate deployments and state mismatches.

[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0125] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are meant to be within the scope of this application and form different embodiments. For example, all the embodiments above can be used in any combination. The information disclosed in this background section is intended only to enhance the understanding of the general background of this application and should not be construed as an admission or in any way implying that such information constitutes prior art known to those skilled in the art.

Claims

1. A DMP data integrated intelligent management method based on dynamic control, characterized in that, include: Step S1: Real-time capture of multi-source user behavior data streams, which include at least user browsing history, interaction events, and device information; preprocessing of the multi-source user behavior data streams and association by user identifier to form an integrated user behavior data view. Step S2: Based on the integrated user behavior data view, the multiple audience segments are defined as multiple audience arms; Step S3: Obtain real-time context features based on the current ad request to be delivered, input the real-time context features into the context multi-armed gambling machine algorithm, and calculate the expected return for each audience arm; Step S4: Based on the expected return of each audience arm, a priority allocation mechanism is used to dynamically adjust the allocation of the advertising budget among multiple audience arms; Step S5: Output the audience targeting strategy corresponding to the advertising budget allocation result to drive real-time ad delivery.

2. The DMP data integrated intelligent management method based on dynamic control as described in claim 1, characterized in that, The real-time context features include at least one of the following: the user's real-time browsing sequence, the current timestamp, geographic location information, and the terminal device type.

3. The DMP data integrated intelligent management method based on dynamic control as described in claim 2, characterized in that, The contextual multi-armed gambling machine algorithm updates the expected return of each audience arm based on real-time feedback signals, which include at least one of click events or conversion events; The step of updating the expected returns of each audience arm based on real-time feedback signals includes: Step S301: Establish an online attribution channel for exposure-click-conversion; the system will monitor the data in real time. Select Audience Arm The click and conversion results of the request are then collected to form a real-time feedback signal that can be used for subsequent updates. Unselected audience arms only undergo time decay processing of historical information. Step S302: Maintain the posterior state of click-through rate and single conversion value estimate for each audience arm, which are used as two components of expected return, and refresh online when new feedback arrives; Step S303: Treat the click as a Bernoulli event and introduce a time discount, applying it to the selected audience arm at time... Posterior parameter incremental update: , in, Indicates the audience arm At any moment Success count parameter, Indicates the audience arm At any moment The failure count parameter, and For the parameters of the previous time step, This represents the time discount factor for click-through rate updates. Indicates time Should audience arm be selected? The indicated quantity, This indicates whether a click occurred during the request. Indicates the audience arm index, Indicates a discrete-time index; Step S304: When a conversion event occurs, update the single conversion value estimate of the audience arm in an online smooth manner. For samples that only have exposure or clicks but no conversion is observed, retain the existing estimate and perform a slight time decay. For conversions with lag attribution, update the corresponding historical moment when the event arrives. Step S305: Multiplicatively combine the posterior mean of the click-through rate with the value of a single conversion to obtain the expected return for real-time decision-making. , in, Indicates the audience arm At any moment Expected returns This represents the mean click-through rate based on the Beta posterior distribution obtained in step S303. Indicates the audience arm At any moment Estimate the value of a single conversion. Indicates the audience arm index, This represents a discrete-time index.

4. The DMP data integrated intelligent management method based on dynamic control as described in claim 3, characterized in that, The contextual multi-armed gambling machine algorithm employs an exploration-exploitation balance mechanism, constructs a probability distribution based on the expected returns of each audience arm, and selects the target audience arm according to the probability distribution.

5. The DMP data integrated intelligent management method based on dynamic control as described in claim 4, characterized in that, It also includes updating the parameters of the contextual multi-armed gambling machine algorithm in real time through an online learning model, wherein the online learning model uses an incremental learning method to iteratively adjust the parameters.

6. The DMP data integrated intelligent management method based on dynamic control as described in claim 5, characterized in that, The real-time ingestion of multi-source user behavior data streams is achieved through a distributed stream processing platform, which is used for distributed access, buffering, and computation of the multi-source user behavior data streams.

7. The DMP data integrated intelligent management method based on dynamic control as described in claim 6, characterized in that, Multiple audience segments are generated based on a dynamic clustering algorithm, which employs a density-based clustering method and incrementally updates the audience segments based on changes in user behavior characteristics.

8. The DMP data integrated intelligent management method based on dynamic control as described in claim 7, characterized in that, Associating multi-source user behavior data streams by user identifier includes integrating entity parsing algorithms, constructing a graph structure for user identifiers from different data sources, and merging nodes belonging to the same entity based on graph similarity calculation; The steps for merging nodes belonging to the same entity based on graph similarity calculation include: Step S308: Using the user identifier as a node, generate weighted undirected edges based on evidence of the same session, same device fingerprint, same account login, and same geographic-temporal co-occurrence, and extract the attribute set and adjacency relationship for each node; Step S309: Measure the attribute overlap based on the attribute set. When numerical stabilization is required for the empty set, a small stabilization term is introduced, and the similarity is written into the node pair. , in, Represents a node With nodes Attribute similarity, Represents a node The set of attributes Represents a node The set of attributes Represents the cardinality of a set. Represents the numerically stable term. Indicates the node index. Indicates the node index; Step S310: Perform time decay accumulation on co-occurrence events on the same screen, in the same session, or in a short time window to obtain the co-occurrence intensity updated over time; , in, Represents a node With nodes At any moment Temporal co-occurrence similarity, This represents the set of timestamps indicating the co-occurrence of the two. , Each represents its own set of related timestamps. Represents a timestamp element. Represents a discrete-time index. Indicates the time decay coefficient. This is the symbol for an exponential function; Step S311: Use adjacency similarity to measure structural proximity, and adopt normalized common neighbors; , in, Represents a node With nodes Topological similarity, Represents a node The neighborhood group, Represents a node The set of neighbors; Step S312: The three components of attribute, temporal co-occurrence and topological structure are fused into graph similarity by convex combination and updated over time. , in, Represents a node With nodes At any moment Fusion similarity, , , These represent the fusion weights of the three components, satisfying... and ; Step S313: Node merging is performed when the fusion similarity reaches the threshold and there is no hard conflict. Consistency is determined by the alignment of key fields. , , in, Indicates at time Merge nodes With nodes The judgment quantity, Indicates an indicator function, Indicates the merger threshold. Indicates a consistency indicator. This represents the set of key field indexes that are included in hard constraints. Indicates the node at the 1st The values ​​of the key fields This indicates that the field is missing. Indicates logical OR; Step S314, when When nodes are merged into the same entity set, the disjoint-set data structure or connected component update transitive closure is used, and the attribute set and adjacency set of the entity representative node are merged and deduplicated.

9. The DMP data integrated intelligent management method based on dynamic control as described in claim 8, characterized in that, The priority allocation mechanism includes: The priority of each audience arm is determined based on its expected returns and preset business metrics; while meeting the overall budget constraints, a base budget is allocated to high-priority audience arms, and a budget share for exploration is reserved for low-priority audience arms.

10. The DMP data integrated intelligent management method based on dynamic control as described in claim 9, characterized in that, The method is executed collaboratively by a data access module, a streaming processing module, an audience management module, a real-time decision-making module, and a strategy output module in a data management platform. The real-time decision-making module is used to execute the contextual multi-armed gambling machine algorithm and generate the audience targeting strategy.