Method and device for crowd selection based on DMP and electronic equipment

By preprocessing and labeling user data from multiple data sources, combined with rule trees and similar population expansion algorithms, the problems of complex rule configuration, insufficient real-time performance, and limited scalability of the DMP population selection system are solved, achieving efficient and accurate population selection and cross-platform data fusion.

CN120780913APending Publication Date: 2025-10-14上海勃池信息技术有限公司
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510941753.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

The existing DMP-based crowd selection system has problems such as complex rule configuration, insufficient real-time performance and limited scalability, making it difficult to adapt to emerging business scenarios and cross-platform data integration needs.

Method used

By preprocessing the original user data from multiple data sources, generating user features and converting them into user tags, using distributed storage to store tags, using the operation interface to display tags and converting them into rule trees for selection, and combining similar crowd expansion algorithms to generate crowd packages, dynamic and static tag management is achieved.

Benefits of technology

It improves the efficiency and accuracy of crowd selection, adapts to emerging business scenarios and cross-platform data integration needs, and enhances the real-time and scalability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780913A_ABST
    Figure CN120780913A_ABST
Patent Text Reader

Abstract

The invention provides a DMP-based crowd selection method and apparatus, and an electronic device, a data management platform pre-processes multi-source user data, generates user features corresponding to the pre-processed data based on a preset rule, converts the obtained user features into user tags based on a pre-constructed tag system, and sends the user tags to the user management platform; the method comprises the following steps: acquiring a user tag, storing the acquired user tag in a distributed storage mode, displaying the stored user tag through an operation interface, converting a circling operation for a target user tag displayed on the operation interface into a rule tree, and circling a target user corresponding to the target user tag based on the rule tree. And associating the circling result with the identification information of the target user corresponding to the circling result to generate a crowd packet, and expanding the obtained crowd packet by adopting a similar crowd expansion algorithm. According to the invention, the problems of complex rule configuration, insufficient real-time performance, limited expansibility and the like of a DMP-based crowd selection system in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method, device and electronic device for selecting a crowd based on a DMP. Background Art

[0002] Data Management Platforms (DMPs) are already very mature in the market. In digital marketing, DMPs can integrate first-party and third-party data to build user tagging systems, allowing advertisers or marketing personnel to filter target audiences based on specific rules (i.e., "audience selection"). These systems provide in-depth data insights and intelligent management, guiding advertisers or marketing personnel in making powerful advertising optimization and placement decisions. However, current DMP-based audience selection systems still have some shortcomings: (1) Complex rule configuration: It requires manual definition of multi-dimensional label combination logic, which is inefficient and prone to missing related features.

[0003] (2) Lack of real-time performance: Data update delays lead to delayed selection results and inability to dynamically respond to market changes.

[0004] (3) Limited scalability: Fixed tag systems are difficult to adapt to emerging business scenarios or cross-platform data integration needs. Summary of the Invention

[0005] In view of this, an object of the present invention is to provide a method, device and electronic device for selecting a crowd based on DMP to alleviate the above-mentioned problems existing in the prior art.

[0006] In a first aspect, an embodiment of the present invention provides a DMP-based crowd selection method, which is applied to a data management platform, comprising: preprocessing raw user data obtained from multiple different data sources, and generating user features corresponding to the preprocessed data based on preset rules; converting the obtained user features into user tags based on a pre-built tag system, and storing the obtained user tags in a distributed storage manner, and then displaying the stored user tags through an operation interface; wherein the user tags include static tags and dynamic tags; converting a selection operation for a target user tag displayed on the operation interface into a rule tree, and selecting target users corresponding to the target user tag based on the rule tree; wherein the selection result is a target user set consisting of multiple target users; associating the selection result with the identification information of the corresponding target user to generate a crowd package, and expanding the obtained crowd package using a similar crowd expansion algorithm; wherein the identification information includes a unique identification of each target user in the target user set.

[0007] In a second aspect, an embodiment of the present invention further provides a DMP-based crowd selection device, which is applied to a data management platform and includes: a feature generation module for preprocessing raw user data obtained from multiple different data sources and generating user features corresponding to the preprocessed data based on preset rules; a label generation module for converting the obtained user features into user labels based on a pre-built label system, storing the obtained user labels in a distributed storage manner, and then displaying the stored user labels through an operation interface; wherein the user labels include static labels and dynamic labels; a selection module for converting a selection operation for a target user label displayed on the operation interface into a rule tree, and selecting target users corresponding to the target user label based on the rule tree; wherein the selection result is a target user set consisting of multiple target users; a crowd package generation module for associating the selection result with the identification information of the corresponding target user to generate a crowd package, and expanding the obtained crowd package using a similar crowd expansion algorithm; wherein the identification information includes the unique identification of each target user in the target user set.

[0008] In a third aspect, an embodiment of the present invention further provides an electronic device comprising a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the DMP-based crowd selection method described in the first aspect above.

[0009] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the DMP-based crowd selection method described in the first aspect above.

[0010] The embodiments of the present invention provide a method, device, and electronic device for crowd selection based on a DMP. The data management platform first preprocesses the original user data obtained from multiple different data sources, and generates user features corresponding to the preprocessed data based on preset rules. The obtained user features are then converted into user tags based on a pre-built tag system, and the obtained user tags are stored in a distributed storage manner. The stored user tags are then displayed through an operation interface, and the selection operation for the target user tag displayed on the operation interface is converted into a rule tree. The target user corresponding to the target user tag is selected based on the rule tree. The selection result is then associated with the identification information of the corresponding target user to generate a crowd package, and the obtained crowd package is expanded using a similar crowd expansion algorithm. The above technology is used to generate user tags by fusing multi-source data, and the selection operation for the user tag on the operation interface is converted into a rule tree to achieve crowd selection. This can adapt to emerging business scenarios or cross-platform data fusion needs, while improving the efficiency and accuracy of crowd selection.

[0011] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the present invention. The purposes and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.

[0012] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0014] Figure 1 Schematic diagram of a flow chart of a method for selecting a group of people based on a DMP in an embodiment of the present invention; Figure 2 Schematic diagram of the structure of a DMP-based crowd selection device in an embodiment of the present invention; Figure 3 Schematic diagram of the structure of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION

[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0016] At present, the existing DMP-based crowd selection system mainly has problems such as complex rule configuration, insufficient real-time performance, and limited scalability.

[0017] Based on this, the present invention provides a DMP-based crowd selection method, device and electronic device, which can alleviate the above-mentioned problems existing in the prior art.

[0018] To facilitate understanding of this embodiment, a method for selecting a crowd based on a DMP disclosed in an embodiment of the present invention is first described in detail. This method can be applied to a DMP. Figure 1 As shown, the method may include the following steps: Step S102 : pre-processing the original user data obtained from a plurality of different data sources, and generating user features corresponding to the pre-processed data based on preset rules.

[0019] Raw user data can be structured and unstructured data such as user behavior data (such as click records, browsing records, purchase records, etc.), transaction data, device fingerprints (IMEI, IDFA), geographic location information, and third-party data (such as credit information, social attributes, and consumption capacity scores) from different data sources (such as apps, official websites, mini-programs, H5 applications, and PC applications).

[0020] After integrating raw user data from various data sources, the integrated data needs to be preprocessed. This preprocessing process primarily involves data cleaning to remove duplicate, erroneous, or invalid data, and standardizing the data to ensure consistency and comparability. Feature extraction is then performed on the preprocessed data using predefined rules to obtain user features.

[0021] Step S104 : converting the obtained user features into user tags based on the pre-built tag system, storing the obtained user tags in a distributed storage manner, and then displaying the stored user tags through an operation interface.

[0022] Among them, user tags can include static tags and dynamic tags, which can specifically be tags covering multiple categories such as basic attributes of the population, geographical distribution, interest preferences, offline scenarios, etc.

[0023] A multi-level label system can be pre-designed: a three-layer label structure (basic attributes -> behavioral characteristics -> labels, such as demographic attributes -> age -> 18-25 years old, behavioral characteristics -> shopping preferences -> beauty product enthusiasts, etc.) is designed based on business needs, and static labels (such as gender, age, etc.) and dynamic labels (such as near-7-day activity, near-30-day purchase frequency, etc.) are mixedly managed.

[0024] User features obtained by extraction, processing, mining, etc. can be aggregated to form more comprehensive user feature descriptions, which can cover user basic attributes, behavioral characteristics, interest preferences, etc. Afterwards, the aggregated user features can be labeled using a pre-designed multi-level label system, i.e., converting user features into labels that are easy to understand and use (for example, converting user purchase frequency features into "high-frequency purchaser" and "low-frequency purchaser" labels, and converting user interest preference features into "technology enthusiast" and "fashionista" labels). The labels converted from user features can be written into a distributed storage architecture for distributed storage. An operation interface can be pre-set to display the stored user labels on the operation interface, thereby facilitating relevant personnel to perform corresponding operations on the operation interface.

[0025] In step S106, the circle selection operation on the target user label displayed on the operation interface is converted into a rule tree, and the target user corresponding to the target user label is circled based on the rule tree.

[0026] The circle selection result can be a target user set composed of multiple target users.

[0027] The circle selection operation can mainly include single-label circle selection and visual multi-label combination operation. Single-label circle selection: a single label (such as "high-value user") is directly selected on the operation interface to generate a people package later. Visual multi-label combination operation: multiple labels are selected on the operation interface and the logical operation (for example, "beauty product enthusiast AND no purchase in the past 30 days") and numerical conditions (such as "purchase frequency greater than 5" and "consumption amount in the past 7 days greater than 500 yuan") between the selected labels are set. The logical operation includes intersection (AND), union (OR), and difference (NOT), etc., to meet various people circle selection needs.

[0028] Circling the target user can be considered as "people circle selection", and "people" is essentially a pre-defined label combination that maps label data to a segmented people. After configuring the circle selection operation, the circle selection operation can be converted into a rule tree, and then the rule tree can be converted into a JSON format to implement the circle selection of the target user later.

[0029] Step S108 : Associating the circled result with the identification information of the corresponding target user to generate a crowd package, and expanding the obtained crowd package using a similar crowd expansion algorithm.

[0030] The identification information may include a unique identification (such as user ID, device ID, mobile phone number, etc.) of each target user in the target user set.

[0031] In actual application, for the same user, the original user identifiers of different platforms (such as cookies, mobile phone numbers, device IDs) can be mapped to a unified anonymous ID through encryption algorithms (such as MD5, SHA-256) as the user's unique identifier and associated with the user's entity to construct a user portrait associated with a unified user identifier.

[0032] An embodiment of the present invention provides a method for selecting a crowd based on a DMP. The data management platform first pre-processes the original user data obtained from multiple different data sources, and generates user features corresponding to the pre-processed data based on preset rules. The obtained user features are then converted into user tags based on a pre-built tag system, and the obtained user tags are stored in a distributed storage manner. The stored user tags are then displayed through an operation interface, and the selection operation for the target user tag displayed on the operation interface is converted into a rule tree. The target user corresponding to the target user tag is selected based on the rule tree. The selection result is then associated with the identification information of the corresponding target user to generate a crowd package, and the obtained crowd package is expanded using a similar crowd expansion algorithm. The above technology is used to generate user tags by fusing multi-source data, and the selection operation for the user tag on the operation interface is converted into a rule tree to achieve crowd selection. This can adapt to emerging business scenarios or cross-platform data fusion needs, while improving the efficiency and accuracy of crowd selection.

[0033] As a possible implementation, generating user features corresponding to the pre-processed data based on preset rules in step S102 may include: In step A1, basic features are extracted from the preprocessed data according to preset rules, and the basic features are statistically analyzed using a preset statistical method to generate a first feature. Then, a preset data mining algorithm is used to analyze the preprocessed data to generate a second feature.

[0034] After the raw user data from different data sources is integrated and preprocessed, the basic features can be extracted from the preprocessed data according to preset rules (for example, the purchase frequency, the purchase amount and other features are extracted from the purchase record, and the browsing time, the browsing page type and other features are extracted from the browsing record). Then, the extracted basic features can be further statistically analyzed by using statistical methods to generate higher-level features (for example, the average purchase amount, the standard deviation of the purchase amount and other statistical quantities of the user are calculated as features to reflect the purchase ability and purchase stability of the user). The machine learning algorithm (such as the clustering analysis algorithm, the association rule mining algorithm and the like) can be used as the data mining algorithm to discover hidden patterns and rules from the preprocessed data, and then new features are constructed (for example, the groups with similar purchase behaviors are discovered by using the clustering analysis, and the corresponding features are constructed for each group with similar purchase behaviors).

[0035] Step A2, analyzing the metadata and data distribution of the preprocessed data to generate the third features.

[0036] In the foregoing example, the metadata and data distribution (microstructure) of the preprocessed data can also be used to automatically discover features. The operation mode of the specific feature discovery can be as follows: the metadata (including the data type, the data format, the data source and other information) of the preprocessed data is analyzed to understand the overall structure and characteristics of the preprocessed data. The information contained in the metadata helps to better understand the preprocessed data and discover the potential features of the preprocessed data. The data distribution of the preprocessed data can be further explored, for example, the data distribution can include the distribution form (such as the normal distribution, the skew distribution) of the data, the association relationship (such as the correlation, the causality) between the data and other microstructure information. By analyzing these microstructure information, the abnormal values, the outliers, the special patterns and other information in the data can be discovered as the basis for constructing new features, and then the new features are constructed according to the analysis results of the microstructure information.

[0037] Step A3, taking the basic features, the first features, the second features and the third features as the user features.

[0038] In the foregoing example, specific automatic feature discovery algorithms (such as the algorithm based on the feature importance ranking, the algorithm based on the feature intersection and the like) can also be used to automatically screen and combine the existing features according to the characteristics of the data and the business requirements, and then new and meaningful features are generated. The existing features and the newly generated features are taken as the user features.

[0039] As a possible implementation, the dynamic label can include a prediction label and a real-time behavior label, and the label system can include a first rule engine, a distributed stream processing engine and a pre-trained machine learning model. Based on this, the step S104 of converting the obtained user features into user labels based on the pre-constructed label system can include: Step a1, generate static labels corresponding to user features using the first rule engine, generate real-time behavior labels corresponding to user features using the distributed stream processing engine, and then generate prediction labels corresponding to user features using the machine learning model.

[0040] Following the previous example, a rule engine can be used to generate a "highly active user" label based on predefined rules such as "login frequency in the last 30 days is not less than 5 times", and Spark SQL can be used to batch process user features to generate static labels such as gender, age, city level, etc., and then store the basic labels in a columnar database such as ClickHouse to form a static label library that supports fast queries. Real-time behavior labels representing short-term user interests can be generated by a stream processing engine such as Flink by calculating user behavior in the last N hours (e.g., "browsed mobile phone category 3 times in the last 1 hour"), and real-time behavior labels reflect the current behavior of the user in real time. Machine learning models such as clustering algorithms, decision trees, etc. can be used to automatically mine user potential features from input user features to predict and generate dynamic labels such as "high potential purchase users" and "churn risk users" as prediction labels, which reflect the user's behavior or characteristics in the future; features input into the machine learning model for prediction can include user behavior sequence features (such as features extracted from the browse-add-purchase conversion chain), text features (such as sentiment analysis of reviews), graph features (such as social relationship networks), etc. User potential feature labels predicted by the machine learning model (i.e., prediction labels) can be stored in the feature platform for subsequent audience selection, precise delivery, and other business use.

[0041] Specifically, the prediction label and the real-time behavior label differ in data source and update frequency: the prediction label is predicted based on historical data and machine learning models, and the update frequency of the prediction label is relatively low, usually updated when the data accumulates to a certain extent or the model is retrained, for example, new prediction labels are generated to update the prediction labels after the model is retrained according to new historical data every week or every month; while the real-time behavior label is generated immediately according to the real-time behavior data of the user, and the update frequency of the real-time behavior label is relatively high (generally higher than the update frequency of the prediction label), for example, when a user has just completed a purchase behavior, the user can be immediately labeled as a "real-time purchase user".

[0042] There are the following differences in the meaning of predictive tags and real-time behavior tags: Predictive tags focus on reflecting the potential characteristics or behavioral trends that users may exhibit in the future. Predictive tags are predictions of users' long-term behavior and potential value. For example, the "high-potential paying user" tag indicates that the user is more likely to make high-value purchases in the future. Real-time behavior tags mainly reflect the user's current behavior status and immediate needs. Real-time behavior tags are records of users' short-term behavior. For example, the "real-time browsing user" tag indicates that the user is currently browsing a product page and may be interested in related products.

[0043] There are the following differences in application scenarios between predictive tags and real-time behavioral tags: Predictive tags are often used in long-term strategy formulation (such as user stratification operations, optimization of personalized recommendation strategies, etc.). By formulating targeted operation strategies for user groups with different predictive tags, user retention and payment conversion rates can be improved. For example, exclusive promotional activities and personalized product recommendations are provided to target users with the "high-potential paying users" label to promote target user consumption; real-time behavioral tags are suitable for real-time marketing (such as pushing relevant promotional information when users browse products in real time) and instant service responses (such as providing timely assistance when users initiate customer service inquiries). For example, when a user browses a mobile phone in real time, discount information and user reviews of the mobile phone can be pushed immediately to increase the user's willingness to buy.

[0044] The combination of predictive tags and real-time behavioral tags provides a more comprehensive understanding of user behavior and needs, enabling more targeted services and marketing. For example, for users tagged with both "high-potential paying users" and "real-time browsing behavior," marketing efforts can be strengthened, with more personalized recommendations and offers, improving conversion rates. Real-time behavioral tags also provide a basis for updating predictive tags. When a user's real-time behavior changes significantly, it may be necessary to reassess their underlying characteristics and update the predictive tags.

[0045] In step a2, static tags, real-time behavior tags, and prediction tags are combined into an initial tag combination, and the initial tag combination is optimized using a preset tag semantic model, and the tags in the optimized tag combination are used as user tags.

[0046] Continuing with the previous example, after static tags, real-time behavior tags, and prediction tags are combined to form an initial tag combination, a preset tag semantic model can be used to perform semantic analysis on the initial tag combination to identify redundant tags in the initial tag combination; then, the preset tag semantic model is used to establish a tag association graph of the initial tag combination based on the redundant tags, in which nodes are used to represent tags and edges between nodes are used to represent the association relationship between tags; then, the preset tag semantic model is used to use the tag association graph to determine the optimal tag in the initial tag combination, and the obtained optimal tags are combined into an optimized tag combination, and then the tags in the optimized tag combination are used as user tags.

[0047] In actual applications, a tag semantic engine can be developed in advance to recommend tag combinations when building a tag system. This recommendation method primarily involves using the tag semantic engine to automatically identify redundant tags (such as synonyms and antonyms) using NLP technology and establish a tag association graph. The tag semantic engine then analyzes the semantic relationships between tags based on this established tag association graph to recommend the optimal tag combination, thereby more accurately describing the characteristics of the target population.

[0048] Recommending optimal tag combinations during the tag system construction phase helps identify and integrate redundant tags in advance, avoiding redundancy and confusion in the tag system and improving the quality and accuracy of the tag system. By recommending optimal tag combinations, we can more accurately describe the characteristics of the target population. This allows us to more accurately select people who meet the target characteristics during the subsequent population selection process, improving the effectiveness and efficiency of population selection and providing strong support for subsequent targeted delivery.

[0049] As a possible implementation method, the above-mentioned DMP-based crowd selection method may also include: obtaining real-time user behavior data, and using a second rule engine to determine whether the real-time user behavior data meets preset conditions; if the real-time user behavior data meets the preset conditions, using the second rule engine to update the user label according to the preset label information corresponding to the preset conditions.

[0050] In actual applications, custom tag rules supported by a rule engine (such as Drools) can be pre-configured (including conditions such as "login device changed at least twice in the past seven days," "product category browsed at least five times in the past 30 days without placing an order," and multiple tag names required to update user tags). Tag rules can be defined in a specific format (such as JSON) through the interface or interface provided by the rule engine, so that the rule engine can parse the tag rules and generate executable logic (such as SQL query instructions and bitmap operation instructions) based on the tag names defined by the tag rules. The rule engine then matches the tag rules against newly received real-time user behavior data (which can be obtained by real-time monitoring of user behavior events through Flink or Kafka). When the real-time user behavior data matches all tag rules (that is, the real-time user behavior data meets the conditions included in all tag rules), the rule engine executes the executable logic converted from all tag rules to update user tags according to the tag names defined by all tag rules. For example, when a user's behavior meets the conditions of the "frequent login device changes" rule, the rule engine will add the "high_risk_device_change" tag to the user. When a user's behavior meets the conditions of the "remove after purchase" rule, the rule engine will remove the "add to cart but not purchased" tag. The rule engine can also persist updated user tags in the database for subsequent business use, such as demographic selection and targeted delivery. As new user behavior data continues to flow in and business needs change, the rule engine can dynamically update user tags to ensure their timeliness and accuracy.

[0051] In actual applications, user tag quality can be monitored by measuring coverage and accuracy, enabling early warnings or deletion of low-quality user tags (e.g., those with coverage <1%). User tag expiration dates can also be set (e.g., setting a "Double Eleven Active User" tag to automatically expire after the event). Expired user tags can then be regularly cleared using tools like Airflow to manage the user tag lifecycle.

[0052] As a possible implementation, the rule tree may adopt an abstract syntax tree. Based on this, selecting the target user corresponding to the target user tag based on the rule tree in step S106 may include: Step B1: If there are high-frequency query tags in the target user tags, an inverted index is created for the high-frequency query tags in the target user tags.

[0053] The inverted index includes each high-frequency query tag in the target user tag and the identification information of the corresponding target user.

[0054] Step B2: Prune the Rule Tree based on the pre-built cost model, and execute the pruned Rule Tree to obtain the target user set.

[0055] Exemplarily, a pre-built cost model may be used to prune the rule tree, and then the pruned rule tree is traversed to obtain an initial target user set whose target user labels correspond to each node of the pruned rule tree, and each obtained initial target user set is converted into a corresponding bitmap; then, the obtained bitmap is subjected to the set operation contained in the pruned rule tree, and the obtained set operation result is converted into a target user set.

[0056] In actual application, the selection operation on the target user corresponding to the target user tag can be converted into an abstract syntax tree (AST). The form of the AST is as follows: AND( TAG("highly active users"), NOT(TAG("Purchased New Product X")), GT(ATTR("Consumption amount in the last 7 days"), 500) ) The specific purpose of a user operation depends on the business scenario. Common scenarios are as follows: Data query scenarios: In data management systems or database query scenarios, users may need to retrieve data based on specific criteria. For example, in an e-commerce system, if you want to query "a list of users whose purchases in the past 30 days exceeded 1,000 yuan and included electronics," you can circle the relevant items to retrieve matching data.

[0057] Data screening and filtering scenarios: During data processing and analysis, users may need to screen and filter data to extract areas of interest. For example, when analyzing user behavior data, users can select users who visited a specific page within a specific time period by entering a circle.

[0058] Rule definition and matching scenarios: In rule engine applications, the circling operation is used to define rules (such as "If a user's login device has changed at least twice in the past seven days, then the user is marked as high-risk") so that the system can match and process data based on these rules.

[0059] Data conversion and calculation scenarios: Relevant personnel may need to convert and calculate data to obtain new data or indicators. For example, they can calculate the average spending amount, purchase frequency, etc. of users by inputting relevant circle operations.

[0060] The purpose of converting the selection operation of the target user corresponding to the target user tag into AST is mainly to: (1) Facilitates syntax analysis and verification: AST can clearly represent the syntax structure of the selection operation. By traversing the AST, you can quickly check whether the selection operation complies with the predefined syntax rules. For example, in data query, check whether the keywords in the query statement are used correctly and whether the brackets match. Based on the AST, you can perform a more in-depth semantic analysis to ensure that the selection operation is semantically reasonable. For example, in the rule definition, check whether the conditions in the rule can correctly reference the data fields and whether the selection operation complies with the business logic.

[0061] (2) Optimize execution plan generation: Based on AST, the system can generate more efficient execution plans. For example, in database queries, the optimizer can analyze query conditions based on AST and select appropriate indexes, connection methods, etc. to improve query performance. For complex selection operations, AST can help the system identify parts that can be executed in parallel, thereby making full use of multi-core processors or distributed computing resources and improving the processing speed of selection operations.

[0062] (3) Support code generation and interpretation execution: AST can be converted into platform-specific code, such as SQL statements, Java code, etc. For example, the user's data query operation AST is converted into SQL statements and executed directly in the database; the system can directly interpret and execute AST and perform selection operations node by node. This method is suitable for some dynamic scenarios that require flexible processing (such as rule matching in rule engines).

[0063] (4) Easy to modify and expand: AST represents the selection operation in a tree structure. When modifying the selection operation, you only need to modify the corresponding node in AST, without rewriting the entire selection operation. For example, in the rule definition, if you need to modify the rule conditions, you only need to find the corresponding condition node in AST and modify it. By expanding the node types and rules of AST, new selection operations and functions can be easily supported. For example, in the data query system, when adding a new query operator or function, you only need to define the corresponding node type and processing logic in AST.

[0064] In actual application, the implementation of crowd selection can be optimized from the perspectives of computing performance, query response time, and resource consumption during the crowd selection process to significantly improve the selection efficiency. The optimization method of crowd selection can mainly include the following steps: (1.1) Regular tree pruning.

[0065] Selection is often based on a series of rules. Rule tree pruning can remove redundant rules and logic, reduce unnecessary computations, and make the selection process more efficient. Specifically, a cost model can be used to calculate evaluation metrics (such as label usage frequency and computational complexity) of the rule tree to prune redundant nodes (for example, merging overlapping label combinations), automatically simplifying redundant logic.

[0066] The calculation process of the cost model mainly includes counting the frequency of label usage, evaluating the node computational complexity, and calculating the node cost value.

[0067] Count tag usage frequency: Count the number of times each tag has been used in historical selection operations. Tags with high usage frequency may be more important in the selection process, while nodes corresponding to tags with low usage frequency in the rule tree may be redundant nodes.

[0068] Evaluate node computational complexity: Evaluate the computational complexity of each node in the rule tree. For example, nodes involving complex operations (such as nested functions and multiple condition combinations) have high computational complexity and may be redundant nodes that cause performance bottlenecks.

[0069] Calculate node cost: Considering the frequency of label usage and computational complexity, a cost can be calculated for each node in the rule tree. The higher the cost, the greater the impact of the node on performance. Nodes with a cost higher than the preset cost threshold can be treated as redundant nodes.

[0070] Rule tree pruning can mainly include the following operations: 1) Traverse the rule tree: Starting from the root node, traverse the rule tree in a depth-first or breadth-first manner.

[0071] 2) Calculate node cost: Calculate the cost value of each node based on the cost model.

[0072] 3) Prune redundant nodes: For each node whose cost is higher than a certain threshold, determine whether the node can be pruned. If pruning the node does not affect the accuracy of the selection result (for example, the condition of the node can be covered by other nodes), then prune the node.

[0073] 4) Merge overlapping label combinations: For nodes with similar conditions, merge their label combinations to simplify the rule logic.

[0074] Rule tree pruning removes redundant rules and logic, reduces unnecessary computation, and makes the selection process more concise and efficient. For example, a user set that originally required the calculation of multiple complex rules may now only require the calculation of a few key rules after pruning.

[0075] (1.2) Index optimization acceleration.

[0076] During the selection process, users need to be screened based on the conditions contained in the tag rules. Index optimization can accelerate the filtering of these conditions, quickly locate users who meet the conditions, and improve the selection speed.

[0077] An inverted index can be created for frequently queried tags (such as "highly active users"), mapping each tag to a list of user IDs containing that tag. For example, for the "highly active users" tag, the index stores the IDs of all users marked as "highly active users." During the selection process, when filtering based on a specific tag is needed, the system can quickly retrieve the list of user IDs containing that tag directly from the inverted index, without having to traverse the entire user dataset. This greatly improves query speed, reducing the time required to traverse all users to the time required to retrieve results directly from the index.

[0078] Through index optimization acceleration, users who meet the conditions can be quickly located, avoiding full data scanning, reducing the amount of calculation and time consumption, and obtaining the set of users who meet the conditions more quickly during the selection process.

[0079] (1.3) Incremental calculation.

[0080] Master data can be pre-defined. Master data refers to core data shared by multiple applications or processes within a business system, primarily including basic user information, behavioral data, and tags. For example, a user's age, gender, registration date, purchase history, browsing history, and tags calculated from this data (such as "high-spending user" or "potential churn user") are all considered master data. When master data changes, incremental computation recalculates only the affected data, avoiding a full scan, reducing computational effort, and lowering computing resource consumption. This accelerates the update of selection results and improves the real-time nature of selection.

[0081] (1.4) Bitmap operations.

[0082] Crowd selection is essentially a set operation on user sets, such as intersection, union, and difference. Bitmap operations can quickly implement these set operations and improve selection efficiency.

[0083] Specifically, the user set can be converted into the RoaringBitmap format. RoaringBitmap is an efficient bitmap data structure that can map user IDs to bits in the bitmap. The state of the bit (0 or 1) indicates whether the user belongs to the corresponding user set. Then, bitmap operations (such as AND, OR, XOR, etc.) are used to implement fast set operations on the user set.

[0084] Since bitmap operations can complete set operations on large user sets in a very short time, converting the user set into the RoaringBitmap format can leverage the efficiency of bitmap operations to speed up the selection process, greatly shortening the selection operation time. When processing selection tasks with multiple condition combinations, the final user set can be obtained efficiently.

[0085] (1.5) Distributed computing.

[0086] For large-scale user data, the selection task can be split into multiple subtasks and distributed to a distributed cluster (such as a Spark cluster) for parallel execution, so as to fully utilize the computing resources of the distributed cluster to achieve a response within seconds, thereby improving the efficiency of selection.

[0087] After the target users corresponding to the target user tags are circled based on the rule tree, the circled user set needs to be bound with the corresponding user ID list to generate a crowd package, so that the crowd package can be subsequently exported and expanded.

[0088] The user ID list is an important component of the crowd package. When generating a crowd package, the user ID list can be converted into a format (such as CSV, JSON, or encrypted file format) that is compatible with the advertising delivery system and supports encrypted transmission (such as HTTPS), and the crowd package can be expanded to increase the scale of the population covered by the crowd package.

[0089] Generated audience packages can also be optimized, such as removing duplicate users and adjusting audience size. Furthermore, audience packages can be further analyzed and mined based on actual business needs to identify potential target user groups. The optimized audience packages can be provided to sales teams or other data users, allowing them to select appropriate audience packages for advertising trials based on their characteristics and objectives, and adjust advertising strategies based on the feedback from these trials.

[0090] When exporting a demographic package, use the API to export the package in a suitable format (e.g., CSV, JSON, etc.) and push it to the target system (e.g., advertising system, CRM platform, etc.) for subsequent business operations. For example, you can export the demographic package to an advertising system for targeted advertising, or to a BI system for data analysis and decision support.

[0091] When expanding the population package, machine learning algorithms such as clustering algorithms and classification algorithms (such as K-means and random forests) can be used to automatically identify the high-value user groups corresponding to the population package, and the Lookalike algorithm can be used to calculate the similarity based on the characteristics of the seed population (such as the high-value user group) (such as user attributes, interests, behaviors, etc.) to expand the similar population and generate potential target populations.

[0092] In actual application, the crowd package can be updated regularly (such as weekly or monthly) to avoid the invalidation of the crowd package due to changes in user behavior, thereby ensuring the timeliness of the crowd package.

[0093] After pushing the audience package to the target system, delivery strategies can be set based on geography, time, frequency, and other dimensions. Dynamic Creative Optimization (DCO) technology can be used to dynamically display ad creatives based on user profiles. DCO is a programmatic advertising technology that uses real-time data and algorithms to automatically generate and optimize ad content, delivering a personalized ad experience. DCO's core logic dynamically combines ad creatives (e.g., text, images, videos, etc.) based on user characteristics (e.g., location, interests, behavioral history) and context (e.g., time, device, weather) to match audience needs, thereby improving ad effectiveness (e.g., click-through rate and conversion rate).

[0094] In summary, the above-mentioned DMP-based crowd selection method alleviates problems such as data delay, label rigidity, and performance bottlenecks in related technologies, achieves more efficient and accurate crowd selection, and provides core technical support for the digital transformation of corporate marketing.

[0095] On the basis of the above-mentioned crowd selection method based on DMP, an embodiment of the present invention further provides a crowd selection device based on DMP, which can be applied to DMP, see Figure 2 As shown, the device may include the following modules: The feature generation module 202 is used to pre-process the original user data obtained from multiple different data sources and generate user features corresponding to the pre-processed data based on preset rules.

[0096] The tag generation module 204 is used to convert the obtained user features into user tags based on a pre-built tag system, store the obtained user tags in a distributed storage manner, and then display the stored user tags through an operation interface; wherein, the user tags include static tags and dynamic tags.

[0097] The selection module 206 is configured to convert the selection operation for the target user tag displayed on the operation interface into a rule tree, and select the target user corresponding to the target user tag based on the rule tree; wherein the selection result is a target user set consisting of multiple target users.

[0098] The crowd package generation module 208 is used to associate the circled results with the identification information of the corresponding target users to generate a crowd package, and expand the obtained crowd package using a similar crowd expansion algorithm; wherein the identification information includes the unique identification of each target user in the target user set.

[0099] The above-mentioned DMP-based crowd selection device is used to generate user tags by fusing multi-source data, and the selection operation of user tags on the operation interface is converted into a rule tree to realize crowd selection. It can adapt to emerging business scenarios or cross-platform data fusion needs, while improving the efficiency and accuracy of crowd selection.

[0100] The feature generation module 202 can also be used to: extract basic features from the preprocessed data according to the preset rules, and use a preset statistical method to perform statistics on the basic features to generate a first feature, and then use a preset data mining algorithm to analyze the preprocessed data to generate a second feature; analyze the metadata and data distribution of the preprocessed data to generate a third feature; and use the basic features, the first feature, the second feature and the third feature as the user features.

[0101] The above-mentioned dynamic tags may include prediction tags and real-time behavior tags, and the above-mentioned tag system may include a first rule engine, a distributed stream processing engine and a pre-trained machine learning model; based on this, the above-mentioned tag generation module 204 can also be used to: use the first rule engine to generate static tags corresponding to the user features, and use the distributed stream processing engine to generate real-time behavior tags corresponding to the user features, and then use the machine learning model to generate prediction tags corresponding to the user features; form an initial tag combination with the static tags, the real-time behavior tags and the prediction tags, and use a preset tag semantic model to optimize the initial tag combination, and use the tags in the optimized tag combination as user tags.

[0102] The above-mentioned label generation module 204 can also be used to: obtain real-time user behavior data, and use a second rule engine to determine whether the real-time user behavior data meets the preset conditions; if the real-time user behavior data meets the preset conditions, the second rule engine is used to update the user label according to the preset label information corresponding to the preset conditions.

[0103] The rule tree may adopt an abstract syntax tree. Based on this, the selection module 206 may further be configured to: if there are high-frequency query tags in the target user tags, establish an inverted index for the high-frequency query tags in the target user tags; wherein the inverted index includes each high-frequency query tag in the target user tags and identification information of the corresponding target user; prune the rule tree based on a pre-built cost model, and execute the pruned rule tree to obtain the target user set.

[0104] The circle selecting module 206 can also be configured to traverse the pruned rule tree to obtain an initial target user set corresponding to each node of the pruned rule tree for the target user label, and convert each obtained initial target user set into a corresponding bitmap; perform a set operation included in the pruned rule tree on the obtained bitmaps, and convert the obtained set operation result into the target user set.

[0105] The label generating module 204 can also be configured to perform semantic analysis on the initial label combination by using the preset label semantic model to identify redundant labels in the initial label combination; establish a label association graph of the initial label combination based on the redundant labels by using the preset label semantic model; wherein a node in the label association graph represents a label, and an edge between nodes represents an association relationship between labels; determine an optimal label in the initial label combination by using the label association graph by using the preset label semantic model, and form the optimized label combination by using the obtained optimal label.

[0106] The circle selecting device based on DMP provided in the embodiments of the present application has the same implementation principle and technical effects as the circle selecting method based on DMP, and for brevity, the parts not mentioned in the circle selecting device based on DMP can refer to the corresponding content in the circle selecting method based on DMP.

[0107] The embodiments of the present application further provide an electronic device, as shown in the accompanying drawings, which is a structural schematic diagram of the electronic device. Figure 3 The electronic device includes a processor 31 and a memory 30, the memory 30 stores computer executable instructions capable of being executed by the processor 31, and the processor 31 executes the computer executable instructions to implement the circle selecting method based on DMP.

[0108] In the embodiments shown in the accompanying drawings, the electronic device further includes a bus 32 and a communication interface 33, wherein the processor 31, the communication interface 33 and the memory 30 are connected through the bus 32. Figure 3

[0109] ​The memory 30 can include a high-speed random access memory (RAM), and can also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 33 (which can be wired or wireless), and the Internet, a wide area network, a local network, a metropolitan area network, etc. can be used. The bus 32 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 32 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one bidirectional arrow is used to represent the system network element and at least one other network element, but it does not mean that there is only one bus or one type of bus.

[0110] The processor 31 can be an integrated circuit chip with signal processing capability. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 31 or the instructions in the form of software. The processor 31 described above can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage medium is located in the memory, and the processor 31 reads the information in the memory and combines the hardware to complete the steps of the above-mentioned embodiment of the DMP-based crowd selection method.

[0111] An embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the above-mentioned DMP-based crowd selection method. For specific implementation, please refer to the aforementioned method embodiment and will not be repeated here.

[0112] The computer program product of the DMP-based crowd selection method, device, and electronic device provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the DMP-based crowd selection method described in the previous method embodiments. For specific implementation, please refer to the method embodiments and will not be repeated here.

[0113] Unless otherwise specifically stated, the relative steps, numerical expressions and values ​​of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0114] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0115] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0116] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A method for selecting a group of people based on DMP, characterized in that: Applied to data management platforms, including: Preprocess the raw user data obtained from multiple different data sources and generate user features corresponding to the preprocessed data based on preset rules; The obtained user features are converted into user tags based on a pre-built tag system, and the obtained user tags are stored in a distributed storage manner, and then the stored user tags are displayed through an operation interface; wherein the user tags include static tags and dynamic tags; Converting the selection operation for the target user label displayed on the operation interface into a rule tree, and selecting the target users corresponding to the target user label based on the rule tree; wherein the selection result is a target user set consisting of multiple target users; The selection result is associated with the identification information of the corresponding target user to generate a crowd package, and the obtained crowd package is expanded using a similar crowd expansion algorithm; wherein the identification information includes the unique identification of each target user in the target user set.

2. The method for selecting a group of people based on DMP according to claim 1, characterized in that: Generate user features corresponding to preprocessed data based on preset rules, including: Extracting basic features from the preprocessed data according to the preset rules, performing statistics on the basic features using a preset statistical method to generate a first feature, and then analyzing the preprocessed data using a preset data mining algorithm to generate a second feature; analyzing metadata and data distribution of the preprocessed data to generate a third feature; The basic feature, the first feature, the second feature, and the third feature are used as the user features.

3. The method for selecting a group of people based on DMP according to claim 2, characterized in that: The dynamic tags include prediction tags and real-time behavior tags, and the tag system includes a first rule engine, a distributed stream processing engine, and a pre-trained machine learning model; The obtained user features are converted into user tags based on the pre-built tag system, including: Using the first rule engine to generate static tags corresponding to the user features, using the distributed stream processing engine to generate real-time behavior tags corresponding to the user features, and then using the machine learning model to generate predicted tags corresponding to the user features; The static tag, the real-time behavior tag and the predicted tag are combined into an initial tag combination, and the initial tag combination is optimized using a preset tag semantic model, and the tags in the optimized tag combination are used as user tags.

4. The method for selecting a group of people based on DMP according to claim 3, characterized in that: Also includes: Acquire real-time user behavior data, and use a second rule engine to determine whether the real-time user behavior data meets a preset condition; If the real-time user behavior data meets the preset condition, a second rule engine is used to update the user tag according to the preset tag information corresponding to the preset condition.

5. The method for selecting a group of people based on DMP according to claim 1, characterized in that: The rule tree adopts an abstract syntax tree; and selecting a target user corresponding to the target user label based on the rule tree includes: If there is a high-frequency query tag in the target user tag, an inverted index is established for the high-frequency query tags in the target user tag; wherein the inverted index includes each high-frequency query tag in the target user tag and identification information of the target user corresponding thereto; The Rule Tree is pruned based on a pre-built cost model, and the pruned Rule Tree is executed to obtain the target user set.

6. The method for selecting a group of people based on DMP according to claim 5, characterized in that: Executing the pruned Rule Tree to obtain the target user set includes: Traversing the pruned Rule-Tree to obtain an initial target user set whose target user label corresponds to each node of the pruned Rule-Tree, and converting each obtained initial target user set into a corresponding bitmap; The obtained bitmap is subjected to the set operation included in the pruned rule tree, and the obtained set operation result is converted into the target user set.

7. The method for selecting a group of people based on DMP according to claim 3, characterized in that: The initial tag combination is optimized using a preset tag semantic model, including: Performing semantic analysis on the initial tag combination using the preset tag semantic model to identify redundant tags in the initial tag combination; The preset tag semantic model is used to establish a tag association graph of the initial tag combination based on the redundant tags; wherein, in the tag association graph, nodes are used to represent tags and edges between nodes are used to represent associations between tags; The preset tag semantic model is used to utilize the tag association graph to determine the optimal tag in the initial tag combination, and the obtained optimal tags are combined into the optimized tag combination.

8. A crowd selection device based on DMP, characterized in that: Applied to data management platforms, including: The feature generation module is used to pre-process the raw user data obtained from multiple different data sources and generate user features corresponding to the pre-processed data based on preset rules; A tag generation module is used to convert the obtained user features into user tags based on a pre-built tag system, store the obtained user tags in a distributed storage manner, and then display the stored user tags through the operation interface; wherein the user tags include static tags and dynamic tags; a selection module, configured to convert a selection operation for a target user tag displayed on the operation interface into a rule tree, and select target users corresponding to the target user tag based on the rule tree; wherein the selection result is a target user set consisting of multiple target users; A crowd package generation module is used to associate the selection results with the identification information of the corresponding target users to generate a crowd package, and expand the obtained crowd package using a similar crowd expansion algorithm; wherein the identification information includes the unique identification of each target user in the target user set.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the DMP-based crowd selection method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions prompt the processor to implement the DMP-based crowd selection method described in any one of claims 1 to 7.

Citation Information

Cited By

  • Method and system for repeated analysis of customer group diagnosis

    CN121660721A

  • A method and system for customer diagnostic reanalysis

    CN121660721B

  • Multi-target reinforcement learning recommendation method and system for long-term user participation degree

    CN121980083A