Efficient user tag generation method and system, electronic equipment and storage medium

By generating a tag tree and performing hierarchical parallel calculations, the performance and real-time problems of user tag generation under massive data in the existing technology are solved, and efficient and accurate tag calculations are achieved.

CN120180286APending Publication Date: 2025-06-20中信证券股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510248163.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When existing user tag generation methods face massive TB data, they have high hardware costs, insufficient data processing capabilities, poor real-time performance and low computing efficiency.

Method used

By identifying the dependencies between each tag, a tag tree is generated, and according to the level of the tag tree, the tag calculation is performed step by step from the leaf node, the nodes at the same level are calculated in parallel, and the parent node performs tag calculation based on the calculation results of all children.

Benefits of technology

It significantly improves label computing performance, reduces hardware costs, improves data processing capabilities and real-time performance, and reduces the risk of calculation errors and data inconsistencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180286A_ABST
    Figure CN120180286A_ABST
Patent Text Reader

Abstract

The invention discloses an efficient user tag generation method and system, electronic equipment and a storage medium. The method comprises the following steps: identifying a dependency relationship among a plurality of labels according to a label rule of each label, and generating a label tree; according to the hierarchy of the label tree, performing label calculation upwards step by step from leaf nodes; wherein the same-level nodes carry out parallel label calculation, and the father node carries out label calculation according to calculation results of all the child nodes. According to the label generation scheme provided by the embodiment of the invention, the dependency relationship among the plurality of labels is identified according to the label rule of each label, and the label tree is generated; according to the levels of the label tree, label calculation is carried out step by step upwards from leaf nodes, and labels at the same level are subjected to parallel calculation, so that the label calculation performance can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the technical field of user profiling, and particularly to an efficient user tag generation method, system, electronic device, and storage medium. Background Art

[0002] In the digital age, data has become one of the most valuable assets of enterprises. With the development of Internet technology and the rise of big data, the user information that enterprises can collect is becoming increasingly rich, including but not limited to users' basic information, behavior habits, consumption records, social interactions, etc. In order to make more effective use of this data and improve marketing efficiency and user experience, user profiling systems have emerged.

[0003] Tag calculation is the core link in constructing user profiles. It is a process of quantifying and classifying information such as users' behaviors, preferences, and attributes through data analysis techniques. The purpose of this process is to extract valuable information from a large amount of user data to form an accurate description of user characteristics for more effective market segmentation, personalized recommendation, precision marketing, etc. Tag calculation based on massive data places increasingly high requirements on the execution performance of the profiling system. Summary of the Invention

[0004] This application provides an efficient user tag generation method, system, electronic device, and storage medium. According to the tag rules of each tag, the dependency relationships between multiple tags are identified to generate a tag tree. According to the levels of the tag tree, tag calculation is performed step by step upward from the leaf nodes, and sibling tags are calculated in parallel, which can significantly improve the tag calculation performance.

[0005] An embodiment of this application provides an efficient user tag generation method, including:

[0006] According to the tag rules of each tag, the dependency relationships between multiple tags are identified to generate a tag tree;

[0007] According to the levels of the tag tree, tag calculation is performed step by step upward from the leaf nodes;

[0008] Among them, the sibling nodes perform parallel tag calculation, and the parent node performs tag calculation based on the calculation results of all child nodes.

[0009] An embodiment of this application also provides an efficient user tag generation system, including:

[0010] A tag tree generation module, configured to identify the dependency relationships between multiple tags according to the tag rules of each tag and generate a tag tree;

[0011] A tag calculation module, configured to perform tag calculation step by step upward from the leaf nodes according to the levels of the tag tree;

[0012] Among them, the sibling nodes perform parallel label calculations, and the parent node performs label calculations based on the calculation results of all child nodes.

[0013] The embodiment of the present application also provides an electronic device, including:

[0014] One or more processors;

[0015] A storage device for storing one or more programs,

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the efficient user label generation method as described in any embodiment of the present application.

[0017] The embodiment of the present application also provides a computer storage medium, in which a computer program is stored, and the computer program is configured to execute the efficient user label generation method as described in any embodiment of the present application when running.

[0018] Other features and advantages of the present application will be described in the subsequent description, and some of them will become obvious from the description, or be understood by implementing the present application. Other advantages of the present application can be realized and obtained through the solutions described in the description and the drawings. Description of the Drawings

[0019] The drawings are used to provide an understanding of the technical solutions of the present application, and constitute a part of the description. Together with the embodiments of the present application, they are used to explain the technical solutions of the present application, and do not constitute a limitation to the technical solutions of the present application.

[0020] Figure 1 It is a flowchart of an efficient user label generation method in an embodiment of the present application;

[0021] Figure 2 It is an example diagram of a label tree in an embodiment of the present application;

[0022] Figure 3 It is another example diagram of a label tree in an embodiment of the present application;

[0023] Figure 4 It is a schematic diagram of another efficient user label generation method in an embodiment of the present application;

[0024] Figure 5 It is a structural diagram of an efficient user label generation system in an embodiment of the present application;

[0025] Figure 6 It is a structural diagram of a user portrait system in an embodiment of the present application;

[0026] Figure 7This is another structural diagram of the user portrait system in the embodiments of the present application. Detailed implementation manners

[0027] The present application describes multiple embodiments, but the description is exemplary rather than restrictive, and it is obvious to those of ordinary skill in the art that there can be more embodiments and implementation solutions within the scope of the embodiments described in the present application. Although many possible feature combinations are shown in the drawings and discussed in the detailed implementation manners, many other combination ways of the disclosed features are also possible. Unless specifically restricted, any feature or element of any embodiment can be combined with any other feature or element in any other embodiment, or can replace any other feature or element in any other embodiment.

[0028] The present application includes and contemplates combinations with features and elements known to those of ordinary skill in the art. The embodiments, features, and elements disclosed in the present application can also be combined with any conventional features or elements to form a unique inventive solution defined by the claims. Any feature or element of any embodiment can also be combined with features or elements from other inventive solutions to form another unique inventive solution defined by the claims. Therefore, it should be understood that any feature shown and / or discussed in the present application can be implemented alone or in any appropriate combination. Therefore, the embodiments are not subject to other limitations except those made according to the appended claims and their equivalent replacements. In addition, various modifications and changes can be made within the scope of protection of the appended claims.

[0029] In addition, when describing representative embodiments, the specification may have presented the method and / or process as a specific sequence of steps. However, to the extent that the method or process does not depend on the specific sequence of the steps described herein, the method or process should not be limited to the specific sequence of steps described. As will be understood by those of ordinary skill in the art, other step sequences are also possible. Therefore, the specific sequence of steps set forth in the specification should not be construed as a limitation on the claims. In addition, the claims directed to the method and / or process should not be limited to performing their steps in the order written, and those skilled in the art can easily understand that these orders can be varied and still remain within the spirit and scope of the embodiments of the present application.

[0030] In the big data era, the rapid growth of user data has made accurate user portraits and label generation an important task for enterprises to conduct market analysis and user behavior prediction. When facing TB-level massive data, the existing traditional user label generation methods have two problems: one is that the hardware cost for second-level return is too high, and the other is that after layers of data processing, problems such as loss of original information, insufficient data processing capacity, poor real-time performance, and low calculation efficiency are likely to occur.

[0031] An embodiment of the present application provides an efficient user tag generation method, as Figure 1 shown, including:

[0032] Step 110: Identify the dependency relationships between multiple tags according to the tag rules of each tag, and generate a tag tree;

[0033] Step 120: Perform tag calculations level by level from the leaf nodes upwards according to the levels of the tag tree;

[0034] Among them, parallel tag calculations are performed for the sibling nodes, and the parent node performs tag calculations based on the calculation results of all child nodes.

[0035] In some exemplary embodiments, step 110 includes: The identifying the dependency relationships between multiple tags according to the tag rules of each tag and generating a tag tree includes:

[0036] Respectively determine the data query components corresponding to each tag rule, and extract the query fields and filtering conditions in the query components;

[0037] Perform correlation analysis based on the query fields and filtering conditions corresponding to all tags in the group to determine the dependency relationships between the tags;

[0038] Use the tags that do not depend on other tags as leaf nodes to generate the tag tree.

[0039] A tag rule is a set of standards or guidelines used to describe or define a tag. In a user portrait system, a tag rule is the logic and conditions used to determine how to assign specific tags based on a user's behavior, attributes, or other characteristics. These rules can be based on simple conditional judgments or complex algorithmic models. In some exemplary embodiments, a tag rule is also referred to as a tag definition.

[0040] In some exemplary embodiments, a tag rule includes: conditions and / or condition weights. For example, for the tag: high-value active user, the rules include: Condition 1: The total asset value exceeds XXX; Condition 2: The number of transactions within 6 months exceeds 10.

[0041] In some exemplary embodiments, for multiple tags, a rule engine is used to respectively determine the data query components corresponding to each rule, and then extract the query fields and filtering conditions; based on the query fields and filtering conditions of each component, perform correlation analysis to determine the dependency relationships between the tags. The rule engine realizes the identification of dependency relationships by deeply analyzing the data lineage relationships between the tag rules. In some exemplary embodiments, by parsing the calculation SQL for tag calculations, the other tag fields referenced therein are extracted, thereby generating a dependency relationship graph between multiple tags. For example, for tags M1 to M6, a tag tree is generated, as Figure 2As shown. For another example, labels N1-N12 generate a label tree, such as Figure 3 As shown in the figure, the corresponding label levels are recorded as L0, L1, ..., Ln, where L0 is the leaf layer and does not depend on other labels.

[0042] Among them, step 120 includes: starting from the leaf node, performing label calculations step by step upward; after the labels of the same level are calculated in parallel, the calculation results are merged and written into the same label table.

[0043] As can be seen, label calculation starts from the leaf node, and the same-level labels are calculated in parallel, which can significantly improve the efficiency of label calculation. On the basis of layered parallel execution of label calculation, a corresponding label table is created for each level, and the results (label data) of the label calculation at the same level are merged and written into the same label table, ensuring that all data in the same label table belongs to the same level, which is helpful for data organization and management and subsequent analysis and utilization.

[0044] For example, the label of L0 layer does not depend on any other label and is the basis of calculation; the label of L1 layer can only depend on the label of L0 layer at most; the label of L2 layer can only depend on the label of L0 and L1 layer at most, and so on. This hierarchical structure ensures that all the required pre-data are ready every time a label is calculated, thereby improving the accuracy of calculation. Figure 2 As shown, labels M1, M2, and M3 are calculated in parallel, followed by labels M4 and M5, and then label M6 is calculated.

[0045] For example, L0 layer: Label 1: user registration information, Label 2: transaction and asset information, Label 3: activity information; L1 layer: Label 4: high-value active users, label rules: total asset value exceeds 10, number of transactions in the past 6 months is greater than 10, it can be seen that label 4 depends on L0 layer labels 2 and 3; L1 layer: Label 5: high-frequency small transaction users, label rules: single transaction amount is less than 500 yuan, average daily transaction number is greater than 2, it can be seen that label 5 depends on L0 layer labels 2 and 3. Then calculate the L0 layer label first, and then further calculate the L1 layer label based on the L0 layer calculation result.

[0046] In some exemplary embodiments, the method further includes: when the tag data corresponding to the tag is updated, automatically updating the downstream tag according to the overall dependency relationship, combining Figure 2 ,The change of label M4 at L1 layer will drive the linkage of label M6 at L2 layer.

[0047] As can be seen, through explicit hierarchical structure and dependency management, it is ensured that each calculation is based on accurate source data. This method avoids calculation errors and data inconsistency problems that may occur in unordered calculations, thereby improving the accuracy of label calculation. There is no dependency relationship between labels at the same level, and they can be calculated concurrently. This concurrent calculation ability greatly reduces the data processing pressure and improves the overall operation performance. Through concurrent calculation, the system can complete the calculation and update of a large number of labels in a shorter time. After the concurrent calculation of labels at the same level is completed, centralized data merging is performed. This data merging process not only simplifies the data processing flow but also improves the efficiency of data processing. By optimizing the data merging algorithm and strategy, the system can integrate and process the calculation results faster, thereby further improving the overall performance.

[0048] In some exemplary embodiments, label calculation is performed for each label according to the following method:

[0049] Parse the label rule of the label to determine the corresponding query statement SQL, where the query statement includes: a SELECT clause and a WHERE condition clause;

[0050] Extract the target query fields in the SELECT clause to form a column requirement set;

[0051] Extract the filtering fields and conditional logic in the WHERE condition clause;

[0052] Generate a query plan corresponding to the query statement SQL, where the query plan includes: a projection node (ProjectNode), a filter node (FilterNode), and a table scan node (TableScanNode); among them, the projection node is determined according to the column requirement set, and the filter node is determined according to the filtering fields and conditional logic;

[0053] Push down the projection node (ProjectNode) to the table scan node (TableScanNode), and push down the filter node (FilterNode) to the table scan node (TableScanNode);

[0054] Execute the query plan to complete the label calculation.

[0055] Among them, extract the columns specified in the SELECT clause of the query statement or the fields explicitly used in the subquery to form a column requirement set. Parse the expressions in the WHERE clause or other filtering conditions to extract the relevant fields and conditional logic.

[0056] Example: event data for the columns of event_content, event_time, and user_id, and for the behavior where userid = "User 1". The filtering condition is userid = "User 1", and the relevant field is userid.

[0057] In some exemplary embodiments, the query optimizer generates a query plan. The initial plan includes a projection node (ProjectNode) (column selection) and a filter node (FilterNode); the projection node (ProjectNode) is called column selection, and the filter node (FilterNode) is also called row filtering logic.

[0058] In some exemplary embodiments, pushing the projection node (ProjectNode) down to the table scan node (TableScanNode) includes: modifying the scan logic of the table scan node (TableScanNode) to read only the columns in the column requirement set and avoid reading redundant data.

[0059] For example, in the scan phase, only read the columns of event_content, event_time, and user_id and skip the unused columns.

[0060] It can be understood that pushing the projection node (ProjectNode) down to the table scan node (TableScanNode) means pushing the column requirement set down to the table scan node (TableScanNode) to complete column filtering directly in the data scan phase. In some exemplary embodiments, pushing the projection node (ProjectNode) down to the table scan node (TableScanNode) includes: modifying the scan logic of the table scan node (TableScanNode) to read only the columns in the column requirement set and avoid reading redundant data.

[0061] In some exemplary embodiments, pushing the filter node (FilterNode) down to the table scan node (TableScanNode) includes: pushing the filter condition logic down to the table scan node (TableScanNode) and using the indexing ability of the data source to complete row filtering, which is also called index filtering.

[0062] Among them, index filtering includes: quickly locating the target data blocks that meet the conditions through an ordered sparse index (such as a B+ tree or skip list index based on userid) and skipping the irrelevant data blocks. When performing fine-grained filtering, mechanisms such as Bloom filters can be used to further narrow the result range. Example: For the condition userid > 100, use the index to directly locate the data block range where userid is greater than 100.

[0063] In some exemplary embodiments, executing the query plan to complete label calculation includes: executing optimized scan logic in a TableScanNode.

[0064] It can be seen that when executing the query plan, the TableScanNode only reads the required columns according to the column requirement set, reducing disk I / O and transmission overhead; when scanning data blocks, it directly skips the rows that do not meet the conditions according to the pushed-down filtering conditions. The rows and data blocks located by the index are directly filtered out of irrelevant data at the storage level; the result data is passed to the upper-level node: the filtered data is passed to the upper-level node (such as a JoinNode or AggregateNode) for further processing. After the ProjectNode is pushed down, only the necessary columns are scanned; after the FilterNode is pushed down, only the rows that meet the conditions are scanned. This reduces the loading and transmission of redundant data, reducing memory and network overhead; in some exemplary embodiments, an ordered sparse index is used to complete data filtering at the storage level, greatly improving the filtering efficiency.

[0065] After pushing down the logic of the ProjectNode to the TableScanNode, column filtering can be completed at the data source level, avoiding the loading and transmission of redundant data and saving computing resources. When the logic of the FilterNode is pushed down to the TableScanNode, data filtering can be directly completed using index or partition information during the scanning phase, avoiding irrelevant data from being loaded into memory or passed to subsequent nodes. Through the combination of column selection and index filtering, data reading is directly optimized at the storage layer, reducing memory occupancy and network transmission costs.

[0066] In some exemplary embodiments, executing the query plan includes: executing the query plan according to one of the following bucketing strategies: event virtual bucketing, user virtual bucketing, and time unit bucketing.

[0067] In some exemplary embodiments, label calculation is performed based on a user's behavioral data and attribute data. These data cover diverse behavioral data (such as transaction records, login behaviors, click events, browsing histories, etc.) and detailed user attribute information (such as gender, age group, occupation category, etc.). The data models corresponding to the behavioral data and attribute data are respectively called the behavioral model and the attribute model.

[0068] For example, the behavioral model includes the following:

[0069]

[0070] Each record represents the occurrence of a behavior.

[0071] The attribute model encompasses all the characteristics of a user in the form of a single record. In some exemplary embodiments, it uses Kudu for data storage and can exhibit excellent query performance even in the scenario of extremely long columns. The attribute model only records the final fixed state of user attributes, and the information of each user is only integrated into one record.

[0072] Among them, the behavior model relies on a custom file format for storage. It uses the combination of user ID, specific date, and behavior type to construct an index system. In this mode, each data record precisely depicts a specific behavior performed by a certain user at a specific moment. To ensure the comprehensiveness and practicality of the data, the behavior model is designed to necessarily include four core components: user ID, timestamp of the behavior occurrence, behavior type, and related behavior attributes, while the location where the behavior occurs is regarded as non-essential information. To further improve the efficiency of data retrieval, each type of behavior and its related attributes are separately designed as a column for storage. The attribute model relies on a format similar to Key-Value for storage and uses the user ID as the primary key to construct an index. Each attribute of the user is assigned to an independent column, and the cross-attribute association query function is realized through view technology. Such a design not only makes the data storage more structured but also greatly facilitates subsequent data analysis and mining work.

[0073] In some exemplary embodiments, a query engine is used to execute a (dynamic) bucketing strategy to connect Kudu and Hive to query behavior data and attribute data. For example, for two tables (such as orders and orders_item) that are bucketed (bucketing) according to the same field (such as orderid) and have the same number of buckets, when joining (Join, a programming language string that returns a string created by concatenating many substrings included in an array) through orderid, since the same orderid in the two tables is assigned to the same bucket, independent join and aggregation calculations can be performed (refer to the partition process of MapReduer). In this way, whenever the calculation of the data in a bucket is completed, the memory occupied by this bucket can be immediately released. Therefore, by controlling the number of buckets processed in parallel, the memory occupancy can be limited.

[0074] It can be seen that adopting the bucketing strategy to execute label calculation can reduce memory occupancy. The optimized memory occupancy = original memory occupancy / (number of buckets of the table * number of buckets processed in parallel). Implementing bucketing query only requires loading the data of this bucket into memory and releasing the memory occupancy after the calculation is completed, which can significantly reduce memory occupancy.

[0075] For example, if the time unit is a day, then it can be bucketed by day; or, if the time unit is a week, then it is bucketed by week. Setting different events according to needs to bucket the data is called event virtual bucketing; for example, an online shopping platform buckets data according to events associated with online shopping behaviors, such as the Double 11 event, the Double 12 event, the New Year's Day promotion event, the Spring Festival promotion event, etc. Bucketing according to user characteristics is called user virtual bucketing; for example, bucketing by the region to which the user's registration location belongs, bucketing by user age, bucketing by user occupation, etc.

[0076] In some exemplary embodiments, executing the query plan includes: using the query engine Presto+ to execute the query plan; the query plan is generated by the query engine Presto+.

[0077] In some exemplary embodiments, the query engine Presto+ connects to at least one data source through at least one data source connector (DB-Link);

[0078] Wherein, the data connector connects to the corresponding data source in one of the following ways: HTTP (Hyper Text Transfer Protocol), JDBC (Java Database Connectivity), and file read-write input-output IO; the data source includes one or more of the following types: relational database, NoSQL database, file storage, and data stream.

[0079] In some exemplary embodiments, the data source includes one or more of the following: Hadoop HDFS, Amazon S3, MySQL, PostgreSQL, oracle, hive, iceberg, hudi, Mongodb, elasticsearch, clickhouse, kudu, Cassandra, etc.

[0080] All data source connectors (DB-Link) implement a standard interface to connect to the query engine Presto+ to ensure the unified access of the query engine to different data sources. Different data source connectors are respectively connected to different types of data sources to perform operations such as initial connection, query execution, and data source metadata acquisition. By customizing the data source connector, users can quickly connect different data sources to the system according to different business requirements and perform flexible queries.

[0081] For example, as Figure 4As shown, the query engine Presto+ uses a standard interface to connect to data sources through a data source connector (DB-Link), including real-time data sources, a common data model, and data federation. It can perform label calculations based on real-time stream data, historical data in a common database, and third-party reference data from data federation. These data sources collect data from various edge / end nodes. For example, these edge / end nodes include: WEB applications, smart terminal APPs, mini-programs, and various Internet of Things applications, etc.

[0082] The embodiment of the present application also provides an efficient user label generation system, as Figure 5 shown, including:

[0083] A label tree generation module 510, configured to identify the dependency relationships between multiple labels according to the label rules of each label and generate a label tree;

[0084] A label calculation module 520, configured to perform label calculations level by level from the leaf nodes upwards according to the levels of the label tree;

[0085] Among them, parallel label calculations are performed for the sibling nodes, and the parent node performs label calculations based on the calculation results of all child nodes.

[0086] In some exemplary embodiments, the efficient user label generation system further includes: a label definition module, configured to provide a visual configuration interface to configure label rules to generate corresponding labels.

[0087] In some exemplary embodiments, the label tree generation module 510 is configured to respectively determine the data query components corresponding to each label rule, extract the query fields and filtering conditions in the query components; perform correlation analysis based on the query fields and filtering conditions corresponding to all labels in the group to determine the dependency relationships between the labels; use the labels that do not depend on other labels as leaf nodes to generate the label tree.

[0088] In some exemplary embodiments, the label calculation module 520 is configured to perform label calculations level by level from the leaf nodes upwards; after parallel calculations of sibling labels, merge the calculation results and write them into the same label table.

[0089] Among them, for each label, label calculations are performed according to the following method:

[0090] Parse the label rule of the label to determine the corresponding query statement SQL, and the query statement includes: a SELECT clause and a WHERE condition clause;

[0091] Extract the target query fields in the SELECT clause to form a column requirement set;

[0092] Extract the filtering fields and conditional logic in the WHERE condition clause;

[0093] Generate a query plan corresponding to the SQL query statement, where the query plan includes: a projection node, a filtering node, and a table scan node; among them, the projection node is determined according to the set of column requirements, and the filtering node is determined according to the filtering fields and conditional logic;

[0094] Push the projection node down to the table scan node, and push the filtering node down to the table scan node;

[0095] Execute the query plan to complete the label calculation.

[0096] The embodiment of the present application also provides a user portrait system, as Figure 6 shown, the label tree generation module 510 and the label calculation module 520 construct a label calculation layer, and the portrait system further includes: a label application layer and a data storage layer. The data storage layer stores the behavior data and user data from each data source, as well as the label data obtained by label calculation. The label application layer provides application functions or application interfaces based on the label data.

[0097] For example, as Figure 7 shown, source data is obtained from multiple data sources and saved in corresponding components in the data storage layer by means of ETL (Extract Transform Load, data extraction, transformation, and loading) import, API extraction, Streaming stream extraction, etc., such as: Kudu, HDFS, Kafka, MangoDb, etc. The label calculation layer connects to the data storage layer through a data DB-Link to obtain behavior data and attribute data, executes scheduling and label calculation, saves the label calculation results in the data storage layer, and provides a label data query interface upward (to the label application layer) to output the label data required by each application.

[0098] The present application also provides an electronic device, including:

[0099] One or more processors;

[0100] A storage device for storing one or more programs,

[0101] When the one or more programs are executed by the one or more processors, the one or more processors implement the efficient user label generation method as described in any embodiment of the present application.

[0102] The embodiment of the present disclosure also provides a computer storage medium, in which a computer program is stored, and the computer program is configured to execute the efficient user label generation method as described in any embodiment of the present application when running.

[0103] In the user label generation method provided by the embodiments of the present application, based on the predefined label rules, the system will automatically analyze and determine the dependency relationships between various labels, and then perform hierarchical parallel processing on the label calculation tasks according to these dependency relationships, ensuring that the tasks at each level are based on the tasks at their lower levels, thereby forming a clear and orderly task hierarchy. The parallel calculation of labels at the same level greatly improves the label calculation efficiency. In some exemplary embodiments, flexible data bucketing is adopted for label calculation, which can significantly reduce the memory occupancy required for label calculation and improve the calculation efficiency. In the label calculation process, the logical push-down method is adopted to push the calculation as much as possible to the data source layer, which can reduce the data transmission volume in the label calculation process and improve the system performance.

[0104] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware implementation, the division of the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation. Some or all components can be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or a non-transitory medium) and a communication medium (or a transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cassette, tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium generally includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

Claims

1. An efficient user tag generation method, characterized in that: include: According to the label rules of each label, identify the dependency relationship between multiple labels and generate a label tree; According to the hierarchy of the label tree, label calculation is performed from the leaf nodes upwards step by step; The nodes at the same level perform label calculation in parallel, and the parent node performs label calculation based on the calculation results of all child nodes.

2. The efficient user tag generation method according to claim 1, characterized in that: The step of identifying the dependency relationship between multiple tags according to the tag rules of each tag and generating a tag tree includes: Determine the data query component corresponding to each tag rule respectively, and extract the query field and filter condition in the query component; Perform correlation analysis based on the query fields and filter conditions corresponding to all tags in the group to determine the dependencies between tags; The label tree is generated by taking labels that are not dependent on other labels as leaf nodes.

3. The efficient user tag generation method according to claim 1 or 2, characterized in that: The label calculation is performed step by step from the leaf node upwards according to the level of the label tree, including: Label calculation is performed from the leaf node upwards step by step; After the same-level tags are calculated in parallel, the calculation results are merged and written into the same tag table.

4. The efficient user tag generation method according to claim 1 or 2, characterized in that: For each label, the label calculation is performed according to the following method: Parse the tag rule of the tag to determine the corresponding query statement SQL, where the query statement includes: a SELECT clause and a WHERE condition clause; Extract the target query fields in the SELECT clause to form a column requirement set; Extract the filter fields and conditional logic in the WHERE conditional clause; Generate a query plan corresponding to the query statement SQL, the query plan including: a projection node, a filter node and a table scan node; wherein the projection node is determined according to the column requirement set, and the filter node is determined according to the filter field and conditional logic; Pushing the projection node down to the table scan node, and pushing the filter node down to the table scan node; The query plan is executed to complete the label calculation.

5. The efficient user tag generation method according to claim 4, characterized in that: The executing the query plan includes: The query plan for the label calculation phase is executed according to one of the following bucketing strategies: Event virtual bucketing, user virtual bucketing, and time unit bucketing.

6. The efficient user tag generation method according to claim 4, characterized in that: The executing the query plan includes: using a query engine Presto+ to execute the query plan; the query plan is generated by the query engine Presto+.

7. The efficient user tag generation method according to claim 6, characterized in that: The query engine Presto+ is connected to at least one data source via at least one data source connector; The data connector connects to the corresponding data source in one of the following ways: Hypertext Transfer Protocol HTTP, Java Database Connection JDBC and file read / write input / output IO; the data source includes one or more of the following types: relational database, NoSQL database, file storage and data stream.

8. An efficient user tag generation system, characterized in that: include: A tag tree generation module is configured to identify dependencies between multiple tags according to tag rules of each tag and generate a tag tree; A label calculation module is configured to perform label calculations from leaf nodes upwards according to the hierarchy of the label tree; The nodes at the same level perform label calculation in parallel, and the parent node performs label calculation based on the calculation results of all child nodes.

9. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the efficient user tag generation method as described in any one of claims 1-7.

10. A computer storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the efficient user tag generation method according to any one of claims 1 to 7 when running.