A method for constructing a user grouping rule engine based on Doris

CN116628020BActive Publication Date: 2026-09-29成都旺小宝科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310180794.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-28
Publication Date
2026-09-29
Estimated Expiration
2043-02-28

AI Technical Summary

Technical Problem

[0004]为解决上述问题,本发明提供了一种基于Doris构建用户分群规则引擎的方法,组件上采用Doris的Bitmap结构来存储计算分群结果,使得企业在降本增效的前提下,能够更加精准的捕获人群,同时解决分群规则引擎成本高,效率低的问题

Benefits of technology

[0029]本发明使用Doris来存储数据,不需要引入沉重的大数据组件,开发人员可根据合适的语言去处理数据进行存储,以及结合Doris内置功能去解析构建规则引擎,灵活性高,降低了构建成本、开发人员的学习成本,同时通过配置规则关系,实现自由组合的分群规则能够大大提高精准度,最后将对应人群采用聚合方式存储起来,资源占用更小,维度更加方便,实现了降本增效。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116628020B_ABST
    Figure CN116628020B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on the method for constructing user group rule engine of Doris, comprising: S1: extraction user basic data and business data, according to the group dimension needed, with user unique ID as unit storage user label table;S2: according to group dimension and behavior data, construct filter item, and configure corresponding enumeration data;S3: according to filter item, judge corresponding basic data type, configure corresponding logic operation, and make mapping relationship in combination with Doris built-in function;S4: configuration rule relationship, according to demand, formulate group rule;S5 will be established group rule into SQL;S6: based on the SQL of group rule conversion, query people who meet the rule, and people who meet the rule are stored in Bitmap type field.The application component is stored in the bitmap structure of Doris to store the calculation group result, solve the problem that group rule engine is high in cost and low in efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method for building a user segmentation rule engine based on Doris. Background Technology

[0002] With the increasing popularity of big data and its widespread use in the internet and other fields, user segmentation is widely applied across various platforms to accurately recommend products to users based on their preferences. However, the components required for current segmentation result parsing engines are too cumbersome, especially when dealing with large amounts of data.

[0003] With current technology, a complete set of big data components and professional big data developers are needed to support rule calculations on various basic and behavioral data of users in order to accurately capture user intentions. However, the construction, procurement, and hiring of professional personnel result in excessively high costs and low efficiency, which, in the current environment, cannot meet the needs of cost reduction and efficiency improvement. Summary of the Invention

[0004] To address the aforementioned issues, this invention provides a method for building a user segmentation rule engine based on Doris. The component uses Doris's Bitmap structure to store the calculated segmentation results, enabling enterprises to capture user groups more accurately while reducing costs and increasing efficiency. This also solves the problems of high cost and low efficiency of segmentation rule engines.

[0005] This invention provides a method for building a user segmentation rule engine based on Doris, the specific technical solution of which is as follows:

[0006] S1: Extract basic user data and business data, and store user tag tables according to the required grouping dimensions, with each user's unique ID as the unit;

[0007] S2: Construct filter items based on cluster dimensions and behavioral data, and configure corresponding enumeration data based on existing filter items;

[0008] S3: Based on the filter items, determine the corresponding basic data type, configure the corresponding logical operation based on the corresponding basic data type, and make a mapping relationship based on the logical operation and Doris built-in functions;

[0009] S4: Configure rule relationships and formulate grouping rules according to requirements;

[0010] S5 converts the established clustering rules into SQL;

[0011] S6: By calling the SQL obtained based on the grouping rules, query the people who meet the rules, and store the people who meet the rules in a Bitmap type field to create a storage data table.

[0012] Furthermore, in step S1, the user tag table includes three fields: a unique tag identifier, a tag value, and a unique user ID.

[0013] Furthermore, in step S2, filter items are constructed based on the cluster dimension, and the corresponding enumeration data is configured, as follows:

[0014] Use the unique identifier of the label as a filter option;

[0015] Configure the corresponding enumeration value for the unique identifier of the label.

[0016] Furthermore, the unique identifier of the label corresponds to at least one enumerated value.

[0017] Furthermore, in step S2, filter items are constructed based on behavioral data, and corresponding enumeration data is configured, as follows:

[0018] Construct a behavior data table, which includes two fields: behavior name and behavior unique identifier;

[0019] Use the unique identifier of the behavior as a filtering option;

[0020] Configure a unique identifier for the behavior and its corresponding enumeration value.

[0021] Furthermore, the behavior unique identifier corresponds to at least one enumeration type, and the enumeration type corresponds to at least one enumeration value.

[0022] Furthermore, both the user tag table and the behavior data table are built and stored based on Doris.

[0023] Furthermore, the mapping table includes four fields: data type, operator, key, and mapping relationship.

[0024] Further, step S5, the specific process is as follows:

[0025] S501: Create a cluster task, select filter items, corresponding enumeration values, corresponding logical operations and rule relationships, and combine them into a long JSON;

[0026] S502: Based on the long JSON, the above-configured filter items, corresponding enumeration values, corresponding logical operation relationships, and rule relationships are parsed into a single SQL statement, ultimately yielding the SQL statements corresponding to all rules.

[0027] Furthermore, the stored data table includes three fields: a unique ID for the subgroup, a subgroup name, and a subgroup result, where the value corresponding to the subgroup result is the user's unique ID.

[0028] The beneficial effects of this invention are as follows:

[0029] This invention uses Doris to store data, eliminating the need for heavy big data components. Developers can process and store data using appropriate languages ​​and combine Doris' built-in functions to parse and build a rule engine, offering high flexibility and reducing construction costs and learning costs for developers. Furthermore, by configuring rule relationships, freely combinable grouping rules can greatly improve accuracy. Finally, the corresponding groups are stored in an aggregated manner, resulting in lower resource consumption, more convenient dimensioning, and cost reduction and efficiency improvement. Attached Figure Description

[0030] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation

[0031] The technical solutions in the embodiments of the present invention are clearly and completely described in the following description. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0032] In the description of the embodiments of the present invention, it should be noted that the indicated orientation or positional relationship is based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of the invention is conventionally placed during use, or the orientation or positional relationship in which those skilled in the art conventionally understand it during use. This is only for the convenience of describing the present invention and simplifying the description, and is not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of the present invention. Furthermore, the terms "first" and "second" are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0033] In the description of the embodiments of the present invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set" and "connection" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.

[0034] Example 1

[0035] Embodiment 1 of the present invention discloses a method for building a user segmentation rule engine based on Doris, such as... Figure 1 As shown, the specific steps are as follows:

[0036] Extract basic user data and business data, and store user tag tables according to the required grouping dimensions, with each user's unique ID as the unit;

[0037] In this embodiment, the user tag table is a Doris regular table, which stores three fields: tag unique identifier (tag name), tag value, and user unique ID.

[0038] The user tag table (tdm_user_tag), in this embodiment, is illustrated using the tag unique identifier (tag_type) as gender and city, and its specific structure is shown in Table 1 below:

[0039] Table 1: User Tag Table

[0040]

[0041] Filter options are constructed based on cluster dimensions and behavioral data, and corresponding enumeration data is configured based on existing filter options;

[0042] Filtering options are constructed based on the clustering dimension, and corresponding enumeration data is configured, as follows:

[0043] Use the unique identifier of the label as a filter option;

[0044] Configure the corresponding enumeration value for the unique identifier of the label.

[0045] In this embodiment, the unique identifier of the label corresponds to at least one enumerated value.

[0046] As shown in Table 2 below:

[0047] Table 2: Tag enumeration value table stored based on tag unique identifier

[0048] gender male gender female city Chengdu city Xi'an

[0049] In this embodiment, the unique identifier of the label "gender" is set with enumeration values ​​"male" and "female", and the unique identifier of the label "city" is set with enumeration values ​​"Chengdu" and "Xi'an". Specifically, the unique identifier of the label and the corresponding enumeration values ​​are not specifically limited here.

[0050] Filter options are constructed based on behavioral data, and corresponding enumeration data is configured, as follows:

[0051] Construct a behavior data table, which includes two fields: behavior name and behavior unique identifier;

[0052] In this embodiment, the behavioral data table is built and stored based on Doris;

[0053] As shown in Table 3 below:

[0054] Table 3: Behavioral Data Table

[0055] Access Page content_read Forwarding Page content_share

[0056] Use the unique identifier of the behavior as a filtering option;

[0057] Configure a unique identifier for the behavior and its corresponding enumeration value.

[0058] In this embodiment, the behavior unique identifier corresponds to at least one enumeration type, and the enumeration type corresponds to at least one enumeration value.

[0059] As shown in Table 4 below:

[0060] Table 4: Table of tag enumeration values ​​stored based on behavior-unique identifiers

[0061] content_read Specified content Zhang San's Travels content_read Duration of stay 5 seconds

[0062] In this embodiment, the behavior unique identifier content_read is configured with an enumeration type, including "specified content" and "stay duration", and the enumeration values ​​"《Zhang San's Travels》" and "5 seconds" are set respectively. This is used as an example for illustration. Specifically, the behavior unique identifier and the corresponding enumeration type and enumeration value are not specifically limited here.

[0063] Based on the filter options, determine the corresponding basic data type, configure the corresponding logical operation based on the corresponding basic data type, and based on the logical operation, combine with Doris built-in functions to make a mapping relationship and build a mapping relationship table.

[0064] As shown in Table 5 below:

[0065] Table 5: Mapping Relationship Table

[0066] STRING equal equal IN STRING not_equal Not equal to NOT IN STRING contain Include LIKE

[0067] In this embodiment, the mapping table includes four fields: data type, operator, key, and mapping relationship.

[0068] The data type is "STRING", the corresponding operators are "equal", "not_equal" and "contain", the corresponding keywords are "equal to", "not equal to" and "contain", and the mapping relationships are "IN", "NOT IN" and "LIKE".

[0069] S4: Configure rule relationships and formulate grouping rules according to requirements;

[0070] In this embodiment, the rule relationship is a rule relationship composed of at least one logical relationship. The specific logical relationship is set according to actual needs and is not specifically limited here. In this embodiment, the logical relationships AND and OR are used as examples for explanation.

[0071] As shown in the table above, the following grouping rules can be specified:

[0072] e1: Query people whose city is Chengdu or whose gender is male.

[0073] e2: Query people whose city is Chengdu and whose gender is male.

[0074] e3: Query people whose city is Chengdu and whose gender is male and who have visited the page "Zhang San's Travels".

[0075] e4: Query people whose city is Chengdu and whose gender is male and who have visited the page "Zhang San's Travels" and stayed for more than 5 seconds.

[0076] e5: Query people whose city is Chengdu and whose gender is male, or whose page "Zhang San's Travels" was visited and whose stay was longer than 5 seconds.

[0077] As can be seen from the above, different combination rules and different relationships result in different target groups, which allows for more accurate targeting and precise reach.

[0078] S5 converts the established clustering rules into SQL;

[0079] Specifically as follows:

[0080] S501: Create a cluster task, select filter items, corresponding enumeration values, corresponding logical operations and rule relationships, and combine them into a long JSON;

[0081] S502: Based on the long JSON, parse the above-configured filter items, corresponding enumeration values, corresponding logical operation relationships, and rule relationships into a single SQL statement, and encapsulate all the resulting SQL statements.

[0082] S6: By calling the SQL transformed based on the grouping rules, query the people who meet the rules, and store the people who meet the rules in a Bitmap type field. The grouping result field type is Bitmap, and it is stored in aggregate.

[0083] As shown in Table 6 below:

[0084] Table 6: Storage Data Table

[0085] 0000001 e1 001,002,003

[0086] The stored data table includes three fields: a unique ID for the subgroup, the name of the subgroup, and the subgroup result. The value corresponding to the subgroup result is the user's unique ID.

[0087] This invention is not limited to the specific embodiments described above. The invention extends to any new feature or combination disclosed in this specification, as well as any new method or process step or combination disclosed herein.

Claims

1. A method for building a user segmentation rule engine based on Doris, characterized in that, include: S1: Extract basic user data and business data, and store user tag tables according to the required grouping dimensions, with each user's unique ID as the unit; S2: Construct filter options based on cluster dimensions and behavioral data, and configure corresponding enumeration data based on existing filter options, as detailed below: Construct a behavior data table, which includes two fields: behavior name and behavior unique identifier; Use the unique identifier of the behavior as a filtering option; Configure a unique identifier for the behavior with its corresponding enumeration value; S3: Based on the filter items, determine the corresponding basic data type, configure the corresponding logical operation based on the corresponding basic data type, and based on the logical operation, combine with Doris built-in functions to make a mapping relationship and build a mapping relationship table; The mapping table includes four fields: data type, operator, key, and mapping relationship. S4: Configure rule relationships and formulate grouping rules according to requirements; S5 converts the established clustering rules into SQL, as follows: S501: Create a cluster task, select filter items, corresponding enumeration values, corresponding logical operations and rule relationships, and combine them into a long JSON; S502: Based on the long JSON, the above-configured filter items, corresponding enumeration values, corresponding logical operation relationships, and rule relationships are parsed into a single SQL statement, ultimately yielding the SQL statements corresponding to all rules; S6: By calling the SQL obtained based on the grouping rules, query the people who meet the rules, and store the people who meet the rules in a Bitmap type field to create a storage data table.

2. The method for constructing a user segmentation rule engine based on Doris according to claim 1, characterized in that, In step S1, the user tag table includes three fields: tag unique identifier, tag value, and user unique ID.

3. The method for constructing a user segmentation rule engine based on Doris according to claim 2, characterized in that, In step S2, filter items are constructed based on the cluster dimension, and the corresponding enumeration data is configured, as follows: Use the unique identifier of the label as a filter option; Configure the corresponding enumeration value for the unique identifier of the label.

4. The method for constructing a user segmentation rule engine based on Doris according to claim 3, characterized in that, The unique identifier of each label corresponds to at least one enumerated value.

5. The method for constructing a user segmentation rule engine based on Doris according to any one of claims 1-4, characterized in that, The behavior unique identifier corresponds to at least one enumeration type, and the enumeration type corresponds to at least one enumeration value.

6. The method for constructing a user segmentation rule engine based on Doris according to any one of claims 1-4, characterized in that, Both the user tag table and the behavior data table are built and stored based on Doris.

7. The method for constructing a user segmentation rule engine based on Doris according to claim 1, characterized in that, The stored data table includes three fields: unique ID of the group, group name, and group result. The value corresponding to the group result is the user's unique ID.

Citation Information

Patent Citations

  • Data retrieval system

    CN112364033A

  • User grouping method and device, electronic equipment and storage medium

    CN115599976A