Data marking method based on off-line total data and all label rules

Through the data marking method based on offline full data and all label rules, the problems of high calculation costs, rule solidification and resource waste in traditional data marking methods are solved, and flexible adjustment and efficient processing of label rules are realized, improving the efficiency and stability of data marking.

CN120277101APending Publication Date: 2025-07-08HANGZHOU ARCWAY TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510347341.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional data marking methods are difficult to meet the rapid development needs of the flexible employment market due to high real-time calculation costs, complex rule adjustments, difficult processing of existing data, solidification of rules and waste of computing resources.

Method used

Using a data marking method based on offline full data and all tag rules, the conditional filtering tag rules are constructed, the filtering conditions are converted using SQL expressions, combined with Hive SQL task scheduling, offline calculation and real-time data processing are realized, and flexible adjustment of tag rules and reasonable resource allocation are supported.

Benefits of technology

It realizes flexible adjustment of label rules, improves the efficiency and maintainability of data tags, reduces computing resource consumption, avoids data coverage conflicts, and meets the stability needs in large-scale scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277101A_ABST
    Figure CN120277101A_ABST
Patent Text Reader

Abstract

The invention discloses a data marking method based on off-line total data and all label rules, which comprises the following steps of: S1, constructing a condition screening label rule which comprises a label name, a marking mode, a marking result coverage strategy and a screening condition formed by attributes of a user group or a post group, a rule expression and an attribute value; s2, the screening conditions are converted into data query logic through an SQL expression, the SQL expression comprises SELECT, FROM and WHERE clauses, and the WHERE clauses comprise default filtering conditions and dynamically generated label rule conditions; and S3, based on all the data groups and all the label rules, generating a label result group through offline calculation. According to the method, flexible adjustment of the label rules, efficient processing of stock data and reasonable allocation of computing resources can be achieved, meanwhile, full-life-cycle management is conducted on the rules through the label rule database, and the flexibility, efficiency and maintainability of data marking are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of part-time job recommendation applications, and particularly to a data marking method based on offline full-volume data and all tag rules. Background Art

[0002] With the rapid development of the flexible employment market, the precise matching of users and jobs has become the core demand for traffic distribution. When traditional data marking methods are used to achieve the matching of users and jobs, they mainly rely on the following technical means: 1. Real-time tagging method based on a rule engine: This method uses a real-time rule engine (such as Drools, Spark Streaming) to perform real-time analysis on user behaviors or job information and customizes the generation of tags. However, this method has the following defects: High real-time calculation cost: It needs to continuously occupy computing resources, has strict requirements for hardware performance, and is difficult to handle large-scale data scenarios. Difficult to process stock data: It can only process real-time data streams and cannot batch-tag historical stock users or jobs. Complex rule adjustment: Rule modification requires code changes and testing processes and cannot meet the needs of rapid business iteration. 2. Offline processing method based on a static tag library: This method pre-defines fixed tag rules (such as "number of recruits > 5") and batch-tags data through offline tasks (such as Hive SQL). However, this method has the following limitations: Rule solidification: The tag rules are strongly bound to the business logic and cannot dynamically adjust the attribute combination (such as adding a "financial attribute" screening condition). Low full-volume matching efficiency: It needs to run an offline task for each tag, resulting in waste of computing resources (for example, 100 tags need to execute 100 tasks). Result coverage conflict: When multiple tasks write to the same result table at the same time, it is easy to cause data coverage errors. Summary of the Invention

[0003] The purpose of the present invention is to provide a data marking method based on offline full-volume data and all tag rules. The present invention can achieve flexible adjustment of tag rules, efficient processing of stock data, and reasonable allocation of computing resources. At the same time, through the tag rule database, the rules are managed throughout the life cycle, improving the flexibility, efficiency, and maintainability of data marking.

[0004] The technical solution provided by the present invention is as follows: A data marking method based on offline full-volume data and all tag rules, comprising the following steps:

[0005] S1. Construct conditional screening tag rules, including tag names, tagging methods, tag result coverage strategies, and screening conditions composed of the attributes, rule expressions, and attribute values of user groups or job groups;

[0006] S2. Convert the filtering conditions into data query logic through SQL expressions. The SQL expressions include SELECT, FROM, and WHERE clauses, where the WHERE clause includes default filtering conditions and dynamically generated tag rule conditions;

[0007] S3. Generate a tag result population through offline calculation based on the entire data population and all the tag rules.

[0008] In the above data marking method based on offline full - volume data and all tag rules, in the step of generating conditional screening tag rules:

[0009] The tag name input box is used to define the rule identifier;

[0010] The tagging methods include two modes: offline operation and real - time calculation;

[0011] The tag result coverage strategy includes two modes: overwrite and accumulation;

[0012] The tag rules contain an attribute set of user portraits or job portraits, and the attribute set includes four types of fields: basic attributes, value attributes, behavior attributes, and financial attributes.

[0013] In the above - mentioned data marking method based on offline full - volume data and all tag rules, in the SQL conversion step:

[0014] Set dynamic rule expressions for different data types:

[0015] When the selected field is of string type, four operators are provided: equal to, not equal to, fuzzy inclusion, and precise inclusion;

[0016] When the selected field is of numeric type, six operators are provided: greater than, greater than or equal to, less than, less than or equal to, not equal to, and range;

[0017] When the selected field is of date and / or time type, a date range operator is provided.

[0018] In the above - mentioned data marking method based on offline full - volume data and all tag rules, the offline operation mode is batch processing based on the latest snapshot data; the real - time calculation mode is incremental processing by docking real - time data streams.

[0019] In the above - mentioned data marking method based on offline full - volume data and all tag rules, the overwrite mode is to update and replace historical marking records each time; the accumulation mode is to retain all historical hit records.

[0020] In the above-mentioned data marking method based on offline full-volume data and all tag rules, in step S2, during the conversion process, placeholders for tag IDs and tag names are embedded in the SELECT field list, the portrait data table generated in step S1 is associated in the FROM data source table, and the system-predefined validity filtering conditions and user-defined rule conditions are merged in the WHERE condition set.

[0021] In the above-mentioned data marking method based on offline full-volume data and all tag rules, the offline calculation to generate the tag result population includes the following sub-steps:

[0022] Connect to the tag rule database to obtain all rules;

[0023] Group the rules according to preset rules and convert them into data processing tasks;

[0024] Use UNION ALL to splice each tag rule SQL task within each group. For the first group, use INSERT OVERWRITE to write, and for subsequent groups, use INSERT INTO to append;

[0025] Finally, through the Hive SQL task scheduling, achieve the full matching calculation of full-volume data and all rules, and generate a tag result population that meets the rules.

[0026] In the above-mentioned data marking method based on offline full-volume data and all tag rules, the grouping strategy satisfies:

[0027] The number of rules in each group is dynamically balanced to ensure the SQL execution efficiency;

[0028] Adopt different insertion methods for the same result table between groups to avoid data overwrite;

[0029] Support dynamic expansion of the number of groups to adapt to different computing resource environments.

[0030] Compared with the prior art, the present invention has the following effects:

[0031] 1. The dynamic rule configuration ability is significantly enhanced: The present invention supports the free combination of multi-dimensional attributes through a visual interface, can dynamically adjust tag rules in real time, solves the pain points of rule solidification and inability to adapt to business changes in traditional solutions, and greatly improves the rule update response speed.

[0032] 2. The offline full-volume calculation mode optimizes resource utilization: The present invention uses the Hive SQL engine to batch process full-volume data, combined with the task grouping strategy, significantly reduces the consumption of computing resources, and avoids the high cost problem of real-time calculation. The present invention supports the efficient processing of TB-level data and meets the stability requirements in large-scale scenarios.

[0033] 3. The distributed execution mechanism ensures stability in high-concurrency scenarios: Through the rule grouping and result merging strategy (the first group overwrites and writes, and subsequent appends), the present invention avoids the data bottleneck of single-threaded processing and the SQL length limit problem, significantly improves the label coverage rate, and there is no risk of data loss or overwrite conflict.

[0034] 4. Automated data processing reduces the risk of manual intervention: The present invention can automatically generate standardized SQL query logic, avoiding logical loopholes when manually writing rules, improving the accuracy of label definition and data consistency, and reducing the manual maintenance cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a schematic diagram of creating job label rules;

[0036] Figure 2 is a schematic diagram of the process of running a Hive SQL task. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The present invention will be further described below in conjunction with embodiments, but it is not used as a basis for limiting the present invention.

[0038] Embodiment: A data marking method based on offline full-volume data and all label rules includes the following steps:

[0039] S1. Construct a conditional screening label rule, including a label name, a marking method, a marking result coverage strategy, and a screening condition composed of the attributes of user groups or job groups, rule expressions, and attribute values;

[0040] In the step of generating the conditional screening label rule:

[0041] The label name input box is used to define the rule identifier;

[0042] The marking methods include two modes: offline operation and real-time calculation; the offline operation mode performs batch processing based on the latest snapshot data; the real-time calculation mode performs incremental processing by docking real-time data streams;

[0043] The marking result coverage strategies include two modes: overwrite and accumulation; the overwrite mode updates and replaces historical marking records each time; the accumulation mode retains all historical hit records.

[0044] The label rule contains an attribute set of user portraits or job portraits, and the attribute set contains four types of fields: basic attributes, value attributes, behavior attributes, and financial attributes. By selecting attributes, specific group users and specific job parts can be circled.

[0045] S2. Convert the filtering condition into data query logic through an SQL expression, where the SQL expression includes SELECT, FROM, and WHERE clauses, and the WHERE clause includes a default filtering condition and a dynamically generated tag rule condition;

[0046] In the SQL conversion step:

[0047] Set dynamic rule expressions for different data types:

[0048] When the selected field is of string type, provide four operators: equal to, not equal to, fuzzy inclusion, and precise inclusion;

[0049] When the selected field is of numerical type, provide six operators: greater than, greater than or equal to, less than, less than or equal to, not equal to, and range;

[0050] When the selected field is of date and / or time type, provide a date range operator.

[0051] During the conversion process, embed placeholders for tag IDs and tag names in the SELECT field list, associate with the portrait data table generated in step S1 in the FROM data source table, and merge the system-predefined validity filtering conditions and user-defined rule conditions in the WHERE condition set.

[0052] S3. Generate a tag result population through offline calculation based on the entire data population and the entire set of tag rules.

[0053] The offline calculation to generate the tag result population includes the following sub-steps:

[0054] Connect to the tag rule database to obtain all rules;

[0055] Group the rules according to preset rules and convert them into data processing tasks; where the grouping strategy satisfies: the number of rules in each group is dynamically balanced to ensure SQL execution efficiency; different insertion methods for the same result table are used between groups to avoid data overwrite; support dynamic expansion of the number of groups to adapt to different computing resource environments.

[0056] Concatenate each tag rule task within each group using UNION ALL, write the first group using INSERT OVERWRITE, and append subsequent groups using INSERT INTO;

[0057] Finally, implement a full-match calculation of the entire data and all rules through Hive SQL task scheduling to generate a tag result population that conforms to the rules.

[0058] In this embodiment, as Figure 1As shown in the figure, this figure shows the creation of a job label rule. Among them, the label condition selects the "number of recruits" attribute, the rule expression selects "greater than", and the attribute value fills in "10".

[0059] The following takes the job label as an example to introduce the label condition, rule expression, and attribute value in detail. The label condition comes from the job portrait. The job portrait is a collection table of the basic attributes, value attributes, behavioral attributes, financial attributes, etc. corresponding to the job. Each job contains 120 fields. Based on these rich attributes, it is convenient for label users to flexibly, dynamically, and efficiently modify the required label rules according to business rules to circle the job group.

[0060] The rule expression is dynamically displayed according to the type corresponding to the above label condition. Among them, for the String type, four screening methods of equal to, not equal to, fuzzy inclusion, and precise inclusion can be selected; for the bigint type, six screening methods of greater than, greater than or equal to, less than, less than or equal to, not equal to, and range can be selected; for the date type, equal to and date range can be selected; for the datetime type, six screening methods of greater than, greater than or equal to, less than, less than or equal to, not equal to, and range can be selected.

[0061] The attribute value is combined and used based on the selected label condition and rule expression, including enumeration type and non-enumeration type. For the enumeration type, click and select according to the prompt here. For the non-enumeration type, the label user needs to manually input here.

[0062] The label rule generated after this rule configuration is as follows:

[0063]

[0064] Through this rule SQL expression, the job group corresponding to this condition can be circled, and the selected label rule condition takes effect in the WHERE condition of this SQL expression.

[0065] Now, a detailed introduction to this SQL expression is given. This SQL expression contains three main parts: SELECT, FROM, and WHERE. The SELECT part contains job_id, the label type with a fixed type of "2", and the label id and label name of the two input parameters '${tagId}' and '${tagName}' at runtime; the FROM part is the data source table, that is, the job portrait table containing 120 fields. The WHERE part contains the default conditions status and verified to retrieve valid jobs, and the last AND represents the rule set selected by the label user.

[0066] The introduction based on the marking results of all data groups and all label rules is as follows:

[0067] Based on the generated rule SQL expressions above, the tag result population that conforms to this rule can be generated by running a Hive SQL task. The running process is as Figure 2 shown. The main process is as follows: First, connect to the database where the tag rules are located to obtain all tag rules. By directly connecting through the jdbc method, the latest rules during task execution can be obtained, and all the obtained tag rules are converted for data processing. All tag rules are divided into 5 groups to improve the execution efficiency and avoid errors caused by overly long SQL lengths. The rules in each group are concatenated in the form of UNION ALL. Among them, the SQL executed in the first group is written into the result table in the form of INSERT OVERWRITE, and the remaining groups are written into the result table in the form of INSERT INTO to avoid the results of each group from overwriting each other. Finally, the concatenated SQL is submitted to run online to generate the results.

[0068] In summary, the present invention can achieve flexible adjustment of tag rules, efficient processing of existing data, and reasonable allocation of computing resources. At the same time, through the tag rule database, full life cycle management of the rules is carried out, improving the flexibility, efficiency, and maintainability of data marking.

Claims

1. A data marking method based on offline full - volume data and full - volume label rules, characterized in that, It includes the following steps: S1. Construct condition screening label rules, including label names, marking methods, marking result coverage strategies, and screening conditions composed of the attributes of user groups or job groups, rule expressions, and attribute values; S2. Convert the screening conditions into data query logic through SQL expressions. The SQL expressions include SELECT, FROM, and WHERE clauses, where the WHERE clause includes default filtering conditions and dynamically generated label rule conditions; S3. Generate a label result group through offline calculation based on the entire data group and the entire label rules.

2. The data labeling method based on offline full-volume data and all label rules according to claim 1, characterized in that, In the step of generating the condition screening label rules: The label name input box is used to define the rule identifier; The marking methods include two modes: offline operation and real-time calculation; The marking result coverage strategies include two modes: overwrite and accumulation; The label rules include a set of attributes of user portraits or job portraits, and the set of attributes includes four types of fields: basic attributes, value attributes, behavior attributes, and financial attributes.

3. The data marking method based on offline full-volume data and all tag rules according to claim 1, wherein In the SQL conversion step: Set dynamic rule expressions for different data types: When the selected field is of string type, provide four operators: equal to, not equal to, fuzzy inclusion, and precise inclusion; When the selected field is of numerical type, provide six operators: greater than, greater than or equal to, less than, less than or equal to, not equal to, and range; When the selected field is of date and / or time type, provide a date range operator.

4. The data marking method based on offline full-volume data and all tag rules according to claim 2, characterized in that, The offline operation mode performs batch processing based on the latest snapshot portrait data; the real-time calculation mode processes incrementally by docking with real-time data streams.

5. The data labeling method based on offline full-volume data and all tag rules according to claim 2, wherein The overwrite mode replaces historical marked records each time it is updated; the accumulation mode retains all historical hit records.

6. The data marking method based on offline full-volume data and all tag rules according to claim 1, characterized in that In step S2, during the data conversion process, placeholders for label IDs and label names are embedded in the SELECT field list, the portrait data table generated in step S1 is associated in the FROM data source table, and the system-predefined validity filtering conditions and user-defined rule conditions are merged in the WHERE condition set.

7. The data marking method based on offline full-volume data and all label rules according to claim 1, characterized in that, The offline calculation to generate the label result group includes the following sub-steps: Connect to the label rule database to obtain all label rules; Group the rules according to preset rules and convert them into different groups of data processing tasks; Concatenate each label rule SQL task within each group using UNION ALL. The first group is written using INSERT OVERWRITE, and subsequent groups are appended using INSERT INTO; Finally, through the Hive SQL task scheduling, a full match calculation of all data and all rules is achieved to generate a label result group that meets the rules.

8. The data marking method based on offline full-volume data and label rules according to claim 7, characterized in that, The grouping strategy satisfies: The number of rules in each group is dynamically balanced to ensure the SQL execution efficiency; Different insertion methods are used for the same result table between groups to avoid data overwrite; Supports dynamic expansion of the number of groups to adapt to different computing resource environments.

Citation Information

Cited By

  • Supervision-oriented data compliance reconstruction and question and answer platform and retrieval method

    CN121233739A