A data processing method and system based on indicators and tags
By configuring calculation logic and dimensions to generate indicator generation templates, combining actual task requirements to generate scripts to be scheduled, and using data sources to execute scripts to calculate indicators and labels, the problem of long data indicator and label processing processes in existing technologies is solved, and rapid generation and efficient management are achieved.
Patent Information
- Application Number
- CN202411613668.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-13
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2044-11-13
AI Technical Summary
In existing technologies, the processing flow of data indicators and labels is long and relies on professionals, resulting in long output cycles, low timeliness, and chaotic management. The data processing system cannot be generated efficiently, and it is impossible to effectively generate labels and indicators that meet actual task requirements.
By configuring the calculation logic and dimensions, generating indicator generation templates, and combining them with actual task requirements to generate scripts to be scheduled, the scripts are executed using data sources to calculate indicators and labels, enabling rapid generation of indicators and labels that meet task requirements.
Without relying on manual operations, it can quickly generate indicators and labels that meet actual task requirements, improve data processing efficiency, simplify management processes, reduce duplication of development, and improve the flexibility and accuracy of data processing.
Smart Images

Figure CN119624210B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data processing technology, and specifically relates to a data processing method and system based on indicators and labels. Background Art
[0002] In today's big data era, companies need to analyze their operating conditions by mining operational data from various perspectives in their daily operations to facilitate subsequent business decisions and activities. These multi-dimensional indicators and labels from various perspectives often require complex calculations and high flexibility.
[0003] However, the processing of data indicators and labels is lengthy and relies on the expertise of professional data technicians. Consequently, the output of these indicators and labels often suffers from long production cycles, delayed launch times, and low efficiency. Furthermore, the value of certain indicators and labels often decreases with delayed production. Furthermore, the sheer number and complexity of indicators and labels can lead to management chaos, hindering the management and development of the indicator and labeling system. This leads to frequent duplication of development and a reliance on professional data technicians, making it difficult for business personnel to conduct independent exploration. Summary of the Invention
[0004] This application proposes a data processing method and system based on indicators and labels, which can quickly generate labels and indicators that meet actual task requirements through simple operations, thereby improving the efficiency of data processing.
[0005] A first aspect of the present application provides a data processing method based on indicators and labels, the method comprising:
[0006] According to the actual application scenario, the indicator generation template is obtained by configuring the calculation logic and dimensions;
[0007] Associating the actual task requirements with the indicator generation template to generate a script to be scheduled;
[0008] According to the predetermined data source, the script to be scheduled is executed to calculate the indicators and labels, and generate the indicators and labels that meet the actual task requirements;
[0009] According to the indicators and labels, demand analysis results are obtained.
[0010] The above solution first configures the logic and dimensions based on the actual application scenario, completing the mapping between the data source and the indicators and labels to be generated, and obtaining the corresponding indicator generation template. The actual task requirements are then associated with the indicator generation template, and the data analysis tasks required to be completed are mapped to the template, thus implementing task allocation and obtaining a script to be scheduled that can execute the task. Finally, a data source is selected for the script to be scheduled. Based on the mapping between the labels, indicators, and data sources in the script to be scheduled, the indicators and labels are calculated. This method obtains the indicators and labels that meet the actual task requirements without relying on manual operation, and then completes the demand analysis based on the obtained indicators and labels.
[0011] In a possible implementation method of the first aspect, according to an actual application scenario, an indicator generation template is obtained by configuring calculation logic and dimensions, specifically:
[0012] According to the actual application scenario, several calculation targets and the dimensions corresponding to the calculation targets are obtained;
[0013] Generating a first logic unit for each of the calculation targets according to the calculation targets; wherein each of the first logic units can be used to calculate a plurality of indicators and labels;
[0014] Based on the dimension, field configuration is performed for each of the first logical units to obtain an indicator generation template.
[0015] The above solution first determines the computational objectives and the data dimensions required for the specific application scenario. It then generates a corresponding first logical unit for each computational objective, providing technical support for the calculation of indicators and labels. It then configures the dimensions of the first logical unit and sets its fields to provide data support for the calculation of indicators and labels.
[0016] In a possible implementation method of the first aspect, based on the dimension, an indicator is configured for each first logical unit to obtain an indicator generation template, specifically:
[0017] Divide each dimension into several sub-dimensions according to the attributes of the dimension, and determine the associated fields corresponding to the sub-dimensions;
[0018] According to the business logic of the actual application scenario, determine the indicator field corresponding to the associated field from a preset data field library;
[0019] Each of the first logic units is configured according to the indicator field to obtain an indicator generation template.
[0020] The above solution first stratifies the dimension into multiple sub-dimensions. By refining the dimension, the corresponding indicator field for each sub-dimension can be accurately identified. This allows for faster matching of data source data with indicator fields in subsequent applications, improving data extraction efficiency. By configuring the first logical unit using the indicator fields, a pre-configured indicator generation template is generated.
[0021] In a possible implementation method of the first aspect, the actual task requirements are associated with the indicator generation template to generate a script to be scheduled, specifically:
[0022] According to the actual task requirements, set the indicator filtering conditions and dimension matching conditions for the indicator generation template to obtain the script to be scheduled;
[0023] The indicator filtering condition is used to filter out indicators that meet the actual task requirements from the data source; and the dimension matching condition is used to extract tags that meet the actual task requirements from the indicators.
[0024] The above solution sets indicator filtering conditions based on actual task requirements, effectively extracting the required data from a large amount of data for indicator generation. Dimension matching conditions are then used to further identify labels that meet the actual task requirements within the generated indicators, ensuring the accuracy of indicator and label generation.
[0025] In a possible implementation method of the first aspect, the demand analysis result is specifically:
[0026] Obtaining the application frequency of each indicator or label in the demand analysis results;
[0027] When the application frequency exceeds a first threshold, solidifying the corresponding label or indicator;
[0028] The demand analysis task is performed using the solidified labels or indicators at a preset frequency.
[0029] After obtaining indicators and labels, the above solution also monitors their frequency of use to determine their importance. When an indicator or label is used frequently, it is fixed for long-term or regular use in data processing, increasing the breadth of data application.
[0030] A second aspect of the present application provides a data processing system based on indicators and labels, the system comprising: a logic configuration module, a script generation module, an indicator and label generation module, and a demand analysis result generation module;
[0031] The logic configuration module is used to obtain an indicator generation template by configuring calculation logic and dimensions according to actual application scenarios;
[0032] The script generation module is used to associate the actual task requirements with the indicator generation template to generate a script to be scheduled;
[0033] The indicator and label generation module is used to execute the script to be scheduled to calculate indicators and labels according to a predetermined data source, and generate indicators and labels that meet the actual task requirements;
[0034] The demand analysis result generating module is used to obtain the demand analysis result according to the indicators and labels.
[0035] In a possible implementation of the second aspect, the logic configuration module includes: a data configuration unit;
[0036] The template generation unit is used to obtain several computing targets and dimensions corresponding to the computing targets based on an actual application scenario; generate a first logical unit for each computing target based on the computing targets; each of the first logical units can be used to generate several indicators and labels; and based on the dimensions, perform field configuration for each of the first logical units to obtain an indicator generation template.
[0037] In a possible implementation of the second aspect, the logic configuration module includes: a template generation unit;
[0038] Among them, the template generation unit is used to divide each dimension into several sub-dimensions according to the attributes of the dimension, and determine the associated fields corresponding to the sub-dimensions; according to the business logic of the actual application scenario, determine the indicator field corresponding to the associated field from a preset data field library; according to the indicator field, configure each of the first logical units to obtain an indicator generation template.
[0039] In a possible implementation of the second aspect, the script generation module includes: a condition setting unit;
[0040] Among them, the condition setting unit is used to set indicator filtering conditions and dimension matching conditions for the task script according to actual task requirements to obtain the script to be scheduled; wherein, the indicator filtering conditions are used to filter out indicators that meet the actual task requirements from the data source; the dimension matching conditions are used to extract labels that meet the actual task requirements from the indicators.
[0041] In a possible implementation of the second aspect, the demand analysis result generating module includes: a data solidification unit;
[0042] The data solidification unit is used to obtain the application frequency of each indicator or label in the demand analysis result; when the application frequency exceeds a first threshold, the corresponding label or indicator is solidified; and the demand analysis task is performed using the solidified label or indicator at a preset frequency. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for use in the implementation. Obviously, the drawings described below are only some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0044] Figure 1 This is a specific flow chart of a data processing method based on indicators and labels provided in one embodiment of the present application;
[0045] Figure 2 This is a logic and dimensional configuration diagram of a data processing method based on indicators and tags provided in one embodiment of the present application;
[0046] Figure 3 This is a specific structural diagram of a data processing system based on indicators and tags provided in a certain embodiment of the present application. DETAILED DESCRIPTION
[0047] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0048] It should be understood that the step numbers used herein are only for convenience of description and are not intended to limit the order in which the steps are to be executed.
[0049] First embodiment
[0050] Many practical application scenarios rely on data labels, such as user profiling, recommendation algorithms, data analysis and mining, and refined operations. Labels are used to uniformly assign classification information to standardized data in the data warehouse. Data indicators are one of the key tools for measuring goals and, along with labels, play a crucial role in the data analysis process. The processing of existing indicators and labels is lengthy and relies on specialized data technicians, resulting in low efficiency in the output of indicators and labels. Furthermore, when the amount of data for indicators and labels becomes excessively large, the lack of a unified management system necessitates research on how to unify the mapping between labels, indicators, and data sources. This allows for the generation of a large number of usable labels and indicators based on the data sources, facilitating effective data analysis.
[0051] like Figure 1 As shown, Figure 1 A specific flow chart of a data processing method based on indicators and labels is provided for a certain embodiment of the present application. The data processing method based on indicators and labels of this embodiment includes steps S1 to S4, which are detailed as follows:
[0052] Step S1: According to the actual application scenario, the indicator generation template is obtained by configuring the calculation logic and dimensions.
[0053] In the embodiment of the present application, it is based on the typical B / S architecture design, that is, the interaction between the web page end, the application end and the database server. The implementation of the entire embodiment is mainly divided into two parts: the configuration of the scheduling script and the use of the configured script to be executed.
[0054] For the script to be executed, the first step is to configure the calculation logic and dimensions to obtain the corresponding indicator generation template. During the configuration process, it is necessary to perform a unified basic maintenance for the most basic calculation logic involved in the actual application scenario, and simply configure the calculation logic unit with multiple algorithms to obtain multiple first logic units. In an embodiment of the present application, the first logic unit is a configured SQL template. Each first logic unit can be used to generate a number of indicators and labels. At the same time, by specifying the analyzable dimensions, associated keywords and syntax of the calculation logic unit, the indicator and label logic generation of the achievable dimensions based on the calculation logic can be completed, and the algorithm logic of multiple calculations can be optionally specified to support subsequent users to configure complex indicators and label logic generation scenarios.
[0055] Then, based on the SQL code in the first logical unit, the corresponding dimensions are determined. These dimensions are then divided into sub-dimensions of major categories and minor categories, and the associated fields corresponding to these sub-dimensions are determined. Based on the business logic of the actual application scenario, the indicator fields corresponding to these associated fields are determined from a preset data field library. Finally, the indicator fields are associated with the sub-dimensions, completing the mapping relationship between the indicator and the data source.
[0056] As an improvement to the above solution, based on the dimensions and combined with the business logic caliber of the indicators and labels involved in the actual application scenarios, the calculation logic units used for the indicators and the data fields for statistical analysis are defined to confirm the meaning of the indicators and the subsequent calculation logic, and the indicator fields corresponding to the associated fields are determined from the preset data field library.
[0057] Broadly speaking, labels are also a type of dimension, but what makes them unique is that they are derived from indicators. For example, in the fund industry, gender is a common dimension. Total client transactions are typically an indicator, but if a client's total transactions exceed 1,000, then a high-frequency trading label is derived from this total number of transactions. This label can be used to further analyze the client's investment behavior and returns. Furthermore, many of these labels are highly diverse and dynamic. One time, a high-frequency trading client might be considered high-frequency trading, while the next might be considered high-frequency trading clients with more than 5,000 transactions. The next might be considered high-frequency trading clients with more than 100 transactions in the past year, and so on. Labels can also be combined with conventional dimensions, such as analyzing the returns and investment performance of high-frequency traders by gender. Therefore, analyzing data based on these variable labels and conventional, diverse dimensions is difficult to achieve using conventional data analysis tools and data warehouses. Therefore, a flexible and adaptable framework is required to support this type of data analysis.
[0058] For example, in a specific application scenario, it is necessary to summarize the total number of transactions, total transaction amount, and total sales amount of fund users this year. The calculation objective of this application scenario is to count the total number of transactions and total transaction amount of users for the year, and filter out sales operations to calculate the total sales amount. The dimensions of this calculation objective are the data involved in the calculation, such as sales amount, transaction operations, and the amount per transaction.
[0059] Specifically, when obtaining several calculation targets and the dimensions corresponding to them through actual application scenarios, it is necessary to consider the business requirements of the actual application scenarios. Typically, business requirements describe the desired data in text or tabular form, such as the total number of transactions, total transaction amounts, and total sales amounts by customer, by age group, by vendor, and by fund, since the beginning of the year. By converting the business description into data analysis language, with age, vendor, fund, and since the beginning of the year as dimensions and total number of transactions, total transaction amounts, and total sales amounts as indicators, the corresponding dimension categories are specified. For example, age is classified as "other," vendor as "channel," fund as "product," and since the beginning of the year as "time." Predefined dimension categories and dimension templates have corresponding preconfigured underlying SQL database templates. By selecting a dimension, the corresponding dimension category and dimension template are matched, and the underlying SQL database template and fields are also matched, thereby generating an indicator generation template.
[0060] To better demonstrate how to configure logic and dimensions, Figure 2 A logical and dimension configuration diagram is provided. As shown in the figure, the indicator generation template in the figure is named Full Test Execution 1. Its analysis dimensions include time, customer, and channel. The template associates the corresponding indicator fields based on these dimensions. The SQL syntax in the middle is the configured first logical unit. Below is the field configuration completed by this first logical unit based on the dimensions: total number of fund codes and total number of distributors, resulting in the corresponding indicator generation template.
[0061] After obtaining the indicator generation template, you can also configure intermediate grouping and secondary algorithms for the indicator fields to support statistical analysis at the indicator level.
[0062] Step S2: Associating the actual task requirements with the indicator generation template to generate a script to be scheduled.
[0063] In this embodiment, the overall task is broken down into multiple smaller tasks based on actual task requirements. Each smaller task corresponds to a segment of dynamic SQL code in the indicator generation template, and the overall task is actually also a segment of dynamic SQL code. Then, for each smaller task, indicators, tags, and dimensions are linked to meet the calculation scenario of the indicators and tags, and the corresponding script to be scheduled is generated.
[0064] In addition, indicator filtering conditions and dimension matching conditions are set for the script to be scheduled. The indicator filtering conditions are used to filter out indicators that meet the actual task requirements from the data source; and the dimension matching conditions are used to extract tags that meet the actual task requirements from the indicators.
[0065] Step S3: Execute the script to be scheduled to calculate indicators and labels based on the predetermined data source, and generate indicators and labels that meet the actual task requirements.
[0066] In an embodiment of the present application, the backend server parses the parameters in the configured script to be scheduled, accesses a predetermined data source, executes the script to be scheduled to calculate indicators and labels, and generates a large number of indicators and labels that meet the actual task requirements.
[0067] Indicators are calculated based on basic database data, such as customer transaction flow, asset flow, and dividend records, along with dimensional data such as fund product categories. These data are then correlated and calculated using calculation rules such as grouping, summing, accumulation, maximum, and minimum values. Labels are similarly calculated based on basic database data. The aforementioned indicators can be converted into labels. For example, if the indicator above is the number of transactions a customer has made in the past year, then a label could be defined as follows: If a customer has made more than 30 transactions in the past year, the customer is labeled as a frequent trader. The label data in this case consists of a customer ID and a frequent trader label. However, in real-world applications, the definition of indicators and labels often changes, making it impossible to pre-calculate the corresponding indicators and labels. For example, if a customer's transactions in the past year exceed 30, then they are considered frequent traders. Next time, they might be 50, and the next time, they might be 30, excluding redemptions, and so on. Therefore, it is necessary to automatically generate indicators and labels that meet current needs through the script to be scheduled in the embodiment of the present application to reduce the manual estimation of indicators and labels in each calculation.
[0068] Step S4: Obtain demand analysis results based on the indicators and labels.
[0069] In an embodiment of the present application, according to actual business needs, the generated indicators and labels are used to analyze the existing data to obtain corresponding demand analysis results.
[0070] After long-term use, these indicators and labels will be evaluated for their value, and high-value indicators or labels will be selected and solidified. Then, the solidified labels or indicators will be used to perform demand analysis tasks at a preset frequency.
[0071] In the embodiments of the present application, the value of indicators and labels is mainly reflected by their application frequency.
[0072] Specifically, since the generated indicators and labels are actually one-time data, they are usually some indicators that the business suddenly came up with or tried out. They must undergo multiple rounds of business use verification before they can be judged to be of great value. For example, a certain requirement is to calculate the return rate of customers who have more than 30 transactions in each year over the past three years, and observe whether such indicators are indeed helpful for investor education. The business will then try to propose customer return rates from other perspectives, such as the return rate of customers who have more than 30 transactions in each year at Industrial and Commercial Bank of China, and so on. The corresponding indicators and labels are obtained and applied to the corresponding business needs analysis. After business verification and observation, if it is indeed helpful to the customer, the business will propose to solidify this indicator, and transform it from a one-time report into a regular weekly or monthly report, and use it as a reference for actual operations.
[0073] The implementation of the embodiments of the present application has the following beneficial effects:
[0074] The embodiment of the present application first configures the logic and dimensions according to the actual application scenario, completes the configuration of the mapping relationship between the data source and the indicators and labels to be generated, and obtains the corresponding indicator generation template; then the actual task requirements and the indicator generation template are associated, and the data analysis tasks to be completed by the actual task requirements are mapped to the template, thereby realizing the allocation of tasks and obtaining the script to be scheduled for the executable task. Finally, a data source is selected for the script to be scheduled, and the indicators and labels are calculated based on the mapping relationship between the labels, indicators and data sources in the script to be scheduled. Indicators and labels that meet the actual task requirements are quickly obtained without relying on manual operation, and then the obtained indicators and labels are used to perform a comprehensive data analysis on the business of the actual task requirements, and the logical generation of complex indicators, labels and dimensions is quickly completed without the need for manual intervention. Indicators and labels are quickly obtained to fully process the data, thereby improving the efficiency of data processing. In addition, the generated indicators and labels can also be valued to meet more customer needs.
[0075] Second embodiment
[0076] Furthermore, in order to execute the data processing system based on indicators and tags corresponding to the above method embodiment to achieve corresponding functions and technical effects, Figure 3 A structural diagram of a data processing system based on indicators and labels is provided. For ease of explanation, only the parts related to this embodiment are shown. The data processing system based on indicators and labels provided in this embodiment of the application includes:
[0077] The logic configuration module 201 is used to obtain an indicator generation template by configuring calculation logic and dimensions according to actual application scenarios.
[0078] In an embodiment of the present application, the logic configuration module 201 further includes a data configuration unit, which is used to obtain a number of computing targets and dimensions corresponding to the computing targets based on an actual application scenario; generate a first logic unit for each computing target based on the computing targets; wherein each first logic unit can be used to generate a number of indicators and labels; and based on the dimensions, perform field configuration for each first logic unit to obtain an indicator generation template.
[0079] The script generation module 202 is used to associate the actual task requirements with the indicator generation template to generate a script to be scheduled.
[0080] In an embodiment of the present application, the script generation module 202 also includes a condition setting unit, which is used to set indicator filtering conditions and dimension matching conditions for the task script according to actual task requirements to obtain the script to be scheduled; wherein, the indicator filtering conditions are used to filter out indicators that meet the actual task requirements from the data source; the dimension matching conditions are used to extract labels that meet the actual task requirements from the indicators.
[0081] The indicator and label generation module 203 is used to execute the script to be scheduled to calculate indicators and labels according to a predetermined data source, and generate indicators and labels that meet the actual task requirements.
[0082] In an embodiment of the present application, the backend server parses the parameters in the configured script to be scheduled, accesses a predetermined data source, executes the script to be scheduled to calculate indicators and labels, and generates a large number of indicators and labels that meet the actual task requirements.
[0083] Indicators are calculated based on basic database data, such as customer transaction flow, asset flow, and dividend records, along with dimensional data such as fund product categories. These data are then correlated and calculated using grouping, summation, accumulation, maximum, and minimum value calculations. Labels are similarly calculated based on basic database data. We can convert these indicators into labels. For example, if the indicator above is the number of transactions a customer has made in the past year, then the label could be defined as follows: If a customer has made more than 30 transactions in the past year, the customer is labeled as a frequent trader. The label data in this case consists of a customer ID and a frequent trader label. However, in real-world applications, the definition of indicators and labels often changes, making it impossible to pre-calculate the corresponding indicators and labels. For example, if a customer's transactions in the past year exceed 30, then they are considered frequent traders. Next time, they might be defined as 50, and the next time, they might be defined as 30, excluding redemptions. And so on. Therefore, it is necessary to automatically generate indicators and labels that meet current needs through the script to be scheduled in the embodiment of the present application to reduce the manual estimation of indicators and labels in each calculation.
[0084] The demand analysis result generating module 204 is used to obtain the demand analysis result according to the indicators and tags.
[0085] In an embodiment of the present application, the demand analysis result generation module 204 also includes a data solidification unit for obtaining the application frequency of each indicator or label in the demand analysis result; when the application frequency exceeds a first threshold, the corresponding label or indicator is solidified; and the demand analysis task is performed using the solidified label or indicator at a preset frequency.
[0086] In some embodiments, the logic configuration module 201 further includes:
[0087] It is based on a typical B / S architecture design, that is, the interaction between the web page end, the application end and the database server. The implementation of the entire embodiment is mainly divided into two parts: the configuration of the scheduling script and the use of the configured script to be executed.
[0088] For the script to be executed, the first step is to configure the calculation logic and dimensions to obtain the corresponding indicator generation template. During the configuration process, it is necessary to perform a unified basic maintenance for the most basic calculation logic involved in the actual application scenario, and simply configure the calculation logic unit with multiple algorithms to obtain multiple first logic units. In an embodiment of the present application, the first logic unit is a configured SQL template. Each first logic unit can be used to generate a number of indicators and labels. At the same time, by specifying the analyzable dimensions, associated keywords and syntax of the calculation logic unit, the indicator and label logic generation of the achievable dimensions based on the calculation logic can be completed, and the algorithm logic of multiple calculations can be optionally specified to support subsequent users to configure complex indicators and label logic generation scenarios.
[0089] Then, based on the SQL code in the first logical unit, the corresponding dimensions are determined. These dimensions are then divided into sub-dimensions of major categories and minor categories, and the associated fields corresponding to these sub-dimensions are determined. Based on the business logic of the actual application scenario, the indicator fields corresponding to these associated fields are determined from a preset data field library. Finally, the indicator fields are associated with the sub-dimensions, completing the mapping relationship between the indicator and the data source.
[0090] As an improvement to the above solution, based on the dimensions and combined with the business logic caliber of the indicators and labels involved in the actual application scenarios, the calculation logic units used for the indicators and the data fields for statistical analysis are defined to confirm the meaning of the indicators and the subsequent calculation logic, and the indicator fields corresponding to the associated fields are determined from the preset data field library.
[0091] Broadly speaking, labels are also a type of dimension, but what makes them unique is that they are derived from indicators. For example, in the fund industry, gender is a common dimension. Total client transactions are typically an indicator, but if a client's total transactions exceed 1,000, then a high-frequency trading label is derived from this total number of transactions. This label can be used to further analyze the client's investment behavior and returns. Furthermore, many of these labels are highly diverse and dynamic. One time, a high-frequency trading client might be considered high-frequency trading, while the next might be considered high-frequency trading clients with more than 5,000 transactions. The next might be considered high-frequency trading clients with more than 100 transactions in the past year, and so on. Labels can also be combined with conventional dimensions, such as analyzing the returns and investment performance of high-frequency traders by gender. Therefore, analyzing data based on these variable labels and conventional, diverse dimensions is difficult to achieve using conventional data analysis tools and data warehouses. Therefore, a flexible and adaptable framework is required to support this type of data analysis.
[0092] For example, in a specific application scenario, it is necessary to summarize the total number of transactions, total transaction amount, and total sales amount of fund users this year. The calculation objective of this application scenario is to count the total number of transactions and total transaction amount of users for the year, and filter out sales operations to calculate the total sales amount. The dimensions of this calculation objective are the data involved in the calculation, such as sales amount, transaction operations, and the amount per transaction.
[0093] Specifically, when obtaining several calculation targets and the dimensions corresponding to them through actual application scenarios, it is necessary to consider the business requirements of the actual application scenarios. Typically, business requirements describe the desired data in text or tabular form, such as the total number of transactions, total transaction amounts, and total sales amounts by customer, by age group, by vendor, and by fund, since the beginning of the year. By converting the business description into data analysis language, with age, vendor, fund, and since the beginning of the year as dimensions and total number of transactions, total transaction amounts, and total sales amounts as indicators, the corresponding dimension categories are specified. For example, age is classified as "other," vendor as "channel," fund as "product," and since the beginning of the year as "time." Predefined dimension categories and dimension templates have corresponding preconfigured underlying SQL database templates. By selecting a dimension, the corresponding dimension category and dimension template are matched, and the underlying SQL database template and fields are also matched, thereby generating an indicator generation template.
[0094] After obtaining the indicator generation template, you can also configure intermediate grouping and secondary algorithms for the indicator fields to support statistical analysis at the indicator level.
[0095] In some embodiments, the script generation module 202 further includes:
[0096] In this embodiment, the overall task is broken down into multiple smaller tasks based on actual task requirements. Each smaller task corresponds to a segment of dynamic SQL code in the indicator generation template, and the overall task is actually also a segment of dynamic SQL code. Then, for each smaller task, indicators, tags, and dimensions are linked to meet the calculation scenario of the indicators and tags, and the corresponding script to be scheduled is generated.
[0097] In addition, indicator filtering conditions and dimension matching conditions are set for the script to be scheduled. The indicator filtering conditions are used to filter out indicators that meet the actual task requirements from the data source; and the dimension matching conditions are used to extract tags that meet the actual task requirements from the indicators.
[0098] Step S3: Execute the script to be scheduled to calculate indicators and labels based on the predetermined data source, and generate indicators and labels that meet the actual task requirements.
[0099] In an embodiment of the present application, the backend server parses the parameters in the configured script to be scheduled, accesses a predetermined data source, executes the script to be scheduled to calculate indicators and labels, and generates a large number of indicators and labels that meet the actual task requirements.
[0100] Indicators are calculated based on basic database data, such as customer transaction flow, asset flow, and dividend records, along with dimensional data such as fund product categories. These data are then correlated and calculated using calculation rules such as grouping, summing, accumulation, maximum, and minimum values. Labels are similarly calculated based on basic database data. The aforementioned indicators can be converted into labels. For example, if the indicator above is the number of transactions a customer has made in the past year, then a label could be defined as follows: If a customer has made more than 30 transactions in the past year, the customer is labeled as a frequent trader. The label data in this case consists of a customer ID and a frequent trader label. However, in real-world applications, the definition of indicators and labels often changes, making it impossible to pre-calculate the corresponding indicators and labels. For example, if a customer's transactions in the past year exceed 30, then they are considered frequent traders. Next time, they might be 50, and the next time, they might be 30, excluding redemptions, and so on. Therefore, it is necessary to automatically generate indicators and labels that meet current needs through the script to be scheduled in the embodiment of the present application to reduce the manual estimation of indicators and labels in each calculation.
[0101] In some embodiments, the demand analysis result generating module 204 further includes:
[0102] In an embodiment of the present application, according to actual business needs, the generated indicators and labels are used to analyze the existing data to obtain corresponding demand analysis results.
[0103] After long-term use, these indicators and labels will be evaluated for their value, and high-value indicators or labels will be selected and solidified. Then, the solidified labels or indicators will be used to perform demand analysis tasks at a preset frequency.
[0104] In the embodiments of the present application, the value of indicators and labels is mainly reflected by their application frequency.
[0105] The implementation of the embodiments of the present application has the following beneficial effects:
[0106] The embodiment of the present application first configures the logic and dimensions according to the actual application scenario, completes the configuration of the mapping relationship between the data source and the indicators and labels to be generated, and obtains the corresponding indicator generation template; then the actual task requirements and the indicator generation template are associated, and the data analysis tasks to be completed by the actual task requirements are mapped to the template, thereby realizing the allocation of tasks and obtaining the script to be scheduled for the executable task. Finally, a data source is selected for the script to be scheduled, and the indicators and labels are calculated based on the mapping relationship between the labels, indicators and data sources in the script to be scheduled. Indicators and labels that meet the actual task requirements are quickly obtained without relying on manual operation, and then the obtained indicators and labels are used to perform a comprehensive data analysis on the business of the actual task requirements, and the logical generation of complex indicators, labels and dimensions is quickly completed without the need for manual intervention. Indicators and labels are quickly obtained to fully process the data, thereby improving the efficiency of data processing. In addition, the generated indicators and labels can also be valued to meet more customer needs.
[0107] The specific embodiments described above further illustrate the purpose, technical solutions, and beneficial effects of this application. It should be understood that the above description is merely a specific embodiment of this application and is not intended to limit the scope of protection of this application. In particular, it should be noted that for those skilled in the art, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.
Claims
1. A data processing method based on indicators and labels, characterized in that: include: Based on the actual application scenario, the indicator generation template is obtained by configuring the calculation logic and dimensions. Specifically, according to the actual application scenario, several calculation targets and the dimensions corresponding to the calculation targets are obtained; based on the calculation targets, a first logical unit is generated for each of the calculation targets; each of the first logical units can be used to calculate several indicators and labels; based on the dimensions, fields are configured for each of the first logical units to obtain the indicator generation template; Associating the actual task requirements with the indicator generation template to generate a script to be scheduled; According to the predetermined data source, the script to be scheduled is executed to calculate the indicators and labels, and generate the indicators and labels that meet the actual task requirements; Based on the indicators and labels, a demand analysis result is obtained; wherein the demand analysis result is specifically: obtaining the application frequency of each indicator or label in the demand analysis result; when the application frequency exceeds a first threshold, solidifying the corresponding label or indicator; and using the solidified label or indicator to perform the demand analysis task at a preset frequency.
2. The data processing method based on indicators and labels according to claim 1 is characterized in that: Based on the dimension, an indicator configuration is performed for each first logical unit to obtain an indicator generation template, specifically: Divide each dimension into several sub-dimensions according to the attributes of the dimension, and determine the associated fields corresponding to the sub-dimensions; According to the business logic of the actual application scenario, determine the indicator field corresponding to the associated field from a preset data field library; Each of the first logic units is configured according to the indicator field to obtain an indicator generation template.
3. The data processing method based on indicators and labels according to claim 1 is characterized in that: The actual task requirements are associated with the indicator generation template to generate a script to be scheduled, specifically: According to the actual task requirements, set the indicator filtering conditions and dimension matching conditions for the indicator generation template to obtain the script to be scheduled; The indicator filtering condition is used to filter out indicators that meet the actual task requirements from the data source; and the dimension matching condition is used to extract tags that meet the actual task requirements from the indicators.
4. A data processing system based on indicators and labels, characterized in that: include: Logic configuration module, script generation module, indicator and label generation module, and demand analysis result generation module; The logic configuration module is configured to obtain an indicator generation template by configuring calculation logic and dimensions according to actual application scenarios. The logic configuration module includes: a data configuration unit configured to obtain a number of calculation targets and dimensions corresponding to the calculation targets according to the actual application scenarios; generate a first logic unit for each calculation target according to the calculation targets; each first logic unit can be used to calculate a number of indicators and labels; and perform field configuration for each first logic unit based on the dimensions to obtain an indicator generation template. The script generation module is used to associate the actual task requirements with the indicator generation template to generate a script to be scheduled; The indicator and label generation module is used to execute the script to be scheduled to calculate indicators and labels according to a predetermined data source, and generate indicators and labels that meet the actual task requirements; The demand analysis result generation module is used to obtain the demand analysis results based on the indicators and labels; wherein, the demand analysis result generation module includes a data solidification unit, which is used to obtain the application frequency of each indicator or label in the demand analysis result; when the application frequency exceeds a first threshold, the corresponding label or indicator is solidified; and the demand analysis task is performed using the solidified label or indicator at a preset frequency.
5. The data processing system based on indicators and tags according to claim 4, characterized in that: The logic configuration module includes: a template generation unit; Among them, the template generation unit is used to divide each dimension into several sub-dimensions according to the attributes of the dimension, and determine the associated fields corresponding to the sub-dimensions; according to the business logic of the actual application scenario, determine the indicator field corresponding to the associated field from a preset data field library; according to the indicator field, configure each of the first logical units to obtain an indicator generation template.
6. The data processing system based on indicators and tags according to claim 4, characterized in that: The script generation module includes: a condition setting unit; Among them, the condition setting unit is used to set indicator filtering conditions and dimension matching conditions for the indicator generation template according to actual task requirements to obtain the script to be scheduled; wherein, the indicator filtering conditions are used to filter out indicators that meet the actual task requirements from the data source; the dimension matching conditions are used to extract labels that meet the actual task requirements from the indicators.
Citation Information
Patent Citations
Report-oriented multi-dimensional management analysis method and system
CN112183379A
Index rule generation method and device, electronic equipment and storage medium
CN116090867A