Configuration-based multi-level index calculation system and method
By building a configured multi-level indicator computing system, using metadata model and rule configuration modules, dynamic generation and real-time calculation of multi-level indicators are realized, solving the problems of long development cycles and poor scalability in traditional methods, and improving the real-time and scalability of the index system.
Patent Information
- Application Number
- CN202510756753.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-09
AI Technical Summary
Traditional indicator calculation methods are difficult to cope with the frequently emerging new indicator demands in large-scale real-time simulation scenarios, resulting in long development cycles, high labor costs, and the multi-level indicator calculation process is fragmented, making it impossible to build an associated network to mine deep data value.
A configuration-based multi-level indicator calculation system is built, including a collection module, a rule configuration module and an indicator calculation module. The metadata model is used to automatically extract the field metadata of heterogeneous data. By configuring statistical fields, field constraints, field constraint values and a rule model for calculation priority, the configuration generation of atomic indicators is realized, and the priority streaming calculation architecture is used for multi-level indicator calculation.
It realizes dynamic generation of the entire process from raw data to multi-level indicators, improves the real-time and scalability of the indicator system, eliminates field heterogeneity and real-time conflicts, and supports the configuration generation of atomic indicators and the dynamic correlation calculation of composite indicators.
Smart Images

Figure CN120256496A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a configuration-based multi-level index calculation system and method. Background Art
[0002] In the information simulation scenario driven by big data, the efficient generation and dynamic update of situation indicators are the core links to support real-time decision-making. Traditional index calculation methods usually adopt a hierarchical processing mode: atomic indicators are extracted from heterogeneous source data through customized code, composite indicators are combined with atomic indicators through a rule engine, and derivative indicators are calculated secondarily based on preset formulas.
[0003] However, with the expansion of the simulation scale and the improvement of real-time requirements, the customized development for the unique format of each data source in related technologies cannot meet the new index requirements that frequently emerge during the simulation process. Due to the field heterogeneity and real-time conflicts of multi-source data, it is necessary to repeatedly write data cleaning, association mapping, and calculation logic code during the development process, resulting in a long development cycle and high labor costs. Therefore, traditional index configuration or calculation methods are difficult to adapt to the dynamic changes of index requirements in large-scale simulation scenarios, leading to significant bottlenecks in the expansion of the index system. Summary of the Invention
[0004] Embodiments of the present invention provide a configuration-based multi-level index calculation system and method, aiming to solve the problems existing in the above background art.
[0005] To solve the above technical problems, the present invention is implemented as follows: In a first aspect, embodiments of the present invention provide a configuration-based multi-level index calculation system, which includes: a collection module, a rule configuration module, and an index calculation module; The collection module is configured with a metadata model; the collection module is used to obtain multiple heterogeneous source data and extract the field metadata of each heterogeneous source data by using the metadata model; The rule configuration module is used to configure at least one calculation rule for the target index based on the index type of the target index, and the calculation rule includes a statistical field, a field constraint operator, a field constraint value, a statistical calculation method, and a calculation priority; The index calculation module is used to match the field metadata of the multiple heterogeneous source data according to the statistical field, the field constraint operator, and the field constraint value of each calculation rule, and determine the heterogeneous source data with successfully matched field metadata as the index calculation data required for calculating the target index; and sequentially execute the statistical calculation methods of the corresponding calculation rules on the index calculation data in descending order of calculation priority to obtain the index result data of the target index.
[0006] Optionally, the metadata model is used to extract the field encoding, field name, and field description of each heterogeneous source data, and use the field encoding, field name, and field description of each heterogeneous source data as the field metadata of the heterogeneous source data; The metric calculation module is configured to: According to the type of the field constraint operator of each calculation rule, parse the field constraint value of each calculation rule to obtain the target constraint interval of each calculation rule. The types of the field constraint operators include equal to, numerical interval, contains, greater than, less than, greater than or equal to, less than or equal to, and not equal to; Among the multiple heterogeneous source data, determine the heterogeneous source data whose field encoding or field name contains the statistical field as candidate calculation data, and determine the data value of each candidate calculation data according to the field description of each candidate calculation data; Determine the heterogeneous source data whose data value is within the target constraint interval as the heterogeneous source data with successful field metadata matching.
[0007] Optionally, the system further includes a metric definition module and a metric recommendation module; The metric definition module is configured to obtain the multiple heterogeneous source data, and generate metric definition data corresponding to the target metric according to the multiple heterogeneous source data. The metric definition data is used to describe the target metric in different dimensions; The metric recommendation module is configured to obtain the metric definition data corresponding to each of the multiple metrics, and the metric result data corresponding to each of the multiple metrics; determine the association relationship between the metrics according to the metric definition data and the metric result data, and output a metric recommendation result representing the association relationship between the metrics.
[0008] Optionally, the metric recommendation module is configured with a graph database and a graph convolutional network model. Each node in the graph database represents a corresponding metric, and the graph convolutional network model is pre-trained according to the metric definition data corresponding to each of the multiple metrics and the metric result data corresponding to each of the multiple metrics; The metric definition module is configured to associate the metric definition data corresponding to each of the multiple metrics with the calculation rules of each of the multiple metrics to obtain the association calculation rule identifiers corresponding to each of the multiple metrics; The metric recommendation module is configured to determine the association relationship between the metrics according to the association calculation rule identifiers corresponding to each of the multiple metrics; calculate the association weights between the metrics according to the association relationship between the metrics, and establish the edges between the nodes in the graph database based on the association weights between the metrics; The index recommendation module is used to calculate the association strength between various indexes according to the nodes and edges in the graph database through the graph convolutional network model, and determine at least one set of associated indexes with an association strength greater than a preset threshold; output a recommended index result including at least one set of associated indexes.
[0009] Optionally, the index types include atomic indexes, composite indexes, and derivative indexes; The rule configuration module is used to, when the index type of the target index is an atomic index, configure a first-type calculation rule set including one calculation rule for the target index. The statistical calculation method of the first-type calculation rule set is configured as an atomic-level operator, and the atomic-level operator is any one of the original value, count, maximum value, minimum value, and average value; The index calculation module is used to, when the index type of the target index is an atomic index, perform the statistical calculation method of the corresponding calculation rule on the index calculation data to obtain the index result data of the target index.
[0010] Optionally, the rule configuration module is used to, when the index type of the target index is a composite index, configure a second-type calculation rule set including at least two calculation rules for the target index. Among them, the statistical calculation method of each calculation rule is configured as an atomic-level operator or a composite-level operator, and the composite-level operator includes at least one of addition, subtraction, multiplication, and division; The rule configuration module is further used to configure the corresponding calculation priority for each calculation rule in the second-type calculation rule set, and bind the calculation result of the calculation rule with the previous calculation priority as the input data of the calculation rule with the next calculation priority; The index calculation module is used to, when the index type of the target index is a composite index, sequentially perform the statistical calculation methods of each calculation rule on the index calculation data in the order from high to low of the calculation priority, and cache the intermediate result data into the memory mapping structure until the index result data of the target index is generated.
[0011] Optionally, the rule configuration module is used to, when the index type of the target index is a derivative index, configure a third-type calculation rule set including at least two calculation rules for the target index. Among them, the statistical calculation method of at least one calculation rule is configured as a derivative-level operator, and the derivative-level operator includes distinct count; The rule configuration module is further used to configure the corresponding calculation priority for each calculation rule in the third-type calculation rule set, and configure an additional field list, and the additional field list is used to extract specified additional field information from the multiple heterogeneous source data; The index calculation module is used to, when the index type of the target index is a derived index, sequentially perform statistical calculation methods of each calculation rule on the index calculation data in the order from the highest to the lowest calculation priority, and cache the intermediate result data and additional field information in the form of key-value pairs into a memory mapping structure until the index result data of the target index is generated.
[0012] Optionally, the system further includes a data preprocessing module, and the data preprocessing module includes a data cleaning sub-module, a data parsing sub-module, and a data mapping sub-module; The data cleaning sub-module is used to obtain multiple original heterogeneous source data from multiple data sources; perform data cleaning on the multiple original heterogeneous source data, filter out outlier data and duplicate data in the multiple original heterogeneous source data, and filter out non-critical feature data in the multiple original heterogeneous source data; The data parsing sub-module is used to parse the multiple original heterogeneous source data output by the data cleaning sub-module according to a predefined format, and the predefined format includes a table format and a JSON format; The data mapping sub-module is used to perform data mapping on the multiple original heterogeneous source data output by the data parsing sub-module according to the data format matched by the metadata model to generate multiple heterogeneous source data in a standardized format.
[0013] Optionally, the acquisition module is used to, when detecting that multiple heterogeneous source data in a standardized format are generated, obtain the multiple heterogeneous source data, and extract field metadata from the multiple heterogeneous source data through the metadata model; The acquisition module is further used to, when detecting that the index definition data corresponding to the target index is generated, compare the field metadata of the target index with the field metadata in the metadata model; in the case where the field metadata of the target index is inconsistent with the field metadata in the metadata model, perform partition replacement on the metadata model based on the version label of the field metadata.
[0014] In a second aspect, an embodiment of the present invention provides a configuration-based multi-level index calculation method, which is applied to the system as described in the first aspect, and the method includes: Obtain multiple heterogeneous source data, and extract field metadata of each heterogeneous source data by using a metadata model; Based on the index type of the target index, configure at least one calculation rule for the target index, and the calculation rule includes a statistical field, a field constraint operator, a field constraint value, a statistical calculation method, and a calculation priority; Match the field metadata of the multiple heterogeneous source data according to the statistical fields, field constraint operators, and field constraint values of each calculation rule, and determine the heterogeneous source data with successfully matched field metadata as the metric calculation data required for calculating the target metric; Sequentially perform the statistical calculation methods of the corresponding calculation rules on the metric calculation data in descending order of calculation priority to obtain the metric result data of the target metric.
[0015] The technical solutions provided in the embodiments of the present invention at least bring the following beneficial effects: Through the metadata model built in the acquisition module, the present invention can automatically extract the field metadata of heterogeneous data sources and uniformly map the original data to a standardized field identification system. It eliminates the burden of repeatedly writing cleaning and parsing codes for different data sources in the traditional method, and only requires dynamic updating of metadata to complete adaptation when adding new data sources, solving the conflict problem between field heterogeneity and real-time performance. The rule configuration module defines a general rule model including statistical fields, constraint conditions, calculation operators, and priorities, and for the first time realizes the configuration-based generation of atomic metrics without custom-developing data filtering logic. For composite metrics, it supports dynamic association calculations between atomic metrics through a priority rule chain, breaking through the bottleneck of the fragmented calculation process of multi-level metrics in the traditional solution. It can be seen that through the configuration-based architecture with multi-module collaboration, the present invention realizes the full-process dynamic generation from the original data to multi-level metrics, significantly improving the real-time performance and scalability of the metric system. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0017] Figure 1 is a block diagram of the architecture of a configuration-based multi-level metric calculation system provided by an embodiment of the present invention; Figure 2 is a schematic diagram of the complete application process of a configuration-based multi-level metric calculation system provided by an embodiment of the present invention; Figure 3 is a schematic diagram of the steps of a configuration-based multi-level metric calculation method provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. In the description of the embodiments of the present invention, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; herein, "and / or" is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c may mean: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c may be single or multiple.
[0019] In a big data - driven simulation scenario, dynamically changing situation indicators need to quickly respond to complex and changeable decision - making requirements. The traditional index calculation methods have the following core defects: First, atomic indicators rely on customized development, and the calculation logic is deeply coupled with the data source, resulting in high development costs and poor scalability; second, there is no unified calculation framework for composite / derived indicators, and it is difficult to flexibly combine multi - level indicators through configuration; finally, indicators are calculated in isolation, and it is impossible to build an association network to mine deep - level data value. These defects are particularly prominent in large - scale real - time simulation scenarios, seriously restricting situation awareness and decision - making support. The core idea of the present invention is to construct a multi - level index calculation system that runs through data processing, rule decoupling, and hierarchical calculation. The embodiments of the present invention realize the plug - and - play of multi - source data through a metadata - driven heterogeneous data dynamic adaptation mechanism. Design a three - level configuration rule model covering atomic, composite, and derived indicators, abstract the calculation logic into an orchestratable rule component, and adopt a priority - based streaming calculation architecture to support the multi - level derivation of complex indicators, effectively solving the technical bottlenecks existing in traditional methods.
[0020] Figure 1 is an architecture block diagram of a configuration - based multi - level index calculation system provided by an embodiment of the present invention, as Figure 1 shown. The system includes: a collection module, a rule configuration module, and an index calculation module; The collection module is configured with a metadata model; the collection module is used to obtain multiple heterogeneous source data and extract the field metadata of each heterogeneous source data by using the metadata model.
[0021] The acquisition module supports real-time acquisition of multiple heterogeneous source data from various data sources such as databases, APIs, and log files. Considering the defect of relevant technologies that there are calculation ambiguities caused by field conflicts or inconsistencies, the present invention pre-configures a metadata model in the acquisition module. The metadata model manages the heterogeneous source data collected by the acquisition module according to multiple dimensions. Specifically, the metadata model can be managed according to six dimensions: field code, field name, field description, field type, field location, and data source. Among them, the field location is recorded in the form of a path to generate a unique identifier for the field and eliminate the ambiguity in the calculation process.
[0022] The rule configuration module is used to configure at least one calculation rule for the target indicator based on the indicator type of the target indicator. The calculation rule includes a statistical field, a field constraint operator, a field constraint value, a statistical calculation method, and a calculation priority.
[0023] The calculation rule is used to indicate how to screen and calculate the indicator calculation data required for the target indicator, and is also used to indicate the calculation logic of the target indicator. The calculation rule is pre-configured by the rule configuration module according to the indicator type of the specified target indicator. In this embodiment, the indicator types include atomic indicators, composite indicators, and derived indicators. Next, these three indicator types will be described: An atomic indicator is the most basic indicator and is an indicator without cross-rule dependencies, such as the daily turnover and the actual number of people present on the same day; a composite indicator is an indicator generated by arithmetic operations depending on more than 2 atomic indicators, such as the quarterly turnover growth rate; a derived indicator is an indicator further obtained by arithmetic operations or conditional screening depending on more than 2 atomic indicators or composite indicators, such as the quarterly profit margin, which is calculated depending on the quarterly turnover (composite indicator) and the cost (atomic indicator), or the customer satisfaction index, which depends on multiple atomic indicators (such as customer ratings, service timeliness, customer feedback, etc.) and involves weighted averaging or conditional screening of different indicators.
[0024] Such as Figure 1As shown, the field metadata collected by the collection module is equivalent to providing a template for configuration items for the rule configuration module, indicating at which level of configuration items the rule configuration module specifically configures the calculation rules. Each calculation rule includes the following parts of configuration items: statistical field, field constraint operator, field constraint value, statistical calculation method, and calculation priority. Among them, the statistical field refers to the specific data field that needs to be statistically calculated during the calculation process. For example, in a sales dataset, the statistical field can be "sales amount" or "sales quantity". If you want to calculate the total sales amount, then "sales amount" is the statistical field; the field constraint operator is a symbol or keyword used to define the limiting conditions for the statistical field; the field constraint value is the specific value used in conjunction with the field constraint operator to define the filtering condition. If the field constraint operator is "greater than", then the field constraint value can be "1000". When performing statistical filtering on multiple heterogeneous source data, only the orders with a sales amount greater than 1000 will be statistically counted; the statistical calculation method refers to the specific statistical method or algorithm adopted during the calculation of the target indicator; the calculation priority refers to the number that determines the execution order among multiple calculation rules. The smaller the number of the calculation priority, the higher the priority. If there are two rules, the priority of rule A is 1, and the priority of rule B is 2. When performing calculations, rule A will be executed first, and then rule B will be executed.
[0025] The said indicator calculation module is used to match the field metadata of the multiple heterogeneous source data according to the statistical field, field constraint operator, and field constraint value of each calculation rule, and determine the heterogeneous source data with successfully matched field metadata as the indicator calculation data required for calculating the target indicator; in the order from high to low of the calculation priority, sequentially execute the statistical calculation methods of the corresponding calculation rules on the indicator calculation data to obtain the indicator result data of the target indicator.
[0026] The indicator calculation module is used to implement data filtering and data operation. According to the calculation rules provided by the rule configuration module, perform statistical calculations on multiple heterogeneous source data and generate the indicator result data of the target indicator. The indicator calculation module first filters out the qualified field metadata from the multiple heterogeneous source data as at least one indicator calculation data required for calculating the target indicator according to the statistical field, field constraint operator, and field constraint value specified in the calculation rule. The indicator calculation data serves as the calculation factor participating in the subsequent operation.
[0027] After the data screening is completed, according to the definition of the calculation rules, the index calculation module sequentially performs the statistical calculation methods in the corresponding calculation rules on at least one index calculation data in the order from high to low of the calculation priority. For example, operations such as summing, averaging, weighted averaging, maximum value, minimum value, etc. are performed on the index calculation data. The streaming calculation based on the calculation priority not only ensures the correct order of execution of multiple calculation rules but also avoids calculation conflicts or errors caused by complex dependencies between rules.
[0028] After all the calculation rules are executed, the index calculation module finally generates the index result data. The index result data includes not only the calculation results of each level of indexes (atomic indexes, composite indexes, and derivative indexes) but also the intermediate result data during the calculation process, which is convenient for subsequent decision support and data visualization work. The index result data can be output in a predetermined format and transmitted to the decision support system or the visualization platform for further analysis and display.
[0029] Through the metadata model built in the acquisition module, the present invention can automatically extract the field metadata of heterogeneous data sources and uniformly map the original data to a standardized field identification system. It eliminates the burden of repeatedly writing cleaning and parsing codes for different data sources in the traditional method, enabling only dynamic update of the metadata to complete the adaptation when adding new data sources and solving the problem of conflicts between field heterogeneity and real-time performance. The rule configuration module realizes the configuration-based generation of atomic indexes for the first time by defining a general rule model including statistical fields, constraint conditions, calculation operators, and priorities, without the need for custom development of data filtering logic. For composite indexes, it supports dynamic association calculation between atomic indexes through a priority rule chain, breaking through the bottleneck of the fragmented calculation process of multi-level indexes in the traditional solution. It can be seen that through the configuration-based architecture with multi-module collaboration, the present invention realizes the full-process dynamic generation from the original data to multi-level indexes, significantly improving the real-time performance and scalability of the index system.
[0030] In an optional implementation manner, the metadata model is used to extract the field code, field name, and field description of each heterogeneous source data, and the field code, field name, and field description of each heterogeneous source data are used as the field metadata of the heterogeneous source data.
[0031] In the management dimension of the metadata model, it may include field encoding, field name, and field description. Specifically, the field encoding is the unique identifier for each field of heterogeneous source data, such as "F001", which is used to locate the field and avoid conflicts of fields with the same name in different data sources (for example, the "ID" field in different databases may have different meanings); the field name is used to record the business name of the field (such as "sales amount", "timestamp"), which is convenient for users to intuitively understand the meaning of the field; the field description is used to supplement the business background, value range, or calculation logic of the field (such as "the timestamp format is UTC milliseconds", "the unit of the sales amount is RMB"), providing semantic support for subsequent data screening and calculation.
[0032] The said index calculation module is used for: according to the type of the field constraint operator of each calculation rule, parsing the field constraint value of each calculation rule to obtain the target constraint interval of each calculation rule, and the types of the field constraint operators include equal to, numerical interval, contains, greater than, less than, greater than or equal to, less than or equal to, and not equal to.
[0033] The types of field constraint operators include equal to, numerical interval, contains, greater than, less than, greater than or equal to, less than or equal to, and not equal to. The field constraint operator and the field constraint value jointly determine the target constraint interval for screening index calculation data. For example, when the type of the field constraint operator is equal to (=), the field constraint value is directly used as the matching condition (such as "entity classification = ship"); when the type of the field constraint operator is numerical interval (∈), the field constraint value is parsed into an interval range (such as "damage percentage ∈ (0, 20]"); when the type of the field constraint operator is contains (CONTAINS), the field constraint value is parsed into a set (such as "target ID contains A001, B002").
[0034] As described above, the logic for parsing the field constraint value according to the type of the field constraint operator is: dynamically converting the format of the constraint value according to the type of the constraint operator to obtain the target constraint interval of each calculation rule, which is equivalent to determining the matching range of each calculation rule.
[0035] Among the said multiple heterogeneous source data, the heterogeneous source data whose field encoding or field name contains the statistical field is determined as candidate calculation data, and the data value of each candidate calculation data is determined according to the field description of each candidate calculation data.
[0036] For example, if the statistical field is "damage percentage", traverse all heterogeneous source data, match the field code (such as "F005") or field name (such as "damage_percent"), and determine the heterogeneous source data with the field code or field name matching the statistical field as candidate calculation data, obtaining at least one candidate calculation data. This embodiment allows users to locate fields in either way of code or name, reducing the risk of configuration failure caused by inconsistent naming of data sources, and improving the success rate of field location with a dual-path matching mechanism. The present invention can adapt to the naming specifications of different data sources without forcing the unified field names.
[0037] Determine the heterogeneous source data with the data value within the target constraint interval as the heterogeneous source data with successful field metadata matching.
[0038] For example, if the field description defines that "the timestamp format is UTC milliseconds", parse "1630454400000" in the candidate calculation data into the time value "2021-09-01 00:00:00". Compare the parsed data value with the target constraint interval. For example, if the statistical field is "timestamp" and the target constraint interval is [9:00, 10:00], then only retain the candidate calculation data with the timestamp value within [9:00, 10:00], that is, filter out the heterogeneous source data with successful timestamp matching within [9:00, 10:00] from at least one candidate calculation data.
[0039] Through the precise screening and effective calculation of data, the present invention significantly improves the efficiency of situation awareness and decision support, providing users with more accurate and timely analysis results.
[0040] In an alternative embodiment, the system further includes an index definition module and an index recommendation module; The index definition module is configured to obtain the multiple heterogeneous source data and generate index definition data corresponding to the target index according to the multiple heterogeneous source data, where the index definition data is used to describe the target index in different dimensions.
[0041] Figure 2 It is a schematic diagram of the complete application process of a configuration-based multi-level index calculation system provided by an embodiment of the present invention. Please refer to Figure 2, the indicator definition module is used to generate indicator definition data for target indicators based on multiple heterogeneous source data. The indicator definition data is structured information describing the indicators, aiming to provide a comprehensive and accurate description of the indicators. Through a configuration-based method, the indicator definition module can dynamically generate indicator definition data for each indicator without manually writing complex code. The indicator definition data includes definition items in eight dimensions: indicator unique code, indicator name, short indicator name, measurement unit, indicator grouping, associated calculation rule identifier, indicator type, and source data. Among them, the associated calculation rule identifier represents which calculation rules the indicator is associated with, to indicate at least one indicator related to this indicator.
[0042] The indicator recommendation module is used to obtain the indicator definition data corresponding to each of the multiple indicators, and the indicator result data of each of the multiple indicators; determine the association relationships between the indicators according to the indicator definition data and the indicator result data, and output an indicator recommendation result representing the association relationships between the indicators.
[0043] The indicator recommendation module is used to generate a highly relevant indicator recommendation result based on the existing indicator definition data and indicator result data during the calculation and analysis process of the target indicator, to assist users in obtaining the information required for decision-making more quickly.
[0044] Based on the association calculation rule identifier of the data according to the metric definition of each metric, at least one calculation rule corresponding to the metric is determined. According to at least one calculation rule corresponding to the metric, the calculation process of the metric and the relevant data involved in the calculation process are determined, and the mutual associations between multiple metrics are analyzed in turn. For example, quarterly sales may be associated with other metrics such as customer satisfaction and market share. By analyzing the metric definition data and the metric result data obtained from historical calculations, these association relationships are identified. Specifically, based on the association calculation rule identifier of the data according to the metric definition of each metric, it is determined which calculation rules the metric is obtained by. Each calculation rule defines the relationship and calculation order between metrics. If a derived metric (such as quarterly profit margin) depends on multiple atomic metrics (such as quarterly turnover and cost), the metric structure data of this metric is directly associated with these atomic metrics. By performing a retrospective analysis on the metric definition data of the target metric and the metric result data obtained from historical calculations, it is identified which metrics often change together or have statistical correlations in a certain decision-making scenario. For example, assume that the user queries the quarterly profit margin metric. According to the association calculation rule identifier, it is determined that the quarterly profit margin depends on the quarterly turnover (composite metric) and cost (atomic metric) for calculation. And according to the metric result data, it is determined that there is a strong correlation between the quarterly turnover and customer satisfaction, and there is also an association relationship between the cost and production efficiency. Then relevant metrics are recommended for the quarterly profit margin metric, such as "quarterly turnover", "quarterly cost" and "customer satisfaction", to help the user further analyze these metrics in depth.
[0045] Finally, the metric recommendation result presented in the form of a table or report is output, clearly listing all relevant metrics and their calculation results, and sorting them according to the degree of association. Through the metric recommendation result, the user can quickly identify and access other key metrics closely related to the target metric, so as to achieve more efficient and comprehensive decision support.
[0046] In an alternative implementation, the metric recommendation module is configured with a graph database and a graph convolutional network model. Each node in the graph database represents a corresponding metric, and the graph convolutional network model is pre-trained according to the metric definition data corresponding to each of the multiple metrics and the metric result data corresponding to each of the multiple metrics.
[0047] In the metric recommendation module, the graph database represents the association relationships between metrics. Each metric will have a corresponding node in the graph database, and the association relationship between them is represented by an edge between the nodes. Specifically, each metric corresponds to a node in the graph database, and the attributes of the node include each definition item of the metric definition data of this metric.
[0048] The index definition module is used to associate the index definition data corresponding to each of the multiple indexes with the calculation rules of each of the multiple indexes, so as to obtain the associated calculation rule identifiers for each of the multiple indexes.
[0049] As described above, the index definition module is also used to associate the index definition data of multiple indexes with their corresponding calculation rules, and obtain the associated calculation rule identifier for each index. The associated calculation rule identifier is used as one of the definition items of the index definition data.
[0050] The index recommendation module is used to determine the association relationship between each index according to the associated calculation rule identifiers of each of the multiple indexes; calculate the association weights between each index according to the association relationship between each index, and establish the edges between each node in the graph database based on the association weights between each index.
[0051] The edges between index nodes represent the association relationship between them. The weight of the edge reflects the degree of correlation between the indexes. Through the associated calculation rule identifier, the dependency relationship chain between each calculation rule is analyzed, and the association weights between each index are further calculated, and these association relationships are reflected in the graph database through the weights of the edges. For example, if there is a strong association between the quarterly turnover and the quarterly profit margin (for example, the calculation of the profit margin depends on the turnover), then in the graph database, there will be a connecting edge between these two nodes, and the weight of the edge will be very high.
[0052] The index recommendation module is used to calculate the association strength between each index according to the nodes and edges in the graph database through the graph convolutional network model, and determine at least one set of associated indexes whose association strength is greater than a preset threshold; output a recommended index result including at least one set of associated indexes.
[0053] The graph convolutional network model is a method of deep learning based on the data structure of the graph database. The graph convolutional network model learns the relationship between nodes through graph-structured data, so as to be able to spread information between nodes and perform prediction and recommendation. In the embodiment of the present invention, the graph convolutional network model calculates the association strength between each index according to the nodes and edges in the graph database. Specifically, the input of the graph convolutional network model is the nodes and edges in the graph database. Through the graph convolutional network model, the interaction between each node in the graph is learned, and the information between nodes is spread, and then the association strength between each index is calculated. The graph convolutional network aggregates the information of adjacent nodes through multiple convolutional operations, so as to enhance the relationship characteristics between nodes. For example, if there is an association relationship between the quarterly turnover and the quarterly profit margin, the graph convolutional network can determine the association strength between the two and give a higher weight in subsequent calculations.
[0054] After the graph convolutional network model calculates the association strength between all metrics, through the operations of the graph convolutional network, the association strength between each metric is calculated, which characterizes the tightness of the association between two or more metrics. The higher the association strength, the more reference value the mutually associated metrics have in decision support. Only when the association strength between two or more metrics is greater than a preset threshold will it be considered that there is a significant association relationship between them. Based on the above judgment, one or more sets of associated metrics are determined, and each set of associated metrics includes two or more metrics.
[0055] Through the metric recommendation module, a recommended metric result including at least one set of associated metrics is output.
[0056] The present invention provides an intelligent and dynamically adaptable metric recommendation function by combining a graph database and a graph convolutional network model. It can identify the association relationships between metrics based on the defined data, calculation rules, and calculation results of the metrics, and perform accurate metric recommendations based on this. It not only improves the decision-making efficiency in a large-scale data environment but also provides advanced intelligent decision support means for various industries.
[0057] In an optional implementation manner, the metric types include atomic metrics, composite metrics, and derived metrics.
[0058] The rule configuration module is used to, when the metric type of the target metric is an atomic metric, configure a first-type calculation rule set containing one calculation rule for the target metric. The statistical calculation method of the first-type calculation rule set is configured as an atomic-level operator, and the atomic-level operator is any one of the original value, count, maximum value, minimum value, and average value.
[0059] When the target metric is an atomic metric, the rule configuration module configures a first-type calculation rule set containing only one calculation rule. The metric calculation module will perform the corresponding statistical calculation method on the metric calculation data of the target metric according to the first-type calculation rule set. During the calculation process, basic atomic-level operators such as the original value, count, maximum value, minimum value, and average value are executed on the metric calculation data to obtain the final metric result data. As mentioned above, the atomic-level operator is the basic calculation unit, which can quickly and efficiently process the data of atomic metrics. The original value is applicable to metrics that do not require calculation, such as the quantity of commodity inventory; the count is applicable to count-type data, for example, calculating the number of users visiting a website within a certain period; the maximum value and minimum value are applicable to scenarios where the extreme values of the data range need to be found, such as the highest temperature or the lowest temperature; the average value is applicable to scenarios where the data fluctuates less and the overall trend needs to be understood, such as the average monthly sales amount.
[0060] The index calculation module is configured to perform a statistical calculation method of a corresponding calculation rule on the index calculation data when the index type of the target index is an atomic index, so as to obtain the index result data of the target index.
[0061] The index calculation module generates the index result data of the target index by executing the corresponding calculation rule, that is, by executing the calculation rule in the first type of data set on the index calculation data.
[0062] In an alternative embodiment, the rule configuration module is configured to, when the index type of the target index is a composite index, configure a second type of calculation rule set including at least two calculation rules for the target index, where the statistical calculation method of each calculation rule is configured as an atomic-level operator or a composite-level operator, and the composite-level operator includes at least one of addition, subtraction, multiplication, and division.
[0063] In this embodiment, for the case where the target index is a composite index, the rule configuration module configures a second type of calculation rule set including at least two calculation rules. The statistical calculation method adopted by each calculation rule can be an atomic-level operator or a composite-level operator, which is used to implement the combined calculation of multiple atomic indexes or composite indexes, and supports more complex business requirements and analyses. The composite-level operator includes basic mathematical operations such as addition (+), subtraction (-), multiplication (×), and division (÷). The composite-level operator can flexibly combine multiple atomic indexes or composite indexes to calculate a new index result. For example, the results of multiple atomic indexes can be summed by addition, or the ratio of a certain index can be calculated by division to meet complex calculation requirements.
[0064] The rule configuration module is further configured to configure a corresponding calculation priority for each calculation rule in the second type of calculation rule set, and bind the calculation result of the calculation rule with the previous-level calculation priority as the input data of the calculation rule with the next-level calculation priority; Since multiple calculation rules are involved in the calculation process of the composite index, the execution order of the calculation rules also needs to be strictly specified. The rule configuration module also configures a calculation priority for each calculation rule to indicate the execution order of each calculation rule in the calculation process. The calculation priority of each calculation rule is identified by a number, and the smaller the priority number, the higher the execution priority of the rule. During execution, the calculation rules will be executed in order from the highest to the lowest calculation priority. In addition, the rule configuration module binds the calculation result of each level of calculation rule with the input data of the next-level calculation rule. Specifically, the calculation result of the previous-level calculation rule will be used as the input data of the next-level calculation rule.
[0065] For example, assume that the target metric is the sales growth rate for a certain quarter, which depends on two composite metrics: the sales amount for this quarter and the sales amount for the previous quarter. A second type of calculation rule set containing three calculation rules is pre-configured. Among them, Rule A is used to calculate the sales amount for this quarter, Rule B is used to calculate the sales amount for the previous quarter, and Rule C is used to calculate ((sales amount for this quarter - sales amount for the previous quarter) ÷ sales amount for the previous quarter) × 100%. In this example, the calculation priorities of Rule A and Rule B can be the same or different, which does not affect the final calculation result. However, the calculation priority of Rule C should be higher than that of Rule A and Rule B. The calculation results of Rule A and Rule B will be used as the input data for Rule C.
[0066] The metric calculation module is used to, when the metric type of the target metric is a composite metric, sequentially perform the statistical calculation methods of each calculation rule on the metric calculation data in the order from high to low of the calculation priority, and cache the intermediate result data into the memory mapping structure until the metric result data of the target metric is generated.
[0067] To improve the calculation efficiency and ensure that subsequent calculations are not affected by previous calculation errors, the metric calculation module caches the intermediate result data after each calculation into the memory mapping structure. This can effectively reduce the time for repeated data calculations and ensure data integrity and consistency during the calculation process. The cached intermediate result data will be continuously used by subsequent calculation rules according to the dependency order. After all calculation rules are sequentially executed, the metric calculation module will generate the metric result data of the target metric, which not only includes the final calculation result of the composite metric but also includes all intermediate result data during the calculation process.
[0068] In an optional implementation manner, the rule configuration module is used to, when the metric type of the target metric is a derived metric, configure a third type of calculation rule set containing at least two calculation rules for the target metric, where the statistical calculation method of at least one calculation rule is configured as a derived-level operator, and the derived-level operator includes distinct count.
[0069] In the calculation of derived metrics, the rule configuration module configures a third type of calculation rule set containing at least two calculation rules according to the specific requirements of the target metric. The calculation rules in the third type of calculation rule set involve complex data processing, including derived-level operators such as distinct count. The distinct count operation is used to count the number of unique values in a field. For example, when analyzing the number of unique purchasing customers for a certain product, the distinct count can ensure that each customer is only counted once, rather than being counted multiple times due to multiple purchases. In this case, the role of the derived-level operator is not limited to common mathematical operations but also extends to operations such as further filtering, grouping, and statistical operations on data, such as de-duplication and conditional screening.
[0070] The rule configuration module is further configured to configure corresponding calculation priorities for each calculation rule in the third type of calculation rule set, and configure an additional field list, where the additional field list is used to extract specified additional field information from the multiple heterogeneous source data.
[0071] For the calculation of derived metrics, the rule configuration module not only needs to configure the calculation priority, but also needs to configure the additional field list. The additional field list specifies which additional data fields (additional field information) need to be extracted from the multiple heterogeneous source data and used as auxiliary data during the calculation process. The additional field information involves more context information or provides important background data support for the derived metrics. For example, when calculating the derived metric of customer satisfaction, the additional field list may include additional field information such as "customer feedback date" and "customer service rating", which are used for auxiliary calculation or as filtering conditions to affect the calculation results. The additional field information makes the calculation process more accurate, while providing more data levels and perspectives, enhancing the interpretability of the results.
[0072] The metric calculation module is configured to, when the metric type of the target metric is a derived metric, sequentially perform the statistical calculation methods of each calculation rule on the metric calculation data in descending order of the calculation priority, and cache the intermediate result data and the additional field information in the form of key-value pairs into a memory mapping structure until the metric result data of the target metric is generated.
[0073] Similar to the calculation logic of composite atomic metrics, during the calculation process where the target metric is a derived metric, the metric calculation module sequentially executes the statistical calculation methods of each calculation rule in descending order of the configured calculation priority. When each calculation rule is executed, in addition to processing the main calculation data, the additional field information related to the rule is also processed. The additional field information is stored in the form of key-value pairs and cached into the memory mapping structure, where the intermediate result data serves as the key of the key-value pair and the additional field information serves as the value of the key-value pair. The intermediate result data and the additional field information after each calculation are stored according to the execution order of the calculation rules. When all the calculation rules in the third type of calculation rule set corresponding to the target metric of the derived metric type are executed, the final target metric result data is generated, which includes the intermediate result data and the attachment field information, facilitating subsequent decision support and data display.
[0074] In an alternative embodiment, the system further includes a data preprocessing module, and the data preprocessing module includes a data cleaning sub-module, a data parsing sub-module, and a data mapping sub-module; The data cleaning sub-module is used to obtain multiple original heterogeneous source data from multiple data sources; perform data cleaning on the multiple original heterogeneous source data, filter out outlier data and duplicate data in the multiple original heterogeneous source data, and filter out non-critical feature data in the multiple original heterogeneous source data.
[0075] Please refer to Figure 2 , in this embodiment, the data cleaning sub-module is used to obtain original heterogeneous data from multiple data sources. The data sources include various devices or platforms that generate data required for calculating metrics, such as flight logs, sensors, and background monitoring programs. The data cleaning sub-module identifies and filters out outlier data in the multiple original heterogeneous source data, such as input errors, data points outside the normal range, or obviously unreasonable data points. In addition, the data cleaning sub-module eliminates duplicate data to prevent the same data from being calculated multiple times, thereby ensuring the accuracy and effectiveness of the calculation results. At the same time, the data cleaning sub-module deletes non-critical feature data from the multiple original heterogeneous source data, that is, those data fields that have no substantial impact on subsequent calculations and decision support, thereby reducing the interference of irrelevant information and improving the efficiency of subsequent data analysis. Optionally, the processing method for outliers can be a statistical method, such as the standard deviation method, the quartile method, etc.; the removal of duplicate data can be achieved by matching the unique identifiers of the data to ensure that each data record is only processed once; the screening of non-critical feature data can be implemented based on feature selection techniques, and the present invention does not limit this.
[0076] The data parsing sub-module is used to parse the multiple original heterogeneous source data output by the data cleaning sub-module according to a predefined format, and the predefined format includes a table format and a JSON format.
[0077] After data cleaning, the multiple original heterogeneous source data output by the data cleaning sub-module is input into the data parsing sub-module. The data parsing sub-module is used to perform format conversion on the cleaned data to make it conform to a predefined standard format. The predefined formats include a table format (such as CSV, Excel table, etc.) and a JSON format. Convert the irregular data formats (such as log files, XML formats, binary data, etc.) in the multiple original heterogeneous source data into a standard format that is easy to process and calculate, thereby realizing the structured management of data.
[0078] The data mapping sub-module is used to perform data mapping on the multiple original heterogeneous source data output by the data parsing sub-module according to the data format matched by the metadata model, and generate multiple heterogeneous source data in a standardized format.
[0079] After the data parsing is completed, the data mapping sub-module takes over and further converts the parsed data into a data format that conforms to a unified standard. According to the metadata model described above, mapping operations are performed on the parsed data to generate heterogeneous source data in a standardized format. For example, assume that the original heterogeneous source data contains fields such as "User ID" and "Purchase Amount", but in different original heterogeneous source data, these two fields may use different naming rules or units (such as "User_ID" and "Amount"). The data mapping sub-module maps these fields to standard field names according to the dimensions of the metadata model and processes the conversion of different units to ensure data consistency and comparability. In this way, the multi-level indicator calculation system provided in this embodiment can process heterogeneous source data from different data sources. In a specific implementation, for example, in a simulation scenario that relies on a large amount of simulation data, one frame of original heterogeneous source data is collected per second following the simulation process, and the originally collected heterogeneous source data forms standardized and configurable heterogeneous source data through the above steps of data cleaning, data parsing, and data mapping.
[0080] In an alternative implementation manner, the acquisition module is configured to, when detecting the generation of a plurality of heterogeneous source data in a standardized format, acquire the plurality of heterogeneous source data and extract field metadata from the plurality of heterogeneous source data through the metadata model.
[0081] It can be understood that the acquisition process of the acquisition module in this embodiment for a plurality of heterogeneous source data is dynamic, and the acquisition timing is triggered at different times according to different acquisition sources. On the one hand, when detecting the generation of a plurality of heterogeneous source data in a standardized format, the heterogeneous source data is automatically acquired (such as Figure 2 the path of active metadata acquisition shown).
[0082] The acquisition module is further configured to, when detecting the generation of the indicator definition data corresponding to the target indicator, compare the field metadata of the target indicator with the field metadata in the metadata model; in the case where the field metadata of the target indicator is inconsistent with the field metadata in the metadata model, perform partition replacement on the metadata model based on the version label of the field metadata.
[0083] On the other hand, when detecting the generation of the definition data of the target indicator, the acquisition module automatically acquires the definition data of the target indicator and further compares it with the field metadata in the metadata model based on the target indicator data.
[0084] If the acquisition module discovers an inconsistency between the field metadata of the target metric and the field metadata in the metadata model during the comparison process, it initiates further correction operations. Specifically, the acquisition module partitions and replaces the metadata model based on the version tag in the field metadata. The version tag of the field metadata is used to identify different versions of the field. In the case of an inconsistency between the field metadata of the target metric and the field metadata in the metadata model, the acquisition module determines whether to replace or update the corresponding field according to the indication of the version tag. At this time, only the mismatched field metadata is replaced, rather than reconstructing the entire metadata model.
[0085] It can be seen that the acquisition module can not only accurately extract and process heterogeneous source data, but also ensure that when new target metrics are introduced, it guides the metadata model to adapt and update in a timely manner, thus maintaining the efficient operation of the overall system and data consistency.
[0086] Figure 3 It is a schematic diagram of the steps of a configuration-based multi-level metric calculation method provided by an embodiment of the present invention, which is applied to a configuration-based multi-level metric calculation system as described above. The method includes: Step S101, obtain multiple heterogeneous source data, and use the metadata model to extract the field metadata of each heterogeneous source data.
[0087] The acquisition module obtains heterogeneous source data from multiple data sources. The data sources can include, but are not limited to, various forms such as databases, APIs, log files, etc. The built-in metadata model of the acquisition module automatically processes the multiple heterogeneous source data collected, and extracts the field metadata in each heterogeneous source data.
[0088] Step S102, based on the metric type of the target metric, configure at least one calculation rule for the target metric. The calculation rule includes a statistical field, a field constraint operator, a field constraint value, a statistical calculation method, and a calculation priority.
[0089] Based on the type of the target metric, the rule configuration module pre-designs a calculation rule for the target metric. The calculation rule is used to define how to screen the metric calculation data required for calculating the target metric from multiple heterogeneous source data, and how to perform operations on the metric calculation data, that is, to specify the calculation logic of the metric calculation data.
[0090] Step S103, according to the statistical field, the field constraint operator, and the field constraint value of each calculation rule, match the field metadata of the multiple heterogeneous source data, and determine the heterogeneous source data with successfully matched field metadata as the metric calculation data required for calculating the target metric.
[0091] The indicator calculation module filters out eligible indicator calculation data from multiple heterogeneous source data according to the calculation rules configured by the rule configuration module. First, according to the statistical fields, field constraint operators, and field constraint values of each calculation rule, the field metadata of the collected multiple heterogeneous source data is matched. By parsing the combined conditions of the field constraint operators and field constraint values, the eligible data is used as the indicator calculation data. The successfully matched heterogeneous source data is used as the data available for the target indicator calculation and enters the subsequent calculation stage.
[0092] Step S104, in the order from the highest to the lowest calculation priority, sequentially perform the statistical calculation methods of the corresponding calculation rules on the indicator calculation data to obtain the indicator result data of the target indicator.
[0093] The indicator calculation module performs calculation processing on the data according to the selected indicator calculation data and calculation rules to obtain the final result of the target indicator. According to the calculation priority defined in the rule configuration module, all calculation rules are sorted, and the rules with higher priority are executed first. According to the statistical calculation methods defined in the calculation rules, statistical operations are performed on the eligible data. After sequential calculations, the indicator result data of the target indicator is output. The indicator result data not only includes the calculation results of each level of indicators (atomic indicators, composite indicators, and derived indicators), but also contains the intermediate result data during the calculation process, which is convenient for further decision-making support and data analysis.
[0094] In the embodiment of the present invention, through the metadata model built in the acquisition module, the field metadata of heterogeneous data sources can be automatically extracted, and the original data is uniformly mapped to a standardized field identification system. It eliminates the burden of repeatedly writing cleaning and parsing codes for different data sources in the traditional method, and only requires dynamic update of metadata when adding new data sources to complete adaptation, solving the conflict problem between field heterogeneity and real-time performance. The rule configuration module realizes the configuration-based generation of atomic indicators for the first time by defining a general rule model including statistical fields, constraint conditions, calculation operators, and priorities, without custom development of data filtering logic. For composite indicators, it supports dynamic association calculation between atomic indicators through a priority rule chain, breaking through the bottleneck of the fragmented calculation process of multi-level indicators in the traditional solution. It can be seen that the present invention realizes the full-process dynamic generation from raw data to multi-level indicators through a configuration-based architecture with multi-module collaboration, significantly improving the real-time performance and scalability of the indicator system.
[0095] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, devices, electronic devices, and storage media. Therefore, the embodiments of the present invention can take the form of all-hardware embodiments, all-software embodiments, or embodiments combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0096] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods and devices according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, such that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0097] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
[0098] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or terminal device. Without further limitation, elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article, or terminal device including the said elements. The above has introduced in detail a configuration-based multi-level index calculation system and method provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A configuration-based multi-level index calculation system, characterized in that The system includes: a collection module, a rule configuration module, and a metric calculation module; The collection module is configured with a metadata model; the collection module is used to obtain multiple heterogeneous source data and extract field metadata of each heterogeneous source data by using the metadata model; The rule configuration module is used to configure at least one calculation rule for the target metric based on the metric type of the target metric, and the calculation rule includes a statistical field, a field constraint operator, a field constraint value, a statistical calculation method, and a calculation priority; The metric calculation module is used to match the field metadata of the multiple heterogeneous source data according to the statistical field, the field constraint operator, and the field constraint value of each calculation rule, and determine the heterogeneous source data with successfully matched field metadata as the metric calculation data required for calculating the target metric; and sequentially execute the statistical calculation methods of the corresponding calculation rules on the metric calculation data in descending order of the calculation priority to obtain the metric result data of the target metric.
2. The system according to claim 1, wherein The metadata model is used to extract the field code, field name, and field description of each heterogeneous source data, and use the field code, field name, and field description of each heterogeneous source data as the field metadata of the heterogeneous source data; The metric calculation module is used to: Parse the field constraint value of each calculation rule according to the type of the field constraint operator of each calculation rule to obtain the target constraint interval of each calculation rule, and the types of the field constraint operators include equal to, numerical interval, contains, greater than, less than, greater than or equal to, less than or equal to, and not equal to; Among the multiple heterogeneous source data, determine the heterogeneous source data whose field code or field name contains the statistical field as candidate calculation data, and determine the data value of each candidate calculation data according to the field description of each candidate calculation data; Determine the heterogeneous source data with successfully matched field metadata as the candidate calculation data whose data value is within the target constraint interval.
3. The system according to claim 2, wherein The system further includes a metric definition module and a metric recommendation module; The metric definition module is used to obtain the multiple heterogeneous source data and generate metric definition data corresponding to the target metric according to the multiple heterogeneous source data, and the metric definition data is used to describe the target metric in different dimensions; The metric recommendation module is used to obtain the metric definition data corresponding to each of the multiple metrics and the metric result data corresponding to each of the multiple metrics; Determine the association relationship between the metrics according to the metric definition data and the metric result data, and output a metric recommendation result representing the association relationship between the metrics.
4. The system according to claim 3, wherein The metric recommendation module is configured with a graph database and a graph convolutional network model, each node in the graph database represents a corresponding metric, and the graph convolutional network model is pre-trained according to the metric definition data corresponding to each of the multiple metrics and the metric result data corresponding to each of the multiple metrics; The metric definition module is used to associate the metric definition data corresponding to each of the multiple metrics with the calculation rules of each of the multiple metrics to obtain the association calculation rule identifiers corresponding to each of the multiple metrics. The index recommendation module is used to determine the association relationships among the various indexes according to the respective association calculation rule identifications of the multiple indexes; calculate the association weights among the various indexes according to the association relationships among the various indexes, and establish the edges between the nodes in the graph database based on the association weights among the various indexes; The index recommendation module is used to calculate the association strengths among the various indexes through the graph convolutional network model according to the nodes and edges in the graph database, and determine at least one set of associated indexes whose association strengths are greater than a preset threshold; output a recommended index result including at least one set of associated indexes.
5. The system according to claim 1, wherein The index types include atomic indexes, composite indexes, and derivative indexes; The rule configuration module is used to, when the index type of the target index is an atomic index, configure a first-type calculation rule set including one calculation rule for the target index, and the statistical calculation method of the first-type calculation rule set is configured as an atomic-level operator, and the atomic-level operator is any one of a raw value, a count, a maximum value, a minimum value, and an average value; The index calculation module is used to, when the index type of the target index is an atomic index, perform the statistical calculation method of the corresponding calculation rule on the index calculation data to obtain the index result data of the target index.
6. The system according to claim 5, wherein The rule configuration module is used to, when the index type of the target index is a composite index, configure a second-type calculation rule set including at least two calculation rules for the target index, wherein the statistical calculation method of each calculation rule is configured as an atomic-level operator or a composite-level operator, and the composite-level operator includes at least one of addition, subtraction, multiplication, and division; The rule configuration module is further used to configure corresponding calculation priorities for each calculation rule in the second-type calculation rule set, and bind the calculation result of the calculation rule with the previous-level calculation priority as the input data of the calculation rule with the next-level calculation priority; The index calculation module is used to, when the index type of the target index is a composite index, sequentially perform the statistical calculation methods of the respective calculation rules on the index calculation data in the order from high to low of the calculation priorities, and cache the intermediate result data into a memory mapping structure until the index result data of the target index is generated.
7. The system according to claim 5, wherein The rule configuration module is used to, when the index type of the target index is a derivative index, configure a third-type calculation rule set including at least two calculation rules for the target index, wherein the statistical calculation method of at least one calculation rule is configured as a derivative-level operator, and the derivative-level operator includes a distinct count; The rule configuration module is further used to configure corresponding calculation priorities for each calculation rule in the third-type calculation rule set, and configure an additional field list, and the additional field list is used to extract specified additional field information from the multiple heterogeneous source data; The index calculation module is used to, when the index type of the target index is a derived index, sequentially perform the statistical calculation methods of each calculation rule on the index calculation data in the order from high to low of the calculation priority, and cache the intermediate result data and additional field information in the form of key-value pairs into the memory mapping structure until the index result data of the target index is generated.
8. The system according to claim 1, characterized in that, The system further includes a data preprocessing module, and the data preprocessing module includes a data cleaning sub-module, a data parsing sub-module, and a data mapping sub-module; The data cleaning sub-module is used to obtain a plurality of original heterogeneous source data from multiple data sources; Perform data cleaning on the plurality of original heterogeneous source data, filter out the outlier data and duplicate data in the plurality of original heterogeneous source data, and filter out the non-critical feature data in the plurality of original heterogeneous source data; The data parsing sub-module is used to parse the plurality of original heterogeneous source data output by the data cleaning sub-module according to a predefined format, and the predefined format includes a table format and a JSON format; The data mapping sub-module is used to perform data mapping on the plurality of original heterogeneous source data output by the data parsing sub-module according to the data format matched by the metadata model to generate a plurality of heterogeneous source data in a standardized format.
9. The system according to claim 3 or 8, wherein The acquisition module is used to, when detecting the generation of a plurality of heterogeneous source data in a standardized format, obtain the plurality of heterogeneous source data, and extract field metadata from the plurality of heterogeneous source data through the metadata model; The acquisition module is further used to, when detecting the generation of the index definition data corresponding to the target index, compare the field metadata of the target index with the field metadata in the metadata model; In the case where the field metadata of the target index is inconsistent with the field metadata in the metadata model, perform partition replacement on the metadata model based on the version label of the field metadata.
10. A configuration-based multi-level index calculation method, characterized in that, Applied to the system according to any one of claims 1-9, the method includes: Obtain a plurality of heterogeneous source data, and extract the field metadata of each heterogeneous source data by using the metadata model; Based on the index type of the target index, configure at least one calculation rule for the target index, and the calculation rule includes a statistical field, a field constraint operator, a field constraint value, a statistical calculation method, and a calculation priority; According to the statistical field, the field constraint operator, and the field constraint value of each calculation rule, match the field metadata of the plurality of heterogeneous source data, and determine the heterogeneous source data with successfully matched field metadata as the index calculation data required for calculating the target index; Sequentially perform the statistical calculation methods of the corresponding calculation rules on the index calculation data in the order from high to low of the calculation priority to obtain the index result data of the target index.
Citation Information
Patent Citations
Real-time index calculation method, system and equipment based on streaming data and medium
CN116049285A
Index calculation method and device based on multiple data tables
CN117827841A
Calculation method and device of index data, equipment, storage medium and program product
CN118569733A
Business data processing method and device, equipment, storage medium and program product
CN118861378A
Container microservice-oriented performance monitoring and alarm method and alarm system
WO2023142054A1
Cited By
Building block model-based index generation method, execution method and program product
CN121070335A
A block model-based index generation method, execution method, and program product
CN121070335B