Data warehouse index monitoring method and device, electronic equipment and storage medium
By rating and prioritizing the correlation and importance of data tables and fields in the data warehouse, we generate indicator monitoring recommendation results, solving the problems of fuzzy and unpredictable characteristics in data warehouse indicator monitoring, and achieving efficient and accurate indicator monitoring.
Patent Information
- Application Number
- CN202510330275.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-20
AI Technical Summary
Data warehouses often experience blurred indicator characteristics and unpredictable indicators during indicator monitoring, resulting in a high failure rate of monitoring indicators.
By determining the relevance score and importance score of the data table based on the information of the data warehouse and the content of the data table, determining the priority score of the field based on these scores, inputting a pre-trained indicator monitoring recommendation model, generating indicator monitoring recommendation results, and then conducting indicator monitoring of the data warehouse.
The efficiency and quality of indicator monitoring tasks are optimized, the accuracy of monitoring indicators is improved, redundant monitoring and resource waste are reduced, labor costs are reduced, and dynamic adjustment and optimization of indicator monitoring tasks are achieved.
Smart Images

Figure CN120180038A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, particularly to data analysis, data mining, etc., and can be used in application scenarios such as information recommendation. Specifically, it relates to a method, apparatus, electronic device, and storage medium for monitoring data warehouse metrics. Background Art
[0002] With the enterprise's need for data analysis and operation management, the data warehouse has gradually become an essential system module for enterprise operation and development, helping the enterprise collect various upstream data, generating rich-dimensional data metrics through technical means, providing management and operation data from different perspectives and dimensions, and assisting the enterprise in making management decisions. The data warehouse usually has a large amount of data, and situations such as fuzzy metric features and unpredictable metrics often occur during metric monitoring, resulting in a high failure rate of monitored metrics. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, electronic device, and storage medium for monitoring data warehouse metrics.
[0004] According to a first aspect of the present disclosure, there is provided a method for monitoring data warehouse metrics. The data warehouse includes at least one data table, and each data table includes at least one field. The method includes: determining the association score and importance score of a data table according to the information of the data warehouse and the content of the data table; determining the priority score of a field in any data table according to the information of the data warehouse, the association score, and the importance score; inputting the association score, importance score, and priority score into a pre-trained metric monitoring recommendation model to generate a metric monitoring recommendation result; and monitoring the data warehouse based on the metric monitoring recommendation result.
[0005] According to a second aspect of the present disclosure, there is provided a data warehouse metric monitoring apparatus applied to a data warehouse. The data warehouse includes at least one data table, and each data table includes at least one field. The apparatus includes: a table scoring module for determining the association score and importance score of a data table according to the information of the data warehouse and the content of the data table; a field scoring module for determining the priority score of a field in any data table according to the information of the data warehouse, the association score, and the importance score; a monitoring recommendation module for inputting the association score, importance score, and priority score into a pre-trained metric monitoring recommendation model to generate a metric monitoring recommendation result; and a metric monitoring module for monitoring the data warehouse based on the metric monitoring recommendation result.
[0006] According to a third aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any method in the embodiments of the present disclosure.
[0007] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.
[0008] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program which, when executed by a processor, implements any method in the embodiments of the present disclosure.
[0009] Adopting the solution of the present disclosure can optimize the efficiency and quality of the index monitoring task.
[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0012] Figure 1 is a schematic flowchart of a data warehouse index monitoring method according to an embodiment of the present disclosure;
[0013] Figure 2 is a schematic structural diagram of a data warehouse index monitoring device according to an embodiment of the present disclosure;
[0014] Figure 3 is a schematic scenario diagram of a data warehouse index monitoring method according to an embodiment of the present disclosure;
[0015] Figure 4 is a structural diagram of an electronic device for implementing the data warehouse index monitoring method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to help understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, the description below omits descriptions of well-known functions and structures.
[0017] As used herein, the term "and / or" is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set composed of A, B, and C. The terms "first" and "second" as used herein refer to and distinguish multiple similar technical terms, and do not mean to limit the order or limit to only two. For example, the first feature and the second feature refer to two categories / two features, the first feature can be one or more, and the second feature can also be one or more.
[0018] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following specific embodiments. Those skilled in the art should understand that the present disclosure can also be implemented without certain specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0019] Before introducing the technical solutions of the embodiments of the present disclosure, further explanations are made on the technical terms that may be used in the present disclosure:
[0020] Data warehouse: A system that stores a large amount of structured data, usually used to support data analysis and decision-making. It integrates data from different business systems, and after cleaning and transformation, stores them according to themes, facilitating users to query and analyze. In particular, the data warehouse is the core of organizing data resources, especially suitable for the big data environment.
[0021] Data table: The basic unit for storing data in a data warehouse, similar to a two-dimensional table, composed of rows and columns. In particular, each table usually represents a theme or entity.
[0022] Field: An element that makes up a data table, that is, a column in the table. Each field represents a certain type of data, and the types of fields can usually include numbers, texts, dates, etc.
[0023] In the related art, for a data warehouse with a large amount of data, when using a common table creation and layering system, thousands of tables are often maintained in the data warehouse at the same time. And the table with the largest proportion among a large number of tables is the wide table. That is, in order to accelerate the query speed, data from multiple business domains is associated into a large wide table, and the final wide table may have thousands of fields at most. For a large number of tables with a large number of fields, it is too much work to establish monitoring for each data warehouse indicator, and it is almost completely infeasible. At the same time, even if it is known that a certain indicator of a certain table is important and needs to be monitored, but it is not known how to monitor it, and it is impossible to adjust and optimize the monitoring plan in combination with the real-time evolution and changes of the indicator. Further, it is impossible to implement hierarchical management of data warehouse indicators, automatically adjust monitoring strategies for indicators with different characteristics and management strategies, and it is also impossible to predict faulty indicators in advance.
[0024] In order to at least partially solve one or more of the above problems and other potential problems, the present disclosure proposes a data warehouse indicator monitoring method, which can optimize the efficiency and quality of indicator monitoring tasks.
[0025] An embodiment of the present disclosure provides a data warehouse indicator monitoring method. Figure 1 It is a schematic flowchart of the data warehouse indicator monitoring method according to an embodiment of the present disclosure. This data warehouse indicator monitoring method can be applied to a data warehouse indicator monitoring device. The data warehouse indicator monitoring device is located in an electronic device. The electronic device includes but is not limited to a fixed device and / or a mobile device. For example, the fixed device includes but is not limited to a server, and the server can be a cloud server or a general server. For example, the mobile device includes but is not limited to an information processing device, and the information processing device can be a mobile phone, a tablet computer, etc. In some possible implementation manners, this data warehouse indicator monitoring method can also be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 1 shown, this data warehouse indicator monitoring method includes:
[0026] S101. Determine the correlation score and importance score of the data table according to the information of the data warehouse and the content of the data table.
[0027] S102. Determine the priority score of the fields in any data table according to the information of the data warehouse, the correlation score, and the importance score.
[0028] S103. Input the correlation score, importance score, and priority score into a pre-trained indicator monitoring recommendation model to generate an indicator monitoring recommendation result.
[0029] S104. Monitor the data warehouse based on the indicator monitoring recommendation result.
[0030] Here, the correlation score can measure the correlation between data tables, that is, whether any data table has a strong connection with other data tables, and can reflect whether the data table can provide additional information through the association relationship. In the embodiments of the present disclosure, the correlation score can be calculated by analyzing the foreign key relationship between data tables, the relevance of field values, or the logical relationship between data tables defined by business rules.
[0031] Here, the importance score can measure the criticality of any data table in business analysis or data monitoring tasks, and can reflect the influence of the data table on business decisions. In the embodiments of the present disclosure, the importance score can be calculated according to criteria such as business rules, the usage frequency of the data table, and the value of the data table in the business scenario.
[0032] In the embodiments of the present disclosure, the information in the data warehouse can be used to analyze and calculate the correlation score, importance score, priority score, etc. Exemplarily, the information in the data warehouse can include the structural information of the data warehouse such as table structure, inter-table relationship, field attributes, partition information, etc., and can also include metadata information such as metadata of data tables and index information. Further, the correlation score of the data table can be determined according to the information in the data warehouse and the content of the data table. Exemplarily, the correlation score can be calculated based on the table structure by using the inter-table relationship in the data warehouse; the correlation score can also be calculated based on data statistics by using the correlation between data table quality inspections; the correlation score can also be preset based on the logical relationship defined by the business. Particularly, the calculation method of the correlation score can be adjusted according to the actual situation and in accordance with the real-time business requirements. Furthermore, the importance score of the data table can be determined according to the information in the data warehouse and the content of the data table. Exemplarily, the frequency of the data table appearing in various query or analysis tasks can be counted, and the importance score can be calculated based on the usage frequency; the influence of the data table on the business objective can also be analyzed by methods such as decision tree or feature importance analysis, and the importance score can be calculated based on the business impact. Particularly, the calculation method of the importance score can be adjusted according to the actual situation and in accordance with the real-time business requirements.
[0033] Here, the priority score combines the correlation score and the importance score, and is used to determine the priority order of fields in the indicator monitoring recommendation. In the embodiments of the present disclosure, the priority score can reflect the necessity of the field being recommended in the analysis.
[0034] In the embodiments of the present disclosure, for any data table, the priority score can be determined according to the information in the data warehouse, the correlation score, and the importance score. Exemplarily, the priority score can be generated by using a weighted or standardized method according to the correlation score and the importance score; a model can also be trained using the scoring historical data to predict the priority score of the field. Particularly, the calculation method of the priority score can be adjusted according to the actual situation and in accordance with the real-time business requirements.
[0035] Here, the index monitoring recommendation model is a machine learning or deep learning model that has been pre-trained and can generate recommendation results for index monitoring based on the input scores. In the embodiments of the present disclosure, the index monitoring recommendation model can be a model based on regression or classification algorithms, or a complex recommendation model such as collaborative filtering or content-based recommendation model.
[0036] In the embodiments of the present disclosure, the relevance score, importance score, and priority score can be first converted into an input format adapted to the model, and then the model recommends the most suitable metrics to be monitored according to the preset recommendation algorithm based on the input scores.
[0037] The technical solution of the embodiments of the present disclosure can make the recommended monitoring metrics more accurate by using data metrics such as relevance and importance scores, improve the decision-making quality, and enhance the practicality of the monitoring solution. Through intelligent recommendation, redundant monitoring and resource waste can be reduced, and at the same time, the labor cost of monitoring metrics can be lowered. With scoring and model recommendation, flexible hierarchical management of metrics can be achieved, and the monitoring metrics can be dynamically adjusted in actual application scenarios to adapt to business changes, so as to optimize the efficiency and quality of the index monitoring task. Based on the index monitoring recommendation results to monitor the data warehouse, not only can an index monitoring system for the data warehouse be constructed, but also the self-adaptive optimization of the monitoring dimension can be realized, and finally the accuracy and timeliness of index monitoring can be improved.
[0038] In some embodiments, the data warehouse index monitoring method further includes: obtaining historical index monitoring records and historical index monitoring results; constructing training samples according to the historical index monitoring records and historical index monitoring results; and using the training samples to train an initial large model to generate a pre-trained index monitoring recommendation model.
[0039] Here, the historical index monitoring record is the detailed data of the index monitoring tasks that have been carried out before, including the specific fields monitored, the time range monitored, the monitoring methods used, and the data content involved in the monitoring process, etc. In the embodiments of the present disclosure, the historical index monitoring record can indicate which metrics have been monitored with emphasis, how the monitoring method is, and the historical relevance of the monitoring data, providing a basis for model training.
[0040] Here, the historical index monitoring result is the output result of the index monitoring task, including whether the monitoring target is completed, whether the monitoring data meets the expected effect, and whether any abnormalities or trend changes are found. In the embodiments of the present disclosure, the historical index monitoring result can help evaluate the effectiveness of the monitoring task and provide feedback data for the training of the subsequent recommendation model.
[0041] In the embodiments of the present disclosure, the data of historical monitoring tasks can be first extracted from a database or a data warehouse, including monitoring records and analysis results of monitoring tasks. In particular, if the metadata of the monitoring tasks is stored in the data warehouse, relevant information can be obtained by querying the log or the task record table. Further, the monitoring results can be cleaned and extracted to ensure the integrity and consistency of the data.
[0042] In the embodiments of the present disclosure, data annotation can be first performed. Exemplarily, historical monitoring records and corresponding monitoring results can be mapped to form training data for supervised learning. Subsequently, feature extraction can be performed on the annotated data. Exemplarily, features can be extracted from the monitoring records, such as the correlation score of fields, the importance score, the monitoring time range, and the monitoring frequency. Then, label generation can be performed. Exemplarily, labels can be extracted from the monitoring results, such as whether an anomaly is found, the success rate of the monitoring task, and the trend change situation. Finally, based on the features and labels, training samples can be established. Exemplarily, the features and labels can be integrated to form a sample set for model training.
[0043] In the embodiments of the present disclosure, a suitable initial large model can be first selected. Exemplarily, the initial large model can use a deep learning-based model, a decision tree model, or a recommendation system model, etc. The above is only an exemplary illustration and does not limit all possible cases of the initial large model. Here, it is not exhausted. Subsequently, model training can be performed. Exemplarily, based on the training samples, the initial large model can be trained using the supervised learning method so that it can learn the field correlation, importance, and priority score rules in the historical data. Then, model verification can be performed. Exemplarily, some samples that did not participate in the training can be used to verify the effect of the model and evaluate its prediction accuracy. Finally, model optimization can be performed to generate a pre-trained metric monitoring recommendation model. Exemplarily, according to the verification results, the model parameters or structure can be adjusted to generate the final pre-trained model.
[0044] In this way, training samples are constructed using historical monitoring records and results, enabling the model to learn the rules from historical data and generate more accurate recommendation results. By automatically recommending monitoring metrics through the pre-trained model without manual intervention, the efficiency of the monitoring task is greatly improved. By continuously updating the historical monitoring records and results, a dynamic training mechanism can be established to allow the recommendation model to gradually evolve and become more intelligent.
[0045] In some embodiments, training samples are constructed based on historical indicator monitoring records and historical indicator monitoring results, including: determining historical monitoring process parameters according to historical indicator monitoring records; the historical monitoring process parameters include: historical correlation scores, historical importance scores, and historical priority scores; determining historical monitoring result parameters according to historical indicator monitoring results; the historical monitoring result parameters include: historical monitoring algorithm types, historical monitoring training parameters, and historical expected monitoring effects; based on prior knowledge, using the historical monitoring process parameters and historical monitoring result parameters, training samples are constructed.
[0046] In the embodiments of the present disclosure, process information of a monitoring task can be extracted from historical records. Exemplarily, historical records can be first extracted from a data warehouse or monitoring logs, and the field score information therein is parsed. Further, historical correlation scores, historical importance scores, and historical priority scores can be calculated through the process information. Exemplarily, the historical correlation score can be calculated by analyzing the historical relationships between data tables; based on the usage frequency of data tables in historical tasks and the contribution degree to monitoring results, the historical importance score is calculated; by synthesizing the correlation score and importance score, the priority order during task execution is extracted. At the same time, it can also be implemented by directly extracting the historical correlation score, historical importance score, and historical priority score from the process information. Finally, the historical correlation score, historical importance score, and historical priority score can be recorded as historical monitoring process parameters.
[0047] In the embodiments of the present disclosure, result information of a monitoring task can be extracted from historical results. Exemplarily, historical results can be first extracted from a data warehouse or monitoring logs, and the field score information therein is parsed. Further, various parameters can be extracted through the result information. Exemplarily, the algorithm type of the monitoring task, task training parameters, and the expected effect of the monitoring task can be extracted and used as the historical monitoring algorithm type, historical monitoring training parameters, and historical expected monitoring effect respectively. Finally, the historical monitoring algorithm type, historical monitoring training parameters, and historical expected monitoring effect can be recorded as historical monitoring result parameters.
[0048] Here, prior knowledge refers to the empirical, regular, or theoretical knowledge about the problem domain that already exists before data analysis or model training, which helps to guide model construction and optimization. In the embodiments of the present disclosure, prior knowledge may include the business background of the indicator monitoring task, data rules, field correlations, selection criteria for monitoring methods, etc.
[0049] In the embodiments of the present disclosure, sample annotation can be first performed using prior knowledge. Exemplarily, prior knowledge can be used to generate classification labels or regression target values for training samples based on historical monitoring process parameters and historical monitoring result parameters. Finally, the historical monitoring process parameters can be used as input features, and the historical monitoring result parameters can be used as the target values for model training to form complete training samples.
[0050] In this way, using the process parameters and result parameters of historical monitoring data can provide rich and high-quality training samples for the model, enabling the model to have stronger prediction capabilities. At the same time, incorporating prior knowledge can help filter out irrelevant information, focus on key features, and improve the adaptability of the model to actual business scenarios. Model training is based on real historical data and empirical rules, which can more accurately recommend monitoring indicators and reduce incorrect or low-value recommendations. The automated processing of historical records and results reduces the time for manual annotation and sample construction, improving the scale and efficiency of training samples. The model can continuously absorb new historical monitoring records and results over time, continuously update training samples, and optimize model performance.
[0051] In some embodiments, determining the association score and importance score of a data table according to the information in the data warehouse and the content of the data table includes: determining the business attribute of any data table according to the information in the data warehouse and the content of the data table; and determining the association score and importance score of the data table according to the content and business attribute of the data table.
[0052] Here, the business attribute refers to the meaning, function, or role of a data table or field in an actual business scenario, as well as the relevance of the data table or field to business logic. In particular, the business attribute can reflect the role of the data table or field in business activities, processes, or goals. In the embodiments of the present disclosure, the business attribute can help determine the value and relevance of a data table or field, thereby affecting the association score and importance score.
[0053] In the embodiments of the present disclosure, the metadata of the data warehouse can be first extracted. Exemplarily, structural information such as the name of the data table, field names, field data types, and inter-table relationships can be extracted from the data warehouse, and then the subject domain information to which the data table belongs can be obtained. Further, the usage records of the data table can be analyzed. Exemplarily, the usage frequency of the data table in historical queries and monitoring tasks can be analyzed, or the usage situation of the field in the business scenario can be viewed, and then the usage frequency or usage situation can be used as the usage record. Finally, the business attribute of any data table can be determined by combining business rules and semantic information. Exemplarily, domain knowledge can be used to judge the business attribute of the table, such as the relevance between the customer table and the order table in the sales process; or the business meaning of the field can be extracted through keywords in the field name and description, and then the business attribute can be determined.
[0054] In the embodiments of the present disclosure, the correlation degree between a data table and other tables can be calculated, and then a correlation degree score can be obtained. Specifically, the calculation can be based on the relationship between tables. Exemplarily, the direct correlation between tables or fields can be calculated according to the relationship between the primary key and the foreign key, or the connection paths and frequencies between tables can be analyzed using the table relationship diagram of the data warehouse. The shorter the correlation path, the higher the correlation degree. Further, the calculation process of the correlation degree can also be based on the correlation calculation of field values. Exemplarily, statistical methods can be used to calculate the correlation between field values. In particular, the correlation degree can also be determined according to the business domain to which the data table belongs and in combination with the logical correlation of the business scenario. Exemplarily, in sales analysis, the logical correlation between "order amount" and "payment status" may be stronger than that between "order amount" and "customer address".
[0055] In the embodiments of the present disclosure, the importance score of a data table can be determined by calculating the importance of the data table or field in the business scenario. Exemplarily, the number of times a table or field is queried and monitored in the business scenario can be counted, and the higher the usage frequency, the higher the importance score; or domain knowledge can be used to classify the table or field. For example, in financial analysis, the importance of the income statement may be higher than that of the employee table; weights can also be assigned to field attributes. For example, the importance scores of primary key fields and key indicator fields are relatively high; the scores of auxiliary descriptive fields are relatively low.
[0056] In this way, through the extraction of business attributes, the roles of data tables and fields in the business scenario can be understood more accurately, so that the calculation of the correlation degree score and the importance score is more in line with the actual requirements. By calculating the correlation degree score and the importance score, it is possible to help prioritize the recommendation of key data tables, thereby reducing the interference of irrelevant data in the monitoring task and optimizing the use of computing resources. By deeply integrating data and business, not only the effect of the monitoring recommendation task is optimized, but also the efficiency and quality of data-driven decision-making are improved, which has a significant positive effect on intelligent monitoring recommendation.
[0057] In some embodiments, according to the information of the data warehouse and the content of the data table, the business attributes of any data table are determined, including: determining the business theme according to the information of the data warehouse; determining the business classification of any data table according to the business theme; and based on the business classification, performing field analysis on at least one data table to determine the business attributes of any data table.
[0058] Here, the business theme refers to the high-level classification organized around a specific business domain or business goal in the data warehouse, which can make a macro-division of business content, be used to aggregate relevant data tables, and reflect the overall structure and logic of the data in the data warehouse.
[0059] In the embodiments of the present disclosure, the structural information and metadata in the data warehouse can be used to identify the business theme to which a data table belongs. Exemplarily, the theme-related information can be extracted from the description information of the data warehouse table and the table name naming rule; the business theme to which the table belongs can also be inferred based on the relevance between data tables; or the domain knowledge can be used to map the data table to a specific theme.
[0060] Here, the business classification is a further refined division of the data tables or fields within a business theme, and the classification is performed according to the business logic or the specific function of the data table. In particular, the business classification is a sub-classification under the business theme, which helps to organize and understand the content of the data table at a finer granularity.
[0061] In the embodiments of the present disclosure, the function of the data table can be further refined under the framework of the business theme, and the business classification can be determined according to its function. Exemplarily, the business classification can be determined according to the field content of the table, the meaning of the table name, and the use of the fields; the historical usage records of the data table can also be analyzed to obtain the business classification of the table.
[0062] In the embodiments of the present disclosure, the business attributes of the data table can be further clarified through field analysis. Exemplarily, the name, data type, semantics of each field in the data table, and the role of each field in the table can be checked first, then the role of each field can be judged in combination with the business logic, and finally the function of the data table can be judged according to the roles of all fields, and thus the business attributes of the data table can be obtained.
[0063] In this way, through the division of the business theme and the business classification, the structure of the data warehouse can be clearly organized and described, which is convenient for accurately locating the required business data tables. Through the business theme, the high-level business domain can be quickly located, and then further refined to the specific function scenario through the business classification, thereby reducing the data search range and improving the analysis efficiency. Through the division of the business theme and the classification, the business logic relevance between data tables can be better reflected, providing logical support for subsequent data monitoring tasks.
[0064] In some embodiments, according to the content and business attributes of the data table, the correlation score and importance score of the data table are determined, including: performing data lineage analysis on the data table based on the content of the data table to determine the dependency relationship of any data table; determining the importance score of any data table according to the business attributes and the dependency relationship; and determining the correlation score of any data table according to the dependency relationship.
[0065] Here, data lineage refers to the process of data flowing and transforming from the source to the target table in the system, including the paths of data generation, transmission, processing, and storage, which can reflect the source, destination, and intermediate processing steps of the data. In particular, data lineage can be used to trace the source of data and its dependencies, helping to understand the hierarchical structure and correlations between data.
[0066] Furthermore, data lineage analysis is to analyze the dependencies, correlation degrees, and the impact of data on business by tracing the source and transfer process of data, which can reveal the logical relationships, processing flows, and data dependency structures of data tables or fields.
[0067] Here, the dependency relationship refers to the logical connection or correlation between data tables or fields. In the embodiments of the present disclosure, the dependency relationship can be determined through data lineage analysis, which can reflect the direct or indirect dependencies of data tables in data calculation, query, or business logic.
[0068] In the embodiments of the present disclosure, the source and transfer relationships between data tables can be analyzed to discover the dependency relationships of data tables. Exemplarily, the dependencies between data tables can be identified from the structural information of the data warehouse first, then the calculation formulas or derived logics within the data tables can be checked, and then the query associations between data tables can be discovered by analyzing historical query records. Finally, a data lineage visualization tool can be used to draw a dependency relationship diagram between tables to visually display the dependency paths, and then the dependency relationships can be determined.
[0069] In the embodiments of the present disclosure, the importance of a data table in a business scenario can be calculated by combining the business attributes and the dependency relationships of the data table. Exemplarily, the business attribute weights can be determined for the data table according to its business attributes first. Subsequently, the dependency relationship weights of the data table can be judged. When the data table is highly dependent on other data tables or is located upstream in the data lineage graph, a higher dependency relationship weight can be assigned to it. Finally, the business attribute weights and the dependency relationship weights can be comprehensively calculated to obtain the importance score of the data table.
[0070] In the embodiments of the present disclosure, the correlation degree score of a data table can be calculated according to the dependency relationships between data tables. Exemplarily, if there is a direct dependency relationship between this data table and any other data table, the correlation degree weight is higher; if this data table is indirectly associated with any other data table through an intermediate table, the correlation degree score is calculated according to the length of the dependency path. The shorter the path, the higher the correlation degree weight. Finally, the correlation degree weights existing between this data table and each other data table can be integrated to obtain the correlation degree score of the data table.
[0071] Thus, through data lineage analysis, the source and transfer relationships between data tables can be revealed, which helps to clarify the source and dependency relationships of data. Through importance scoring, the key data tables in the business scenario can be identified, which helps to clarify the core data. Through correlation scoring, strong correlation relationships between tables can be discovered, providing a basis for data query optimization and table structure adjustment. The importance scoring and correlation scoring of data tables provide high-quality feature inputs for monitoring the recommendation model, improving the intelligence and accuracy of recommendation results.
[0072] In some embodiments, according to the information of the data warehouse, the correlation score, and the importance score, the priority score of the fields in any data table is determined, including: determining the sub-budget of the data table according to the information of the data warehouse, the correlation score, and the importance score; and determining the priority score of any field according to the sub-budget and the content of the data table.
[0073] Here, the sub-budget refers to the division of the resource allocation amount or priority of the data table according to the correlation score and importance score of the data table in the overall business monitoring task. In the embodiments of the present disclosure, the sub-budget can represent the resources, attention, or priority that a certain data table can be allocated in business monitoring under the condition of limited resources.
[0074] In the embodiments of the present disclosure, the resource allocation amount of the data table in the monitoring task, that is, the sub-budget, can be determined according to the correlation score and importance score of the data table. Exemplarily, the comprehensive score of each data table can be obtained from the previously calculated correlation score and importance score, and then based on the total budget, the sub-budget can be allocated proportionally according to the comprehensive score of the data table. In particular, the sub-budget can be dynamically adjusted according to the business requirements or monitoring task objectives in the actual scenario.
[0075] In the embodiments of the present disclosure, the priority score of the field can be calculated based on the sub-budget and field content of the data table, which is used to determine the attention or resource allocation of the field in the monitoring task. Exemplarily, the attribute analysis of the fields in the data table can be performed first to determine the business value of the fields, then the weight of the fields can be calculated according to the attributes of the fields, and finally, according to the weight of the fields and the sub-budget of the data table where the field is located, the sub-budget is allocated to the fields, and thus the priority score of the fields is obtained. In particular, the field priority can be dynamically adjusted according to the objectives of the monitoring task in the actual scenario.
[0076] In this way, the sub-budget allocates resources based on the comprehensive scores of the data tables to ensure that key data tables receive sufficient resources. Through field priority scoring, the resource allocation can be further refined to ensure that core fields receive sufficient attention and reduce excessive attention to low-value fields. The sub-budget and field priority scoring help monitor tasks focus on high-value data tables and fields, reduce interference from irrelevant data, and thus improve the efficiency of monitoring tasks. Through weight calculation and priority scoring, field analysis can highlight key fields in the business scenario and provide data for the monitoring tasks to focus on.
[0077] In some embodiments, based on the information of the data warehouse, the correlation score, and the importance score, the sub-budget of the data table is determined, including: determining the total budget according to the information of the data warehouse; determining the table weight of any data table according to the importance score and the correlation score; and determining the sub-budget of any data table according to the total budget and the table weight.
[0078] Here, the total budget refers to the total amount of resources allocated to all data tables as a whole during the monitoring task or resource allocation process. In the embodiments of the present disclosure, the total budget can be the total amount of computing power, storage space, monitoring weight, attention, funds, etc., used for unified resource allocation for all data tables.
[0079] In the embodiments of the present disclosure, the total amount of resources for the monitoring task can be determined according to the information of the data warehouse, and then the total budget can be determined. Exemplarily, the total amount of resources to be allocated can be first evaluated according to the objectives and scope of the monitoring task, and then the upper limit of the total budget can be set according to the available resources. In particular, the total budget can be dynamically adjusted according to real-time business requirements or the priority of the monitoring task.
[0080] Here, the table weight refers to the relative importance of a certain data table in the monitoring task, which is comprehensively calculated according to the correlation score and the importance score of the data table. In the embodiments of the present disclosure, the table weight can reflect the contribution degree of a single data table in the business scenario and can be used to guide the resource allocation of the sub-budget.
[0081] In the embodiments of the present disclosure, the correlation score and the importance score of the data table can be combined to calculate the weight that should be allocated to a single data table in the total budget. Exemplarily, the table weight can be calculated using a weighted formula according to the correlation score and the importance score of the data table. Further, all table weights can be normalized to ensure that the sum of the weights is 100% or within the total budget range.
[0082] In the embodiments of the present disclosure, by combining the total budget and the table weight, the resources of each data table are allocated, and then the sub-budget is determined. Exemplarily, the sub-budget can be calculated using a formula according to the total budget and the table weight of each data table. In particular, the sub-budget can be dynamically adjusted according to real-time monitoring requirements.
[0083] In this way, the total budget can provide global constraints, ensure overall coordination of resource allocation, and reduce resource waste. At the same time, table weights are based on importance scores and relevance scores, which can ensure that the allocation of sub-budgets meets business needs and prioritizes resources for key indicators. Sub-budgets can be used to highlight key data tables and reduce resource allocation to low-value data tables, thereby improving the efficiency of monitoring tasks. By calculating field priority scores, resources can be further allocated to key fields in the data table to highlight core data.
[0084] In some embodiments, the priority score of any field is determined according to the content of the sub-budget and the data table, including: based on a preset field accuracy expectation, according to the information of the sub-budget and the data table, determining the field weight of any field in any data table; according to the field weight, prioritizing the fields to determine the priority score.
[0085] Here, the expected value of field accuracy refers to the ideal accuracy level or quality requirement that a field needs to achieve in a data monitoring or analysis task. In the disclosed embodiment, the expected value of field accuracy can reflect the importance of the field data in the business scenario and the degree of dependence on the business logic, which is usually determined by business needs.
[0086] Here, field weight refers to the importance of a field relative to other fields in the data table to which it belongs. In the disclosed embodiment, field weight can reflect the contribution or criticality of the field to the business goal, and can be comprehensively determined by the attributes of the field, the budget of the data table, and the expected value of field accuracy.
[0087] In the disclosed embodiment, the accuracy expectation of the field, the sub-budget of the data table, and the business attributes of the field can be comprehensively considered to determine the weight of the field in the data table. For example, the attributes of the field can be first analyzed, the type, purpose and business value of the field can be checked, and the field weight ratio of the field can be determined; then, based on the field weight ratio, the initial field weight can be assigned to each field according to the sub-budget of the data table; further, the weight coefficient can be set according to the needs of the actual scenario, and then the field weight can be calculated according to the weight coefficient, the initial field weight and the pre-set accuracy expectation.
[0088] In the disclosed embodiment, the fields can be sorted according to the field weights and assigned priority scores to highlight the key fields. For example, the fields can be sorted from high to low according to the field weights, and then the priority scores can be assigned according to the sorting results. Among them, the process of assigning priority scores can be adjusted according to the number of fields, and each field can be assigned a priority score, or multiple fields can be assigned the same priority score. In particular, the priority score can be adjusted according to the real-time business needs in the actual scenario.
[0089] In this way, through the calculation of field weights, key fields can be highlighted, ensuring that resources are preferentially allocated to fields with higher business value and reducing resource waste. By determining the priority score, the computing resources of the monitoring task can be concentrated on the core fields, reducing the monitoring interference of low-value fields, thereby improving the task efficiency. By determining the field weights and priority scores, the core fields closely related to business goals can be highlighted, improving the accuracy and effectiveness of the monitoring task, helping the data analysis task to highlight key fields, and providing high-quality analysis data for business decision-making.
[0090] In some embodiments, the process of calculating the priority score can also be performed using an algorithm model. Exemplarily, an algorithm model for calculating the priority of indicator monitoring can be established, and this algorithm model can be expressed by the following formula:
[0091] rank = cost × accurate
[0092] In the formula, rank represents the priority score, cost represents the calculation cost, and accurate represents the accuracy.
[0093] Specifically, the larger the value of rank, the more sufficient the cost budget of the table, and the greater the expected deviation of the indicator, and the higher the priority of the corresponding monitoring plan.
[0094] Specifically, the larger the cost, the higher the calculation cost, and the more the indicators in the table should receive the investment of the monitoring plan. In particular, in order to involve the table cost budget in the same formula calculation, normalization calculation is required to remove the unit information, and the normalization parameter is the global average uninvested budget. Exemplarily, cost can be calculated by the following formula:
[0095]
[0096] In the formula, A represents the total budget, A′ represents the invested part of the total budget, a represents the sub-budget, and a′ represents the invested part of the sub-budget.
[0097] Specifically, the larger the accurate, the greater the expected deviation of the accuracy, and the more the field should receive the investment of the monitoring plan compared to other indicators. Exemplarily, accurate can be calculated by the following formula:
[0098]
[0099] In the formula, C represents the expected value of field accuracy, and C′ represents the evaluated value of field accuracy.
[0100] Here, the field accuracy evaluation value refers to the actual measurement result or evaluation result of the current data accuracy of a certain field, corresponding to the expected field accuracy expectation value. In the embodiments of the present disclosure, the field accuracy evaluation value can reflect the reliability of the field data and the degree of compliance with business requirements, and is a specific measurement index of the field data quality.
[0101] In this way, by establishing an algorithm model, millions of metrics in the data warehouse can be directly calculated and sorted, and then the field priority scores can be calculated by combining the budget and accuracy, so as to screen out the monitoring metrics that most need to be trained by the algorithm model.
[0102] In some embodiments, after sorting the fields according to the field weights and determining the priority scores, it further includes: obtaining the life cycle stage of the field; generating a time weight according to the life cycle stage; and adjusting the priority score according to the time weight to generate an adjusted priority score.
[0103] Here, the life cycle stage refers to the current state of the field in the entire life cycle such as data creation, use, storage, and archiving. The life cycle stage of the field reflects the activity, usage frequency, and importance of the field in the current business task. Different life cycle stages will affect the priority and time weight of the field. In the embodiments of the present disclosure, the life cycle stage may include a creation stage, a use stage, a maintenance stage, an archiving stage, a discard stage, etc.
[0104] In the embodiments of the present disclosure, the life cycle stage of the field can be determined according to the current state of the data to which the field belongs. Exemplarily, the timestamp and business status of the data to which the field belongs can be checked first, and then the field stage can be judged according to the life cycle rules defined by the business.
[0105] Here, the time weight is a dynamic weight generated according to the life cycle stage of the field, used to reflect the relative importance of the field in the current time period.
[0106] In the embodiments of the present disclosure, the time weight can be generated according to the activity and importance of the life cycle stage to which the field belongs. Exemplarily, the time weight rules can be preset, and the default time weight can be defined for each life cycle stage.
[0107] In the embodiments of the present disclosure, the time weight is combined with the initial priority score to generate the final adjusted priority score. Exemplarily, the weighted calculation can be performed according to the time weight and the priority score, and then the adjusted priority score can be obtained.
[0108] In this way, by combining the life cycle stage and time weight, the importance of the field can be dynamically adjusted, which can more accurately reflect the actual value of the field in the current task. The adjusted priority score helps the monitoring task allocate resources more precisely, focus on the fields in the active stage, and reduce the waste of resources on fields with low importance. Through the dynamic adjustment of time weight, the monitoring task can quickly adapt to real-time business needs, improve computing efficiency and optimize the monitoring effect.
[0109] In some embodiments, when a new business area is added and several corresponding data tables are added, due to the small amount of initial data, it will lead to insufficient understanding of its internal indicators by business operation personnel and it is also difficult for technical personnel to identify the patterns therein. At this time, the time weight can be lowered, and less cost budget and accuracy expectation are allocated in the initial stage, so that the priority score of the new indicator becomes lower, thereby reducing the cost incurred. As the life cycle of the indicator progresses, when the trainable data in the data warehouse reaches a certain scale, the time weight can be increased, and then the corresponding budget and accuracy expectation can be increased, so that the priority score of the new indicator becomes higher, thereby increasing the cost incurred to obtain the relevant data of the corresponding indicator. Further, based on the relevant data of the corresponding indicator, business decisions can be made, and the required business effects can be achieved by increasing or decreasing the time weight.
[0110] In some embodiments, the association score, importance score, and priority score are input into a pre-trained indicator monitoring recommendation model to generate an indicator monitoring recommendation result, including: constructing a structured input field according to the association score, importance score, and priority score; inputting the structured input field into the indicator monitoring recommendation model to generate a monitoring algorithm type, monitoring training parameters, and monitoring expected effect; and using the monitoring algorithm type, monitoring training parameters, and monitoring expected effect as the indicator monitoring recommendation result.
[0111] Here, the structured input field refers to the input data that has been sorted and organized, which is used to standardize the input of the model so that the model can effectively understand and process the data. In the embodiments of the present disclosure, the structured input field is the key information extracted from the association score, importance score, and priority score, and is constructed according to a specific format or structure so that the model can efficiently generate an indicator monitoring recommendation result.
[0112] In the embodiments of the present disclosure, the relevance score, importance score, and priority score can be converted into structured inputs that can be understood by the model. Exemplarily, the relevance score, importance score, and priority score of each field can be obtained first, and then the score values are normalized to ensure that all score values are normalized to a specific range, so as to ensure that the input dimensions of the model are consistent. Finally, the score values can be organized into structured field information, including field identifiers, score values, weight information, and time-related information. In particular, the structured field information can be in the data interchange format (JavaScript Object Notation, JSON), table, or matrix format.
[0113] In the embodiments of the present disclosure, the organized structured data can be used as input and passed to a pre-trained metric monitoring recommendation model, and the model generates recommendation results. Exemplarily, the structured input fields can be passed to the model to start the inference process. The model generates recommendation results based on the relevance, importance, and priority scores of the input fields. Further, according to the inference results of the model, recommendation results suitable for the current monitoring task can be output. Exemplarily, according to the scoring characteristics of the input fields, the most suitable monitoring algorithm can be selected from a pre-set algorithm set for recommendation, and then monitoring parameters can be recommended for the monitoring algorithm according to the weights and score values of the input fields. Finally, according to the characteristics of the recommendation algorithm and the scores of the input fields, the monitoring effect can be estimated.
[0114] In this way, through the combination of structured input fields and pre-trained models, monitoring recommendation results can be automatically generated, reducing manual intervention and making the recommendation results more accurate, and being able to adapt to complex business scenarios. The structured input fields make the scoring information more standardized and improve the model processing efficiency. Automatically generating monitoring algorithms and parameters reduces the manual configuration time and improves the task response speed. The model can recommend different types of monitoring algorithms according to the scoring characteristics of the input fields, covering a variety of business scenarios.
[0115] The solution of the present disclosure addresses the problems of fuzzy features and unpredictable scenarios in a large number of data warehouse metrics. By mining the data rules and applicable model algorithms of important data warehouse metrics within a controllable cost range, the effect of "monitoring all that need to be monitored, all that can be monitored, and even those that cannot be monitored by humans" is achieved. This solution can not only reduce the labor cost of monitoring metrics, but also improve the practicality of the monitoring system, support flexible hierarchical management of metrics, thereby optimizing the utilization rate of cost resources, and significantly reducing the failure rate caused by unmonitored metrics, providing an efficient and reliable solution for data warehouse data management and quality monitoring.
[0116] The embodiments of the present disclosure provide a data warehouse metric monitoring device, as Figure 2As shown, the device can monitor metrics for a data warehouse, where the data warehouse includes at least one data table, and the data table includes at least one field. The device may include: a table scoring module 201 for determining the correlation score and importance score of a data table according to the information of the data warehouse and the content of the data table; a field scoring module 202 for determining the priority score of a field in any data table according to the information of the data warehouse, the correlation score, and the importance score; a monitoring recommendation module 203 for inputting the correlation score, importance score, and priority score into a pre-trained metric monitoring recommendation model to generate a metric monitoring recommendation result; and a metric monitoring module 204 for monitoring metrics for the data warehouse based on the metric monitoring recommendation result.
[0117] In some embodiments, the data warehouse metric monitoring device further includes: a historical data module ( Figure 2 not shown in the figure) for obtaining historical metric monitoring records and historical metric monitoring results; a sample construction module ( Figure 2 not shown in the figure) for constructing training samples according to the historical metric monitoring records and historical metric monitoring results; and a model training module ( Figure 2 not shown in the figure) for training an initial large model using the training samples to generate a pre-trained metric monitoring recommendation model.
[0118] In some embodiments, the sample construction module includes: a process parameter determination sub-module for determining historical monitoring process parameters according to the historical metric monitoring records; the historical monitoring process parameters include: historical correlation score, historical importance score, and historical priority score; a result parameter determination sub-module for determining historical monitoring result parameters according to the historical metric monitoring results; the historical monitoring result parameters include: historical monitoring algorithm type, historical monitoring training parameters, and historical monitoring expected effect; and a training sample sub-module for constructing training samples based on prior knowledge using the historical monitoring process parameters and historical monitoring result parameters.
[0119] In some embodiments, the table scoring module 201 includes: a business determination sub-module for determining the business attribute of any data table according to the information of the data warehouse and the content of the data table; and a scoring calculation sub-module for determining the correlation score and importance score of the data table according to the content of the data table and the business attribute.
[0120] In some embodiments, the business determination sub-module is used to: determine the business theme according to the information of the data warehouse; determine the business classification of any data table according to the business theme; and perform field analysis on at least one data table based on the business classification to determine the business attribute of any data table.
[0121] In some embodiments, the scoring calculation sub-module is configured to: perform data lineage analysis on a data table based on the content of the data table to determine the dependency relationship of any data table; determine the importance score of any data table according to the business attribute and the dependency relationship; and determine the association degree score of any data table according to the dependency relationship.
[0122] In some embodiments, the field scoring module 202 includes: a sub-budget calculation sub-module, configured to determine the sub-budget of a data table according to the information of the data warehouse, the association degree score, and the importance score; and a priority scoring sub-module, configured to determine the priority score of any field according to the sub-budget and the content of the data table.
[0123] In some embodiments, the sub-budget calculation sub-module is configured to determine the total budget according to the information of the data warehouse; determine the table weight of any data table according to the importance score and the association degree score; and determine the sub-budget of any data table according to the total budget and the table weight.
[0124] In some embodiments, the priority scoring sub-module is configured to: determine the field weight of any field in any data table according to the sub-budget and the information of the data table based on a preset field accuracy expectation value; and perform priority sorting on the fields according to the field weight to determine the priority score.
[0125] In some embodiments, the priority scoring sub-module is further configured to: obtain the life cycle stage of a field; generate a time weight according to the life cycle stage; and adjust the priority score according to the time weight to generate an adjusted priority score.
[0126] In some embodiments, the monitoring and recommendation module 203 includes: an input construction sub-module, configured to construct a structured input field according to the association degree score, the importance score, and the priority score; a model inference sub-module, configured to input the structured input field into an index monitoring and recommendation model to generate a monitoring algorithm type, monitoring training parameters, and monitoring expected effects; and a result generation sub-module, configured to use the monitoring algorithm type, the monitoring training parameters, and the monitoring expected effects as index monitoring and recommendation results.
[0127] For the specific functions and example descriptions of the modules and sub-modules of the device according to the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the foregoing method embodiments, which will not be elaborated herein.
[0128] The data warehouse metric monitoring device in the embodiments of the present disclosure can utilize data metrics such as correlation degree and importance score, enabling more accurate recommended monitoring metrics, improving decision-making quality, and enhancing the practicality of the monitoring solution. Through intelligent recommendation, redundant monitoring and resource waste are reduced, and at the same time, the labor cost of monitoring metrics is lowered. With scoring and model recommendation, flexible hierarchical management of metrics can be achieved, and the monitoring metrics can be dynamically adjusted in actual application scenarios to adapt to business changes, optimizing the efficiency and quality of the metric monitoring task.
[0129] The embodiments of the present disclosure provide a schematic diagram of the scenario of a data warehouse metric monitoring method, as Figure 3 shown.
[0130] As mentioned above, the data warehouse metric monitoring method provided by the embodiments of the present disclosure is applied to an electronic device. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers.
[0131] Specifically, the electronic device can specifically perform the following operations:
[0132] Determine the correlation score and importance score of a data table according to the information of the data warehouse and the content of the data table; determine the priority score of the fields in any data table according to the information of the data warehouse, the correlation score, and the importance score; input the correlation score, importance score, and priority score into a pre-trained metric monitoring recommendation model to generate a metric monitoring recommendation result; perform metric monitoring on the data warehouse based on the metric monitoring recommendation result.
[0133] It should be understood that Figure 3 the shown scenario diagram is merely illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 3 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.
[0134] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0135] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0136] Figure 4FIG. 0 is a schematic block diagram of an exemplary electronic device 400 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital assistants, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0137] As Figure 4 shown, the device 400 includes a computing unit 401 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0138] A plurality of components in the device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0139] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above, such as the data warehouse metric monitoring method. For example, in some embodiments, the data warehouse metric monitoring method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the data warehouse metric monitoring method described above can be executed. Alternatively, in other embodiments, the computing unit 401 can be configured to execute the data warehouse metric monitoring method in any other suitable manner (e.g., by means of firmware).
[0140] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application-specific standard products (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0141] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0142] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0143] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0144] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: a Local Area Network (LAN), a Wide Area Network (WAN), and the Internet.
[0145] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact via a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.
[0146] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0147] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A data warehouse indicator monitoring method, wherein the data warehouse includes at least one data table, and the data table includes at least one field, and the method comprises: Determining a relevance score and an importance score of the data table according to the information of the data warehouse and the content of the data table; Determine the priority score of the field in any of the data tables according to the information of the data warehouse, the relevance score and the importance score; Inputting the relevance score, the importance score and the priority score into a pre-trained indicator monitoring recommendation model to generate an indicator monitoring recommendation result; The data warehouse is subjected to indicator monitoring based on the indicator monitoring recommendation result.
2. The method according to claim 1, wherein: The method further comprises: Obtain historical indicator monitoring records and historical indicator monitoring results; Constructing training samples according to the historical indicator monitoring records and the historical indicator monitoring results; The training samples are used to train the initial large model to generate the pre-trained indicator monitoring recommendation model.
3. The method according to claim 2, wherein: The constructing of training samples according to the historical indicator monitoring records and the historical indicator monitoring results includes: Determine historical monitoring process parameters according to the historical indicator monitoring records; the historical monitoring process parameters include: historical relevance score, historical importance score and historical priority score; Determine historical monitoring result parameters according to the historical indicator monitoring results; the historical monitoring result parameters include: historical monitoring algorithm type, historical monitoring training parameters and historical monitoring expected effect; Based on prior knowledge, the training samples are constructed using the historical monitoring process parameters and the historical monitoring result parameters.
4. The method according to claim 1, wherein: Determining the relevance score and importance score of the data table according to the information of the data warehouse and the content of the data table includes: Determine the business attribute of any of the data tables according to the information of the data warehouse and the content of the data tables; The relevance score and importance score of the data table are determined according to the content of the data table and the business attributes.
5. The method according to claim 4, wherein: Determining the business attribute of any of the data tables according to the information of the data warehouse and the content of the data table includes: Determine a business topic based on the information in the data warehouse; Determine the business classification of any of the data tables according to the business subject; Based on the business classification, field analysis is performed on the at least one data table to determine the business attributes of any of the data tables.
6. The method according to claim 4, wherein: Determining the relevance score and importance score of the data table according to the content of the data table and the business attribute includes: Based on the content of the data table, performing data lineage analysis on the data table to determine the dependency relationship of any of the data tables; Determining the importance score of any of the data tables according to the business attributes and the dependency relationship; The relevance score of any of the data tables is determined according to the dependency relationship.
7. The method according to claim 1, wherein: Determining the priority score of the field in any of the data tables according to the information of the data warehouse, the relevance score and the importance score includes: Determining a sub-budget for the data table according to the information of the data warehouse, the relevance score and the importance score; A priority score for any field is determined based on the sub-budgets and the content of the data table.
8. The method according to claim 7, wherein: Determining the sub-budget of the data table according to the information of the data warehouse, the relevance score and the importance score includes: Determining a total budget based on the information in the data warehouse; Determining a table weight of any of the data tables according to the importance score and the relevance score; The sub-budget of any of the data tables is determined according to the total budget and the table weight.
9. The method according to claim 7, wherein: Determining the priority score of any field according to the sub-budget and the content of the data table includes: Based on a preset field accuracy expectation value, according to the sub-budget and information of the data table, determining a field weight of any field in any of the data tables; The fields are prioritized according to the field weights to determine the priority scores.
10. The method according to claim 9, wherein: After determining the priority score, the method further includes: Get the life cycle phase of the field; generating a time weight according to the life cycle stage; The priority score is adjusted according to the time weight to generate an adjusted priority score.
11. The method according to claim 1, wherein: The step of inputting the relevance score, the importance score and the priority score into a pre-trained indicator monitoring recommendation model to generate an indicator monitoring recommendation result includes: constructing a structured input field according to the relevance score, the importance score, and the priority score; Inputting the structured input field into the indicator monitoring recommendation model to generate a monitoring algorithm type, monitoring training parameters and monitoring expected effects; The monitoring algorithm type, the monitoring training parameters and the expected monitoring effect are used as the indicator monitoring recommendation result.
12. A data warehouse indicator monitoring device, wherein the data warehouse includes at least one data table, the data table includes at least one field, and the device comprises: A table scoring module, used to determine the relevance score and importance score of the data table according to the information of the data warehouse and the content of the data table; A field scoring module, used to determine the priority score of the field in any of the data tables according to the information of the data warehouse, the relevance score and the importance score; A monitoring recommendation module, used for inputting the relevance score, the importance score and the priority score into a pre-trained indicator monitoring recommendation model to generate an indicator monitoring recommendation result; An indicator monitoring module is used to perform indicator monitoring on the data warehouse based on the indicator monitoring recommendation result.
13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to at least one processor; wherein, The memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor to enable the at least one processor to perform the method of any one of claims 1-11.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are for causing a computer to perform a method according to any one of claims 1-11.
15. A computer program product comprising a computer program stored on a storage medium, the computer program implementing the method according to any one of claims 1 to 11 when executed by a processor.