Method and device for determining data asset generation duration and computing device cluster
By constructing an associated link between data assets and jobs, estimating the data asset generation time based on the historical execution time, and dynamically adjusting resource allocation, the problem of difficult to determine the data asset generation time in the big data field is solved, and accurate SLA guarantee is achieved.
Patent Information
- Application Number
- CN202410067181.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2025-07-18
AI Technical Summary
In the field of big data, it is difficult for the existing technology to accurately estimate the generation time of data assets, which makes it difficult to determine the actual generation time of data assets, and it is impossible to ensure that the service level agreement (SLA) of data assets is reached on time.
By determining the association between data assets and jobs, a job link is built, and the data asset generation time is automatically estimated based on the historical execution time in the job link, and when the job link cannot complete on time, the computing resource allocation is dynamically adjusted to ensure the implementation of SLA.
Accurate estimates of the generation time of data assets are achieved, the need for manual intervention is reduced, and the flexibility of data assets is improved and the guarantee efficiency of SLA is improved.
Smart Images

Figure CN120336310A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method, an apparatus, and a cluster of computing devices for determining the generation duration of data assets. Background Art
[0002] In the related big data field, data assets can be obtained after processing data, and data assets are usually applied to data analysis and exploration or data applications. Enterprise data analysts, operators, and management decision-makers daily rely on data assets produced by processing data on time every day in order to carry out daily work based on the data assets. Therefore, it is crucial to know the SLA (data availability guarantee) time for data asset production (generally taking the data output time expected by users as the SLA), and to ensure that the SLA of data assets is achieved on time (that is, the generation of data assets is completed before the time expected by users). In the related art, data is generally processed by multiple jobs with dependencies. During the processing, problems such as job execution failure and execution delay are likely to occur, resulting in difficulty in determining the true generation time of the data assets provided for users to use, and thus unable to provide users with accurate data asset generation times. Therefore, how to accurately estimate the SLA time of data assets is an urgent problem to be solved. Summary of the Invention
[0003] Embodiments of the present invention provide a method, an apparatus, and a cluster of computing devices for determining the generation duration of data assets, which can realize the association between data assets and jobs to obtain a job link, and based on the job link, can more accurately estimate the generation duration of data assets.
[0004] In a first aspect, an embodiment of the present invention provides a method for determining the generation duration of data assets, including: determining a job link corresponding to the data asset to be generated based on the table associated with the data asset to be generated, the table associated with the job, and other jobs on which the job depends; the table associated with the data asset to be generated is the table in the data warehouse required to generate the data asset to be generated; the table associated with the job is the table stored in the data warehouse after the job is processed; the job link indicates the sequence of jobs required to generate the table associated with the data asset to be generated; determining the duration required to generate the data asset based on the historical execution durations of the jobs in the job link of the data asset to be generated.
[0005] In this solution, the association between data assets and jobs can be realized to obtain a job link, and based on the job link, the generation duration of data assets can be more accurately estimated.
[0006] In a possible implementation, based on the tables associated with the data assets to be generated, the tables associated with the jobs, and the other jobs that the jobs depend on, determine the job link corresponding to the data assets to be generated, including: based on the tables associated with the data assets to be generated and the tables associated with the jobs, determine the jobs associated with the data assets to be generated, and the tables associated with the jobs associated with the data assets to be generated are used to generate the data assets to be generated; based on the other jobs that the jobs depend on and the jobs associated with the data assets to be generated, determine the job link corresponding to the data assets to be generated.
[0007] In an example of this implementation, based on the tables associated with the data assets to be generated and the tables associated with the jobs, determine the jobs associated with the data assets to be generated, including: based on the first position information of the table associated with the data assets to be generated, determine the first identifier of the table associated with the data assets to be generated; the first position information indicates the position of the table associated with the data assets to be generated in the data warehouse; based on the second position information of the table associated with the job, determine the second identifier of the table associated with the job; the second position information indicates the position of the table associated with the job in the data warehouse; determine the jobs associated with the table where the first identifier and the second identifier are the same and the data assets to be generated are associated.
[0008] In this solution, the association between jobs and data assets is realized through the identifiers of the tables.
[0009] Exemplarily, the first position information includes the name of the table associated with the data assets to be generated and the address of the data warehouse where it is located;
[0010] Exemplarily, the second position information includes the name of the table associated with the job and the address of the data warehouse where it is located;
[0011] Exemplarily, when the data warehouse where the table associated with the job is located and the data warehouse indicated by the table associated with the data assets to be generated are the same, the first identifier and the second identifier are the same.
[0012] In a possible implementation, based on the historical execution duration of the jobs in the job link of the data assets to be generated, determine the duration required to generate the data assets, including: based on the historical execution duration of each job in the job link, determine the critical job path in the job link; the critical job path is the path formed by the critical jobs that determine the generation time of the data assets to be generated; based on the historical execution duration of the jobs in the critical job path, determine the duration required to generate the data assets.
[0013] In this solution, analyze the critical job path in the job link through the historical execution duration, and the duration required to generate the data assets can be estimated more accurately through the critical job path.
[0014] In an example of this implementation manner, the method further includes: for each job in the job link, removing outliers from the execution durations at different collection time points of the job; and determining the historical execution duration of the job based on the execution durations after outlier removal.
[0015] In this solution, by removing outliers, the historical execution duration of the job can be estimated more accurately.
[0016] Exemplarily, the historical execution duration is the target execution duration among the execution durations after outlier removal, and the proportion of the execution durations after outlier removal that are less than or equal to the target execution duration is a preset proportion, such as 90%.
[0017] In this solution, considering that the execution duration of the job will change continuously, by selecting the relatively larger execution duration after outlier removal as the historical execution duration of the job, the historical execution duration of the job can be estimated more accurately.
[0018] In a possible implementation manner, the job link is formed by sequentially connecting jobs in multiple layers, where the data asset to be generated is obtained after the job in the last layer of the multiple layers is executed; for each layer other than the last layer in the multiple layers, each job in the layer is connected to at least one job in the next layer; the critical job link is the path formed by the jobs with the longest historical execution duration in each layer of the multiple layers.
[0019] In a possible implementation manner, the method further includes: updating the computing resource allocation of the job link when the job link cannot be completed on time.
[0020] In this solution, by updating the computing resource allocation of the job, the job link can be completed on time as much as possible.
[0021] In a possible implementation manner, the method further includes: for at least some of the jobs in the job link, determining the deterioration situation of the job, and when the deterioration situation of the job indicates that the job has deteriorated, giving an alarm and / or proposing a governance solution for the job, where the governance solution is used to reduce the deterioration situation of the job.
[0022] In this solution, by reminding or governing the deteriorated job, the job is processed.
[0023] In a second aspect, an embodiment of the present invention provides an apparatus for determining the generation duration of a data asset. The apparatus for determining the generation duration of a data asset includes a number of modules, and each module is used to execute each step in the method for determining the generation duration of a data asset provided in the first aspect of the embodiment of the present invention. The division of the modules is not limited herein. For the specific functions executed by each module of the apparatus for determining the generation duration of a data asset and the beneficial effects achieved, please refer to the functions of each step in the method for determining the generation duration of a data asset provided in the first aspect of the embodiment of the present invention, which will not be elaborated herein.
[0024] Exemplarily, the apparatus for determining the generation duration of a data asset includes:
[0025] An association module, configured to determine a job link corresponding to the data asset to be generated based on the table associated with the data asset to be generated, the table associated with the job, and other jobs on which the job depends; the table associated with the data asset to be generated is the table in the data warehouse required to generate the data asset to be generated; the table associated with the job is the table stored in the data warehouse after the job is processed; the job link indicates the sequence of jobs required to generate the table associated with the data asset to be generated.
[0026] A duration determination module, configured to determine the duration required to generate the data asset based on the historical execution durations of the jobs in the job link of the data asset to be generated.
[0027] In a possible implementation, the association module includes: a job association unit and a link association unit; where
[0028] The job association unit is configured to determine the job associated with the data asset to be generated based on the table associated with the data asset to be generated and the table associated with the job, and the table associated with the job associated with the data asset to be generated is used to generate the data asset to be generated.
[0029] The link association unit is configured to determine the job link corresponding to the data asset to be generated based on other jobs on which the job depends and the job associated with the data asset to be generated.
[0030] In an example of this implementation, the job association unit is configured to determine the first identifier of the table associated with the data asset to be generated based on the first position information of the table associated with the data asset to be generated; the first position information indicates the position of the table associated with the data asset to be generated in the data warehouse; determine the second identifier of the table associated with the job based on the second position information of the table associated with the job; the second position information indicates the position of the table associated with the job in the data warehouse; determine the job associated with the table whose first identifier and second identifier are the same and the data asset to be generated.
[0031] Exemplarily, the first location information includes the name of the table associated with the data asset to be generated and the address of the data warehouse where it is located;
[0032] Exemplarily, the second location information includes the name of the table associated with the job and the address of the data warehouse where it is located;
[0033] Exemplarily, when the data warehouse indicated by the data warehouse where the table associated with the job is located is the same as the data warehouse where the table associated with the data asset to be generated is located, the first identifier and the second identifier are the same.
[0034] In a possible implementation manner, the duration determination module is configured to determine the critical job path in the job chain based on the historical execution duration of each job in the job chain; the critical job path is the path formed by the critical jobs that determine the generation time of the data asset to be generated; based on the historical execution duration of the jobs in the critical job path, determine the duration required to generate the data asset.
[0035] In an example of this implementation manner, the duration determination module is further configured to, for each job in the job chain, remove outliers from the execution durations at different collection time points of the job; based on the execution durations after outlier removal, determine the historical execution duration of the job.
[0036] Exemplarily, the historical execution duration is the target execution duration among the execution durations after outlier removal, and the proportion of the execution durations after outlier removal that are less than or equal to the target execution duration is a preset proportion.
[0037] In a possible implementation manner, the job chain is formed by sequentially connecting jobs in multiple layers, where the data asset to be generated is obtained after the jobs in the last layer of the multiple layers are executed; for each layer other than the last layer in the multiple layers, each job in the layer is connected to at least one job in the next layer; the critical job link is the path formed by the jobs with the longest historical execution duration in each layer of the multiple layers.
[0038] In a possible implementation manner, the apparatus further includes: an allocation module, configured to update the calculation resource allocation of the job chain when the job chain cannot be completed on time.
[0039] In a possible implementation manner, the apparatus further includes: a governance module, configured to determine the deterioration situation of at least some of the jobs in the job chain, and when the deterioration situation of the job indicates job deterioration, give an alarm and / or propose a governance solution for the job, and the governance solution is used to reduce the deterioration situation of the job.
[0040] In a third aspect, an embodiment of the present invention provides an apparatus for determining the generation duration of data assets, including: at least one memory for storing programs; at least one processor for executing the programs stored in the memory, and when the programs stored in the memory are executed, the processor is used to execute the method provided in the first aspect.
[0041] In a fourth aspect, an embodiment of the present invention provides an apparatus for determining the generation duration of data assets, characterized in that the apparatus runs computer program instructions to execute the method provided in the first aspect. Exemplarily, the apparatus may be a chip or a processor.
[0042] In one example, the apparatus may include a processor, which may be coupled to a memory, read instructions in the memory, and execute the method provided in the first aspect according to the instructions. Among them, the memory may be integrated in the chip or the processor, or may be independent of the chip or the processor.
[0043] In a fifth aspect, an embodiment of the present invention provides a computing device cluster, including: at least one computing device, each computing device including a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the method provided in the first aspect.
[0044] In a sixth aspect, an embodiment of the present invention provides a computer storage medium, in which instructions are stored, and when the instructions run on a computer, the computer is made to execute the method provided in the first aspect.
[0045] In a seventh aspect, an embodiment of the present invention provides a computer program product containing instructions, and when the instructions run on a computer, the computer is made to execute the method provided in the first aspect. Description of the Drawings
[0046] Figure 1 is a system architecture diagram of a data processing system provided by an embodiment of the present invention;
[0047] Figure 2 is a schematic diagram of the architecture of a scheduling system provided by an embodiment of the present invention;
[0048] Figure 3 is a flowchart of the method for determining the generation duration of data assets provided by an embodiment of the present invention;
[0049] Figure 4 is a schematic diagram of the scenario of the association between data assets and jobs provided by an embodiment of the present invention;
[0050] Figure 5 is a schematic diagram of the scenario of the job link associated with the generated data assets provided by an embodiment of the present invention;
[0051] Figure 6 It is a schematic diagram of an operation link provided by an embodiment of the present invention;
[0052] Figure 7a It is a scenario schematic diagram of a method for determining the generation duration of data assets provided by an embodiment of the present invention Figure 1 ;
[0053] Figure 7b It is a scenario schematic diagram of a method for determining the generation duration of data assets provided by an embodiment of the present invention Figure 2 ;
[0054] Figure 8 It is a schematic structural diagram of a device for determining the generation duration of data assets provided by an embodiment of the present invention;
[0055] Figure 9 It is a schematic structural diagram of a computing device provided by an embodiment of the present invention;
[0056] Figure 10 It is a schematic structural diagram of a computing device cluster provided by an embodiment of the present invention;
[0057] Figure 11 It is a schematic diagram of the connection between two computing devices provided by an embodiment of the present invention. Detailed implementation manners
[0058] In order to make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be described below with reference to the accompanying drawings.
[0059] In the description of the embodiments of the present invention, words such as "exemplary", "for example", or "for illustration" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for illustration" in the embodiments of the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for illustration" is intended to present relevant concepts in a specific manner.
[0060] In the description of the embodiments of the present invention, the term "and / or" is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, B exists alone, and A and B exist simultaneously. In addition, unless otherwise specified, the meaning of the term "plural" refers to two or more. For example, multiple systems refer to two or more systems, and multiple terminals refer to two or more terminals.
[0061] In addition, the terms "first" and "second" are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the technical features indicated. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0062] The following explains some terms in this embodiment. It should be noted that these explanations are for the convenience of those skilled in the art and do not limit the scope of protection required by the present invention.
[0063] Data asset: refers to data assets that are owned or controlled by an entity (such as an enterprise, an organization), can bring future benefits, and are recorded in a physical or electronic manner. Such data assets can be, for example, document materials or electronic data. The asset entity can include databases, data tables, directories, jobs, nodes, logical entities, business attributes, fields, column lineage, insights. Column lineage refers to data lineage at the column level. Data lineage is also called data provenance or data pedigree. Data lineage is usually defined as a life cycle, mainly including the source of data and the location where the data moves over time.
[0064] Extract-Load-Transform (ETL): is used to describe the process of extracting, transposing, and loading data from the source end to the destination end (data warehouse). Transform usually describes the pre-data processing process in the data warehouse.
[0065] SLA (Service Level Agreement): Service Level Agreement, which is a guarantee of website service availability for Internet companies. Data SLA, that is, data availability guarantee, generally takes the data output time as the SLA.
[0066] Enterprise Resources Planning (ERP): is based on advanced enterprise management concepts and uses information technology to achieve integrated management of the entire enterprise resources. ERP is an enterprise management information system that can provide real-time information integration across regions, departments, and even companies. On the premise of optimizing the allocation of enterprise resources, it integrates the main or all business activities within the enterprise, including main functional modules such as financial accounting, management accounting, production planning and management, material management, sales and distribution, etc., to achieve the goal of efficient operation.
[0067] Kanban: An easy-to-use tool for visualizing and managing workflows. It is characterized by having columns representing the various stages of the workflow. Kanban cards are used to track the progress of individual tasks and activities through the stages. The two main types of kanban are physical kanban and digital kanban. Physical kanban is most suitable for in-office, co-located teams. Digital kanban is more suitable for remote and hybrid teams, especially those working on complex projects.
[0068] Data Lake: A data storage concept, a system or repository for data stored in its natural format, usually object blobs or files. Specifically, it can bring together different types of data for data analysis without a predefined mathematical model. A data lake stores raw data that has not been processed for a specific purpose.
[0069] Data Warehouse: A data warehouse is a subject-oriented, integrated, relatively stable, historical-changing data collection used to support management decision-making. A data warehouse is for supporting decision-making and analytical data processing; it effectively integrates multiple heterogeneous data sources, reorganizes them by subject after integration, contains historical data, and the data stored in the data warehouse is generally not modified. A data warehouse can only store processed and refined data, while a data lake stores raw data that has not been processed for a specific purpose. Therefore, a data lake requires a much larger storage capacity than a data warehouse, and the data is flexible and the analysis is rapid, making it very suitable for machine learning. A data warehouse does not serve business information systems; it serves analytical applications. More often, it is accessed through various business intelligence (BI) front-end visualization analysis tools or reporting tools, and ultimately is for report query and data analysis services.
[0070] Database (DB): A database is a repository for organizing, storing, and managing data according to a data structure. There are many types of databases, ranging from the simplest table storing various data to large database systems capable of storing massive amounts of data, which have been widely applied in various aspects. Databases usually serve business, while data warehouses usually serve analysis. Generally, the databases mentioned are usually for business application software, regardless of whether these software are in B / S architecture or C / S architecture. For example, the commonly used ERP systems and OA systems in enterprises, or the ordering food APPs and online ticket-purchasing APPs on our mobile phones, etc. The characteristics of these business systems are that users operate on these software systems, such as logging in, filling in personal information, modifying personal profiles, querying a record, etc. Data interacts through these software programs and the underlying databases, and operations such as adding, deleting, modifying, and querying are performed on the underlying data tables. Therefore, usually these databases serve various business systems and application software running on the operating system, more oriented towards business processes and business management. The data sources of databases come from the data generated by various business system software programs or the data generated by users interacting with these business system software. The data sources of data warehouses are directly one or more databases or files of these business systems, such as SQL Server, Oracle, MySQL, Excel, text files, etc. It can also be simply understood that the databases of many business systems send data to the data warehouse, which is a collection of various databases, a larger database. The establishment of a data warehouse is to connect the data of these basic databases.
[0071] Schema: In a database, a schema is a way to organize and manage data. It defines various objects in the database, such as tables, views, indexes, etc., and the relationships between them. A schema can be regarded as the blueprint of a database, which stipulates the structure and organization method of data, enabling data to be effectively stored and retrieved. A schema usually consists of multiple tables, and each table contains multiple columns. The table defines the structure of data, and each column defines the data type and constraints. By defining tables and columns, a schema can ensure data consistency and integrity. In a database, there can be multiple schemas, and each schema can contain multiple tables and other objects. Different database management systems support different types of schemas. Common ones are: Single schema: All tables and objects are in the same schema. This is the simplest type of schema, suitable for small databases or simple applications. Multiple schemas: Multiple independent schemas can be created in a database, and each schema contains a set of related tables and objects. This way can better organize and manage data, improving the maintainability and scalability of the database.
[0072] Engine: It is the core component for developing programs or systems on an electronic platform. With the engine, developers can quickly establish and lay out the functions required by the program or use it to assist the operation of the program. Generally speaking, an engine is the supporting part of a program or a system. Common program engines include game engines, search engines, graph engines, etc.
[0073] Next, an introduction will be given to the data processing system to which the method for determining the generation duration of data assets provided by the embodiments of the present invention may be applied. Figure 1 The architecture example diagram of a data processing system provided by the embodiments of the present invention is shown. The method for determining the generation duration of data assets provided by the embodiments of the present invention can be applied to the Figure 1 system architecture diagram as shown. As Figure 1 shown, the data processing system includes a data job platform 110 and a terminal 120.
[0074] Among them, the data job platform 110 includes a data lake 111, a scheduling system 112, a data warehouse system 113, a data asset system 114, a data application 115, and an asset management device 116. The scheduling system 112 is used to process the tables in the data lake 111, store the processed tables, and store the processed tables in the data warehouse 113; the data asset system 114 generates data assets through one or more tables in the data warehouse 113.
[0075] The architecture of the scheduling system 112 is the ODS (Operational Data Store) layer, the DWD (Data Warehouse Detail) layer, the DWS (Data Warehouse Summary) layer, and the ADS (Application Data Store) layer. Among them, the OSD layer, the DWD layer, the DWS layer, and the ADS layer are only examples and do not constitute specific limitations. In some other possible implementation manners, the architecture of the scheduling system 112 may be the OSD layer, the CDM (Common Dimensional Model) layer, and the ADS layer; among them, the CDM layer is the most core and crucial layer in the data warehouse, mainly used to provide a standardized and shared dimensional model to facilitate data analysis. The CDM layer usually includes two parts, namely the DWD layer and the DWS layer. The embodiments of the present invention are described by taking the OSD layer, the DWD layer, the DWS layer, and the ADS layer as examples.
[0076] Among them, the ODS layer is the layer closest to business operations in the data warehouse. This layer is mainly used to store raw data and complete data accumulation, usually reflecting the latest operations in the enterprise's business system. At the same time, it is also the basis for building the data warehouse. Its main task is to integrate data from different data sources into a unified data model. In this way, business users can obtain real-time and accurate data information through the ODS layer, thus better supporting business decisions. In practical applications, the ODS layer is mainly responsible for obtaining data from data sources (such as MySQL, OBS, etc.) of various business systems (such as transaction systems, billing systems, etc.) and performing data processing. Among them, data processing can include data extraction, data cleaning, data transformation, and data loading. Data extraction is used to extract data from various data sources and convert it into a unified data format; data cleaning is used to clean data and remove bad data such as duplicates, missing values, and errors; data transformation is used to transform data to meet the construction standards of the data warehouse; data loading: load the processed data into the data warehouse. The ODS layer is also called the operational data source layer and is a core component of the data warehouse. In practical applications, the ODS layer usually uses reliable data warehouse ETL tools to provide data for the data warehouse, so as to keep the source data and the data warehouse synchronized. At the same time, the data in the ODS layer of the data warehouse is stored on disk, directly reflecting a characteristic of the data warehouse - non-volatility, that is, in the case of downtime or crashes, the data will not be lost.
[0077] The DWD layer is the detailed layer in the data warehouse. Its main task is to receive the raw data from the ODS layer and perform operations such as cleaning, standardization, dimensionality degradation, and abnormal data elimination for unified processing to ensure data quality and consistency. In this way, business users can obtain accurate and clean data information through the DWD layer, thus ensuring data quality and consistency and better supporting business decisions. Among them, the main functions of the DWD layer include: data cleaning, data transformation, and data loading; data cleaning is used to clean data and remove bad data such as duplicates, missing values, and errors; data transformation is used to transform data to meet the construction standards of the data warehouse; data loading: load the processed data into the data warehouse. In some possible cases, the DWD layer generally models according to business themes, including multiple dimensions and fact tables. The dimension tables can be used to describe the characteristics of business data, while the fact tables contain key data indicators (such as sales volume, price, etc.).
[0078] The DWS layer is the service layer in the data warehouse, mainly responsible for providing data services and data analysis for business users. The main task of the DWS layer is to provide data services and data analysis for business users, so as to support business decision-making. Business users can obtain the data information and analysis results they need through the DWS layer, so as to better understand the business situation and make wise decisions. Among them, the main functions of the DWS layer include: data processing, data analysis, and data integration; data processing is used to process data to meet the needs of business users; data analysis is used to analyze data to provide data support and suggestions for business users; data integration is used to integrate data from different data sources into a unified data model. In some possible cases, the main role of the DWS layer is to aggregate and summarize the data in the DWD layer by theme to form a wide table, thereby improving data analysis performance. The DWS layer usually contains multiple wide tables, and each wide table is generated by aggregating and grouping multiple fact tables and dimension tables. The wide tables in the DWS layer can meet the analysis needs of specific themes and different dimensions, reduce the operations on other tables, and improve data analysis performance.
[0079] ODS, DWD, and DWS are the three basic levels in the hierarchical construction of the data warehouse. The ODS layer is mainly responsible for data extraction, cleaning, transformation, and loading; the DWD layer is mainly responsible for data cleaning, transformation, and loading to ensure data quality and consistency; the DWS layer is mainly responsible for providing data services and data analysis for business users to support business decision-making. The construction and operation of these three levels are the key to the success of data warehouse construction, and can provide accurate, clean, and valuable data information and analysis results for business users.
[0080] The main function of the ADS layer is to save the result data, provide query interfaces for external systems, provide value-added applications for enterprises based on the data in the data warehouse, and apply the data in the data warehouse to fields such as enterprise decision-making, reports, analysis, and control. The ADS layer usually adopts OLAP (Online Analytical Processing) technology for fast access and query of data. The ADS layer generally includes multiple wide tables to support operations such as querying, analyzing, reporting, controlling, and making decisions related to enterprise applications. These wide tables can generally be queried and accessed through BI tools or custom applications to meet the various data needs of enterprises. To improve access and query speeds, the ADS layer usually uses technologies such as data indexing, caching, and pre-aggregation.
[0081] The scheduling system 112 can complete various jobs based on the tasks in the ODS layer, DWD layer, DWS layer, and ADS layer. Exemplarily, the ODS layer includes several access tasks; the DWD layer includes several computing tasks; the DWS includes several service tasks; the ADS layer includes several application tasks; for each of the ODS layer, DWD layer, DWS layer, and ADS layer, the tasks in this layer can form multiple jobs; for any two or more adjacent layers among the ODS layer, DWD layer, DWS layer, and ADS layer, the combination of tasks in these layers can form multiple jobs; in addition, the access tasks, computing tasks, service tasks, and application tasks can be combined for processing to obtain the tables required for generating data assets. For example, as Figure 2 shown, the ODS layer includes 4 access tasks, the DWD layer includes 3 computing tasks, the DWS layer includes 5 service tasks, the ADS layer includes 3 application tasks, and the tasks in the ODS layer, DWD layer, DWS layer, and ADS layer can obtain 3 tables required for generating data assets: Table 1, Table 2, and Table 3; for Table 1: After access task 1 and access task 2 are completed simultaneously, computing task 1 is executed, and after computing task 1 is completed, service task 1 is executed; after access task 3 and access task 4 are completed simultaneously, computing task 3 is executed, and after computing task 3 is completed, service task 3 is executed. After service task 1 and service task 3 are completed simultaneously, application task 1 is executed, thereby obtaining Table 1; for Table 2: After access task 1 and access task 2 are completed simultaneously, computing task 1 is executed, and after access task 1 and access task 3 are completed simultaneously, computing task 2 is executed; after computing task 1 and computing task 2 are completed simultaneously, service task 2 is executed, and after computing task 2 is completed, service task 4 is executed. After service task 2 and service task 4 are completed simultaneously, application task 2 is executed, thereby obtaining Table 2; for Table 3: After access task 1 and access task 3 are completed simultaneously, computing task 2 is executed, and after computing task 2 is completed, service task 4 is executed; after access task 3 and access task 4 are completed simultaneously, computing task 3 is executed, and after computing task 3 is completed, service task 5 is executed. After service task 4 and service task 5 are completed simultaneously, application task 3 is executed, thereby obtaining Table 3. It should be noted that multiple tasks in the scheduling system 112 can form multiple jobs, and a table can be obtained after each job is executed. These tables can include the tables required for generating data assets. There is a dependency relationship between different jobs. For example, job 2 depends on the table obtained after job 1 is executed, and job 3 depends on the table obtained after job 2 is executed; thus, it can be seen that data asset generation depends on tables, table generation depends on jobs, and there is a dependency relationship between jobs and other jobs. It should be noted that the scheduling system 112 can send jobs to the data lake 111, and the data lake 111 allocates computing resources to execute the jobs.
[0082] As Figure 1As shown, by way of example, the data warehouse system 113 may include multiple data warehouses, such as a data warehouse in a public cloud, a data warehouse in a private cloud, a columnar database management system, and Doris (an open-source distributed columnar storage database focused on real-time analysis and interactive queries). The data asset system 114 may generate multiple data assets through the tables in the data warehouse system 113. For example, data interface assets, dataset assets, data object assets, and data behavior assets. The data application 115 is used to visualize the data assets in the data asset system 114 and may include dashboard applications, analysis applications, and AI applications.
[0083] Among them, the asset management device 116 is connected to the data lake 111, the scheduling system 112, the data warehouse system 113, and the data asset system 114 to implement the method provided in the embodiments of the present invention. For detailed content, see the description of Figure 3 below, which will not be elaborated here. It should be noted that in some possible scenarios, the asset management device 116 may be integrated into the scheduling system 112 or the data asset system 114.
[0084] Among them, the terminal 120 may be, but is not limited to, various personal computers, laptop computers, smartphones, tablets, and portable wearable devices. Exemplary embodiments of the terminal devices involved in this solution include, but are not limited to, electronic devices running iOS, android, Windows, Harmony OS, or other operating systems. The embodiments of the present invention do not specifically limit the type of electronic device. In the embodiments of the present invention, the terminal 120 may run a data application to facilitate users to view data.
[0085] Among them, the terminal 120 communicates with the data operation platform 110 via a network. The network can be a wired network or a wireless network. By way of example, the wired network can be a cable network, an optical fiber network, a Digital Data Network (DDN), etc., and the wireless network can be a telecommunications network, an internal network, the Internet, a Local Area Network (LAN), a Wide Area Network (WAN), a Wireless Local Area Network (WLAN), a Metropolitan Area Network (MAN), a Public Service Telephone Network (PSTN), a Bluetooth network, a ZigBee network, a Global System for Mobile Communications (GSM), a Code Division Multiple Access (CDMA) network, a General Packet Radio Service (GPRS) network, etc. or any combination thereof. It can be understood that the network can use any known network communication protocol to implement communication between different client layers and gateways. The above network communication protocols can be various wired or wireless communication protocols, such as Ethernet, universal serial bus (USB), fire wire, global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), new radio (NR), Bluetooth, wireless fidelity (Wi-Fi), etc.
[0086] In the relevant big data field, the data results processed by the scheduling system 112 in the data operation platform 110 are usually applied to data analysis and exploration or data applications. Enterprise data analysts, operation personnel, and management decision-makers daily rely on the data assets processed by the data operation platform 110 on time every day, so as to carry out daily work based on the data assets. Therefore, it is crucial to know the SLA time for the production of data assets and ensure the timely achievement of the data asset SLA. Often, during the scheduling and processing of big data jobs, problems such as job failures and execution delays are likely to occur, making it difficult to determine and guarantee the actual generation time of the data assets provided for users to use. As a result, it is impossible to provide users with the accurate data asset generation time. Therefore, how to accurately estimate the SLA time of data assets and ensure the timely generation of data assets is an important technical ability and business pain point.
[0087] Currently, the solutions for the timely achievement of the SLA (used to describe the moment when the data asset is generated) of data assets generally include the following two types:
[0088] The first type: According to the business requirements, identify the jobs that need to be key-guaranteed. To avoid the operation of the tasks required for job execution (tasks in the scheduling system 112) being affected by resources and the operation of upstream tasks, add the tasks to the intelligent baseline, and calculate the estimated completion time of each task added to the baseline on a daily or hourly basis. The priority of the baseline ensures that the baseline tasks are preferentially allocated resources.
[0089] The second type: Complete the signing of the current job SLA (used to describe the moment when the job needs to be completed) by signing the SLA (used to describe the moment when the task needs to be completed) for all upstream tasks, and calculate the tasks that need to be signed through the Database Availability Group (DAG) to reduce the signing cost. Further reduce the cost through the automatic signing of the recommended SLA algorithm. After all SLAs for the processing links that need to be guaranteed are signed in a semi-automatic + manual manner throughout the entire link, then ensure the achievement of the SLA based on the signed SLA. If the SLA cannot be achieved, generate SLA governance tasks to ensure the achievement of the SLA. In specific implementation, the applicant needs to fill out a declaration form, and then the applicant conducts link analysis and analysis of key tasks (generally tasks that require relatively large resource consumption or long execution duration). Through the key tasks and the SAL recommendation algorithm, judge whether the key tasks need to be automatically signed with the SLA. If so, automatically sign the SLA and notify and broadcast. Otherwise, provide the key tasks for recommended SLA signing to the user for self-signing.
[0090] For the above two solutions, manual participation is required to identify key jobs.
[0091] To solve the above problems, an embodiment of the present invention proposes a method for determining the generation duration of data assets, which is implemented by the asset management device 116.
[0092] By extending the data lineage that could originally only be applied between jobs and between tables to the scenario of data assets, this method realizes the association between data assets and jobs, and can quickly deduce the job link affecting the generation of data assets; based on the historical execution duration of each job in the job link, the critical job path is determined from the job link, and the critical job path can determine the generation time of the data asset; in summary, by automatically establishing the association between the to-be-generated data asset and the job, the job link of the to-be-generated data asset is obtained, and subsequently, based on the job link of the data asset, the generation duration of the data asset can be estimated more accurately; in addition, based on the job link, the jobs that need to be key-guaranteed can also be identified. There is no need to manually calibrate the job guarantee baseline or manually sign the guarantee link, which reduces the input cost of the SLA guarantee critical link and improves the flexibility of identifying critical jobs. This is only a brief description of the method. For the detailed content of this method, please refer to the following description of Figure 3 .
[0093] Next, in combination with the above-provided data processing system, a method for determining the generation duration of data assets provided by an embodiment of the present invention will be introduced in detail. Figure 3 It is a schematic flowchart of the method for determining the generation duration of data assets provided by an embodiment of the present invention. This embodiment can be applied to the asset management device 116.
[0094] As Figure 3 shown, the method for determining the generation duration of data assets provided by an embodiment of the present invention at least includes the following steps:
[0095] Step 301: The asset management device 116 determines the tables associated with the to-be-generated data asset based on the information of the to-be-generated data asset; the tables associated with the to-be-generated data asset are the tables in the data warehouse required to generate the data asset.
[0096] In a possible implementation manner, the asset management device 116 can interface with the data asset system 114. For example, it can interface through the API interface method or through the configuration library of the open data job platform 110. When interfacing, network intercommunication needs to be ensured. After configuration, the asset management device 116 can periodically or real-time collect the information of the to-be-generated data asset in the data asset system 114, and after parsing, obtain the tables associated with the to-be-generated data asset.
[0097] Exemplarily, the information of the data asset to be generated may include the name of the data asset to be generated, the definition of the data asset to be generated, and the data warehouse connection information of the data asset to be generated, such as a URL. Among them, the definition of the data asset is used to describe the identifier of the table required to generate the data asset to be generated, such as a name (used to distinguish it from other tables in the data warehouse). It should be noted that the table required to generate the data asset to be generated may be one or more. The URL is used to locate the data warehouse required for the data asset to be generated. Specifically, the URL may indicate the IP address of the data warehouse, the interface (port), and the name of the database (DB). The IP address can describe the network where the data warehouse is located, and the interface (port) can describe the server port of the data warehouse. It should be noted that the data warehouse may consist of multiple databases. Therefore, the URL needs to indicate the name of the database (DB) to locate the database in the data warehouse.
[0098] In a specific implementation, the asset management device 116 may determine the table associated with the data asset to be generated based on the data warehouse connection information of the data asset to be generated and the name of the table described in the definition of the data asset to be generated. This table is the table indicated by the name of the table described in the definition of the data asset to be generated in the data warehouse indicated by the data warehouse connection information of the data asset to be generated.
[0099] Step 302: The asset management device 116 determines the table associated with the job based on the information of the job. The table associated with the job is the table stored in the data warehouse after the job is processed.
[0100] In a possible implementation, the asset management device 116 may interface with the scheduling system 112. For example, it may interface through an API interface or through the configuration library of the open data job platform 110. The interface requires ensuring network intercommunication. After configuration, the asset management device 116 may periodically or real-time collect the information of the jobs of the scheduling system 112 and parse it to obtain the table associated with the job.
[0101] Exemplarily, the information of the job may include job information, computing task information, and transmission task information.
[0102] Among them, the job information may include the job name, other jobs that the job depends on before execution (which may be one or more), information about several tasks inside the job (such as tasks in any one or more of the ODS layer, DWD layer, DWS layer, and ADS layer) (such as the name of the task), the dependency situation of the tasks inside the job. The dependency situation of the tasks is used to describe other tasks that the task needs to depend on. Specifically, this task can only be executed after the other tasks it depends on are processed and based on the processed results.
[0103] The computing task information may include the name of the computing task, the data warehouse connection information of the computing task such as the URL (used to locate the data warehouse), and the content of the computing task (used to describe the identifiers of the tables required for the computing, such as the name and the processing method of the tables).
[0104] The transfer task information may include the name of the transfer task, the source connection information (used to describe the address of the source), the source library name, the identifier of the source table such as the name, the target connection information (used to connect to the address of the target), the target library name, and the identifier of the target table such as the name. It should be noted that, in order to meet the business requirements, the data in an existing data warehouse can be migrated to another data warehouse. In the embodiments of the present invention, the data warehouse where the data to be migrated is located is referred to as the source library, and the data warehouse to which the data is migrated is referred to as the target library. Exemplarily, the source may store the data after DWS layer and DWD layer calculations, and the target may be a data warehouse for querying, and the query efficiency of the target is higher than that of the source.
[0105] In specific implementation, the asset management device 116 may determine the tables associated with the computing script based on the data warehouse connection information of the computing task and the name of the table in the content of the computing task. This table is the table indicated by the name of the table in the content of the computing task in the data warehouse indicated by the data warehouse connection information of the computing script; based on the target connection information and the name of the target table in the transfer task information, determine the table associated with the transfer task. This table is the table indicated by the name of the target table in the data warehouse indicated by the target connection information; associate the tables associated with the computing task and the transfer task with the job where the computing task and the transfer task are located, and establish the association between the job and the table.
[0106] Step 303: The asset management device 116 determines the job link for generating the data asset to be generated based on the tables associated with the job, the tables associated with the data asset to be generated, and other jobs on which the job depends; the job link indicates the sequence of the jobs required to generate the tables associated with the data asset to be generated.
[0107] Among them, other jobs on which the job depends can be determined based on the information of the job; specifically, as described above, the information of the job describes other jobs on which the job depends beforehand (which can be one or more); in this way, the asset management device 116 can obtain other jobs on which each job depends based on the information of each job.
[0108] According to a feasible implementation manner, the asset management device 116 determines the jobs associated with the data asset to be generated based on the tables associated with the data asset to be generated and the tables associated with the job. The tables associated with the jobs associated with the data asset to be generated are used to generate the data asset to be generated; based on the jobs associated with the data asset to be generated, the dependency relationship between the job and other jobs, determine the job link of the data asset to be generated.
[0109] In specific implementation, for each job, if the table associated with the job is the same as the table associated with the data asset to be generated, it indicates that the table required for generating the data asset is generated by this job; then, based on the dependency relationship between jobs (used to describe other jobs that each job depends on) and the jobs associated with the data asset to be generated, the job link for generating the data asset to be generated is determined.
[0110] It should be noted that a data warehouse can have several backups. At this time, there is a primary-backup relationship between the data warehouse and its backup data warehouse, and these data warehouses can be considered as the same data warehouse. At this time, if the table related to the job and the table associated with the data asset to be generated are located in different data warehouses, but there is a primary-backup relationship between these two data warehouses, or they are backed up from the same data warehouse, it indicates that the table associated with the job and the table associated with the data asset to be generated come from the same data warehouse, and the job is related to the data asset to be generated. For example, Figure 4 , assume that data warehouse A1 and data warehouse A2 have a primary-backup relationship, that is, they are the same data warehouse. The data asset to be generated is associated with table a in data warehouse A1, and the job is associated with table a in data warehouse A2. Since data warehouse A1 and data warehouse A2 are the same data warehouse, the data asset to be generated is associated with table a.
[0111] Among them, the job link is used to describe the job for generating the table associated with the data asset to be generated and other jobs that this job depends on; the other jobs that this job depends on are used to describe the jobs required from the initial job to the end of generating the table associated with the data asset to be generated, and the relationships between these jobs. Exemplarily, the job link can be formed by sequentially connecting jobs in multiple levels. The table generated by the job in the last level is associated with the data asset to be generated, and the jobs in other levels are other jobs that the table associated with the data asset to be generated depends on; for each level among the multiple levels, each job in this level connects at least some of the jobs in the next level. Exemplarily, as Figure 6 shown, assume that the job link consists of 3 levels in total. The first level includes jobs 1 to 4, the second level includes jobs 5 to 7, and the third level includes job 8 (for calculating table A), job 9 (for calculating table B), and job 10 (for calculating table C); among them, jobs 1 to 4 connect to job 5, jobs 6 and 7 connect to job 9, and the tables 1, 2, and 3 calculated by jobs 8, 9, and 10 are used to generate the data asset.
[0112] In some possible implementations, the asset management device 116 may determine the first identifier of the table associated with the data asset to be generated based on the first location information of the table associated with the data asset to be generated; determine the second identifier of the table associated with the job based on the second location information of the table associated with the job; when the first identifier and the second identifier are the same, it indicates that the table associated with the data asset to be generated is the same as the table associated with the job, and the data asset to be generated is related to the job. Therefore, it may be determined that the job associated with the table with the same first identifier and second identifier is associated with the data asset to be generated; wherein, the first location information may indicate the identifier of the table associated with the data asset to be generated, such as the name, the address of the data warehouse where it is located, such as the URL, and further may indicate the name of the schema. The URL may indicate the IP address of the data warehouse, the interface (Port), and the name of the database (DB). Exemplarily, the first location information may be the data warehouse connection information of the data asset to be generated, such as the URL, and the name of the table described in the definition of the data asset to be generated; the second location information of the table associated with the job is similar to the first location information and will not be elaborated here. The difference is that the second location information may be the data warehouse connection information of the computing task, such as the URL, and the name of the table in the content of the computing task, as well as the target end connection information and the name of the target end table.
[0113] It should be noted that one data warehouse may have several backups. The addresses of the data warehouse and its backup data warehouses, or the addresses of multiple backup data warehouses of the data warehouse, such as the URL, are different. Therefore, the URLs indicated by the data warehouse connection information of the job and the data asset to be generated may be different, but the data warehouses indicated by the data warehouse connection information of the job and the data asset to be generated may have a primary-backup relationship, or be backed up from the same data warehouse; therefore, in order to realize the association between the data asset and the job, when the data warehouse indicated by the address of the data warehouse where the table associated with the job is located is the same as the data warehouse indicated by the address of the data warehouse where the table associated with the data asset to be generated is located, such as having a primary-backup relationship, or being backed up from the same data warehouse, the first identifier and the second identifier are the same. For example, Figure 4 assuming that data warehouse A1 and data warehouse A2 have a primary-backup relationship, that is, they are the same data warehouse. The data asset to be generated is associated with table a in data warehouse A1, and table a in data warehouse A1 has the first identifier. The job is associated with table a in data warehouse A2, and table a in data warehouse A1 has the second identifier. Since there is a primary-backup relationship between data warehouse A1 and data warehouse A2, the first identifier and the second identifier are the same, and the data asset to be generated is associated with table a.
[0114] Subsequently, after associating the jobs with the same first identifier and second identifier with the data assets to be generated, based on the dependencies between the jobs associated with the data assets to be generated, the job link for generating the data assets to be generated can be determined.
[0115] Exemplarily, as Figure 5 shown, after the job is parsed, a transmission task and a calculation task are obtained. Based on the transmission task and the calculation task, a table associated with the transmission task and the calculation task is obtained. Through an identifier generation algorithm, a virtual ID of the table associated with the transmission task and the calculation task is determined. The identifier generation algorithm is used to assign the same identifier to tables with a primary-backup relationship or tables with the same name in multiple data warehouses backed up from the same data warehouse. After the data asset to be generated is parsed, a table associated with the data asset to be generated is obtained. Through the identifier generation algorithm, a virtual ID of the table associated with the data asset to be generated is determined. The data assets to be generated and the jobs associated with the tables with the same virtual ID are associated, and the jobs associated with the data assets to be generated are stored in a graph engine (mainly used for relationship analysis, abstracting the relationship network into an image-like graph structure data and then performing queries and analyses). The graph engine can also store the dependencies between jobs. Subsequently, the graph engine can generate the job link for the data assets to be generated based on the dependencies between jobs and the jobs associated with the data assets to be generated.
[0116] Step 304, the asset management device 116 determines the duration required to generate the data asset based on the job link of the data asset to be generated.
[0117] In a possible implementation, the asset management device 116 determines the critical job path in the job link based on the historical execution duration of each job in the job link. The critical job path is the path formed by the critical jobs that determine the generation time of the data asset to be generated, and the critical job path is used to determine the time required to generate the data asset to be generated.
[0118] In an example, the asset management device 116 can interface with the scheduling system 112. For example, it can interface through an API interface or through the execution history database of the open data job platform 110. The interface requires ensuring network intercommunication. After configuration, the asset management device 116 can periodically or real-time collect the execution duration of each job in the scheduling system 112 at different collection time points. After removing the outliers from the execution durations at different collection time points, the historical execution duration is determined. In summary, through outlier removal, the possibility that the execution duration fluctuation of a single job affects the accuracy of the entire job execution duration is reduced, and the prediction accuracy is improved.
[0119] Among them, outlier rejection is used to remove abnormal data, such as data that is too large or too small. Any method in the prior art can be adopted for outlier rejection, and the embodiments of the present invention do not make specific descriptions thereon.
[0120] Among them, the historical execution duration is the target execution duration in the execution durations after outlier rejection, and the proportion of the execution durations less than or equal to the target execution duration in the execution durations after outlier rejection is a preset proportion, such as 95% or 90%.
[0121] In specific implementation, for each job in the job link, the asset management device 116 can perform outlier rejection on the execution durations of the job at different collection time points; based on the execution durations after outlier rejection, determine the historical execution duration of the job. For example, the execution durations after outlier rejection can be sorted in ascending order, and the maximum execution duration among the top 90% can be used as the historical execution duration, which can also be understood as the execution duration at TP90. TP90 represents the execution duration such that 90% of the execution durations are less than it, specifically, the maximum execution duration among the first 90% of the execution durations sorted in ascending order.
[0122] When estimating the generation time of data assets, the critical job path can be determined based on the historical execution durations of each job in the job link. The job link is formed by sequentially connecting jobs in multiple levels, and the jobs in the last level are associated with the data assets to be generated; for each level among the multiple levels, each job in this level is connected to at least some of the jobs in the next level; the critical job path is the path formed by the jobs with the longest historical execution time in each level.
[0123] Exemplarily, as Figure 6 shown, assume that the job link consists of 3 levels in total. The first level includes jobs 1 to 4, the second level includes jobs 5 to 7, and the third level includes job 8 (for calculating Table A), job 9 (for calculating Table B), and job 10 (for calculating Table A); among them, jobs 1 to 4 are connected to job 5, and jobs 6 and 7 are connected to job 9. For the first level, the job with the longest historical execution duration among jobs 1 to 4 is job 3. For the second level, the job with the longest historical execution duration is job 5. For the third level, the job with the highest execution duration is job 9. Then the critical job path is job 3 → job 5 → job 9.
[0124] In specific implementation, the asset management device 116 can store the jobs associated with the data assets to be generated and the dependency relationships between the jobs into the graph engine, and the graph engine can output the critical job path based on the historical execution durations of the jobs.
[0125] In some possible implementation manners, the time required for generating the data asset to be generated is the sum of the historical execution durations of the jobs in the critical job path, or the result after correcting the sum of the historical execution durations of the jobs in the critical job path. The correction method can be to weight the sum of the execution durations of the jobs in the critical job path. Exemplarily, the weighting coefficient can be greater than 1 and less than 1.2.
[0126] In this solution, by automatically establishing the association between the data asset to be generated and the jobs, the job link of the data asset to be generated is obtained. Subsequently, based on the job link of the data asset, the generation duration of the data asset can be estimated more accurately. In addition, based on the job link, the jobs that need to be key-guaranteed can also be identified.
[0127] Furthermore, when the job link cannot be completed on time, the asset management device 116 updates the computing resource allocation of the job link. It should be noted that once the scheduling system 112 obtains sufficient data, for example, starting from early this morning to obtain all the sales data of the previous day, it starts to execute the job. In addition, the scheduling system 112 stores the SLA commitment of the data asset (generally set by the user and the user can perceive it), and the SLA commitment is used to indicate the time when the data asset needs to be generated, for example, 9:30 every morning. Correspondingly, the asset management device 116 can, based on the SLA commitment of the data asset to be generated and the execution situation of the job link of the data asset to be generated, determine whether the completion time of the job link of the data asset to be generated meets the SLA commitment of the data asset to be generated. If it meets, it means that the job link can be completed on time; otherwise, the job link cannot be completed on time.
[0128] When updating the computing resource allocation of the job link, the asset management device 116 can send the pre-configured guarantee policy to the computing resource allocation module of the data lake. The computing resource allocation module is used to allocate the computing resources of the data lake 111 to the tasks in the scheduling system 112 to complete the job. The computing resources are generally the number of CPU cores and the size of the memory. Among them, the pre-configured guarantee policy is used to improve the execution priority of the jobs in the job link. Multiple guarantee policies can be pre-configured. In some possible scenarios, the asset management device 116 can send at least some of the pre-configured guarantee policies to the computing resource allocation module of the data lake. Exemplarily, the guarantee policy can be any one of the following: Let the tasks required by the jobs in the job link execute first according to the remaining computing resources; Suspend other jobs that occupy computing resources (the importance level of this job is relatively low) and / or jobs that cannot be completed; The computing resources are elastically scaled to achieve resource expansion.
[0129] Further, for each job in the job link, the asset management device 116 determines the deterioration condition of the job. When the deterioration condition of the job indicates job deterioration, an alarm is generated, and / or a governance solution for the job is provided.
[0130] Among them, the deterioration condition can be a deterioration rate. When the deterioration rate is greater than or equal to a preset threshold, it can indicate job deterioration. The deterioration rate can be calculated by any one of multiple calculation methods:
[0131] Calculation method 1: The deterioration rate indicates the ratio of the difference between the current execution duration of the job and the standard execution duration to the standard execution duration. Among them, the standard execution duration can be the TP50 among the execution durations at different collection time points within the collection period with the current moment as the end point. TP50 means that the proportion of the execution duration of the job at different collection time points that is less than or equal to the standard execution duration is 50%. Specifically, the collection period can be designed according to actual needs. Exemplarily, the end moment of the collection period is the current moment, the collection duration is 1 month, and the collection interval is 1 hour.
[0132] Calculation method 2: The deterioration rate indicates the ratio of the number of occurrences of the execution duration that differs from the standard execution duration by more than a preset threshold to the total number of occurrences.
[0133] Calculation method 3: The deterioration rate indicates the ratio of the number of consecutive occurrences of the execution duration that differs from the standard execution duration by more than a preset threshold to the total number of occurrences.
[0134] It should be noted that the preset threshold used to determine whether a job deteriorates can be flexibly set in combination with the calculation method of the deterioration rate.
[0135] Among them, the asset management device 116 can perform diagnostic analysis based on the execution status of the job, generate a governance plan for the job, and the governance plan for the job is used to reduce the deterioration of the job, such as returning the deteriorated job to a normal state. The governance plan for the job can be used to modify the computer program related to the job, such as the code of the job, the storage plan of the table (such as in what type of data warehouse it is stored and in what structure it is stored, such as a tree structure), and the computing engine (used to process data), such as the parameters of hive and spark; among them, the asset management device 116 can determine the governance plan for the job based on the execution status of the job. In some possible ways, the asset management device 116 can perform at least some of the following parameter analyses based on the execution status of the job: computing resource consumption, characteristics of the dependent tables, execution plan of the job, etc. The computing resource consumption can be the number of CPU cores occupied, the size of the memory, the network bandwidth, etc.; the characteristics of the table are used to describe the meaning, type, data distribution (used to describe the distribution of data for different values), and table distribution (dividing and storing the data of the table in different places) of the data stored in the table, etc.; the execution plan of the job is used to describe the actual execution status of the job (which can be understood as the logic executed after code parsing).
[0136] It should be noted that the alarm can enable the data engineer to understand that there is a problem with the job, so that the data engineer can focus on ensuring and optimizing the job with problems, and ensure the generation of the data assets to be generated.
[0137] Based on the method for determining the data asset generation duration provided above, a specific application of the method for determining the data asset generation duration is described.
[0138] Figure 7a This is a schematic diagram of a specific application of a method for determining the data asset generation duration provided for the implementation of the present invention. As Figure 7a shown, the specific content includes:
[0139] Step 701, the asset management device 116 analyzes the relationship between the data asset to be generated and the table.
[0140] For the detailed content, refer to the description of step 301.
[0141] Step 702, the asset management device 116 analyzes the relationship between the job and the computing task and the transmission task, and the relationship between the computing task and the transmission task and the table.
[0142] For the detailed content, refer to the description of step 302.
[0143] Step 703, the asset management device 116 generates a unique virtual ID for the table according to the ip, port, db name, and table name of the table.
[0144] For details, see the description of the first position information, the second position information in the table in step 303, and Figure 5 the description of
[0145] Step 704: The data asset to be generated by the asset management device 116 is associated with the job through the unique virtual ID of the table, and the job associated with the data asset to be generated is obtained.
[0146] For details, see the description of Figure 5 and step 303.
[0147] Step 705: The asset management device 116 determines the job link of the data asset to be generated based on the job associated with the data asset to be generated and the dependency relationship between jobs.
[0148] For details, see the relevant description of the job link in step 303.
[0149] Step 706: For each job in the job link, the asset management device 116 determines whether the historical execution duration of the job jumps. If so, step 709 is executed; if not, step 707 is executed.
[0150] It should be noted that the jump of the historical execution duration is used to indicate that there are abnormalities in the historical execution duration, such as excessive values.
[0151] Step 707: The asset management device 116 determines whether the historical execution duration deteriorates. If so, step 708 is executed.
[0152] To determine whether the historical execution duration deteriorates, refer to the description of determining the deterioration situation of the job above.
[0153] Step 708: The asset management device 116 generates a job governance task.
[0154] The job governance task can be a task generated based on the governance solution of the job and is used to optimize the execution of the job.
[0155] Step 709: The asset management device 116 eliminates the abnormal values in the historical execution duration and takes the historical execution duration of TP95.
[0156] For details, see the description of obtaining the historical execution duration in step 304.
[0157] Step 710: The asset management device 116 determines the critical job path of the data asset to be generated based on the historical execution duration of each job in the job link.
[0158] For details, see the description of obtaining the critical job path in step 304.
[0159] Step 711: The asset management device 116 determines whether the data asset SLA is achieved based on the historical execution duration of each job in the critical operation path. If not, step 712 is executed.
[0160] For the detailed content, refer to the above description of the situation where the operation link cannot be completed on time.
[0161] Step 712: The asset management device 116 controls the data lake 111 to execute the policy for computing resource allocation.
[0162] For the detailed content, refer to the above description of the computing resource allocation for updating the operation link. The policy for computing resource classification can be referred to the above description of the guarantee policy.
[0163] Figure 7b This is a schematic diagram of a specific application of a method for determining the generation duration of data assets provided for the implementation of the present invention. As Figure 7b shown, the asset management device 116 may include an asset operation parsing module 1161, an asset generation estimation module 1162, an asset SAL guarantee magic armor 1163, and an asset guarantee governance module 1164.
[0164] Among them, the asset operation parsing module 1161 can interface with the data asset system 114 and the scheduling system 112. The asset management device 116 can periodically or real-time collect the information of the data assets and the operation information of the data asset system 114. The asset operation parsing module 1161 may include a graph engine to implement multiple functions, such as operation parsing, asset parsing, and relationship parsing. Among them, operation parsing can parse the operation information to obtain the tables associated with the operation and the dependency relationship between operations; asset parsing can parse the information of the data assets to obtain the tables associated with the data assets; relationship parsing can determine the operations associated with the data assets based on the description information of the tables associated with the data assets and the description information of the tables associated with the operations; store the operations associated with the data assets and the dependency relationship between operations into the graph engine; the graph engine outputs the operation link of the data assets. For the detailed content, refer to the description of the above steps 301 to 303 and will not be repeated here.
[0165] Among them, the asset generation estimation module 1162 can interface with the scheduling system 112, and the asset management device 116 can periodically or real-time collect the execution information of the job. The asset generation estimation module 1162 can implement multiple functions, such as noise processing, sequence smoothing, critical path generation, and time estimation. Noise processing is used to remove outliers from the job execution information. Sequence smoothing is used to smooth the job execution information after outlier removal. For example, the execution durations after outlier removal are sorted in descending order, and the execution durations of the preset percentage, such as TP95, with the highest ranking are taken to obtain the historical execution durations. Critical path generation is used to determine the critical job path based on the historical execution durations of the jobs in the job chain. Time estimation is used to estimate the time required to generate the data asset based on the execution durations of the jobs in the critical job path. For the detailed content, please refer to the description of step 304 above and will not be elaborated here.
[0166] Among them, the asset SAL guarantee module 1163 can implement multiple functions, such as risk diagnosis, risk warning, and guarantee strategy generation. Among them, risk diagnosis; among them, risk diagnosis is used to judge whether the job chain can be completed on time. Risk warning is used to give an alarm when the job chain cannot be completed on time. Guarantee strategy generation is used to generate a guarantee strategy under the condition that the job chain cannot be completed on time. The guarantee strategy is used to update the resource allocation of the job chain and improve the execution priority of the jobs in the job chain.
[0167] Among them, the asset guarantee governance module 1164 can implement multiple functions, such as job deterioration assessment, job deterioration suggestion, and generation of job governance tasks. Among them, job deterioration assessment is used to determine the deterioration situation of each job in the job chain based on the historical execution situation of the job. Job deterioration suggestion is used to give a governance plan for the job when the deterioration situation of the job indicates job deterioration. Generation of job governance tasks is used to generate a governance task for the governance plan of the job. The governance task is used to implement the governance plan of the job.
[0168] The present invention also provides a device for determining the generation duration of a data asset, as Figure 8 shown, including:
[0169] An association module 801, configured to determine a job chain corresponding to the data asset to be generated based on the table associated with the data asset to be generated, the table associated with the job, and other jobs on which the job depends; the table associated with the data asset to be generated is the table in the data warehouse required to generate the data asset to be generated; the table associated with the job is the table stored in the data warehouse after job processing; the job chain indicates the sequence of jobs required to generate the table associated with the data asset to be generated.
[0170] The duration determination module 802 is configured to determine the duration required to generate a data asset based on the historical execution duration of jobs in the job link of the data asset to be generated.
[0171] Among them, both the association module 801 and the duration determination module 802 can be implemented by software or by hardware. Exemplarily, next, taking the association module 801 as an example, the implementation manner of the association module 801 is introduced. Similarly, the implementation manner of the duration determination module 802 can refer to the implementation manner of the association module 801.
[0172] As an example of a software functional unit, the association module 801 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the association module 801 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code may be distributed in the same region, or may be distributed in different regions. Further, the multiple hosts / virtual machines / containers for running the code may be distributed in the same availability zone (AZ), or may be distributed in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Among them, generally one region may include multiple AZs.
[0173] Similarly, the multiple hosts / virtual machines / containers for running the code may be distributed in the same virtual private cloud (VPC), or may be distributed in multiple VPCs. Among them, generally one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is realized through the communication gateway.
[0174] As an example of a hardware functional unit, the associated module 801 may include at least one computing device, such as a server, etc. Alternatively, the associated module 801 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0175] The multiple computing devices included in the associated module 801 may be distributed in the same region or in different regions. The multiple computing devices included in the associated module 801 may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the associated module 801 may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0176] It should be noted that in other embodiments, the associated module 801 may be used to execute any step in the method for determining the generation duration of data assets, and the duration determination module 802 may be used to execute any step in the method for determining the generation duration of data assets. The steps implemented by the associated module 801 and the duration determination module 802 can be specified as needed. The entire function of the apparatus for determining the generation duration of data assets is realized by the associated module 801 and the duration determination module 802 respectively implementing different steps in the method for determining the generation duration of data assets.
[0177] The present invention also provides a computing device 900. As Figure 9 shown, the computing device 900 includes: a bus 902, a processor 904, a memory 906, and a communication interface 908. The processor 904, the memory 906, and the communication interface 908 communicate with each other through the bus 902. The computing device 900 may be a server or a terminal device. It should be understood that the present invention does not limit the number of processors and memories in the computing device 900.
[0178] The bus 902 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 9 only one line is used in Figure 9 , but it does not mean that there is only one bus or one type of bus. The bus 902 can include a path for transmitting information between various components of the computing device 900 (such as the memory 906, the processor 904, and the communication interface 908).
[0179] The processor 904 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0180] The memory 906 can include volatile memory, such as random access memory (RAM). The processor 904 can also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0181] The memory 906 stores executable program code, and the processor 904 executes the executable program code to respectively implement the functions of the foregoing associated module 801 and duration determination module 802, thereby implementing the method for determining the data asset generation duration. That is, the memory 906 stores instructions for executing the method for determining the data asset generation duration.
[0182] The communication interface 908 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 900 and other devices or a communication network.
[0183] An embodiment of the present invention also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0184] As Figure 10 shown, the computing device cluster includes at least one computing device 900. Instructions for executing the method for determining the generation duration of data assets can be stored in the memories 906 of one or more of the computing devices 900 in the computing device cluster.
[0185] In some possible implementation manners, partial instructions for executing the method for determining the generation duration of data assets can also be stored separately in the memories 906 of one or more of the computing devices 900 in the computing device cluster. In other words, a combination of one or more computing devices 900 can jointly execute the instructions for executing the method for determining the generation duration of data assets.
[0186] It should be noted that the memories 906 in different computing devices 900 in the computing device cluster can store different instructions, respectively for executing partial functions of the apparatus for determining the generation duration of data assets. That is, the instructions stored in the memories 906 of different computing devices 900 can implement the functions of one or more of the associated module 801 and the duration determination module 802. It is worth noting that the computing device cluster can be deployed Figure 1 in the data job platform 110.
[0187] In some possible implementation manners, one or more computing devices in the computing device cluster can be connected through a network. Wherein, the network can be a wide area network or a local area network, etc. Figure 11 shows a possible implementation manner. As Figure 11 shown, two computing devices 900A and 900B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation manners, instructions for executing the function of the associated module 801 are stored in the memory 906 of the computing device 900A. At the same time, instructions for executing the function of the duration determination module 802 are stored in the memory 906 of the computing device 900B.
[0188] Figure 11 The connection manner between the computing device clusters shown can be considered because the method for determining the generation duration of data assets provided by the present invention needs to parse jobs and data assets, and then analyze the historical operation data of the jobs. Therefore, it is considered to hand over the associated module 801 and the duration determination module 802 to different computing devices 900 for execution.
[0189] It should be understood that Figure 11 the functions of the computing device 900A shown in FIG. may also be completed by multiple computing devices 900. Similarly, the functions of the computing device 900B may also be completed by multiple computing devices 900.
[0190] Embodiments of the present invention also provide a computer program product including instructions. The computer program product may be software or a program product including instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes at least one computing device to execute a method for determining the generation duration of data assets, or a method for determining the generation duration of data assets.
[0191] Embodiments of the present invention also provide a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center including one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive), etc. The computer-readable storage medium includes instructions that instruct the computing device to execute a method for determining the generation duration of data assets, or instruct the computing device to execute a method for determining the generation duration of data assets.
[0192] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0193] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0194] The basic principles of the present invention have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present invention are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-disclosed specific details are only for the purposes of illustration and easy understanding, rather than limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details to implement.
[0195] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including," "comprising," "having," etc. are open-ended terms that mean "including but not limited to" and can be used interchangeably with each other. The words "or" and "and" used herein mean "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The phrase "such as" used herein means "such as but not limited to" and can be used interchangeably with it.
[0196] It should also be noted that in the apparatuses, equipment, and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.
[0197] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.
[0198] It can be understood that the various numerical numbers involved in the embodiments of the present invention are only for the convenience of description and are not used to limit the scope of the embodiments of the present invention.
Claims
1. A method for determining the generation duration of data assets, characterized in that, Including: Determine the job link corresponding to the data asset to be generated based on the tables associated with the data asset to be generated, the tables associated with the job, and other jobs upon which the job depends. The tables associated with the data asset to be generated are the tables in the data warehouse required to generate the data asset to be generated. The tables associated with the job are the tables stored in the data warehouse after the job is processed; the job link indicates the sequence of precedence among the jobs required to generate the tables associated with the data asset to be generated. Determine the duration required to generate the data asset based on the historical execution durations of the jobs in the job link of the data asset to be generated.
2. The method according to claim 1, wherein The determining of the job link corresponding to the data asset to be generated based on the tables associated with the data asset to be generated, the tables associated with the job, and other jobs upon which the job depends includes: Determine the jobs associated with the data asset to be generated based on the tables associated with the data asset to be generated and the tables associated with the job, where the tables associated with the jobs associated with the data asset to be generated are used to generate the data asset to be generated. Determine the job link corresponding to the data asset to be generated based on the other jobs upon which the job depends and the jobs associated with the data asset to be generated.
3. The method according to claim 2, wherein The determining of the jobs associated with the data asset to be generated based on the tables associated with the data asset to be generated and the tables associated with the job includes: Determine the first identifier of the table associated with the data asset to be generated based on the first position information of the table associated with the data asset to be generated; the first position information indicates the position of the table associated with the data asset to be generated in the data warehouse. Determine the second identifier of the table associated with the job based on the second position information of the table associated with the job; the second position information indicates the position of the table associated with the job in the data warehouse. Determine the jobs associated with the table where the first identifier and the second identifier are the same and the data asset to be generated.
4. The method according to claim 3, characterized in that, The first position information includes the name of the table associated with the data asset to be generated and the address of the data warehouse where it is located. The second position information includes the name of the table associated with the job and the address of the data warehouse where it is located. When the data warehouse indicated by the data warehouse where the table associated with the job is located and the data warehouse where the table associated with the data asset to be generated is located is the same, the first identifier and the second identifier are the same.
5. The method according to any one of claims 1 to 4, characterized in that The determining of the duration required to generate the data asset based on the historical execution durations of the jobs in the job link of the data asset to be generated includes: Determine the critical job path in the job link based on the historical execution durations of each job in the job link; the critical job path is the path formed by the critical jobs that determine the generation time of the data asset to be generated. Determine the duration required to generate the data asset based on the historical execution durations of the jobs in the critical job path.
6. The method according to claim 5, characterized in that The method further includes: For each job in the job link, eliminate the outliers from the execution durations at different collection time points of the job; based on the execution durations after outlier elimination, determine the historical execution duration of the job.
7. The method according to claim 6, wherein The historical execution duration is the target execution duration among the execution durations after outlier removal, and the proportion of the execution durations after outlier removal that are less than or equal to the target execution duration is a preset proportion.
8. The method according to any one of claims 1 to 7, characterized in that The job chain is formed by sequentially connecting jobs in multiple layers. Among them, the data asset to be generated is obtained after the jobs in the last layer of the multiple layers are executed; for each layer other than the last layer in the multiple layers, each job in the layer is connected to at least one job in the next layer; the critical job chain is the path formed by the jobs with the longest historical execution duration in each layer of the multiple layers.
9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: Updating the computing resource allocation of the job chain in the case where the job chain cannot be completed on time.
10. The method according to any one of claims 1 to 9, characterized in that, The method further includes: For at least some of the jobs in the job chain, determining the deterioration condition of the job, and when the deterioration condition of the job indicates job deterioration, giving an alarm and / or proposing a governance solution for the job, where the governance solution is used to reduce the deterioration condition of the job.
11. An apparatus for determining the generation duration of a data asset, characterized in that, The device includes: An association module, configured to determine the job chain corresponding to the data asset to be generated based on the tables associated with the data asset to be generated, the tables associated with the jobs, and the other jobs on which the jobs depend; the tables associated with the data asset to be generated are the tables in the data warehouse required to generate the data asset to be generated; the tables associated with the jobs are the tables stored in the data warehouse after the jobs are processed; the job chain indicates the sequence of jobs required to generate the tables associated with the data asset to be generated. A duration determination module, configured to determine the duration required to generate the data asset based on the historical execution durations of the jobs in the job chain of the data asset to be generated.
12. A cluster of computing devices, characterized in that, It includes at least one computing device, and each computing device includes a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 10.
13. A computer program product comprising instructions, characterized in that, When the instructions are run by the computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 10.
14. A computer-readable storage medium, characterized in that, It includes computer program instructions, and when the computer program instructions are executed by the computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 10.