A visualization processing method and a visualization processing device based on metadata management

By registering each target task in the data processing process metadata relationship, forming a metadata traceability diagram, realizing the visual effect of data processing, solving the problem that traditional methods cannot manage complex data processing tasks.

CN115168457BActive Publication Date: 2025-07-01CHINA TOBACCO ANHUI IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210473258.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-07-01
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

Traditional SQL statements cannot achieve visualization of data processing and cannot view the blood relationship of data, resulting in difficulty in data processing and management.

Method used

By registering each target task in the data processing process metadata relationship, a metadata traceability diagram of the data processing process is formed, and related target tasks are connected together to form a traceability topological relationship diagram of the corresponding data and tasks, the visualization effect of data processing is realized, and the full-link task scheduling and monitoring functions are integrated in the data governance life cycle.

Benefits of technology

It realizes the visual effect of data processing, facilitates users to manage tasks, and solves the problem of chaotic and ineffective governance and maintenance of massive tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115168457B_ABST
    Figure CN115168457B_ABST
Patent Text Reader

Abstract

The present application provides a visualization processing method and a visualization processing device based on metadata management. The visualization processing method includes: receiving a data processing flow set by a user; for each target task in the data processing flow, based on the task type and initial configuration information of the target task, performing metadata relationship registration on the target task to generate a metadata relationship diagram corresponding to the target task in a visualization interface; based on the execution order of each target task in the data processing flow and the metadata relationship diagrams of each target task, generating a metadata traceability diagram corresponding to the data processing flow in the visualization interface. According to the visualization processing method and the visualization processing device, related target tasks are connected in series to form a traceability topology diagram of corresponding data and tasks, realizing the visualization effect of data processing and solving the problem that a large number of tasks are chaotic and cannot be effectively managed and maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing. Specifically, it relates to a visualization processing method and a visualization processing device based on metadata management. Background Art

[0002] With the rapid development of Internet technology, the speed and quantity of data generation have also increased. In order to effectively utilize this data with hidden value, there is usually a need for secondary processing during or before data transmission, such as encryption / decryption and desensitization of sensitive data, parsing of semi-structured data, and secondary calculation of data.

[0003] However, the processing of data through traditional SQL statements has problems such as the inability to achieve a visualization effect and the inability to view the data lineage. Therefore, how to achieve the visualization effect of data processing has become an urgent problem to be solved. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a visualization processing method and a visualization processing device based on metadata management. By registering metadata relationships for each target task in the data processing flow, a metadata traceability graph of the data processing flow is formed, and the related target tasks are connected in series to form a corresponding traceability topology relationship graph of data and tasks, so that users can effectively manage tasks according to the traceability topology relationship graph, achieving the visualization effect of data processing. And the data processing flow integrates data tasks from different task platforms and at different stages throughout the data governance life cycle, providing unified data task scheduling and monitoring functions, realizing the scheduling and monitoring of data tasks throughout the data governance life cycle, and solving the problem in the prior art that a large number of tasks are chaotic and cannot be effectively governed and maintained.

[0005] In a first aspect, an embodiment of this application provides a visualization processing method based on metadata management. The visualization processing method includes:

[0006] Receiving a data processing flow set by a user; wherein, the data processing flow includes at least one target task from different task platforms that can implement different functional logics;

[0007] For each target task in the data processing flow, based on the task type and initial configuration information of the target task, registering metadata relationships for the target task to generate a metadata relationship graph corresponding to the target task in a visualization interface; wherein, the initial configuration information is used to represent the processing method of the target task, and the metadata relationship graph is used to represent the data lineage of the target task;

[0008] Generate a metadata traceability graph corresponding to the data processing flow in the visualization interface based on the execution order of each target task in the data processing flow and the metadata relationship graph of each target task.

[0009] Further, registering metadata relationships for the target task based on the task type and initial configuration information of the target task to generate a metadata relationship graph corresponding to the target task in the visualization interface includes:

[0010] Obtain at least one initial configuration information determined by the user when setting the target task;

[0011] Obtain at least one target configuration information from the at least one initial configuration information according to the task type of the target task; wherein, the target configuration information is used to generate the metadata relationship graph;

[0012] For each target configuration information, determine the connection relationship between the target configuration information and the topological node corresponding to the target task in the metadata relationship graph according to the data attribute to which the target configuration information belongs; wherein, the connection relationship includes a connection line and the arrow direction on the connection line;

[0013] Add the target configuration information and the topological node corresponding to the target task to the visualization interface, and draw according to the connection relationship between the target configuration information and the topological node corresponding to the target task in the metadata relationship graph to generate a metadata relationship graph corresponding to the target task.

[0014] Further, generating a metadata traceability graph corresponding to the data processing flow in the visualization interface based on the execution order of each target task in the data processing flow and the metadata relationship graph of each target task includes:

[0015] Determine the execution order of each target task based on the data processing flow;

[0016] For each target task, based on the execution order of the target task, determine the adjacent tasks executed before and / or after the target task from the data processing flow;

[0017] Connect the metadata relationship graph of the target task with the metadata relationship graphs of the adjacent tasks to generate a metadata traceability graph corresponding to the data processing flow.

[0018] Further, after generating the metadata traceability graph corresponding to the data processing flow, the visualization processing method further includes:

[0019] When executing the task instance corresponding to the data processing flow, monitor the execution status of each target task;

[0020] For each target task, based on the execution status of the target task, render the display color corresponding to the execution status at the position of the topology node corresponding to the target task in the metadata traceability graph.

[0021] Furthermore, the visualization processing method further includes:

[0022] For each target task in the data processing flow, determine the target execution node corresponding to the target task from at least one pre-configured execution node according to a pre-set load policy, so that the target execution node executes the target task.

[0023] Furthermore, the step of selecting the corresponding target execution node for the target task according to the pre-set load policy includes:

[0024] When the load policy is a random load policy, determine the target execution node corresponding to the target task through the following steps:

[0025] Randomly select from at least one execution node to determine the target execution node corresponding to the target task;

[0026] When the load policy is a weighted round-robin load policy, determine the target execution node corresponding to the target task through the following steps:

[0027] For each execution node, determine the available memory information and available load information of the execution node;

[0028] Use the available memory information and available load information of the execution node for weighted calculation to obtain the available resource value of the execution node;

[0029] Sort each execution node using the available resource value of each execution node, and use the execution node with the highest available resource value among the multiple execution nodes as the target execution node;

[0030] When the load policy is a concurrent load policy, determine the target execution node corresponding to the target task through the following steps:

[0031] For each execution node, determine the number of concurrent task instances of the execution node;

[0032] Sort each execution node using the number of concurrent task instances of each execution node, and use the execution node with the fewest number of concurrent task instances among the multiple execution nodes as the target execution node;

[0033] When the load policy is a resource-weighted load policy, the target execution node corresponding to the target task is determined through the following steps:

[0034] For each execution node, determine the memory free information and task instance free information of the execution node;

[0035] Use the memory free information and task instance free information of the execution node to perform weighted calculation to obtain the free resource value of the execution node;

[0036] Sort each execution node using the free resource value of each execution node, and use the execution node with the highest free resource value among the multiple execution nodes as the target execution node.

[0037] Further, after determining the target execution node corresponding to each target task, the visualization processing method further includes:

[0038] For each target task, determine whether the task execution time of the target execution node corresponding to the target task is greater than or equal to the execution time threshold;

[0039] If so, determine that the target task has a running exception, and re-determine the corresponding target execution node for the target task based on the load policy;

[0040] According to a preset notification method, send the abnormal operation information of the target task to a preset notified user, so that the notified user can determine the running situation of the target task according to the abnormal operation information.

[0041] In a second aspect, an embodiment of the present application further provides a visualization processing device based on metadata management. The visualization processing device includes:

[0042] A receiving module, configured to receive a data processing flow set by a user; wherein, the data processing flow includes at least one target task from different task platforms that can implement different functional logics;

[0043] A metadata relationship diagram determination module, configured to, for each target task in the data processing flow, perform metadata relationship registration on the target task based on the task type and initial configuration information of the target task, so as to generate a metadata relationship diagram corresponding to the target task in a visualization interface; wherein, the initial configuration information is used to characterize the processing method of the target task, and the metadata relationship diagram is used to characterize the data lineage of the target task;

[0044] A metadata traceability graph determination module, configured to generate a metadata traceability graph corresponding to the data processing flow in the visualization interface based on the execution order of each target task in the data processing flow and the metadata relationship graph of each target task.

[0045] In a third aspect, an embodiment of the present application further provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the visualization processing method based on metadata management as described above are executed.

[0046] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the visualization processing method based on metadata management as described above are executed.

[0047] The visualization processing method based on metadata management provided by the embodiments of the present application first receives a data processing flow set by a user; then, for each target task in the data processing flow, based on the task type and initial configuration information of the target task, metadata relationship registration is performed on the target task to generate a metadata relationship graph corresponding to the target task in the visualization interface; finally, based on the execution order of each target task in the data processing flow and the metadata relationship graph of each target task, a metadata traceability graph corresponding to the data processing flow is generated in the visualization interface. Compared with the methods in the prior art, the present application performs metadata relationship registration on each target task in the data processing flow, analyzes the data lineage relationship of the entire process, forms a metadata traceability graph of the data processing flow, and connects related target tasks in series to form a corresponding traceability topology relationship graph of data and tasks, so that users can effectively manage tasks according to the traceability topology relationship graph, realizing the visualization effect of data processing. And the data processing flow integrates data tasks at all stages from different task platforms in the entire data governance life cycle, provides a unified data task scheduling and monitoring function, realizes the scheduling and monitoring of data tasks in the entire data governance life cycle, and solves the problem that a large number of tasks in the prior art are chaotic and cannot be effectively governed and maintained.

[0048] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. Description of the Drawings

[0049] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0050] Figure 1 It is a flowchart of a visualization processing method based on metadata management provided by an embodiment of the present application;

[0051] Figure 2 It is a flowchart of a method for generating a metadata relationship diagram provided by an embodiment of the present application;

[0052] Figure 3 It is a schematic structural diagram of a visualization processing device based on metadata management provided by an embodiment of the present application;

[0053] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, every other embodiment obtained by those skilled in the art without creative efforts falls within the scope of protection of the present application.

[0055] With the rapid development of Internet technology, the speed and quantity of data generation have also increased. To effectively utilize this data with hidden value, there is usually a need for secondary processing during or before data transmission, such as encryption / decryption and desensitization of sensitive data, parsing of semi-structured data, and secondary calculation of data. However, the processing of data through traditional SQL statements has problems such as inability to achieve a visualization effect and inability to view the data lineage. Therefore, how to achieve the visualization effect of data processing has become an urgent problem to be solved.

[0056] In the context of the "digitalization" of enterprise informatization construction, a series of digital technologies such as data middle platform methodology, big data, and artificial intelligence have gradually matured and achieved certain results in various industries. The digital transformation of enterprises has gradually become the trend of future development. Under the new background of digital construction, the way of enterprise informatization construction has gradually changed, shifting from the traditional chimney-style application construction mode to a data operation mode centered on the enterprise-level data platform. The business analysis and processing logic of data has gradually shifted from traditional chimney-style applications to the enterprise-level data platform. That is, in the traditional enterprise application construction, each application collects data separately, processes data and conducts business analysis within the application. The service capabilities of each application cannot be effectively reused, shared, or extended, resulting in the problem of "data islands" in enterprises. In the new mode, the enterprise-level data center is responsible for collecting all data, formulating data standard specifications, sorting out data assets, and conducting business analysis and data processing according to business needs, and opening and sharing the processed data results to enterprise applications. In this mode, business needs can be quickly, agilely, and efficiently responded to, and data capabilities can be highly reused and shared. Moreover, compared with the traditional application analysis mode, the data processing mode based on the data platform makes the data processing and analysis methods more convenient, adjustable, extensible, and manageable due to various tool platforms.

[0057] To promote the digital transformation of enterprises, build an enterprise-level data foundation, and construct a new data processing platform for the entire domain to solve the problems of complex and changing massive data processing scenarios. The following aspects need to be considered in the design: In terms of data processing capabilities, in the face of the sinking of massive business application analysis requirements, the data platform needs to have comprehensive data processing capabilities to solve complex and changing data processing requirements. It includes traditional application function processing capabilities, script processing capabilities such as Shell and Python, data processing methods based on database SQL and stored procedures, real-time and offline processing capabilities of big data, AI algorithm analysis capabilities of data, etc., and needs to provide the expansion ability of corresponding processing capabilities according to the development of industry technologies. In terms of scalability, since it is a data processing system of the enterprise-level data platform, facing the processing and analysis of massive businesses, the task volume of data processing will continuously expand with the expansion of business and data operation. Therefore, considering from the architecture, the processing nodes of data support dynamic expansion. In terms of task management capabilities, because it is necessary to support a large number of data processing tasks, there are certain requirements for task management scheduling, monitoring and alarming, and task link tracking. For example, when an abnormal situation occurs in a certain data indicator, the data object and the ETL link of the task can be quickly traced according to the metadata, and the problem can be accurately and efficiently located and solved.

[0058] To implement the above-mentioned scalable distributed scheduling multi-data processing system, the mainstream data scheduling engines in the industry are investigated, mainly including xxl-job, Ooize, and Apache's dolphinscheduler. Apache DolphinScheduler is an open-source system for distributed and easily extensible visual DAG workflow task scheduling. It solves the intricate dependency relationships in data research and development ETL and the problem of not being able to intuitively monitor the health status of tasks.

[0059] However, when it comes to implementation in enterprise-level data platforms, DolphinScheduler cannot perform analysis and registration of task and data lineage relationships. In the scenario of enterprise-scale data tasks, without data lineage analysis, task management becomes chaotic, and data or metrics cannot be effectively traced and traced along the task chain. DolphinScheduler is a visual distributed scheduling system designed for big data analysis and processing. In terms of data processing tasks, whether it is big data's MR, Flink, Spark, or SQL and Produce based on databases, they are basically based on scripts or programs, requiring very high professional knowledge from users. Currently, DolphinScheduler lacks a visual and easy-to-operate data processing task construction method based on dragging and configuring processing components. DolphinScheduler focuses on the data processing stage, while in the enterprise-level data platform architecture, task scheduling needs to cover the entire life cycle of data exploration, data integration, data development and processing, data quality detection, data sharing, etc. DolphinScheduler lacks in the integrated scheduling with other third-party data tasks and cannot achieve one-stop construction and scheduling of data tasks throughout the life cycle.

[0060] Based on this, the embodiment of the present application provides a visual processing method based on metadata management. By registering metadata relationships for each target task in the data processing flow, analyzing the data lineage relationship of the entire process, forming a metadata traceability graph of the data processing flow, and connecting related target tasks in series to form a corresponding traceability topology graph of data and tasks, so that users can effectively manage tasks according to the traceability topology graph, achieving the visual effect of data processing. And the data processing flow integrates data tasks from different task platforms and at all stages throughout the data governance life cycle, providing unified data task scheduling and monitoring functions, realizing the scheduling and monitoring of data tasks throughout the data governance life cycle, and solving the problem of chaotic and ineffective governance and maintenance of a large number of tasks in the prior art.

[0061] Please refer to Figure 1 , Figure 1The flowchart of a visualization processing method based on metadata management provided by an embodiment of the present application. As Figure 1 shown, the visualization processing method based on metadata management provided by an embodiment of the present application includes:

[0062] S101, receiving the data processing flow set by the user.

[0063] It should be noted that the data processing flow refers to the data processing process for processing data. Among them, the data processing flow includes at least one target task from different task platforms that can implement different functional logics. Here, the target task refers to the task configured in the data processing flow for different processing of data. According to the embodiment provided by the present application, the target task can come from different task platforms and be used to implement different functional logics. As an example, the target task can include data exploration tasks, data quality inspection tasks, data reconciliation tasks, data processing tasks, data integration tasks, data sharing tasks, etc. including data integration platforms, data control platforms, and data publishing platforms.

[0064] Here, it should be noted that the above examples of target tasks are only examples. In practice, the target tasks are not limited to the above examples.

[0065] In the landing scenario of the data platform, in the face of a business requirement scenario, it is necessary to build data tasks with corresponding functions in each subsystem of the data platform. According to the life cycle, data tasks include different stages such as data exploration, data integration, data development and processing, data quality inspection, data reconciliation, and data sharing. The data governance tasks for the same data requirement are often distributed in different systems, scheduled and monitored separately, which is extremely inconvenient, and it is impossible to build a data governance link and dependency relationship, and it is impossible to effectively implement the dependency scheduling between tasks. To solve such problems, the present application uses a unified one-stop task integration and scheduling system. The system integrates data tasks in all stages of the entire data governance life cycle, including data exploration tasks, data integration tasks, data quality inspection tasks, data reconciliation tasks, and data sharing tasks including data integration platforms, data control platforms, and data publishing platforms, and provides unified data task scheduling and monitoring functions accordingly. The system realizes cross-platform scheduling and monitoring capabilities through the integration interfaces provided by each platform, and realizes the scheduling and monitoring of data tasks in the entire data governance life cycle.

[0066] According to the embodiments provided by the present application, when setting up a data processing flow, the user can integrate data tasks from different systems to build the data task dependencies for the entire data governance life cycle by drawing a job DAG (Directed Acyclic Graph) workflow according to the data governance logic, so as to achieve one-stop management and scheduling of jobs for the entire data governance life cycle. The present application extends a method for constructing a stream-batch integrated data task based on visual drag-and-drop of DIPE engine components on top of the data processing platform. This method constructs a data integration, processing, or sharing task quickly and intuitively by dragging and dropping data extraction, processing, and loading components on the task construction page and drawing the task processing topology, and configuring component parameters according to component requirements. In addition to the simple and easy-to-use feature of visual drag-and-drop, the DIPE engine has its unique advantages compared with data processing methods such as SQL, Hive, and Flink: the independence of the DIPE engine's operation. Similar to SQL, Hive, Flink, etc., they all need to rely on databases, Yarn, etc. for operation, while the DIPE engine tasks run independently on the engine nodes with less dependence. DIPE supports the processing of multi-source heterogeneous data. Different from SQL, Hive, etc., which basically only support the calculation and processing of data within the database, the data sources supported by DIPE rely on the comprehensiveness of the extraction components and basically support various data sources such as data files, message middleware, databases, big data storage components, and distributed object storage. Another reason is that DIPE extracts and then processes the data from various data sources, so DIPE's powerful multi-source data fusion calculation and processing ability is also very strong. The DIPE engine has strong scalability. The DIPE engine supports customizing various components according to requirements, including data source extraction and loading components. For example, according to business expansion, data extraction components such as Http interfaces and WebSocket services are customized, and the extraction and loading components can be customized according to business characteristics, making the DIPE engine highly adaptable to business production scenarios.

[0067] At the same time, the DIPE engine supports customizing processing components according to the business and performing dynamic data processing in combination with extensible custom functions, which greatly improves the processing ability of visual components and greatly enriches the data processing ability of the DIPE engine. This method greatly reduces the usage threshold of the data processing platform, enabling non-technical business personnel to quickly and efficiently conduct data governance. And it can efficiently integrate and uniformly manage and schedule the tasks of the entire data governance life cycle scattered in each sub-platform. By associating these scattered tasks in the way of constructing a job DAG workflow, unified management and scheduling are carried out, thus solving the problem of dependent scheduling between tasks.

[0068] Regarding the above step S101, in specific implementation, a data processing flow constructed by a user is received. The data processing flow includes at least one target task from different task platforms that can implement different functional logics.

[0069] S102. For each target task in the data processing flow, based on the task type and initial configuration information of the target task, metadata relationship registration is performed on the target task to generate a metadata relationship diagram corresponding to the target task in the visualization interface.

[0070] It should be noted that the task type refers to the type corresponding to the target task. According to the embodiments provided in the present application, different target tasks correspond to different task types. The task types may include data exploration tasks, data quality detection tasks, data reconciliation tasks, data integration tasks, data processing tasks, data sharing tasks, etc. The initial configuration task refers to some configuration parameters determined by the user when setting the target task. The initial configuration information is used to characterize the processing method of the target task. For example, when the user wants to configure a data integration component, the initial configuration information may be at least one of the names of the required data tables and the data fields in the data tables. The name of the data table may be the name of an existing data table such as a customer table or a user table. The fields corresponding to the data attributes in the data table may be data attributes in the existing data table such as name and gender. Metadata refers to data that describes data, describes data and information resources, and is a higher-level abstraction of data. Metadata records the data included in the system, the representation of the data, the source of the data, and the flow relationship in the system. The application of metadata is extensive, and it can be used to construct business terms, data standards, data dictionaries, data asset catalogs, data lineage, and data maps, etc. Metadata relationship registration refers to using metadata to construct the data lineage corresponding to the target task. The metadata relationship diagram is the result obtained after performing metadata relationship registration and is used to characterize the data lineage in the target task. Data lineage can represent the relationship between data and reflect the production and processing process of data in the system, mainly including cluster lineage, system lineage, table-level lineage, and field-level lineage.

[0071] Here, it should be noted that the above examples of the initial configuration information are only examples, and in practice, the initial configuration information is not limited to the above examples.

[0072] Regarding the above step S102, in specific implementation, for each target task in the obtained data processing flow, according to the task type of the target task and the initial configuration information determined by the user when configuring the target task, metadata relationship registration is performed on the target task, and a metadata relationship diagram corresponding to the target task is generated in the user's visualization interface.

[0073] Please refer to Figure 2 , Figure 2 which is a flowchart of the method for generating a metadata relationship diagram provided by an embodiment of the present application. As shown in Figure 2 , registering metadata relationships for the target task based on the task type and initial configuration information of the target task to generate a metadata relationship diagram corresponding to the target task in a visual interface includes:

[0074] S201, obtaining at least one initial configuration information determined by the user when setting the target task.

[0075] Regarding the above step S201, in specific implementation, since the user can select a component template corresponding to the target task and fill in the initial configuration information in the component template when configuring the target task in the data processing flow, so that the target task can meet the user's business requirements. Therefore, it is necessary to obtain at least one initial configuration information determined by the user when setting the target task.

[0076] S202, obtaining at least one target configuration information from at least one initial configuration information according to the task type of the target task.

[0077] It should be noted that the target configuration information refers to one or several configuration information in the initial configuration information, which is used to generate the metadata relationship diagram of the target task.

[0078] Regarding the above step S202, in specific implementation, since the task types of each target task in the data processing flow are different, the configuration information required for drawing the metadata relationship diagram is also different, and the representation forms of the metadata relationship diagrams of target tasks with different task types are not very consistent. Therefore, it is necessary to determine which configuration information is required for drawing the metadata relationship diagram according to the task type of the target task, that is, to obtain at least one target configuration information from at least one initial configuration information according to the task type of the target task. Continuing the embodiment in step S101, when the target task includes data profiling tasks, data quality inspection tasks, data reconciliation tasks, data processing tasks, data integration tasks, and data sharing tasks, data profiling tasks, data quality detection tasks, and data reconciliation tasks belong to one category. The metadata relationship diagrams of these tasks express the relationship between data tasks and one or more data objects or object attributes. For example, for a data profiling task, the metadata diagram of this type of task expresses the profiling relationship for one or more data objects or object attributes. Therefore, this type of task needs to register the relationship between the task and the data object. That is, the required target configuration information can be the data object. And the three types of data integration tasks, data processing tasks, and data sharing tasks are another category. This category expresses an ETL (Extract-Transform-Load) task, which means extracting, processing, and loading data from one or more data objects into one or more target objects. This type of relationship diagram not only needs to register the relationship between the task and the data object, but also needs to register the mapping relationship from the source object attribute to the target object attribute. That is, the required target configuration information also includes the source data and the target data. As an optional implementation manner, the target configuration information may also include a certain or certain data fields in a certain table. For example, for a data integration task, the user needs to obtain all the data under a certain data field in data table A. For example, to obtain the customer name and customer gender in the customer table, the user needs to fill in "source data" as data table A and the data fields customer name and customer gender in data table A when configuring the data integration task. At this time, data table A, customer name, and customer gender are all the target configuration information required for generating the metadata relationship diagram.

[0079] S203. For each target configuration information, determine the connection relationship between the target configuration information and the topological node corresponding to the target task in the metadata relationship diagram according to the data attribute to which the target configuration information belongs.

[0080] It should be noted that the data attribute refers to the data type corresponding to the target configuration information. Specifically, the data attribute to which the target configuration information belongs can be determined according to the configuration method set by the user in the component template. The topological node corresponding to the target task refers to the position node corresponding to the target task in the visualization interface. The connection relationship between the target configuration information and the topological node includes the connection line between the target configuration information and the topological node and the arrow direction on the connection line.

[0081] For the above-mentioned step S203, in specific implementation, for each target configuration information of the target task, according to the data attribute to which the target configuration information belongs, determine the connection relationship between the target configuration information and the topological node corresponding to the target task in the metadata relationship diagram. As an example, when the user configures a data processing task, it is necessary to configure which data tables or which data are the source data, and what kind of data tables or data are the target data obtained after data processing. Therefore, there are configuration information determination boxes for "source data" and "target data" in the data processing component template. The user needs to input or select the required source table, such as data table A, in the configuration information determination box for "source data", and input or select the target table that they want to obtain, such as data table B, in the configuration information determination box for "target data". In this way, the target configuration information of the data processing task obtained in step S203 can be "data table A" and "data table B". Then, determine that the data attribute of data table A is source data, and the data attribute of data table B is target data. According to the normal data processing process, it can be known that the source data needs to be processed to obtain the target data. At this time, the connection relationship between the target configuration information and the topological node corresponding to the target task in the metadata relationship diagram can be determined. "Data table A" needs to be connected to the topological node through connection line 1, and then the topological node is connected to "data table B" through connection line 2. The arrow on connection line 1 should point to the topological node, and the arrow on connection line 2 should point to "data table B".

[0082] S204, add the target configuration information and the topological node corresponding to the target task to the visualization interface, and draw according to the connection relationship between the target configuration information and the topological node corresponding to the target task in the metadata relationship diagram, so as to generate the metadata relationship diagram corresponding to the target task.

[0083] For the above step S204, in specific implementation, after the target configuration information required for drawing the metadata relationship diagram and the connection relationship between each target configuration information and the topological node corresponding to the target task in the metadata relationship diagram are both determined, each target configuration information and the topological node corresponding to the target task are respectively added to the visualization interface, and the connection relationship between each target configuration information and the topological node corresponding to the target task in the metadata relationship diagram determined in step S203 is used for drawing to obtain the metadata relationship diagram corresponding to the target task.

[0084] S103. Based on the execution order of each target task in the data processing flow and the metadata relationship diagram of each target task, generate a metadata traceability diagram corresponding to the data processing flow in the visualization interface.

[0085] It should be noted that the metadata traceability diagram refers to the finally generated schematic diagram used to represent the data lineage relationship within the entire data processing flow, and is used to reflect the production and processing process of the entire data processing flow.

[0086] For the above step S103, in specific implementation, since the execution order of each target task is configured in the data processing flow, the metadata traceability diagram corresponding to the data processing flow can be generated in the visualization interface based on the execution order of each target task in the data processing flow and the metadata relationship diagram of each target task, so that users can view the data lineage relationship of the data in the data processing flow according to this metadata traceability diagram.

[0087] Regarding the above step S103, the generating the metadata traceability diagram corresponding to the data processing flow in the visualization interface based on the execution order of each target task in the data processing flow and the metadata relationship diagram of each target task includes:

[0088] Step 1031. Based on the data processing flow, determine the execution order of each target task.

[0089] For the above step 1031, in specific implementation, based on the data processing flow set by the user, determine the execution order of each target task. As an example, when target task A is connected to target task B and target task B is connected to target task C in the data processing flow, the execution order of the target tasks is to execute target task A first, then execute target task B, and finally execute target task C.

[0090] Step 1032. For each target task, based on the execution order of the target task, determine the adjacent tasks that are executed before and / or after the target task from the data processing flow.

[0091] Regarding the above step 1032, in specific implementation, after the execution order of each target task is determined, for each target task, based on the execution order of this target task, adjacent tasks that are executed before and / or after this target task are determined from the data processing flow. Continuing with the embodiment in step 1031, when the execution order of the target tasks is to first execute target task A, then execute target task B, and finally execute target task C, for target task A, the adjacent task is target task B; for target task B, the adjacent tasks are target task A and target task C; for target task C, the adjacent task is target task B.

[0092] Step 1033, connect the metadata relationship graph of this target task with the metadata relationship graph of the adjacent task to generate the metadata traceability graph corresponding to the data processing flow.

[0093] Regarding the above step 1033, in specific implementation, after the adjacent tasks of each target task are determined, the metadata relationship graphs of each target task have been drawn in the visualization interface. At this time, for each target task, connect the metadata relationship graph of this target task with the metadata relationship graphs of the adjacent tasks, and the metadata traceability graph corresponding to the data processing flow can be generated.

[0094] This application designs and implements a visualization processing method based on metadata management. In the data processing platform, different-stage target tasks are constructed through the task management page. After the construction is completed, metadata relationship registration is performed for each target task. The background of the processing platform will automatically parse the metadata according to the task type of the target task and the configuration information of the target task. If the task supports automatic parsing, that is, use the steps provided in the above embodiment for metadata relationship registration, and the page will automatically render the source data and target data involved in the task, and automatically draw the metadata relationship graph of this target task according to the parsed relationship data. After that, it can also be submitted for registration after manual confirmation and adjustment. If it is a task type that the platform does not support automatic parsing, the data objects under the data source need to be manually selected on the registration page, and the metadata relationship is manually registered by dragging and drawing lines. In order to build an easy-to-manage and easy-to-maintain full-life-cycle data task scheduling platform, the data tasks in the entire data governance life cycle, including data exploration, data integration, data development and processing, data quality inspection, data reconciliation, data sharing, etc., are uniformly registered and managed for metadata relationships, and the data governance link is analyzed and maintained based on a unified metadata system.

[0095] As an optional implementation manner, after generating the metadata traceability graph corresponding to the data processing flow, the visualization processing method further includes:

[0096] (1) When executing the task instance corresponding to the data processing flow, monitor the execution status of each target task.

[0097] (2) For each target task, based on the execution status of the target task, render the display color corresponding to the execution status at the position of the topological node corresponding to the target task in the metadata traceability graph.

[0098] It should be noted that the execution status refers to the execution status of the target task. Here, the execution status may include "executing", "execution successful", "execution failed", etc., and the present application does not make specific limitations thereto.

[0099] For the above steps (1) and (2), in specific implementation, when executing the task instance corresponding to the data processing flow, monitor the execution status of each target task in the data processing flow. For each target task, based on the execution status of the target task, render the display color corresponding to the execution status at the position of the topological node corresponding to the target task in the metadata traceability graph. For example, when the execution status of the target task is "executing", the color yellow can be rendered at the position of the topological node corresponding to the target task in the metadata traceability graph; when the execution status of the target task is "execution successful", the color green can be rendered at the position of the topological node corresponding to the target task in the metadata traceability graph; when the execution status of the target task is "execution failed", the color red can be rendered at the position of the topological node corresponding to the target task in the metadata traceability graph. This facilitates the user to directly judge the execution status of the target task according to the display color at the position of the topological node corresponding to the target task in the metadata traceability graph.

[0100] Here, it should be noted that the above examples of the display colors corresponding to different execution statuses are only examples. In practice, the display colors corresponding to different execution statuses are not limited to the above examples.

[0101] According to the embodiments provided by the present application, since the data platform finally provides data resource services to the outside and inside in the form of data object tables, data metrics or data services. When anomalies occur in the data provided by the data services, metrics, and data object tables, traceability analysis of the services, metrics, or tables can be quickly performed through metadata, and the analysis is presented in the form of a visual metadata traceability graph. Different colors are used in the metadata traceability graph to mark the status of the metadata or target tasks. When a fault anomaly occurs in the data processing flow, the data link fault can be quickly located and analyzed through the color of the topology nodes, and by clicking on the topology nodes, jumps can be made to the data tables, tasks, metrics, and services to view their details, data, and operation logs, quickly analyze and locate the fault link, and perform fault recovery. Additionally, it supports quickly generating a fault impact report to facilitate the rapid operation and maintenance of data faults. According to the visualization processing method based on metadata management provided by the present application, based on the metadata management method, massive data tasks, data objects, metrics, and services in different business scenarios are associated, and they are managed and maintained in the most efficient way. The status of the task (such as execution success, execution failure, execution in progress, etc.) is identified through the color marking of the target task on the metadata traceability relationship graph. The problem of the data link is quickly located in a visual manner, and it supports quickly locating the metadata node of the problem and jumping to the task for detailed operation troubleshooting, quickly analyzing and solving the fault problem.

[0102] As an alternative implementation, the visualization processing method further includes:

[0103] For each target task in the data processing flow, according to a preset load policy, a target execution node corresponding to the target task is determined from at least one pre-configured execution node, so that the target execution node executes the target task.

[0104] It should be noted that the load policy refers to a policy that needs to be set in advance and is used to select the required execution node for the target task. The execution node refers to a pre-configured node for executing tasks. The target execution node refers to the execution node that executes the task instance corresponding to the target task in the data processing flow.

[0105] For the above steps, in specific implementation, for each target task in the data processing flow, according to the preset load policy, determine the target execution node corresponding to the target task among at least one pre-configured execution node, so that the target execution node executes the target task. According to the embodiments provided by the present application, after all execution nodes go online, a registration task will be created, and the node and task instance operation information will be periodically obtained for information registration of distributed nodes. The scheduling node will monitor the distributed nodes. When the scheduling node needs to perform distributed scheduling on tasks, it will select an execution node according to the real-time registration information of the obtained execution nodes combined with the configured load policy, and send the task to the specified execution node for scheduling and running.

[0106] According to the embodiments provided by the present application, the preset load policies can include four types: random load policy, weighted round-robin load policy, concurrent load policy, and resource-weighted load policy. Here, the random load policy randomly selects among all registered execution nodes; the weighted round-robin load policy calculates according to the resource values of the execution nodes, sorts them according to high and low, and then selects execution nodes by round-robin; the concurrent load policy selects execution nodes according to the number of concurrent task instances of the obtained execution nodes according to the lowest concurrency principle; the resource-weighted load policy performs weighted calculation according to the resource idle situation of the obtained execution nodes and the number of plugin concurrent task instances, and selects the execution node with the optimal resources for scheduling. Selecting the corresponding target execution node for the target task according to the preset load policy includes:

[0107] When the load policy is the random load policy, the target execution node corresponding to the target task is determined through the following steps:

[0108] Randomly select among at least one execution node to determine the target execution node corresponding to the target task.

[0109] For the above steps, when the load policy is the random load policy, for at least one configured execution node, randomly select among at least one execution node to determine the target execution node corresponding to the target task.

[0110] When the load policy is the weighted round-robin load policy, the target execution node corresponding to the target task is determined through the following steps:

[0111] A: For each execution node, determine the available memory information and available load information of the execution node.

[0112] B: Use the available memory information and available load information of the execution node to perform weighted calculation to obtain the available resource value of the execution node.

[0113] C: Sort each execution node according to the available resource value of each execution node, and use the execution node with the highest available resource value among the multiple execution nodes as the target execution node.

[0114] For the above steps A to C, in specific implementation, when the load policy is the weighted round-robin load policy, for each configured execution node, determine the current available memory information and available load information of this execution node, and then perform weighted calculation using the available memory information and available load information to obtain the available resource value of this execution node. Statistically calculate the available resource values of each execution node, sort each execution node according to the available resource value of each execution node, and use the execution node with the highest available resource value among the multiple execution nodes as the target execution node corresponding to this target task.

[0115] When the load policy is the concurrent load policy, determine the target execution node corresponding to this target task through the following steps:

[0116] a: For each execution node, determine the number of concurrent task instances of this execution node.

[0117] b: Sort each execution node according to the number of concurrent task instances of each execution node, and use the execution node with the fewest number of concurrent task instances among the multiple execution nodes as the target execution node.

[0118] Here, the number of concurrent task instances refers to the number of task instances mounted on the execution node.

[0119] For the above steps a and b, in specific implementation, for each execution node, determine the number of concurrent task instances of this execution node. Statistically calculate the number of concurrent task instances of each execution node, sort each execution node according to the number of concurrent task instances of each execution node, and use the execution node with the fewest number of concurrent task instances among the multiple execution nodes as the target execution node corresponding to this target task.

[0120] When the load policy is the resource weighted load policy, determine the target execution node corresponding to this target task through the following steps:

[0121] I: For each execution node, determine the memory free information and task instance free information of this execution node.

[0122] II: Perform weighted calculation using the memory free information and task instance free information of this execution node to obtain the free resource value of this execution node.

[0123] III: Sort each execution node according to the free resource value of each execution node, and use the execution node with the highest free resource value among the multiple execution nodes as the target execution node.

[0124] For the above steps I to III, in specific implementation, when the load policy is the resource weighted load policy, for each configured execution node, determine the current memory free information and task instance free information of the execution node, and then perform weighted calculation using the memory free information and task instance free information to obtain the free resource value of the execution node. Statistically analyze the free resource values of each execution node, sort each execution node using the free resource values of each execution node, and use the execution node with the highest free resource value among multiple execution nodes as the target execution node corresponding to the target task.

[0125] According to the embodiments provided by the present application, an execution node is often selected and task scheduling is performed using a scheduling node. Therefore, the present application also provides a master-slave election mechanism, which is mainly designed for the master-slave architecture of the scheduling component in distributed scheduling. The main functions of the scheduling component are to provide a rest task scheduling interface and perform periodic task scheduling using the Quartz timer of the component itself. The idea of the master-slave design of this scheduling component is to deploy and run multiple scheduling components during implementation, one Master and multiple Slaves. The Master provides a rest scheduling service externally and starts its own scheduling capabilities (including timed Quzrtz, etc.), while all capabilities of the Slave are dormant and its scheduling capabilities are stopped. When the Master fails, the Slave automatically runs for election to become the Leader, and the Slave component that becomes the Leader opens the rest service and starts the scheduling capabilities, realizing the high availability of the scheduling component. The Leader election mechanism of the scheduling component is implemented through the LeaderLatch mechanism, mainly by creating a distributed node lock and selecting the master node through an election mechanism.

[0126] After determining the target execution node corresponding to each target task, the visualization processing method further includes:

[0127] i: For each target task, determine whether the task execution time of the target execution node corresponding to the target task is greater than or equal to the execution time threshold.

[0128] It should be noted that the task execution time refers to the time executed by the target execution node corresponding to the target task. The execution time threshold refers to a preset time threshold used to determine whether the target task is abnormal. For example, the execution time threshold can be set to 5 minutes, and the present application does not make specific limitations on this.

[0129] For step i above, in specific implementation, for each target task in the data processing flow, first obtain the task execution time of the target execution node corresponding to the target task, and then determine whether the task execution time is greater than or equal to the execution time threshold. If so, execute step ii.

[0130] ii: If so, determine that the target task has a running exception, and re-determine the corresponding target execution node for the target task based on the load policy.

[0131] For step ii above, in specific implementation, when it is determined that the task execution time of the target execution node corresponding to the target task is greater than or equal to the preset execution time threshold, it is considered that the target task has a running exception, and the corresponding target execution node for the target task is re-determined according to the preset load policy, and the task instance corresponding to the target task is scheduled to the re-selected target execution node for execution. Specifically, the method of re-determining the corresponding target execution node for the target task according to the preset load policy is the same as the method provided in the above embodiments of the present application, and will not be elaborated here. According to the embodiments provided in the present application, the task timeout mechanism is uniformly designed and implemented based on the worker execution component, and is applicable to all task plugins including the DIPE engine. The main function is to specify the timeout duration of the task in the task scheduling JSON configuration. When the task runs beyond the specified timeout duration, it is determined that the task has a running exception, and the task termination operation is automatically performed. This mechanism is mainly to avoid the task running in a false deadlock or exceeding the expected running duration due to configuration or operation problems, and at the same time solve the problem of inconsistent monitoring status caused by the task false deadlock, and avoid misleading the monitoring personnel with abnormal status. This mechanism also avoids the situation where subsequent tasks in a processing job composed of multiple tasks cannot run in time due to the false deadlock or timeout of the previous task. This part of the mechanism is implemented based on the distributed load policy of the scheduling node. When configuring the task, the task exception retry times and the exception retry interval can be specified. These two configuration indicators indicate that after the task is started and an exception occurs within the specified number of times, it will be restarted after the specified exception retry interval time (min minutes), that is, the scheduling node will select the running node according to the distributed load policy and re-schedule the execution node of the task to run.

[0132] According to the embodiments provided by this application, in distributed scheduling of the DIPE engine or partial task plugins, for the same task at different times, the generated task instances are scheduled on different nodes, and some configuration information needs to be shared. For example, in the scenario of CDC incremental synchronization of data, each time a task instance starts, incremental synchronization is performed according to the latest incremental value, and the final incremental value is stored when the task ends for CDC incremental extraction during the next scheduling. In single-node scheduling, the usual solution is to store these configuration values in a configuration file generated by the task ID in the local working directory. In distributed scheduling, configuration sharing is required. In the overall scheduling solution, in order to achieve weak coupling between components, reduce dependencies between components, and reduce the connection pressure on the database, the execution nodes in the overall design do not communicate with the database. Therefore, configuration sharing between distributed nodes is adopted. Configuration nodes are constructed according to distributed nodes, and configuration information is registered and shared according to business needs.

[0133] iii: Send the abnormal operation information of the target task to the pre-set notified user according to the pre-set notification method, so that the notified user can determine the operation status of the target task according to the abnormal operation information.

[0134] It should be noted that the notification method refers to the pre-set method for sending abnormal operation information, such as sending by email, sending by text message, or sending by DingTalk, etc. This application does not make specific limitations on this. The abnormal operation information may include key information such as abnormal time, abnormal task, abnormal information, and abnormal data.

[0135] Regarding the above step iii, in specific implementation, when an abnormal operation occurs for a certain target task, the abnormal operation information of the target task is sent to the pre-set notified user according to the pre-set notification method, so that the notified user can determine the operation status of the target task according to the abnormal operation information. According to the embodiments provided by this application, this part of the mechanism is implemented based on the monitoring of the meta-database of task scheduling. The implementation process mainly involves monitoring the task status change information of the meta-database, combining the notification strategies specified during task configuration (no notification, only failure notification, only success notification, all notifications), notification methods (email, text message, DingTalk, WeChat official account), and configured notified users, etc. When the task status change meets the specified notification strategy, a notification is sent to the notifier according to the specified notification type. This mechanism realizes the full-round monitoring of task operation, timely and effectively notifies task abnormalities, enables monitoring personnel to quickly track and recover abnormal tasks, and more effectively guarantees the continuous and stable operation of tasks.

[0136] The visualization processing method based on metadata management provided by the embodiments of the present application first receives a data processing flow set by a user; then, for each target task in the data processing flow, based on the task type and initial configuration information of the target task, metadata relationship registration is performed on the target task to generate a metadata relationship graph corresponding to the target task in a visualization interface; finally, based on the execution order of each target task in the data processing flow and the metadata relationship graph of each target task, a metadata traceability graph corresponding to the data processing flow is generated in the visualization interface.

[0137] Compared with the methods in the prior art, in the present application, by performing metadata relationship registration on each target task in the data processing flow, data lineage analysis is performed on the entire process to form a metadata traceability graph of the data processing flow, and related target tasks are connected in series to form a traceability topology graph of corresponding data and tasks, so that users can effectively manage tasks according to the traceability topology graph, realizing the visualization effect of data processing. And the data processing flow integrates data tasks at all stages from different task platforms in the entire data governance life cycle, providing unified data task scheduling and monitoring functions, realizing the scheduling and monitoring of data tasks in the entire data governance life cycle, and solving the problems of chaotic and ineffective governance and maintenance of a large number of tasks in the prior art.

[0138] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a visualization processing device provided by the embodiments of the present application. As Figure 3 shown in

[0139] The receiving module 301 is configured to receive a data processing flow set by a user; wherein, the data processing flow includes at least one target task from different task platforms that can implement different functional logics;

[0140] The metadata relationship graph determination module 302 is configured to, for each target task in the data processing flow, perform metadata relationship registration on the target task based on the task type and initial configuration information of the target task to generate a metadata relationship graph corresponding to the target task in a visualization interface; wherein, the initial configuration information is used to represent the processing method of the target task, and the metadata relationship graph is used to represent the data lineage of the target task;

[0141] The metadata traceability graph determination module 303 is configured to generate a metadata traceability graph corresponding to the data processing flow in the visualization interface based on the execution order of each target task in the data processing flow and the metadata relationship graph of each target task.

[0142] Further, when the metadata relationship graph determination module 302 performs metadata relationship registration for the target task based on the task type and initial configuration information of the target task to generate a metadata relationship graph corresponding to the target task in the visualization interface, the metadata relationship graph determination module 302 is further configured to:

[0143] Obtain at least one initial configuration information determined by the user when setting the target task;

[0144] According to the task type of the target task, obtain at least one target configuration information from the at least one initial configuration information; wherein, the target configuration information is used to generate the metadata relationship graph;

[0145] For each target configuration information, determine the connection relationship between the target configuration information and the topological node corresponding to the target task in the metadata relationship graph according to the data attribute to which the target configuration information belongs; wherein, the connection relationship includes a connection line and the arrow direction on the connection line;

[0146] Add the target configuration information and the topological node corresponding to the target task to the visualization interface, and draw according to the connection relationship between the target configuration information and the topological node corresponding to the target task in the metadata relationship graph to generate a metadata relationship graph corresponding to the target task.

[0147] Further, when the metadata traceability graph determination module 303 is used to generate a metadata traceability graph corresponding to the data processing flow in the visualization interface based on the execution order of each target task in the data processing flow and the metadata relationship graph of each target task, the metadata traceability graph determination module 303 is further configured to:

[0148] Based on the data processing flow, determine the execution order of each target task;

[0149] For each target task, based on the execution order of the target task, determine the adjacent tasks that are executed before and / or after the target task from the data processing flow;

[0150] Connect the metadata relationship graph of the target task with the metadata relationship graph of the adjacent task to generate a metadata traceability graph corresponding to the data processing flow.

[0151] Further, the visualization processing device 300 further includes a monitoring module. After generating the metadata traceability graph corresponding to the data processing flow, the monitoring module is configured to:

[0152] Monitor the execution situation of each target task when executing the task instance corresponding to the data processing flow;

[0153] For each target task, based on the execution status of the target task, render the display color corresponding to the execution status at the position of the topological node corresponding to the target task in the metadata traceability graph.

[0154] Further, the visualization processing device 300 further includes an execution node determination module, and the execution node determination module is used for:

[0155] For each target task in the data processing flow, determine the target execution node corresponding to the target task from at least one pre-configured execution node according to a pre-set load policy, so that the target execution node executes the target task.

[0156] Further, when the execution node determination module is used to select the corresponding target execution node for the target task according to the pre-set load policy, the execution node determination module is used for:

[0157] When the load policy is a random load policy, determine the target execution node corresponding to the target task through the following steps:

[0158] Randomly select from at least one execution node to determine the target execution node corresponding to the target task;

[0159] When the load policy is a weighted round-robin load policy, determine the target execution node corresponding to the target task through the following steps:

[0160] For each execution node, determine the available memory information and available load information of the execution node;

[0161] Perform weighted calculation using the available memory information and available load information of the execution node to obtain the available resource value of the execution node;

[0162] Sort each execution node using the available resource value of each execution node, and use the execution node with the highest available resource value among the multiple execution nodes as the target execution node;

[0163] When the load policy is a concurrent load policy, determine the target execution node corresponding to the target task through the following steps:

[0164] For each execution node, determine the number of concurrent task instances of the execution node;

[0165] Sort each execution node using the number of concurrent task instances of each execution node, and use the execution node with the fewest number of concurrent task instances among the multiple execution nodes as the target execution node;

[0166] When the load policy is a resource weighted load policy, the target execution node corresponding to the target task is determined through the following steps:

[0167] For each execution node, determine the memory free information and task instance free information of the execution node;

[0168] Use the memory free information and task instance free information of the execution node for weighted calculation to obtain the free resource value of the execution node;

[0169] Sort each execution node using the free resource value of each execution node, and use the execution node with the highest free resource value among the multiple execution nodes as the target execution node.

[0170] Furthermore, the visualization processing device 300 further includes an exception monitoring module. After determining the target execution node corresponding to each target task, the exception monitoring module is used for:

[0171] For each target task, determine whether the task execution time of the target execution node corresponding to the target task is greater than or equal to the execution time threshold;

[0172] If so, determine that the target task has a running exception, and re-determine the corresponding target execution node for the target task based on the load policy;

[0173] Send the abnormal running information of the target task to the pre-set notified user according to the pre-set notification method, so that the notified user can determine the running status of the target task based on the abnormal running information.

[0174] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 4 shown in

[0175] the electronic device 400 includes a processor 410, a memory 420, and a bus 430. Figure 1 and Figure 2 The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 runs, the processor 410 communicates with the memory 420 through the bus 430. When the machine-readable instructions are executed by the processor 410, the steps of the visualization processing method based on metadata management in the method embodiments as described above

[0176] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it can execute the steps of the visualization processing method based on metadata management in the method embodiment as described above. For the specific implementation manner, reference can be made to the method embodiment, which will not be elaborated here. Figure 1 and Figure 2 shown in the method embodiment. For the specific implementation manner, reference can be made to the method embodiment, which will not be elaborated here.

[0177] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated here.

[0178] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0179] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0180] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0181] When the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0182] It should be noted that: similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.

[0183] Finally, it should be noted that: the above-mentioned embodiments are only specific implementation manners of this application, used to illustrate the technical solutions of this application, rather than limiting it. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by this application can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A visualization processing method based on metadata management, characterized in that The visualization processing method includes: Receiving a data processing flow set by a user; wherein, at least one target task capable of implementing different functional logics from different task platforms is included in the data processing flow; For each target task in the data processing flow, based on the task type and initial configuration information of the target task, perform metadata relationship registration on the target task to generate a metadata relationship diagram corresponding to the target task in the visualization interface; wherein, the initial configuration information is used to represent the processing method of the target task, and the metadata relationship diagram is used to represent the data lineage relationship of the target task; Based on the execution order of each target task in the data processing flow and the metadata relationship diagrams of each target task, generate a metadata traceability diagram corresponding to the data processing flow in the visualization interface; The performing metadata relationship registration on the target task based on the task type and initial configuration information of the target task to generate a metadata relationship diagram corresponding to the target task in the visualization interface includes: Obtaining at least one initial configuration information determined by the user when setting the target task; According to the task type of the target task, obtaining at least one target configuration information from the at least one initial configuration information; wherein, the target configuration information is used to generate the metadata relationship diagram; For each target configuration information, determine the connection relationship between the target configuration information and the topological node corresponding to the target task in the metadata relationship diagram according to the data attribute to which the target configuration information belongs; wherein, the connection relationship includes a connection line and the arrow direction on the connection line; Adding the target configuration information and the topological node corresponding to the target task to the visualization interface, and drawing according to the connection relationship between the target configuration information and the topological node corresponding to the target task in the metadata relationship diagram to generate a metadata relationship diagram corresponding to the target task; The generating a metadata traceability diagram corresponding to the data processing flow in the visualization interface based on the execution order of each target task in the data processing flow and the metadata relationship diagrams of each target task includes: Based on the data processing flow, determining the execution order of each target task; For each target task, based on the execution order of the target task, determining adjacent tasks that are executed before and / or after the target task from the data processing flow; Connecting the metadata relationship diagram of the target task with the metadata relationship diagrams of the adjacent tasks to generate a metadata traceability diagram corresponding to the data processing flow.

2. The visualization processing method according to claim 1, wherein After generating the metadata traceability diagram corresponding to the data processing flow, the visualization processing method further includes: When executing the task instance corresponding to the data processing flow, monitoring the execution situation of each target task; For each target task, rendering the display color corresponding to the execution situation at the position of the topological node corresponding to the target task in the metadata traceability diagram based on the execution situation of the target task.

3. The visualization processing method according to claim 1, wherein The visualization processing method further includes: For each target task in the data processing flow, according to a preset load policy, determine a target execution node corresponding to the target task from at least one pre-configured execution node, so that the target execution node executes the target task.

4. The visualization processing method according to claim 3, wherein Select a corresponding target execution node for the target task according to a preset load policy, including: When the load policy is a random load policy, determine the target execution node corresponding to the target task through the following steps: Randomly select from at least one execution node to determine the target execution node corresponding to the target task; When the load policy is a weighted round-robin load policy, determine the target execution node corresponding to the target task through the following steps: For each execution node, determine the available memory information and available load information of the execution node; Perform weighted calculation using the available memory information and available load information of the execution node to obtain the available resource value of the execution node; Sort each execution node using the available resource value of each execution node, and use the execution node with the highest available resource value among multiple execution nodes as the target execution node; When the load policy is a concurrent load policy, determine the target execution node corresponding to the target task through the following steps: For each execution node, determine the number of concurrent task instances of the execution node; Sort each execution node using the number of concurrent task instances of each execution node, and use the execution node with the fewest number of concurrent task instances among multiple execution nodes as the target execution node; When the load policy is a resource-weighted load policy, determine the target execution node corresponding to the target task through the following steps: For each execution node, determine the memory free information and task instance free information of the execution node; Perform weighted calculation using the memory free information and task instance free information of the execution node to obtain the free resource value of the execution node; Sort each execution node using the free resource value of each execution node, and use the execution node with the highest free resource value among multiple execution nodes as the target execution node.

5. The visualization processing method according to claim 3, wherein After determining the target execution node corresponding to each target task, the visualization processing method further includes: For each target task, determine whether the task execution time of the target execution node corresponding to the target task is greater than or equal to the execution time threshold; If so, determine that the target task has a running exception, and re-determine the corresponding target execution node for the target task based on the load policy; According to a preset notification method, send the abnormal running information of the target task to a preset notified user, so that the notified user can determine the running situation of the target task according to the abnormal running information.

6. A visualization processing device based on metadata management, characterized in that, The visualization processing device includes: A receiving module, configured to receive a data processing flow set by a user; wherein, the data processing flow includes at least one target task from different task platforms that can implement different functional logics; A metadata relationship graph determination module, which is used for each target task in the data processing flow. Based on the task type and initial configuration information of the target task, it performs metadata relationship registration on the target task to generate a metadata relationship graph corresponding to the target task in the visualization interface. Wherein, the initial configuration information is used to characterize the processing method of the target task, and the metadata relationship graph is used to characterize the data lineage relationship of the target task; A metadata traceability graph determination module, which is used to generate a metadata traceability graph corresponding to the data processing flow in the visualization interface based on the execution order of each target task in the data processing flow and the metadata relationship graph of each target task; When the metadata relationship graph determination module is used to perform metadata relationship registration on the target task based on the task type and initial configuration information of the target task to generate a metadata relationship graph corresponding to the target task in the visualization interface, the metadata relationship graph determination module is further used for: Obtain at least one initial configuration information determined by the user when setting the target task; According to the task type of the target task, obtain at least one target configuration information from the at least one initial configuration information. Wherein, the target configuration information is used to generate the metadata relationship graph; For each target configuration information, determine the connection relationship between the target configuration information and the topological node corresponding to the target task in the metadata relationship graph according to the data attribute to which the target configuration information belongs. Wherein, the connection relationship includes a connection line and the arrow direction on the connection line; Add the target configuration information and the topological node corresponding to the target task to the visualization interface, and draw according to the connection relationship between the target configuration information and the topological node corresponding to the target task in the metadata relationship graph to generate a metadata relationship graph corresponding to the target task; When the metadata traceability graph determination module is used to generate a metadata traceability graph corresponding to the data processing flow in the visualization interface based on the execution order of each target task in the data processing flow and the metadata relationship graph of each target task, the metadata traceability graph determination module is further used for: Based on the data processing flow, determine the execution order of each target task; For each target task, based on the execution order of the target task, determine the adjacent tasks that are executed before and / or after the target task from the data processing flow; Connect the metadata relationship graph of the target task with the metadata relationship graph of the adjacent task to generate a metadata traceability graph corresponding to the data processing flow.

7. An electronic device, characterized in that, It includes: A processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are run by the processor, they execute the steps of the visualization processing method based on metadata management as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it executes the steps of the visualization processing method based on metadata management according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Metadata query method and system and electronic equipment

    CN113297139A

  • Systems, methods and apparatus for analysis and visualization of metadata information

    US20090138279A1