NiFi task scheduling method and device, equipment and storage medium

By parsing the directed acyclic graph defined by the target user, the dependencies between NiFi tasks are obtained and scheduled, solving the problem of inter-process dependency scheduling of NiFi tasks in the Apache Airflow system and realizing the flexibility and stability of cross-cluster task scheduling.

CN120950217APending Publication Date: 2025-11-14SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511107316.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In existing technologies, NiFi tasks in the Apache Airflow system cannot meet the scheduling mechanism for inter-process dependencies, resulting in inflexible task scheduling.

Method used

By parsing the target user-defined directed acyclic graph, the dependencies between NiFi tasks are obtained. Using the preset NiFi interface and address information table, the target processor is controlled to perform dependency scheduling. Combined with real-time monitoring and load adjustment, cross-cluster task scheduling is achieved.

Benefits of technology

It enables flexible dependency scheduling between NiFi tasks, improves the flexibility and stability of task scheduling, shortens the task failure recovery time, and ensures the stability of the data synchronization process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950217A_ABST
    Figure CN120950217A_ABST
Patent Text Reader

Abstract

The invention discloses a NiFi task scheduling method and device, equipment and a storage medium, and relates to the technical field of distributed scheduling, and the method comprises the steps: analyzing a directed acyclic graph to obtain task operators corresponding to each to-be-scheduled NiFi task and a dependency relationship between the tasks; querying a preset cluster configuration table and a preset processor address information table according to a cluster ID and a processor ID in each task operator to obtain a first address corresponding to each target cluster and a second address of each target processor; and sending a control instruction to each target processor based on a preset NiFi interface, each first address, each second address and the dependency relationship, so as to control each target processor to sequentially schedule each NiFi task to be scheduled based on the dependency relationship. By analyzing the directed acyclic graph, the dependency relationship among different NiFi tasks can be obtained, so that the NiFi tasks are scheduled according to the dependency relationship among the NiFi tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of distributed scheduling technology, and in particular to a NiFi task scheduling method, apparatus, device, and storage medium. Background Technology

[0002] NiFi is a high-performance ETL (Extract-Transform-Load) tool open-sourced by Apache. It uses a visual data flow graph to extract, transform, and load data, and is widely used in data pipeline construction. Its advantages include low latency, high reliability, and flexible data flow configuration. Apache Airflow is an open-source distributed task scheduling system from Apache that supports task orchestration using a DAG (Directed Acyclic Graph) approach, providing features such as timed scheduling, resource management, and fault tolerance.

[0003] Currently, in the Apache Airflow system, each NiFi process group runs independently, which cannot satisfy the scheduling mechanism for multiple processes with dependencies. Therefore, scheduling dependencies between multiple NiFi tasks has become a technical problem that needs to be solved. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a NiFi task scheduling method, apparatus, device, and storage medium that can obtain the dependencies between different NiFi tasks by parsing a directed acyclic graph, and thus schedule NiFi tasks according to these dependencies. The specific solution is as follows:

[0005] Firstly, this application provides a NiFi task scheduling method applied to a distributed task scheduling system, comprising:

[0006] The target directed acyclic graph defined by the target user through the local user interface is read, and the target directed acyclic graph is parsed to obtain the target task operators corresponding to each target NiFi task to be scheduled and the dependency relationships between each target NiFi task to be scheduled; wherein, the target task operator includes the target cluster ID and the target processor ID corresponding to each target NiFi task to be scheduled.

[0007] The preset cluster configuration table is queried according to each target cluster ID to obtain the first address corresponding to each target cluster, and the preset processor address information table is queried according to each target processor ID to obtain the second address of each target processor in the corresponding target cluster.

[0008] Based on the preset NiFi interface, each of the first addresses, each of the second addresses, and the dependency relationship, control instructions are sent to each of the target processors to control each of the target processors to schedule each of the target NiFi tasks to be scheduled in sequence based on the dependency relationship.

[0009] Optionally, before sending control instructions to each of the target processors based on the preset NiFi interface, each of the first addresses, each of the second addresses, and the dependency relationship, the method further includes:

[0010] Real-time monitoring of each target processor is performed to obtain the load status of each target processor.

[0011] The computing resources corresponding to each target processor are adjusted in real time according to the load status of each target processor, so as to utilize each target processor to schedule and process each target NiFi task to be scheduled.

[0012] Optionally, the target processor schedules each of the target NiFi tasks to be scheduled sequentially based on the dependency relationship, including:

[0013] The task scheduling result corresponding to the previous target NiFi task to be scheduled is obtained using the preset NiFi interface, and the current target NiFi task to be scheduled is determined from each target NiFi task to be scheduled based on the task scheduling result and the dependency relationship.

[0014] The target processor corresponding to the current target NiFi task to be scheduled is used to schedule the current target NiFi task.

[0015] Optionally, the NiFi task scheduling method further includes:

[0016] The target NiFi tasks to be scheduled are monitored in real time to obtain the task status corresponding to each target NiFi task to be scheduled, and a target statistical report is generated based on the task status; wherein, the target statistical report includes the task execution time, resource utilization rate and resource throughput corresponding to each target NiFi task to be scheduled.

[0017] Optionally, after controlling each of the target processors to schedule each of the target NiFi tasks to be scheduled sequentially based on the dependency relationship, the method further includes:

[0018] Determine whether each of the target NiFi tasks to be scheduled has been successfully scheduled. If each of the target NiFi tasks to be scheduled has not been successfully scheduled, then proceed to the step of sending control instructions to each of the target processors based on the preset NiFi interface, each of the first addresses, each of the second addresses and the dependency relationship, so as to reschedule each failed task.

[0019] Optionally, after the target processor schedules each of the target NiFi tasks to be scheduled sequentially based on the dependency relationship, the process further includes:

[0020] Determine whether the number of retries corresponding to the currently scheduled failed task is less than a preset retry threshold. If the number of retries corresponding to the currently scheduled failed task is not less than the preset retry threshold, then roll back each of the target NiFi tasks to be scheduled according to the task type of the currently scheduled failed task.

[0021] Optionally, the task type of the currently scheduled failed task may be used to roll back each of the target NiFi tasks to be scheduled, including:

[0022] Determine whether the currently scheduled failed task is a non-critical task. If the currently scheduled failed task is a non-critical task, obtain a new scheduled failed task, identify the new scheduled failed task as the current scheduled failed task, and jump to the step of determining whether the number of retries corresponding to the current scheduled failed task is less than a preset retry threshold.

[0023] If the currently scheduled failed task is a critical task, then all the target NiFi tasks to be scheduled will be rolled back.

[0024] Secondly, this application provides a NiFi task scheduling device for use in a distributed task scheduling system, comprising:

[0025] The dependency determination module is used to read the target directed acyclic graph defined by the target user through the local user interface, and parse the target directed acyclic graph to obtain the target task operators corresponding to each target NiFi task to be scheduled and the dependency relationships between each target NiFi task to be scheduled; wherein, the target task operator includes the target cluster ID and the target processor ID corresponding to each target NiFi task to be scheduled.

[0026] The address acquisition module is used to query the preset cluster configuration table according to each target cluster ID to obtain the first address corresponding to each target cluster, and to query the preset processor address information table according to each target processor ID to obtain the second address of each target processor in the corresponding target cluster.

[0027] The task scheduling module is used to send control instructions to each of the target processors based on the preset NiFi interface, each of the first addresses, each of the second addresses, and the dependency relationship, so as to control each of the target processors to schedule each of the target NiFi tasks to be scheduled in sequence based on the dependency relationship.

[0028] Thirdly, this application provides an electronic device, comprising:

[0029] Memory, used to store computer programs;

[0030] A processor for executing the computer program to implement the aforementioned NiFi task scheduling method.

[0031] Fourthly, this application provides a computer-readable storage medium for storing a computer program that, when executed by a processor, implements the aforementioned NiFi task scheduling method.

[0032] This application first reads the target directed acyclic graph (DAG) defined by the target user through a local user interface, and parses the DAG to obtain the target task operators corresponding to each target NiFi task to be scheduled and the dependencies between the target NiFi tasks to be scheduled. The target task operators include the target cluster ID and target processor ID corresponding to each target NiFi task to be scheduled. Then, a preset cluster configuration table is queried based on each target cluster ID to obtain the first address corresponding to each target cluster, and a preset processor address information table is queried based on each target processor ID to obtain the second address of each target processor in the corresponding target cluster. Finally, control commands are sent to each target processor based on the preset NiFi interface, each of the first addresses, each of the second addresses, and the dependencies to control each target processor to schedule each target NiFi task to be scheduled sequentially based on the dependencies. Therefore, this application can obtain the dependency relationships between different NiFi tasks by parsing the directed acyclic graph constructed by the target user, and thus schedule NiFi tasks according to the dependency relationships between NiFi tasks; by obtaining the target cluster ID and target processor ID corresponding to the scheduled NiFi tasks from the target task operators, and querying the cluster address and processor address corresponding to the task from the preset cluster configuration table and the preset processor address information table according to the target cluster ID and target processor ID, cross-cluster task scheduling is realized. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0034] Figure 1 This is a flowchart of a NiFi task scheduling method disclosed in this application;

[0035] Figure 2 This is a schematic diagram of the structure of a NiFi task scheduling system disclosed in this application;

[0036] Figure 3 This is a schematic diagram of the structure of a NiFi task scheduling device disclosed in this application;

[0037] Figure 4 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0039] Current NiFi task scheduling processes cannot satisfy the mechanism for scheduling dependencies between multiple processes. To address this, this application provides a NiFi task scheduling method that, by parsing a directed acyclic graph, can obtain the dependencies between different NiFi tasks, and thus schedule NiFi tasks based on these dependencies.

[0040] See Figure 1 As shown, this embodiment of the invention discloses a NiFi task scheduling method, applied to a distributed task scheduling system, comprising:

[0041] Step S11: Read the target directed acyclic graph defined by the target user through the local user interface, and parse the target directed acyclic graph to obtain the target task operators corresponding to each target NiFi task to be scheduled and the dependency relationships between each target NiFi task to be scheduled; wherein, the target task operators include the target cluster ID and target processor ID corresponding to each target NiFi task to be scheduled.

[0042] The system architecture in this embodiment is as follows: Figure 2As shown, it includes four core modules: NiFi task adaptation layer, Airflow extension engine, resource management module, and visualization interaction layer.

[0043] The system's operational flow is as follows:

[0044] (1) NiFi resource initialization: When a user registers a NiFi cluster in Airflow, the system automatically synchronizes the process group and processor list.

[0045] (2) Workflow definition phase: Define the DAG containing NiFiOperator through Airflow UI (User Interface) or code, and configure dependencies, timing strategies and resource requirements.

[0046] (3) Scheduling and execution phase: Airflow Scheduler triggers DAG according to rules, the parser generates execution plan, calls NiFi API (Application Programming Interface), i.e., the preset NiFi interface controls task execution, monitors status in real time and handles exceptions.

[0047] (4) Feedback on execution results: Generate statistical reports (execution time, throughput, resource utilization), and support export or integration into BI (Business Intelligence) systems.

[0048] In this embodiment, the target user extends the task through the system's NiFi task adaptation layer: a new NiFiOperator (i.e., the target task operator) is added in Airflow, inheriting from BaseOperator. The structure of the new NiFiOperator is shown below:

[0049] class NiFiOperator(BaseOperator):

[0050] template_fields = ('nifi_instance_id', 'process_group_id', 'processor_id')

[0051] def __init__(

[0052] self,

[0053] nifi_instance_id: str,

[0054] process_group_id: str,

[0055] processor_id: str,

[0056] auth_config: dict,

[0057] retry_config: dict,

[0058] *args, **kwargs, ):

[0060] super().__init__(*args, **kwargs);

[0061] self.nifi_instance_id = nifi_instance_id;

[0062] self.process_group_id = process_group_id;

[0063] self.processor_id = processor_id;

[0064] self.auth_config = auth_config;

[0065] self.retry_config = retry_config;

[0066] As can be seen from the above data structure, the target task operator includes information such as the target cluster ID (nifi_instance_id) and the target processor ID (processor_id) corresponding to the target NiFi task to be scheduled.

[0067] In addition, the NiFi task adaptation layer is also used for the design of API gateways: encapsulating NiFi REST APIs (i.e., default NiFi interfaces) such as starting / stopping processors, querying status, and obtaining statistical metrics, supporting unified access to multiple NiFi instances, and solving cross-domain and authentication differences issues.

[0068] In this embodiment, the user's configuration of the directed acyclic graph is completed in the system's visualization interaction layer; the visualization interaction layer includes:

[0069] NiFi Task Binding Interface: A NiFi task configuration panel has been added to the Airflow UI (i.e., the target user interaction interface), which supports selecting registered NiFi instances, process groups and processors via drop-down lists;

[0070] Dependency visualization: A DAG editor based on AntV G6 is implemented, which supports dragging and dropping NiFi task nodes, connecting dependencies with arrows, and checking circular dependencies in real time and prompting errors.

[0071] In this embodiment, the target directed acyclic graph is a directed acyclic graph predefined by the target user according to the task scheduling requirements. The graph consists of several target task operators, and the connection order between the task operators represents the execution dependency relationship between the tasks. After the target user has completed the directed acyclic graph, the system can obtain the dependency relationship between the tasks by parsing the directed acyclic graph, thereby realizing the dependency scheduling of tasks.

[0072] The code for the target user's task orchestration using the workflow orchestration engine in the above process is shown below: DAG construction algorithm: Adjacency list is used to store task dependencies, supporting circular dependency detection, such as DFS (Depth First Search) and topological sorting, to generate parallel execution plans;

[0073] Conditional branch handling:

[0074] Using Airflow's BranchPythonOperator or by defining conditional dependencies directly in the DAG, branch selection is based on the NiFi execution result:

[0075] def decide_branch(**context):

[0076] task_status = context['task_instance'].xcom_pull(task_ids='nifi_task');

[0077] if task_status == 'success':

[0078] return 'task_b';

[0079] else:

[0080] return 'task_c';

[0081] with DAG(dag_id='nifi_workflow') as dag:

[0082] nifi_task = NiFiOperator(task_id='nifi_task', ...);

[0083] task_b = PythonOperator(task_id='task_b', ...);

[0084] task_c = PythonOperator(task_id='task_c', ...);

[0085] branch = BranchPythonOperator(task_id='branch', python_callable=decide_branch);

[0086] nifi_task >> branch >> [task_b, task_c];

[0087] By orchestrating tasks to construct a directed acyclic graph, NiFi tasks can be scheduled based on dependencies, thus improving the flexibility of task scheduling.

[0088] Step S12: Query the preset cluster configuration table according to each target cluster ID to obtain the first address corresponding to each target cluster, and query the preset processor address information table according to each target processor ID to obtain the second address of each target processor in the corresponding target cluster.

[0089] In this embodiment, after obtaining the target cluster ID and target processor ID corresponding to the NiFi task to be scheduled from the target task operator, it is also necessary to look up the target cluster URL (Uniform Resource Locator), i.e., the first address, based on the target cluster ID. Then, the preset processor address information table is queried through the target processor ID to obtain the second address of each target processor in the corresponding target cluster.

[0090] The aforementioned preset cluster configuration table and preset processor address information table are tables that the user pre-configures during the task configuration phase. Specifically:

[0091] (1) NiFi resource registration:

[0092] Users configure NiFi cluster information in the Airflow management backend, which is stored in the nifi_cluster_config table (i.e., the default cluster configuration table):

[0093] CREATE TABLE nifi_cluster_config (

[0094] id BIGINT PRIMARY KEY AUTO_INCREMENT,

[0095] cluster_name VARCHAR(100) UNIQUE,

[0096] nifi_url VARCHAR(200) NOT NULL,

[0097] auth_type VARCHAR(50) DEFAULT 'API_KEY',

[0098] auth_params TEXT,

[0099] create_time TIMESTAMP DEFAULT CURRENT_TIMESTAMP, );

[0101] The system periodically synchronizes the NiFi process group list and caches it in the nifi_process_group table, establishing a three-level resource directory of "cluster - process group - processor".

[0102] (2) Task node binding:

[0103] During the workflow definition phase, after the user selects the NiFi task type, they specify the target processor through the cascading selector. The system automatically generates the binding relationship and stores it in the workflow_nifi_binding table (i.e., the preset processor address information table):

[0104] CREATE TABLE workflow_nifi_binding (

[0105] task_instance_id BIGINT,

[0106] nifi_cluster_id BIGINT,

[0107] process_group_id VARCHAR(50),

[0108] processor_id VARCHAR(50),

[0109] FOREIGN KEY (task_instance_id) REFERENCES task_instance(id),

[0110] FOREIGN KEY (nifi_cluster_id) REFERENCES nifi_cluster_config(id), );

[0112] The orderly scheduling of tasks can be ensured by querying the preset cluster configuration table and preset processor address information table based on the task ID and cluster ID.

[0113] Step S13: Based on the preset NiFi interface, each of the first addresses, each of the second addresses, and the dependency relationship, send control instructions to each of the target processors to control each of the target processors to schedule each of the target NiFi tasks to be scheduled in sequence based on the dependency relationship.

[0114] In this embodiment, the operation of verifying the dependency logic between tasks is handled by the Airflow extension engine. The Airflow extension engine has enhanced workflow parser functionality: it adds NiFi task dependency verification logic during the Airflow DAG parsing stage to ensure that: when the preceding NiFi task is in a "successful" state, the subsequent task is triggered; and NiFi processor resource conflict detection for parallel branch tasks (such as processor concurrency limits within the same cluster).

[0115] In addition, the Airflow extension engine is also responsible for the transformation of the scheduler executor: when a TaskInstance is created, a unique NiFiExecutionId is generated. The NiFiExecutionId is used to associate the Airflow task instance with the NiFi process execution log to achieve bidirectional status tracking.

[0116] In this embodiment, the NiFi task lifecycle control includes:

[0117] (1) Startup: Call the NiFi API to set the processor status to ENABLED.

[0118] (2) Stop: Send a STOPPED request and poll for a change in status.

[0119] (3) Status monitoring: Update the Airflow task status via NiFi WebSocket or timed polling (default 10s).

[0120] Two-way log association: Add an Airflow task instance ID tag to the NiFi log, and Airflow stores NiFi error information and displays it in the UI.

[0121] In this embodiment, before sending control instructions to each target processor based on the preset NiFi interface, each first address, each second address, and the dependency relationship, the method further includes: real-time monitoring of each target processor to obtain the load status corresponding to each target processor; and real-time adjustment of the computing resources corresponding to each target processor according to the load status corresponding to each target processor, so as to utilize each target processor to schedule and process each target NiFi task to be scheduled.

[0122] Specifically, this embodiment pre-configures a dynamic resource allocation algorithm: The dynamic resource allocation algorithm dynamically adjusts resource allocation based on the NiFi processor load (such as queue data volume and CPU utilization) and the Airflow resource pool mechanism.

[0123] Logic code:

[0124] def adjust_resources(processor_id: str, load: float):

[0125] if load > threshold_high:

[0126] airflow_api.request_pool_resources(processor_id, cpu=1,memory=512);

[0127] elif load < threshold_low:

[0128] airflow_api.release_pool_resources(processor_id, cpu=1,memory=256);

[0129] Cross-cluster load balancing: The list of NiFi instances is maintained through ZooKeeper (a distributed coordination service), and tasks are allocated using a consistent hashing algorithm to avoid single-point overload.

[0130] In this embodiment, the process of controlling each target processor to schedule each target NiFi task sequentially based on the dependency relationship can specifically include: obtaining the task scheduling result corresponding to the previous target NiFi task using a preset NiFi interface; determining the current target NiFi task from each target NiFi task according to the task scheduling result and the dependency relationship; and scheduling the current target NiFi task using the target processor corresponding to the current target NiFi task.

[0131] Furthermore, this embodiment can also perform real-time monitoring of each target NiFi task to be scheduled, to obtain the task status corresponding to each target NiFi task, and generate a target statistical report based on the task status. The target statistical report includes the task execution time, resource utilization, and resource throughput corresponding to each target NiFi task to be scheduled. The above process involves real-time monitoring of status and handling of anomalies, generating statistical reports (execution time, throughput, resource utilization), which can be exported or integrated into a BI system.

[0132] In this embodiment, after controlling each target processor to schedule each target NiFi task to be scheduled in sequence based on the dependency relationship, the method further includes: determining whether each target NiFi task to be scheduled has been successfully scheduled; if each target NiFi task to be scheduled has not been successfully scheduled, the method jumps to the step of sending control instructions to each target processor based on the preset NiFi interface, each first address, each second address and the dependency relationship, so as to reschedule each failed task.

[0133] In addition, in this embodiment, after controlling each target processor to schedule each target NiFi task to be scheduled in sequence based on the dependency relationship, the method further includes: determining whether the number of retries corresponding to the currently scheduled failed task is less than a preset retry number threshold. If the number of retries corresponding to the currently scheduled failed task is not less than the preset retry number threshold, then the target NiFi tasks to be scheduled are rolled back according to the task type of the currently scheduled failed task.

[0134] The process of rolling back each target NiFi task to be scheduled based on the task type of the currently scheduled failed task can specifically include: determining whether the currently scheduled failed task is a non-critical task; if the currently scheduled failed task is a non-critical task, obtaining a new scheduled failed task, identifying the new scheduled failed task as the current scheduled failed task, and jumping to the step of determining whether the number of retries corresponding to the current scheduled failed task is less than a preset retry threshold; if the currently scheduled failed task is a critical task, rolling back each target NiFi task to be scheduled.

[0135] The above process is achieved through the system's fault tolerance and alarm mechanisms:

[0136] (1) Multi-level fault tolerance strategy:

[0137] Retry mechanism: Configure retries and retry_delay to support exponential backoff.

[0138] Skip strategy: When a task fails and fail_strategy='skip' is configured, mark it as "skip" and continue executing the subsequent task.

[0139] Global rollback: When a critical task fails, all tasks are stopped and rolled back via Airflow's TaskInstanceCallback.

[0140] (2) Intelligent alarm system:

[0141] It integrates alarm channels such as email, and the triggering conditions include continuous task failures, queue backlog exceeding the threshold, and excessive resource utilization.

[0142] It should be noted that the task scheduling in this embodiment is cross-cluster task scheduling:

[0143] (1) Cross-cluster task routing: Use Airflow's Executor (such as Kubernetes Executor) to distribute tasks to different NiFi clusters, supporting sequential cross-cluster and parallel cross-cluster modes.

[0144] (2) Data consistency guarantee: Add a data verification node between cross-cluster tasks. After the data is transmitted through the NiFi processor, call the MD5 (Message-Digest Algorithm 5, a cryptographic hash function) verification tool to ensure integrity. If it fails, a rollback will be triggered.

[0145] By employing multi-level fault tolerance strategies and real-time alarm mechanisms, the task failure recovery time is reduced by 50%, ensuring the stability of the data synchronization process.

[0146] Therefore, this application can obtain the dependency relationships between different NiFi tasks by parsing the directed acyclic graph constructed by the target user, and thus schedule NiFi tasks according to the dependency relationships between NiFi tasks; by obtaining the target cluster ID and target processor ID corresponding to the scheduled NiFi tasks from the target task operators, and querying the cluster address and processor address corresponding to the task from the preset cluster configuration table and the preset processor address information table according to the target cluster ID and target processor ID, cross-cluster task scheduling is realized.

[0147] See Figure 3 As shown, this embodiment of the invention discloses a NiFi task scheduling device applied to a distributed task scheduling system, comprising:

[0148] The dependency determination module 11 is used to read the target directed acyclic graph defined by the target user through the local user interaction interface, and parse the target directed acyclic graph to obtain the target task operators corresponding to each target NiFi task to be scheduled and the dependency relationships between each target NiFi task to be scheduled; wherein, the target task operator includes the target cluster ID and the target processor ID corresponding to each target NiFi task to be scheduled.

[0149] Address acquisition module 12 is used to query a preset cluster configuration table based on each target cluster ID to obtain the first address corresponding to each target cluster, and to query a preset processor address information table based on each target processor ID to obtain the second address of each target processor in the corresponding target cluster.

[0150] The task scheduling module 13 is used to send control instructions to each of the target processors based on the preset NiFi interface, each of the first addresses, each of the second addresses and the dependency relationship, so as to control each of the target processors to schedule each of the target NiFi tasks to be scheduled in sequence based on the dependency relationship.

[0151] Therefore, this application can obtain the dependency relationships between different NiFi tasks by parsing the directed acyclic graph constructed by the target user, and thus schedule NiFi tasks according to the dependency relationships between NiFi tasks; by obtaining the target cluster ID and target processor ID corresponding to the scheduled NiFi tasks from the target task operators, and querying the cluster address and processor address corresponding to the task from the preset cluster configuration table and the preset processor address information table according to the target cluster ID and target processor ID, cross-cluster task scheduling is realized.

[0152] In some specific embodiments, the task scheduling module 13 further includes:

[0153] The processor monitoring unit is used to monitor each of the target processors in real time to obtain the load status of each target processor.

[0154] The first task scheduling unit is used to adjust the computing resources corresponding to each target processor in real time according to the load status of each target processor, so as to use each target processor to schedule and process each target NiFi task to be scheduled.

[0155] In some specific embodiments, the task scheduling module 13 may specifically include:

[0156] The task determination unit is used to obtain the task scheduling result corresponding to the previous target NiFi task to be scheduled using the preset NiFi interface, and determine the current target NiFi task to be scheduled from each target NiFi task to be scheduled according to the task scheduling result and the dependency relationship.

[0157] The second task scheduling unit is used to schedule the current target NiFi task using the target processor corresponding to the current target NiFi task.

[0158] In some specific embodiments, the NiFi task scheduling device further includes:

[0159] The report generation module is used to monitor each of the target NiFi tasks to be scheduled in real time, so as to obtain the task status corresponding to each target NiFi task to be scheduled, and generate a target statistical report based on the task status; wherein, the target statistical report includes the task execution time, resource utilization rate and resource throughput corresponding to each target NiFi task to be scheduled.

[0160] In some specific embodiments, the task scheduling module 13 further includes:

[0161] The first step jump unit is used to determine whether each of the target NiFi tasks to be scheduled has been successfully scheduled. If each of the target NiFi tasks to be scheduled has not been successfully scheduled, the jump is made to the step of sending control instructions to each of the target processors based on the preset NiFi interface, each of the first addresses, each of the second addresses and the dependency relationship, so as to reschedule each failed task.

[0162] In some specific embodiments, the task scheduling module 13 further includes:

[0163] The task rollback submodule is used to determine whether the number of retries corresponding to the currently scheduled failed task is less than a preset retry threshold. If the number of retries corresponding to the currently scheduled failed task is not less than the preset retry threshold, then each of the target scheduled NiFi tasks is rolled back according to the task type of the currently scheduled failed task.

[0164] In some specific embodiments, the task rollback submodule may specifically include:

[0165] The second step jump unit is used to determine whether the currently scheduled failed task is a non-critical task. If the currently scheduled failed task is a non-critical task, a new scheduled failed task is obtained, the new scheduled failed task is identified as the current scheduled failed task, and the process jumps to the step of determining whether the number of retries corresponding to the current scheduled failed task is less than a preset retry threshold.

[0166] The task rollback unit is used to roll back each of the target scheduled NiFi tasks if the currently scheduled failed task is a critical task.

[0167] Furthermore, embodiments of this application also disclose an electronic device, Figure 4 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0168] Figure 4This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the NiFi task scheduling method disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0169] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0170] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0171] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the NiFi task scheduling method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.

[0172] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned NiFi task scheduling method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0173] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0174] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0175] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0176] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0177] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A NiFi task scheduling method, characterized in that, Applications in distributed task scheduling systems include: The target directed acyclic graph defined by the target user through the local user interface is read, and the target directed acyclic graph is parsed to obtain the target task operators corresponding to each target NiFi task to be scheduled and the dependency relationships between each target NiFi task to be scheduled; wherein, the target task operator includes the target cluster ID and the target processor ID corresponding to each target NiFi task to be scheduled. The preset cluster configuration table is queried according to each target cluster ID to obtain the first address corresponding to each target cluster, and the preset processor address information table is queried according to each target processor ID to obtain the second address of each target processor in the corresponding target cluster. Based on the preset NiFi interface, each of the first addresses, each of the second addresses, and the dependency relationship, control instructions are sent to each of the target processors to control each of the target processors to schedule each of the target NiFi tasks to be scheduled in sequence based on the dependency relationship.

2. The NiFi task scheduling method according to claim 1, characterized in that, Before sending control instructions to each of the target processors based on the preset NiFi interface, each of the first addresses, each of the second addresses, and the dependency relationship, the method further includes: Real-time monitoring of each target processor is performed to obtain the load status of each target processor. The computing resources corresponding to each target processor are adjusted in real time according to the load status of each target processor, so as to utilize each target processor to schedule and process each target NiFi task to be scheduled.

3. The NiFi task scheduling method according to claim 1, characterized in that, The control of each target processor to schedule each target NiFi task to be scheduled sequentially based on the dependency relationship includes: The task scheduling result corresponding to the previous target NiFi task to be scheduled is obtained using the preset NiFi interface, and the current target NiFi task to be scheduled is determined from each target NiFi task to be scheduled based on the task scheduling result and the dependency relationship. The target processor corresponding to the current target NiFi task to be scheduled is used to schedule the current target NiFi task.

4. The NiFi task scheduling method according to claim 1, characterized in that, Also includes: The target NiFi tasks to be scheduled are monitored in real time to obtain the task status corresponding to each target NiFi task to be scheduled, and a target statistical report is generated based on the task status; wherein, the target statistical report includes the task execution time, resource utilization rate and resource throughput corresponding to each target NiFi task to be scheduled.

5. The NiFi task scheduling method according to any one of claims 1 to 4, characterized in that, After the control of each target processor to schedule each target NiFi task to be scheduled is based on the dependency relationship, the method further includes: Determine whether each of the target NiFi tasks to be scheduled has been successfully scheduled. If each of the target NiFi tasks to be scheduled has not been successfully scheduled, then proceed to the step of sending control instructions to each of the target processors based on the preset NiFi interface, each of the first addresses, each of the second addresses and the dependency relationship, so as to reschedule each failed task.

6. The NiFi task scheduling method according to claim 5, characterized in that, After the control of each target processor to schedule each target NiFi task to be scheduled is based on the dependency relationship, the method further includes: Determine whether the number of retries corresponding to the currently scheduled failed task is less than a preset retry threshold. If the number of retries corresponding to the currently scheduled failed task is not less than the preset retry threshold, then roll back each of the target NiFi tasks to be scheduled according to the task type of the currently scheduled failed task.

7. The NiFi task scheduling method according to claim 6, characterized in that, The step of rolling back each of the target NiFi tasks to be scheduled based on the task type of the currently scheduled failed task includes: Determine whether the currently scheduled failed task is a non-critical task. If the currently scheduled failed task is a non-critical task, obtain a new scheduled failed task, identify the new scheduled failed task as the current scheduled failed task, and jump to the step of determining whether the number of retries corresponding to the current scheduled failed task is less than a preset retry threshold. If the currently scheduled failed task is a critical task, then all the target NiFi tasks to be scheduled will be rolled back.

8. A NiFi task scheduling device, characterized in that, Applications in distributed task scheduling systems include: The dependency determination module is used to read the target directed acyclic graph defined by the target user through the local user interface, and parse the target directed acyclic graph to obtain the target task operators corresponding to each target NiFi task to be scheduled and the dependency relationships between each target NiFi task to be scheduled; wherein, the target task operator includes the target cluster ID and the target processor ID corresponding to each target NiFi task to be scheduled. The address acquisition module is used to query the preset cluster configuration table according to each target cluster ID to obtain the first address corresponding to each target cluster, and to query the preset processor address information table according to each target processor ID to obtain the second address of each target processor in the corresponding target cluster. The task scheduling module is used to send control instructions to each of the target processors based on the preset NiFi interface, each of the first addresses, each of the second addresses, and the dependency relationship, so as to control each of the target processors to schedule each of the target NiFi tasks to be scheduled in sequence based on the dependency relationship.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the NiFi task scheduling method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, implements the NiFi task scheduling method as described in any one of claims 1 to 7.