Data processing method and computing device

By acquiring and storing the NiFi processor's scheduled operation data, monitoring the processor's operation status in real time and triggering subsequent job processes, the problem that the NiFi cluster cannot perceive the processor's operation status is solved, and the job processing efficiency and reliability of data flow management are improved.

CN120066698APending Publication Date: 2025-05-30HENAN QINWEI DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411958599.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The NiFi cluster cannot sense the processor's operation, resulting in users being unable to monitor the processor's status or performance, affecting the processing efficiency of the job process.

Method used

Provides a data processing method, by acquiring and storing operation data of the NiFi processor, monitoring the operation status of the processor in real time, and triggering subsequent operation processes based on the operation status.

Benefits of technology

Real-time monitoring of NiFi processors is realized, job processing efficiency is improved, and data flow management is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066698A_ABST
    Figure CN120066698A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method, which is applied to a NiFi cluster, the NiFi cluster comprises a plurality of NiFi nodes, the method is executed by any one or more NiFi nodes in the plurality of NiFi nodes, and the method comprises the following steps: obtaining a scheduling task; wherein the scheduling task indicates that the NiFi node schedules a NiFi processor to process a data stream; scheduling the NiFi processor to execute the scheduling task; in the process that the NiFi processor executes the scheduling task, scheduling operation data of the NiFi processor is acquired; and storing the scheduling operation data of the NiFi processor. Thus, in the process of executing the scheduling task through the NiFi processor, the scheduling operation data of the NiFi processor is obtained to monitor the operation condition of the NiFi processor in the process of executing the scheduling task in real time, so that a subsequent operation process is triggered according to the operation condition of the NiFi processor, and the operation processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of servers, and particularly to a data processing method and a computing device. Background Art

[0002] NiFi is a powerful data integration and data flow management tool, specifically a big data framework for automating data flow processing between systems, aiming to automate operations such as data collection, transformation, processing, and transmission. The NiFi Processor in the NiFi cluster is the basic unit for performing actual data processing tasks. The computing nodes in the NiFi cluster execute specific tasks through NiFi Processors and manage the data flow through scheduling rules. For example, after the NiFi cluster starts running, it performs task scheduling according to the scheduling type selected by the user, and the scheduling type determines how the tasks are executed and scheduled.

[0003] Currently, the NiFi cluster cannot sense the running status of the processors, resulting in users being unable to monitor the status or performance of the processors, etc., which is not conducive to the processing of subsequent job flows and there is a situation of low job processing efficiency. Therefore, how to monitor the running status of the processors in the NiFi cluster is one of the problems that need to be solved urgently. Summary of the Invention

[0004] Embodiments of this application provide a data processing method and a computing device, which can monitor the running status of NiFi Processors in the NiFi cluster, thereby facilitating triggering subsequent job flows according to the running status of the NiFi Processors and improving job processing efficiency.

[0005] For this purpose, the embodiments of this application provide the following technical solutions:

[0006] In a first aspect, an embodiment of this application provides a data processing method, which is applied to a NiFi cluster. The NiFi cluster includes multiple NiFi nodes, and the method is executed by any one or more of the multiple NiFi nodes, and includes: obtaining a scheduling task; wherein the scheduling task indicates that the NiFi node schedules a NiFi Processor to process a data flow; scheduling the NiFi Processor to execute the scheduling task; during the process of the NiFi Processor executing the scheduling task, obtaining the scheduling operation data of the NiFi Processor; and storing the scheduling operation data of the NiFi Processor.

[0007] In this way, during the execution of the scheduling task by the NiFi processor, the scheduling operation data of the NiFi processor is obtained to monitor the running status of the NiFi processor in real time when executing the scheduling task, which is conducive to triggering subsequent job processes according to the running status of the NiFi processor and improving job processing efficiency. In addition, by obtaining the scheduling tasks of the NiFi processor, the execution cycle and trigger conditions of each task can be accurately tracked, the running data of the tasks (such as execution time, data volume, etc.) can be recorded, and these data can be stored. In this way, not only can chain analysis be realized to help enterprises obtain business insights and conduct fault diagnosis in a timely manner, but also subsequent processing tasks can be automatically triggered according to the scheduling data in the integrated job to ensure the continuity and efficiency of the workflow. The advantage of this method is that it improves the transparency and controllability of the scheduling task, thereby enhancing the reliability and real-time performance of data flow management.

[0008] As an implementable embodiment, scheduling the NiFi processor to execute the scheduling task includes: starting the NiFi processor and obtaining the scheduling status of the NiFi processor; when the scheduling status of the NiFi processor is STOPED, modifying the scheduling status of the NiFi processor to STARTING and activating the current thread task of the NiFi processor; and processing the data stream according to the scheduling task configured by the NiFi processor.

[0009] As an implementable embodiment, after processing the data stream according to the scheduling task configured by the NiFi processor, it further includes: removing the current thread task of the NiFi processor; and when the scheduling status of the NiFi processor is STARING, using the callback method of the scheduling task to schedule the NiFi processor to process the data stream again.

[0010] As an implementable embodiment, the scheduling operation data includes the start time, end time, data volume of processed data, and execution result of the NiFi processor executing the data processing task. By accurately recording and displaying the start time, end time, processed data volume, and execution result of the NiFi processor executing the task, it provides powerful data flow monitoring and analysis capabilities for data engineers and operation and maintenance personnel. Its advantages include but are not limited to: real-time monitoring of task execution status, improving operation and maintenance efficiency; performance analysis and bottleneck troubleshooting to ensure the efficient operation of the data stream; intelligent alarm and anomaly detection to reduce manual intervention; task scheduling optimization to improve resource utilization and system performance. Generally speaking, this solution can help enterprises optimize the running efficiency of the NiFi data stream, improve job processing efficiency, and provide strong data support for business decision-making.

[0011] As an implementable embodiment, obtaining the scheduling operation data of the NiFi processor for performing data processing tasks includes: obtaining the timestamps when the NiFi processor starts or ends executing a scheduling task, and obtaining the start time or end time of the NiFi processor for performing data processing tasks.

[0012] As an implementable embodiment, obtaining the scheduling operation data of the NiFi processor for performing data processing tasks includes: obtaining the data set of the data types currently scheduled and processed when the NiFi processor executes the scheduling task, and recording the number of processed items in this scheduling based on the data set to record the data volume of the processed data.

[0013] As an implementable embodiment, obtaining the scheduling operation data of the NiFi processor for performing data processing tasks includes: obtaining the execution result of the NiFi processor's scheduling operation according to the successful, failed, timed out or abnormal execution status returned after the NiFi processor executes the scheduling task.

[0014] As an implementable embodiment, after obtaining the scheduling operation data of the NiFi processor, the method further includes: according to the scheduling operation data, performing subsequent data processing tasks on the data stream of the same scheduling task, or removing the current scheduling task and then performing a scheduling task under another scheduling policy.

[0015] As an implementable embodiment, performing subsequent data processing tasks on the data stream of the same scheduling task according to the scheduling operation data includes: when the execution result of the NiFi processor for processing the data stream is successful, transmitting or copying the data stream to a terminal device; and / or when the execution result of the NiFi processor for processing the data stream is failed, pausing the current scheduling task of the NiFi processor; and / or when the execution result of the NiFi processor for processing the data stream is failed, sending a fault warning prompt.

[0016] As an implementable embodiment, the method further includes: analyzing the data processing result of the NiFi processor according to the scheduling operation data.

[0017] In this embodiment, analyzing the "scheduling operation data" is a key step to further enhance the data processing chain and improve the transparency and reliability of data stream management. Specifically, analyzing the data processing result of the NiFi processor according to the scheduling operation data received from the memory can provide support for subsequent decisions and help detect potential problems in a timely manner. Its main advantages are reflected in performance optimization, fast fault response, obtaining business insights, and automatically triggering subsequent job processes, thereby effectively improving the enterprise's data stream management ability and business execution efficiency.

[0018] As an implementable embodiment, analyzing the data processing result of the NiFi processor includes: real-time displaying the data processing result scheduled by the NiFi processor in the form of a report; and / or real-time analyzing the data processing situation scheduled by the NiFi processor in the form of a statistical chart; and / or real-time analyzing the data processing situation scheduled by the NiFi processor in the form of a mathematical model.

[0019] In this embodiment, by means of various methods such as reports, statistical charts, and mathematical models, the scheduling data of the NiFi processor is displayed and analyzed in real time, which not only improves the monitoring and management efficiency of the data stream, but also enhances the intelligent level of the system. Its advantages are reflected in multiple aspects such as real-time performance, visualization, performance optimization, anomaly detection, and decision support, which can help enterprises better cope with the challenges in large-scale data stream processing, optimize system performance, and improve business efficiency.

[0020] In a second aspect, an embodiment of the present application further provides a data processing device, including an acquisition module and a processing module. Among them, the acquisition module is used to acquire a scheduling task; wherein, the scheduling task indicates that the NiFi node schedules the NiFi processor to process the data stream. The processing module is used to schedule the NiFi processor to execute the scheduling task; during the process of the NiFi processor executing the scheduling task, acquire the scheduling operation data of the NiFi processor; and store the scheduling operation data of the NiFi processor.

[0021] In a third aspect, an embodiment of the present application further provides a computing device, including at least one memory for storing a program; at least one processor for executing the program stored in the memory; wherein, the memory is coupled to the processor, and when the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation manner of the first aspect.

[0022] In a fourth aspect, an embodiment of the present application further provides a computing device cluster, including at least one computing device, each computing device includes a processor and a memory; the processor of at least one computing device is used to execute the instructions stored in the memory of at least one computing device, so that the computing device cluster executes the method described in the first aspect or any possible implementation manner of the first aspect.

[0023] In a fifth aspect, an embodiment of the present application further provides a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and when the computer instructions in the computer-readable storage medium are executed by a computing device, the computing device is caused to execute the method involved in the first aspect and its possible implementation manners.

[0024] In a sixth aspect, an embodiment of the present application further provides a computer program product, which includes computer instructions. When the computer instructions are executed by a computing device, the computing device is caused to execute the method described in the first aspect and its possible implementation manners.

[0025] It can be understood that for the beneficial effects of the above second aspect to the sixth aspect, reference can be made to the relevant descriptions of the first aspect above, and details are not repeated here.

[0026] By using one or more of the methods, devices, computing devices, computing device clusters, storage media, and computer program products in the above aspects, it is possible to obtain the start time, end date, amount of data processed, and result of data processing for each scheduling of the processor, and store the above scheduling operation data in a database, which can monitor the operation status of NiFi processors in the NiFi cluster, thereby facilitating triggering subsequent job processes according to the operation status of NiFi processors and improving job processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a schematic hardware architecture diagram for deploying a data processing system provided in an embodiment of the present application;

[0028] Figure 2 It is a schematic software architecture diagram of a data processing system provided in an embodiment of the present application;

[0029] Figure 3 It shows possible ways to perform subsequent analysis on scheduling operation data in a repository provided in an embodiment of the present application;

[0030] Figure 4 It shows a schematic flowchart of a NiFi-based data processing method provided in an embodiment of the present application;

[0031] Figure 5 It shows a schematic flowchart of the operation process of a NiFi processor provided in an embodiment of the present application;

[0032] Figure 6 It shows a schematic flowchart of another data processing method provided in an embodiment of the present application;

[0033] Figure 7 It shows a schematic structural diagram of a data processing device provided in an embodiment of the present application;

[0034] Figure 8 It is a schematic structural diagram of a computing device provided in an embodiment of the present application;

[0035] Figure 9 It is a schematic architecture diagram of a computing device cluster provided in an embodiment of the present application;

[0036] Figure 10 This is a schematic diagram of another architecture of a computing device cluster provided in the embodiments of the present application. Detailed implementation manners

[0037] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.

[0038] In the description of the embodiments of the present application, words such as "exemplary", "for example", or "for instance" are used to mean as an example, illustration, or explanation. Any embodiment or design solution described as "exemplary", "for example", or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for instance" is intended to present relevant concepts in a specific manner.

[0039] In the description of the embodiments of the present application, the term "and / or" only describes the association relationship of associated objects and means that three relationships may exist. For example, A and / or B may represent: A exists alone, B exists alone, and A and B exist simultaneously. In addition, unless otherwise specified, the meaning of the term "plural" refers to two or more. For example, multiple systems refer to two or more systems, and multiple terminals refer to two or more terminals.

[0040] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the technical features indicated. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise particularly emphasized in other ways.

[0041] In the description of the embodiments of the present application, when referring to "some embodiments", it describes a subset of all possible embodiments. However, it can be understood that "some embodiments" may be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.

[0042] In the description of the embodiments of the present application, the terms "first / second / third, etc." or module A, module B, module C, etc. are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that, where permitted, the specific order or sequence can be interchanged so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0043] In the description of the embodiments of the present application, the reference numerals representing steps, such as S101, S102... etc., do not necessarily indicate that these steps will be executed in this order. Under permitted circumstances, the order of the front and back steps can be interchanged, or they can be executed simultaneously.

[0044] Relevant terms involved in the embodiments of the present application:

[0045] OS: Operating System (OS), a set of interrelated system software programs that manage and control computer operations, utilize and run hardware and software resources, and provide public services to organize user interactions. According to the operating environment, operating systems can be divided into desktop operating systems, mobile operating systems, server operating systems, embedded operating systems, etc.

[0046] Cluster: Refers to a group of collaborating servers (or computing nodes) that are connected through a network to form a unified system to work together. Although these servers are independent physical machines, they collaborate through cluster management software or protocols to provide a combined service or resource to the outside world. Generally, for users and application programs, these servers appear to be a single system. Therefore, a cluster can be used as a single system to provide services, process tasks, or store data. The purpose of a cluster is to improve the reliability, availability, and performance of the system by increasing redundancy and scalability. Clusters are widely used in computing and storage tasks that require high performance and high reliability. Through the cluster architecture, the system can maintain stable and efficient operation in the face of failures, increased load, or a sharp increase in data volume.

[0047] A data processing system refers to a software system or a collection of tools used to process massive amounts of data. A data processing system can effectively store, process, analyze, and manage large-scale structured, semi-structured, and unstructured data to extract valuable information and insights. Common data processing systems can include Apache NiFi, Apache Hadoop, Apache Spark, Apache Flink, etc., which provide core components such as distributed storage, distributed computing frameworks, and data processing engines to meet the processing and analysis requirements of large-scale data. A data processing system can be deployed on a large-scale server cluster to complete data processing tasks and provide high-performance, high-reliability, and high-scalability data processing capabilities.

[0048] Technical terms of NiFi:

[0049] NiFi Processor: A component. Each processing unit in NiFi is called a processor, also known as a microprocessor. As an independent processing unit, it is the basic module that makes up the NiFi data stream and mainly completes operations such as creating and retrieving flow files, reading / writing the content of flow files, reading / changing the attributes of flow files, and routing / distributing flow files.

[0050] NiFi Processor Running Modes: Master Node Running Mode and Cluster Running Mode. The Master Node Running Mode means that the NiFi processor runs on one node of the NiFi cluster. The Cluster Running Mode means that the NiFi processor runs on all nodes of the cluster.

[0051] Flowfile: The carrier for data transfer between NiFi processors is called a flow file. As each data object flowing in the system, the flow file represents a piece of data in a NiFi task and consists of an attribute part and a content part. The content is the data represented by the flow file, and the attributes are a set of key-value pairs that provide metadata related to the content.

[0052] Connection: NiFi uses a connector, Connection, to connect each processor and transfer data through the connector.

[0053] The canvas is a visual display of the NiFi framework. On the canvas, NiFi processor instances and task links can be created and configured, and the status information of processors and connections can be viewed.

[0054] Unless otherwise defined, all technical and scientific terms used in this document have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in this document are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0055] NiFi is a streaming data processing tool that performs a series of operations on data by defining and executing NiFi processors. Each NiFi processor completes a specific function, such as reading data, transforming data, filtering data, pushing data to a target system, etc. In enterprise data integration jobs, the data flow is usually complex, possibly involving multiple steps, different data sources, and different processing operations. These operations often need to rely on the results of the previous operation to determine whether to continue with the subsequent tasks. For example, after data is read from the source system and cleaned, it needs to be inserted into the target database. If the read or cleaning step fails, the subsequent insertion operation cannot be executed. Or, when a NiFi processor completes a task, the data is passed through a connector (Connection) to the next NiFi processor. If a NiFi processor fails to complete a task (e.g., fails or the data does not meet the requirements), the data will not continue to be passed to the downstream NiFi processor. In some integration jobs, multiple tasks can be executed in parallel. But even for parallel tasks, they need to wait for the completion of the previous tasks at certain critical nodes (e.g., the final writing or synchronization of data).

[0056] NiFi focuses on a streaming data transfer, which is real-time transferred to a specified data processing system instead of waiting for all data processing to be completed before transmission. This approach can improve data processing efficiency and reduce data processing costs. Currently, the NiFi cluster cannot sense the running status of the processors, resulting in users being unable to monitor the status or performance of the processors, etc., which is not conducive to the subsequent job process handling and there is a situation of low job processing efficiency.

[0057] The embodiments of this application provide a data processing method and a computing device. Among them, this method can monitor the running status of NiFi processors in the NiFi cluster, which is beneficial to triggering subsequent job processes according to the running status of NiFi processors and improving job processing efficiency. For example, by adding a scheduling status awareness module in the NiFi cluster, the running status of NiFi processors in the NiFi cluster can be obtained. The running status of NiFi processors includes but is not limited to obtaining the start time before scheduling execution, obtaining the end time after scheduling execution, obtaining the number of data currently being scheduled and processed during execution, and updating the status of this scheduling after scheduling ends. Based on the running status of NiFi processors known by the scheduling status awareness module, it is convenient to perform chain analysis to obtain fault diagnosis, performance monitoring, and business insight analysis, and solve the problem that the scheduling execution situation of NiFi processors cannot be sensed and subsequent job processes cannot be triggered.

[0058] Next, an introduction to the architecture of a data processing system provided by the embodiments of this application will be given.

[0059] Figure 1This is a schematic diagram of the hardware architecture for deploying a data processing system provided in an embodiment of the present application. As Figure 1 shown, the data processing system 100 includes a server cluster (such as a NiFi cluster, etc.) 101, a memory DB 102, and a terminal device 103. Among them, the DB 102 can be directly or indirectly communicatively connected to the server cluster 101 and the terminal device 103 through a wired network or a wireless network for data transmission, and the embodiments of the present application do not limit this.

[0060] Exemplarily, the server cluster 101 may include multiple computing devices (or called computing nodes). The number of computing devices is at least two, and can be more or less, and the embodiments of the present application do not limit this. It should be noted that the computing device can be a device such as a server, a computer, or a smart phone. For example, if the computing device is a server, the server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, and can also be a cloud server that provides basic cloud computing services such as cloud services, cloud business libraries, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, security services, and big data and artificial intelligence platforms.

[0061] Optionally, the multiple computing devices can be one type of device or a combination of multiple types of devices. The multiple computing devices can be deployed with an application program for running the data processing system to execute the functions of the data processing system, such as data cleaning, data storage and management, data conversion and integration, data analysis and mining, real-time data processing, security and privacy protection, and data synchronization. For example, the server cluster in the embodiments of the present application can be a NiFi cluster, and the NiFi cluster can perform the above various operations and processes on the data.

[0062] The computing device can be provided with a storage unit. The storage unit can be a memory for temporarily storing data, such as a cache, a dynamic random access memory (DRAM), a static random access memory (SRAM), etc., and is used to store the data that needs to be temporarily stored during the operation of the data processing system for access or operation by the processor or other hardware. In the embodiments of the present application, the storage unit can store the data corresponding to the computing tasks to be executed by the data processing system, the data corresponding to the unexecuted computing tasks in the interrupted computing tasks, the intermediate data during the processing, the processed data, and other cached data.

[0063] A computing device may be provided with a communication interface to enable data transmission with other computing devices, memories, and other devices. The communication interface may be an interface for wired transmission, such as a Compute Express Link (CXL) interface, a Peripheral Component Interconnect Express (PCIe) interface, a Universal Serial Bus (USB) interface, etc. The communication interface may be an interface for wireless transmission, such as a Bluetooth (BT) module, a Wireless Fidelity (WI-FI) module, a wireless communication module, etc.

[0064] Memory DB 102 is an independent storage device. Exemplarily, DB 102 may be a hard disk drive (HDD), a solid state drive (SSD), a NAND flash memory, a magnetic disk, etc. Exemplarily, DB 102 is used to store the overflow cache data when multiple computing devices run a data processing system, as well as other data, such as the programs for running the data processing system and the data processed by the data processing system. It should be noted that Figure 1 the connection of DB 102 to the computing device in [[ ]] is not limited to establishing a communication connection between DB 102 and the computing device, but also means that DB 102 establishes a communication connection with the data processing system deployed on the computing device, enabling DB 102 to establish a communication connection with any computing device.

[0065] DB 102 may be provided with a communication interface to enable data transmission with multiple computing devices and other devices. The communication interface may be an interface for wired transmission, such as a CXL interface, a PCIe interface, a USB interface, etc. The communication interface may be an interface for wireless transmission, such as a BT module, a WI-FI module, a wireless communication module, etc.

[0066] Exemplarily, the above-mentioned wired or wireless networks may use standard communication technologies and / or protocols, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), mobile, wired networks, private networks, or virtual private networks.

[0067] The terminal device 103 can be an entity on the user side for receiving or transmitting signals. The terminal device 103 can be referred to as a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal device, industrial control terminal device, UE unit, UE station, mobile station, remote station, remote terminal device, mobile device, wireless communication device, UE agent, or UE device, etc. The terminal device 103 can be fixed or mobile. It should be noted that the terminal device 103 can support at least one wireless communication technology, such as Long Term Evolution (LTE), NR, 6th-generation (6G) mobile communication system, or next-generation wireless communication technology, etc.

[0068] For example, the terminal device 103 can be a mobile phone, tablet (pad), desktop computer, laptop computer, all-in-one computer, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, wireless terminal in smart home, cellular phone, cordless phone, Session Initiation Protocol (SIP) phone, Wireless Local Loop (WLL) station, Personal Digital Assistant (PDA), handheld device with wireless communication function, computing device, or other processing device connected to a wireless modem, wearable device, terminal device in a future mobile communication network, or terminal device in a future evolved Public Land Mobile Network (PLMN), etc. The embodiments of the present application do not limit the specific technologies and specific device forms adopted by the terminal device 103.

[0069] In one embodiment, when the server cluster is a NiFi cluster, the NiFi cluster is composed of multiple NiFi nodes (computing devices), and each node is a NiFi instance. They cooperate with each other to process data flows (FlowFile). Exemplarily, multiple NiFi processors are deployed on a single NiFi node. Figure 1(not shown). After the NiFi node runs, task scheduling will be performed according to the scheduling type selected by the user. The scheduling type determines how tasks are executed and scheduled, while task scheduling indicates how to manage and allocate the execution of data flow tasks (i.e., processors, process groups, etc.) among multiple nodes. In this embodiment, the NiFi node can monitor the running status of the NiFi processor. By monitoring and tracking the execution status of each processor in the NiFi data flow, for example, the start and end times of the NiFi processor scheduling execution, the amount of data processed, and the status after the scheduling ends, and storing the obtained running status of the NiFi processor processor in the memory DB 102, the running record is stored in the database, so as to help understand the performance, latency, and potential problems of the entire data flow, and further provide support for business insight and fault diagnosis.

[0070] In one embodiment, the terminal device 103 can obtain the scheduling execution status of the NiFi processor stored in the memory DB 102, such as the start and end times of the scheduling execution, the amount of data processed, and the status after the scheduling ends, and trigger other program processes according to the scheduling execution status of the NiFi processor. For example, the terminal device 103 determines whether the NiFi processor fails according to the scheduling execution status of the NiFi processor stored in the memory DB 102, and executes corresponding processes according to the judgment result. For example, in the case of a NiFi processor failure, a failure alarm is issued, or the business process related to the faulty processor is suspended, etc., to ensure the efficient cooperation and flexible scheduling of the data flow and processors in the NiFi cluster among different nodes, and at the same time be able to handle task failures or abnormal states, ensuring the stability and reliability of the entire data flow.

[0071] It can be understood that the data processing system described in the embodiments of the present application is to more clearly illustrate the technical solutions of the embodiments of the present application, and does not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0072] To facilitate the understanding of the technical solutions of the present application, next, the data processing principle of the embodiments of the present application will be introduced in detail.

[0073] Figure 2 This is a schematic diagram of the software architecture of a data processing system provided in the embodiments of the present application. As Figure 2As shown, the data processing system 200 can be deployed on server nodes (such as physical servers like rack servers and cabinet servers, or virtual machines, etc.). For example, the data processing system 200 can run within a virtual machine on a host operating system (OS). From a software perspective, the data processing system 200 can include a Web service module 201 (Web Serve), an authentication module 202 (Authorzation), a flow controller module 203 (Flow Controller), a processor module 204, and a repository module 212. Optionally, the processor module 204 includes a timing scheduling module 205 (Quartz Schedule) and a scheduling operation status awareness module 206 (ScheduleRunCondation). The repository module 212 (Repository / DB) can include: a flow file repository 208 (FlowFile Repository), a content repository 209 (Content Repository), an origin repository 210 (Provenance Repository210), and a data warehouse 211 (datebase, DB) 211.

[0074] Among them, the Web service module 201 is used to host the HTTP-based command and control API (application programming interface) of NiFi. The data processing system 200 provides a browser-based user interface (UI) through the Web Server 201, and this user interface can include a canvas. Users can perform system configuration, monitoring, and management through this interface. When a user accesses the Web interface, they actually communicate with the Web Server through the HTTP protocol. The Web interface allows users to: configure the data flow, including adding, configuring, and connecting processors; monitor the status of the data flow, view the running conditions, throughput, failure records, etc. of each processor; view the system health status, including the resource usage of the server, the latency of the data flow, etc.; manage user permissions, configure access control lists (ACL), etc.

[0075] The authentication module 202 provides identity authentication and authorization authentication for NiFi. Among them, identity authentication is used to verify the identity of users to ensure that the system allows legitimate users to access. Authorization authentication is used to verify whether a user has sufficient permissions to perform certain operations. Specifically, through identity authentication, the data processing system 200 can verify the identity of users, usually by means of username and password, certificates, OAuth, etc. For example, when logging in to NiFi through the web interface, the data processing system 200 will perform identity authentication to ensure that the username and password you provide are correct. If the authentication is successful, you can enter the data processing system 200 and perform operations. Once the user's identity is authenticated, the data processing system 200 will check whether the user is authorized to perform specific operations. Authorization authentication ensures that only users with sufficient permissions can perform certain sensitive or important operations. For example, even if you successfully log in to the data processing system 200, if you do not have sufficient permissions, you cannot modify the data flow configuration or start the NiFi processor. The data processing system 200 determines whether a user can perform an operation based on the configured role permissions.

[0076] The flow controller module 203 is the brain of NiFi operations. It provides running threads for the extension program 207 and manages the schedule for the extension program 207 to receive resources for execution. Specifically, as the core component of NiFi, the flow controller module 203 is responsible for managing the life cycle of the entire data flow. Exemplarily, the flow controller module 203 can control elements such as multiple NiFi processors, connectors, queues, and storage to form the data flow in NiFi. The flow controller module 203 coordinates the operation of these elements to ensure that data can flow and be processed correctly and in a timely manner. That is to say, the flow controller determines how data flows, when it flows, how much flows, and what operations to perform at what time. It is the scheduling center of NiFi and is responsible for scheduling the operation of each module.

[0077] It should be noted that the extension program 207 (Extension) is various functional modules in the data processing system 200. There are multiple types of NiFi extensions. For example, it can include NiFi processors, controller services (Controller Service), reporting tasks (Reporting Task), etc., which are used to extend the functions of NiFi. These extension modules are usually written by developers and executed as part of the NiFi process. Since NiFi is written in Java, all extension modules are also developed based on Java. These modules will run in the Java virtual machine (JVM), and the JVM provides cross-platform support, enabling NiFi extensions to run consistently on different operating systems. Therefore, the key point here is that the extension runs and executes in the JVM. Each extension program performs a specific function. The running thread is the "unit of work" that executes the operations of the extension program. For example, when a data stream needs to be processed by a certain NiFi processor, the flow controller module 203 will allocate a thread for the NiFi processor and start the thread to execute the work of the NiFi processor. The number and allocation method of the threads depend on the system load and configuration. The flow controller module 203 manages these threads to ensure that each extension program 207 can run in parallel and efficiently. For example, if a NiFi processor needs to process a large amount of data, it may be allocated more threads to improve the processing speed. Resource management means that the flow controller module 203 will decide which extension programs 207 (such as NiFi processors) use system resources (such as CPU, memory, network bandwidth, etc.) to execute tasks at what time. The execution schedule (Schedule) indicates that the flow controller module 203 will decide when to execute certain operations according to the set time. For example, a certain data processing task is scheduled for cyclic processing every 20 minutes. This scheduling method helps to improve the efficiency of the system and the utilization rate of resources. For example, when the traffic is large, the flow controller can allocate processing tasks according to priorities and schedules to avoid system overload.

[0078] The processor module 204, as a collection of different NiFi processors in NiFi, is the core component of the data processing system 200 and is responsible for processing and transforming the data stream. The processor module 204 is the most important part of NiFi. They are responsible for the inflow, outflow, routing, and operation of data. By dragging various processors in the processor module 204 onto the canvas, users can select the types of processors to be used and connect them in series into a data stream through connection lines, thereby realizing various operations and processing of the data.

[0079] Each processor performs a specific function and can be used alone or in combination with other processors to complete complex data processing tasks. For example, if a user wishes to read data from a database and convert it to CSV format, they can drag a database connection processor (such as ExecuteSQL) onto the canvas to extract data from the database. After that, drag a conversion processor (such as ConvertRecord) to convert the data from SQL results to CSV format. Finally, drag an output processor (such as PutFile) to save the converted data to the file system, such as HDFS (Hadoop Distributed File System).

[0080] In one embodiment, the processor module 204 may include processors such as ListenTCP, ListenUDP, ListenSyslog, ExecuteGroovyScript, RouteOnAttribute, EvaluateJsonPath, EvaluateXPath, ExtractGrok, and ReplaceText. This application does not strictly limit them here. Through different NiFi processor types, various operations and processing of data are achieved. For example, assume a data flow that needs to obtain data in JSON (JavaScript object notation) format from an HTTP interface, then convert it to CSV (comma-separated values) format, and finally save it to the file system. This task can be completed through the following processors: Obtain data (GetHTTP or ListenHTTP processor) to obtain JSON format data from the HTTP interface. For example, the NiFi instance will regularly go to a certain REST API endpoint to obtain new JSON data. Convert data format (ConvertRecord processor): Convert the obtained JSON data to CSV format. The ConvertRecord processor of NiFi can perform data format conversion according to the configured Record Reader and Record Writer. Store data (PutFile processor): Write the CSV format data to the local file system. This processor allows specifying the target folder, and NiFi will automatically save the processed data to the specified location. The entire data flow drags processors through the graphical interface (canvas) to complete the operations from data reception, data conversion to data storage.

[0081] Exemplarily, a timing scheduling module 205 (Quartz Schedule) can be deployed on the processor module 204. Optionally, a timing scheduling module 205 can be deployed for each NiFi processor, or multiple NiFi processors of the same type (such as NiFi processors of the ExecuteSQL type) can share a timing scheduling module 205. In this embodiment, according to one or more timing scheduling modules 205, a scheduling operation status awareness module 206 (ScheduleRunCondation) can be deployed on the processor module 204. The scheduling operation status awareness module 206 can obtain the start time, end time, amount of data processed, and status after scheduling of the NiFi processors scheduled by the timing scheduling module 205.

[0082] Specifically, the timing scheduling module 205 provides the function of running the NiFi processor at regular intervals according to the specified time interval, which is usually used for tasks that need to be executed periodically or with a delay. Specifically, the timing scheduling module 205 is used to trigger tasks or schedule a part of the tasks at regular intervals. It provides an accurate time scheduling mechanism to help control the execution timing of the NiFi processor. It can be understood that the time interval can be a specific duration, such as every 5 minutes, 1 hour, etc. The timing scheduling module 205 uses a standard timing expression (such as a Cron expression) to set the scheduling rules. A Cron expression is a time-based string format used to represent the scheduling rules of timing tasks, such as defining the execution period of tasks. A Cron expression usually consists of 6 or 7 fields, and each field represents a specific time unit. These fields include seconds, minutes, hours, date, month, day of the week, and year (optional). By combining these fields, a Cron expression can accurately describe the execution time of the task. Exemplarily, Quartz supports 6 fields, and standard Unix / Linux cron usually has 5 fields. The specific meanings of each field are as follows:

[0083] Seconds Minutes Hours Day of month Month Day of week Year (optional)

[0084] *******

[0085] Exemplarily, assume there is a NiFi processor whose task is to obtain the latest data from an API interface at regular intervals and store it in a local file. To ensure that it executes once every hour, the execution rule can be set through the timing scheduling module 205. Configuration steps: In the timing scheduling module 205 of NiFi, configure the timing task and set its trigger time to execute once every hour. Select the NiFi processor and associate it with the timing scheduling module 205. Configure the Quartz scheduling expression as 0 0 * * *? (Cron expression: seconds: 0; minutes: 0; hours: * (any hour of the day); date: * (any day of the month); month: * (every month); day of week:? (irrelevant to the day of the week)), and this expression indicates that the task is triggered at the first minute of each hour (i.e., on the hour). Execution result: Every hour on the hour, the timing scheduling module 205 will automatically trigger the NiFi processor, then pull data from the API interface and store it in the specified folder. Thus, through the timing scheduling module 205, users can set a specific time interval so that the NiFi processor can execute tasks regularly according to the set time rules.

[0086] The scheduling operation status awareness module 206 is used to record the operation status of the NiFi processor during scheduling execution. For example, the scheduling operation status awareness module 206 can obtain the start time before scheduling execution, the end time after scheduling execution, the number of data processed in the current scheduling during the execution of the NiFi processor when it starts, update the status of this scheduling after scheduling ends, and save the start and end times, data volume, and status of the scheduling execution to the data warehouse 211 (DB, database), realizing the storage of operation records in the database, thereby helping to understand the performance, latency, and potential problems of the entire data flow, and further providing support for business insights and fault diagnosis.

[0087] It should be noted that the scheduling operation status perception module 206 can obtain the start time, end time, and scheduling status in the following ways: NiFi provides a Schedule property that allows the timed scheduling module 205 to start a NiFi processor at specific time intervals or conditions. When NiFi starts the processor, the start time of the task can be obtained through the life cycle event of the NiFi processor. Alternatively, in the onTrigger method of the NiFi processor, the scheduling operation status perception module 206 can record the start time of the task. The onTrigger method is triggered during scheduling execution. Each time the NiFi processor starts execution, the scheduling operation status perception module 206 can call a time function (such as System.currentTimeMillis()) to record the start time. Exemplarily, the function code can be: long startTime = System.currentTimeMillis(); / / Record the start time. When the onTrigger method of the NiFi processor is completed, that is, after the task execution ends, the scheduling operation status perception module 206 can obtain the end time. When recording the end time, the execution duration can be calculated by comparing with the start time, and the scheduling status can be marked as successful or failed. For example, if there is no exception, the status is set to successful, otherwise it is failed.

[0088] The scheduling operation status perception module 206 can obtain the amount of data during execution in the following ways: Each NiFi processor in NiFi processes a batch of data, usually using FlowFile as the data unit. In the onTrigger method, the number and size of the currently processed FlowFile can be accessed. The scheduling operation status perception module 206 can obtain the amount of data currently being processed by accessing the FlowFile in the ProcessSession (process session). The ProcessSession allows developers to obtain information such as the number and size of FlowFile. If more detailed information is needed (such as the byte size of the data), the size() can be obtained from the properties of the FlowFile.

[0089] The dispatching operation status perception module 206 saves the operation status during dispatching execution to the data warehouse 211: This can be achieved by connecting to the database and using SQL statements to insert information into the corresponding tables. Alternatively, JDBC (Java Database Connectivity) can be used to connect to the database and execute the INSERT operation to store the execution details (start time, end time, data volume, status). In this way, the dispatching operation status perception module 206 can accurately record and monitor the execution of dispatching tasks, providing data support for subsequent performance analysis, fault troubleshooting, and optimization.

[0090] Optionally, the dispatching operation status perception module 206 implements the collection and storage of dispatching operation status data based on Aop. It should be noted that AOP (Aspect-Oriented Programming) is a programming paradigm used to improve the modularity of code by separating concerns. In AOP, the core business logic of the program is separated from auxiliary functions (such as logging, transaction management, etc.). This allows some common functions (such as monitoring, exception handling, logging, etc.) to be "weaved" into different parts of the program without modifying the core business code. The goal of aspect-oriented programming is to decouple cross-cutting concerns (such as monitoring and logging of dispatching tasks) from the business logic, making the code clearer and easier to maintain. In traditional programming, we usually add logging or status recording code at each key point of the task scheduling code (such as task start, task end, task failure, etc.). In AOP, we can handle these operations uniformly through aspects. An aspect will be automatically triggered at specific "join points" (such as before method execution, after method execution, or when an exception is thrown), thus realizing the collection of the running status of dispatching tasks. This process does not require modifying the original dispatching task code, but instead "weaves" into the execution process of the dispatching task through AOP. In this embodiment, by using an AOP framework such as Spring AOP, functions such as the collection of the execution status of task scheduling and logging can be achieved without modifying the core business code. In the specific implementation, Spring AOP intercepts the invocation of the target method by using the proxy pattern and executes additional operations before, after, or when an exception is thrown during the method invocation. Specific implementation steps:

[0091] Define the aspect: The aspect is the core of AOP, which defines when (such as before method execution, after method execution, or when method execution throws an exception) we want to execute certain operations. For example, we can record the start time when the task starts to execute, record the end time when the task is completed, and save this data to the database.

[0092] Configure AOP Proxy: AOP intercepts target methods through a proxy (usually a dynamic proxy). In Java, we typically use Spring AOP to achieve this functionality. By configuring aspects, data collection operations can be automatically executed before and after the execution of the task scheduling module. During the execution of the scheduling task, we will collect relevant status information, such as: Start Time of the task: Records the start time of the task. End Time of the task: Records the time when the task is completed. Execution Status of the task: Records whether the task is successfully completed, failed, timed out, etc. Error Message: If the task execution fails, the reason for failure or exception information can be recorded. Duration: Calculates the execution duration of the task, i.e., the time difference from start to end. This information will be obtained through method parameters, return values, or exception information in the AOP aspect method.

[0093] Data Storage: In the aspect, when the task execution status is captured, the relevant data will be inserted into the database. For example, we can insert information such as the task execution status, start time, and end time into the scheduling task status table.

[0094] Specifically, after collecting the execution status information of the task, these data can be stored in the database through the JDBC (Java Database Connectivity) or ORM (Object Relational Mapping) framework (such as Hibernate, MyBatis). For example, a database table (such as the task_run_status table) is used to store the execution status data of the task. The fields of the table may include: task_id (task ID): uniquely identifies a scheduling task. task_name (task name): the name of the task. start_time (start time): the time when the task starts to execute. end_time (end time): the time when the task ends. status (status): the execution status of the task (such as success, failure). error_message (error message): if the task fails, record the error message. duration (duration): the duration of the task execution. In other words, each table in the database has predefined fields (columns), and these fields represent the data attributes to be stored. Just like the above database table has fields such as "start_time, end_time, status", each field corresponds to a specific type of data. After that, the originally collected data (scheduling operation data) is converted into data that conforms to the database field type and format according to the requirements of the database table structure, and classified and stored in the database according to business requirements, and the records are saved to the data warehouse 211 for subsequent data analysis. In this embodiment, through AOP, the scheduling operation status perception module 206 can automatically collect the start, end time, data volume, and status data of the task during the task scheduling execution process, and store these data in the data warehouse 211.

[0095] The flow file repository 208 is a component used by NiFi to track the status of all FlowFiles in the current process. A FlowFile is the core data unit in NiFi, representing a data stream. Each FlowFile stores the content of the data (such as a file, message, or other data) and related metadata (such as file name, size, attributes, etc.). Specifically, the flow file repository 208 is responsible for storing and managing the status information related to each FlowFile, including their positions in the NiFi process and different stages of the life cycle. It can be understood that the implementation of the flow file repository 208 is to persistently store the data in the form of a frontend log on a specified disk partition.

[0096] The Content Repository 209 is the actual content bytes of a given FlowFile. That is, the Content Repository 209 is where the actual data (such as file content or messages) in the FlowFile is stored. Exemplarily, the Content Repository 209 stores data blocks in the file system and allows configuring multiple storage locations to disperse the storage load. For example, multiple file system storage locations can be specified to obtain different physical partitions to reduce contention on any single storage volume.

[0097] The Provenance Repository 210 is a place for storing Provenance Event Data related to a data flow (FlowFile). Provenance here refers to the historical tracking and source recording of data. That is, the Provenance Repository records the whole process information of data from generation to processing and then to final storage, helping users trace the origin, transfer path, processing situation, etc. of the data. Event Data is the detailed information of each operation or event related to a certain data flow or FlowFile. For example, when a FlowFile is transferred from one NiFi processor to another, or when a FlowFile is modified, deleted, or passed to an external system, event data will be generated. The Provenance Repository 210 records all event data (i.e., source events) related to the data flow processing process. This repository enables users to view, analyze, and audit the history of the data flow and its processing situation.

[0098] The Data Warehouse 211 is a place for storing the scheduling execution status data of the Storage Scheduling Health Awareness Module 206. In this database, there will be one or more tables for storing this data. Each piece of data usually represents an execution record of a scheduling task, including information such as the start time, end time, and execution status of the task. Thus, the data is transformed into database entity fields, classified and stored in the database according to business requirements, and the records are saved to the Data Warehouse 211 for subsequent data analysis.

[0099] In this embodiment, the use of each repository in the Repository Module 212 enables the execution status of the scheduling task to be saved and queried for a long time, helping users understand the history of task execution, facilitating chain analysis to obtain business insights or conduct fault diagnosis.

[0100] It should be noted that the physical structure of each repository in the Repository Module 212 is pluggable. The default implementation at the physical level of each repository is to use one or more physical disk volumes, and within each location, the event data is indexed and searchable.Figure 3 Shows possible ways to perform subsequent analysis on the scheduled operation data in the repository in the embodiments of the present application. As Figure 3 shown, in this embodiment, according to the data saved in the repository during the scheduling execution of the NiFi processor, the detailed results of the data processing for each scheduling of each NiFi processor can be analyzed in real time to ensure the efficiency and accuracy of the scheduling process. For example, the data processing situation of the scheduling can be displayed in real time in the form of a report; or, the data processing situation of the scheduling can be analyzed in real time in the form of statistical charts; or, the data processing situation of the scheduling can be analyzed in real time in the form of a mathematical model; or, in an integrated job, other process jobs can be triggered according to the running situation of the scheduling.

[0101] In one embodiment, the scheduling execution results can be queried in real time through SQL queries and reporting tools (such as Power BI, Tableau), and a report that is updated in real time can be generated for analysis. Through graphical reports (such as tables, charts, graphs, etc.), the detailed situation of the scheduling execution can be displayed in an intuitive manner. For example, a report can be generated to show the execution situation of each scheduling, including the execution status, start time, end time, execution time, and the amount of data processed by each NiFi processor. These data will be presented to the user in the form of a table, and the user can judge which scheduling tasks are abnormal based on the table data. This way can help the user quickly understand the overall running situation of the scheduling and identify potential abnormalities or bottlenecks.

[0102] In one embodiment, the scheduling data is visually displayed through charts (such as bar charts, line charts, pie charts, etc.). Charts can help users quickly identify data trends, abnormal patterns, and bottlenecks. Common analysis dimensions include processing time, failure rate, execution efficiency of each NiFi processor, etc. Exemplarily, chart tools (such as Excel, Tableau, Power BI, Grafana) are used to visualize the scheduling data. For example, the change trend of the execution time of the scheduling task can be shown through a line chart, or the success / failure ratio of each NiFi processor can be shown through a pie chart. Real-time data stream processing: Combine real-time data stream processing tools (such as Apache Kafka, Apache Flink) to analyze the real-time scheduling data and generate charts in real time. Illustration: If we want to analyze the execution situation of all NiFi processors within a certain time period, we can use PowerBI to draw a bar chart to show the number of records executed or the processing time of each NiFi processor, so as to find out which NiFi processors execute slowly or have a small record processing volume.

[0103] In one embodiment, by applying mathematical models (such as statistical models, machine learning models, etc.) to analyze scheduling data, it helps to discover potential patterns or abnormal patterns. For example, using regression analysis, clustering analysis or anomaly detection models to analyze performance bottlenecks, reasons for task failures, etc. during the scheduling process. Exemplarily, statistical analysis: statistical methods (such as mean, variance, standard deviation, etc.) can be used to analyze the distribution of scheduling results and identify outliers in the data. Machine learning models: using machine learning algorithms (such as decision trees, random forests, K-means clustering, etc.), by training historical scheduling data, predict future scheduling results or discover potential problems. For example: Suppose we find that some NiFi processors have extended execution times during peak processing periods by analyzing historical scheduling data. Through mathematical models (such as regression models or time series analysis), we can predict the execution situation during peak periods and optimize scheduling tasks in advance.

[0104] In one embodiment, after the scheduling task is completed, other jobs can be triggered according to the running results of the scheduling (such as success, failure, timeout, etc.). For example, when a certain scheduling task fails, it can trigger the re-execution of the task or trigger an email to notify relevant personnel. Exemplarily, use Figure 2 the timing scheduling module 205 in to manage and schedule jobs. The timing scheduling module 205 can decide whether to trigger other jobs according to the status of the scheduling (such as success, failure). Or trigger subsequent processes according to specific conditions (such as failure retry, success notification, etc.). For example: Suppose in a data processing schedule, a certain NiFi processor fails to execute. If the system detects the failure status, the NiFi processor can trigger a job for failure retry or notify the administrator to handle the problem.

[0105] Through the above several methods, the data processing system 200 can achieve comprehensive monitoring and analysis of the data processing process. Each analysis method has its specific advantages, which can help users discover problems in a timely manner and optimize the data processing flow. Database storage and real-time query help us track the detailed execution results of each NiFi processor. Reports and charts display show the running situation of scheduling tasks in an intuitive way. Mathematical model analysis helps us discover potential patterns or anomalies through data mining and modeling. Workflow management and automation trigger subsequent processes or alarms according to the scheduling execution results to ensure the smooth progress of scheduling tasks. Combining these technical means can greatly improve the efficiency and reliability of the scheduling process, ensure that data processing tasks are completed smoothly as expected, and is also beneficial to triggering subsequent job processes according to the running situation of NiFi processors and improving job processing efficiency.

[0106] The above is the introduction to the principle and process of the data processing system provided by the embodiments of the present application. Based on the above content, the technical solutions of the present application will be described in detail below with specific embodiments. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0107] Figure 4 FIG. shows a schematic flow chart of a data processing method provided by an embodiment of the present application. As Figure 4 shown, it can be understood that this method can be executed by any device, equipment, platform, or equipment cluster with computing and processing capabilities. Exemplarily, this method can be executed by a data processing device, where the device can be implemented by software and / or hardware, and can be, but is not limited to, configured in a computing device cluster including at least one computing device. Typically, it can be configured in any NiFi node in the NiFi cluster. For ease of description, in this embodiment, any NiFi node in the NiFi cluster will be used as the execution subject for description. As Figure 4 shown, the data processing method may include the following steps:

[0108] S101. Obtain a scheduling task; wherein, the scheduling task indicates that the NiFi node schedules a NiFi processor to process a data stream.

[0109] In this implementation step, the scheduling task is created according to a scheduling policy, and the scheduling policy indicates the execution period and trigger condition of the processor. The scheduling task indicates that under the drive of the scheduling policy, the NiFi node schedules the execution of the NiFi processor according to the specified policy.

[0110] Exemplarily, in the NiFi node, the NiFi processor is the basic unit for executing data stream tasks. Each NiFi processor can be configured with a scheduling policy (for example, timed trigger, on-demand trigger, etc.). There are various types of NiFi processors. For example, a Processor of the ExetuceSQL type is used to execute SQL queries. Of course, the NiFi processor can also be of types such as PutSQL, ConsumeMQTT, SplitJSON, etc., and the present application is not limited thereto. It can be understood that PutSQL is used to write data into a database; ConsumeMQTT is used to consume MQTT messages; SplitJSON is used to split JSON data. Each NiFi processor can execute periodic or timed tasks according to the set scheduling policy.

[0111] In one embodiment, the scheduling policy refers to how often the NiFi processor executes. For example, create a Scheduler task for scheduling. Exemplarily, the timing interval can be specified by setting a cron expression on the page, and the scheduling control is throughFigure 2 It is controlled by the timing scheduling module 205 (Quartz Schedule) in []. For example, for a NiFi processor of the ExetuceSQL type, the scheduling policy can be set to execute once every 30 minutes through a cron expression. Then, the timing scheduling module 205 will execute the onTrigger() method (scheduling task method) of the NiFi processor once every 30 minutes for data processing. It can be understood that when using the timing scheduling module 205, the onTrigger() method is the core method combined with the scheduling trigger mechanism of the timing scheduling module 205 to start the actual task of the processor. For a custom scheduling task (such as a NiFi processor of the ExecuteSQL type), the timing scheduling module 205 will call this method to execute the actual task logic, and the onTrigger() will contain code for executing SQL queries, data updates, or other database operations.

[0112] Optionally, in this embodiment, before obtaining the scheduling task, the NiFi processor needs to be started first, and then it is mobilized according to the scheduling policy configured in the NiFi processor scheduling task. Specifically, Figure 5 shows a schematic diagram of the running process of a NiFi processor provided in an embodiment of the present application. As Figure 5 shown, when the NiFi processor is running, it includes the following steps:

[0113] S201. Start the NiFi processor.

[0114] Exemplarily, in the code, first obtain an instance of the NiFi processor, and then call the start() method through the Processor API of the NiFi processor to start the NiFi processor. The NiFi processor will decide whether to start according to the pre-set scheduling policy.

[0115] S202. Obtain the current scheduling status of the NiFi processor.

[0116] At startup, we can obtain the current scheduling status of the NiFi processor, such as whether it is in the Running, Stopped, or Disabled state. Determine whether to continue the scheduling process based on the current scheduling status of the NiFi processor. Exemplarily, determine the current scheduling status of the NiFi processor and perform different operations according to different statuses. In one possible implementation, when the current scheduling status of the NiFi processor is Running, step S203-1 is executed. In another implementation, when the current scheduling status of the NiFi processor is Disabled, step S203-2 is executed. In yet another implementation, when the current scheduling status of the NiFi processor is Stopped, step S203-3 is executed.

[0117] S203-1. The process is completed and returns.

[0118] If the processor scheduling status is Running, it means the process has been completed and there is no need to continue execution, so return.

[0119] S203-2. Terminate the process and prompt "Processor is disabled".

[0120] If the status is Disabled, then terminate the process and prompt that the NiFi processor is disabled. For example, remind the user through logs or a prompt box.

[0121] S203-3. Use CAS to modify the scheduling status to: STARTING.

[0122] If the status is Stopped, then use the CAS (compare and swap) mechanism to modify the scheduling status to Starting to prepare to start the NiFi processor. It should be noted that the CAS mechanism is an atomic operation widely used in concurrent programming to avoid data races and ensure thread safety. It is usually used to modify shared variables in a multi-threaded environment to ensure that there are no conflicts when updating variables. In the NiFi node, the CAS mechanism is usually used to ensure that in a multi-threaded or distributed environment, the modification of the NiFi processor status is safe and to avoid conflicts caused by multiple threads modifying the same status simultaneously. For example, Compare: CAS will compare a value in memory (e.g., the status of the NiFi processor) with the provided expected value (e.g., "Stopped"). Swap: If the comparison result is true (i.e., the current value and the expected value are equal), it will update the current value to the new value (e.g., change the status to “Starting”). If the comparison fails (i.e., the current value and the expected value do not match), no modification will be made and the CAS operation will return a failure.

[0123] Review Figure 4 After step S101, step S102 is executed:

[0124] S102. Schedule the NiFi processor to execute the scheduling task.

[0125] In this implementation step, according to the execution period and trigger conditions of the NiFi processor, the scheduling task of the NiFi processor is executed for data processing.

[0126] Exemplarily, when the specified time is reached, the timed scheduling module 205 in the NiFi node will trigger the corresponding scheduling task. Taking the execution of SQL as an example, when the scheduling task is triggered, ExecuteSQL will execute the corresponding SQL query to obtain and process data. Specifically, refer to Figure 5 The execution process of the NiFi processor scheduling task is as follows:

[0127] S204. Activate the current thread task. Ensure that the scheduling task can be executed in the current thread. Once the scheduling status of the processor is updated to Starting, a thread needs to be activated to process the scheduling task. Exemplarily, the scheduling task is triggered by the timed scheduling module 205, and the onTrigger() method of the corresponding NiFi processor (such as ExecuteSQL) is called to activate the execution of the NiFi processor. In some embodiments, the NiFi node can also create and execute a new thread task through ExecutorService.

[0128] S205. Reflectively obtain the current processor scheduling method and call it.

[0129] In this step, the scheduling method of the NiFi processor is called through reflection. Reflection is an important feature in the Java language. It can obtain the attributes, methods, and other class information of a class through the fully qualified name of the class, and then call these attributes and methods. That is to say, the Java reflection mechanism allows us to dynamically obtain class information (such as class attributes, methods, etc.) at runtime and dynamically call these methods. It should be noted that in Java, the fully qualified class name (FQCN) of a class refers to the complete name of the class, including the package name and the class name. It is the way to uniquely identify a class.

[0130] In this embodiment, the scheduling policy of the NiFi processor is obtained through reflection, and the scheduling task is executed. Specifically, through the reflection mechanism, the system can dynamically call the scheduling-related callback methods (such as the trigger() method) without knowing the specific type of the NiFi processor. Exemplarily, the class object of the processor is obtained through Class.forName(), and then the scheduling policy method of this class is obtained and called. In this way, it can be applied to scheduling different types of processors. The NiFi processor processes or operates on data according to the scheduling policy configuration.

[0131] Optionally, in S206, remove the current thread task.

[0132] In this step, the current thread can be removed after the scheduling task is executed to ensure resource release. Exemplarily, after the scheduling task is executed, if it is necessary to end the thread, the current thread can be removed through the shutdown() method of the ExecutorService or other tool components to release resources.

[0133] Review Figure 4 , in the data processing method, during the execution of step S102, step S103 is synchronously executed:

[0134] S103. During the process of the NiFi processor executing the scheduling task, obtain the scheduling operation data of the NiFi processor and store the scheduling operation data of the NiFi processor.

[0135] In one embodiment, the scheduling operation data includes the start time, end time, data volume processed, and execution result of the NiFi processor executing the data processing task.

[0136] It can be understood that the start time of the NiFi processor executing the data processing task indicates the start time of each data processing task, which is used to calculate the execution duration of the task and provide a time reference for scheduling analysis. The end time indicates the completion time of each data processing task, and the execution period and duration of the task can be calculated in combination with the start time. The data volume processed by each task can be the number of data records, the number of bytes, or other unit measurement methods, which is used to measure the load and complexity of the task. The execution result indicates the execution status of the data processing task, including success, failure, timeout, or exception, etc., which is used to reflect whether the task is completed as expected and whether there are any problems.

[0137] Exemplarily, according to the scheduling operation status perception module 206 deployed in the NiFi node (see Figure 2)Obtain the scheduling operation data. Specifically, the scheduling operation status perception module 206 first interacts with the timed scheduling module 205 in the NiFi node to collect the scheduling status of all current NiFi processors. For each scheduling task, the scheduling operation perception module 206 records information such as its scheduling period, the time points of task start and end, task execution result (success / failure), and the amount of data processed. Execution information of the scheduling task: By interacting with the interface of the NiFi node (such as the API or internal log system of the NiFi node), the scheduling operation perception module 206 can obtain the execution details of each NiFi processor task in real time. Data aggregation and analysis: The scheduling operation perception module 206 not only collects the individual execution data of each task, but also aggregates the execution situations of all tasks and sends them to the memory (such as the database DB) for subsequent data analysis and scheduling optimization. For example, visual statistical reports and charts can be generated for reference by operation and maintenance personnel and developers.

[0138] In one embodiment, the scheduling operation status perception module 206 realizes the acquisition and storage of scheduling operation data in the memory based on Aop. In the scheduling operation status perception module, AOP (Aspect-Oriented Programming) is used to implement the acquisition and storage of scheduling operation data. AOP allows additional functions to be dynamically inserted into the existing business logic during system operation without modifying the original code structure. In this way, the monitoring of scheduling tasks and the acquisition of performance data can be carried out without disturbing the main business logic. Taking the NiFi node as an example, assume that we want to monitor the execution process of a certain scheduling task: Define an aspect using AOP to intercept the execution method related to this task (such as the process() method). Before the method execution, record the task start time through the before advice and record the metadata related to the task (such as task ID, scheduling period, etc.). After the method execution, record the task end time, execution result, data volume, etc. through the after advice. If an exception occurs during the execution, record the error information through the exception advice and store the exception type and stack information in the memory. AOP decouples the monitoring and data acquisition logic of scheduling tasks from the core business code, avoiding directly modifying the business code, thereby reducing the system coupling degree. This makes the system more modular, maintainable, and easy to expand. When new monitoring functions need to be introduced, only the corresponding aspect needs to be written without modifying the core scheduling logic.

[0139] In one embodiment, at the code level of the scheduling operation status perception module 206, the onTrigger method will be triggered during the scheduling execution. Each time the processor starts to execute, the system can call a time function (such as System.currentTimeMillis()) to record the start time. For example, the current timestamp (in milliseconds) can be obtained by calling the System.currentTimeMillis() method. In this way, the purpose of recording the start and end times of the task can be achieved, which is convenient for subsequent analysis of the task execution duration or performance metrics.

[0140] In one embodiment, when obtaining the NiFi processor to execute the scheduling task, a data set of the data type of the currently scheduled processed data is returned, and the number of processed items of this scheduling is recorded based on the data set to record the amount of processed data. This data set contains various attributes of the currently processed data, such as data type, number of records, data size, etc. It should be noted that during the execution of the scheduling task, the data set is not only a collection of data, but also contains meta-information about the data, such as: data type being processed: for example, text data, JSON, XML or other formats. The number of processed items refers to how many records the data set contains. The size of the processed data: represents the overall size of the data set or the average size of each piece of data. The processing status: records whether the task is successfully executed, whether there are errors or exceptions, etc. Each time the NiFi processor executes a scheduling task, the number of items in the data set is recorded. In some embodiments, the number of items in the data set can also be obtained through built-in attributes of the NiFi node such as flowFileCount or custom attributes. flowFileCount is an attribute used by the NiFi node to represent the number of data files in the data stream. It can be used as an indicator of the number of processed items of the task. For more complex tasks (such as batch data processing), custom attributes can also be used to record the specific number of processed items. For example, when executing ExecuteSQL, the sql data set currently scheduled for processing will be returned. The returned data set can be obtained through ResultSet, and the number of processed items of this scheduling can be recorded based on the data set to record the amount of processed data.

[0141] In one embodiment, according to the actual running situation of the NiFi processor scheduling execution, the execution result of the scheduling is updated and recorded. Each NiFi processor will return an execution result when executing, usually including status information such as task success or failure, error logs, etc., to record the scheduling operation result.

[0142] Optionally, in this step, the scheduling operation data can be converted into database entity fields and classified and stored in the database according to business requirements for subsequent centralized analysis of the data. Business database entity fields: refer to the fields defined in the business system, which are the specific units for storing data. These fields usually map to columns in the database table. The conversion process means converting the scheduling operation data into a format and fields suitable for storage in the business system according to the requirements of business logic. For example, in the scheduling data, there is information such as the start time and end time of a task. These data need to be mapped to the table structure of the business system and converted into specific fields such as "task start time" and "task end time". Exemplarily, data conversion can be achieved by writing a conversion program or using an ETL (Extract, Transform, Load) tool. An ETL tool can extract data from different data sources, perform necessary conversions (such as formatting, unit conversion, field mapping, etc.), and then load the converted data into the target database. Business requirements refer to different classification and storage methods for data according to business scenarios and needs. For example, it may be necessary to classify by task type, time period, execution status, etc. Classifying and storing in the database means storing the data in the database after classifying it according to specific rules. This process usually involves techniques such as data partitioning, tagging, and classification to manage and query more efficiently. For example, it can be achieved through database table design, such as designing different tables for different data types or using fields in the same table to identify data types (for example, using the field type to identify the task type). Centralized analysis means that after archiving, these data are centralized for analysis. The archived data can be further processed through means such as data analysis, data mining, and machine learning to obtain valuable business insights or fault diagnosis.

[0143] For example: Suppose we have a scheduling operation data, including task ID, task start time, task end time, execution status, resources used, etc. Scheduling operation data: Task ID: 12345; Start time: 2024-11-27 10:00; End time: 2024-11-27 10:30; Execution status: Success; Resources used: 2-core CPU, 4GB memory.

[0144] Converted into business database entity fields:

[0145] Suppose we have a business database table task_execution, which contains the following fields:

[0146] task_id (Task ID); start_time (Start Time); end_time (End Time); status (Execution Status); cpu_cores (Number of CPU Cores); memory_gb (Memory Size). After transforming the scheduling data, the data may be inserted into a table through a program or an ETL tool.

[0147] Optionally, Figure 4 In the data processing method, after S103, S104 may also be executed. That is to say, S104 is an optional step and S104 may not be executed. Figure 4 The steps shown by the dashed line in are optional steps.

[0148] S104. Execute the subsequent data processing tasks of the data flow of the same scheduling task according to the scheduling operation data, or execute the scheduling tasks under another scheduling policy after removing the current scheduling task.

[0149] It should be noted that in a scheduling system of a NiFi node, multiple scheduling tasks may be executed in series or in parallel. The execution result of each scheduling task serves as the input for the subsequent task, and the system needs to dynamically execute the subsequent tasks according to the task dependency relationship. Queue Mechanism: The triggering of subsequent tasks can be achieved through Message Queue technology. When a task is completed, it will put the result into the queue and wait for the downstream task to process it. Conditional Trigger: After the task is executed, it is determined whether to execute the subsequent task based on the execution result (for example, whether the task is successful, the execution time, etc.).

[0150] For example, assume we have a data processing task flow: Task 1: Data preprocessing (successfully completed); Task 2: Data analysis (depends on the output of Task 1); Task 3: Result storage (depends on the output of Task 2). If Task 1 is successfully completed, then Task 2 will be automatically executed, and Task 3 depends on the output of Task 2 to be executed. If Task 1 fails, then neither Task 2 nor Task 3 will be executed. In this way, by obtaining the scheduling execution status of the NiFi processor Processor, the subsequent job process can be triggered in a timely manner.

[0151] In one embodiment, it is also possible to trigger an external program process or task according to requirements (for example, by executing an external script, calling an API interface, or starting an external application). These external processes are other systems or services related to data stream processing. For example, when a certain data processing task is completed, the NiFi node can call an external service through an HTTP (hypertext transfer protocol) request, or trigger the next-stage task in an ETL (Extract, Transform, Load) process. For example, when the execution result of the NiFi processor for processing the data stream is successful, the subsequent task can be to transmit or copy the data stream to a terminal device to complete data synchronization processing and ensure data consistency between different systems.

[0152] In one embodiment, if the execution result of the NiFi processor for processing the data stream is a failure, the current scheduling task of the NiFi processor can also be paused.

[0153] In one embodiment, when the execution result of the NiFi processor for processing the data stream is a failure, a fault warning prompt, log record, etc. are issued. In addition, if the processor fails or the execution fails, the task can also be automatically retried, and a notification is sent to the administrator to perform an error handling process.

[0154] And to execute the scheduling task under another scheduling policy after removing the current scheduling task can be achieved through the following steps:

[0155] Continue to refer to Figure 5 , during the running of the processor, step S207 is executed after step S206.

[0156] S207. Judge the current scheduling status.

[0157] In one case, if the scheduling status is Stopped or Stopping, then step S208 is executed; if the status is Starting, then step S209 is executed.

[0158] S208. Terminate the scheduling task.

[0159] In this step, if the current scheduling status is STOPED / STOPING, it means that the scheduling task of the NiFi processor has been completed, or continue to wait for the next scheduling cycle, then terminate the scheduling task. S209. Execute the callback method trigger().

[0160] In this embodiment, the secondary judgment of the scheduling status: Determine whether the current scheduling status is STARING. If it is STARING, it means that the NiFi processor is preparing to start the scheduling task at this time and has completed some initialization, resource or environment configuration work. At this time, the scheduling task is in the transition state of starting, so start the scheduling and adopt the callback method of the scheduling task, that is, execute the callBack callback method trigger(). Then schedule the NiFi processor again to process the data stream. It can be understood that trigger() is the callBack callback method, and its function is to let the NiFi processor perform actual work according to the scheduling task. That is, the callback method trigger() can be regarded as the trigger for the scheduling task to schedule the NiFi processor to execute its main function.

[0161] S210. Execute the scheduling task regularly according to the scheduling policy.

[0162] In this step, to create a scheduling task, quartz Schedule (timed scheduling module) can be used for timed scheduling. For example, the scheduling policy of a NiFi processor can be set as timed scheduling using the timed scheduling module. For example, if the scheduling interval is set to 10 minutes, the timed scheduling module will call the onTrigger() method of this processor every 10 minutes to execute the scheduling task regularly according to the scheduling policy.

[0163] It should be noted that during the execution of the scheduling task, ScheduleRunCondation (scheduling operation status perception module) is used to record the running situation and save it into the database. In this way, the ScheduleRunCondation module persists the running situation into the data warehouse to ensure the consistency and persistence of the data.

[0164] Optionally, Figure 4 In the data processing method, after S104, S105 can also be executed. That is to say, S105 is an optional step and S105 can be not executed. Figure 4 The steps shown by the dotted line in are optional steps.

[0165] S105. Analyze the data processing result of the NiFi processor based on the scheduling operation data.

[0166] In this step, the data processing result of the NiFi processor scheduled can be displayed in real time through a report; and / or the data processing situation of the NiFi processor scheduled can be analyzed in real time through a statistical chart; and / or the data processing situation of the NiFi processor scheduled can be analyzed in real time through a mathematical model. Specifically, reference can be made to the above for Figure 3A description of possible ways to perform subsequent analysis on the scheduling operation data in the repository will not be elaborated here.

[0167] Next, a data processing method provided in a specific embodiment of the present application will be described with reference to the accompanying drawings.

[0168] Embodiment 1

[0169] Exemplarily, Figure 6 shows a flowchart of another data processing method provided in an embodiment of the present application. As Figure 4 shown, this method is applied to a NiFi cluster and includes:

[0170] S301. Start the processor.

[0171] For detailed content, refer to the description of step S101.

[0172] S302. Create a Scheduler task schedule.

[0173] In this step, the creation of the schedule is started and executed through the scheduling policy configured by the NiFi processor. The scheduling policy refers to how often the processor executes. It can be specified by setting a cron expression on the page to specify the timing interval, and the control of the schedule is carried out through a third-party component Quartz.

[0174] For example: For a Processor of the ExetuceSQL type, if the set scheduling policy is to execute once every 30 minutes, then the Quartz component will execute the onTrigger() method (scheduling task method) of the Processor once every 30 minutes for data processing.

[0175] Other processors include PutSQL, ConsumeMQTT, SplitJSON, etc. The scheduling execution method is the same as above. For example, create a Scheduler task schedule.

[0176] S303. Execute the Scheduler scheduling task.

[0177] In this step, when the execution time of the scheduling trigger arrives, the schedule is executed for data processing. For example, execute the Scheduler scheduling task.

[0178] S304. Record the start time.

[0179] In this step, the real-time time at the start and end of the current scheduling execution is obtained through the code System.currentTimeMillis(). For example, record the start time and the end time.

[0180] S305. Record the data processing volume.

[0181] In this step, obtain the number of data processed to record the data processing volume. Exemplarily, the total number of data processed can be obtained by performing real-time recording and summarization of the data processing process. Specifically, when executing a scheduling task, such as a Processor of the ExetuceSQL type, it will return the sql data set currently being scheduled and processed, and the number of processes scheduled this time can be recorded based on the data set to record the data processing volume.

[0182] S306. Record the end time.

[0183] For the detailed content, please refer to the description of step S304.

[0184] S307. Record the scheduling operation result.

[0185] In this step, update the scheduling execution result according to the actual operation situation of the scheduling execution to record the scheduling operation result.

[0186] S308. Save the record to the database.

[0187] In this step, convert the data into database entity fields, classify and store them in the database according to business requirements, and save the record to the database for subsequent data analysis.

[0188] It should be noted that steps S304 - S308 can be implemented through the ScheduleRunCondation module. The ScheduleRunCondation module realizes the acquisition and storage of scheduling operation status data based on Aop.

[0189] In this way, when NiFi starts the processor, obtain the start time before the scheduling execution, obtain the end time after the scheduling execution, obtain the number of data currently being scheduled and processed during the execution, update the status of this scheduling after the scheduling ends, and save the start time, end time, data volume, and status of the scheduling execution to the data warehouse, which can monitor the operation of NiFi processors in the NiFi cluster, thereby facilitating the triggering of subsequent job processes according to the operation of NiFi processors and improving the job processing efficiency.

[0190] It should be understood that the sequence numbers of the steps in the above embodiments do not indicate the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. In addition, in some possible implementation manners, the steps in the above embodiments may be selectively executed according to actual situations, may be partially executed, or may be fully executed, which is not limited herein. Additionally, all or part of any feature of any of the above embodiments may be freely combined arbitrarily without conflict; the combined technical solutions are also within the scope of the present application.

[0191] As Figure 7 shown, an embodiment of the present application further provides a data processing device 500, including: an acquisition module 501 and a processing module 502. Among them, the acquisition module 501 is used to acquire a scheduling task; wherein, the scheduling task indicates that the NiFi node schedules the NiFi processor to process the data stream. The processing module 502 is used to schedule the NiFi processor to execute the scheduling task. The acquisition module 501 is further used to acquire the scheduling operation data of the NiFi processor during the process of the NiFi processor executing the scheduling task; and store the scheduling operation data of the NiFi processor.

[0192] In a possible implementation manner, the processing module 502 is further used to start the NiFi processor and acquire the scheduling state of the NiFi processor; when the scheduling state of the NiFi processor is STOPED, modify the scheduling state of the NiFi processor to STARTING, and activate the current thread task of the NiFi processor; and process the data stream according to the scheduling task configured by the NiFi processor.

[0193] In a possible implementation manner, the processing module 502 is further used to remove the current thread task of the NiFi processor; and when the scheduling state of the NiFi processor is STARING, use the callback method of the scheduling task to schedule the NiFi processor to process the data stream again.

[0194] In a possible implementation manner, the processing module 502 is further used to analyze the data processing result of the NiFi processor according to the scheduling operation data.

[0195] In a possible implementation manner, the processing module 502 analyzes the data processing result of the NiFi processor, specifically used for: real-time displaying the data processing result scheduled by the NiFi processor in the form of a report; and / or real-time analyzing the data processing situation scheduled by the NiFi processor in the form of a statistical chart; and / or real-time analyzing the data processing situation scheduled by the NiFi processor in the form of a mathematical model.

[0196] In a possible implementation, the scheduling operation data includes the start time, end time, data volume of the processed data, and execution result of the NiFi processor executing the data processing task. The obtaining module 501 is further configured to obtain the scheduling operation data of the NiFi processor executing the data processing task, specifically: obtain the time stamp when the NiFi processor starts or ends executing the scheduling task, and obtain the start time or end time of the NiFi processor executing the data processing task.

[0197] In a possible implementation, the obtaining module 501 is further configured to obtain the scheduling operation data of the NiFi processor executing the data processing task, specifically: obtain the data set of the data type currently scheduled and processed returned when the NiFi processor executes the scheduling task, and record the number of processed items of this scheduling based on the data set to record the data volume of the processed data.

[0198] In a possible implementation, the obtaining module 501 is further configured to obtain the scheduling operation data of the NiFi processor executing the data processing task, specifically: obtain the execution result of the NiFi processor scheduling operation according to the successful, failed, timed out or abnormal execution status returned after the NiFi processor executes the scheduling task, and update and record the execution result of the scheduling according to the actual operation situation of the NiFi processor scheduling execution.

[0199] In a possible implementation, the processing module 502 is further configured to execute subsequent data processing tasks of the data stream of the same scheduling task according to the scheduling operation data, or execute a scheduling task under another scheduling policy after removing the current scheduling task.

[0200] In a possible implementation, the processing module 502 is further configured to execute subsequent data processing tasks of the data stream of the same scheduling task according to the scheduling operation data, specifically: when the execution result of the NiFi processor processing the data stream is successful, transmit or copy the data stream to the terminal device; and / or when the execution result of the NiFi processor processing the data stream is failed, pause the current scheduling task of the NiFi processor; and / or when the execution result of the NiFi processor processing the data stream is failed, issue a fault warning prompt.

[0201] It should be understood that both the obtaining module 501 and the processing module 502 can be implemented by software or can be implemented by hardware. Exemplarily, next, taking the obtaining module 501 as an example, the implementation manner of the obtaining module 501 is introduced. Similarly, the implementation manner of the processing module 502 can refer to the implementation manner of the obtaining module 501.

[0202] As an example of a software functional unit, the obtaining module 501 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the computing instance may be one or more. For example, the obtaining module 501 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running the code may be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Usually, one region may include multiple AZs.

[0203] Similarly, the multiple hosts / virtual machines / containers for running the code may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is achieved through the communication gateway.

[0204] As an example of a hardware functional unit, the obtaining module 501 may include at least one computing device, such as a server, etc. Alternatively, the obtaining module 501 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0205] The multiple computing devices included in the obtaining module 501 may be distributed in the same region or in different regions. The multiple computing devices included in the obtaining module 501 may be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the obtaining module 501 may be distributed in the same VPC or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0206] Based on the same inventive concept, the principle and beneficial effect of the server provided in the embodiments of the present application for solving problems can be referred to the principle and beneficial effect of the method implementation. For the sake of brevity of description, it will not be repeated here.

[0207] Such as Figure 8As shown in the figure, an embodiment of the present application further provides a computing device 600. Exemplarily, the computing device 600 may be a server or a terminal. When the computing device 600 runs, the computing device 600 may execute the method in the above embodiment.

[0208] The computing device includes: a bus 601, a processor 602, a memory 603, and a communication interface 604. The processor 602, the memory 603, and the communication interface 604 communicate with each other through the bus 601. The computing device 600 may be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 600.

[0209] The bus 601 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only one line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. The bus 601 may include a path for transmitting information between various components of the computing device 600 (for example, the processor 602, the memory 603, and the communication interface 604).

[0210] The processor 602 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0211] The memory 603 may include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0212] The memory 603 stores executable program instructions, and the processor 102 executes the executable program instructions to respectively implement the test methods involved in the above embodiments.

[0213] The communication interface 604 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement the communication between the computing device 600 and other devices or communication networks.

[0214] The embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0215] As Figure 9 shown, the computing device cluster includes at least one computing device 600. Instructions for executing a data processing method that are the same can be stored in the memories 603 of one or more of the computing devices 600 in the computing device cluster.

[0216] In some possible implementation manners, partial instructions for executing a data processing method can also be stored separately in the memories 603 of one or more of the computing devices 600 in the computing device cluster. In other words, a combination of one or more computing devices 600 can jointly execute instructions for executing a data processing method.

[0217] It should be noted that the memories 603 in different computing devices 600 in the computing device cluster can store different instructions, which are respectively used to execute partial functions of the above-mentioned multiple modules. That is, the instructions stored in the memories 603 of different computing devices 600 can implement the functions of one or more of the above-mentioned multiple modules.

[0218] In some possible implementation manners, one or more computing devices in the computing device cluster can be connected through a network. Among them, the network can be a wide area network or a local area network, etc. Figure 10 Shows a possible implementation manner. As Figure 10 shown, two computing devices, namely computing device 600A and computing device 600B, are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation manner, instructions for executing the functions of partial modules among the above-mentioned multiple modules are stored in the memory 603 of computing device 600A. At the same time, instructions for executing the functions of another part of the above-mentioned multiple modules are stored in the memory 603 of computing device 600B.

[0219] Figure 10 The connection manner between the computing device clusters shown can be considered because the data processing method provided in the present application requires a large amount of data storage. Therefore, it is considered to hand over the functions implemented by another part of the above-mentioned multiple modules to computing device 600B for execution.

[0220] It should be understood that Figure 10The functions of the computing device 600A shown can also be completed by multiple computing devices 600. Similarly, the functions of the computing device 600B can also be completed by multiple computing devices 600.

[0221] The embodiments of the present application also provide another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similarly referred to Figure 9 and Figure 10 the connection method of the computing device cluster. The difference is that the same instructions for executing the data processing method can be stored in the memory 603 of one or more of the computing devices 600 in the computing device cluster.

[0222] In some possible implementation manners, partial instructions for executing the data processing method can also be separately stored in the memory 603 of one or more of the computing devices 600 in the computing device cluster. In other words, the combination of one or more computing devices 600 can jointly execute the instructions for executing the data processing method.

[0223] It should be noted that the memories 603 in different computing devices 600 in the computing device cluster can store different instructions for executing partial functions of the computing device 600. That is, the instructions stored in the memories 603 of different computing devices 600 can implement the functions of one or more of the above-mentioned multiple modules.

[0224] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium is used to store computer program instructions. When the computer program instructions run on a computing device, the computing device is enabled to execute the data processing method involved in the above embodiments. The computer-readable storage medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc.

[0225] The embodiments of the present application also provide a computer program product containing instructions. The computer program product can be a software or program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, at least one computing device is enabled to execute the data processing method.

[0226] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application. Those of ordinary skill in the art should understand that although the present application has been described in detail with reference to the foregoing embodiments, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions in the embodiments of the present application.

Claims

1. A data processing method, characterized in that: Applied to a NiFi cluster, the NiFi cluster includes multiple NiFi nodes, and the method is executed by any one or more NiFi nodes among the multiple NiFi nodes, including: Obtain a scheduling task; wherein the scheduling task instructs the NiFi node to schedule the NiFi processor to process the data stream; Scheduling the NiFi processor to execute the scheduling task; In the process of the NiFi processor executing the scheduling task, obtaining the scheduling operation data of the NiFi processor; and Stores the scheduling and running data of the NiFi processor.

2. The method according to claim 1, characterized in that The scheduling of the NiFi processor to execute the scheduling task includes: Start the NiFi processor and obtain the scheduling status of the NiFi processor; When the scheduling state of the NiFi processor is STOPED, the scheduling state of the NiFi processor is changed to STARTING, and the current thread task of the NiFi processor is activated; The data stream is processed according to the scheduling task configured by the NiFi processor.

3. The method according to claim 2, characterized in that After the scheduling task configured according to the NiFi processor executes the processing of the data stream, it also includes: Remove the current thread task of the NiFi processor; And when the scheduling state of the NiFi processor is STARING, the callback method of the scheduling task is adopted to schedule the NiFi processor again to process the data flow.

4. The method according to any one of claims 1 to 3, characterized in that: The scheduling operation data includes the start time, end time, amount of processed data and execution result of the data processing task executed by the NiFi processor.

5. The method according to claim 4, characterized in that The obtaining of the scheduling operation data of the NiFi processor to perform the data processing task includes: Get the timestamp when the NiFi processor starts or ends executing the scheduling task, and get the start time or end time when the NiFi processor executes the data processing task.

6. The method according to claim 4, characterized in that The obtaining of the scheduling operation data of the NiFi processor to perform the data processing task includes: Get the data set of the data type currently scheduled and processed returned when the NiFi processor executes the scheduling task, and record the number of processed items of this scheduling based on the data set to record the amount of processed data.

7. The method according to claim 1, characterized in that The obtaining of the scheduling operation data of the NiFi processor to perform the data processing task includes: According to the execution status of success, failure, timeout or exception returned by the NiFi processor after executing the scheduling task, the execution result of the NiFi processor scheduling operation is obtained.

8. The method according to any one of claims 1 to 7, characterized in that: After obtaining the scheduling operation data of the NiFi processor, the method further includes: According to the scheduling operation data, subsequent data processing tasks of the data stream of the same scheduling task are executed, or a scheduling task under another scheduling strategy is executed after removing the current scheduling task.

9. The method according to claim 8, characterized in that The subsequent data processing task of executing the data stream of the same scheduling task according to the scheduling operation data includes: When the NiFi processor successfully processes the data stream, the data stream is transmitted or copied to the terminal device; and / or When the execution result of the NiFi processor processing the data stream is a failure, suspending the current scheduling task of the NiFi processor; and / or When the NiFi processor fails to process the data flow, a fault alarm is issued.

10. A computing device, characterized in that: include: at least one memory for storing a program; at least one processor, configured to execute the program stored in the memory; The memory is coupled to the processor, and when the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1 to 9.