Multi-source heterogeneous data task arrangement method and system based on process engine

Through the multi-source heterogeneous data task orchestration method based on the process engine, the problems of complex multi-source heterogeneous data access and insufficient task orchestration flexibility are solved, the hybrid orchestration and full-link tracking of batch and stream processing are realized, and the reusability and maintainability of data services are improved.

CN120762867AActive Publication Date: 2025-10-10CIVIL AVIATION CARES OF XIAMEN LTD

Patent Information

Application Number
CN202511278112.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-10-10
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

When processing multi-source heterogeneous data, existing data integration and development platforms have problems such as complex multi-source heterogeneous data access, insufficient task orchestration flexibility, limited service and governance, and poor scalability and maintainability.

Method used

It adopts a multi-source heterogeneous data task orchestration method based on a process engine, accesses multi-source heterogeneous data through pluggable data source adapter components, supports batch and stream processing tasks, has flexible task orchestration and full-link tracking capabilities, and realizes result service output. It uses BPMN process modeling, Groovy scripts and plug-in extension mechanisms to achieve low-code development.

Benefits of technology

It realizes hybrid orchestration and dynamic control, supports hybrid orchestration of batch processing and stream processing in the same process model, improves the flexibility and observability of task modeling, and enhances the reusability and maintainability of data services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120762867A_ABST
    Figure CN120762867A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source heterogeneous data task arrangement method and system based on a process engine, and the method comprises the steps: accessing multi-source heterogeneous data through a pluggable data source adaption assembly, and converting the data into a unified internal data model; process modeling and task arrangement, wherein a process model comprises a start event node, a batch processing task node, a stream processing task node, a gateway node and an end event node; the limitation of a traditional static FGA model is broken through, and mixed arrangement of batch processing and stream processing in the same process model and complex process control based on parallel and containing gateways are supported; task whole-process tracking is realized through a traceId mechanism, so that fault positioning and operation monitoring are facilitated; flexible self-definition of processing logic is realized through SDK plug-in and Groomy script support, and developers do not need to write codes on a large scale; and any task process result can be quickly packaged into REST API, so that the data service reusability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data integration and development technology, and in particular to a multi-source heterogeneous data task orchestration method and system based on a process engine, which is suitable for accessing, converting, scheduling, tracking and service-oriented output of multi-source heterogeneous data. Background Art

[0002] In the civil aviation information technology sector (such as security inspection systems and baggage systems), large amounts of data from heterogeneous sources and complex structures exist, requiring batch processing, stream processing, cleaning, association, storage, and service-oriented packaging. However, existing data integration and development platforms (such as Apache NiFi and SeaTunnel) typically use static directed acyclic graphs (DAGs) as task scheduling models, which have the following shortcomings: Complex access to multi-source heterogeneous data: Different types of data sources vary greatly, often requiring developers to manually write a large amount of adaptation code.

[0003] Insufficient task scheduling flexibility: The static DAG model cannot support complex branching, parallelism, conditional judgment, and mixed scheduling.

[0004] Limited service-oriented and governance: The data processing process lacks a full-link tracking mechanism, and the interface encapsulation and operation monitoring capabilities are insufficient.

[0005] Poor scalability and maintainability: The lack of low-code or plug-in extension mechanisms makes it difficult to adapt to the rapid changes in complex businesses. Summary of the Invention

[0006] The purpose of the present invention is to provide a multi-source heterogeneous data task orchestration method based on a process engine, which can simultaneously support batch processing and stream processing tasks, has flexible task orchestration and full-link tracking capabilities, and can realize data task orchestration method and system for service-oriented output of results.

[0007] To achieve the above objectives, the present invention provides a solution: a multi-source heterogeneous data task orchestration method based on a process engine, including the following process: Data source access and adaptation: Access multi-source heterogeneous data through pluggable data source adapter components and convert the data into a unified internal data model; Process modeling and task orchestration: The process model includes a start event node, a batch task node, a stream task node, a gateway node, and an end event node, wherein the start event node is configured as a timer event or a stream event to trigger subsequent batch tasks or stream tasks respectively; In the processing task node, the data output configuration is as follows: directly output the batch data to the next node or to the user-configured target data source; or, first use Groovy scripts to process the batch data according to the agreed interface, and then output it to the next node or target data source; The gateway nodes include parallel gateways and inclusive gateways, which are used to perform parallel distribution, synchronous aggregation or conditional routing on task execution paths; Link tracking: A unique trace ID (traceId) is generated at the beginning of each process. The traceId is transparently transmitted along the nodes throughout the entire process, enabling full-link tracking and monitoring. Task execution and flow: The process engine schedules task nodes and gateway nodes for execution according to the process model, and transmits data between nodes through the system's internal message mechanism; Result output and service-oriented: After the process is completed, the task processing results are output to the specified target data source and then published as a REST API service through configuration.

[0008] Furthermore, the data source adapter component supports JDBC database and MQ message queue.

[0009] Furthermore, the batch task node supports JDBC data source type, MQ-PULL data source type and batch plug-in type: JDBC data source type: configure JDBC data source and SQL statement; MQ data source type: Configure the MQ data source that supports the PULL mode, as well as the topic, queue, and batch size parameters. Batch plug-in type: Users develop and upload Java plug-ins based on the SDK provided by the system to configure batch plug-in task nodes.

[0010] Furthermore, if the output target is a JDBC data source, you need to configure the JDBC data source information and write an insert statement to support O{json path} expression fills the JSON attribute value; If the output target is an MQ data source, configure the MQ data source information and the topic or queue name.

[0011] Furthermore, the stream processing task node supports the MQ-PUSH data source type and stream processing plug-in type: MQ-PUSH data source type: configure the MQ data source, topic, and queue parameters; Stream processing plug-in type: Users develop Java plug-ins based on the SDK provided by the system.

[0012] Furthermore, the gateway nodes specifically include the following types: Parallel branch gateway: receives one task and outputs it to multiple parallel tasks; Parallel Merging Gateway: connects to multiple parallel tasks, aggregates all task results, and then outputs them to the next task; Inclusive branch gateway: receives a task, divides it into multiple branches, and selects a branch to execute the task based on conditions; Inclusive merging gateway: accesses multiple tasks, aggregates the actual executed task results from the inclusive branch, and then enters the subsequent tasks.

[0013] Furthermore, the traceId supports operation logs, alarms and lineage analysis.

[0014] Furthermore, after the process is completed, an operation report and status events are automatically generated for users to query.

[0015] The present invention also provides a multi-source heterogeneous data task orchestration system based on a process engine, comprising: Data access layer, used to adapt to multiple heterogeneous data sources and convert data into internal data models; The process orchestration and control layer is used to receive the process model configured by the user and schedule and control the task nodes and gateway nodes in the process; The task processing and computing layer is used to execute the batch processing tasks and stream processing tasks scheduled by the process orchestration and control layer; The data flow and monitoring layer is used to transmit data between task nodes, provide full-link tracking and monitoring functions, and generate operation reports; The result servitization layer is used to output task processing results and publish them as service interfaces.

[0016] Furthermore, the data access layer includes: Data access module: responsible for the adaptation and abstraction of multi-source heterogeneous data sources, uniformly converting them into internal data models, and shielding the underlying differences; Plugin / extension interface: Provides SDK interface for users to extend custom data sources or processing logic; The process orchestration and control layer includes: Process engine module: Based on the BPMN process modeling concept, it realizes the arrangement of task nodes, branch gateways, and conditional controls; Scheduling and execution control submodule: supports batch processing, stream processing, concurrent execution, and retry fault tolerance; The task processing and computing layer includes: Task executor module: includes batch task executors and stream task executors, and supports outputting task processing results to specified target data sources; Dynamic script and expression module: supports expression and script parsing; The data flow and monitoring layer includes: Message middleware module: responsible for data transmission between nodes, result caching and transparent transmission of traceId; Monitoring and governance module: provides full-link traceId tracking, operation log collection, alarm detection, and data lineage analysis; The result servitization layer includes: Service-oriented configuration module: supports outputting data to a specified target data source after task processing, and automatically publishing it as a REST API service through configuration; Report and Notification submodule: automatically generates operation reports and task completion notifications.

[0017] After adopting the above solution, the beneficial effects of the present invention are: 1. Hybrid orchestration and dynamic control: This breaks the limitations of the traditional static DAG model and supports hybrid orchestration of batch and stream processing in the same process model, as well as complex process control based on parallel and inclusive gateways, improving the flexibility of task modeling.

[0018] 2. Full-link observability: The traceId mechanism enables full-process task tracking, facilitating fault location and operation monitoring, and improving observability and reliability.

[0019] 3. Low-code extension: Supports SDK plug-ins and Groovy scripting to achieve flexible customization of processing logic, without the need for developers to write large-scale code.

[0020] 4. Enhanced output of service capabilities: Any task process result can be quickly encapsulated as a REST API to improve the reusability of data services.

[0021] 5. Governance and lineage analysis: It enables cross-task data lineage tracking, facilitating data quality management and compliance auditing. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a structural diagram of a multi-source heterogeneous data task orchestration system based on a process engine according to an embodiment of the present invention; Figure 2 This is a diagram of each gateway node in an embodiment of the present invention; Figure 3 This is an example diagram of a parallel branch gateway according to an embodiment of the present invention; Figure 4 This is an example diagram of a parallel merging gateway according to an embodiment of the present invention; Figure 5 This is an example diagram of an inclusive branch gateway according to an embodiment of the present invention; Figure 6 This is an example diagram of an inclusive merging gateway according to an embodiment of the present invention; Figure 7This is an example diagram of a multi-source heterogeneous data task orchestration process based on a process engine according to an embodiment of the present invention; Figure 8 This is an example flow chart of the first embodiment of the present invention; Figure 9 This is an example flow chart of the second application embodiment of the present invention. DETAILED DESCRIPTION

[0023] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0024] The present invention provides a multi-source heterogeneous data task scheduling system based on a process engine, including the following modules (refer to Figure 1 ): Data access layer Data access module: responsible for the adaptation and abstraction of multi-source heterogeneous data sources (JDBC database, message queue MQ, etc.), unified conversion into internal data models, and shielding the underlying differences.

[0025] Plugin / extension interface: Provides SDK interface, allowing users to extend custom data sources or processing logic.

[0026] Process orchestration and control layer Process engine module: core module, based on BPMN / process modeling ideas, realizes the arrangement of task nodes, branch gateways, conditional control, etc. Scheduling and execution control submodule: supports batch processing tasks (CRON timing), stream processing tasks (real-time triggering), concurrent execution and retry fault tolerance.

[0027] Task processing and computing layer Task executor module: contains two types of task executors: batch task executor (JDBC / MQ-PULL / plug-in task); stream processing task executor (MQ-PUSH / plug-in task); Dynamic script and expression module: support {parameter}, Expressions such as O{json path} and Groovy script parsing are used for flexible data processing rules.

[0028] Data flow and monitoring layer Message middleware module: responsible for data transmission between nodes, result caching and transparent transmission of traceId; Monitoring and governance module: provides full-link traceId tracking, operation log collection, alarm detection, and data lineage analysis.

[0029] Result servitization layer Service-oriented configuration module: supports outputting task processing results to databases and MQ message queues, and automatically publishing them as REST API services through configuration; Report and Notification submodule: automatically generates operation reports and task completion notifications.

[0030] The present invention is based on a multi-source heterogeneous data task orchestration method based on a process engine of the above system, which includes the following process: <Data source access and adaptation> Multi-source access: Provides pluggable data source adapter components. The system currently supports a variety of common data sources such as JDBC databases (such as Oracle, SqlServer, MySQL, OpenGauss), MQ message queues (such as RabbitMQ, RocketMQ, Pulsar), and can continue to develop and adapt new data sources according to project requirements.

[0031] Data standardization: Standardize and abstract the original heterogeneous data sources to shield the underlying differences, that is, convert the data formats of different connected data sources into a unified internal data model.

[0032] <Process Modeling and Dynamic Task Orchestration> The present invention introduces a graphical modeling method based on BPMN (Business Process Model and Notation) to provide users with visual process design capabilities to support dynamic and hybrid orchestration of complex data tasks.

[0033] Visual low-code modeling: The system provides an intuitive visual designer interface, allowing users to build data processing processes by dragging and dropping without having to write underlying scheduling code, significantly lowering the threshold for process design and achieving a low-code development experience.

[0034] Process model components: The process model consists of multiple nodes, and the icons of each node are as follows: Figure 2 These nodes are the basis for hybrid orchestration and complex process control. The specific node types and operating mechanisms are as follows: 1) Start event node (StartEvent): Configure the start event in the designer interface according to actual business needs.

[0035] Start events support timer events and stream events.

[0036] If it is a timer event, you need to configure the timer's CRON (a time rule expression) to trigger the timer event. The timer event can only be connected to the 2) batch event node; If it is a stream event, no need to configure CRON, and the subsequent event can only be connected to the 3) stream processing event node.

[0037] The further subsequent task can be any batch or stream processing event node, that is, only the type of the task node that can be connected to the subsequent start event node is limited.

[0038] 2) Batch processing event node (batch processing Task node): Support JDBC data source type, MQ-PULL data source type and batch processing plug-in type: ①JDBC data source type: configure JDBC data source and SQL statement (select, insert, select into) (query / insert / select and write). The SQL execution result can be specified to be directly output to the next node (suitable for select) or output to the target data source configured by the user (suitable for insert / select into). Before the SQL execution result is output, the SQL execution result can be configured to be first processed using a Groovy script according to the agreed interface, and then output.

[0039] ②MQ data source type: configure MQ data source supporting PULL mode, topic / queue (topic / queue) and batch size (batch size) parameters, which can be understood as: select MQ in PULL mode (such as Kafka consumer, RabbitMQBasic.Get), specify the consumption target through topic / queue, and balance throughput and real-time performance through batchSize.

[0040] The pulled data can be directly output to the next node or output to the target data source configured by the user; similarly, the pulled data can be first processed using a Groovy script according to the agreed interface, and then output.

[0041] ③Batch processing plug-in type: users can develop and upload Java plug-ins based on the SDK provided by the system to configure batch processing plug-in task nodes and expand batch processing capabilities; similarly, support output to the next node or output to the target data source configured by the user.

[0042] If the above output target is a JDBC data source, the JDBC data source information needs to be configured, and the insert statement needs to be written, supporting O{json path} expression fills JSON property value; if the output target is an MQ data source, configure the MQ data source information and the topic or queue name.

[0043] Except for the first task, any processing task node in the process that has a predecessor task will use the output of the predecessor task as the input data source.

[0044] 3) Stream processing event node (stream processing Task node): Support MQ-PUSH data source type, configure MQ data source and topic / queue parameters; The stream data can be configured to be processed using Groovy scripts according to the agreed interface before being output, or it can be directly output to the next node or user-defined target data source.

[0045] Stream processing plug-in type: Users can develop Java plug-ins based on the SDK provided by the system to expand stream processing capabilities.

[0046] Dynamic orchestration capabilities: The process engine dynamically parses and executes the model at runtime. Based on the logical definitions of the gateway nodes (such as parallelism and conditional judgments), it schedules and coordinates the execution order and paths of batch and stream processing tasks in real time, thereby addressing the complex and ever-changing data processing requirements in different business scenarios and achieving true dynamic task orchestration.

[0047] The present invention addresses the problems of poor scalability and maintainability of existing platforms by taking the following measures: Low-code / scripting: Task nodes support the use of Groovy scripts for custom logic processing, without the need for compilation and deployment, and are highly flexible.

[0048] Pluginization: An SDK is provided, allowing users to develop custom batch and stream processing plugins to meet extremely complex customization requirements. The system gains powerful scalability through the dynamic running of plugins.

[0049] Visual orchestration: Drag-and-drop process design through a visual interface reduces the cost of business logic changes.

[0050] 4) Gateway Node: The primary function of a gateway node is to logically control the execution results of batch and stream tasks, determining the direction of subsequent processes. A gateway is a decision point or synchronization point in the process, responsible for dynamically adjusting process branches based on task execution results or data content.

[0051] Gateway nodes include parallel gateways and inclusive gateways, as follows: Parallel Fork supports parallel distribution of tasks: it connects to one task and outputs it to multiple parallel tasks. Figure 3 .

[0052] Parallel Join Gateway: Connects to multiple parallel tasks, aggregates all task results and then outputs them to the next task. Figure 4 In addition, the gateway can merge data streams based on traceId and support timeout strategies to prevent processes from being stuck due to infinite waiting.

[0053] Inclusive Fork Gateway: accepts a task and divides it into multiple branches. It selects branches based on conditions, but does not execute all of them. Figure 5 , supports SELECT + HAVING syntax judgment, for example, Select Max(a)From #dataset[0] Having Max(a)>300; Inclusive Join Gateway: Connects to multiple tasks, aggregates the results of tasks actually executed by all inclusive branches, and then outputs them to the next task. Figure 6 .

[0054] 5) End event node (EndEvent): When the process is marked as finished, the system automatically sends the task completion status to the built-in message queue.

[0055] Below through Figure 7 The following example explains the relationship between the nodes: Start event: timer triggers, enters "batch task 1"; The parallel gateway connects to "batch task 1" and outputs three parallel tasks (stream processing task 2, batch processing task 3, and batch processing task 4). After "Stream Processing Task 2" is completed, enter "Stream Processing Plug-in Task 5"; The inclusive gateway accesses two tasks (batch task 3 and batch task 4) and outputs "batch task 6"; The parallel gateway accesses "stream processing plug-in task 5" and "batch processing task 6" and outputs "batch processing task 7". After "batch processing task 7" is processed, the entire process ends.

[0056] <traceId全链路追踪> Each time a process instance starts executing, that is, when the start event node (StartEvent) is triggered, the system automatically generates a globally unique traceId, which is used to uniquely identify the specific process execution instance.

[0057] When the start event is a timer event, each time the timer triggers, a traceId is generated, identifying a batch processing job instance. When the start event is a stream event, the arrival of each message (or batch) triggers the process and generates a traceId, identifying the processing instance of that message (or batch).

[0058] The generated traceId is transparently transmitted along with the data across all task nodes and gateway nodes throughout the lifecycle of the process instance, enabling full-link tracing. Regardless of whether the process logic involves parallel branching, conditional routing, or synchronous aggregation, all task nodes belonging to the same process instance share the same traceId.

[0059] The traceId, as a context identifier throughout, provides core support for system observability and supports: Run log association: The logs printed by all nodes carry this traceId, which can be used to aggregate and query all logs of a complete process execution, greatly improving troubleshooting efficiency.

[0060] Alarm event location: Any alarm information generated by the system will carry the traceId at that time, so operation and maintenance personnel can quickly locate the specific process instance and its context that caused the alarm.

[0061] Data lineage analysis: By recording traceId, data operations (read / write), and data objects, the system can accurately map the complete flow and transformation of data in the process, supporting upstream traceability and downstream impact analysis.

[0062] The traceId mechanism integrates a decentralized, potentially concurrently executed task processing process into a traceable entity with a unified identifier, achieving transparency and observability of the entire data processing chain.

[0063] <Data Flow and Execution> 1) Data transfer mechanism: Data transmission between task nodes is handled by the system's built-in message middleware. This middleware provides a unified communication channel for nodes. Sending nodes publish processing results (encapsulated as internal data models) to the middleware, from which downstream consuming nodes subscribe and retrieve the data. This mechanism shields the complexities of physical deployment and communication protocols between nodes, achieving decoupled, asynchronous, and buffered data transmission, making it transparent to users.

[0064] 2) Task scheduling and execution mechanism: The process engine has the following responsibilities: Drive process instances: According to the definition of the process model (such as BPMN process), follow the path logic of start events, task nodes, and gateways to schedule the execution order of task nodes.

[0065] Coordinated message consumption: This works in conjunction with the built-in message middleware. When the engine schedules a task node for execution, the node pulls upstream data from the middleware. After execution, it pushes the results back to the middleware for consumption by downstream nodes. The engine indirectly controls the data flow by coordinating the node's message consumption.

[0066] Support concurrent execution: When the process executes to the parallel gateway, the engine will simultaneously schedule the execution of task nodes on multiple branches, making full use of system resources to improve processing efficiency.

[0067] Fault tolerance and retry support: The engine monitors the execution status of each task node. If a node fails, the engine automatically initiates a retry based on pre-set policies (such as fixed intervals and exponential backoff). For tasks that consistently fail, the engine marks the process instance as failed and issues an alert, ensuring system reliability.

[0068] <Result Output and Service> 1) After the process is completed, the system will output the processing results to the target of JDBC data source, MQ data source, etc., and then publish it as a REST API service through configuration.

[0069] 2) After the process is completed, the system automatically generates an operation report and generates corresponding status events for users to query.

[0070] Application Example 1: Fusion Processing of Security Inspection Logs and Alarm Data In airport security business scenarios, there are two types of key data: passenger channel log data from the security system database (including the number of passengers passing through the channel and the time information) and real-time alarm data from the passenger channel alarm message queue.

[0071] In order to timely and accurately detect abnormal operation of the security inspection channel, this embodiment uses the task scheduling method based on the process engine provided by the present invention to fuse the above two types of data, generate fused alarm data, and store it as an alarm report. The process is as follows, refer to the process diagram Figure 8 : 1) StartEvent: Type: CRON timer event (for example, triggered every 10 minutes).

[0072] The system generates a unique traceId to identify a process execution instance.

[0073] 2) Batch Task (JDBC): Configure data source: security system database (Oracle).

[0074] Execute SQL: SELECT FROM passenger_log WHERE ts>last_time; Output result: passenger channel log dataset dataset[0].

[0075] 3) Batch processing Task (MQ-PULL): Configure MQ: alarm message queue (RocketMQ).

[0076] Pull data: configure topic = "SECURITY_ALARM"; batchSize=10.

[0077] Process data using Groovy script, output dataset[1].

[0078] 4) Gateway: merge two-way data; Input: two-way data from dataset[0] and dataset[1].

[0079] Aggregation condition: aggregate results by traceId, generate merged dataset dataset[2].

[0080] 5) Stream processing event (abnormal detection and new alarm generation): Real-time calculation and abnormal detection on dataset[2], for example, judge whether the number of passengers in a certain time period exceeds the threshold; If an exception is detected, a new alarm event is generated; The output result is the merged alarm dataset dataset[3], written to the database alarm_report table.

[0081] 6) EndEvent: mark the end of the process.

[0082] Application Example Two: Airport Passenger Flow Prediction and Security Channel Dynamic Allocation With a 5-minute prediction interval, merge historical channel logs, recent flight plans, external features (weather / holidays), and real-time passenger flow to predict the throughput and expected queue length of each security channel in the next 15-30 minutes, and automatically generate alarms (such as opening / closing channels and personnel allocation) when exceeding SLA (Service Level Agreement, such as the maximum acceptable queue length for passengers passing through security is not more than 10 minutes) or channel occupancy threshold, and output through MQ service.

[0083] Main data sources JDBC: passenger_log (historical channel logs for the past 60 days, including time granularity, number of passengers passing through, etc.) JDBC: flight_schedule (flight schedule for the next two hours, including departure time, gate, load factor, and other fields) JDBC: ext_features (external features such as weather, holidays, major events, etc.) MQ: PASSENGER_TURNSTILE (real-time gate / counter passenger flow events) The execution process is as follows, refer to the process diagram Figure 9 : 1) StartEvent: Type: CRON timer (trigger every 5 minutes); When triggered, a unique traceId is generated to identify this process instance.

[0084] 2) Batch Task (JDBC: Historical Log): Query the passenger_log data for the past 60 days, aggregate it to 5-minute granularity, and generate basic statistical features (time period / weekdays, holidays, flight peak windows, etc.); Output dataset[hist].

[0085] 3) Batch Task (JDBC: flight planning): Extract the flight_schedule for the next two hours and map the flight passenger flow contribution (which can be estimated by seat capacity or aircraft type) by time window (t, t+15, t+30); Output dataset[flt].

[0086] 4) Batch Task (JDBC: external characteristics): Obtain external influencing factors such as weather (rainfall / visibility) and holidays / large-scale events; Output dataset[ext].

[0087] 5) Parallel Join: Merge in parallel to 6) for feature construction; 6) Batch processing plug-in (feature fusion + prediction inference): Perform feature engineering fusion on dataset[hist] / [flt] / [ext]; Load the offline trained model and make predictions (throughput and waiting time) for the next 15 / 30 minutes. Output dataset[pred].

[0088] 7) Batch processing plug-in (real-time passenger flow + prediction correction): Pull the passenger flow count dataset [pred] within a recent window (for example, the last 5 minutes) from MQ: PASSENGER_TURNSTILE to perform nowcasting correction on the prediction results. Align the time window and channel dimensions of dataset [pred] and dataset [rt], and fuse them to obtain the corrected prediction.

[0089] Output dataset[fused].

[0090] 8) Inclusive branch (conditional routing): If the expected waiting time is greater than the target SLA or the channel occupancy rate is greater than 85%, an alarm task will be initiated.

[0091] 9) Batch processing tasks (alarm tasks): Generate alerts and send them to the message queue. 10) Batch processing tasks (persistent result sets) Persist dataset[fused] into security_lane_forecast table for other uses such as dashboards / BI To further illustrate various embodiments, the present invention is provided with accompanying drawings. These drawings form part of the present disclosure and are primarily used to illustrate the embodiments and, in conjunction with the relevant description in the specification, to explain the operating principles of the embodiments. By referring to these drawings, one of ordinary skill in the art will understand other possible embodiments and the advantages of the present invention. The components in the figures are not drawn to scale, and similar reference numerals are generally used to represent similar components.

[0092] At the same time, the directions such as front, back, left, and right involved in this embodiment are only used as a reference for directions and do not represent directions in actual use. In addition, the terms "first", "second", and "third" are used for descriptive purposes only and should not be understood as indicating or implying relative importance.

[0093] The above description is only a preferred embodiment of the present invention and is not intended to limit the design of this case. Any equivalent changes made based on the key design of this case shall fall within the scope of protection of this case.

Claims

1. A multi-source heterogeneous data task orchestration method based on a process engine, characterized in that: The following processes are included: Data source access and adaptation: Access multi-source heterogeneous data through pluggable data source adapter components and convert the data into a unified internal data model; Process modeling and task orchestration: Build processes based on the BPMN model. The process model includes a start event node, a batch task node, a stream task node, a gateway node, and an end event node. The start event node is configured as a timer event or a stream event to trigger subsequent batch tasks or stream tasks, respectively. In the processing task node, data output is configured to: directly output batch data to the next node or to a user-configured target data source; alternatively, Groovy scripts are used to process the data according to the agreed interface before outputting it to the next node or target data source. The gateway nodes include parallel gateways and inclusive gateways, which are used to perform parallel distribution, synchronous aggregation or conditional routing on task execution paths; Link tracking: A unique trace ID (traceId) is generated at the beginning of each process. The traceId is transparently transmitted along the nodes throughout the entire process, enabling full-link tracking and monitoring. Task execution and flow: The process engine schedules task nodes and gateway nodes for execution according to the process model, and transmits data between nodes through the system's internal message mechanism; Result output and service: After the process is completed, the task processing results are output to the specified target data source and can be published as a REST API service through configuration.

2. The multi-source heterogeneous data task orchestration method based on a process engine according to claim 1, characterized in that: The data source adapter component supports JDBC database and MQ message queue.

3. The multi-source heterogeneous data task orchestration method based on a process engine according to claim 1, characterized in that: The batch task node supports JDBC data source type, MQ-PULL data source type and batch plug-in type: JDBC data source type: configure JDBC data source and SQL statement; MQ data source type: Configure the MQ data source that supports the PULL mode, as well as the topic, queue, and batch size parameters. Batch plug-in type: Users develop and upload Java plug-ins based on the SDK provided by the system to configure batch plug-in task nodes.

4. The multi-source heterogeneous data task orchestration method based on a process engine as claimed in claim 3, characterized in that: If the output target is a JDBC data source, you need to configure the JDBC data source information and write an insert statement to support O{json path} expression fills the JSON attribute value; If the output target is an MQ data source, configure the MQ data source information and the topic or queue name.

5. The multi-source heterogeneous data task orchestration method based on a process engine according to claim 1, characterized in that: The stream processing task node supports the MQ-PUSH data source type and stream processing plug-in type: MQ-PUSH data source type: configure the MQ data source, topic, and queue parameters; Stream processing plug-in type: Users develop Java plug-ins based on the SDK provided by the system.

6. The multi-source heterogeneous data task orchestration method based on a process engine according to claim 1, characterized in that: The gateway nodes specifically include the following types: Parallel branch gateway: receives one task and outputs it to multiple parallel tasks; Parallel Merging Gateway: connects to multiple parallel tasks, aggregates all task results, and then outputs them to the next task; Inclusive branch gateway: receives a task, divides it into multiple branches, and selects a branch to execute the task based on conditions; Inclusive merging gateway: accesses multiple tasks, aggregates all task structures executed in the inclusive branch, and then outputs them to the next task.

7. The multi-source heterogeneous data task orchestration method based on a process engine according to claim 1, characterized in that: The traceId supports operation logs, alarms, and lineage analysis.

8. The multi-source heterogeneous data task orchestration method based on a process engine according to claim 1, characterized in that: After the process is completed, an operation report and status events are automatically generated for users to query.

9. A multi-source heterogeneous data task orchestration system based on a process engine, characterized by: include: Data access layer, used to adapt to multiple heterogeneous data sources and convert data into internal data models; The process orchestration and control layer is used to receive the process model configured by the user and schedule and control the task nodes and gateway nodes in the process; The task processing and computing layer is used to execute the batch processing tasks and stream processing tasks scheduled by the process orchestration and control layer; The data flow and monitoring layer is used to transmit data between task nodes, provide full-link tracking and monitoring functions, and generate operation reports; The result servitization layer is used to publish the output of task processing results as a service interface through configuration.

10. The multi-source heterogeneous data task orchestration system based on a process engine according to claim 9, characterized in that: The data access layer includes: Data access module: responsible for the adaptation and abstraction of multi-source heterogeneous data sources, converting them into internal data models and shielding underlying differences; Plugin / extension interface: Provides SDK interface for users to extend custom data sources or processing logic; The process orchestration and control layer includes: Process engine module: Based on the BPMN process modeling concept, it realizes the arrangement of task nodes, branch gateways, and conditional controls; Scheduling and execution control submodule: supports batch processing, stream processing, concurrent execution, and retry fault tolerance; The task processing and computing layer includes: Task executor module: includes batch task executor and stream task executor; Dynamic script and expression module: supports expression and script parsing; The data flow and monitoring layer includes: Message middleware module: responsible for data transmission between nodes, result caching and transparent transmission of traceId; Monitoring and governance module: provides full-link traceId tracking, operation log collection, alarm detection, and data lineage analysis; The result servitization layer includes: Service-oriented configuration module: supports outputting task processing results to the target data source and automatically publishing them as REST API services through configuration; Report and Notification submodule: automatically generates operation reports and task completion notifications.

Citation Information

Patent Citations

  • Workflow scheduling method and device

    CN114169801A

  • Business arrangement method based on Flowable workflow engine

    CN115185496A

  • Data processing task cooperative control scheduling method and system

    CN115509721A

  • Data lake system based on collaborative scheduling processing of stream data and batch data

    CN115599524A

  • Data processing method and device, equipment and storage medium

    CN116719513A

Cited By

  • Heterogeneous data source processing method and device and electronic equipment

    CN121187566A