A real-time stream computing and service scheduling method and system based on a distributed architecture platform

By integrating Apache Airflow, StreamEngineParser and DolphinDB on a distributed architecture platform, efficient real-time stream computing and service scheduling are achieved, solving the problems of low resource utilization and large computing latency in the existing technology, and improving data processing efficiency and service response speed.

CN119718582BActive Publication Date: 2025-06-27HARBIN AEROSPACE STAR DATA SYST TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411869281.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-06-27
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

In the prior art, the distributed architecture platform has low resource utilization, single scheduling strategy, lack of dynamic adjustment mechanism in terms of service scheduling and real-time flow calculation, and the inability of traditional computing models to effectively deal with large-scale continuous arrival data flows, resulting in increased computing delays and unable to meet real-time requirements.

Method used

Using real-time stream computing and service scheduling methods based on the distributed architecture platform, the efficient, flexible and scalable real-time data processing and service scheduling solutions are realized by integrating workflow management of Apache Airflow components, stream processing and parsing capabilities of StreamEngineParser components, and high-performance timing databases of DolphinDB.

Benefits of technology

It effectively solves the problems of latency, uneven resource allocation and system complexity in large-scale data stream processing, and significantly improves data processing efficiency and service response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119718582B_ABST
    Figure CN119718582B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time stream computing and service scheduling method and system based on a distributed architecture platform, belonging to the technical field of real-time stream computing and scheduling strategies. It solves the problems of low resource utilization rate and low computing efficiency in the traditional real-time stream computing and service scheduling methods in the prior art; the present invention builds a real-time stream computing and service scheduling system based on a distributed architecture platform; in the custom task scheduling center of the real-time stream computing and service scheduling system, a workflow scheduling and management platform is integrated, a directed acyclic graph workflow is defined, and scheduling rules and strategies are formulated; in the custom stream computing processing center, a stream computing engine integrating a parser for real-time stream data processing is used for stream computing, and the calculation results are stored in the DolphinDB database to complete real-time stream computing and service scheduling. The present invention effectively improves the data processing efficiency and service response speed and can be applied to large-scale data stream processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a real-time stream computing and service scheduling method and system, and in particular to a real-time stream computing and service scheduling method and system based on a distributed architecture platform, belonging to the technical field of real-time stream computing and scheduling strategies. Background Art

[0002] In the current distributed architecture platform, service scheduling and real-time stream computing are two core functions. However, with the continuous explosion of data volume and the improvement of real-time requirements, existing technologies and methods have gradually revealed some significant problems. In terms of platform service scheduling, in the existing technologies, there are often problems such as low resource utilization rate, single scheduling strategy, and lack of dynamic adjustment mechanism in the scheduling strategy. In terms of real-time stream computing, traditional computing models often cannot effectively handle large-scale continuously arriving data streams, and bottlenecks are likely to occur during the transmission and processing of data between computing nodes, resulting in an increase in computing latency and inability to meet real-time requirements.

[0003] In summary, a real-time stream computing and service scheduling method and system based on a distributed architecture platform are needed. Summary of the Invention

[0004] A brief overview of the present invention is given below to provide a basic understanding of certain aspects of the present invention. It should be understood that this overview is not an exhaustive overview of the present invention. It is not intended to identify the key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is only to present certain concepts in a simplified form as a prelude to the more detailed description to be discussed later.

[0005] In view of this, to solve the problems of low resource utilization rate and low computing efficiency in the traditional real-time stream computing and service scheduling methods in the prior art, the present invention provides a real-time stream computing and service scheduling method and system based on a distributed architecture platform.

[0006] The first technical solution is as follows: A real-time stream computing and service scheduling method based on a distributed architecture platform includes the following steps:

[0007] S1. Build a real-time stream computing and service scheduling system based on a distributed architecture platform;

[0008] S2. Integrate the Apache Airflow component in the custom task scheduling center of the real-time stream computing and service scheduling system, define a directed acyclic graph workflow, and formulate scheduling rules and strategies;

[0009] S3. In the custom streaming computing processing center of the real-time stream computing and service scheduling system, integrate the stream computing engine of the StreamEngineParser component to perform streaming computing, store the computing results in the DolphinDB database, and complete real-time stream computing and service scheduling.

[0010] Furthermore, in S2, it specifically includes the following steps:

[0011] S21. In the custom task scheduling center, use Java to call the Apache Airflow component to build a directed acyclic graph workflow, where each node represents a task and the edge represents the dependency relationship between tasks;

[0012] S22. Add task nodes to the directed acyclic graph workflow. The task nodes include data source acquisition, data preprocessing, feature extraction, and real-time analysis. Define the execution order and dependency relationship of the tasks;

[0013] S23. Set the properties of the directed acyclic graph workflow to ensure that the directed acyclic graph workflow starts execution according to the predetermined period and time point;

[0014] S24. According to the constructed directed acyclic graph workflow and the dependency relationship of the tasks, use the scheduler in the custom task scheduling center to schedule the task execution, introduce comprehensive priority calculation, and obtain the scheduling comprehensive priority;

[0015] S25. According to the scheduling comprehensive priority, design an exception handling and fault tolerance mechanism to obtain the scheduling rules and strategies, and implement the automatic retry mechanism for tasks. When a task execution fails, automatically reschedule the task according to the preset number of retries and delay time interval;

[0016] In S24, the process of comprehensive priority calculation is specifically as follows:

[0017] S241. Use the normalization method to represent the urgency of the task as the relative value of the difference between the current time and the deadline, and calculate the task urgency priority P u,i ;

[0018] The task urgency priority P u,i is expressed as:

[0019]

[0020] where D i represents the deadline of the current task i, T represents the current time, and D j represents the deadline of any task;

[0021] S242. Calculate the resource requirement priority P by comparing the task resource requirement with the maximum resource requirement in the task set. r,i ;

[0022] The resource requirement priority P r,i is expressed as:

[0023]

[0024] where R i represents the amount of resources required for the current task i, and R j represents the amount of resources required for any task;

[0025] S243. Use the in-degree and out-degree of the task as a measure of the task dependency priority, and calculate the task dependency priority P d,i ;

[0026] The task dependency priority P d,i is expressed as:

[0027]

[0028] where De i represents the dependency factor corresponding to the current task i, and De j represents the dependency factor corresponding to any task;

[0029] S244. Calculate the comprehensive priority P t,i ;

[0030] P t,i = a·P u,i + b·P r,i + c·P d,i

[0031] where a represents the first weight coefficient, which is used to adjust the proportion of urgency in the priority calculation, b represents the first weight coefficient, which is used to adjust the proportion of resource requirements in the priority calculation, c represents the first weight coefficient, which is used to adjust the proportion of task dependencies in the priority calculation, and a + b + c = 1.

[0032] Furthermore, in S3, the following steps are specifically included:

[0033] S31. Configure optional stream computing engines for the DolphinDB database, including a reactive state engine, a cross-sectional engine, and a time series engine, and integrate the StreamEngineParser component into the DolphinDB database environment by configuring a JSON file;

[0034] S32. Define the stream computing parameters including an input table, an output table, and a custom algorithm in the DolphinDB database;

[0035] S33. Write a script in the DolphinDB database to subscribe to the input table, call the custom algorithm, and publish the results to the output table;

[0036] S34. Build a pipeline through the StreamEngineParser component to convert the user-defined algorithm into computing instructions that can be directly understood and efficiently executed by the underlying hardware. The computing instructions include data reading instructions, data processing instructions, and data output instructions, and optimize the computing instructions, execute the optimized computing instructions, and obtain the computing results;

[0037] S35. Store the computing results in the DolphinDB database in real time, or send the results to the specified output destination according to business requirements;

[0038] In S32, first define the input table, which includes the metric name, metric value, and timestamp, then define the output table, which includes the calculated predicted value and the corresponding timestamp, and finally define the custom algorithm for the user to set the algorithm logic according to business requirements.

[0039] The second technical solution is as follows: A real-time stream computing and service scheduling system based on a distributed architecture platform for implementing the real-time stream computing and service scheduling method based on a distributed architecture platform described in the first technical solution, including a custom task scheduling center, a custom stream computing processing center, a distributed storage center, and a data acquisition center;

[0040] The custom task scheduling center is respectively connected to the custom stream computing processing center and the distributed storage center, and the custom stream computing processing center is respectively connected to the distributed storage center and the data acquisition center.

[0041] The beneficial effects of the present invention are as follows: The present invention proposes a real-time stream computing and service scheduling method, relying on a distributed architecture platform, and realizes an efficient, flexible, and scalable real-time data processing and service scheduling solution by integrating the workflow management of the Apache Airflow component, the stream processing parsing ability of the StreamEngineParser component, and the high-performance time series database of DolphinDB; The present invention effectively solves the problems of latency, uneven resource allocation, and system complexity in large-scale data stream processing, and significantly improves the data processing efficiency and service response speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The accompanying drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation of the present invention. In the drawings:

[0043] Figure 1 It is a schematic flowchart of a real-time stream computing and service scheduling method based on a distributed architecture platform;

[0044] Figure 2 It is a schematic structural diagram of a real-time stream computing and service scheduling system based on a distributed architecture platform;

[0045] Figure 3 It is a schematic diagram of an embodiment of a real-time stream computing and service scheduling method and system based on a distributed architecture platform.

[0046] Reference numerals: 1. Custom task scheduling center; 2. Custom stream computing processing center; 3. Distributed storage center; 4. Data acquisition center. Detailed implementation manners

[0047] In order to make the technical solutions and advantages in the embodiments of the present invention clearer and more understandable, the following further details the exemplary embodiments of the present invention with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0048] Embodiment 1: Refer to Figures 1-3 This embodiment is described in detail. A real-time stream computing and service scheduling method based on a distributed architecture platform specifically includes the following steps:

[0049] S1. Build a real-time stream computing and service scheduling system based on a distributed architecture platform;

[0050] S2. In the custom task scheduling center of the real-time stream computing and service scheduling system, integrate the Apache Airflow component (workflow scheduling and management platform), define a directed acyclic graph workflow (DAG), and formulate scheduling rules and strategies;

[0051] S3. In the stream computing processing center of the real-time stream computing and service scheduling system, integrate the stream computing engine of the StreamEngineParser component (parser for real-time stream data processing) to perform stream computing, and store the calculation results in the DolphinDB database for subsequent data analysis, business decision-making, or visual display.

[0052] Further, in S2, it specifically includes the following steps:

[0053] In the custom task scheduling center, a directed acyclic graph (DAG) workflow is constructed by calling the Apache Airflow component in Java, where each node represents a task and the edges represent the dependencies between tasks.

[0054] S22. Add task nodes to the DAG workflow. The task nodes include data source acquisition, data preprocessing, feature extraction, real-time analysis, etc. Define the execution order and dependencies of the tasks.

[0055] S23. Set the properties of the DAG workflow, such as task owner, start date, scheduling interval, etc., to ensure that the DAG workflow starts execution according to the predetermined cycle and time point.

[0056] S24. According to the constructed DAG workflow and the dependencies of the tasks, use the scheduler in the custom task scheduling center to schedule the task execution, introduce comprehensive priority calculation, and obtain the scheduling comprehensive priority.

[0057] S25. According to the scheduling comprehensive priority, design an exception handling and fault tolerance mechanism to obtain the scheduling rules and strategies, and implement the automatic retry mechanism for tasks. When a task execution fails, the task is automatically rescheduled according to the preset number of retries and delay time interval. Integrate a monitoring and alarm system into the real-time stream computing and service scheduling system to monitor the task execution status and system resource usage in real time, and issue an alarm in a timely manner for abnormal situations and trigger the corresponding processing process.

[0058] In S24, the process of comprehensive priority calculation is as follows:

[0059] S241. Use the normalization method to represent the urgency of the task as the relative value of the difference between the current time and the deadline, and calculate the task urgency priority P u,i ;

[0060] The task urgency priority P u,i is expressed as:

[0061]

[0062] where D i represents the deadline of the current task i, T represents the current time, and D j represents the deadline of any task;

[0063] S242. By comparing the task resource requirements with the maximum resource requirements in the task set, calculate the resource requirement priority P r,i ;

[0064] The resource requirement priority P r,i is expressed as:

[0065]

[0066] Among them, R i represents the amount of resources required for the current task i, and R j represents the amount of resources required for any task;

[0067] S243. Use the in-degree (i.e., the number of other tasks that depend on this task) and out-degree (i.e., the number of other tasks that this task depends on) of the task as a measure of the dependency relationship priority, and calculate the dependency relationship priority P of the task d,i ;

[0068] The dependency relationship priority P of the task d,i is expressed as:

[0069]

[0070] Among them, De i represents the dependency relationship factor corresponding to the current task i, and De j represents the dependency relationship factor corresponding to any task;

[0071] S244. Calculate the comprehensive priority P t,i ;

[0072] P t,i = a·P u,i + b·P r,i + c·P d,i

[0073] Among them, a represents the first weight coefficient, which is used to adjust the proportion of the urgency level in the priority calculation, b represents the first weight coefficient, which is used to adjust the proportion of the resource requirement in the priority calculation, c represents the first weight coefficient, which is used to adjust the proportion of the task's dependency relationship in the priority calculation, and a + b + c = 1.

[0074] Furthermore, in the said S3, it specifically includes the following steps:

[0075] S31. Configure optional stream computing engines for the DolphinDB database, including the ReactiveStateEngine, CrossSectionalEngine, TimeSeriesEngine, etc., and integrate the StreamEngineParser component into the DolphinDB database environment by means of configuring a JSON file;

[0076] S32. Define stream computing parameters including input tables, output tables, and custom algorithms in the DolphinDB database;

[0077] S33. Write a script in the DolphinDB database to subscribe to the input table, call the custom algorithm, and publish the results to the output table;

[0078] S34. Build a pipeline through the StreamEngineParser component to convert the user-defined algorithm into computational instructions that can be directly understood and executed efficiently by the underlying hardware. The computational instructions include data reading instructions (data source access, data flow control, etc.), data processing instructions (computational instructions, data conversion instructions, window operation instructions, etc.), data output instructions (result storage instructions, data flow end instructions), etc., and optimize the computational instructions, execute the optimized computational instructions, and obtain the computational results;

[0079] S35. Store the computational results in the distributed time-series database of DolphinDB in real time, or send the results to the specified output destination according to business requirements, such as message queues, database tables, visualization interfaces, etc., for subsequent analysis and applications;

[0080] In S32, first define the input table, which is used to receive real-time data streams. The input table includes fields such as the metric name (name), metric value (value), and timestamp (timestamp). Then define the output table, which is used to store the results of the stream calculation. The output table includes the calculated predicted value (MAP) and the corresponding timestamp. Finally, define the custom algorithm for the user to set the algorithm logic according to business requirements;

[0081] Specifically, in this embodiment, in step S32, the specific settings of the custom algorithm are as follows:

[0082] The Internet of Things system contains a large number of temperature sensors and humidity sensors deployed at different locations. It is necessary to calculate the comprehensive environmental index CEI for each location in real time. This index comprehensively considers the changes in temperature and humidity and has complex non-linear relationships;

[0083] The comprehensive environmental index CEI is expressed as:

[0084]

[0085] where T is the current temperature, H is the current humidity, T0 is the reference value of temperature, T0 = 25 °C, H0 is the reference value of humidity, H0 = 50%RH, α0 is the first model parameter, α0 = 1.2, α1 is the second model parameter, α1 = 1.5, α2 is the third model parameter, α2 = 0.8, σ is the fourth model parameter, σ = 0.5, and β is the fifth model parameter, β = 0.01;

[0086] In step S34, the optimization method of the computational instructions is as follows:

[0087] S341. By adopting loop unrolling and instruction parallelism, reduce the overhead of loop control instructions and increase instruction-level parallelism. For a loop instruction that needs to iterate m times to perform a certain complex calculation task;

[0088] The iterative process of performing a certain complex calculation task is expressed as:

[0089]

[0090] where A represents the first matrix, B represents the second matrix, C represents the third matrix, m is the number of iterations, and j is the task other than the current task i;

[0091] During the calculation process, through loop unrolling, duplicate the code in the loop body to reduce the number of loop iterations; multiple elements can be calculated simultaneously in each iteration to improve the calculation efficiency;

[0092] The loop unrolling process is expressed as:

[0093] C[i][j] = A[i][0] × B[0][j] + A[i][1] × B[1][j]

[0094] + A[i][2] × B[2][j] + A[i][3] × B[3][j] +...

[0095] S342. Adopt the vectorized single instruction multiple data (SIMD) instruction set to process multiple instructions simultaneously;

[0096] For a function that needs to operate on an array, through vectorization, divide the array into multiple vectors, each vector contains multiple elements, and then use the single instruction multiple data (SIMD) instruction to perform parallel processing on the vectors;

[0097] The process of parallel processing is as follows:

[0098] First, divide the array X into multiple vectors, for example, each vector contains 4 elements;

[0099] The vector V containing 4 elements K is expressed as:

[0100] V K = (X[4k], X[4k + 1], X[4k + 2], X[4k + 3])

[0101] Use the SIMD instruction to perform parallel summation on the vector V containing 4 elements K to obtain the processing results S of multiple instructions, that is, the optimized calculation instructions;

[0102] The processing results S of multiple instructions are expressed as:

[0103]

[0104] Among them, SIMD_SUM represents the operation of summing the vector V using single instruction multiple data instructions. k for summing.

[0105] Embodiment 2: Refer to Figures 2-3 This embodiment will be described in detail. A real-time stream computing and service scheduling system based on a distributed architecture platform is used to implement the real-time stream computing and service scheduling method based on the distributed architecture platform described in Embodiment 1, including a custom task scheduling center 1, a custom stream computing processing center 2, a distributed storage center 3, and a data acquisition center 4;

[0106] The custom task scheduling center 1 is respectively connected to the custom stream computing processing center 2 and the distributed storage center 3, and the custom stream computing processing center 2 is respectively connected to the distributed storage center 3 and the data acquisition center 4.

[0107] Specifically, the custom stream computing processing center 2 relies on the cooperation of each node to execute complex real-time computing tasks;

[0108] The custom task scheduling center 1 is used to divide into master and slave nodes and intelligently allocate tasks according to the system load and resource status;

[0109] The distributed storage center 3 is responsible for the persistent storage of data, and the distributed storage center 3 is the DolphinDB database;

[0110] Each component node realizes collaborative work through an efficient data transmission and scheduling mechanism.

[0111] Although the present invention has been described according to a limited number of embodiments, those skilled in the art in this technical field will understand that other embodiments can be envisioned within the scope of the present invention as thus described. In addition, it should be noted that the language used in this specification is mainly selected for the purpose of readability and teaching, rather than for the purpose of explaining or limiting the subject matter of the present invention. Therefore, many modifications and variations will be obvious to those of ordinary skill in the art in this technical field without departing from the scope and spirit of the appended claims. For the scope of the present invention, the disclosure of the present invention is illustrative rather than restrictive, and the scope of the present invention is defined by the appended claims.

Claims

1. A real-time stream computing and service scheduling method based on a distributed architecture platform, characterized in that: The following steps are involved: S1. Build a real-time stream computing and service scheduling system based on a distributed architecture platform; S2. Integrate Apache Airflow components in the custom task scheduling center of the real-time stream computing and service scheduling system, define directed acyclic graph workflows, and formulate scheduling rules and strategies; S3. In the custom stream computing processing center of the real-time stream computing and service scheduling system, the stream computing engine of the integrated StreamEngineParser component performs stream computing and stores the calculation results in the DolphinDB database to complete the real-time stream computing and service scheduling; The S2 specifically includes the following steps: S21. In the custom task scheduling center, Apache Airflow components are called through Java to build a directed acyclic graph workflow, where each node represents a task and the edge represents the dependency relationship between tasks; S22. Add task nodes of the directed acyclic graph workflow, which include data source acquisition, data preprocessing, feature extraction, real-time analysis, and define the execution order and dependencies of tasks; S23. Setting the properties of the directed acyclic graph workflow to ensure that the directed acyclic graph workflow starts executing according to the predetermined cycle and time point; S24. According to the constructed directed acyclic graph workflow and the dependency relationship of the tasks, the scheduler in the custom task scheduling center is used to schedule the task execution, and the comprehensive priority calculation is introduced to obtain the scheduling comprehensive priority; S25. According to the comprehensive scheduling priority, design the exception handling and fault tolerance mechanism, obtain the scheduling rules and strategies, and implement the automatic retry mechanism of the task. When the task fails to execute, the task is automatically rescheduled according to the preset number of retries and delay time interval; In S24, the process of calculating the comprehensive priority is as follows: S241. Using the normalization method, the urgency of the task is expressed as the relative value of the difference between the current time and the deadline, and the task urgency priority P is calculated. u,i ; Task urgency priority u,i It is expressed as: Among them, D i represents the deadline of the current task i, T represents the current time, and D j Indicates the deadline of any task; S242. By comparing the task resource requirement with the maximum resource requirement in the task set, the resource requirement priority P is calculated. r,i ; Resource requirement priority P r,i It is expressed as: Among them, R i represents the amount of resources required for the current task i, R j Indicates the amount of resources required for any task; S243. Use the in-degree and out-degree of the task as the measure of the task dependency priority to calculate the task dependency priority P d,i ; The dependency priority of the task P d,i It is expressed as: Among them, De i Denotes the dependency factor corresponding to the current task i, De j Represents the dependency factor corresponding to any task; S244. Calculate the comprehensive priority P t,i ; P t,i =a·P u,i +b·P r,i +c·P d,i Wherein, a represents the first weight coefficient, which is used to adjust the proportion of urgency in the priority calculation, b represents the first weight coefficient, which is used to adjust the proportion of resource demand in the priority calculation, c represents the first weight coefficient, which is used to adjust the proportion of task dependency in the priority calculation, and a+b+c=1; The S3 specifically includes the following steps: S31. Configure optional stream computing engines for the DolphinDB database, including the responsive state engine, cross-section engine, and time series engine, and integrate the StreamEngineParser component into the DolphinDB database environment by configuring a JSON file; S32. Define the stream computing parameters including input table, output table and custom algorithm in DolphinDB database; S33. Write a script in the DolphinDB database that subscribes to the input table, calls the custom algorithm, and publishes the results to the output table; S34. Build a pipeline through the StreamEngineParser component to convert the user-defined algorithm into efficient computing instructions that can be directly understood and executed by the underlying hardware. The computing instructions include data reading instructions, data processing instructions, and data output instructions. The computing instructions are optimized and the optimized computing instructions are executed to obtain the computing results. S35. Store the calculation results in the DolphinDB database in real time, or send the results to a specified output destination according to business needs; In S32, an input table is first defined, the input table includes an indicator name, an indicator value and a timestamp, and then an output table is defined, the output table includes a calculated predicted value and a corresponding timestamp, and finally a custom algorithm is defined in which the user sets the algorithm logic according to business needs.

2. A real-time stream computing and service scheduling system based on a distributed architecture platform, characterized in that: Used to implement the real-time stream computing and service scheduling method based on a distributed architecture platform as described in claim 1, comprising a custom task scheduling center (1), a custom stream computing processing center (2), a distributed storage center (3) and a data acquisition center (4); The custom task scheduling center (1) is connected to the custom stream computing processing center (2) and the distributed storage center (3) respectively, and the custom stream computing processing center (2) is connected to the distributed storage center (3) and the data collection center (4) respectively.

Citation Information

Patent Citations

  • Real-time data calculation and storage method based on flow and message scheduling

    CN112597205A

  • Distributed stream processing task scheduling method and device

    CN117806781A