Multi-system dynamic scheduling processing method for public service real-time streaming data
Through distributed data acquisition and real-time streaming batch integrated processing, combined with machine learning scheduling strategies, data processing problems among multiple systems in the public service field are solved, and efficient and flexible data flow and collaborative processing are achieved.
Patent Information
- Application Number
- CN202510560301.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-04-30
AI Technical Summary
Traditional data processing architectures are difficult to meet the fast and accurate processing needs of massive real-time streaming data in the modern public service field, especially lacking high performance and flexibility in the flow and processing between multiple systems.
The distributed data acquisition framework is adopted to collect data and preprocess it, combined with a real-time stream computing engine and batch processing system, and dynamically allocate tasks to different subsystems through machine learning load prediction and multi-objective optimization scheduling model, and realize multi-system collaborative work through system adaptation interfaces and data routing mechanisms.
It realizes the integration of real-time stream computing and batch processing, improves the real-time, reliability and flexibility of data processing, adapts to complex business needs, and supports efficient processing of multiple data types and cross-domain data fusion.
Smart Images

Figure CN120471365A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data processing technology in the field of public services, and in particular to, but not limited to, a multi-system dynamic scheduling processing method for real-time streaming data of public services. Background Art
[0002] In modern society, public service sectors such as traffic management, healthcare, and urban security face the need to process massive amounts of real-time streaming data. This data comes from a wide range of sources, in diverse formats, and is frequently updated. It must be quickly and accurately transferred and processed across multiple systems to support real-time decision-making and efficient services. However, traditional data processing architectures and scheduling methods struggle to meet the high-performance requirements of these complex scenarios. Summary of the Invention
[0003] In view of this, an embodiment of the present invention provides a multi-system dynamic scheduling and processing method for real-time streaming data of public services, aiming to overcome the shortcomings of the existing technology and improve the real-time performance, reliability and flexibility of data processing.
[0004] The technical solutions of the embodiments of the present invention are as follows:
[0005] An embodiment of the present invention provides a multi-system dynamic scheduling processing method for real-time streaming data of public services, the method comprising:
[0006] Data collection and preprocessing layer: collects real-time streaming data from multiple data sources through a distributed data collection framework and performs preprocessing and transmission; wherein, the multiple data sources include various public service data sources;
[0007] Processing and analysis layer: Use the real-time stream computing engine to perform real-time stream computing on pre-processed data. At the same time, combine with the batch processing system to conduct in-depth analysis of historical data and trigger data processing tasks.
[0008] Multi-system scheduling and collaboration layer: The task scheduling center uses a dynamic scheduling strategy that combines machine learning-based load prediction and a multi-objective optimization scheduling model to dynamically allocate data processing tasks to different subsystems. The system adaptation interface and data routing and forwarding mechanism enable collaborative work between multiple subsystems.
[0009] Application and service layer: Based on the processing and analysis results of each subsystem, it provides business applications for specific public service areas, as well as decision support systems and data sharing and open platforms.
[0010] In some embodiments, the real-time streaming data is collected from multiple data sources and pre-processed through a distributed data collection framework, including: using Kafka as a message queue to collect real-time streaming data from different data sources, and performing preliminary format unification and standardization processing; designing an intelligent edge computing task allocation algorithm based on the characteristics and service requirements of the data source, and dynamically deploying edge computing nodes to perform real-time pre-processing of part of the data; caching the pre-processed data into different cache systems according to the data type, and uploading it to the data center.
[0011] In some embodiments, the real-time stream computing engine is used to perform real-time stream computing on the preprocessed data, and at the same time, the batch processing system is combined to perform in-depth analysis of historical data to trigger data processing tasks, including: using Flink's CEP library to define event patterns, and detecting key events that meet specific patterns from the continuous data streams transmitted by the data acquisition and preprocessing layers; using Flink's window functions to perform aggregation statistics within the time window on the real-time streaming data corresponding to the key events; using Hadoop / Hive to analyze historical accident data, predict risk areas, and obtain batch processing results; based on the Kappa architecture, the real-time stream and batch processing results are unified through the middleware to obtain a stream-batch integrated data processing task.
[0012] In some embodiments, an intelligent edge computing task allocation algorithm is designed based on the characteristics and service requirements of the data source, and edge computing nodes are dynamically deployed to perform real-time preprocessing of some data, including: for first-level tasks, processing is completed at the edge computing nodes first; for second-level tasks, the results or key data are uploaded to the central system for further analysis after preliminary processing at the edge computing nodes; for third-level tasks, they are directly uploaded to the central system for processing; wherein, the real-time requirements of the first-level tasks, the second-level tasks, and the third-level tasks decrease in sequence and the computational complexity increases in sequence.
[0013] In some embodiments, the method further includes: integrating a machine learning library and AI model training and reasoning tools in edge computing nodes, and using historical data and real-time data to predict, classify, and detect anomalies in public service scenarios.
[0014] In some embodiments, the dynamic scheduling strategy that combines machine learning-based load prediction and a multi-objective optimization scheduling model dynamically allocates the data processing tasks to different subsystems, including: training a load prediction model and predicting the load trend of each system based on historical task execution data and real-time streaming data; constructing a multi-objective optimization scheduling model with the goals of minimizing response time, maximizing resource utilization, load balancing index, and task priority, while satisfying resource capacity, task dependency, and QoS constraints, and using the non-dominated sorting genetic algorithm NSGA-III to generate a Pareto optimal scheduling solution; through a sliding window mechanism, dynamically adjusting the allocation of data processing tasks according to the predicted load.
[0015] In some embodiments, the method further includes: real-time monitoring of environmental changes and task execution status, and triggering model updates and policy adjustments.
[0016] In some embodiments, the collaborative work between multiple subsystems is achieved through the system adaptation interface and data routing and forwarding mechanism, including: developing standardized interfaces for different subsystems to support task submission, status query, and resource application operations; using Kafka message queues to achieve asynchronous transmission of task metadata and intermediate results; describing task dependencies through directed acyclic graphs, and using the Airflow scheduling framework to achieve task cascade triggering.
[0017] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0018] 1) A stream-batch integrated data processing model is proposed, combining real-time stream computing with batch processing. A real-time stream computing engine performs complex event processing and window aggregation calculations on real-time streaming data to unlock the real-time value of the data. Simultaneously, a batch processing system is used to conduct in-depth analysis of historical data, achieving stream-batch integrated processing and providing more comprehensive data support for public service decision-making.
[0019] 2) A multi-system collaborative architecture was constructed, including the data source layer, data acquisition and preprocessing layer, data processing and analysis layer, multi-system scheduling and collaboration layer, and application and service layer. Each layer has a clear division of labor. Through system adaptation interfaces and data routing and forwarding mechanisms, seamless flow and interaction between heterogeneous systems are achieved, ensuring that data can be efficiently processed and utilized across multiple systems.
[0020] 3) A dynamic scheduling strategy combines machine learning-based load forecasting with a multi-objective optimization scheduling model. This strategy uses machine learning to accurately predict future system loads. It then considers multiple objectives, including real-time data processing, resource utilization, system throughput, and task priority, to build a multi-objective optimization scheduling model. It then applies intelligent optimization algorithms to determine the optimal task scheduling solution, achieving dynamic resource allocation and efficient task execution.
[0021] 4) This method can process a variety of data types, including structured, semi-structured, and unstructured data generated by diverse data sources such as traffic sensors, medical IoT devices, and video surveillance cameras. Furthermore, the proposed solution is not only applicable to specific public service areas such as traffic management, healthcare, and urban security, but can also be expanded and upgraded based on actual needs, adapting to evolving business demands and technological developments. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive work, among which:
[0023] Figure 1 A flowchart of a multi-system dynamic scheduling processing method for public service real-time streaming data provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0025] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0026] It should be pointed out that the terms "first\second\third" involved in the embodiments of the present invention are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present invention described here can be implemented in an order other than that illustrated or described here.
[0027] Those skilled in the art will understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art in the art to which the embodiments of the present invention pertain. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with those in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0028] Figure 1 A flowchart of a multi-system dynamic scheduling processing method for real-time streaming data of public services provided by an embodiment of the present invention is shown as follows: Figure 1 As shown, the method comprises at least the following steps:
[0029] Step S110, data collection and preprocessing layer: collect real-time streaming data from multiple data sources through a distributed data collection framework and perform preprocessing and transmission.
[0030] This step aims to extract high-quality streaming data from heterogeneous data sources. These multiple data sources include various public service data sources, i.e., various systems in the public service sector, such as traffic surveillance cameras, environmental monitoring sensors, government service interfaces, weather stations, medical devices, and social media APIs. These data sources continuously generate real-time streaming data in various data formats (structured, semi-structured, and unstructured) and transmission protocols. These data sources are characterized by large data volumes, diverse formats, and high real-time requirements.
[0031] The technical implementation mainly includes the following key points:
[0032] 1) Distributed Collection Framework: This framework is used to build the data collection and preprocessing layer, supporting high-throughput, low-latency data access. Furthermore, a "data diversion" strategy is employed during collection: high-priority data (such as emergency alarms) is transmitted via low-latency channels, while general data is transmitted via batch queues.
[0033] Common distributed data collection frameworks include:
[0034] Apache Kafka: A high-throughput, distributed, and scalable message queue that supports real-time data stream processing. Apache Flume: A distributed, reliable, and highly available log collection system suitable for real-time collection of log data. Apache Nifi: A visual tool that supports data stream management and is easy to configure and manage. Logstash: An open source data collection engine that supports real-time collection and preprocessing of multiple data sources.
[0035] During implementation, select appropriate tools based on business needs to build the data collection and preprocessing layer and processing framework (e.g., Kafka for high-throughput scenarios, Flink for complex computing scenarios). Also, select the appropriate method to access data based on the type of data source:
[0036] Log files (including server logs, application logs, access logs, etc.): Use Flume or Logstash to collect log files, supporting multiple log formats (such as JSON and CSV).
[0037] Message queues (Kafka, RabbitMQ, Pulsar, etc.): Use Kafka as a data source or target through Kafka Connect or other client libraries.
[0038] Databases (relational and NoSQL): Use tools such as Debezium to capture database change logs (CDC) for real-time data collection.
[0039] API interface: Get data in real time through HTTP requests or WebSocket.
[0040] 2) Preprocessing operations: After data collection, the following real-time preprocessing is performed to meet the needs of subsequent analysis or storage.
[0041] Data cleaning: removes noise, outliers, and other invalid data (such as null values, duplicate values, and GPS drift data), and formats data (such as timestamp conversion and field mapping). Data transformation: aggregates data (by time window, such as traffic statistics in a certain area within 5 minutes to reduce downstream computing pressure), and calculates derived fields (such as average and maximum values). Data filtering: filters data according to business rules (such as retaining only specific types of events). For example, the urban traffic monitoring system collects data from 1,000+ cameras through Kafka, reduces the amount of raw data by 30% after preprocessing, and marks accident-prone sections in real time.
[0042] Efficiently connecting to diverse data sources, including but not limited to sensor networks, log files, and database change logs, while automatically identifying and parsing different data formats and transmission protocols to ensure the integrity and accuracy of collected data, is fundamental to subsequent dynamic multi-system scheduling. This invention aims to use a distributed data acquisition framework to collect data from multiple heterogeneous data sources with low latency and high throughput, while supporting subsequent real-time processing and analysis.
[0043] For example, scenario 1: traffic flow monitoring, the corresponding data source is traffic sensors.
[0044] Processing: Sensor data (such as GPS coordinates, traffic flow, and speed) is pushed to Kafka via MQTT or CoAP. Flink consumes the Kafka data and calculates minute-by-minute traffic flow. The results are stored in PostgreSQL and displayed as heatmaps using Grafana.
[0045] Scenario 2: Patient health monitoring, the corresponding data source is medical Internet of Things devices.
[0046] Processing: Patient vital signs (heart rate, blood pressure) and device status are pushed to Kafka via HTTP APIs or WebSockets. Flink consumes the Kafka data and detects abnormal heart rates or blood pressure. When an abnormality is detected, an alarm is triggered to notify the doctor.
[0047] Scenario 3: Video surveillance analysis, the corresponding data source is a video surveillance camera.
[0048] Processing flow: The camera runs the object detection model and pushes keyframes to Kafka. Flink consumes the Kafka data and counts the number of vehicles detected per minute. The results are stored in Elasticsearch, supporting historical queries.
[0049] Step S120, processing and analysis layer: use the real-time stream computing engine to perform real-time stream computing on the pre-processed data, and combine the batch processing system to perform in-depth analysis on historical data to trigger data processing tasks.
[0050] This step aims to jointly analyze real-time streaming data with historical data to achieve integrated stream and batch processing. The preprocessed data is fed into a real-time stream computing engine, such as Apache Flink or Spark Streaming, for stream processing to unlock the data's real-time value. For tasks requiring in-depth analysis of historical data, batch processing systems (such as Hadoop and Hive) are combined to analyze historical data, uncovering long-term trends or correlations, working in conjunction with the stream computing engine. Finally, the Kappa architecture unifies real-time and offline computing logic, reducing data redundancy and achieving stream and batch integration. Custom middleware is used to merge real-time and batch processing results into a unified view, and the analysis results are pushed to downstream applications (such as early warning systems and decision-making platforms).
[0051] The technical implementation involves the following key points:
[0052] 1) Real-time stream computing engine: Use Flink / Spark Streaming to process complex events (such as "three consecutive speeding events trigger an alert") and window aggregation (such as "calculating the peak traffic flow within 1 hour").
[0053] 2) Batch Processing System: Combined with Hadoop / Hive, this system analyzes historical accident data and predicts risk areas. For example, it analyzes user behavior logs from the past month to identify user preferences and predict churn risk. Furthermore, it trains machine learning models (such as user profiling models and recommendation systems) based on batch processing tasks and feeds the results into a real-time stream processing system to improve the accuracy of real-time decision-making.
[0054] 3) Stream-batch integration: Based on the Kappa architecture, real-time stream and batch processing logic are unified, and data fusion is achieved through middleware (such as Kafka Streams) to reduce data redundancy. The working principle is as follows:
[0055] The data processing model in existing technologies is relatively simple. For example, some studies only focus on the processing of real-time streaming data, or only focus on batch data processing. When faced with complex and changeable data processing needs in the public service field, this single processing model may not be able to fully utilize the value of data. The present invention proposes a stream-batch integrated data processing model that combines real-time stream computing with batch processing. A real-time stream computing engine is used to perform stream computing operations on real-time streaming data to mine the real-time value of the data; at the same time, a batch processing system is used to conduct in-depth analysis of historical data to achieve stream-batch integrated processing and provide more comprehensive data support for public service decision-making.
[0056] Step S130, multi-system scheduling and collaboration layer: The task scheduling center adopts a dynamic scheduling strategy that combines machine learning-based load prediction and a multi-objective optimization scheduling model to dynamically allocate data processing tasks to different subsystems, and realizes collaborative work between multiple systems through system adaptation interfaces and data routing and forwarding mechanisms.
[0057] This step aims to dynamically allocate tasks based on resources and achieve efficient multi-system collaboration. A dynamic scheduling strategy combining machine learning-based load forecasting and a multi-objective optimization scheduling model is employed to achieve dynamic resource allocation and efficient task execution. The technical implementation involves the following key points:
[0058] 1) Task Scheduling Center: Responsible for dynamically allocating data processing tasks based on factors such as data characteristics, and monitoring and managing the task execution process. Preferably, the present invention implements resource-aware scheduling based on Kubernetes / YARN, taking into account system load and task priority.
[0059] 2) Machine learning-based load forecasting: Collect historical system load data and business traffic patterns, and use time series analysis, deep learning and other technologies to accurately predict system load in the future, providing a forward-looking basis for task scheduling.
[0060] 3) Multi-objective optimization scheduling model: Taking into account multiple objectives such as real-time data processing, resource utilization, system throughput, and task priority, a mathematical optimization model is constructed, and the optimal task scheduling solution is solved through an intelligent optimization algorithm.
[0061] 3) System adaptation interface: Provides RESTful API / gRPC interface to support rapid access to subsystems (such as early warning systems, data analysis platforms, and historical databases).
[0062] 4) Data routing and forwarding: Determine the data transmission path based on the task scheduling strategy and transmit the data to the corresponding subsystem to ensure the security and reliability of data transmission. Preferably, use NSQ / RabbitMQ for data distribution to ensure low latency (<100ms).
[0063] Step S140, application and service layer: based on the processing and analysis results of each subsystem, business applications oriented to specific public service fields, as well as decision support systems and data sharing and open platforms are provided.
[0064] This step aims to transform data value into business capabilities. Based on the data processing and analysis results obtained in the previous steps, various business applications for specific public service areas can be developed, such as intelligent traffic control systems, remote medical diagnosis platforms, and urban security monitoring and command systems, providing users with intuitive and convenient service interfaces.
[0065] Among them, business applications in specific public service fields, such as real-time traffic diversion systems, can push congestion warnings to navigation apps via WebSocket; medical equipment abnormality monitoring systems can trigger automatic alarms to hospital systems; and pollution warning systems can automatically trigger traffic restrictions when PM2.5 exceeds the standard.
[0066] The decision support system provides public service managers and decision makers with tools such as data visualization, report generation, and trend analysis to assist them in making scientific decisions and optimizing resource allocation. For example, the traffic flow forecasting model uses LSTM to predict traffic flow over the next hour.
[0067] The data sharing and open platform will implement the secure sharing and openness of public service data in accordance with relevant laws, regulations, and data security standards, promoting cross-departmental and cross-domain data integration and innovation. It will also provide standardized APIs to support data access by third-party applications (such as urban planning departments).
[0068] Through the above steps, the embodiment of the present invention achieves efficient processing of real-time streaming data of public services and multi-system collaboration, providing technical support for smart city management. Specifically, the integrated stream-batch processing in the processing and analysis layer, combining the real-time stream computing engine and the batch processing system, can reduce data redundancy and improve analysis timeliness; dynamic scheduling in the multi-system scheduling and collaboration layer adapts to changes in business priorities and provides scheduling flexibility; the application and service layer provides users with an intuitive and convenient service interface, assisting them in making scientific decisions and optimizing resource allocation, and promoting cross-departmental and cross-domain data integration innovation.
[0069] In some embodiments, the above step S110 "collecting real-time streaming data from multiple data sources and preprocessing it through a distributed data collection framework" is further implemented through the following process: using Kafka as a message queue to collect real-time streaming data from different data sources, and performing preliminary format unification and standardization processing; designing an intelligent edge computing task allocation algorithm based on the characteristics and service requirements of the data source, and dynamically deploying edge computing nodes to perform real-time preprocessing of part of the data; according to the data type, caching the preprocessed data into different cache systems, and uploading it to the data center.
[0070] Here, we first identify and connect to various data sources, automatically parsing data in different formats and transmission protocols. Traffic sensors and medical IoT devices use JSON or Avro formats. Video streams from video surveillance systems are directly stored in object storage (such as S3), and key frame data (such as motion detection results) are pushed to Kafka in JSON format.
[0071] To unify data in different formats into a standard format for easier processing and analysis, the following strategies can be adopted: 1) Define a unified data model: Based on business requirements, define a unified data model, including data structure, field types, and naming conventions. 2) Data format conversion tools: Develop or adopt existing data format conversion tools to convert data in different formats into a unified data model. 3) Data quality verification: Incorporate data quality verification into the data conversion process to ensure data accuracy and integrity.
[0072] Then, based on the characteristics and service requirements of the data source, an intelligent edge computing task allocation algorithm is designed and deployed close to the data source to optimize overall system performance. For data with high real-time requirements, the deployed edge computing nodes perform real-time preprocessing such as data cleaning, filtering, and compression, thereby reducing invalid data transmission and alleviating processing pressure on the central system.
[0073] Among them, the characteristics of the data source mainly include: data type (such as sensor data, video stream data, log data, etc.); data generation frequency: high frequency (such as hundreds of times per second) or low frequency (such as once an hour); data volume (ranging from KB to GB), data real-time requirements (such as real-time monitoring, early warning, control, etc.), and data sensitivity (data involving privacy, security or commercial secrets requires special processing).
[0074] Service requirements primarily include: Response time requirements, such as millisecond, second, or minute responses. Processing accuracy requirements, such as high-precision and approximate calculations. Resource consumption limits, such as CPU, memory, and storage usage limits. Reliability requirements, such as no data loss and accurate calculation results.
[0075] Finally, we utilize high-performance caching technologies (such as Hadoop HDFS, Elasticsearch, relational databases, and NoSQL databases) to temporarily cache collected data to address scenarios like data spikes and system failure recovery. Hadoop HDFS is suitable for large-scale data storage; Elasticsearch supports full-text search and real-time analysis; relational databases such as MySQL and PostgreSQL are suitable for structured data; and NoSQL databases such as MongoDB and Cassandra are suitable for semi-structured or unstructured data.
[0076] In this embodiment, distributed message queue technology is used to efficiently collect data from different data sources and perform preliminary format unification and standardization. Deploying edge computing nodes reduces invalid data transmission and reduces processing pressure on the central system. A data caching mechanism is also provided to address scenarios such as data surges and system failure recovery.
[0077] In some embodiments, the above-mentioned step S120 "using a real-time stream computing engine to perform real-time stream computing on preprocessed data, and combining the batch processing system to perform in-depth analysis of historical data to trigger data processing tasks" is further implemented through the following process: using Flink's CEP library to define event patterns, and detecting key events that meet specific patterns from the continuous data streams transmitted by the data acquisition and preprocessing layers; using Flink's window functions to perform aggregation statistics within the time window on the real-time streaming data corresponding to the key events; using Hadoop / Hive to analyze historical accident data, predict risk areas, and obtain batch processing results; based on the Kappa architecture, the real-time stream and batch processing results are unified through the middleware (Kafka Streams) to obtain a stream-batch integrated data processing task.
[0078] Here, in stream processing, complex events are processed through tools such as Flink / Spark Streaming. For example, in a traffic monitoring scenario, if a vehicle experiences three consecutive speeding incidents (speed exceeding 100 km / h) with an interval of less than 5 minutes, a "speeding driving" warning is triggered; in a financial risk control scenario, if an account has three remote logins within 5 minutes, an abnormal warning is triggered.
[0079] Aggregate and analyze streaming data based on time windows (such as sliding windows and session windows). For example, in traffic monitoring scenarios, use Flink's window functions (such as TumblingWindow) to aggregate and analyze video key frames and calculate the peak traffic flow within an hour. In e-commerce scenarios, count the number of clicks and conversion rate of a product page per minute and dynamically update the real-time large screen data.
[0080] Combined with batch processing systems such as Hadoop / Hive, historical accident data can be analyzed to uncover long-term trends or correlations. For example, when analyzing real-time traffic data, historical accident data (e.g., T+1 data) can be combined to dynamically adjust warning thresholds (e.g., lowering the speeding threshold to 80 km / h on rainy days). In financial risk control scenarios, user behavior logs from the past month can be analyzed to uncover user preferences and predict churn risk.
[0081] In this example, complex event processing and window aggregation computing enable low-latency real-time computing, while batch processing enables in-depth analysis of historical data. This integrated stream-batch data processing model combines real-time stream computing with batch processing results, achieving a balance between real-time response and deep insight, while ensuring business timeliness and unlocking the long-term value of data.
[0082] In some embodiments, the above-mentioned step S130 "designs an intelligent edge computing task allocation algorithm based on the characteristics and service requirements of the data source, and dynamically deploys edge computing nodes to perform real-time pre-processing of some data", includes: for first-level tasks, processing is completed at the edge computing node first; for second-level tasks, the results or key data are uploaded to the central system for further analysis after preliminary processing at the edge computing node; for third-level tasks, they are directly uploaded to the central system for processing; wherein, the real-time requirements of the first-level tasks, the second-level tasks, and the third-level tasks decrease in sequence and the amount of computation increases in sequence.
[0083] Here, the first-level tasks are those with high real-time requirements and low computational workload, such as real-time monitoring and early warning of sensor data; the second-level tasks are those with moderate real-time requirements and high computational workload, such as preliminary analysis of video stream data; the third-level tasks are those with low real-time requirements and high computational workload, such as in-depth analysis of large-scale log data.
[0084] Based on task classification and edge node capability assessment, the following intelligent task allocation strategy is designed: first-level tasks such as traffic congestion warning and accident detection are completed first at edge nodes and immediately trigger corresponding control actions to reduce network latency and central system load; second-level tasks such as traffic flow statistics and vehicle behavior analysis can be preliminarily processed at edge nodes, and then the results or key data can be uploaded to the central system for further analysis and visualization; third-level tasks such as long-term traffic trend prediction and road maintenance plan formulation can be directly uploaded to the central system for processing to fully utilize the computing resources and data storage capabilities of the central system.
[0085] In this embodiment, by designing an intelligent edge computing task allocation algorithm, it is possible to dynamically determine which data processing operations are completed at the edge node and which need to be uploaded to the central system based on the characteristics of the data source and service requirements, thereby achieving optimal overall system performance.
[0086] In some embodiments, the method further includes: integrating a machine learning library and AI model training and reasoning tools in edge computing nodes, and using historical data and real-time data to predict, classify, and detect anomalies in public service scenarios.
[0087] Machine learning libraries such as TensorFlow and PyTorch are used here. By integrating advanced machine learning libraries and efficient AI model training and inference tools in edge computing nodes, an intelligent data processing and analysis system is built. This system can fully leverage accumulated historical data and the continuous influx of real-time data to conduct in-depth mining and analysis of various public service scenarios. Specifically, it can accurately predict various situations in public services, such as predicting peak passenger flow periods for public transportation and fluctuating trends in energy consumption. At the same time, it can scientifically classify different service objects or events, such as classifying the types of failures in urban facilities, so that targeted maintenance measures can be taken. In addition, it can efficiently complete anomaly detection tasks and promptly identify abnormal conditions in public services, such as abnormal water quality and traffic flow, thereby providing strong data support for the optimization and decision-making of public services.
[0088] In some embodiments, the dynamic scheduling strategy that combines machine learning-based load prediction and a multi-objective optimization scheduling model dynamically allocates the data processing tasks to different subsystems, including: training a load prediction model and predicting the load trend of each system based on historical task execution data and real-time streaming data; constructing a multi-objective optimization scheduling model with the goals of minimizing response time, maximizing resource utilization, load balancing index, and task priority, while satisfying resource capacity, task dependency, and QoS constraints, and using the NSGA-III algorithm to generate a Pareto optimal scheduling solution; through a sliding window mechanism, dynamically adjusting the allocation of data processing tasks according to the predicted load.
[0089] Here, the load forecasting model training and trend forecasting process includes:
[0090] Data collection: Two types of core data are collected. The first is historical task execution data, which covers task type (such as compute-intensive, I / O-intensive), execution time (second-level accuracy), and resource consumption (CPU / memory / storage usage per unit time); the second is real-time streaming data, namely system status data, including CPU / memory utilization (samples per second) and network bandwidth (Mbps-level dynamic monitoring).
[0091] Feature Engineering: During data preprocessing, we perform one-hot encoding on task types and normalize execution time and resource consumption. We also aggregate statistics for real-time streaming data by time window (e.g., 5 minutes).
[0092] Model construction and training: The load forecasting model uses an LSTM-GRU hybrid architecture. Preprocessed data is fed into an LSTM layer (50 units) to capture time series dependencies, and then into a GRU layer (30 units) to enhance short-term dynamic features. Features are mapped to the load trend space using a two-layer fully connected network (ReLU activation). The model outputs load forecasts (expressed as a percentage of resource utilization) for each subsystem (such as computing, storage, and network) within the next 15 minutes. The mean squared error (MSE) loss function is used during training:
[0093]
[0094] Among them, y i For real load, is the predicted value.
[0095] The multi-objective optimization scheduling model construction process includes:
[0096] The objective function of minimizing response time Where M is the number of tasks, is the delay time of task j; the objective function of maximizing resource utilization Where K is the number of subsystems, is the used resources of subsystem k, is the total resource; the objective function of the load balancing index Among them L k is the load of subsystem k, is the average load; the objective function of task priority w j is the priority weight, P j The task completion status (0 or 1). Constraints include: resource capacity constraints Quality of Service Constraints: Task dependency constraint: If task j depends on task i, then the start time of task j is later than the end time of task i.
[0097] The process of generating the Pareto optimal scheduling solution using the NSGA-III algorithm is as follows:
[0098] S1, population initialization: randomly generate N scheduling schemes, each of which contains a task allocation matrix (M×K);
[0099] S2, non-dominated sorting: stratify the solutions according to the objective function values f1, f2, f3, and f4;
[0100] S3, reference point generation: Generate uniformly distributed reference points using the Das-Dennis method;
[0101] S4, environmental selection: select the next generation population based on the reference point and retain the Pareto frontier solution.
[0102] The final output is a Pareto optimal scheduling solution that includes task allocation, resource allocation, and delay prediction, and supports dynamic adjustment to cope with sudden load fluctuations.
[0103] For example, the sliding window mechanism updates the scheduling plan every 5 minutes. When the predicted load exceeds a threshold (such as 80%), the scheduling plan is re-evaluated, and tasks of the high-load subsystem are migrated to the low-load subsystem. At the same time, the load prediction model parameters are fine-tuned according to the actual task execution results (delay, resource utilization), forming a closed-loop optimization.
[0104] Through this embodiment, the multi-system dynamic scheduling method provided by the present invention can dynamically balance response time, resource utilization and load balancing, while meeting task priorities and resource constraints, and realize efficient multi-system collaborative scheduling.
[0105] In some embodiments, the method further includes: real-time monitoring of environmental changes and task execution status, and triggering model updates and policy adjustments.
[0106] Here, the system automatically identifies key thresholds or pattern deviations by real-time monitoring of external environmental changes (such as dynamic data fluctuations, abnormal event triggering, etc.) and internal task execution status (such as progress deviation, resource utilization, etc.), and immediately triggers model parameter optimization and strategy dynamic adjustment mechanisms to ensure the timeliness and adaptability of system responses.
[0107] In some embodiments, the collaborative work between multiple subsystems is achieved through the system adaptation interface and data routing and forwarding mechanism, including: developing standardized interfaces for different subsystems to support task submission, status query, and resource application operations; using Kafka message queues to achieve asynchronous transmission of task metadata and intermediate results; describing task dependencies through directed acyclic graphs, and using the Airflow scheduling framework to achieve task cascade triggering.
[0108] A set of standardized interfaces has been developed for different subsystems. This interface is highly versatile and compatible, fully supporting key operations such as task submission, status query, and resource request. The task submission interface ensures that each subsystem can easily and accurately transmit task information to the system core; the status query interface allows subsystems to monitor task processing progress and current status in real time; and the resource request interface ensures that subsystems can obtain the required computing, storage, and other resources in a timely manner during operation.
[0109] For data transmission, the Kafka message queue was introduced. Kafka, with its advantages of high throughput, low latency, and scalability, enables asynchronous transmission of task metadata and intermediate results. Task metadata includes basic task information and configuration parameters, while intermediate results are the interim data generated during task execution. Through the Kafka message queue, this data can flow efficiently and stably between subsystems, avoiding data congestion and delays during data transmission and improving the overall system responsiveness.
[0110] To clearly describe the dependencies between tasks, a directed acyclic graph (DAG) is used for modeling. A DAG can intuitively display the order and dependency logic between tasks, making task scheduling and execution more reasonable and efficient. On this basis, the Airflow scheduling framework is used. Airflow has powerful task scheduling and monitoring capabilities, and can implement cascading task triggering based on the task dependencies described by the directed acyclic graph. When a task is completed, Airflow automatically triggers the subsequent tasks it depends on, ensuring that the entire task process proceeds in an orderly manner according to the predetermined logical sequence, thereby achieving efficient collaboration between multiple subsystems.
[0111] The multi-system dynamic scheduling processing method for the above-mentioned public service real-time streaming data is described below with reference to a specific embodiment. However, it should be noted that this specific embodiment is only for better illustrating the present invention and does not constitute an improper limitation to the present invention.
[0112] The specific embodiment of this invention constructs an overall architecture for multi-system collaboration, including a data source layer, a data acquisition and preprocessing layer, a processing and analysis layer, a multi-system scheduling and collaboration layer, and an application and service layer. Each layer has a clear division of labor. Through system adaptation interfaces and data routing and forwarding mechanisms, seamless flow and interaction between heterogeneous systems are achieved, ensuring that data can be efficiently processed and utilized across multiple systems.
[0113] First, build the experimental environment.
[0114] Hardware platform: Build an experimental cluster consisting of multiple servers, edge computing devices, and simulated data sources. The servers are equipped with high-performance CPUs, large-capacity memory, and high-speed network interfaces. The edge computing devices use embedded hardware with certain computing capabilities.
[0115] Software environment: Deploy various open source software and technical frameworks involved in the above overall architecture, including message queues, stream computing engines, batch processing systems, machine learning platforms, etc., to build a complete public service real-time streaming data processing experimental platform.
[0116] Data simulation generator: Develop data simulation generators for different public service scenarios, which can generate simulated real-time streaming data according to preset data format, generation rate and business logic for experimental testing and performance evaluation.
[0117] Secondly, design the experimental plan.
[0118] Performance indicator selection: Determine a series of key performance indicators, such as data collection latency, data processing throughput, task scheduling response time, system resource utilization, data consistency accuracy, etc., to quantitatively evaluate the performance of the proposed solution.
[0119] Comparative experiment setup: The solution of the present invention is compared with traditional static scheduling, single system processing and other methods. Under the same experimental environment, the performance indicator data is recorded for different data scales, data types and business load conditions, and the advantages and improvement effects of this solution are analyzed.
[0120] Finally, the experimental results and analysis are presented.
[0121] In terms of improving data processing performance: Experimental results show that the solution of the present invention has significant effects on reducing data acquisition delays, improving data processing throughput, and shortening task scheduling response time compared with traditional methods. This is mainly due to the reasonable allocation and optimization of resources by the dynamic scheduling algorithm and the acceleration of data preprocessing by edge computing.
[0122] Efficient utilization of system resources: Through the flexible resource allocation strategy, the system resource utilization rate has increased by an average of 70%. When the business load fluctuates greatly, resource allocation can be automatically adjusted to avoid resource waste and performance bottlenecks, thereby reducing operating costs.
[0123] This invention provides a multi-system dynamic scheduling and processing method for real-time streaming data from public services. This method can handle a variety of data types, including structured, semi-structured, and unstructured data generated by diverse data sources such as traffic sensors, medical IoT devices, and video surveillance cameras. Furthermore, this solution is not only applicable to specific public service areas such as traffic management, healthcare, and urban security, but also allows for functional expansion and performance upgrades based on actual needs, adapting to evolving business demands and technological developments.
[0124] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present invention. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention. The serial numbers of the above-mentioned embodiments of the present invention are for description only and do not represent the advantages and disadvantages of the embodiments.
[0125] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0126] In the several embodiments provided herein, it should be understood that the disclosed methods can be implemented in other ways. The methods disclosed in the several method embodiments provided herein can be combined arbitrarily, unless they conflict, to produce new method embodiments. The features disclosed in the several method embodiments provided herein can be combined arbitrarily, unless they conflict, to produce new method embodiments.
[0127] The above description is merely an embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A multi-system dynamic scheduling and processing method for real-time streaming data of public services, characterized in that: include: Data collection and preprocessing layer: collects real-time streaming data from multiple data sources through a distributed data collection framework and performs preprocessing and transmission; wherein, the multiple data sources include various public service data sources; Processing and analysis layer: Use the real-time stream computing engine to perform real-time stream computing on pre-processed data. At the same time, combine with the batch processing system to conduct in-depth analysis of historical data and trigger data processing tasks. Multi-system scheduling and collaboration layer: The task scheduling center uses a dynamic scheduling strategy that combines machine learning-based load prediction and a multi-objective optimization scheduling model to dynamically allocate data processing tasks to different subsystems. The system adaptation interface and data routing and forwarding mechanism enable collaborative work between multiple subsystems. Application and service layer: Based on the processing and analysis results of each subsystem, it provides business applications for specific public service areas, as well as decision support systems and data sharing and open platforms.
2. The multi-system dynamic scheduling processing method according to claim 1, characterized in that: The distributed data collection framework is used to collect real-time streaming data from multiple data sources and perform pre-processing, including: Use Kafka as a message queue to collect real-time streaming data from different data sources and perform preliminary format unification and standardization; Design an intelligent edge computing task allocation algorithm based on the characteristics of the data source and service requirements, and dynamically deploy edge computing nodes to perform real-time preprocessing of some data; According to the data type, the preprocessed data is cached in different cache systems and uploaded to the data center.
3. The multi-system dynamic scheduling processing method according to claim 1, characterized in that: The real-time stream computing engine is used to perform real-time stream computing on the pre-processed data, and the batch processing system is combined to conduct in-depth analysis of historical data to trigger data processing tasks, including: Use Flink's CEP library to define event patterns and detect key events that match specific patterns from the continuous data stream transmitted by the data acquisition and preprocessing layers. Use Flink's window function to aggregate statistics within the time window of real-time streaming data corresponding to key events; Use Hadoop / Hive to analyze historical accident data, predict risk areas, and obtain batch processing results; Based on the Kappa architecture, real-time stream and batch processing results are unified through middleware to obtain stream-batch integrated data processing tasks.
4. The multi-system dynamic scheduling processing method according to claim 2, characterized in that: According to the characteristics of the data source and service requirements, an intelligent edge computing task allocation algorithm is designed to dynamically deploy edge computing nodes to perform real-time preprocessing of some data, including: For first-level tasks, priority is given to completing processing at edge computing nodes; For the second-level tasks, after preliminary processing at the edge computing node, the results or key data are uploaded to the central system for further analysis; For the third-level tasks, they are directly uploaded to the central system for processing; wherein, the real-time requirements of the first-level tasks, the second-level tasks, and the third-level tasks decrease in sequence and the amount of calculation increases in sequence.
5. The multi-system dynamic scheduling processing method according to claim 4, characterized in that: The method further comprises: Integrate machine learning libraries and AI model training and inference tools in edge computing nodes, and use historical and real-time data to predict, classify, and detect anomalies in public service scenarios.
6. The multi-system dynamic scheduling processing method according to any one of claims 1 to 5, characterized in that: The dynamic scheduling strategy, which combines machine learning-based load forecasting with a multi-objective optimization scheduling model, dynamically allocates the data processing tasks to different subsystems, including: Based on historical task execution data and real-time streaming data, train load forecasting models and predict the load trends of each system; A multi-objective optimization scheduling model is constructed with the goals of minimizing response time, maximizing resource utilization, load balancing index, and task priority, while satisfying resource capacity, task dependency, and QoS constraints. The non-dominated sorting genetic algorithm NSGA-III is used to generate a Pareto optimal scheduling solution. Through the sliding window mechanism, the allocation of data processing tasks is dynamically adjusted according to the predicted load.
7. The multi-system dynamic scheduling processing method according to claim 6, characterized in that: The method further comprises: Monitor environmental changes and task execution status in real time, triggering model updates and strategy adjustments.
8. The multi-system dynamic scheduling processing method according to any one of claims 1 to 5, characterized in that: The collaborative work between multiple subsystems is achieved through the system adaptation interface and data routing and forwarding mechanism, including: Develop standardized interfaces for different subsystems to support task submission, status query, and resource application operations; Use Kafka message queue to realize asynchronous transmission of task metadata and intermediate results; Task dependencies are described through a directed acyclic graph, and the Airflow scheduling framework is used to implement task cascade triggering.
Citation Information
Patent Citations
Factory electric energy management and control system and method based on side-cloud cooperation
CN111144715A
Big data support management system for city-level data center station
CN112148718A
Stream batch integrated distributed method and system for nuclear engineering big data analysis
CN117950804A
Multi-source computing power data integration and intelligent scheduling system and method
CN118916147A
Stream batch data rapid fusion method
CN119026070A