Multi-engine collaborative data processing and task scheduling system and method
Through the multi-engine coordinated data processing and task scheduling system, parallel processing, real-time monitoring and automatic cleaning are realized, which solves the problems of inefficiency and low resource utilization of a single engine system, improves the stability and user experience of data processing, and ensures efficient data management and analysis.
Patent Information
- Application Number
- CN202411668949.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2044-11-21
AI Technical Summary
In the prior art, a single engine system cannot fully utilize the performance of multi-core processors, resulting in slow processing speed, limited scalability, low efficiency of task scheduling and monitoring relying on manual management, inaccurate data processing, lack of unified data management portal, independent systems cannot work together, low resource utilization, unintuitive user interface, low data cleaning relying on manual efficiency, and inability to dynamically adjust task scheduling.
A multi-engine collaborative data processing and task scheduling system is adopted, including a task scheduler, a data processing engine, a monitoring module, a data cleaning module and a unified data management unit, to realize parallel processing, real-time monitoring, automatic data cleaning and unified management, and provide a visual operation interface through an abnormal feature recognition library and self-healing mechanism.
It significantly improves data processing efficiency, ensures the stability and reliability of task processing, improves data quality and user experience, enhances the flexibility and scalability of the system, can quickly respond to abnormal situations, reduce human errors, and improve resource utilization.
Smart Images

Figure CN119597423B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of collaborative data processing, and in particular relates to a multi-engine collaborative data processing and task scheduling system and method. Background Art
[0002] Traditional single-engine systems use a single data processing engine to perform all tasks. This approach may not fully utilize the performance of multi-core processors, resulting in slow processing speeds and limited scalability. With the surge in data volumes and the increasing complexity of computing tasks, manual task management can no longer meet the needs of modern businesses. Current technologies have several shortcomings:
[0003] Manual task scheduling and monitoring: This relies on operations personnel to manually start, monitor, and adjust tasks. This approach is error-prone, inefficient, and difficult to manage when there are a large number of tasks. Batch data processing: This processes data on a regular basis (such as nightly or weekly) rather than in real time or near real time. This can lead to data processing delays and is not suitable for scenarios with high timeliness requirements. No data cleansing or a simplified data cleansing process: This involves not performing data cleansing after processing, or implementing a relatively simple data cleansing process. This can result in low data quality and affect the final data analysis results. Independent data import and quality check tools: Independent tools are used to import data and check data quality rather than operating through a unified interface. This increases operational complexity and learning costs, and reduces work efficiency. Rule-based system: This uses fixed rules for task scheduling and data processing rather than dynamically adjusting based on the real-time status of the system. Such a system is not flexible enough to adapt to changing loads and demands.
[0004] No monitoring module: No task execution status monitoring is implemented, or only basic status viewing functions are provided. This will make the system lack the ability to automatically intervene and recover when an anomaly occurs. Restrictive user interface: Providing a user interface with limited functions or non-intuitiveness makes it difficult for operation and maintenance personnel to efficiently manage data and check quality. Non-automated data cleaning: Relying on manual data cleaning, which is not only inefficient but also prone to human errors. Independent systems without collaborative work: Each system runs independently and cannot work together, resulting in low resource utilization and low data processing efficiency. Script-based system: Relying on scripts to manage task scheduling and data processing, which requires operation and maintenance personnel to have programming skills and is difficult to adapt to rapidly changing business needs. Summary of the Invention
[0005] This invention proposes a multi-engine collaborative data processing and task scheduling system and method, addressing the existing problem of delayed data processing caused by untimely task initiation. The lack of automated monitoring during task execution prevents timely detection and resolution of anomalies. After task completion, manual data cleaning is required, resulting in low efficiency. Operations and maintenance personnel lack a unified data management portal, making data quality checks cumbersome. Interfaces and manually curated data must be manually imported into the database, impacting work efficiency.
[0006] The technical solution of the present invention is implemented as follows: a multi-engine collaborative data processing and task scheduling system includes a task scheduler, a data processing engine, a monitoring module, a data cleaning module and a unified data management unit. The task scheduler receives uploaded task requests, allocates scheduling tasks according to the task requests, and schedules the tasks to the corresponding data processing engines according to preset scheduling rules;
[0007] There are several data processing engines, and the several data processing engines perform parallel processing at the same time. Each data processing engine processes the corresponding task independently, and each data processing engine is associated with at least three other data processing engines;
[0008] The monitoring module monitors the tasks in real time and sets an abnormal feature recognition library in the monitoring module. The status of the monitored tasks is monitored through the pre-stored data in the abnormal feature recognition library. When abnormal features appear, the monitoring data upload frequency is increased and marked. When the abnormal features exceed the set threshold, automatic repair is started and the tasks are reset and updated synchronously.
[0009] The data cleaning module is used to summarize the task information after the data processing engine completes the task, automatically clean and remove noise from the data of the completed task, and import the cleaned data into the unified data management unit;
[0010] The unified data management unit is provided with an operation interface, which displays the task data after cleaning by the data cleaning module and the cleaning parameters of the data cleaning module, compares the monitoring data and cleaning data of each task, generates a quality inspection report in the operation interface, and adjusts the threshold parameters in the monitoring module according to the quality inspection report.
[0011] Compared with existing technologies, this multi-engine collaborative data processing and task scheduling system shows significant advantages and innovations.
[0012] Existing data processing systems often rely on a single data processing engine or linear processing model, lacking parallel processing capabilities. This results in slow processing speed and low efficiency when dealing with large amounts of data. However, this system utilizes multiple data processing engines to enable parallel processing of tasks, with each engine independently executing its corresponding task, fundamentally improving data processing efficiency. This parallel processing design not only significantly reduces data processing time but also effectively addresses the challenges of big data environments and improves system scalability.
[0013] The introduction of a task scheduler enables the system to intelligently allocate tasks based on real-time task requests. Existing technologies often rely on static scheduling strategies and are unable to flexibly respond to dynamically changing task demands. However, our system's task scheduler can flexibly adjust task allocation based on uploaded task requests and pre-set scheduling rules, improving resource utilization and ensuring that different tasks are efficiently executed on the appropriate engine.
[0014] The design of a monitoring module is relatively uncommon in existing technologies. Many systems lack real-time monitoring of task status, making it difficult to promptly identify and address anomalies. This system, by establishing an anomaly signature recognition library, monitors task status in real time and dynamically adjusts the frequency of monitoring data uploads when anomalies emerge, ensuring a timely response to potential issues. When anomalies exceed a set threshold, the system automatically initiates a repair mechanism and resets the task. This self-healing capability significantly enhances system stability and reliability.
[0015] After the task is completed, the data cleaning module aggregates and removes noise from the data to ensure the quality of the output data. In existing technologies, data cleaning is often an independent step that lacks close integration with the data processing process. However, this system closely integrates data cleaning with task processing to ensure that the cleaned data can be seamlessly connected to the unified data management unit, providing a high-quality data foundation for subsequent data analysis and decision-making. Finally, the design of the unified data management unit provides a visual operation interface, allowing users to intuitively view the cleaned data and monitoring data and generate quality inspection reports. This design improves the system's operability and user experience, allowing users to more conveniently perform data analysis and adjust monitoring parameters.
[0016] As a preferred embodiment, the task scheduler performs preliminary classification according to task priority, type and required resources, calls the current load status of the system, allocates tasks according to the load status, and sends classification information to the data processing engine for associated engine mobilization.
[0017] As a preferred embodiment, the data processing engine performs real-time data transmission through data nodes when performing association. A processing range interval is separately set for each data processing engine. When the processing range exceeds the set interval, the excess data is split and calculated by the associated data processing engine using the data node, and feedback information from the associated data processing engine is received.
[0018] As a preferred embodiment, the abnormal feature recognition library in the monitoring module records the abnormal feature values corresponding to different tasks. When the data processing engine receives the task, the monitoring module calibrates the specific data interval of the data processing engine. When the data processing engine shows a characteristic data interval, the monitoring data upload frequency is increased and marked. When the abnormal feature exceeds the set threshold, it is repaired and reset.
[0019] As a preferred embodiment, the data cleaning module obtains the data generated after the task is completed from the data processing engine. The data includes task status, execution time and resource usage. The data cleaning module performs a preliminary evaluation on the obtained data, identifies the noise data and outliers therein, and filters out the expected data points according to the set rules, eliminates the abnormal data, and fills in the missing values. At the same time, it corrects the erroneous data and synchronously completes the data format unification.
[0020] A multi-engine collaborative data processing and task scheduling method, the method comprising the following steps:
[0021] S1: Receive uploaded task requests through the task scheduler, assign scheduling tasks according to the task requests, and schedule the tasks to the corresponding data processing engine according to the preset scheduling rules;
[0022] S2: Parallel processing is performed on several data processing engines simultaneously. Each data processing engine processes the corresponding task independently, and each data processing engine is associated with at least three other data processing engines.
[0023] S3: The monitoring module monitors tasks in real time and sets up an abnormal feature recognition library in the monitoring module. The status of the monitored tasks is monitored using the pre-stored data in the abnormal feature recognition library. When abnormal features appear, the monitoring data upload frequency is increased and marked. When the abnormal features exceed the set threshold, automatic repair is initiated and the task is reset and updated synchronously.
[0024] S4: The data cleaning module receives the task completion information from the data processing engine, summarizes the task information, automatically cleans and removes noise from the data of the completed tasks, and imports the cleaned data into the unified data management unit;
[0025] S5: The unified data management unit is provided with an operation interface, which displays the task data after cleaning by the data cleaning module and the cleaning parameters of the data cleaning module, compares the monitoring data and the cleaning data of each task, generates a quality inspection report in the operation interface, and adjusts the threshold parameters in the monitoring module according to the quality inspection report.
[0026] After adopting the above technical solution, the beneficial effects of the present invention are: through parallel processing and intelligent scheduling, the system significantly improves data processing efficiency and can complete large-scale data processing tasks in a short period of time. This high efficiency enables enterprises to gain data insights more quickly and make timely decisions, thereby gaining an advantage in the highly competitive market. The combination of real-time monitoring and automatic repair mechanisms ensures the stability and reliability of the task processing process. When the system detects an anomaly, it can respond quickly, reducing potential risks and losses, and enhancing system security. The effective integration of the data cleaning module ensures the high quality of output data, reduces decision-making errors caused by data problems, and thus improves the credibility of data analysis. The visual operation interface provided by the unified data management unit enhances user control and understanding of the system, making it easier for users to manage data and adjust monitoring parameters. All of this makes the system show greater flexibility and scalability in practical applications, providing strong support for enterprises in data processing and management. In short, this innovative data processing and task scheduling system has promoted the development of data management technology and improved the efficiency and security of enterprises in the process of digital transformation. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0030] Example:
[0031] A multi-engine collaborative data processing and task scheduling system includes a task scheduler, a data processing engine, a monitoring module, a data cleaning module, and a unified data management unit. The task scheduler receives uploaded task requests, allocates scheduling tasks based on the task requests, and schedules the tasks to the corresponding data processing engine according to preset scheduling rules.
[0032] There are several data processing engines, and the several data processing engines perform parallel processing at the same time. Each data processing engine processes the corresponding task independently, and each data processing engine is associated with at least three other data processing engines;
[0033] The monitoring module monitors the tasks in real time and sets an abnormal feature recognition library in the monitoring module. The status of the monitored tasks is monitored through the pre-stored data in the abnormal feature recognition library. When abnormal features appear, the monitoring data upload frequency is increased and marked. When the abnormal features exceed the set threshold, automatic repair is started and the tasks are reset and updated synchronously.
[0034] The data cleaning module is used to summarize the task information after the data processing engine completes the task, automatically clean and remove noise from the data of the completed task, and import the cleaned data into the unified data management unit;
[0035] The unified data management unit is provided with an operation interface, which displays the task data after cleaning by the data cleaning module and the cleaning parameters of the data cleaning module, compares the monitoring data and cleaning data of each task, generates a quality inspection report in the operation interface, and adjusts the threshold parameters in the monitoring module according to the quality inspection report.
[0036] The core of the system is the task scheduler, which receives uploaded task requests. Users or other systems submit task requests to the task scheduler, which then allocates processing tasks based on the nature, priority, and resource requirements of the task. This process follows pre-set scheduling rules, such as task type, data volume, and processing time, ensuring that tasks are efficiently assigned to the appropriate data processing engine. This intelligent scheduling mechanism improves resource utilization and ensures the system's flexibility and efficiency in handling diverse tasks.
[0037] The system is equipped with multiple data processing engines. These engines can work in parallel, each responsible for handling specific tasks, quickly completing data processing. By introducing multiple data processing engines, the system can achieve high concurrency and improve overall processing capabilities. Furthermore, the interconnectedness between data processing engines increases system flexibility. For example, intermediate results completed by one engine can be directly used by other engines, reducing data transmission time and improving processing efficiency.
[0038] During data processing, the monitoring module monitors the status of tasks in real time. By setting up a library of abnormality signatures, the monitoring module can identify anomalies in task execution, such as excessive processing time or unexpected results. When abnormal signatures are detected, the system automatically increases the frequency of monitoring data uploads and marks these tasks for subsequent tracking and analysis. This real-time monitoring capability ensures that the system can promptly identify and address potential issues, reducing the risk of failure.
[0039] If detected abnormal characteristics exceed a set threshold, the system initiates an automatic repair mechanism. This may involve rescheduling tasks, adjusting resource allocation, or reprocessing data to ensure the system can quickly return to normal. Furthermore, task resetting and updating ensures data accuracy and consistency, preventing data loss due to interruptions or errors.
[0040] The introduction of a data cleaning module is another key component of the system. Its primary task is to aggregate task information after the data processing engine completes a task and automatically clean the processed data, removing noise and redundant information. This data cleaning process ensures the quality of the final data and enhances the credibility of subsequent analysis. The cleaned data is then imported into a unified data management unit, providing a reliable foundation for subsequent operations and analysis.
[0041] The unified data management unit serves as the control center for the entire system. It features a user interface that displays task data and corresponding cleaning parameters after completion of the data cleaning module. This interface provides intuitive data monitoring and management capabilities. Users can compare monitoring and cleaning data for each task and generate quality inspection reports. These reports provide clear feedback on task execution, data quality, and processing efficiency. Based on these quality inspection reports, the system intelligently adjusts threshold parameters in the monitoring module. This feedback mechanism enables the system to self-optimize based on historical execution and data quality, continuously improving overall performance and reliability.
[0042] The setting of this structure and operation mode is mainly reflected in the following aspects: Resource optimization: through the combination of task scheduler and multi-data processing engine, efficient resource utilization is achieved, and idle time and resource waste are reduced; Flexible processing: The intelligent task scheduling enables the system to adapt to tasks of different types and priorities, and improves processing capabilities; Real-time monitoring: The introduction of the monitoring module ensures the transparency and timeliness of task processing, and can quickly respond to abnormal situations; Data quality assurance: The setting of the data cleaning module makes the final data analysis results more reliable and reduces the possibility of data errors; User-friendliness: The unified data management unit provides an intuitive operation interface, allowing users to easily monitor and manage data; Self-optimization: Through feedback from monitoring data and quality inspection reports, the system can continuously adjust and optimize its own performance, making it more efficient in actual applications.
[0043] The task scheduler performs preliminary classification according to task priority, type and required resources, calls the current load status of the system, allocates tasks according to the load status, and sends classification information to the data processing engine for correlation engine mobilization.
[0044] Task scheduling often uses static or simple priority mechanisms, lacking dynamic monitoring and adjustment of the overall system load status. This leads to uneven task distribution under high load conditions, affecting system efficiency and response time. However, this task scheduler design closely integrates task scheduling with the actual usage of system resources by invoking the system load status in real time. This allows for flexible adjustment of task allocation strategies under varying load conditions, thereby optimizing resource utilization and improving system performance.
[0045] The multi-dimensional classification mechanism of this task scheduler is also a significant difference from the existing technology. The existing technology usually only focuses on the priority or type of the task, and lacks comprehensive consideration of the resources required for the task, which may lead to resource bottlenecks in the task allocation process and affect the overall processing capacity of the system. However, this application comprehensively considers the priority, type and required resources of the task to ensure that each task can obtain the most appropriate processing resources when allocated, thereby improving the intelligence level and flexibility of task scheduling. This comprehensive classification method enables the task scheduler to better cope with complex and dynamically changing workloads, ensuring that the system can maintain an efficient operating state in different scenarios.
[0046] After completing task assignment, this task scheduler sends classification information to the data processing engine to mobilize the association engine. The design of this process also reflects the difference from existing technologies. Traditional task scheduling systems often lack subsequent tracking and adjustment mechanisms after task assignment, which may cause delays or waste of resources during task execution. However, this design can mobilize the association engine in a timely manner by sending classification information in real time, ensuring that the execution of tasks can adapt to the dynamic changes in resources. This mechanism not only improves the system's response speed, but also enhances the reliability and accuracy of task execution, allowing the system to maintain efficient operation in complex working environments.
[0047] When performing association, the data processing engine performs real-time data transmission through the data nodes. A processing range interval is independently set for each data processing engine. When the processing range exceeds the set interval, the excess data is split and calculated by the associated data processing engine using the data nodes, and feedback information from the associated data processing engine is received.
[0048] The data processing engine is designed to realize real-time data transmission through data nodes when performing data association, which shows obvious innovation and advantages compared with the existing technology. Traditional data processing systems usually adopt a centralized data flow method, lacking dynamic monitoring of data flow and processing capacity, resulting in bottlenecks in the system when facing large-scale data, and reduced processing efficiency. The data processing engine of the present application can effectively monitor and manage the load of each engine by setting an independent processing range interval for each engine, ensuring that the data always remains within the optimal range during the processing process, greatly improving the flexibility and response speed of data processing.
[0049] When the processing range exceeds the set interval, this data processing engine uses the design of splitting the excess data for calculation with the help of data nodes, which reflects the effective optimization of data processing capabilities. When processing excess data, the existing technology often adopts a simple discarding or queuing strategy, lacks effective use of data, and causes data delays or loss. In contrast, this application distributes excess data to the associated data processing engine for parallel processing by splitting the operation, which significantly enhances the system's ability and efficiency when processing large-scale data. This design enables the data processing engine to fully utilize the computing resources of each node, thereby maintaining the stability and efficiency of the system under high load conditions.
[0050] In terms of receiving feedback information, existing technologies generally lack an effective feedback mechanism, resulting in a weak correlation between the processing results and the input data, making it difficult to achieve dynamic adjustment. However, after completing the splitting operation, the data processing engine of the present application can promptly receive feedback information from the associated data processing engine. This design enables the entire data processing system to have adaptive capabilities. By analyzing the feedback information, the system can adjust the subsequent data processing strategy in real time to cope with the changing input data and processing requirements, thereby optimizing the overall data flow efficiency. This real-time feedback mechanism is relatively rare in the existing technology and further enhances the intelligence level of the system.
[0051] The abnormal feature recognition library in the monitoring module records the abnormal feature values corresponding to different tasks. When the data processing engine receives the task, the monitoring module calibrates the specific data interval of the data processing engine. When the data processing engine shows a characteristic data interval, the monitoring data upload frequency is increased and marked. When the abnormal feature exceeds the set threshold, it is repaired and reset.
[0052] Traditional monitoring systems often rely on static threshold settings and one-time data analysis, and are unable to identify and respond to abnormal situations in real time during task execution. This approach is prone to delayed response when faced with complex and dynamically changing environments, resulting in greater system losses. The monitoring module proposed in this application can record the abnormal feature values corresponding to different tasks in real time. By calibrating the specific data intervals of the data processing engine, the system can respond quickly when abnormal features appear.
[0053] When the data processing engine receives a task, the monitoring module immediately calibrates it to a specific data interval. This process ensures that the system can accurately monitor the range of possible abnormal data during operation, avoiding the monitoring blind spots caused by the lack of dynamic adjustment in existing technologies. By setting specific data intervals, the monitoring module can effectively improve its ability to identify abnormal characteristics, thereby enhancing the security and stability of the system.
[0054] When the monitoring module detects that the data processing engine's characteristic data enters a set abnormal range, it automatically increases the upload frequency of the monitoring data and marks it. This mechanism is uncommon in existing technologies. Traditional monitoring systems typically use a fixed upload frequency and lack the ability to quickly respond to abnormal situations. By dynamically adjusting the upload frequency, critical data can be recorded and analyzed promptly when an anomaly occurs, providing important information for subsequent troubleshooting and system optimization.
[0055] More importantly, when the monitoring module identifies anomaly characteristics exceeding a set threshold, it immediately initiates repair and reset mechanisms. This proactive repair measure is relatively rare in existing technologies. Many systems simply issue alerts when faced with anomalies and lack effective automatic repair capabilities. By integrating repair and reset functions, this application can take timely action after detecting a problem, reducing system downtime and improving system reliability and stability.
[0056] The data cleaning module obtains the data generated after the task is completed from the data processing engine. The data includes task status, execution time and resource usage. The data cleaning module performs a preliminary evaluation on the obtained data, identifies the noise data and outliers, and filters out the data points that meet the expectations according to the set rules, eliminates outliers, and fills in the missing values. At the same time, it corrects the erroneous data and simultaneously completes the data format unification.
[0057] The data cleaning module in this application is designed with full consideration for the efficiency and accuracy of data processing, aiming to improve data quality through systematic steps. In modern data processing and analysis, raw data often contains noise, outliers, and missing values. These inaccurate or incomplete data can directly affect the results of subsequent analysis and the effectiveness of decision-making. Therefore, designing an efficient data cleaning module is particularly important.
[0058] The data cleaning module obtains the raw data generated after completing tasks from the data processing engine, ensuring that the module can obtain the latest data relevant to actual operations. This design concept emphasizes the real-time and accuracy of data acquisition, enabling the cleaning module to process newly generated data in a timely manner, improving the responsiveness and flexibility of the entire system. Compared with existing technologies, many data cleaning solutions often rely on static data sets and lack the ability to update in real time. However, this application ensures the timeliness of data cleaning through dynamic data acquisition.
[0059] The data cleaning module performs a preliminary assessment of the acquired data and achieves basic control of data quality by identifying noisy data and outliers. This process is carried out through set rules and algorithms, ensuring the scientific nature and accuracy of the cleaning operation. In the existing technology, many data cleaning methods often lack systematic evaluation criteria, resulting in ineffective identification and processing of abnormal data. This application uses a clear evaluation mechanism to more accurately identify data points that do not meet expectations, thereby improving the overall quality of the data.
[0060] After selecting the expected data points, the data cleaning module also fills in missing values and corrects erroneous data. This design not only improves data integrity but also enhances the reliability of data analysis. Missing value filling and error correction are crucial steps in data cleaning. Existing technologies typically use simple mean filling or deletion methods, which can lead to information loss or deviation. This application, however, ensures the comprehensiveness and accuracy of the data through more complex and reasonable filling and correction methods.
[0061] The data cleansing module also simultaneously standardizes data formats, a crucial step in improving data usability. Data from different sources often have inconsistent formats, making subsequent processing and analysis difficult. By standardizing data formats, the module ensures consistency and compatibility during subsequent analysis, improving the efficiency of the entire data processing system. This feature is often overlooked in existing technologies, leading to unnecessary complications and obstacles in data use.
[0062] like Figure 1 As shown, a multi-engine collaborative data processing and task scheduling method includes the following steps:
[0063] S1: Receive uploaded task requests through the task scheduler, assign scheduling tasks according to the task requests, and schedule the tasks to the corresponding data processing engine according to the preset scheduling rules;
[0064] S2: Parallel processing is performed on several data processing engines simultaneously. Each data processing engine processes the corresponding task independently, and each data processing engine is associated with at least three other data processing engines.
[0065] S3: The monitoring module monitors tasks in real time and sets up an abnormal feature recognition library in the monitoring module. The status of the monitored tasks is monitored using the pre-stored data in the abnormal feature recognition library. When abnormal features appear, the monitoring data upload frequency is increased and marked. When the abnormal features exceed the set threshold, automatic repair is initiated and the task is reset and updated synchronously.
[0066] S4: The data cleaning module receives the task completion information from the data processing engine, summarizes the task information, automatically cleans and removes noise from the data of the completed tasks, and imports the cleaned data into the unified data management unit;
[0067] S5: The unified data management unit is provided with an operation interface, which displays the task data after cleaning by the data cleaning module and the cleaning parameters of the data cleaning module, compares the monitoring data and the cleaning data of each task, generates a quality inspection report in the operation interface, and adjusts the threshold parameters in the monitoring module according to the quality inspection report.
[0068] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A multi-engine collaborative data processing and task scheduling system, characterized in that: It includes a task scheduler, a data processing engine, a monitoring module, a data cleaning module and a unified data management unit. The task scheduler receives uploaded task requests, allocates scheduling tasks according to the task requests, and schedules the tasks to the corresponding data processing engine according to preset scheduling rules; There are several data processing engines, and the several data processing engines perform parallel processing at the same time. Each data processing engine processes the corresponding task independently, and each data processing engine is associated with at least three other data processing engines; The monitoring module monitors the tasks in real time and sets an abnormal feature recognition library in the monitoring module. The status of the monitored tasks is monitored through the pre-stored data in the abnormal feature recognition library. When abnormal features appear, the monitoring data upload frequency is increased and marked. When the abnormal characteristics exceed the set threshold, automatic repair is initiated and the task is reset and updated synchronously; The data cleaning module is used to summarize the task information after the data processing engine completes the task, automatically clean and remove noise from the data of the completed task, and import the cleaned data into the unified data management unit; The unified data management unit is provided with an operation interface, which displays the task data after cleaning by the data cleaning module and the cleaning parameters of the data cleaning module, compares the monitoring data and the cleaning data of each task, generates a quality inspection report in the operation interface, and adjusts the threshold parameters in the monitoring module according to the quality inspection report; When performing association, the data processing engine transmits data in real time through the data nodes. Each data processing engine is individually set with a processing range interval. When the processing range exceeds the set interval, the excess data is split and calculated by the associated data processing engine using the data nodes, and feedback information from the associated data processing engine is received. The abnormal feature recognition library in the monitoring module records the abnormal feature values corresponding to different tasks. When the data processing engine receives the task, the monitoring module calibrates the specific data interval of the data processing engine. When the data processing engine shows a characteristic data interval, the monitoring data upload frequency is increased and marked. When the abnormal feature exceeds the set threshold, it is repaired and reset.
2. The multi-engine collaborative data processing and task scheduling system according to claim 1, characterized in that: The task scheduler performs preliminary classification according to task priority, type and required resources, calls the current load status of the system, allocates tasks according to the load status, and sends classification information to the data processing engine for correlation engine mobilization.
3. The multi-engine collaborative data processing and task scheduling system according to claim 1, characterized in that: The data cleaning module obtains the data generated after the task is completed from the data processing engine. The data includes task status, execution time and resource usage. The data cleaning module performs a preliminary evaluation on the obtained data, identifies the noise data and outliers, and filters out the data points that meet the expectations according to the set rules, eliminates outliers, and fills in the missing values. At the same time, it corrects the erroneous data and simultaneously completes the data format unification.
4. A multi-engine collaborative data processing and task scheduling method, applied to the multi-engine collaborative data processing and task scheduling system according to claim 1, characterized in that: The method comprises the following steps: S1: Receive uploaded task requests through the task scheduler, assign scheduling tasks according to the task requests, and schedule the tasks to the corresponding data processing engine according to the preset scheduling rules; S2: Parallel processing is performed on several data processing engines simultaneously. Each data processing engine processes the corresponding task independently, and each data processing engine is associated with at least three other data processing engines. S3: The monitoring module monitors tasks in real time and sets up an abnormal feature recognition library in the monitoring module. The status of the monitored tasks is monitored using the pre-stored data in the abnormal feature recognition library. When abnormal features appear, the monitoring data upload frequency is increased and marked. When the abnormal features exceed the set threshold, automatic repair is initiated and the task is reset and updated synchronously. S4: The data cleaning module receives the task completion information from the data processing engine, summarizes the task information, automatically cleans and removes noise from the data of the completed tasks, and imports the cleaned data into the unified data management unit; S5: The unified data management unit is provided with an operation interface, which displays the task data after cleaning by the data cleaning module and the cleaning parameters of the data cleaning module, compares the monitoring data and the cleaning data of each task, generates a quality inspection report in the operation interface, and adjusts the threshold parameters in the monitoring module according to the quality inspection report.
Citation Information
Patent Citations
Data processing task dispatching system and method
CN104407919A
Task scheduling method and device of multi-engine computing system and storage medium
CN116820730A