Artificial intelligence task scheduling method and scheduling system based on big data analysis

By using an AI-powered task scheduling system based on big data analytics, performance parameters, resource status parameters, and load data are collected and processed in real time. Principal component analysis and support vector machine algorithms are combined to predict task execution time, and the isolated forest algorithm is used to detect anomalies. This solves the problems of unreasonable scheduling schemes and insufficient anomaly detection in traditional scheduling systems, thereby improving task execution efficiency and system stability.

CN120973502APending Publication Date: 2025-11-18NANJING SIRUILI TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511496930.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Traditional task scheduling systems cannot collect performance parameters and resource status parameters during task execution in real time and comprehensively. This makes it difficult for scheduling schemes to adapt to complex and ever-changing operating environments, unable to accurately predict task execution time, and lacking anomaly detection mechanisms, thus affecting system stability and resource utilization.

Method used

An AI-powered task scheduling system based on big data analytics is adopted, comprising a data acquisition module, a data processing module, a prediction module, and an anomaly detection module. It collects and processes performance parameters, resource status parameters, and real-time load data in real time. It predicts task execution time using principal component analysis and support vector machine algorithms, and uses the isolated forest algorithm to detect anomalies and dynamically adjust the scheduling scheme.

Benefits of technology

It enables accurate prediction of task execution time and anomaly detection, generates scientific and reasonable scheduling schemes, improves task execution efficiency and resource utilization, and enhances system stability and anti-interference capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973502A_ABST
    Figure CN120973502A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence scheduling, in particular to an artificial intelligence task scheduling method and scheduling system based on big data analysis. The system comprises a data acquisition module, a data processing module, a data prediction module, a scheduling generation module and an anomaly detection module. The data acquisition module acquires performance parameters and resource state parameters of task execution and real-time load data of a task queue in real time; the data processing module performs feature extraction on the acquired data to form a task feature matrix; the prediction module predicts task execution time according to the matrix and outputs a prediction execution time parameter; the scheduling generation module generates an initial task scheduling scheme in combination with the prediction parameters and the real-time load data; and the anomaly detection module monitors task execution anomaly indexes, and dynamically adjusts the initial scheme according to the indexes to obtain an optimized scheduling scheme. The system can adapt to task and resource states in real time, improve scheduling rationality, enhance system operation stability and adapt to complex artificial intelligence application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence scheduling technology, and in particular to an artificial intelligence task scheduling method and scheduling system based on big data analysis. Background Technology

[0002] In the context of rapid digitalization and intelligentization, various artificial intelligence application scenarios are constantly expanding, from large-scale data processing in the cloud to real-time task response on edge devices. The efficiency and rationality of task scheduling directly affect the overall system performance. As task types become increasingly complex and data volumes grow exponentially, traditional task scheduling methods are gradually revealing many shortcomings. Traditional scheduling systems often rely on preset fixed rules or simple priority ranking, lacking consideration for dynamic factors during task execution, such as performance fluctuations and real-time changes in hardware resource load. This makes scheduling schemes difficult to adapt to complex and ever-changing operating environments.

[0003] In practical applications, performance parameters (such as computing speed and data transfer rate) and resource status parameters (such as CPU utilization, memory usage, and remaining storage space) during task execution are constantly changing dynamically. Traditional scheduling systems cannot collect these parameters in real time and comprehensively, nor can they accurately analyze the real-time load data of the task queue (such as the number of tasks to be executed and the average complexity of tasks), thus failing to accurately grasp the matching relationship between the actual needs of task execution and the resource supply capacity.

[0004] Due to a lack of effective data processing and analysis methods, traditional scheduling systems are unable to extract key information reflecting the essential characteristics of tasks and resource operation patterns from massive amounts of parameter data, leading to significant deviations in task execution time estimations. This deviation directly affects the rationality of scheduling schemes; for example, assigning long-running tasks to nodes nearing resource saturation, or assigning resource-intensive tasks to weakly performing nodes, ultimately results in task execution delays, low resource utilization, and even serious problems such as task failures and system crashes.

[0005] Traditional scheduling systems lack effective anomaly detection mechanisms, making it impossible to promptly identify and handle abnormal situations during task execution, such as sudden resource failures, task execution errors, and data transmission interruptions. When anomalies occur, the system cannot quickly adjust its scheduling scheme and can only continue execution according to the original scheduling logic, leading to a wider impact of the anomaly and further reducing the system's stability and reliability. As artificial intelligence applications increasingly demand real-time performance, stability, and resource utilization, these shortcomings of traditional task scheduling systems have become significant bottlenecks restricting the further development and application of artificial intelligence technology. Summary of the Invention

[0006] The main objective of this invention is to provide an artificial intelligence task scheduling method and system based on big data analysis, aiming to solve the technical problems in the prior art.

[0007] This invention proposes an artificial intelligence task scheduling system based on big data analysis, comprising:

[0008] The data acquisition module is used to collect performance parameters and resource status parameters during task execution and obtain real-time load data of the task queue.

[0009] The data processing module is used to extract features from the performance parameters, resource status parameters, and real-time load data to obtain a task feature matrix;

[0010] The prediction module is used to predict the task execution time based on the task feature matrix and obtain the predicted execution time parameters.

[0011] The scheduling generation module is used to generate an initial task scheduling scheme based on the predicted execution time parameters and real-time load data.

[0012] The anomaly detection module is used to monitor abnormal indicators during task execution and dynamically adjust the initial task scheduling scheme based on the abnormal indicators to obtain an optimized scheduling scheme.

[0013] Preferably, the data acquisition module collects performance parameters and resource status parameters during task execution and obtains real-time load data of the task queue, including:

[0014] Perform performance monitoring on multiple computing nodes to obtain a set of performance parameters;

[0015] Collect resource utilization data to obtain a set of resource status parameters;

[0016] Real-time query of task queue length and processing speed to obtain real-time load data.

[0017] Preferably, the data processing module performs feature extraction on the performance parameters, resource status parameters, and real-time load data to obtain a task feature matrix, including:

[0018] The performance parameter set, resource status parameter set, and real-time load data are normalized to form standardized data;

[0019] Principal component analysis algorithm is used to reduce the dimensionality of standardized data and extract the main feature components;

[0020] The main feature components are combined into a task feature matrix.

[0021] Preferably, the prediction module predicts the task execution time based on the task feature matrix to obtain predicted execution time parameters, including:

[0022] Based on historical task execution data, collect sample task feature matrices and sample execution time parameters;

[0023] A task execution time predictor is constructed based on the support vector machine algorithm.

[0024] The task execution time predictor is trained and tested using the sample task feature matrix and sample execution time parameters;

[0025] The task feature matrix is ​​input into the task execution time predictor, which predicts the execution time parameters.

[0026] Preferably, the scheduling generation module generates an initial task scheduling scheme based on the predicted execution time parameters and real-time load data, including:

[0027] By combining the predicted execution time parameters and real-time load data, a priority score is calculated for each task;

[0028] The tasks are sorted according to their priority scores to obtain the task queue order;

[0029] Based on resource status parameters, computing nodes are allocated to tasks to form an initial task scheduling scheme.

[0030] Preferably, the anomaly detection module monitors anomaly indicators during task execution and dynamically adjusts the initial task scheduling scheme based on these indicators to obtain an optimized scheduling scheme, including:

[0031] Collect task execution logs in real time and extract abnormal indicators, including error rate and latency.

[0032] Use the isolated forest algorithm to detect outliers in anomaly indicators;

[0033] When an anomaly is detected, a dynamic adjustment mechanism is triggered;

[0034] Based on the severity of the anomalies, the task priority scores are recalculated, the task queuing order and resource allocation are adjusted, and an optimized scheduling scheme is generated.

[0035] Preferably, the anomaly detection module uses the isolated forest algorithm to detect outliers in the anomaly indicators, including:

[0036] Construct an isolated forest model using historical normal task execution data as training data;

[0037] Input real-time anomaly indicators into the isolated forest model and output anomaly scores;

[0038] Set an anomaly threshold; when the anomaly score exceeds the threshold, it is marked as an anomaly.

[0039] Preferably, the dynamic adjustment mechanism recalculates the task priority score based on the severity of the anomalies, including:

[0040] Assign severity weights based on the type and scope of impact of the anomalies;

[0041] The priority score is recalculated using a weighted formula by combining the predicted execution time parameter and real-time load data.

[0042] Task scheduling is adjusted based on the new priority scores.

[0043] Preferably, the system further includes:

[0044] The user interface module is used to display the initial task scheduling scheme and the optimized scheduling scheme, and to receive user feedback;

[0045] The feedback learning module updates the task execution time predictor and anomaly detection model based on user feedback.

[0046] Preferably, the present invention also includes an artificial intelligence task scheduling method based on big data analysis, the method comprising all the modules and method flow of the artificial intelligence task scheduling system based on big data analysis as described above.

[0047] The beneficial effects of this invention are as follows:

[0048] By setting up a data acquisition module, performance parameters and resource status parameters during task execution can be collected in real time and comprehensively. Simultaneously, real-time load data of the task queue can be obtained, overcoming the limitations of traditional scheduling systems where data collection is untimely and incomplete. This provides a sufficient and accurate data foundation for subsequent data analysis and scheduling decisions. Through continuous collection of this dynamically changing data, the system can monitor the actual progress of task execution and the real-time operating status of resources, avoiding scheduling decision deviations caused by missing or delayed data. This allows the scheduling scheme to better adapt to current task requirements and resource supply conditions.

[0049] The data processing module effectively extracts features from collected performance parameters, resource status parameters, and real-time load data to form a task feature matrix. This process filters key features that are representative and reflect the essential attributes of tasks and the operational patterns of resources from massive amounts of raw data, transforming scattered and disordered data into structured and valuable information. By constructing the task feature matrix, the system can more clearly grasp the differences between different tasks and the compatibility between tasks and resources, providing accurate analytical basis for subsequent task execution time prediction. This avoids the problem of inaccurate understanding of task characteristics caused by the inability of traditional scheduling systems to effectively process data, resulting in a deeper and more comprehensive understanding of tasks.

[0050] The prediction module predicts task execution time based on the task feature matrix, obtaining predicted execution time parameters. Compared to traditional scheduling systems that rely on experience or simple estimations, this significantly improves the accuracy of task execution time prediction. Accurate predicted execution time parameters help the scheduling generation module more rationally plan the execution order of tasks and resource allocation schemes. For example, based on the predicted execution time of different tasks, the module can rationally arrange the start order of tasks, avoiding situations where task queuing times are too long or resources are wasted due to incorrect execution time predictions. This results in a more orderly task execution rhythm and more efficient resource utilization.

[0051] The scheduling generation module combines predicted execution time parameters and real-time load data to generate an initial task scheduling scheme, ensuring the scientific validity and rationality of the initial scheme. This initial scheme fully considers the execution requirements of the tasks and the real-time load status of resources, and can find a good balance between task execution efficiency and resource utilization. It avoids the unreasonable scheduling schemes caused by traditional scheduling systems that ignore differences in task execution time or real-time resource load. For example, it avoids concentrating a large number of high-load tasks on the same resource node, or the unbalanced phenomenon of some resource nodes being idle while others are overloaded.

[0052] The anomaly detection module can monitor abnormal indicators during task execution in real time and dynamically adjust the initial task scheduling plan based on these indicators to form an optimized scheduling plan. This function effectively compensates for the lack of anomaly handling capabilities in traditional scheduling systems. When resource failures or task execution anomalies occur during task execution, the system can promptly detect the anomalies and quickly adjust the scheduling strategy. For example, it can reassign tasks on the faulty node to other normal nodes or adjust the execution order of tasks to avoid abnormal resources, thereby effectively reducing the impact of anomalies on task execution, reducing the probability of task execution delays or failures, ensuring the continuous and stable operation of the system, and improving the system's reliability and anti-interference capabilities. Attached Figure Description

[0053] Figure 1 This is a timing diagram of the artificial intelligence task scheduling system based on big data analysis described in this invention;

[0054] Figure 2 This is a flowchart of the data processing module;

[0055] Figure 3 This is a flowchart of the prediction module.

[0056] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0057] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0058] like Figure 1 As shown, this application provides an artificial intelligence task scheduling system based on big data analysis, including: a data acquisition module, a data processing module, a prediction module, a schedule generation module, and an anomaly detection module. Specific implementation details are as follows:

[0059] The data acquisition module is responsible for collecting performance parameters and resource status parameters during task execution, as well as acquiring real-time load data of the task queue. These parameters and data are collected in real time through sensors and monitoring tools. The data processing module performs feature extraction on the collected performance parameters, resource status parameters, and real-time load data to generate a task feature matrix. Feature extraction involves data cleaning and transformation. The prediction module uses the task feature matrix to predict task execution time and outputs predicted execution time parameters. This prediction is based on historical data trained by a machine learning model. The scheduling generation module combines the predicted execution time parameters and real-time load data to generate an initial task scheduling scheme, which includes task sorting and resource allocation strategies. The anomaly detection module monitors abnormal indicators during task execution, such as error rate and latency, and dynamically adjusts the initial task scheduling scheme based on these indicators to obtain an optimized scheduling scheme. The adjustment mechanism responds to system changes in real time to ensure that the scheduling scheme adapts to the task execution environment.

[0060] In one embodiment, Example 1: The data acquisition module collects performance parameters and resource status parameters during task execution and obtains real-time load data of the task queue. This module is deployed in a distributed computing environment to monitor multiple computing nodes in real time through performance monitoring tools. These monitoring tools use an agent model to deploy lightweight acquisition programs on each computing node. These programs run as daemons at the operating system level and obtain underlying hardware and runtime performance data through system calls and kernel interfaces. The performance parameter set includes core indicators such as CPU utilization, memory usage, and network bandwidth utilization. CPU utilization is obtained by parsing the / proc / stat file or using performance counters, which can reflect the processor's busyness and computing power allocation. Memory usage is recorded by monitoring the / proc / meminfo file or using the statistical information of the memory management unit, recording the usage of physical memory and virtual memory, as well as the allocation status of cache and buffer. Network bandwidth utilization is calculated by analyzing the packet transmission rate and error rate of the network interface card to calculate network throughput and latency indicators. These parameters are continuously collected in the form of time series and marked with precise timestamps and node identifiers. When collecting resource utilization data, the module accesses the database and application programming interface of the resource management system. The resource management system typically uses a container orchestration platform or cluster management framework such as Kubernetes or YARN. These systems provide rich metric export functions and data query interfaces. The module obtains cluster-level resource status information through RESTful API or gRPC calls. The set of resource status parameters includes the number of CPU cores allocated, memory capacity limits, disk storage space, and IOPS performance metrics. These parameters reflect the overall health status and available capacity of computing resources. The module also collects accelerator resource metrics such as GPU utilization and video memory usage to adapt to the special needs of artificial intelligence workloads. The resource status data is stored in a structured format and undergoes preliminary verification and filtering during the collection process to remove invalid or abnormal data points and ensure data quality and reliability.

[0061] When querying the length and processing speed of the task queue in real time, the module calls the query interface of the task queue management system. The task queue is usually implemented based on message middleware or distributed queue system such as RabbitMQ or Apache Kafka. These systems provide real-time monitoring and statistics functions. The module obtains the number of tasks to be processed in the queue, message backlog and consumer processing rate through the management interface. Real-time load data includes queue length indicators that reflect the current workload and processing backlog of the system. The processing speed indicator is derived by calculating the number of tasks completed per unit time, reflecting the system's throughput and processing efficiency. These data are sampled at a high frequency, usually reaching the second or even millisecond level update rate, to ensure that instantaneous changes and peak conditions of the load can be captured. The module integrates performance parameter sets, resource status parameter sets, and real-time load data into a unified structured dataset. The integration process includes data alignment, timestamp synchronization, and format standardization. Data alignment matches metrics from different sources within a unified timeframe using time windows. Timestamp synchronization uses the NTP protocol to ensure time consistency across all nodes. Format standardization converts outputs from different data sources into a unified JSON or Protocol Buffers format for easier subsequent processing and transmission. The integrated dataset contains complete contextual information such as task identifiers, node identifiers, acquisition time points, and data quality markers. These datasets are transmitted to the data processing module for further analysis via a message bus or data pipeline. Data transmission employs an asynchronous, non-blocking mechanism to avoid impacting system performance. The module uses high-performance message queues such as Apache Pulsar or RedisStreams as data transmission channels. These channels provide persistent storage and flow control to ensure no data loss or backlog during peak system loads. During transmission, data is fragmented and compressed to reduce network bandwidth consumption and improve transmission efficiency. The receiving end uses batch processing to increase data ingestion throughput. The entire acquisition process has fault tolerance and retry mechanisms; when a data source is temporarily unavailable, the module automatically caches data and resends it upon recovery.

[0062] The module's error handling mechanism includes data validation, anomaly detection, and automatic recovery. Data validation checks the integrity and rationality of the collected data, such as numerical range verification and data type checks. Anomaly detection identifies and marks or filters obviously erroneous data points using simple threshold rules. The automatic recovery function can restart and restore the previous collection state when the collection process crashes or the network is interrupted, ensuring the continuity and consistency of data collection. The module also provides configuration management functions that allow dynamic adjustment of collection frequency, data source address, and transmission parameters to adapt to different scales and types of computing environments. The collection of performance parameter sets adopts an efficient event-driven architecture to reduce resource overhead. The collection agent uses asynchronous I / O mechanisms such as epoll or IOCP to handle concurrent requests from multiple data sources. The collection of resource status parameters utilizes a caching mechanism to reduce frequent queries to the resource management system. The collection of real-time load data adopts a publish-subscribe model, actively pushing updates when the queue status changes. These optimization measures ensure that the data collection module operates efficiently while minimizing its impact on system performance. The module's output data contains rich metadata such as data collection confidence, time precision, and data source description, providing sufficient context for subsequent processing stages.

[0063] In one embodiment, Example 2: See Figure 2 and Figure 3The data processing module performs feature extraction operations on performance parameters, resource status parameters, and real-time load data. This module receives a structured dataset from the data acquisition module and first performs normalization processing to convert the raw data into standardized data. The normalization uses Min-Max scaling technology to perform linear transformations on parameters of different dimensions, so that all values ​​fall within a unified range of [0,1]. During the processing, abnormal values ​​that exceed the reasonable range are automatically identified and excluded. At the same time, missing data points are filled by interpolation between adjacent time points. The standardized data is organized into a two-dimensional table and stored in an in-memory database for fast access. The rows of the table correspond to the acquisition time points, the columns correspond to different parameter types, and each cell contains a standardized value and a timestamp index. Principal component analysis (PCA) is used to perform dimensionality reduction on standardized data. The algorithm first calculates the covariance matrix of the standardized data, then obtains eigenvectors and their corresponding eigenvalues ​​through eigenvalue decomposition. The eigenvectors are arranged in descending order of their eigenvalues, and the top k eigenvectors whose cumulative contribution rate meets the preset condition are selected as the principal feature components. During the dimensionality reduction process, the data reconstruction error is automatically calculated and the number of principal components is dynamically adjusted to ensure that sufficient information is retained while reducing dimensionality. The principal feature components are encapsulated as transformation matrices and stored in the model file. These transformation matrices can be applied to subsequent real-time data stream processing. The dimensionality-reduced data points are mapped to a new feature space to form a low-dimensional representation, and these data points carry the main variation patterns of the original data. The main feature components are combined into a task feature matrix. The matrix construction process is arranged according to the time sequence of task execution. Each task instance corresponds to a row in the matrix, and each main feature component corresponds to a column in the matrix. The matrix cells are filled with the dimensionality-reduced feature values. The task feature matrix is ​​stored in memory as a multidimensional array, with task identifiers and timestamp metadata attached. During the matrix generation stage, data alignment is performed to ensure that the feature dimensions of different tasks are completely consistent. For multidimensional data generated by parallel tasks, the number of matrix columns is expanded by feature concatenation. The final output task feature matrix is ​​transmitted to the prediction module in binary serialization format, while a text-formatted snapshot is retained for debugging and analysis.

[0064] The prediction module predicts task execution time based on the task feature matrix. This module collects sample data from a historical task database. The sample task feature matrix contains representative task records from the past six months. The sample execution time parameters are accurate to the millisecond level and have undergone data cleaning. The historical database uses a distributed storage architecture to support high-concurrency queries. The data extraction process uses time range filters and task type selectors to ensure the relevance and timeliness of the sample data. The extracted sample set is divided into training and testing subsets according to time order, with the division ratio following preset rules and maintaining a balanced data distribution. A task execution time predictor is built based on the Support Vector Machine (SVM) algorithm. The algorithm implementation uses an optimized version provided by an open-source machine learning library. The kernel function type is configured as a radial basis function, and initial hyperparameter values ​​are set. During model initialization, computational resources, including memory buffers and processor cores, are allocated. The SVM model uses batch training mode to handle large-scale datasets. During training, the learning rate and regularization parameters are dynamically adjusted to optimize convergence speed. After each iteration, the loss function value is calculated, and overfitting is checked. Training automatically stops when the validation set accuracy no longer improves for several consecutive cycles. During model testing, an independent sample set is used to evaluate prediction accuracy, and the prediction error distribution and statistical characteristics are recorded. The task execution time predictor is trained and tested using the sample task feature matrix and sample execution time parameters. The training process uses mini-batch gradient descent to update the model parameters. Each batch processes a fixed number of sample matrix rows. The training loop includes three main stages: forward propagation to calculate the predicted value, loss function to calculate the error, and backpropagation to update the weights. In the testing stage, the model parameters are frozen and the input is the test sample feature matrix. The predicted execution time is compared with the actual value to calculate the evaluation index. The evaluation results generate a detailed report including statistics such as mean absolute error, root mean square error, and correlation coefficient. The model is marked as usable after its performance reaches a preset threshold. The real-time generated task feature matrix is ​​input into the task execution time predictor. The predictor loads the pre-trained model file and initializes the inference environment. The input data undergoes the same preprocessing process as the training data, including dimension verification and format conversion. The prediction execution process adopts a single-record streaming processing mode. Each task feature vector is independently calculated by the model. The output layer generates floating-point values ​​to represent the prediction execution time parameters. The prediction parameters are supplemented with confidence scores to reflect the model's certainty about the prediction. The entire prediction process is executed on a dedicated computing unit with hardware acceleration enabled. The prediction results are bound to a unique task identifier and then transmitted to the scheduling generation module. The predictor has a built-in monitoring function to record the latency and resource consumption of each prediction.

[0065] Taking the AI ​​task scheduling of an e-commerce platform during a promotional event as an example, the system needs to handle a large number of real-time recommendation tasks and inventory update tasks. The data acquisition module has collected performance parameters from 200 computing nodes, including CPU utilization fluctuating between 65% and 85%, memory usage ranging from 50% to 70%, network bandwidth utilization varying between 40% and 60%, and resource status parameters showing the overall resource utilization level of the cluster. Among them, the GPU nodes generally have high memory utilization, reaching over 75%. Real-time load data shows that there are about 1,500 tasks waiting to be processed in the task queue, with an average processing speed of 120 tasks per second. After receiving the raw data, the data processing module initiates a normalization process. The original CPU utilization rate of 85% is converted to 0.85, the memory utilization rate of 70% is converted to 0.70, and the network bandwidth utilization rate of 60% is mapped to 0.60. After linear transformation, all values ​​fall into the range [0,1]. During the normalization process, it is found that the memory utilization rate of some nodes is abnormally high, reaching 95%. These outliers are automatically marked and replaced with the average value of adjacent time points, 0.72. The processed standardized data forms a two-dimensional matrix and is stored in the memory database. The number of rows in the matrix corresponds to the number of time series points, and the number of columns represents the number of parameter types. Principal component analysis (PCA) is used to reduce the dimensionality of standardized data. The algorithm calculates the covariance matrix in a 200-dimensional parameter space, obtains eigenvectors and eigenvalues ​​through eigenvalue decomposition, and selects the top 20 principal feature components with a cumulative contribution rate of 88%. These components capture key patterns of system operation, including computationally intensive task features and memory access patterns. The dimensionality-reduced data points are compressed from 200 dimensions to 20 dimensions, with each data point representing a summary of the system state within a time window. The principal feature components are saved as a transformation matrix file for real-time data processing. When combining the principal feature components into a task feature matrix, the rows of the matrix correspond to 500 task instances to be analyzed, and the columns correspond to 20 feature components. Each cell is filled with the projection value of a specific task onto the corresponding feature component. During matrix generation, it was found that some task feature vectors had missing values, which were imputed using the feature values ​​of the nearest neighbor tasks. The final task feature matrix has a size of 500×20 and is serialized in binary format before being transmitted to the prediction module. The prediction module collects sample data from the historical database, selects task execution records of similar promotional activities in the past three months, the sample task feature matrix contains feature data of 50,000 task instances, the sample execution time parameter accurately records the actual runtime of each task, the data cleaning process removes test tasks with abnormally short execution times and faulty tasks with abnormally long execution times, the sample set is divided in chronological order, the first 40,000 records are used for training, and the last 10,000 records are used for testing.

[0066] A task execution time predictor was built based on the Support Vector Machine (SVM) algorithm. The algorithm uses a radial basis function kernel, with a regularization parameter C set to 1.0 and the kernel parameter gamma set to the default value of 1 / feature dimension. Mini-batch gradient descent was used for training, with a batch size of 256 records. The training process lasted for 100 iterations. After each iteration, the mean absolute error on the validation set was calculated. Training was terminated early when the error stopped decreasing for 10 consecutive iterations. The trained model showed good predictive consistency on the test set. The model was trained using the sample task feature matrix and sample execution time parameters. During training, the loss function converged after 50 iterations. The final prediction error distribution on the test set exhibited a normal distribution, with most predictions deviating from the actual values ​​by less than 15%. The model was saved as a binary file containing support vectors and decision function parameters. The real-time generated task feature matrix is ​​input into the trained predictor. The predictor loads the model file and initializes the inference environment. Each task feature vector undergoes the same feature scaling process. The prediction process adopts a batch processing mode, processing 100 tasks at a time and outputting the prediction execution time parameters for 500 tasks. The parameter values ​​range from 0.5 seconds to 180 seconds, reflecting the differences in computational complexity among different tasks. A confidence score is added to the prediction results. The prediction values ​​of tasks with high confidence scores (above 0.8) are marked as reliable, while the prediction values ​​of tasks with low confidence scores need to be verified later. All prediction parameters are bound to the task ID and then transmitted to the scheduling generation module.

[0067] In one embodiment, Example 3: The scheduling generation module generates an initial task scheduling scheme based on the predicted execution time parameters and real-time load data. This module receives the predicted execution time parameters from the prediction module and the real-time load data from the data acquisition module, and calculates the priority score for each task using a comprehensive evaluation model. This model considers the urgency of the task and the system load status. The priority score calculation formula is expressed as follows:

[0068]

[0069] in: Indicates priority score, This represents the parameter for predicting execution time. This indicates the length of the task queue in the real-time load data. It is a queue length factor function (as...) (increases and decreases) and These are dynamic weighting coefficients (default values ​​0.6 and 0.4), which are automatically adjusted based on the system's operational phase, increasing during peak business periods. The weighting process involves traversing all tasks to be scheduled to generate a score list, and the scores are normalized to a percentage range for easy comparison.

[0070] When sorting tasks based on priority scores, an improved quicksort algorithm is used. This algorithm optimizes the partitioning strategy based on the characteristics of score distribution. The sorting process maintains a min-heap data structure and updates task priorities in real time. When a new task is added or the score changes, a local reordering is triggered. The sorting result forms a task queue sequential linked list. The linked list nodes contain task ID, score value, and dependency relationship markers. High-priority tasks are located at the head of the linked list and are allocated resources first. The sorting algorithm sets an upper limit on time complexity to avoid scheduling delays. The linked list structure supports fast insertion and deletion operations to adapt to dynamic changes.

[0071] The initial task scheduling scheme is formed by allocating computing nodes based on resource status parameters. The resource allocation adopts a multi-objective optimization strategy, with objectives including load balancing, resource utilization, and task affinity. The module obtains the available CPU cores, remaining memory capacity, and network bandwidth resource status of the computing nodes in real time. The allocation process adopts a two-stage matching mechanism: the first stage filters the set of qualified nodes according to task resource requirements (such as GPU type or memory size), and the second stage uses a greedy algorithm to select the node with the fewest resource fragments. The allocation results are recorded in the node mapping table. The entries in the table include task ID, node ID, allocation timestamp, and expected start time. The output of the initial scheme is a structured document containing task sequences and resource allocation details. The anomaly detection module monitors anomaly indicators during task execution. This module captures task execution logs in real time through a distributed log collection system. Log sources include container runtimes and task execution engines on compute nodes. The log parser extracts key anomaly indicator fields, including error code frequency, timeout event count, and actual execution time deviation. The error rate is calculated as the ratio of the number of failed tasks to the total number of tasks within a period. The latency indicator is obtained by comparing the actual execution time and the predicted execution time. The indicator data is aggregated and stored in a circular buffer in units of time windows (default 5 minutes). The isolated forest algorithm is used to detect anomalies in the anomaly indicators. The algorithm implementation adopts an incremental learning version to adapt to streaming data. The model input is a multi-dimensional anomaly indicator vector (error rate, latency, resource volatility). The maximum depth of each isolated tree is set to automatic adjustment mode. The path length is calculated using a binary tree traversal algorithm. The anomaly score output function is normalized to the [0,1] interval. The threshold setting module dynamically calculates the interquartile range of historical data. When the anomaly score exceeds the upper limit of the threshold, an anomaly event is generated. The event record includes the anomaly type, associated task ID, and initial severity assessment. When an anomaly is detected, a dynamic adjustment mechanism is triggered. This mechanism comprises two components: an event analyzer and a policy executor. The event analyzer analyzes the impact range of the anomaly, identifies the affected task set and resource nodes, and assesses severity using a three-level classification standard (mild, moderate, severe). The assessment criteria include the duration of the anomaly, the number of affected tasks, and the type of resource failure. Severity weighting coefficients are then assigned based on the assessment results. (Value range 0.3-0.9), the weight mapping table is preset in the configuration file and supports dynamic updates.

[0072] The priority correction formula is used when recalculating task priority scores:

[0073]

[0074] in: This is the adjusted priority score. The original priority score. It is a severity weighting, The compensation factor (calculated based on the overall system load status) The value depends on the available resource balance. When resources are scarce, the compensation is reduced. After recalculation, the task queue sequence list is updated. Affected tasks are repositioned according to the new score. During the resource reallocation phase, faulty nodes are marked as isolated. Affected tasks are migrated to standby nodes. The scheduling scheme is optimized, a difference report is generated to record the change details, and the scheme is published to the execution engine in real time through the message bus. The whole process is completed within 200 milliseconds to ensure timely response.

[0075] In a real-time task scheduling scenario of a financial risk analysis system, the system needs to handle a large number of transaction risk control tasks and data analysis tasks. The prediction module has output predicted execution time parameters for 500 tasks, ranging from 2 seconds to 300 seconds. Real-time load data shows that the current task queue contains 1200 tasks to be processed, with an average processing speed of 80 tasks per second. Resource status parameters indicate that the CPU utilization of the 50 computing nodes in the cluster fluctuates between 60% and 90%, and the memory utilization varies between 55% and 75%. The scheduling generation module uses a multi-dimensional evaluation model to calculate the priority score of each task. The model comprehensively considers the urgency of the task and the system resource status. Transaction risk control tasks with high risk levels receive a basic weight increase, while data analysis tasks are weighted according to their business importance. The score calculation process traverses all tasks, generating a priority score list with scores ranging from 0 to 100. High-risk tasks generally have scores higher than 70, while ordinary analysis tasks mostly have scores in the range of 30-60. Tasks are sorted based on priority scores using a stable merge sort algorithm. This algorithm preserves the original order of tasks with the same score, forming a task execution sequence. The sequence consists of urgent risk control tasks with scores above 75, regular analysis tasks with scores between 50 and 75, and batch processing tasks with scores below 50. The sorted results are stored in a doubly linked list, supporting fast insertion and deletion operations. A load-aware strategy is used when allocating computing nodes based on resource status parameters. This strategy prioritizes the real-time load level and resource availability of nodes. The allocation algorithm first selects healthy nodes with CPU utilization below 80% and memory availability above 20%. Then, it matches the most suitable node based on task resource requirements. Risk control tasks are assigned to nodes with dedicated encryption accelerators, and data analysis tasks are assigned to nodes with larger memory capacities, forming an initial task scheduling scheme containing a complete task-node mapping relationship. The anomaly detection module collects task execution data in real time through a log collection agent. This agent is deployed on each computing node to monitor task execution status. The log parser extracts key metrics including task error codes, execution time deviations, and resource usage anomalies. The error rate calculation period is adjusted to a 3-minute window, counting the number of failures and retries for each task. Latency is obtained by comparing the actual completion time with the predicted time; tasks with a deviation exceeding 30% are marked for review. When using the Isolation Forest algorithm to detect anomalies, a sliding window mechanism is employed. The algorithm analyzes the metric data stream over the past 5 minutes. The model input includes three dimensions: error rate, latency rate, and resource volatility. Each isolation tree is limited to a depth of 15 layers. Anomaly scores are calculated using a path length normalization method. When the anomaly score exceeds a dynamic threshold of 0.7, an anomaly event is generated. The event record includes the anomaly time point, associated tasks, and a preliminary classification.When an anomaly is detected, a dynamic adjustment mechanism is triggered. The mechanism first analyzes the scope of the anomaly's impact, identifying 13 affected tasks and 3 computing nodes. The severity assessment considers the anomaly's duration of 8 minutes, affecting multiple critical risk control tasks, classifying it as a severe anomaly. A severity weight of 0.8 is assigned. When recalculating task priority scores, an emergency adjustment formula is used, appropriately reducing the scores of affected tasks based on the anomaly's severity, while maintaining high priority for critical tasks. During task queuing, affected tasks are re-inserted into appropriate positions in the queue, ensuring critical tasks are executed first while preventing the anomaly from spreading. In the resource reallocation phase, the 3 anomaly nodes are marked as isolated, and their tasks are migrated to backup nodes. The migration process ensures data consistency and task continuity. An optimized scheduling plan is generated, reducing the expected completion time by 15% and improving resource utilization by approximately 12%. The entire adjustment process is completed within 150 milliseconds.

[0076] In one embodiment, Example 4: The anomaly detection module uses the Isolation Forest algorithm to detect anomalies in the anomaly indicators. This algorithm is based on the random forest principle and is specifically designed for anomaly detection scenarios. When constructing the Isolation Forest model, historical normal task execution data is used as training data. This data comes from three months of log records collected during the stable operation of the system, containing approximately 2 million normal task execution samples. The training data is strictly screened to exclude any known abnormal records. The data fields include task execution duration, resource utilization, and system load indicators. The model initialization setting is 100 trees, and each tree uses a random subspace method to select features. The sample subset size of each tree is set to 256 records, and the maximum depth limit of the tree is automatically calculated based on the subset size. The model training process uses a recursive partitioning method to construct binary isolation trees. Each tree randomly selects features and split values ​​until the data points are completely isolated. After training, the model saves the tree structure and split point information to persistent storage. Model version management records the training timestamp and data feature distribution. When real-time anomaly indicators are input into the isolated forest model, the data undergoes the same preprocessing process as the training data. The input vector contains the error rate, latency in seconds, and resource fluctuation percentage. The model traverses each isolation tree to calculate the path length of the data points. The path length represents the average number of splits required to isolate the data points. The anomaly score is calculated based on the path length normalization. The score ranges from 0 to 1. The closer the value is to 1, the higher the probability of an anomaly.

[0077] A dynamic adjustment mechanism is used when setting the anomaly threshold. The initial threshold value is set to 0.65 based on the historical distribution of the training data. During system operation, the distribution changes of anomaly scores are continuously monitored. When the score distribution deviates, the threshold is automatically recalculated. The threshold calculation uses statistical methods to consider the mean and standard deviation of the score distribution. When anomaly scores exceed the threshold, they are marked as anomalies. Anomaly records include task ID, anomaly type, score, and timestamp. These records are stored in the anomaly event database for subsequent analysis. The dynamic adjustment mechanism recalculates the task priority score based on the severity of the anomalies. The mechanism first analyzes the type of anomaly, distinguishing between resource anomalies, performance anomalies, and system anomalies. Resource anomalies include memory leaks and CPU overload; performance anomalies involve execution timeouts and response delays; and system anomalies cover node failures and service interruptions. Each anomaly type has a preset impact coefficient ranging from 0.1 to 0.9. The impact range assessment considers the number of tasks and resource nodes affected by the anomaly. The assessment uses a graph traversal algorithm to analyze task dependencies. The severity weight is calculated by combining the anomaly type's impact coefficient and the impact range. Priority scores are recalculated by combining predicted execution time parameters and real-time load data. A weighted adjustment method is used to account for the timeliness of anomaly impacts, assigning higher weights to recently occurring anomalies and recognizing that historical anomalies have diminishing impact over time. The recalculation process iterates through all affected tasks, updating their priority scores and queue positions. Resource allocation is adjusted accordingly to avoid anomaly nodes, migrating tasks to healthy nodes for execution. See Table 1 for data records from the anomaly detection process.

[0078] Table 1: Record of Abnormal Events

[0079] Task ID Exception types Error rate Delay time (seconds) Abnormal scores Severity weighting Processing status T10045 Resource anomaly 0.12 45.6 0.78 0.6 Processed T10087 Performance abnormality 0.08 89.2 0.82 0.7 Processing T10123 System malfunction 0.25 120.5 0.91 0.9 Pending processing T10156 Resource anomaly 0.15 56.8 0.75 0.5 Processed

[0080] The anomaly handling process includes two stages: automatic response and manual confirmation. Upon detecting an anomaly, the system first attempts automatic repair, such as restarting tasks or migrating nodes. If automatic repair fails, it escalates to manual handling. The handling status is updated in real-time on the monitoring interface. Detailed operation logs and effectiveness evaluations are recorded during the handling process. After the anomaly is resolved, an event report is generated, including root cause analysis and preventative measures. These reports are used to improve the anomaly detection model and adjust threshold parameters. The mechanism maintains an anomaly knowledge base storing historical anomaly cases and handling solutions. The knowledge base uses a graph database to store the relationships between anomaly events. When a new anomaly occurs, pattern matching is used to recommend a handling solution. The knowledge base is regularly updated to include new anomaly types and handling experience. During system operation, feedback data is continuously collected to optimize anomaly detection accuracy and adjust mechanism parameters to adapt to environmental changes.

[0081] In one embodiment, Example 5: The user interface module is used to display the initial task scheduling scheme and the optimized scheduling scheme and to receive user feedback. This module adopts a web-based interactive interface design, and the interface layout is divided into three main areas: a scheme visualization area, a parameter adjustment area, and a feedback input area. The scheme visualization area uses the D3.js library to render a task scheduling Gantt chart. Different colored bars in the chart represent different types of tasks. The X-axis represents the timeline, and the Y-axis displays a list of computation nodes. Users can view detailed information about each task, including the predicted execution time and actual resource allocation, by hovering the mouse over it. The parameter adjustment area provides sliders and drop-down menus that allow users to manually adjust the scheduling priority and resource allocation strategy. The feedback input area contains structured forms and free text fields. Users can select predefined feedback types or enter custom suggestions. All user operations are recorded with timestamps and user identity information.

[0082] The interface updates in real time to display changes in the scheduling scheme. When the system generates a new optimized scheduling scheme, the interface receives the updated data via a WebSocket connection. The Gantt chart is dynamically redrawn to show the differences in the schemes, with changed parts marked with highlighted borders. Users can use a time slider to review historical scheduling schemes and compare the effects of schemes at different points in time. The interface supports switching between multiple views, including a resource utilization heatmap, a task dependency graph, and a performance indicator trend graph. The heatmap uses color depth to represent the node load level, the dependency graph shows the dependencies between tasks and the data flow, and the trend graph shows the historical changes of key indicators. User feedback is transmitted to the server through an encrypted channel. Feedback data is stored in a dedicated table in a relational database. The table structure includes fields such as feedback ID, user ID, feedback time, feedback type, associated task ID, detailed description, and processing status. Feedback types are divided into four categories: scheduling strategy adjustment, exception confirmation, resource allocation suggestion, and system parameter optimization. Each type of feedback has a corresponding processing flow and response mechanism. The system performs integrity checks on feedback data, checks required fields, and validates data formats to ensure the availability of feedback information. The feedback learning module updates the task execution time predictor and anomaly detection model based on user feedback. The module periodically scans the feedback database to obtain newly generated feedback records, parses the feedback content to extract key information, and for feedback related to scheduling strategy adjustments, the module analyzes the user's modified priority rules and resource allocation preferences, extracts the pattern features in these decisions, and uses the confirmed correct scheduling decisions as new training samples. After standardization, the sample data is added to the historical task database, which uses a version control mechanism to manage data updates.

[0083] When updating the task execution time predictor, the module first checks the data distribution changes of the new sample data and calculates the statistical difference between the new sample and the original training data. When the difference exceeds a threshold, the model retraining process is triggered. The retraining process adopts an incremental learning approach, using the new samples to fine-tune the model based on the original model, retaining the model's existing knowledge while adapting to new data features. After training, the model performance is evaluated using a validation set. Only after confirming improved accuracy can the model be deployed to the production environment. The model version management records the time, data changes, and performance metrics of each update. The anomaly detection model update focuses on user-confirmed anomaly cases. The module collects user-labeled and processed anomaly events and adds these cases to the anomaly sample library. When retraining the isolated forest model, normal and anomaly samples are used simultaneously to adjust model parameters and improve the detection sensitivity for known anomaly types. Special attention is paid to sample balance during the model update process to avoid too many anomaly samples affecting the recognition of normal patterns. The updated model is only put into use after being validated on a test set. The system establishes a closed-loop feedback processing mechanism, ensuring that every user feedback receives a response. The module automatically generates a feedback processing report, indicating whether the feedback was adopted, the model update status, and the expected results. This report is sent to users via email and in-app messages. The system periodically summarizes feedback analysis results, generating user feedback statistical reports that display common problem types and improvement suggestions. These reports guide the system's long-term optimization direction. The module implements quality monitoring measures to track the feedback learning effect. Monitoring indicators include changes in prediction accuracy after model updates, improvements in anomaly detection accuracy, and user satisfaction metrics. When model performance deteriorates after an update, the system automatically rolls back to the previous version and notifies the system administrator for manual intervention. The entire feedback learning process adheres to strict data security standards to ensure the confidentiality and integrity of user feedback information.

[0084] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. An artificial intelligence task scheduling system based on big data analysis, characterized in that, The system includes: The data acquisition module is used to collect performance parameters and resource status parameters during task execution and obtain real-time load data of the task queue. The data processing module is used to extract features from the performance parameters, resource status parameters, and real-time load data to obtain a task feature matrix; The prediction module is used to predict the task execution time based on the task feature matrix and obtain the predicted execution time parameters. The scheduling generation module is used to generate an initial task scheduling scheme based on the predicted execution time parameters and real-time load data. The anomaly detection module is used to monitor abnormal indicators during task execution and dynamically adjust the initial task scheduling scheme based on the abnormal indicators to obtain an optimized scheduling scheme.

2. The artificial intelligence task scheduling system based on big data analysis according to claim 1, characterized in that, The data acquisition module collects performance parameters and resource status parameters during task execution and obtains real-time load data of the task queue, including: Perform performance monitoring on multiple computing nodes to obtain a set of performance parameters; Collect resource utilization data to obtain a set of resource status parameters; Real-time query of task queue length and processing speed to obtain real-time load data.

3. The artificial intelligence task scheduling system based on big data analysis according to claim 2, characterized in that, The data processing module extracts features from the performance parameters, resource status parameters, and real-time load data to obtain a task feature matrix, including: The performance parameter set, resource status parameter set, and real-time load data are normalized to form standardized data; Principal component analysis algorithm is used to reduce the dimensionality of standardized data and extract the main feature components; The main feature components are combined into a task feature matrix.

4. The artificial intelligence task scheduling system based on big data analysis according to claim 3, characterized in that, The prediction module predicts the task execution time based on the task feature matrix to obtain predicted execution time parameters, including: Based on historical task execution data, collect sample task feature matrices and sample execution time parameters; A task execution time predictor is constructed based on the support vector machine algorithm. The task execution time predictor is trained and tested using the sample task feature matrix and sample execution time parameters; The task feature matrix is ​​input into the task execution time predictor, which predicts the execution time parameters.

5. The artificial intelligence task scheduling system based on big data analysis according to claim 4, characterized in that, The scheduling generation module generates an initial task scheduling scheme based on the predicted execution time parameters and real-time load data, including: By combining the predicted execution time parameters and real-time load data, a priority score is calculated for each task; The tasks are sorted according to their priority scores to obtain the task queue order; Based on resource status parameters, computing nodes are allocated to tasks to form an initial task scheduling scheme.

6. The artificial intelligence task scheduling system based on big data analysis according to claim 5, characterized in that, The anomaly detection module monitors abnormal indicators during task execution and dynamically adjusts the initial task scheduling scheme based on these indicators to obtain an optimized scheduling scheme, including: Collect task execution logs in real time and extract abnormal indicators, including error rate and latency. Use the isolated forest algorithm to detect outliers in anomaly indicators; When an anomaly is detected, a dynamic adjustment mechanism is triggered; Based on the severity of the anomalies, the task priority scores are recalculated, the task queuing order and resource allocation are adjusted, and an optimized scheduling scheme is generated.

7. The artificial intelligence task scheduling system based on big data analysis according to claim 6, characterized in that, The anomaly detection module uses the Isolation Forest algorithm to detect outliers in the anomaly indicators, including: Construct an isolated forest model using historical normal task execution data as training data; Input real-time anomaly indicators into the isolated forest model and output anomaly scores; Set an anomaly threshold; when the anomaly score exceeds the threshold, it is marked as an anomaly.

8. The artificial intelligence task scheduling system based on big data analysis according to claim 7, characterized in that, The dynamic adjustment mechanism recalculates the task priority score based on the severity of the anomalies, including: Assign severity weights based on the type and scope of impact of the anomalies; The priority score is recalculated using a weighted formula by combining the predicted execution time parameter and real-time load data. Task scheduling is adjusted based on the new priority scores.

9. The artificial intelligence task scheduling system based on big data analysis according to claim 8, characterized in that, The system also includes: The user interface module is used to display the initial task scheduling scheme and the optimized scheduling scheme, and to receive user feedback; The feedback learning module updates the task execution time predictor and anomaly detection model based on user feedback.

10. An artificial intelligence task scheduling method based on big data analysis, characterized in that, It includes all modules and method flows of the AI ​​task scheduling system based on big data analysis as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Task scheduling method and system based on resource demand prediction

    CN114579270A

  • Intelligent task scheduling system and method based on machine learning

    CN116909712A

  • Adaptive intelligent task planning and scheduling system based on convolutional neural network

    CN119358931A

  • Task scheduling optimization analysis method and system based on big data

    CN119883580A

  • Server load state evaluation method based on dynamic evaluation algorithm

    CN119902905A

Cited By

  • Adjusting method and system for signal transmission of industrial vibration monitoring sensor

    CN122137804A

  • Artificial intelligence application system scheduling method and system based on multi-source service data fusion

    CN122526753A