A reinforcement learning-driven intelligent scaling and tuning approach for big data pipelines
The LinUCB algorithm dynamically adjusts the number of threads and resource configuration of the Apache NiFi data pipeline, solves the problems of insufficient resource utilization and performance bottlenecks, and achieves efficient intelligent scaling and tuning, which improves the stability and throughput of the system.
Patent Information
- Application Number
- CN202510808471.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The resource configuration and parallelism adjustment of the existing Apache NiFi data pipeline is difficult to adapt to real-time changing traffic requirements, resulting in insufficient resource utilization, performance bottlenecks and inefficient tuning efficiency, and lack of effective dynamic adjustment and rapid response mechanisms.
A dynamic adjustment strategy based on the context slot machine algorithm (LinUCB) is adopted, and the configuration of the number of threads, memory allocation and CPU cores is optimized to achieve intelligent scaling and tuning by monitoring the system status characteristics in real time and combining the online learning mechanism.
It improves the throughput and resource utilization of data pipelines, reduces queue backlog and performance bottlenecks, reduces operation and maintenance costs, and maintains the stability and efficiency of the system in a dynamic load environment.
Smart Images

Figure CN120315903B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data processing and computing technology, and in particular to a reinforcement learning-driven intelligent scaling and tuning method in a big data pipeline. Background Art
[0002] With the rapid development of big data, cloud computing, and artificial intelligence technologies, global data volumes are exploding, including vast amounts of data streams that require real-time processing. To efficiently aggregate and process this massive amount of real-time data, big data pipeline technology is widely used to support real-time business decision-making and data analysis. Apache NiFi, an open-source big data processing framework, is widely used in real-time data pipelines for its powerful data stream management and processing capabilities. In NiFi data pipelines, processor parallelism (number of threads) and system resource allocation (such as CPU and memory) directly impact data processing throughput and system performance.
[0003] However, in real-world applications, resource allocation and parallelism tuning for Apache NiFi data pipelines present significant challenges. Due to the dynamic nature of data traffic and the variability of system load, the number of threads in Apache NiFi's processors, memory allocation within Apache NiFi nodes, and the number of CPU cores are typically configured statically or manually adjusted based on experience. This approach struggles to adapt to changing traffic demands, leading to the following issues:
[0004] (1) Insufficient resource utilization: When the system load is low, excessive processor thread count or resource allocation will cause resource competition, resulting in performance degradation. Idle resources will also lead to waste of CPU and memory resources, increasing operation and maintenance costs;
[0005] (2) Performance bottleneck: In sudden high-load scenarios, insufficient number of processor threads or insufficient system resource allocation can lead to queue backlogs, task delays, and even system performance degradation or service interruptions.
[0006] (3) Low tuning efficiency: Manually adjusting the number of processor threads and system resource configuration relies on the experience of operation and maintenance personnel. The tuning cycle is long and it is difficult to achieve the optimal configuration, and it is difficult to strike a balance between throughput and resource utilization.
[0007] In addition, existing reinforcement learning methods are rarely used in NiFi data pipelines, and it is difficult to effectively handle high-dimensional context features and dynamic load changes. They also lack hierarchical adjustment and rapid response mechanisms, resulting in limited optimization effects.
[0008] Therefore, a method is needed to dynamically adjust the number of processor threads and system resource configuration based on system status to improve the execution efficiency and stability of Apache NiFi data pipelines. By analyzing multi-dimensional metrics during system runtime (such as queue length, processor utilization, system load, etc.), adaptive optimization of parallelism and resource allocation can not only reduce resource waste but also effectively respond to sudden load changes. However, traditional static configuration or rule-based adjustment strategies cannot learn the mapping relationship between system status and optimal configuration in real time, and often lack adaptability when faced with complex and changing operating environments. In addition, frequent adjustments to the number of threads or resource configuration may cause system fluctuations, affecting the stability of data processing.
[0009] To address these challenges, an intelligent approach is needed that can learn and optimize thread counts and system resource allocation strategies in real time in a dynamic environment, while maintaining system stability and efficiency. This paper aims to address the challenge of balancing resource utilization and throughput in Apache NiFi data pipelines by introducing reinforcement learning technology, providing a new solution for intelligent scaling and tuning of big data pipelines. Summary of the Invention
[0010] In response to the shortcomings of the existing technology, the present invention proposes a big data pipeline intelligent scaling and tuning method based on reinforcement learning.
[0011] In this method, the present invention proposes a dynamic adjustment strategy based on the contextual slot machine algorithm (LinUCB). By collecting the system status characteristics of the Apache NiFi data pipeline in real time and combining it with an online learning mechanism, the configuration of the number of threads, memory allocation and the number of CPU cores is gradually optimized. Based on this, the intelligent scaling and tuning of the data pipeline is realized to improve throughput, resource utilization and load balancing capabilities.
[0012] The technical solution of the present invention is:
[0013] A reinforcement learning-driven intelligent scaling and tuning method for big data pipelines, including:
[0014] Step 1: Data pipeline operation and status monitoring;
[0015] Run data processing tasks in the Apache NiFi data pipeline, including: data collection, transformation, and distribution; perform data stream processing through processors; monitor the system operation status in real time; the system refers to the distributed computing environment running Apache NiFi and other components of the data pipeline; collect multi-dimensional status characteristics, including: processor busyness, input queue length, output queue length, system load, memory utilization, and CPU utilization;
[0016] Step 2: Resource indicators and feature extraction;
[0017] Define resource indicators, including dynamic and static resource indicators; extract multidimensional feature vectors based on the collected multidimensional state characteristics, including queue change rate, queue pressure, processor efficiency, and resource interaction characteristics;
[0018] Step 3: Data preprocessing and model initialization;
[0019] Preprocess the collected multi-dimensional state features, including normalization, initialization of the LinUCB model, setting feature dimensions, exploration parameters, and initial arm set;
[0020] Step 4: Run the dynamic adjustment algorithm based on LinUCB;
[0021] The following steps are involved:
[0022] Initialize the action space: define the combination of number of threads, memory size, and number of CPU cores as an arm, and set the optional range;
[0023] Arm selection and UCB calculation: Based on the current context features, the LinUCB algorithm is used to calculate the expected return and confidence bound for each arm and select the optimal configuration. A hierarchical strategy is adopted: the high-level layer periodically adjusts the memory and CPU core count, and the low-level layer adjusts the number of threads in real time.
[0024] Set the reward function: Design a weighted reward function that takes into account queue processing efficiency, processor efficiency, system load, and resource costs. Dynamically adjust the weights for different scenarios and introduce switching penalties to reduce frequent adjustments.
[0025] Model update and optimization: Execute the selected configuration, observe system performance, calculate rewards, and update LinUCB model parameters to achieve online learning;
[0026] Step 5: Periodic reset and rapid response mechanism;
[0027] Periodically reset the LinUCB model's prior parameters to encourage re-exploration of potential optimal configurations. In extreme scenarios, when queue backlogs exceed a threshold or system load exceeds 80%, a rapid response strategy is triggered to limit the range of selectable arms and quickly alleviate system pressure.
[0028] Step 6: Application and evaluation of resource allocation results;
[0029] Apply the optimal configuration selected by the LinUCB algorithm to the Apache NiFi data pipeline, adjusting processor parallelism and node resource allocation in real time. Evaluate the adjusted system performance, including throughput, queue backlog, and resource utilization, to verify the optimization effect.
[0030] Step 7: Save the tuning results and continue to optimize;
[0031] Record the configuration, context features, and reward values of each adjustment and save the LinUCB model state. By loading historical models, it supports rapid recovery after system restart and continuously optimizes system scalability.
[0032] According to the preferred embodiment of the present invention, in step 2, resource index and feature extraction includes:
[0033] During the monitoring period, Apache NiFi's built-in API periodically collects and analyzes system status indicators to obtain dynamic resource indicators for each node, including the number of threads, processor busyness, input queue length, output queue length, memory usage, and CPU usage. The dynamic resource indicators are then saved in real time to local log files.
[0034] Obtain static resource indicators of each node through Apache NiFi's built-in API, including the maximum number of threads, total memory, and number of CPU cores, and save the static resource indicators to the configuration file.
[0035] According to the preferred embodiment of the present invention, in step 3, data preprocessing and model initialization include:
[0036] Step 3-1: Extract dynamic resource indicators and static resource indicators from local log files and configuration files during the monitoring period to form a feature data set;
[0037] Step 3-2: For each status record, i.e., the dynamic resource index and static resource index of each node, calculate the resource utilization; processor busyness , input queue backlog , output queue backlog , System Busyness , thread number normalization , queue change rate , queue pressure , processor efficiency ; is the output queue length; is the input queue length, is the queue capacity;
[0038] Step 3-3: Normalize the CPU usage, processor busyness, and number of threads;
[0039] Step 3-4: Calculate resource interaction characteristics, including: the product of processor busyness and queue length The product of system busyness and processor busyness, and the product of queue pressure and thread number ratio;
[0040] Step 3-5: The feature vector includes processor busyness, input queue length, output queue length, system load, memory usage, CPU usage, queue pressure, processor efficiency, and resource interaction features, forming a multi-dimensional context input.
[0041] According to a preferred embodiment of the present invention, in step 4, initializing the action space includes:
[0042] When the LinUCB algorithm starts, the action space is initialized and the number of threads is , memory size and number of CPU cores The combination of is defined as an arm, and the ranges are 、 、 ; is the maximum and minimum number of threads, is the maximum memory and minimum memory, Is the maximum number of CPU cores and the minimum number of CPU cores;
[0043] Each arm represents a resource configuration, including , i represents the number of threads selected by the current arm, Represents the number of threads, memory size and CPU cores of the arm at initialization, and assigns initial parameters to each arm, including the covariance matrix and the bias vector .
[0044] Further preferably, in step 3-2, the system busyness : ; ; ;
[0045] in, is the CPU busyness, is the memory busyness, 、 The CPU value and memory value collected from the system;
[0046] Queue change rate : .
[0047] Memory usage and queue pressure Calculated by the following formulas:
[0048] (1);
[0049] (2);
[0050] in, is the memory usage, is the total memory, is the input queue length, is the queue capacity, is the hyperbolic tangent function.
[0051] According to a preferred embodiment of the present invention, in step 4, arm selection and UCB calculation include:
[0052] According to the current feature vector , for each arm Calculating expected returns and confidence bound , as shown in formulas (3) and (4):
[0053] (3);
[0054] (4);
[0055] in, is the parameter estimate, for exploration-exploitation balance parameters;
[0056] choose The largest arm , and adopt a hierarchical adjustment strategy: the number of threads Adjust the memory size every half a minute and number of CPU cores Adjust every hour;
[0057] implement After that, observe the system performance and calculate the reward ,renew and , as shown in formulas (5) and (6):
[0058] (5);
[0059] (6);
[0060] If the system state is stable, that is, the reward change is less than the threshold, the current configuration is maintained; otherwise, iterative selection and update are continued.
[0061] Further preferably, in step 4, the reward function Calculation; including:
[0062] Reward Function Comprehensive queue processing efficiency , processor efficiency and system load factor , as shown in formula (7):
[0063] (7);
[0064] in, is the dynamic weight, For switching penalties;
[0065] Queue processing efficiency The calculation is shown in formula (8):
[0066] (8);
[0067] in, Indicates the last input queue backlog, Indicates the current input queue backlog;
[0068] Processor efficiency The calculation is shown in formula (9):
[0069] (9);
[0070] in, It refers to the processor busyness;
[0071] System load factor The calculation is shown in formula (10):
[0072] (10);
[0073] in, For system load.
[0074] According to a preferred embodiment of the present invention, in step 5, the periodic reset and rapid response mechanism includes:
[0075] At fixed intervals, the parameters of the LinUCB model are reset to decay, as shown in formula (11):
[0076] (11);
[0077] If the queue backlog exceeds the threshold and the system load is less than 0.7, the arm selection range is limited to high thread count and high resource configuration. High thread count and high resource configuration means: increasing the number of memory and CPU cores. If the system load exceeds 0.8, low thread count configuration is preferred. Low thread count configuration means: reducing the number of threads and memory allocation.
[0078] Preferably, according to the present invention, in step 6, real-time adjustment of processor parallelism and node resource allocation includes:
[0079] Processor adjustment: Calculate rewards based on throughput, queue backlog, and resource utilization, and call NIFI APIs based on the rewards to adjust parallelism;
[0080] Resource allocation: Call the k8s API based on the reward to adjust the CPU and memory of the NiFi node.
[0081] A computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements the steps of the above-mentioned reinforcement learning-driven intelligent scaling and tuning method in a big data pipeline.
[0082] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the above-mentioned reinforcement learning-driven intelligent scaling and tuning method in a big data pipeline.
[0083] Compared with the prior art, the above technical solution conceived by the present invention can achieve the following beneficial effects:
[0084] 1. Dynamic Resource Optimization: By monitoring and analyzing the operational status of the Apache NiFi data pipeline in real time, the system can dynamically adjust the number of threads, memory allocation, and CPU core configuration based on multi-dimensional contextual characteristics (such as queue length, processor busyness, and memory utilization). This not only improves the matching of resources with load requirements, improves the throughput and execution efficiency of the data pipeline, avoids queue backlogs or performance bottlenecks, but also reduces resource waste and lowers operation and maintenance costs.
[0085] 2. LinUCB-based Intelligent Scaling: We designed dynamic scaling based on the contextual slot machine algorithm (LinUCB), which uses an online learning mechanism to gradually optimize resource allocation. Compared with traditional static configuration or manual tuning, this method effectively addresses the problems of experience-based reliance and insufficient adaptability. Furthermore, through periodic resets and rapid response mechanisms, the algorithm avoids falling into local optimality, ensuring long-term optimization capabilities under dynamic load environments.
[0086] 3. Adaptive Reward Function Design: The LinUCB algorithm integrates queue processing efficiency, processor efficiency, and system load factor as a reward function, dynamically adjusting the weights based on scenarios such as queue backlogs or resource constraints. The key idea is that higher reward values indicate more optimal resource allocation and better system performance. This design provides a reliable basis for intelligent scaling of thread count and resource allocation, ensuring a balance between performance and resource utilization under varying load conditions.
[0087] 4. Broad Applicability: The resource allocation optimization results of this invention are not only applicable to real-time tuning of Apache NiFi data pipelines, but can also be extended to parallelism and resource management in other big data processing frameworks (such as Apache Kafka and Flink). Furthermore, the online learning capabilities of the LinUCB algorithm are not limited to data pipeline optimization but can also be applied to other scenarios requiring dynamic decision-making, such as cloud computing resource scheduling, real-time recommendation systems, and automated operations and maintenance. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] Figure 1 A flowchart of the reinforcement learning-driven intelligent scaling and tuning method in big data pipelines;
[0089] Figure 2 This is a schematic diagram of the resource indicator monitoring and feature extraction architecture;
[0090] Figure 3 Schematic diagram of the LinUCB algorithm dynamic adjustment framework;
[0091] Figure 4 Schematic diagram of the reward function calculation and scene adaptation process;
[0092] Figure 5 This is a schematic diagram of periodic reset and quick response effects;
[0093] Figure 6 is a schematic diagram of the average reward of the arm;
[0094] Figure 7 Schematic diagram of the relationship between the number of threads and performance indicators;
[0095] Figure 8 Schematic diagram of the change of arm rewards over time;
[0096] Figure 9 Schematic diagram of the relationship between queue backlog and processor busyness;
[0097] Figure 10 A diagram showing how system resource usage changes over time. DETAILED DESCRIPTION
[0098] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.
[0099] Explanation of terms:
[0100] 1. Apache NiFi: Apache NiFi is an open source data stream processing framework used to automate and manage the flow of data between systems. It uses processors, queues, and connectors to form data pipelines, supporting real-time data collection, transformation, and distribution.
[0101] 2. Processor: In Apache NiFi, the processor is the core component of the data pipeline, responsible for performing specific data processing tasks (such as data transformation, routing, or storage). Each processor can be configured with multiple threads to process data streams in parallel.
[0102] 3. Thread count: This refers to the number of concurrently executing threads assigned to a single processor, controlling the processor's parallel processing capabilities. The number of threads directly impacts data processing throughput and resource utilization, and typically needs to be adjusted dynamically based on system load.
[0103] 4. Contextual Bandit (LinUCB): LinUCB is an online learning algorithm that combines linear regression and multi-armed bandit learning. By observing the environment state (context), it estimates the reward for each action (arm) and calculates a confidence bound, balancing the needs of exploration and exploitation. In this paper, the number of threads and system resource configuration are considered "arms", and the system state characteristics are the context.
[0104] 5. Queue processing efficiency: refers to the rate of reduction of queue backlog tasks per unit time, used to measure the impact of thread number configuration on data processing speed.
[0105] 6. Processor efficiency: This refers to the degree to which the processor's busyness is close to the ideal working state (for example, 0.85 or 0.9). It is used to evaluate the optimization effect of the number of threads on processor resource utilization.
[0106] 7. System load: refers to the resource usage of the entire system (including CPU, memory, etc.), which is used to determine whether the current configuration causes resource shortage or waste.
[0107] 8. Recovery period mechanism: A time window (e.g., 5 minutes) is set after switching from a high thread count to a low thread count. By reducing the weight of the switching penalty and adjusting the reward calculation, the system operation is stabilized to avoid performance fluctuations caused by frequent switching.
[0108] Example 1
[0109] A reinforcement learning-driven intelligent scaling and tuning method for big data pipelines, including:
[0110] Step 1: Data pipeline operation and status monitoring;
[0111] Run data processing tasks in the Apache NiFi data pipeline, which include data collection, conversion, and distribution. For example, in the experimental environment, the data processing task is to obtain video frames from the surveillance video stream, transmit the data to the image recognition module through NiFi, and then transmit it to the live streaming end through NiFi to complete real-time surveillance video recognition. In the pipeline, NiFi routes and modifies the data attributes, and the remaining processing modules are corresponding.
[0112] Data stream processing is performed through the processor. In the experimental environment, the data stream is a continuous frame image extracted from the monitoring source, which is processed through the data pipeline and finally flows to the user. The processing flow is roughly as follows: frame extraction → NIFI (establishing a path and modifying data attributes) → identification module → NIFI (establishing a path and modifying data attributes) → live streaming module.
[0113] Monitor the system's operating status in real time. The system refers to the distributed computing environment that runs Apache NiFi and other components that make up the data pipeline. It is usually deployed on a Kubernetes cluster and includes NiFi nodes (instances running data pipelines), the underlying operating system, and hardware resources (such as CPU and memory).
[0114] Collect multi-dimensional status characteristics, including processor busyness, input queue length, output queue length, system load, memory usage, and CPU usage; as a basis for subsequent dynamic adjustments.
[0115] Processor busyness is an indicator used to measure the processor load in the Apache NiFi system. It is defined as a weighted average based on the backlog of the input queue and the output queue. It reflects how busy the processor is when processing data flow tasks. The calculation of processor busyness is:
[0116] ;
[0117] Indicates how busy the NiFi processor is in processing data streams, used for thread optimization. High busyness indicates that the processor is saturated with tasks and may be facing a backlog, resulting in slower response times. Low busyness indicates that the processor is idle and has few tasks. Ideally, the busyness is expected to be between 0.7 and 0.9, indicating that the processor is using resources efficiently without being overloaded.
[0118] Step 2: Resource indicators and feature extraction;
[0119] Define resource indicators, including dynamic resource indicators (such as the number of threads, queue length, and processor busyness) and static resource indicators (such as the maximum number of threads, memory capacity, and the total number of CPU cores). Based on the collected multidimensional state features, extract multidimensional feature vectors, including queue change rate, queue pressure, processor efficiency, and resource interaction characteristics (such as the product of the number of threads and queue length). Provide contextual input for the LinUCB algorithm.
[0120] Step 3: Data preprocessing and model initialization;
[0121] The collected multidimensional state features are preprocessed, including normalization (e.g., queue length divided by maximum capacity, number of threads divided by maximum number of threads) to ensure that the feature values are on a uniform scale. The LinUCB model is initialized, setting the feature dimensions, exploration parameter (alpha), and initial arm set (a range of combinations of thread count, memory, and CPU core count). Initialization sets the feature dimensions (12), exploration parameter (default 1.0), and thread count range. Each thread count is assigned a matrix A and a reward vector b. This provides a blank starting point for the LinUCB model, which dynamically selects the optimal thread count based on system status (e.g., queue backlog).
[0122] Step 4: Run the dynamic adjustment algorithm based on LinUCB;
[0123] The following steps are involved:
[0124] Initialize the action space: define the combination of number of threads, memory size, and number of CPU cores as an "arm" and set the optional range (such as minimum number of threads to maximum number of threads, minimum memory to maximum memory);
[0125] Arm selection and UCB calculation: Based on the current contextual features, the LinUCB algorithm is used to calculate the expected return and upper confidence bound (UCB) for each arm and select the optimal configuration. A hierarchical strategy is adopted: high-level periodic adjustments to memory and CPU core count (e.g., hourly) and low-level real-time adjustments to thread count (e.g., minutely) are used.
[0126] Setting the reward function: Design a weighted reward function that takes into account queue processing efficiency, processor efficiency, system load, and resource costs. Dynamically adjust the weights for different scenarios (such as severe queue backlogs or resource constraints), and introduce switching penalties to reduce frequent adjustments.
[0127] Model update and optimization: Execute the selected configuration, observe system performance (such as queue reduction and resource utilization), calculate rewards, and update LinUCB model parameters (covariance matrix and bias vector) to achieve online learning;
[0128] Step 5: Periodic reset and rapid response mechanism;
[0129] Periodically (for example, at fixed intervals) the prior parameters of the LinUCB model are reset to encourage re-exploration of potential optimal configurations. In extreme scenarios, when the queue backlog exceeds a threshold or the system load exceeds 80%, a rapid response strategy is triggered to limit the range of available arms (for example, by simply increasing the number of threads and resource allocation) to quickly alleviate system pressure.
[0130] Step 6: Application and evaluation of resource allocation results;
[0131] Apply the optimal configuration (number of threads, memory size, and number of CPU cores) selected by the LinUCB algorithm to the Apache NiFi data pipeline, adjusting processor parallelism and node resource allocation in real time. Evaluate the adjusted system performance, including throughput, queue backlog, and resource utilization, to verify the optimization effect.
[0132] Step 7: Save the tuning results and continue to optimize;
[0133] Record the configuration, context features, and reward values of each adjustment, and save the LinUCB model status (including parameters and performance history). By loading the historical model, the historical model refers to the collection of LinUCB models and their historical data, including model parameters (feature dimensions, exploration parameters, thread number range) and operation records (performance, selection count, reward, efficiency). This algorithm is an online learning algorithm, and the model is continuously updated, so there will be reading and writing of historical models. This supports rapid recovery after system restart and continuously optimizes system scalability.
[0134] Example 2
[0135] The difference between the reinforcement learning-driven intelligent scaling and tuning method in a big data pipeline described in Example 1 is that:
[0136] In step 2, resource indicators and features are extracted; Figure 2 Shown, including:
[0137] During the monitoring period, Apache NiFi's built-in API periodically collects and analyzes system status indicators to obtain dynamic resource indicators for each node, including the number of threads, processor busyness, input queue length, output queue length, memory usage, and CPU usage. The dynamic resource indicators are then saved in real time to local log files.
[0138] Obtain static resource indicators of each node through Apache NiFi's built-in API, including the maximum number of threads, total memory, and number of CPU cores, and save the static resource indicators to the configuration file.
[0139] In step 3, data preprocessing and model initialization, including:
[0140] Step 3-1: Extract dynamic resource indicators and static resource indicators from the local log files and configuration files during the monitoring period to form a feature dataset; this is used as context input for the LinUCB algorithm.
[0141] Step 3-2: For each status record, i.e., the dynamic resource index and static resource index of each node, calculate the resource utilization; processor busyness , input queue backlog , output queue backlog , System Busyness , thread number normalization , queue change rate , queue pressure , processor efficiency ; is the output queue length; is the input queue length, is the queue capacity;
[0142] Step 3-3: Normalize the CPU usage, processor busyness, and number of threads. The normalization of the number of threads is as follows:
[0143] ;
[0144] in, is the current number of threads, is the maximum number of threads, ensuring that the eigenvalue is in the range [0, 1], Refers to the thread number ratio, which indicates the ratio of the current number of threads to the maximum number of threads. It is used as a feature input into the LinUCB model to optimize the thread number configuration.
[0145] Step 3-4: Calculate resource interaction characteristics, including: the product of processor busyness and queue length ; The product of system busyness and processor busyness, as well as the product of queue pressure and thread number ratio; enhance the expressiveness of context.
[0146] Step 3-5: The feature vector includes processor busyness, input queue length, output queue length, system load, memory usage, CPU usage, queue pressure, processor efficiency, and resource interaction features, forming a multi-dimensional context input.
[0147] In step 4, the action space is initialized, including:
[0148] When the LinUCB algorithm starts, the action space is initialized and the number of threads is , memory size and number of CPU cores The combination of is defined as an arm, and the ranges are 、 、 ; is the maximum and minimum number of threads, is the maximum memory and minimum memory, The maximum and minimum number of CPU cores; used to specify the upper and lower limits that the algorithm can try during the exploration phase;
[0149] Each arm represents a resource configuration, including , i represents the number of threads (parallelism) selected by the current arm, Represents the number of threads, memory size and CPU cores of the arm at initialization, and assigns initial parameters to each arm, including the covariance matrix (identity matrix) and the bias vector .
[0150] In step 3-2, the system busyness : ; ; ;
[0151] in, is the CPU busyness, is the memory busyness, 、 The CPU value and memory value collected from the system;
[0152] Queue change rate : .
[0153] Memory usage and queue pressure Calculated by the following formulas:
[0154] (1);
[0155] (2);
[0156] in, is the memory usage, is the total memory, is the input queue length, is the queue capacity, Refers to the hyperbolic tangent function. In the queue pressure calculation, the normalized backlog ratio (magnified by 5 times) is mapped to a range of 0 to 1 to highlight the impact of high backlog scenarios. This is used as a feature input for the LinUCB model.
[0157] In step 4, arm selection and UCB calculation; Figure 3 Shown, including:
[0158] According to the current feature vector , for each arm Calculating expected returns and confidence bound , is the context vector of the current time step, which contains 12 features and is used by the LinUCB model to calculate the expected reward and select the optimal number of threads. As shown in formulas (3) and (4):
[0159] (3);
[0160] (4);
[0161] in, is the parameter estimate, for exploration-exploitation balance parameters;
[0162] choose The largest arm , and adopt a hierarchical adjustment strategy: the number of threads Adjust the memory size every half a minute and number of CPU cores Adjust every hour;
[0163] implement After that, observe the system performance and calculate the reward ,renew and , as shown in formulas (5) and (6):
[0164] (5);
[0165] (6);
[0166] If the system state is stable, that is, the reward change is less than the threshold, the current configuration is maintained; otherwise, iterative selection and update are continued.
[0167] In step 4, the reward function Calculation; such as Figure 4 Shown, including:
[0168] First, adjust the weights based on the scenario (focus on queues when there is a backlog, and focus on loads when there is a shortage), and then calculate the reward R:
[0169] Reward Function Comprehensive queue processing efficiency , processor efficiency and system load factor , as shown in formula (7):
[0170] (7);
[0171] in, is the dynamic weight, is the switching penalty (0.1 if the number of threads or resource configuration changes, otherwise 0);
[0172] Model Update: Execute Selected Arm , observe the actual reward , update the parameters as follows:
[0173] ;
[0174] Queue processing efficiency The calculation is shown in formula (8):
[0175] (8);
[0176] in, Indicates the last input queue backlog, Indicates the current input queue backlog; used to reflect the processing performance of the processor under the parallelism. A positive value indicates a decrease in the backlog, and a negative value indicates an increase in the backlog.
[0177] Processor efficiency The calculation is shown in formula (9):
[0178] (9);
[0179] in, It refers to the processor busyness;
[0180] System load factor The calculation is shown in formula (10):
[0181] (10);
[0182] in, is the system load (the average of memory and CPU usage). The weight is adjusted according to the scenario, such as queue backlog hour ,otherwise .
[0183] In step five, periodic reset and rapid response mechanism; Figure 5 Shown, including:
[0184] Every fixed time (such as 3600 seconds), the LinUCB model parameters are reset to decay, as shown in formula (11):
[0185] (11);
[0186] If the queue backlog exceeds the threshold (such as ) and the system load is less than 0.7, the arm selection range is limited to high thread count and high resource configuration. High thread count and high resource configuration means: increasing memory and number of CPU cores; if the system load exceeds 0.8, low thread count configuration is preferred. Low thread count configuration means: reducing the number of threads and memory allocation.
[0187] In step 6, processor parallelism and node resource allocation are adjusted in real time, including:
[0188] Processor adjustment: Calculate rewards based on throughput, queue backlog, and resource utilization, and call NIFI APIs based on the rewards to adjust parallelism;
[0189] Resource allocation: Call the k8s API based on the reward to adjust the CPU and memory of the NiFi node.
[0190] The selected optimal configuration Adjust the number of processor threads and resource allocation by sending it to the data pipeline node through the NiFi API. After running for a period of time, evaluate the system throughput (data processed per unit time) and resource utilization (actual resources used / allocated resources) to verify the effectiveness of the adjustments.
[0191] Save each adjusted configuration, feature vector, and reward value to a local file and serialize the LinUCB model parameters After the system restarts, fast recovery and continuous optimization can be achieved by loading the saved models and historical data.
[0192] Figure 6 is a schematic diagram of the average reward of the arm; Figure 7 Schematic diagram of the relationship between the number of threads and performance indicators; Figure 8 Schematic diagram of the cumulative reward of an arm changing over time; Figure 9Schematic diagram of the relationship between queue backlog and processor busyness; Figure 10 A diagram showing how system resource usage changes over time.
[0193] Example 3
[0194] A computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements the steps of the reinforcement learning-driven intelligent scaling and tuning method in a big data pipeline described in Example 1 or 2.
[0195] Example 4
[0196] A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the reinforcement learning-driven intelligent scaling and tuning method in a big data pipeline described in Example 1 or 2 are implemented.
[0197] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A reinforcement learning-driven intelligent scaling and tuning method for big data pipelines, characterized by: include: Step 1: Data pipeline operation and status monitoring; Run data processing tasks in the Apache NiFi data pipeline, including: data collection, transformation, and distribution; perform data stream processing through processors; monitor the system operation status in real time; the system refers to the distributed computing environment running Apache NiFi and other components of the data pipeline; collect multi-dimensional status characteristics, including: processor busyness, input queue length, output queue length, system load, memory utilization, and CPU utilization; Step 2: Resource indicators and feature extraction; Define resource indicators, including dynamic and static resource indicators; extract multidimensional feature vectors based on the collected multidimensional state characteristics, including queue change rate, queue pressure, processor efficiency, and resource interaction characteristics; Step 3: Data preprocessing and model initialization; Preprocess the collected multi-dimensional state features, including normalization, initialization of the LinUCB model, setting feature dimensions, exploration parameters, and initial arm set; Step 4: Run the dynamic adjustment algorithm based on LinUCB; Step 5: Periodic reset and rapid response mechanism; Periodically reset the LinUCB model's prior parameters to encourage re-exploration of potential optimal configurations. In extreme scenarios, when queue backlogs exceed a threshold or system load exceeds 80%, a rapid response strategy is triggered to limit the range of selectable arms and quickly alleviate system pressure. Step 6: Application and evaluation of resource allocation results; Apply the optimal configuration selected by the LinUCB algorithm to the Apache NiFi data pipeline, adjusting processor parallelism and node resource allocation in real time. Evaluate the adjusted system performance, including throughput, queue backlog, and resource utilization, to verify the optimization effect. Step 7: Save the tuning results and continue to optimize; Record the configuration, context features, and reward values of each adjustment and save the LinUCB model state. By loading historical models, it supports rapid recovery after system restart and continuously optimizes system scalability.
2. The method for intelligent scaling and tuning driven by reinforcement learning in a big data pipeline according to claim 1, characterized in that: Run the LinUCB-based dynamic adjustment algorithm; including the following steps: Initialize the action space: define the combination of number of threads, memory size, and number of CPU cores as an arm, and set the optional range; Arm selection and UCB calculation: Based on the current context features, the LinUCB algorithm is used to calculate the expected return and confidence bound for each arm and select the optimal configuration. A hierarchical strategy is adopted: the high-level layer periodically adjusts the memory and CPU core count, and the low-level layer adjusts the number of threads in real time. Set the reward function: Design a weighted reward function that takes into account queue processing efficiency, processor efficiency, system load, and resource costs. Dynamically adjust the weights for different scenarios and introduce switching penalties to reduce frequent adjustments. Model update and optimization: Execute the selected configuration, observe system performance, calculate rewards and update LinUCB model parameters to achieve online learning.
3. The method for intelligent scaling and tuning driven by reinforcement learning in a big data pipeline according to claim 1, characterized in that: In step 2, resource indicators and features are extracted, including: During the monitoring period, Apache NiFi's built-in API periodically collects and analyzes system status indicators to obtain dynamic resource indicators for each node, including the number of threads, processor busyness, input queue length, output queue length, memory usage, and CPU usage. The dynamic resource indicators are then saved in real time to local log files. Obtain static resource indicators of each node through Apache NiFi's built-in API, including the maximum number of threads, total memory, and number of CPU cores, and save the static resource indicators to the configuration file.
4. The method for intelligent scaling and tuning driven by reinforcement learning in a big data pipeline according to claim 1, characterized in that: In step 3, data preprocessing and model initialization, including: Step 3-1: Extract dynamic resource indicators and static resource indicators from local log files and configuration files during the monitoring period to form a feature data set; Step 3-2: For each status record, i.e., the dynamic resource index and static resource index of each node, calculate the resource utilization; processor busyness , input queue backlog , output queue backlog , System Busyness , thread number normalization , queue change rate , queue pressure , processor efficiency ; is the output queue length; is the input queue length, is the queue capacity; Step 3-3: Normalize the CPU usage, processor busyness, and number of threads; Step 3-4: Calculate resource interaction characteristics, including: the product of processor busyness and queue length The product of system busyness and processor busyness, and the product of queue pressure and thread number ratio; Step 3-5: The feature vector includes processor busyness, input queue length, output queue length, system load, memory usage, CPU usage, queue pressure, processor efficiency, and resource interaction features, forming a multi-dimensional context input.
5. The method for intelligent scaling and tuning driven by reinforcement learning in a big data pipeline according to claim 2, characterized in that: In step 4, the action space is initialized, including: When the LinUCB algorithm starts, the action space is initialized and the number of threads is , memory size and number of CPU cores The combination of is defined as an arm, and the ranges are 、 、 ; is the maximum and minimum number of threads, is the maximum memory and minimum memory, Is the maximum number of CPU cores and the minimum number of CPU cores; Each arm represents a resource configuration, including , i represents the number of threads selected by the current arm, Represents the number of threads, memory size and CPU cores of the arm at initialization, and assigns initial parameters to each arm, including the covariance matrix and the bias vector .
6. The method for intelligent scaling and tuning driven by reinforcement learning in a big data pipeline according to claim 4, characterized in that: In step 3-2, the system busyness : ; ; ; in, is the CPU busyness, is the memory busyness, 、 The CPU value and memory value collected from the system; Queue change rate : ; Memory usage and queue pressure Calculated by the following formulas: (1); (2); in, is the memory usage, is the total memory, is the input queue length, is the queue capacity, is the hyperbolic tangent function.
7. The method of intelligent scaling and tuning driven by reinforcement learning in a big data pipeline according to claim 2, characterized in that: In step 4, arm selection and UCB calculation; including: According to the current feature vector , for each arm Calculating expected returns and confidence bound , as shown in formulas (3) and (4): (3); (4); in, is the parameter estimate, for exploration-exploitation balance parameters; choose The largest arm , and adopt a hierarchical adjustment strategy: the number of threads Adjust the memory size every half a minute and number of CPU cores Adjust every hour; implement After that, observe the system performance and calculate the reward ,renew and , as shown in formulas (5) and (6): (5); (6); If the system state is stable, that is, the reward change is less than the threshold, the current configuration is maintained; otherwise, iterative selection and update are continued.
8. The method of intelligent scaling and tuning driven by reinforcement learning in a big data pipeline according to claim 2, characterized in that: In step 4, the reward function Calculation; including: Reward Function Comprehensive queue processing efficiency , processor efficiency and system load factor , as shown in formula (7): (7); in, is the dynamic weight, For switching penalties; Queue processing efficiency The calculation is shown in formula (8): (8); in, Indicates the last input queue backlog, Indicates the current input queue backlog; Processor efficiency The calculation is shown in formula (9): (9); in, It refers to the processor busyness; System load factor The calculation is shown in formula (10): (10); in, For system load.
9. The method of intelligent scaling and tuning driven by reinforcement learning in a big data pipeline according to claim 1, characterized in that: In step 5, the periodic reset and rapid response mechanism includes: At fixed intervals, the parameters of the LinUCB model are reset to decay, as shown in formula (11): (11); If the queue backlog exceeds the threshold and the system load is less than 0.7, the arm selection range is limited to high thread count and high resource configuration. High thread count and high resource configuration means: increasing the number of memory and CPU cores. If the system load exceeds 0.8, low thread count configuration is preferred. Low thread count configuration means: reducing the number of threads and memory allocation.
10. A reinforcement learning driven intelligent scaling and tuning method in a big data pipeline according to any one of claims 1 to 9, characterized in that: In step 6, processor parallelism and node resource allocation are adjusted in real time, including: Processor adjustment: Calculate rewards based on throughput, queue backlog, and resource utilization, and call NIFI APIs based on the rewards to adjust parallelism; Resource allocation: Call the k8s API based on the reward to adjust the CPU and memory of the NiFi node.
Citation Information
Patent Citations
Edge computing resource allocation optimization method and system based on reinforcement learning
CN119862029A
Method for automatically regulating explicit congestion notification of data center network based on multi-agent reinforcement learning
US20240080270A1