Method and System for Dynamically Optimizing Operator Chains Based on a Computational Engine Model
By building a multi-dimensional feature perception network and a hierarchical optimization decision-making mechanism, the problem of unstable performance in traditional computing engines in the face of data flow changes and resource state changes is solved, and efficient dynamic optimization of operator chains and improvement of resource utilization is achieved.
Patent Information
- Application Number
- CN202510241629.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-03
AI Technical Summary
When traditional computing engines face changes in data flow and resource state changes, it is difficult for traditional computing engines to achieve dynamic optimization, resulting in unstable performance and low resource utilization.
By building a multi-dimensional feature perception network, obtaining the operation data of the operator chain, training performance prediction models, and building a hierarchical optimization decision-making mechanism based on reinforcement learning and heuristic algorithms to realize dynamic optimization of the operator chain.
The performance prediction accuracy and dynamic optimization capabilities of the operator chain are improved, the computing efficiency and resource utilization are improved, and the adaptive evolution of optimization strategies is realized.
Smart Images

Figure CN119739534B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to drive engine technology, and in particular to a method and system for dynamically optimizing an operator chain based on a computing engine model. Background Art
[0002] A computing engine is a core component of a modern data processing system, responsible for executing various complex data processing tasks, such as data cleaning, transformation, aggregation, and analysis. With the continuous growth of data scale and the increasing complexity of application scenarios, higher requirements are put forward for the performance and efficiency of the computing engine. Traditional computing engines usually adopt static optimization strategies to optimize the operator chain at the compilation stage or deployment stage, and cannot adapt to changes in the runtime environment.
[0003] Unable to dynamically adapt to changes in data flow: Traditional optimization methods usually optimize based on static data characteristics and resource status, and it is difficult to handle dynamic scenarios such as data traffic fluctuations and load distribution changes. This results in the performance of the computing engine not being able to meet expectations during actual operation, and even performance bottlenecks may occur.
[0004] Lack of collaborative optimization between operators: Most existing optimization methods focus on the optimization of individual operators and lack consideration of the overall performance of the operator chain. The mutual influence and collaborative optimization opportunities between operators are ignored, resulting in limited improvement in overall performance.
[0005] The resource allocation strategy is not flexible enough: Traditional resource allocation strategies usually rely on pre-set rules or empirical values and lack the ability to sense the resource status in real time and adjust dynamically. This may lead to low resource utilization and the inability to fully exert the performance potential of the computing engine. Summary of the Invention
[0006] Embodiments of the present invention provide a method and system for dynamically optimizing an operator chain based on a computing engine model, which can solve the problems in the prior art.
[0007] In the first aspect of the embodiments of the present invention,
[0008] A method for dynamically optimizing an operator chain based on a computing engine model is provided, including:
[0009] Construct a multi-dimensional feature perception network. The multi-dimensional feature perception network obtains operator chain operation data through a distributed acquisition layer, and the distributed acquisition layer is provided with a time-series data acquisition unit, a load feature acquisition unit, and a resource status acquisition unit; the time-series data acquisition unit dynamically tracks the data flow characteristics between operators based on a sliding window mechanism to construct a data traffic fluctuation curve; the load feature acquisition unit uses an adaptive sampling algorithm to obtain operator calculation load characteristics and generate a load distribution heat map; the resource status acquisition unit records the system resource occupancy status based on resource profiling technology to form a multi-dimensional resource utilization rate matrix; input the data traffic fluctuation curve, the load distribution heat map, and the multi-dimensional resource utilization rate matrix into a deep learning model to train and generate an operator chain performance prediction model;
[0010] Construct a hierarchical optimization decision-making mechanism based on the operator chain performance prediction model. In the policy layer, use a reinforcement learning algorithm to train an optimization agent model. The optimization agent model calculates the affinity coefficient between adjacent operators according to the prediction result; when the affinity coefficient is higher than the dynamic threshold, trigger the automatic operator merging operation; in the execution layer, use a heuristic algorithm to optimize the parallelism of the merged operators and dynamically adjust the number of instances based on the computational complexity score; in the resource layer, construct an elastic scaling unit, and the elastic scaling unit automatically generates a fine-grained resource allocation plan according to the change in parallelism; integrate the automatic operator merging operation, the parallelism optimization, and the fine-grained resource allocation plan into an end-to-end optimization strategy;
[0011] Deploy the end-to-end optimization strategy. During the execution process, start an adaptive performance monitoring module. The adaptive performance monitoring module continuously optimizes the monitoring index weights based on an online learning method; calculate the performance gain gradient through the backpropagation algorithm. When the performance gain gradient shows a decreasing trend, activate the policy tuning engine; the policy tuning engine uses a genetic algorithm to dynamically adjust the optimization parameters and simultaneously synchronize the optimization records to the knowledge graph in real time; the knowledge graph continuously improves the optimization rule library through association analysis to provide knowledge support for subsequent optimization decisions and realize the adaptive evolution of the optimization strategy.
[0012] Construct a hierarchical optimization decision-making mechanism based on the operator chain performance prediction model. In the policy layer, use a reinforcement learning algorithm to train an optimization agent model. The optimization agent model calculates the affinity coefficient between adjacent operators according to the prediction result, including:
[0013] Construct a hierarchical optimization decision-making mechanism based on the operator chain performance prediction model. The hierarchical optimization decision-making mechanism includes a policy layer, an execution layer, and a resource layer. The policy layer receives the performance prediction results output by the operator chain performance prediction model, performs principal component analysis and dimensionality reduction processing on the performance prediction results, selects the principal components with feature contribution rates higher than the preset contribution threshold, and maps the dimensionality-reduced features to the interval from zero to one through the Min-Max normalization method to generate a standardized feature matrix.
[0014] In the policy layer, a reinforcement learning algorithm is used to train and optimize the proxy model. The optimization proxy model is based on the double Q-network architecture. The standardized feature matrix is input into the optimization proxy model to construct a state space including operator computing load, data flow characteristics, and resource utilization rate, define an action space for operator merging, splitting, and migration, and design a reward function based on performance gain.
[0015] The optimization proxy model calculates the affinity coefficient of adjacent operators based on the training results. The affinity coefficient is comprehensively calculated through three dimensions: data dependence intensity, resource complementarity degree, and performance improvement expectation.
[0016] In the execution layer, a heuristic algorithm is used to optimize the parallelism of the merged operators and dynamically adjust the number of instances based on the computational complexity score. In the resource layer, an elastic scaling unit is constructed. The elastic scaling unit automatically generates a fine-grained resource allocation plan according to the change in parallelism. The operator automatic merging operation, the parallelism optimization, and the fine-grained resource allocation plan are integrated into an end-to-end optimization strategy, including:
[0017] In the execution layer, a heuristic algorithm is used to optimize the parallelism of the merged operators. A computational complexity scoring model is constructed to analyze the computational density, memory access pattern, and data dependence relationship of the operators to generate a static feature evaluation result, monitor the runtime processor usage rate, memory occupancy, and input / output throughput to generate a dynamic load evaluation result, and calculate the weighted sum of the static feature evaluation result and the dynamic load evaluation result to generate an operator complexity scoring matrix.
[0018] The operator complexity scoring matrix is input into the simulated annealing algorithm to dynamically adjust the search temperature parameter based on the degree of performance improvement, search for the optimal parallelism configuration in the search space, perform local optimization on the optimal parallelism configuration to obtain an optimized parallelism configuration plan, input the optimized parallelism configuration plan into the time series analysis model to generate the short-term load change trend of the system, and set the instance scaling threshold according to the short-term load change trend of the system to realize the dynamic adjustment of the number of instances.
[0019] Build an elastic scaling unit at the resource layer. The elastic scaling unit establishes a mapping relationship model between parallelism and resource requirements according to the optimized parallelism configuration scheme and the instance scaling threshold. The mapping relationship model analyzes the operator complexity scoring matrix to identify the intensive load feature set, which includes compute-intensive load features, memory-intensive load features, and input / output-intensive load features. Input the intensive load feature set into a deep learning model to predict the resource usage trend. The elastic scaling unit automatically generates a fine-grained resource allocation scheme based on the resource usage trend, implements resource isolation for resource quotas, and executes resource policies; Integrate the merged operators, the optimized parallelism configuration scheme, and the fine-grained resource allocation scheme into an end-to-end optimization strategy.
[0020] Deploy the end-to-end optimization strategy. During the execution process, start the adaptive performance monitoring module. The adaptive performance monitoring module continuously optimizes the monitoring metric weights based on the online learning method; Calculate the performance gain gradient through the backpropagation algorithm. When the performance gain gradient shows a decaying trend, activate the policy tuning engine, including:
[0021] Deploy the end-to-end optimization strategy and start the adaptive performance monitoring module. Build a multi-layer perceptron performance prediction model based on the adaptive performance monitoring module. The input layer of the multi-layer perceptron performance prediction model receives the basic monitoring metrics of processor utilization, memory occupancy, and network throughput. The hidden layer of the multi-layer perceptron performance prediction model uses the ReLU activation function for non-linear feature extraction. The output layer of the multi-layer perceptron performance prediction model generates the system performance prediction result;
[0022] The adaptive performance monitoring module continuously collects performance data using the sliding time window mechanism, calculates the Pearson correlation coefficient between the basic monitoring metrics and the system performance prediction result, and uses the Pearson correlation coefficient as the monitoring metric weight to perform weighted aggregation on the basic monitoring metrics to obtain the weighted performance metric; Divide the weighted performance metric into training batches according to the time series, and use the online learning method to update the monitoring metric weight based on the stochastic gradient descent algorithm. The stochastic gradient descent algorithm uses L2 regularization to constrain the weight parameters and uses the exponential moving average method to smooth the parameter update process to obtain the optimized monitoring metric weight;
[0023] Apply the optimized monitoring metric weight to real-time performance monitoring, calculate the prediction deviation between the current system performance and the system performance prediction result, and calculate the performance difference between adjacent time windows for historical performance data segmented by time window to obtain the performance change trend;
[0024] Input the prediction deviation and the performance change trend into a backpropagation algorithm. The backpropagation algorithm calculates the end-to-end delay gradient and the throughput gradient, and synthesizes a performance gain gradient by weighting the end-to-end delay gradient and the throughput gradient. Detect the change trend of the performance gain gradient. When the performance gain gradient decays in three consecutive time windows, activate the policy tuning engine.
[0025] Divide the weighted performance metrics into training batches according to the time series, and update the weights of the monitoring metrics based on the stochastic gradient descent algorithm in an online learning manner. The stochastic gradient descent algorithm uses L2 regularization to constrain the weight parameters and uses the exponential moving average method to smooth the parameter update process. The optimized weights of the monitoring metrics obtained include:
[0026] Use a double-layer sliding window structure to divide the time series of the weighted performance metrics. The double-layer sliding window structure includes an inner time window. Each inner time window collects 300 performance data sampling points to form a training batch.
[0027] Calculate the cosine distance between each sampling point in the training batch and the current system state based on the attention mechanism to obtain a similarity value. Input the similarity value into the softmax function to obtain a sample weight vector. The sample weight vector is used to weight the training batch to obtain weighted training data.
[0028] Input the weighted training data into the stochastic gradient descent algorithm. The stochastic gradient descent algorithm calculates the gradient vector of the monitoring metric weights. When the norm of the gradient vector exceeds a preset vector threshold, scale the gradient vector proportionally while keeping the gradient direction unchanged to obtain an adjusted gradient vector.
[0029] Introduce an L2 regularization term in the loss function to constrain the update of the monitoring metric weights. The regularization coefficient of the L2 regularization term decays linearly with the number of training rounds. Smooth the adjusted gradient vector using the exponential moving average method. Weight-average the adjusted gradient vector and the historical gradient vector with a decay coefficient of 0.9 to obtain a smoothed gradient vector.
[0030] Update the weights of the monitoring metrics based on the smoothed gradient vector to obtain updated weights of the monitoring metrics. Calculate the predicted mean squared error in multiple consecutive inner time windows to obtain the prediction accuracy. Calculate the standard deviation of the change of the updated weights of the monitoring metrics to obtain the weight stability index.
[0031] Monitor the loss change trend of consecutive training batches. When the loss change trend continuously decreases, increase the learning rate to accelerate convergence. When the loss change trend fluctuates, decrease the learning rate to improve stability. Use the updated monitoring metric weights and the adjusted learning rate for the next round of training iteration until the prediction accuracy and the weight stability metric simultaneously meet the preset conditions, and then output the optimized monitoring metric weights.
[0032] The policy tuning engine dynamically adjusts the optimization parameters using a genetic algorithm and simultaneously synchronizes the optimization records to the knowledge graph in real time. The knowledge graph continuously improves the optimization rule library through association analysis, provides knowledge support for subsequent optimization decisions, and realizes the adaptive evolution of the optimization strategy, including:
[0033] Construct a multi-layer knowledge graph, which includes an entity layer, a relationship layer, and an attribute layer. The policy tuning engine constructs chromosome encoding using a genetic algorithm. Extract the historical optimal configuration from the multi-layer knowledge graph, generate an initial population by perturbing the parameters of the historical optimal configuration. Construct a fitness function with system latency, throughput, and resource utilization, calculate the fitness distribution of the initial population, and obtain the population convergence degree and diversity index.
[0034] Dynamically adjust the selection pressure coefficient based on the population convergence degree, and use the roulette wheel strategy to select high-quality individuals. Adaptively calculate the crossover probability according to the diversity index, and perform arithmetic crossover on the selected high-quality individuals. Adaptively adjust the mutation probability based on the convergence speed of consecutive generations, and perform Gaussian mutation on the crossed individuals.
[0035] Form a new population with the individual with the best fitness in each generation and the mutated individuals. Use the parameter configuration, performance data, and fitness evaluation results of the new population as optimization records, and synchronize them to the entity layer, the relationship layer, and the attribute layer in the multi-layer knowledge graph in real time.
[0036] Perform association analysis on the optimization records in the multi-layer knowledge graph, use the priori algorithm to mine the parameter configuration rules, use time series analysis to identify the performance change rules, and use causal reasoning to construct the parameter adjustment path. Use the parameter configuration rules, the performance change rules, and the parameter adjustment path as optimization strategies and store them in the rule library.
[0037] Calculate the confidence of the optimization strategies in the rule library, mark the strategies with confidence lower than the preset confidence threshold as strategies to be verified. Perform A / B testing in the verification environment to verify the effectiveness of the strategies to be verified. Update the verified optimization strategies to the rule library to realize the dynamic optimization of the rule library.
[0038] Use a graph neural network to calculate the similarity between the current state and historical scenarios, and extract the optimization strategies corresponding to the similar scenarios from the rule base; perform weighted combination based on the policy confidence to obtain a parameter adjustment plan; use the parameter adjustment plan as prior knowledge to guide the population evolution direction of the policy tuning engine;
[0039] Monitor the performance metrics during the optimization process, and trigger parameter rollback when a performance decline is detected; record the performance fluctuation scenarios into the multi-layer knowledge graph, and continuously improve the rule base through continuous two-way feedback to achieve the adaptive evolution of the optimization strategy.
[0040] Dynamically adjust the selection pressure coefficient based on the population convergence degree, and use the roulette wheel strategy to select high-quality individuals; adaptively calculate the crossover probability according to the diversity index, and perform arithmetic crossover on the selected high-quality individuals; adaptively adjust the mutation probability based on the convergence speed of consecutive generations, and perform Gaussian mutation on the crossed individuals, including:
[0041] Dynamically adjust the selection pressure coefficient based on the population convergence degree. When the population convergence degree is lower than the preset interval threshold, increase the selection pressure coefficient according to the first linear function. When the population convergence degree is higher than the preset interval threshold, decrease the selection pressure coefficient according to the second linear function; calculate the individual selection probability by dividing the selection pressure coefficient power of the individual fitness value by the sum of the population fitness, and use the roulette wheel strategy to select high-quality individuals according to the individual selection probability;
[0042] Take the product of the compensation value of the diversity index and the adjustment coefficient as the dynamic adjustment amount, and add the dynamic adjustment amount to the basic crossover probability to obtain the adaptive crossover probability;
[0043] Calculate the crossover weight coefficient based on the fitness ratio of the high-quality individuals, and perform arithmetic crossover operations on the high-quality individuals according to the crossover weight coefficient and the adaptive crossover probability to generate offspring individuals;
[0044] Calculate the fitness change rate of the optimal individuals in adjacent generations to obtain the population convergence speed; when the population convergence speed is less than the preset convergence threshold, add the basic mutation probability and the mutation compensation amount to obtain the adaptive mutation probability. When the population convergence speed is greater than the preset convergence threshold, multiply the basic mutation probability by the attenuation function to obtain the adaptive mutation probability;
[0045] Calculate the Gaussian mutation step size based on the population convergence speed, take the product of the Gaussian mutation step size and the normal distribution random number as the mutation amount, and apply the mutation amount to the offspring individuals according to the adaptive mutation probability for Gaussian mutation.
[0046] In the second aspect of the embodiments of the present invention,
[0047] Provided is a system for dynamically optimizing an operator chain based on a computational engine model, including:
[0048] A first unit for constructing a multi-dimensional feature perception network. The multi-dimensional feature perception network obtains operator chain operation data through a distributed acquisition layer, and the distributed acquisition layer is provided with a time-series data acquisition unit, a load feature acquisition unit, and a resource status acquisition unit. The time-series data acquisition unit dynamically tracks the data flow characteristics between operators based on a sliding window mechanism and constructs a data flow fluctuation curve. The load feature acquisition unit uses an adaptive sampling algorithm to obtain operator calculation load characteristics and generates a load distribution heat map. The resource status acquisition unit records the system resource occupancy status based on resource profiling technology and forms a multi-dimensional resource utilization rate matrix. The data flow fluctuation curve, the load distribution heat map, and the multi-dimensional resource utilization rate matrix are input into a deep learning model to train and generate an operator chain performance prediction model.
[0049] A second unit for constructing a hierarchical optimization decision-making mechanism based on the operator chain performance prediction model. In the policy layer, a reinforcement learning algorithm is used to train an optimization agent model, and the optimization agent model calculates the affinity coefficient of adjacent operators according to the prediction results. When the affinity coefficient is higher than the dynamic threshold, an automatic operator merging operation is triggered. In the execution layer, a heuristic algorithm is used to optimize the parallelism of the merged operators, and the number of instances is dynamically adjusted based on the computational complexity score. In the resource layer, an elastic scaling unit is constructed, and the elastic scaling unit automatically generates a fine-grained resource allocation plan according to the change in parallelism. The automatic operator merging operation, the parallelism optimization, and the fine-grained resource allocation plan are integrated into an end-to-end optimization strategy.
[0050] A third unit for deploying the end-to-end optimization strategy. During the execution process, an adaptive performance monitoring module is started, and the adaptive performance monitoring module continuously optimizes the monitoring index weights based on an online learning method. The performance gain gradient is calculated through a backpropagation algorithm. When the performance gain gradient shows a decaying trend, a policy tuning engine is activated. The policy tuning engine uses a genetic algorithm to dynamically adjust the optimization parameters and simultaneously synchronizes the optimization records to the knowledge graph in real time. The knowledge graph continuously improves the optimization rule library through association analysis and provides knowledge support for subsequent optimization decisions to achieve the adaptive evolution of the optimization strategy.
[0051] In the third aspect of the embodiments of the present invention,
[0052] Provided is an electronic device, including:
[0053] A processor;
[0054] A memory for storing processor-executable instructions;
[0055] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0056] In the fourth aspect of the embodiments of the present invention,
[0057] a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0058] The beneficial effects of this application are as follows:
[0059] 1. Improve the accuracy of operator chain performance prediction: By constructing a multi-dimensional feature perception network, the present invention comprehensively considers multi-dimensional features such as data traffic, computing load, and resource status, and uses a deep learning model for training, which can more accurately predict the performance of the operator chain and provide a reliable basis for subsequent optimization decisions.
[0060] 2. Achieve end-to-end dynamic optimization of the operator chain: The present invention constructs a hierarchical optimization decision-making mechanism covering the policy layer, execution layer, and resource layer, and realizes end-to-end dynamic optimization of the operator chain through operations such as operator automatic merging, parallelism optimization, and resource elastic scaling, effectively improving the computing efficiency and resource utilization rate.
[0061] 3. Ensure the adaptive evolution of optimization strategies: The present invention deploys an adaptive performance monitoring module and combines technologies such as online learning, policy tuning engines, and knowledge graphs to achieve the adaptive evolution of optimization strategies, which can dynamically adjust optimization parameters and rules according to actual operating conditions, continuously improve the optimization effect, and have the ability of long-term learning and improvement. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 is a schematic flowchart of the method for dynamically optimizing an operator chain based on a computing engine model according to an embodiment of the present invention;
[0063] Figure 2 is a schematic structural diagram of a system for dynamically optimizing an operator chain based on a computing engine model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0065] The technical solution of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0066] Figure 1 FIG. is a schematic flowchart of a method for dynamically optimizing an operator chain based on a computing engine model according to an embodiment of the present invention. As Figure 1 shown, the method includes:
[0067] S11. Construct a multi-dimensional feature perception network. The multi-dimensional feature perception network obtains operator chain operation data through a distributed acquisition layer. The distributed acquisition layer is provided with a time-series data acquisition unit, a load feature acquisition unit, and a resource status acquisition unit. The time-series data acquisition unit dynamically tracks the data flow characteristics between operators based on a sliding window mechanism to construct a data flow fluctuation curve. The load feature acquisition unit uses an adaptive sampling algorithm to obtain operator calculation load characteristics and generate a load distribution heat map. The resource status acquisition unit records the system resource occupancy status based on resource profiling technology to form a multi-dimensional resource utilization rate matrix. The data flow fluctuation curve, the load distribution heat map, and the multi-dimensional resource utilization rate matrix are input into a deep learning model to train and generate an operator chain performance prediction model.
[0068] S12. Construct a hierarchical optimization decision-making mechanism based on the operator chain performance prediction model. In the policy layer, a reinforcement learning algorithm is used to train and optimize an agent model. The optimization agent model calculates the affinity coefficient between adjacent operators according to the prediction result. When the affinity coefficient is higher than the dynamic threshold, an automatic operator merging operation is triggered. In the execution layer, a heuristic algorithm is used to optimize the parallelism of the merged operators, and the number of instances is dynamically adjusted based on the computational complexity score. In the resource layer, an elastic scaling unit is constructed. The elastic scaling unit automatically generates a fine-grained resource allocation plan according to the change in parallelism. The automatic operator merging operation, the parallelism optimization, and the fine-grained resource allocation plan are integrated into an end-to-end optimization strategy.
[0069] S13. Deploy the end-to-end optimization strategy. During the execution process, an adaptive performance monitoring module is started. The adaptive performance monitoring module continuously optimizes the monitoring index weights based on an online learning method. The performance gain gradient is calculated through a backpropagation algorithm. When the performance gain gradient shows a decaying trend, a strategy tuning engine is activated. The strategy tuning engine uses a genetic algorithm to dynamically adjust the optimization parameters, and at the same time synchronizes the optimization records to the knowledge graph in real time. The knowledge graph continuously improves the optimization rule library through association analysis to provide knowledge support for subsequent optimization decisions, realizing the adaptive evolution of the optimization strategy.
[0070] In an alternative embodiment, a hierarchical optimization decision-making mechanism is constructed based on the operator chain performance prediction model. In the policy layer, a reinforcement learning algorithm is used to train an optimization agent model. The optimization agent model calculates the affinity coefficient between adjacent operators according to the prediction result, including:
[0071] A hierarchical optimization decision-making mechanism is constructed based on the operator chain performance prediction model. The hierarchical optimization decision-making mechanism includes a policy layer, an execution layer, and a resource layer. The policy layer receives the performance prediction result output by the operator chain performance prediction model, performs principal component analysis and dimensionality reduction processing on the performance prediction result, selects the principal components with the feature contribution rate higher than the preset contribution threshold, and maps the dimensionality-reduced features to the interval from zero to one through the Min-Max normalization method to generate a standardized feature matrix.
[0072] In the policy layer, a reinforcement learning algorithm is used to train an optimization agent model. The optimization agent model is based on a double Q-network architecture. The standardized feature matrix is input into the optimization agent model to construct a state space including operator computing load, data flow characteristics, and resource utilization rate, define an action space for operator merging, splitting, and migration, and design a reward function based on performance gain.
[0073] The optimization agent model calculates the affinity coefficient between adjacent operators based on the training result. The affinity coefficient is comprehensively calculated through three dimensions: data dependence intensity, resource complementarity degree, and performance improvement expectation.
[0074] First, an operator chain performance prediction model is constructed. This model can predict performance indicators such as the execution time and resource consumption of the operator chain based on information such as historical performance data and operator characteristics. For example, collect data such as the execution time, input data volume, and output data volume of a certain operator chain in the past week, and use machine learning algorithms (such as regression models, neural networks, etc.) to train a performance prediction model.
[0075] Next, a hierarchical optimization decision-making mechanism is constructed. This mechanism includes a policy layer, an execution layer, and a resource layer. The policy layer is responsible for formulating optimization strategies according to the performance prediction results; the execution layer is responsible for executing the optimization strategies; the resource layer is responsible for managing and allocating computing resources.
[0076] Then, in the policy layer, receive the performance prediction results output by the performance prediction model. For example, the prediction model outputs that the estimated execution time of a certain operator chain is 10 minutes and the resource consumption is 100 CPU cores. Perform principal component analysis dimensionality reduction processing on the performance prediction results. For example, reduce multiple performance metrics (execution time, resource consumption, etc.) to a few principal components, retaining the main feature information. Select the principal components with feature contribution rates higher than the preset contribution threshold. For example, set the contribution threshold to 80% and select the principal components with contribution rates exceeding 80%. Map the dimensionality-reduced features to the interval from zero to one through the maximum-minimum normalization method to generate a standardized feature matrix. For example, scale the values of the principal components to between 0 and 1 to form a matrix.
[0077] In the policy layer, use the reinforcement learning algorithm to train and optimize the proxy model. This model is based on the double Q-network architecture. Input the standardized feature matrix into the optimized proxy model. Construct a state space that includes operator computational load, data flow characteristics, and resource utilization. For example, the state space can include information such as the CPU usage rate, memory usage rate, input data volume, and output data volume of each operator. Define the action space for operator merging, splitting, and migration. For example, two adjacent operators can be merged into one operator, one operator can be split into multiple operators, or one operator can be migrated to another machine for execution. Design a reward function based on performance gain. For example, if an action can reduce the execution time of the operator chain, a positive reward is given; otherwise, a negative reward is given.
[0078] The optimized proxy model calculates the affinity coefficient between adjacent operators based on the training results. The affinity coefficient is comprehensively calculated through three dimensions: data dependence intensity, resource complementarity degree, and performance improvement expectation. For example, if there is a strong data dependence relationship between two operators, and their resource requirements are complementary (for example, one operator requires a large amount of CPU resources and the other operator requires a large amount of memory resources), and merging these two operators can significantly improve performance, then their affinity coefficient is relatively high.
[0079] Finally, according to the affinity coefficient and the decision result of the optimized proxy model, the execution layer performs corresponding optimization operations, such as merging operators with high affinity, splitting operators with low affinity, and migrating operators to appropriate resource nodes. The resource layer performs resource allocation and scheduling according to the instructions of the execution layer.
[0080] The solution of this application can:
[0081] Improve the performance of the operator chain: By predicting performance bottlenecks and intelligently optimizing the operator execution strategy, the execution time and resource consumption of the operator chain can be effectively reduced, and the overall performance can be improved. For example, by merging operators with high affinity, the data transfer time and computational overhead can be reduced, thus improving performance. Enhance resource utilization: Through intelligent resource scheduling and allocation, computing resources can be fully utilized, resource waste can be avoided, and resource utilization can be improved. For example, by migrating operators to appropriate resource nodes, the resource load can be balanced, and the situation where some nodes are overloaded while others are idle can be avoided. Automated optimization decision-making: By training an optimization agent model with reinforcement learning algorithms, automated optimization decisions can be achieved without manual intervention, reducing the operation and maintenance costs. For example, the optimization agent model can automatically select the optimal optimization strategy based on the performance prediction results without manual parameter adjustment.
[0082] In an alternative embodiment, a heuristic algorithm is used in the execution layer to optimize the parallelism of the merged operators, and the number of instances is dynamically adjusted based on the computational complexity score; an elastic scaling unit is constructed in the resource layer, and the elastic scaling unit automatically generates a fine-grained resource allocation scheme according to the change in parallelism; integrating the automatic operator merging operation, the parallelism optimization, and the fine-grained resource allocation scheme into an end-to-end optimization strategy includes:
[0083] In the execution layer, a heuristic algorithm is used to optimize the parallelism of the merged operators. A computational complexity scoring model is constructed to analyze the computational density, memory access pattern, and data dependency relationship of the operators to generate a static feature evaluation result. The runtime processor utilization rate, memory occupancy, and input / output throughput are monitored to generate a dynamic load evaluation result. The static feature evaluation result and the dynamic load evaluation result are weighted and calculated to generate an operator complexity scoring matrix;
[0084] The operator complexity scoring matrix is input into the simulated annealing algorithm, and the search temperature parameter is dynamically adjusted based on the degree of performance improvement. The optimal parallelism configuration is searched in the search space, and local optimization is performed on the optimal parallelism configuration to obtain an optimized parallelism configuration scheme; the optimized parallelism configuration scheme is input into the time series analysis model to generate the short-term load change trend of the system, and the instance scaling threshold is set according to the short-term load change trend of the system to achieve dynamic adjustment of the number of instances;
[0085] Build an elastic scaling unit at the resource layer. The elastic scaling unit establishes a mapping relationship model between parallelism and resource requirements according to the optimized parallelism configuration scheme and the instance scaling threshold. The mapping relationship model analyzes the operator complexity scoring matrix to identify intensive load feature sets, which include compute-intensive load features, memory-intensive load features, and input / output-intensive load features. Input the intensive load feature sets into a deep learning model to predict the resource usage trend. The elastic scaling unit automatically generates a fine-grained resource allocation scheme based on the resource usage trend, implements resource isolation for resource quotas, and executes resource policies; integrate the merged operators, the optimized parallelism configuration scheme, and the fine-grained resource allocation scheme into an end-to-end optimization strategy.
[0086] First, perform operator merging. Merge the operators that can be merged in the data flow graph. For example, merge consecutive matrix multiplication operations into a larger matrix multiplication operation to reduce data transfer and scheduling overhead. For instance, merge the originally three independent matrix multiplication operators A*B, B*C, and C*D into an operator (A*B*C)*D.
[0087] Then, perform parallelism optimization. Optimize the parallelism of the merged operators to fully utilize computing resources. This includes building a computational complexity scoring model. This model comprehensively considers the static characteristics of the operators and the results of dynamic load evaluation. Static characteristics include computational density (e.g., the ratio of floating-point operation count to data read count), memory access pattern (e.g., sequential access or random access), and data dependencies. The results of dynamic load evaluation include runtime processor utilization, memory occupancy, and input / output throughput. For example, an operator that requires a large number of floating-point operations but has a small data read volume has a high computational density. By performing weighted calculations on the static characteristics and the results of dynamic load evaluation, an operator complexity scoring matrix is generated. Suppose there are two operators. The static characteristic evaluation result of operator 1 is 0.8, and the dynamic load evaluation result is 0.7, with weights of 0.6 and 0.4 respectively; the static characteristic evaluation result of operator 2 is 0.6, and the dynamic load evaluation result is 0.9, with the same weights. Then the complexity score of operator 1 is 0.8 * 0.6 + 0.7 * 0.4 = 0.76, and the complexity score of operator 2 is 0.6 * 0.6 + 0.9 * 0.4 = 0.72.
[0088] Next, use the simulated annealing algorithm. Input the operator complexity scoring matrix into the simulated annealing algorithm to find the optimal parallelism configuration. The simulated annealing algorithm dynamically adjusts the search temperature parameter according to the degree of performance improvement and searches for the optimal solution in the search space. For example, the initial temperature is set to 100, and the temperature is decreased in each iteration until it is lower than the preset threshold. At each temperature, the algorithm tries different parallelism configurations and evaluates their performance. If the new configuration has better performance, it is accepted; if the performance is worse, it is accepted with a certain probability to avoid being trapped in a local optimal solution. Perform local optimization on the found optimal parallelism configuration, such as fine-tuning the parallelism, to further improve the performance. Assume that the optimal parallelism configuration is 4, and after local optimization, the final parallelism configuration scheme is 3.
[0089] After that, perform dynamic adjustment of the number of instances. Input the optimized parallelism configuration scheme into the time series analysis model to predict the short-term load change trend of the system. Set the instance scaling threshold according to the prediction result to achieve dynamic adjustment of the number of instances. For example, if it is predicted that the load will increase in the next hour, the number of instances is increased in advance to avoid performance degradation caused by excessive load. Assume that it is predicted that the load will increase by 20% in the next 5 minutes, exceeding the preset threshold of 15%, then trigger the instance expansion operation and increase the number of instances from 2 to 3.
[0090] Subsequently, construct an elastic scaling unit. Construct an elastic scaling unit at the resource layer. This unit establishes a mapping relationship model between parallelism and resource requirements according to the optimized parallelism configuration scheme and the instance scaling threshold. This model analyzes the operator complexity scoring matrix and identifies the set of intensive load characteristics, including compute-intensive, memory-intensive, and input / output-intensive load characteristics. For example, if the compute density of an operator is very high, it is considered a compute-intensive load. Input these sets of intensive load characteristics into the deep learning model to predict the resource usage trend. The elastic scaling unit automatically generates a fine-grained resource allocation scheme based on the prediction result, implements resource isolation for resource quotas, and executes resource policies. For example, if it is predicted that a certain operator requires a large amount of memory resources, more memory quotas are allocated to it. Assume that it is predicted that operator 1 requires 2GB of memory and operator 2 requires 1GB of memory, then the elastic scaling unit will allocate 2GB of memory to operator 1 and 1GB of memory to operator 2. To improve resource utilization, resource policies can be implemented for resource quotas, such as allowing the total memory allocation to exceed the actual available memory.
[0091] Finally, integrate the end-to-end optimization strategy. Integrate the merged operators, the optimized parallelism configuration scheme, and the fine-grained resource allocation scheme into an end-to-end optimization strategy to achieve global optimal performance.
[0092] The solution of this application can:
[0093] Improve resource utilization: Through fine-grained resource allocation and resource policies, resource utilization can be maximized to avoid resource waste. Reduce operating costs: By dynamically adjusting the number of instances and optimizing resource allocation, operating costs can be reduced, such as reducing the costs of cloud computing platforms. Improve system performance: Through parallelism optimization and elastic scaling, system throughput can be increased and latency can be reduced, thereby improving the overall system performance.
[0094] In an alternative embodiment, the end-to-end optimization policy is deployed, and the adaptive performance monitoring module is started during execution. The adaptive performance monitoring module continuously optimizes the monitoring metric weights based on an online learning method; the performance gain gradient is calculated through the backpropagation algorithm. When the performance gain gradient shows a decaying trend, activating the policy tuning engine includes:
[0095] Deploy the end-to-end optimization policy, start the adaptive performance monitoring module, and construct a multi-layer perceptron performance prediction model based on the adaptive performance monitoring module. The input layer of the multi-layer perceptron performance prediction model receives basic monitoring metrics such as processor utilization, memory occupancy, and network throughput. The hidden layer of the multi-layer perceptron performance prediction model uses the ReLU activation function for non-linear feature extraction. The output layer of the multi-layer perceptron performance prediction model generates a system performance prediction result;
[0096] The adaptive performance monitoring module continuously collects performance data using a sliding time window mechanism, calculates the Pearson correlation coefficient between the basic monitoring metrics and the system performance prediction result, and uses the Pearson correlation coefficient as the monitoring metric weight to perform weighted aggregation on the basic monitoring metrics to obtain weighted performance metrics; the weighted performance metrics are divided into training batches according to time series, and the monitoring metric weights are updated based on the online learning method using the stochastic gradient descent algorithm. The stochastic gradient descent algorithm uses L2 regularization to constrain the weight parameters and uses the exponential moving average method to smooth the parameter update process to obtain optimized monitoring metric weights;
[0097] Apply the optimized monitoring metric weights to real-time performance monitoring, calculate the prediction deviation between the current system performance and the system performance prediction result, and calculate the performance difference between adjacent time windows for historical performance data segmented by time window to obtain the performance change trend;
[0098] Input the prediction deviation and the performance change trend into the backpropagation algorithm. The backpropagation algorithm calculates the end-to-end delay gradient and the throughput gradient, and synthesizes the performance gain gradient based on the weighted sum of the end-to-end delay gradient and the throughput gradient; detect the change trend of the performance gain gradient. When the performance gain gradient decays for three consecutive time windows, activate the policy tuning engine.
[0099] First, deploy an end-to-end optimization strategy. The purpose of this strategy is to optimize the overall system performance, such as reducing end-to-end latency or increasing throughput. The specific content of the strategy depends on the actual system. For example, it can be optimizations in aspects such as resource allocation, task scheduling, and data transmission. Assume the strategy we deploy is to dynamically adjust the resource allocation ratios of each server in the server cluster to balance the load and maximize the overall throughput.
[0100] Next, start the adaptive performance monitoring module. The core function of this module is to continuously optimize the weights of monitoring metrics based on an online learning approach to more accurately reflect the system performance. This module includes sub-modules such as performance data collection, performance prediction model construction, monitoring metric weight calculation, and weight update.
[0101] In the performance data collection stage, the module uses a sliding time window mechanism to continuously collect the basic monitoring metrics of the system, such as processor utilization, memory occupancy, and network throughput. Assume the size of the sliding time window is set to 60 seconds, and data is collected once per second.
[0102] Then, construct a multi-layer perceptron performance prediction model. The input layer of this model receives the collected basic monitoring metrics, such as processor utilization, memory occupancy, and network throughput. The hidden layer uses the ReLU activation function for non-linear feature extraction. The output layer generates the system performance prediction results, such as predicting the average end-to-end latency in the next minute. Assume this model contains two hidden layers, and each hidden layer contains 10 neurons. The training data of the model comes from historical performance data, such as the performance data of the past week.
[0103] After that, calculate the monitoring metric weights. The module calculates the Pearson correlation coefficient between the basic monitoring metrics and the system performance prediction results. The Pearson correlation coefficient is used to measure the linear correlation between two variables. The calculated Pearson correlation coefficient is used as the monitoring metric weight to perform weighted aggregation on the basic monitoring metrics to obtain the weighted performance metric. For example, assume the Pearson correlation coefficients between processor utilization, memory occupancy, and network throughput and the system performance prediction results are 0.8, 0.5, and 0.7 respectively. Then the weighted performance metric is 0.8 * processor utilization + 0.5 * memory occupancy + 0.7 * network throughput.
[0104] Next, update the monitoring metric weights using online learning. Divide the weighted performance metrics into training batches according to the time series. Adopt the online learning method to update the monitoring metric weights based on the stochastic gradient descent algorithm. The stochastic gradient descent algorithm uses L2 regularization to constrain the weight parameters and uses the exponential moving average method to smooth the parameter update process to obtain the optimized monitoring metric weights. Assume that each update uses the data of one time window as a training batch, the learning rate is set to 0.01, the regularization coefficient is set to 0.1, and the exponential moving average decay rate is set to 0.9.
[0105] Then, apply the optimized monitoring metric weights to real-time performance monitoring. Calculate the prediction deviation between the current system performance and the system performance prediction result. For example, if the current system end-to-end delay is 100 milliseconds and the system performance prediction result is 90 milliseconds, then the prediction deviation is 10 milliseconds. At the same time, segment the historical performance data by time window and calculate the performance difference between adjacent time windows to obtain the performance change trend. For example, if the average end-to-end delays in the past three time windows are 90 milliseconds, 95 milliseconds, and 100 milliseconds respectively, then the performance change trend is upward.
[0106] Next, input the prediction deviation and the performance change trend into the backpropagation algorithm. The backpropagation algorithm calculates the end-to-end delay gradient and the throughput gradient. Based on the end-to-end delay gradient and the throughput gradient, the performance gain gradient is weighted and synthesized. Assume that the end-to-end delay weight is 0.6 and the throughput weight is 0.4, then the performance gain gradient is 0.6 * end-to-end delay gradient + 0.4 * throughput gradient.
[0107] Finally, detect the change trend of the performance gain gradient. When the performance gain gradient shows attenuation for three consecutive time windows, activate the policy tuning engine. The policy tuning engine adjusts the deployed end-to-end optimization policy according to the change trend of the performance gain gradient and other relevant information to further improve the system performance.
[0108] The solution of this application can:
[0109] Improve monitoring accuracy: By adaptively adjusting the monitoring metric weights, the monitoring system pays more attention to the metrics closely related to the system performance, thus more accurately reflecting the system performance status. Improve the performance prediction accuracy: Using the multi-layer perceptron performance prediction model, it can more accurately predict the future performance of the system and provide a more reliable basis for policy adjustment. Achieve automated policy optimization: Guide the policy tuning engine through the performance gain gradient to achieve automated policy adjustment, reduce manual intervention, and improve the system operation and maintenance efficiency.
[0110] In an alternative embodiment, the weighted performance metrics are divided into training batches according to a time series, and an online learning method is used to update the weights of the monitoring metrics based on the stochastic gradient descent algorithm. The stochastic gradient descent algorithm uses L2 regularization to constrain the weight parameters and an exponential moving average method to smooth the parameter update process. The optimized weights of the monitoring metrics obtained include:
[0111] A double-layer sliding window structure is used to divide the weighted performance metrics according to a time series. The double-layer sliding window structure includes an inner time window, and each inner time window collects 300 performance data sampling points to form a training batch;
[0112] Based on the attention mechanism, the cosine distance between each sampling point in the training batch and the current system state is calculated to obtain a similarity value. The similarity value is input into the softmax function to obtain a sample weight vector, and the sample weight vector is used to perform weighted processing on the training batch to obtain weighted training data;
[0113] The weighted training data is input into the stochastic gradient descent algorithm. The stochastic gradient descent algorithm calculates the gradient vector of the weights of the monitoring metrics. When the norm of the gradient vector exceeds a preset vector threshold, the gradient vector is scaled proportionally while keeping the gradient direction unchanged to obtain an adjusted gradient vector;
[0114] An L2 regularization term is introduced into the loss function to constrain the update of the weights of the monitoring metrics. The regularization coefficient of the L2 regularization term decays linearly with the number of training rounds; the adjusted gradient vector is smoothed using the exponential moving average method, and the adjusted gradient vector and the historical gradient vector are weighted averaged according to a decay coefficient of 0.9 to obtain a smoothed gradient vector;
[0115] The weights of the monitoring metrics are updated based on the smoothed gradient vector to obtain updated weights of the monitoring metrics. The predicted mean squared error within multiple consecutive inner time windows is calculated to obtain the prediction accuracy rate, and the change standard deviation of the updated weights of the monitoring metrics is calculated to obtain the weight stability index;
[0116] The change trend of the loss in multiple consecutive training batches is monitored. When the change trend of the loss continues to decline, the learning rate is increased to accelerate convergence. When the change trend of the loss fluctuates, the learning rate is decreased to improve stability; the updated weights of the monitoring metrics and the adjusted learning rate are used for the next round of training iteration until the prediction accuracy rate and the weight stability index simultaneously meet the preset conditions, and then the optimized weights of the monitoring metrics are output.
[0117] First, collect system performance data and construct a performance dataset. For example, collect performance metrics such as CPU utilization, memory occupancy, and disk I / O rate every 1 second, and continuously collect data at 3000 time points to construct a performance dataset containing 3000 sampling points.
[0118] Then, use a double-layer sliding window structure to perform time series partitioning on the performance dataset. The length of the inner time window is set to 300 sampling points, that is, each inner time window contains 300 continuously collected performance data. The length of the outer time window is set according to actual needs, for example, set to 10 inner time windows. First, starting from the beginning of the performance dataset, sequentially intercept inner time windows with a length of 300 sampling points to form the first training batch. Subsequently, slide the inner time window backward by one sampling point and intercept the next inner time window containing 300 sampling points to form the second training batch, and so on until the entire performance dataset is traversed.
[0119] Next, perform weighted processing on each training batch. Obtain the system state at the current moment, such as CPU utilization, memory occupancy, disk I / O rate, etc. at the current moment. Calculate the cosine similarity between each sampling point in the training batch and the current system state, and input the similarity values into the softmax function to obtain the sample weight vector. For example, the similarity between the first sampling point and the current system state is 0.8, the second sampling point is 0.7, and so on. Input these similarity values into the softmax function to obtain a normalized sample weight vector, such as [0.2, 0.18, …]. Multiply the sample weight vector by each sampling point in the training batch to obtain weighted training data.
[0120] Subsequently, input the weighted training data into the stochastic gradient descent algorithm to calculate the gradient vector of the monitoring metric weights. Set a preset vector threshold, for example, 10. If the norm of the calculated gradient vector exceeds 10, then scale the gradient vector proportionally so that its norm is equal to 10 while keeping the gradient direction unchanged. For example, if the calculated gradient vector is [20, 15] and its norm is 25, which exceeds the threshold 10, then scale it to [8, 6], and the norm becomes 10 while the direction remains unchanged.
[0121] Introduce an L2 regularization term into the loss function to constrain the update of the monitoring metric weights. The regularization coefficient of the L2 regularization term decays linearly with the increase in the number of training rounds. For example, the initial regularization coefficient is set to 0.1, and after each round of training, the regularization coefficient decreases by 0.001. Smooth the adjusted gradient vector using the exponential moving average method, and set the decay coefficient to 0.9. For example, if the current gradient vector is [8, 6] and the historical gradient vector is [7, 5], then the smoothed gradient vector is 0.9 * [7, 5] + 0.1 * [8, 6] = [7.1, 5.1].
[0122] Update the monitoring metric weights based on the smoothed gradient vector. For example, if the current monitoring metric weights are [0.2, 0.3], the smoothed gradient vector is [7.1, 5.1], and the learning rate is 0.01, then the updated monitoring metric weights are [0.2 + 0.01 * 7.1, 0.3 + 0.01 * 5.1] = [0.271, 0.351]. Calculate the predicted mean squared error within 10 consecutive inner time windows to obtain the prediction accuracy. Calculate the standard deviation of the changes in the updated monitoring metric weights to obtain the weight stability metric.
[0123] Monitor the loss change trend for multiple consecutive training batches. If the loss continues to decrease, increase the learning rate, for example, increase the learning rate from 0.01 to 0.02. If the loss fluctuates, decrease the learning rate, for example, decrease the learning rate from 0.02 to 0.01. Use the updated monitoring metric weights and the adjusted learning rate for the next round of training iteration.
[0124] Repeat the above steps until both the prediction accuracy and the weight stability metric meet the preset conditions. For example, if the prediction accuracy reaches 95% and the weight stability metric is less than 0.01, then output the optimized monitoring metric weights.
[0125] The solution of this application can:
[0126] Improve the accuracy of monitoring metrics: Through online learning and the double-layer sliding window mechanism, it can dynamically adjust the monitoring metric weights according to the changes in the system state, thus more accurately reflecting the actual performance of the system. Enhance the stability of the model: By using L2 regularization and the exponential moving average method, it can effectively suppress the overfitting of the model and improve the generalization ability and stability of the model. Enhance the adaptability of the model: By dynamically adjusting the learning rate, the model can converge faster and adapt to the performance data changes in different scenarios.
[0127] In an alternative embodiment, the policy tuning engine dynamically adjusts the optimization parameters using a genetic algorithm, and simultaneously synchronizes the optimization records to the knowledge graph in real time; the knowledge graph continuously improves the optimization rule base through association analysis, provides knowledge support for subsequent optimization decisions, and realizes the adaptive evolution of the optimization strategy, including:
[0128] Construct a multi-layer knowledge graph, which includes an entity layer, a relationship layer, and an attribute layer; the policy tuning engine constructs chromosome encoding using a genetic algorithm; extracts the historical optimal configuration from the multi-layer knowledge graph, perturbs the parameters of the historical optimal configuration to generate an initial population; constructs a fitness function with system latency, throughput, and resource utilization, calculates the fitness distribution of the initial population, and obtains the population convergence degree and diversity index;
[0129] Dynamically adjust the selection pressure coefficient based on the population convergence degree, and adopt the roulette wheel strategy to select high-quality individuals; adaptively calculate the crossover probability according to the diversity index, and perform arithmetic crossover on the selected high-quality individuals; adaptively adjust the mutation probability based on the convergence speed of consecutive generations, and perform Gaussian mutation on the crossed individuals;
[0130] Form a new population with the individual with the best fitness in each generation and the mutated individuals; use the parameter configuration, performance data, and fitness evaluation results of the new population as optimization records, and synchronize them to the entity layer, the relationship layer, and the attribute layer in the multi-layer knowledge graph in real time;
[0131] Perform association analysis on the optimization records in the multi-layer knowledge graph, use the prior algorithm to mine parameter configuration rules, use time series analysis to identify performance change rules, and use causal reasoning to construct parameter adjustment paths; use the parameter configuration rules, the performance change rules, and the parameter adjustment paths as optimization strategies and store them in the rule base;
[0132] Calculate the confidence of the optimization strategies in the rule base, mark the strategies with a confidence lower than the preset confidence threshold as strategies to be verified; perform A / B testing in a verification environment to verify the effectiveness of the strategies to be verified; update the verified optimization strategies to the rule base to realize the dynamic optimization of the rule base;
[0133] Use a graph neural network to calculate the similarity between the current state and historical scenarios, and extract the optimization strategies corresponding to similar scenarios from the rule base; perform weighted combination based on the policy confidence to obtain a parameter adjustment plan; use the parameter adjustment plan as prior knowledge to guide the population evolution direction of the policy tuning engine;
[0134] Monitor the performance metrics during the optimization process and trigger parameter rollback when a performance degradation is detected; record the performance fluctuation scenarios into the multi-layer knowledge graph, and continuously improve the rule base through continuous two-way feedback to achieve the adaptive evolution of the optimization strategy.
[0135] First, construct a multi-layer knowledge graph. This knowledge graph consists of three core levels: the entity layer, the relationship layer, and the attribute layer. The entity layer is used to store the key components in the system, such as servers, databases, caches, etc. The relationship layer describes the associations between entities, such as the connection relationship between a server and a database. The attribute layer records the characteristic information of entities and relationships, such as the CPU utilization rate of a server, the response time of a database, etc. Taking an e-commerce platform as an example, the entity layer can include entities such as "orders", "products", "users", etc.; the relationship layer can include relationships such as "users place orders", "products are associated with orders", etc.; the attribute layer can include attributes such as "order amount", "product price", "user activity", etc.
[0136] Next, the strategy optimization engine constructs a chromosome encoding scheme using the genetic algorithm. Map the parameters to be optimized to the gene positions of the chromosome. For example, encode parameters such as the thread pool size of the server and the database connection pool size into the chromosome. Taking the database connection pool size as an example, assuming the value range of the connection pool size is from 10 to 100, a segment of gene positions on the chromosome can be used to represent the connection pool size. For example, it can be represented by a 7-bit binary number, corresponding to a decimal value range of 0 - 127, and then linearly mapped to 10 - 100 through linear mapping.
[0137] Then, extract the historical optimal configuration from the multi-layer knowledge graph. By querying the knowledge graph, obtain the parameter configuration when the system performance was the best in the past period. For example, by querying the database connection pool size, server thread pool size, etc. when the order processing speed of the e-commerce platform was the fastest in the past month, obtain the historical optimal configuration. Assume the historical optimal configuration is that the database connection pool size is 50 and the server thread pool size is 200.
[0138] Perform parameter perturbation on the historical optimal configuration to generate an initial population. Based on the historical optimal configuration, randomly adjust the parameters to generate multiple different parameter combinations to form the initial population. For example, randomly take values for the database connection pool size between 40 and 60, and randomly take values for the server thread pool size between 180 and 220 to generate 100 different parameter combinations.
[0139] After that, construct a fitness function. Combine key performance metrics such as system latency, throughput, and resource utilization into a fitness function to evaluate the quality of parameter configurations. For example, the fitness function can be defined as the weighted value of throughput minus the weighted value of latency, and then minus the opposite of the weighted value of resource utilization.
[0140] Calculate the fitness distribution of the initial population to obtain the population convergence degree and diversity index. Evaluate the fitness of each individual in the initial population, count the distribution of fitness values, and calculate the population convergence degree and diversity. For example, the variance of the fitness values in the population can be calculated to measure the diversity. The larger the variance, the better the diversity. The convergence degree can be measured by calculating the gap between the fitness value in the population and the optimal fitness value. The smaller the gap, the higher the convergence degree.
[0141] Dynamically adjust the selection pressure coefficient based on the population convergence degree, and adopt the roulette wheel strategy to select high-quality individuals. Adjust the selection pressure coefficient according to the size of the population convergence degree to control the probability of high-quality individuals being selected during the selection process. The higher the convergence degree, the greater the selection pressure, and the higher the probability of high-quality individuals being selected. For example, if the population convergence degree is relatively high, the selection pressure coefficient can be increased so that individuals with higher fitness values are more likely to be selected.
[0142] Adaptive calculate the crossover probability according to the diversity index, and perform arithmetic crossover on the selected high-quality individuals. Adjust the crossover probability adaptively according to the size of the population diversity to control the frequency of the crossover operation. The higher the diversity, the smaller the crossover probability. For example, if the population diversity is relatively high, the crossover probability can be reduced to decrease the number of crossover operations. Arithmetic crossover operation means that the gene positions of two parent individuals are weighted averaged to generate new offspring individuals.
[0143] Adaptive adjust the mutation probability based on the convergence speed of consecutive generations, and perform Gaussian mutation on the individuals after crossover. Adjust the mutation probability adaptively according to the change of the population convergence speed to control the intensity of the mutation operation. The faster the convergence speed, the smaller the mutation probability. For example, if the population convergence speed is relatively fast, the mutation probability can be reduced to decrease the amplitude of the mutation operation. Gaussian mutation operation means adding a random number obeying the Gaussian distribution to the gene position to perturb the gene position.
[0144] Form a new population by combining the individual with the best fitness in each generation with the mutated individuals. Combine the individual with the highest fitness in the current generation and the individuals after the mutation operation to form a new population and enter the next generation of evolution.
[0145] Take the parameter configuration, performance data, and fitness evaluation results of the new population as optimization records and synchronize them to the multi-layer knowledge graph in real time. Store the parameter configuration of each individual in the new population, the corresponding performance data, and the fitness evaluation results in the knowledge graph. For example, store information such as the database connection pool size of 55, the server thread pool size of 210, the corresponding throughput of 1000 TPS, the latency of 50 ms, the resource utilization rate of 80%, and the fitness value of 0.95 in the knowledge graph.
[0146] Perform correlation analysis on the optimization records in the multi-layer knowledge graph, use the prior algorithm to mine parameter configuration rules, use time series analysis to identify performance change rules, and use causal reasoning to construct parameter adjustment paths. By analyzing the historical optimization records stored in the knowledge graph, the correlation relationship between parameter configuration and performance indicators is mined. For example, it can be found that when the size of the database connection pool increases, the throughput will increase first and then decrease, thus obtaining a parameter configuration rule.
[0147] Store the parameter configuration rules, performance change rules, and parameter adjustment paths as optimization strategies in the rule library. Store the mined information such as parameter configuration rules, performance change rules, and parameter adjustment paths in the rule library as the basis for subsequent optimization decisions.
[0148] Calculate the confidence of the optimization strategies in the rule library, and mark the strategies with confidence lower than the preset confidence threshold as strategies to be verified. Evaluate the confidence of each optimization strategy in the rule library, and mark the strategies with confidence lower than the preset threshold as strategies to be verified. For example, if a certain rule is mined based on a small amount of data, its confidence may be low and further verification is required.
[0149] Execute A / B testing in the verification environment to verify the effectiveness of the strategies to be verified. Apply the strategies to be verified to the verification environment and compare them with the existing strategies to evaluate their effectiveness. For example, a new parameter configuration scheme can be applied to a part of users to observe whether their performance indicators have improved.
[0150] Update the verified optimization strategies to the rule library to achieve the dynamic optimization of the rule library. Update the strategies that have been verified to be effective to the rule library and increase their confidence.
[0151] Use graph neural network to calculate the similarity between the current state and historical scenarios, and extract the optimization strategies corresponding to similar scenarios from the rule library. According to the current state of the system, such as the current load situation, resource utilization rate, etc., use graph neural network to calculate the similarity between the current state and historical scenarios, and extract the optimization strategies corresponding to similar scenarios from the rule library.
[0152] Based on the strategy confidence, perform weighted combination to obtain a parameter adjustment plan. According to the confidence of the extracted optimization strategies, perform weighted combination to obtain the final parameter adjustment plan. The higher the confidence of the strategy, the greater its weight.
[0153] Use the parameter adjustment plan as prior knowledge to guide the population evolution direction of the strategy tuning engine. Use the obtained parameter adjustment plan as prior knowledge to guide the population evolution direction of the genetic algorithm. For example, the parameter adjustment plan can be used as part of the initial population.
[0154] Monitor the performance metrics during the optimization process and trigger parameter rollback when a performance degradation is detected. Continuously monitor the system performance metrics, and if a performance degradation is detected, roll back the parameters to the previous state.
[0155] Record the performance fluctuation scenarios into a multi-layer knowledge graph, and continuously improve the rule base through continuous two-way feedback to achieve the adaptive evolution of the optimization strategy. Record the performance fluctuation scenarios in the knowledge graph, and analyze the reasons for the performance fluctuations for improving the optimization strategy and the rule base to achieve the adaptive evolution of the optimization strategy.
[0156] The solution of this application can:
[0157] Improve system performance: By dynamically adjusting parameters, keep the system in the best state all the time, thereby improving system throughput, reducing latency, and increasing resource utilization. Reduce manual intervention: Through the adaptive evolution mechanism, automatically learn and optimize parameter configurations, reduce the need for manual intervention, and lower the operation and maintenance costs. Enhance system stability: Through the parameter rollback mechanism and the recording of performance fluctuation scenarios, timely discover and solve performance problems, and enhance the stability and reliability of the system.
[0158] In an optional implementation manner, dynamically adjust the selection pressure coefficient based on the population convergence degree, and adopt the roulette wheel strategy to select high-quality individuals; adaptively calculate the crossover probability according to the diversity index, and perform arithmetic crossover on the selected high-quality individuals; adaptively adjust the mutation probability based on the convergence speed of consecutive generations, and perform Gaussian mutation on the individuals after crossover, including:
[0159] Dynamically adjust the selection pressure coefficient based on the population convergence degree. When the population convergence degree is lower than the preset interval threshold, increase the selection pressure coefficient according to the first linear function. When the population convergence degree is higher than the preset interval threshold, decrease the selection pressure coefficient according to the second linear function; calculate the individual selection probability by dividing the power of the individual fitness value by the sum of the population fitness, and adopt the roulette wheel strategy to select high-quality individuals according to the individual selection probability;
[0160] Take the product of the compensation value of the diversity index and the adjustment coefficient as the dynamic adjustment amount, and add the dynamic adjustment amount to the basic crossover probability to obtain the adaptive crossover probability;
[0161] Calculate the crossover weight coefficient based on the fitness ratio of the high-quality individuals, and perform arithmetic crossover operations on the high-quality individuals according to the crossover weight coefficient and the adaptive crossover probability to generate offspring individuals;
[0162] Calculate the fitness change rate of the optimal individuals in adjacent generations to obtain the population convergence speed; when the population convergence speed is less than the preset convergence threshold, add the basic mutation probability and the mutation compensation amount to obtain the adaptive mutation probability, and when the population convergence speed is greater than the preset convergence threshold, multiply the basic mutation probability by the attenuation function to obtain the adaptive mutation probability;
[0163] Calculate the Gaussian mutation step size based on the population convergence speed, use the product of the Gaussian mutation step size and the normally distributed random number as the mutation amount, and apply the mutation amount to the offspring individuals according to the adaptive mutation probability to perform Gaussian mutation.
[0164] First, initialize the population. Randomly generate a certain number of individuals, each individual representing a potential solution to the problem, and these individuals constitute the initial population. For example, generate 100 individuals, each individual containing 10 genes, and the values of the genes are randomly taken between 0 and 1.
[0165] Then, evaluate the population fitness. Evaluate each individual in the population and calculate its fitness value. The fitness value reflects the quality of the solution of the individual to the problem. The higher the fitness value of an individual, the better the quality of its solution. For example, use the value of the objective function as the fitness value, and the larger the value of the objective function, the higher the fitness value.
[0166] Next, dynamically adjust the selection pressure coefficient based on the population convergence degree. Calculate the convergence degree of the current population. For example, the variance of the individual fitness values in the population can be used to measure it. Preset a convergence degree interval threshold. If the convergence degree of the current population is lower than the lower limit of the threshold, increase the selection pressure coefficient according to the first linear increasing method; if the convergence degree of the current population is higher than the upper limit of the threshold, decrease the selection pressure coefficient according to the second linear decreasing method; if the convergence degree of the current population is within the threshold interval, keep the selection pressure coefficient unchanged. For example, let the convergence degree interval threshold be [0.01, 0.1], the current convergence degree be 0.005, then increase the selection pressure coefficient; the current convergence degree is 0.2, then decrease the selection pressure coefficient; the current convergence degree is 0.05, then keep the selection pressure coefficient unchanged.
[0167] Perform the selection operation according to the adjusted selection pressure coefficient. Calculate the selection probability of each individual. After performing the selection pressure coefficient power operation on the individual fitness value and then dividing by the total population fitness value, obtain the selection probability of this individual. Adopt the roulette wheel strategy to select high-quality individuals. For example, the fitness value of a certain individual is 10, the selection pressure coefficient is 2, and the total population fitness value is 1000, then the selection probability of this individual is (10^2) / 1000 = 0.1.
[0168] Adaptive calculation of the crossover probability based on diversity metrics. Calculate the population diversity metric. For example, it can be measured by the average difference degree of the individual genes in the population. Multiply the compensation value of the diversity metric by the adjustment coefficient to obtain the dynamic adjustment amount. Add the dynamic adjustment amount to the basic crossover probability to obtain the adaptive crossover probability. For example, if the diversity metric is 0.8, the compensation value is 0.2, the adjustment coefficient is 0.5, and the basic crossover probability is 0.7, then the adaptive crossover probability is 0.7 + 0.2 * 0.5 = 0.8.
[0169] Perform the crossover operation. Calculate the crossover weight coefficient based on the fitness ratio of the selected high-quality individuals. Perform an arithmetic crossover operation on the selected high-quality individuals according to the crossover weight coefficient and the adaptive crossover probability to generate offspring individuals. For example, if the fitness values of two parent individuals are 10 and 20 respectively, then their crossover weight coefficients are 10 / (10 + 20) = 0.33 and 20 / (10 + 20) = 0.67 respectively. Assume the adaptive crossover probability is 0.8, then the crossover operation is performed with a probability of 0.8.
[0170] Adaptive adjustment of the mutation probability based on the convergence speed of consecutive generations. Calculate the fitness change rate of the optimal individuals in adjacent generations to obtain the population convergence speed. Preset a convergence speed threshold. If the population convergence speed is less than the convergence threshold, add the basic mutation probability and the mutation compensation amount to obtain the adaptive mutation probability; if the population convergence speed is greater than the convergence threshold, multiply the basic mutation probability by the attenuation function to obtain the adaptive mutation probability.
[0171] Perform the mutation operation. Calculate the Gaussian mutation step size based on the population convergence speed, and multiply the Gaussian mutation step size by a random number obeying the standard normal distribution as the mutation amount. Apply the mutation amount to the offspring individuals according to the adaptive mutation probability for Gaussian mutation.
[0172] Repeat steps three to eight until the termination condition is met, such as reaching the maximum number of iterations or finding a solution that meets the requirements.
[0173] The solution of this application can:
[0174] Improve the optimization ability of the algorithm: By dynamically adjusting the selection pressure, crossover probability, and mutation probability, it can better balance the global search and local search abilities of the algorithm, thereby improving the optimization ability of the algorithm. Enhance the robustness of the algorithm: By adaptively adjusting the parameters, the algorithm can better adapt to the characteristics of different problems, thereby enhancing the robustness of the algorithm. Speed up the convergence speed of the algorithm: By dynamically adjusting the mutation probability according to the population convergence speed, it can speed up the convergence speed of the algorithm.
[0175] Figure 2 This is a schematic structural diagram of the dynamic optimization system of the operator chain based on the computing engine model for the embodiments of the present invention, asFigure 2 As shown, the system includes:
[0176] A first unit for constructing a multi-dimensional feature perception network. The multi-dimensional feature perception network obtains operator chain operation data through a distributed acquisition layer, and the distributed acquisition layer is provided with a time-series data acquisition unit, a load feature acquisition unit, and a resource status acquisition unit; the time-series data acquisition unit dynamically tracks the data flow characteristics between operators based on a sliding window mechanism to construct a data flow fluctuation curve; the load feature acquisition unit uses an adaptive sampling algorithm to obtain operator calculation load characteristics and generate a load distribution heat map; the resource status acquisition unit records the system resource occupancy status based on resource profiling technology to form a multi-dimensional resource utilization rate matrix; the data flow fluctuation curve, the load distribution heat map, and the multi-dimensional resource utilization rate matrix are input into a deep learning model to train and generate an operator chain performance prediction model;
[0177] A second unit for constructing a hierarchical optimization decision mechanism based on the operator chain performance prediction model. In the policy layer, a reinforcement learning algorithm is used to train and optimize an agent model. The optimized agent model calculates the affinity coefficient of adjacent operators according to the prediction result; when the affinity coefficient is higher than the dynamic threshold, an automatic operator merging operation is triggered; in the execution layer, a heuristic algorithm is used to optimize the parallelism of the merged operators, and the number of instances is dynamically adjusted based on the computational complexity score; in the resource layer, an elastic scaling unit is constructed, and the elastic scaling unit automatically generates a fine-grained resource allocation scheme according to the change in parallelism; the automatic operator merging operation, the parallelism optimization, and the fine-grained resource allocation scheme are integrated into an end-to-end optimization strategy;
[0178] A third unit for deploying the end-to-end optimization strategy. During the execution process, an adaptive performance monitoring module is started. The adaptive performance monitoring module continuously optimizes the monitoring index weights based on an online learning method; the performance gain gradient is calculated through a backpropagation algorithm. When the performance gain gradient shows a decaying trend, a strategy tuning engine is activated; the strategy tuning engine uses a genetic algorithm to dynamically adjust the optimization parameters, and at the same time synchronizes the optimization records to the knowledge graph in real time; the knowledge graph continuously improves the optimization rule library through association analysis to provide knowledge support for subsequent optimization decisions and realizes the adaptive evolution of the optimization strategy.
[0179] In the third aspect of the embodiments of the present invention,
[0180] An electronic device is provided, including:
[0181] A processor;
[0182] A memory for storing processor-executable instructions;
[0183] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0184] In the fourth aspect of the embodiments of the present invention,
[0185] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0186] The present invention can be a method, an apparatus, a system, and / or a computer program product. The computer program product can include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are loaded.
[0187] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A dynamic optimization method based on a computing engine model-driven operator chain, characterized in that: include: Construct a multi-dimensional feature perception network, which obtains operator chain operation data through a distributed collection layer, and the distributed collection layer is provided with a time series data collection unit, a load feature collection unit and a resource status collection unit; the time series data collection unit dynamically tracks the data flow characteristics between operators based on a sliding window mechanism and constructs a data flow fluctuation curve; the load feature collection unit obtains the operator calculation load characteristics using an adaptive sampling algorithm and generates a load distribution heat map; the resource status collection unit records the system resource occupancy status based on resource profiling technology to form a multi-dimensional resource utilization matrix; the data flow fluctuation curve, the load distribution heat map and the multi-dimensional resource utilization matrix are input into a deep learning model to train and generate an operator chain performance prediction model; A hierarchical optimization decision-making mechanism is constructed based on the operator chain performance prediction model. The optimization agent model is trained using a reinforcement learning algorithm at the strategy layer. The optimization agent model calculates the affinity coefficients of adjacent operators based on the prediction results. When the affinity coefficient is higher than the dynamic threshold, the automatic merging operation of the operators is triggered. At the execution layer, a heuristic algorithm is used to optimize the parallelism of the merged operators, and the number of instances is dynamically adjusted based on the computational complexity score. Constructing an elastic scaling unit at the resource layer, the elastic scaling unit automatically generates a fine-grained resource allocation plan according to the change of parallelism; integrating the operator automatic merging operation, the parallelism optimization and the fine-grained resource allocation plan into an end-to-end optimization strategy; Deploy the end-to-end optimization strategy, and start an adaptive performance monitoring module during execution, wherein the adaptive performance monitoring module continuously optimizes monitoring indicator weights based on an online learning method; Calculate the performance gain gradient by using a back-propagation algorithm, and activate the strategy tuning engine when the performance gain gradient shows a decay trend; The strategy tuning engine uses genetic algorithms to dynamically adjust optimization parameters and synchronizes optimization records to the knowledge graph in real time; the knowledge graph continuously improves the optimization rule base through association analysis, provides knowledge support for subsequent optimization decisions, and realizes the adaptive evolution of optimization strategies.
2. The method according to claim 1, characterized in that A hierarchical optimization decision-making mechanism is constructed based on the operator chain performance prediction model. The optimization agent model is trained using a reinforcement learning algorithm at the strategy layer. The optimization agent model calculates the affinity coefficients of adjacent operators according to the prediction results, including: A hierarchical optimization decision mechanism is constructed based on the operator chain performance prediction model, and the hierarchical optimization decision mechanism includes a strategy layer, an execution layer, and a resource layer; the strategy layer receives the performance prediction result output by the operator chain performance prediction model, performs principal component analysis and dimensionality reduction processing on the performance prediction result, selects the principal component whose feature contribution rate is higher than the preset contribution threshold, maps the reduced-dimensional features to the interval of zero to one through the Min-Max normalization method, and generates a standardized feature matrix; A reinforcement learning algorithm is used at the strategy layer to train an optimization agent model, wherein the optimization agent model is based on a dual Q network architecture; the standardized feature matrix is input into the optimization agent model, a state space including operator computing load, data flow characteristics, and resource utilization is constructed, an action space of operator merging, splitting, and migration is defined, and a reward function based on performance gain is designed; The optimization agent model calculates affinity coefficients of adjacent operators based on training results, and the affinity coefficients are comprehensively calculated through three dimensions: data dependency intensity, resource complementarity, and performance improvement expectations.
3. The method according to claim 1, characterized in that At the execution layer, a heuristic algorithm is used to optimize the parallelism of the merged operators, and the number of instances is dynamically adjusted based on the computational complexity score. At the resource layer, an elastic scaling unit is built, which automatically generates a fine-grained resource allocation plan based on the change in parallelism. Integrating the operator automatic merging operation, the parallelism optimization and the fine-grained resource allocation scheme into an end-to-end optimization strategy includes: At the execution layer, a heuristic algorithm is used to optimize the parallelism of the merged operators, and a computational complexity scoring model is constructed to analyze the computational density, memory access pattern, and data dependency of the operators to generate static feature evaluation results. The runtime processor usage, memory occupancy, and input and output throughput are monitored to generate dynamic load evaluation results. The static feature evaluation results and the dynamic load evaluation results are weighted to generate an operator complexity scoring matrix. The operator complexity scoring matrix is input into a simulated annealing algorithm, the search temperature parameter is dynamically adjusted based on the degree of performance improvement, the optimal parallelism configuration is found in the search space, and local optimization is performed on the optimal parallelism configuration to obtain an optimized parallelism configuration scheme; the optimized parallelism configuration scheme is input into a time series analysis model to generate a short-term load change trend of the system, and the instance expansion and contraction threshold is set according to the short-term load change trend of the system to achieve dynamic adjustment of the number of instances; An elastic scaling unit is constructed at the resource layer. The elastic scaling unit establishes a mapping relationship model between parallelism and resource requirements according to the optimized parallelism configuration scheme and the instance scaling threshold. The mapping relationship model analyzes the operator complexity scoring matrix to identify an intensive load feature set. The intensive load feature set includes computing intensive load features, memory intensive load features, and input / output intensive load features. The intensive load feature set is input into a deep learning model to predict resource usage trends. The elastic scaling unit automatically generates a fine-grained resource allocation plan based on the resource usage trend, implements resource isolation for resource quotas, and executes resource policies. The merged operator, the optimized parallelism configuration scheme, and the fine-grained resource allocation plan are integrated into an end-to-end optimization strategy.
4. The method according to claim 1, characterized in that: Deploy the end-to-end optimization strategy, and start an adaptive performance monitoring module during execution, wherein the adaptive performance monitoring module continuously optimizes monitoring indicator weights based on an online learning method; The performance gain gradient is calculated by the back propagation algorithm. When the performance gain gradient shows a decay trend, the activation strategy tuning engine includes: Deploy the end-to-end optimization strategy, start the adaptive performance monitoring module, build a multi-layer perceptron performance prediction model based on the adaptive performance monitoring module, the input layer of the multi-layer perceptron performance prediction model receives basic monitoring indicators such as processor utilization, memory occupancy, and network throughput, the hidden layer of the multi-layer perceptron performance prediction model uses a ReLU activation function to perform nonlinear feature extraction, and the output layer of the multi-layer perceptron performance prediction model generates a system performance prediction result; The adaptive performance monitoring module uses a sliding time window mechanism to continuously collect performance data, calculates the Pearson correlation coefficient between the basic monitoring indicator and the system performance prediction result, and uses the Pearson correlation coefficient as the monitoring indicator weight to weightedly aggregate the basic monitoring indicator to obtain a weighted performance indicator; the weighted performance indicator is divided into training batches according to the time series, and the monitoring indicator weight is updated based on the stochastic gradient descent algorithm using an online learning method. The stochastic gradient descent algorithm uses L2 regularization to constrain the weight parameters, and uses the exponential moving average method to smooth the parameter update process to obtain the optimized monitoring indicator weight; Apply the optimized monitoring indicator weights to real-time performance monitoring, calculate the prediction deviation between the current system performance and the system performance prediction result, and calculate the performance difference between adjacent time windows for historical performance data by time window segmentation to obtain the performance change trend; The prediction deviation and the performance change trend are input into a back propagation algorithm, the back propagation algorithm calculates an end-to-end delay gradient and a throughput gradient, and weighted synthesizes a performance gain gradient based on the end-to-end delay gradient and the throughput gradient; the change trend of the performance gain gradient is detected, and when the performance gain gradient decays in three consecutive time windows, a policy tuning engine is activated.
5. The method according to claim 4, characterized in that The weighted performance index is divided into training batches according to the time series, and the monitoring index weight is updated based on the stochastic gradient descent algorithm using an online learning method. The stochastic gradient descent algorithm uses L2 regularization to constrain the weight parameters and uses the exponential moving average method to smooth the parameter update process. The optimized monitoring index weights include: A double-layer sliding window structure is used to divide the weighted performance index into time series, wherein the double-layer sliding window structure includes an inner time window, and each inner time window collects three hundred performance data sampling points to form a training batch; Based on the attention mechanism, the cosine distance between each sampling point in the training batch and the current system state is calculated to obtain a similarity value, and the similarity value is input into the softmax function to obtain a sample weight vector, and the sample weight vector is used to perform weighted processing on the training batch to obtain weighted training data; Inputting the weighted training data into a stochastic gradient descent algorithm, the stochastic gradient descent algorithm calculates a gradient vector of a monitoring indicator weight, and when a norm of the gradient vector exceeds a preset vector threshold, scaling the gradient vector proportionally to keep the gradient direction unchanged to obtain an adjusted gradient vector; An L2 regularization term is introduced into the loss function to constrain the update of the monitoring indicator weight, and the regularization coefficient of the L2 regularization term decays linearly with the training rounds; the adjusted gradient vector is smoothed by an exponential moving average method, and the adjusted gradient vector and the historical gradient vector are weighted averaged according to an attenuation coefficient of 0.9 to obtain a smoothed gradient vector; Based on the smooth gradient vector, the monitoring indicator weight is updated to obtain an updated monitoring indicator weight, the prediction mean square error in a plurality of consecutive inner time windows is calculated to obtain a prediction accuracy, and the change standard deviation of the updated monitoring indicator weight is calculated to obtain a weight stability index; Monitor the loss change trend of multiple consecutive training batches, increase the learning rate to accelerate convergence when the loss change trend continues to decline, and reduce the learning rate to improve stability when the loss change trend fluctuates; use the updated monitoring indicator weights and the adjusted learning rate for the next round of training iterations until the prediction accuracy and the weight stability index simultaneously meet the preset conditions, and output the optimized monitoring indicator weights.
6. The method according to claim 1, characterized in that The strategy tuning engine uses genetic algorithms to dynamically adjust optimization parameters and synchronizes optimization records to the knowledge graph in real time; The knowledge graph continuously improves the optimization rule base through association analysis, provides knowledge support for subsequent optimization decisions, and realizes the adaptive evolution of optimization strategies, including: Construct a multi-layer knowledge graph, wherein the multi-layer knowledge graph includes an entity layer, a relationship layer, and an attribute layer; the strategy tuning engine constructs chromosome encoding using a genetic algorithm; extracts the historical optimal configuration from the multi-layer knowledge graph, and performs parameter perturbations on the historical optimal configuration to generate an initial population; constructs a fitness function using system delay, throughput, and resource utilization, calculates the fitness distribution of the initial population, and obtains population convergence and diversity indicators; Dynamically adjusting the selection pressure coefficient based on the population convergence degree, and adopting the roulette strategy to select high-quality individuals; adaptively calculating the crossover probability based on the diversity index, and performing arithmetic crossover on the selected high-quality individuals; adaptively adjusting the mutation probability based on the convergence speed of consecutive generations, and performing Gaussian mutation on the individuals after the crossover; The individuals with the best fitness in each generation and the mutant individuals form a new population; the parameter configuration, performance data, and fitness evaluation results of the new population are used as optimization records and synchronized in real time to the entity layer, the relationship layer, and the attribute layer in the multi-layer knowledge graph; Performing association analysis on the optimization records in the multi-layer knowledge graph, mining parameter configuration rules using a priori algorithms, identifying performance change laws using time series analysis, and constructing parameter adjustment paths using causal reasoning; storing the parameter configuration rules, the performance change laws, and the parameter adjustment paths as optimization strategies in a rule library; Calculate the confidence of the optimization strategy in the rule base, and mark the strategy with a confidence lower than a preset confidence threshold as a strategy to be verified; perform A / B testing in a verification environment to verify the effectiveness of the strategy to be verified; update the optimization strategy that has passed the verification to the rule base to realize dynamic optimization of the rule base; A graph neural network is used to calculate the similarity between the current state and the historical scenario, and the optimization strategy corresponding to the similar scenario is extracted from the rule base; a parameter adjustment scheme is obtained by weighted combination based on the strategy confidence; and the parameter adjustment scheme is used as prior knowledge to guide the population evolution direction of the strategy tuning engine; Monitor the performance indicators during the optimization process, and trigger parameter rollback when performance degradation is detected; record performance fluctuation scenarios to the multi-layer knowledge graph, continuously improve the rule base through continuous two-way feedback, and realize the adaptive evolution of the optimization strategy.
7. The method according to claim 6, characterized in that Dynamically adjust the selection pressure coefficient based on the population convergence degree, and select high-quality individuals using a roulette strategy; adaptively calculate the crossover probability based on the diversity index, and perform arithmetic crossover on the selected high-quality individuals; Based on the convergence speed of successive generations, the mutation probability is adaptively adjusted, and Gaussian mutation is performed on the individuals after crossover, including: Dynamically adjusting the selection pressure coefficient based on the population convergence, increasing the selection pressure coefficient according to a first linear function when the population convergence is lower than a preset interval threshold, and reducing the selection pressure coefficient according to a second linear function when the population convergence is higher than the preset interval threshold; calculating the individual selection probability by dividing the power of the selection pressure coefficient of the individual fitness value by the sum of the group fitness, and selecting high-quality individuals according to the individual selection probability using a roulette strategy; The product of the compensation value of the diversity index and the adjustment coefficient is used as a dynamic adjustment amount, and the dynamic adjustment amount is added to the basic crossover probability to obtain an adaptive crossover probability; Calculating a crossover weight coefficient based on the fitness ratio of the high-quality individuals, and performing an arithmetic crossover operation on the high-quality individuals according to the crossover weight coefficient and the adaptive crossover probability to generate offspring individuals; The fitness change rate of the best individuals of two adjacent generations is calculated to obtain the population convergence speed; when the population convergence speed is less than a preset convergence threshold, the basic mutation probability is added to the mutation compensation amount to obtain the adaptive mutation probability; when the population convergence speed is greater than the preset convergence threshold, the basic mutation probability is multiplied by the attenuation function to obtain the adaptive mutation probability; The Gaussian variable step length is calculated based on the population convergence speed, the product of the Gaussian variable step length and a normally distributed random number is used as a mutation amount, and the mutation amount is applied to the offspring individual according to the adaptive mutation probability to perform Gaussian mutation.
8. A computing engine model-based operator chain dynamic optimization system, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to build a multi-dimensional feature perception network, which obtains operator chain operation data through a distributed collection layer, and the distributed collection layer is provided with a time series data collection unit, a load feature collection unit and a resource status collection unit; the time series data collection unit dynamically tracks the data flow characteristics between operators based on a sliding window mechanism and constructs a data flow fluctuation curve; the load feature collection unit uses an adaptive sampling algorithm to obtain the operator calculation load characteristics and generates a load distribution heat map; the resource status collection unit records the system resource occupancy status based on resource profiling technology to form a multi-dimensional resource utilization matrix; the data flow fluctuation curve, the load distribution heat map and the multi-dimensional resource utilization matrix are input into a deep learning model to train and generate an operator chain performance prediction model; The second unit is used to build a hierarchical optimization decision-making mechanism based on the operator chain performance prediction model, and use the reinforcement learning algorithm to train the optimization agent model at the strategy layer. The optimization agent model calculates the affinity coefficients of adjacent operators according to the prediction results; when the affinity coefficient is higher than the dynamic threshold, the automatic merging operation of the operators is triggered; at the execution layer, a heuristic algorithm is used to optimize the parallelism of the merged operators, and the number of instances is dynamically adjusted based on the computational complexity score; Constructing an elastic scaling unit at the resource layer, the elastic scaling unit automatically generates a fine-grained resource allocation plan according to the change of parallelism; integrating the operator automatic merging operation, the parallelism optimization and the fine-grained resource allocation plan into an end-to-end optimization strategy; A third unit is used to deploy the end-to-end optimization strategy and start an adaptive performance monitoring module during execution, wherein the adaptive performance monitoring module continuously optimizes the monitoring indicator weights based on an online learning method; Calculate the performance gain gradient by using a back-propagation algorithm, and activate the strategy tuning engine when the performance gain gradient shows a decay trend; The strategy tuning engine uses genetic algorithms to dynamically adjust optimization parameters and synchronizes optimization records to the knowledge graph in real time; the knowledge graph continuously improves the optimization rule base through association analysis, provides knowledge support for subsequent optimization decisions, and realizes the adaptive evolution of optimization strategies.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Robot path planning parallel optimization method based on load balancing and multi-search strategy
CN118605513A
Ray-based lightweight distributed reinforcement learning training platform design method
CN119151019A