Service proxy method and system based on Dores front-end node
By building a service agent system on the Doris front-end node, using technologies such as multi-head recursive attention weight matrix, causal convolution decomposition and graph attention convolution, node state vectors are extracted and probability graph evolution model is constructed, node performance change trends are predicted, hierarchical query plans are generated, and scheduling strategies are optimized. The problem that query scheduling methods in the existing technology cannot adapt to dynamic changes and heterogeneous requests is achieved, and efficient resource scheduling and system performance optimization are achieved.
Patent Information
- Application Number
- CN202510091348.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, query scheduling methods cannot effectively adapt to the dynamic changes in node performance and the heterogeneity of query requests, and do not fully utilize the relevant information between nodes and the experience of historical scheduling strategies.
The service agent method based on Doris front-end nodes is adopted, and node performance indicators are collected, node performance feature matrix is generated by collecting node performance indicators, mapping them into feature vectors, and causal convolution decomposition and graph attention convolution operations are performed to extract node state vectors. Then, a heterogeneous multi-layer connection pool is constructed, and dynamic parameters are calculated using online learning algorithms, an association matrix between nodes is established, a probability graph evolution model is constructed, and the node performance change trend is predicted. Based on this, a hierarchical query plan is generated, resource evaluation is performed, scheduling strategies are optimized, and the global optimal scheduling scheme is determined through the multi-agent reinforcement learning framework.
It improves query efficiency and overall system performance. By monitoring node performance and failure prediction in real time, it ensures the continuity of database services, effectively allocates and utilizes database resources, and avoids resource waste and bottlenecks.
Smart Images

Figure CN119938335A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of database access request processing, and in particular to a service proxy method and system based on Doris front-end nodes. Background Art
[0002] Distributed database systems play a key role in processing large-scale data and high-concurrency queries. To ensure efficient query processing and resource utilization, query scheduling strategies are crucial;
[0003] The goal of the query scheduling strategy is to assign incoming query requests to the most appropriate node to minimize query latency and maximize system throughput. Traditional query scheduling methods usually rely on simple heuristic methods, such as polling or random assignment based on node load, which cannot effectively adapt to the dynamic changes in node performance and the heterogeneity of query requests, nor do they make full use of the association information between nodes and the experience of historical scheduling strategies.
[0004] Therefore, a solution is urgently needed to solve the problems existing in the prior art. Summary of the invention
[0005] The embodiment of the present invention provides a service proxy method and system based on Doris front-end node, which can at least solve some problems existing in the prior art.
[0006] A first aspect of an embodiment of the present invention provides a service proxy method based on a Doris front-end node, comprising:
[0007] Collect the performance indicator sequence of the front-end node, map the performance indicator to a feature vector by calculating the multi-head recursive attention weight matrix, generate a node performance feature matrix and perform causal convolution on the time series data in the node performance feature matrix, extract the long-range time series dependency features, and perform graph attention convolution operations to extract the dynamic association features between nodes. Combine the dynamic association features and the long-range time series dependency features to obtain the node state vector, build a heterogeneous multi-layer connection pool, use a dual gradient-based online learning algorithm to iteratively calculate the dynamic parameters of each layer of the connection pool, establish the node association matrix by maximizing the conditional mutual information between the node representation vectors and introducing contrast noise constraints, perform variational inference to obtain the posterior probability distribution, and build a probabilistic graph evolution model of node performance based on Monte Carlo sampling;
[0008] Predict the performance change trend of each node based on the probabilistic graph evolution model, receive database access requests and parse them into heterogeneous query graphs, iteratively update node representations through the edge attention message passing mechanism, generate a hierarchical query plan, semantically encode the hierarchical query plan through relative position encoding and bidirectional cross attention mechanism, integrate the performance change trend to generate query resource demand representation and build a meta-learning-based policy network, optimize policy parameters through policy gradient and entropy regularization to perform resource evaluation, input the resource evaluation results into the multi-agent reinforcement learning framework, solve the global optimal scheduling solution based on the hierarchical soft maximum and minimum game and perform Monte Carlo tree search to expand the scheduling decision tree, determine the target node through upper confidence bound sampling, extract the discriminative features in the historical scheduling strategy through comparative distillation and migrate them to the online decision model, generate an initial scheduling decision and analyze heterogeneous causal effects based on the dual machine learning method, calibrate the decision results and forward the database access request to the target node;
[0009] The performance deviation value of each node is calculated based on the performance change trend and the real-time execution status. If the performance deviation value is greater than a preset performance deviation threshold, the current node is taken as a faulty node, and the recent work information of the faulty node is obtained to construct a fault feature sequence. The fault feature sequence is segmented according to a preset time window, and the statistical features within each time window are calculated, and a feature matrix is constructed and sorted. The timing pattern is extracted through a long short-term memory network, the fault development trend is predicted, and the unfinished access requests are divided into multiple priority groups through a clustering algorithm, and a weighted directed graph is constructed based on the network bandwidth matrix between nodes and the current load level. The minimum spanning tree algorithm is applied to select the optimal migration path, the performance indicators of the target node are collected, and the performance fluctuations are monitored. The probability graph evolution model is updated according to the status information of the faulty node after the migration is completed.
[0010] In an optional embodiment,
[0011] The performance indicator sequence of the front-end node is collected, and the performance indicator is mapped into a feature vector by calculating the multi-head recursive attention weight matrix, a node performance feature matrix is generated, and the time series data in the node performance feature matrix is causally convolved and decomposed to extract the long-range time series dependency features, and the graph attention convolution operation is performed to extract the dynamic association features between nodes, and the node state vector is obtained by combining the dynamic association features and the long-range time series dependency features, and a heterogeneous multi-layer connection pool is constructed. The dynamic parameters of each layer of the connection pool are iteratively calculated using an online learning algorithm based on dual gradients, and the node association matrix is established by maximizing the conditional mutual information between the node representation vectors and introducing contrast noise constraints. The variational inference is performed to obtain the posterior probability distribution and a probability graph evolution model of the node performance is constructed based on Monte Carlo sampling, including:
[0012] Continuously record the CPU usage rate, memory occupancy rate, disk IO read / write rate and network bandwidth usage rate according to the preset sampling interval, wherein the CPU usage rate includes user state occupancy, system state occupancy and IO waiting time, the memory occupancy rate includes physical memory usage, virtual memory usage and cache usage, the disk IO read / write rate includes read rate, write rate and average response time, and the network bandwidth usage rate includes inbound traffic, outbound traffic and number of TCP connections, to obtain the performance indicator sequence;
[0013] The performance indicator sequence is divided into multiple performance indicator subsequences according to the indicator type, an attention head is constructed for each performance indicator subsequence, an attention weight is recursively calculated by a hyperbolic tangent function to generate a feature vector, the feature vectors output by each attention head are concatenated to form a node performance feature vector, and the node performance feature vectors of each time step are combined to generate a node performance feature matrix;
[0014] The convolution kernel size and step size parameters are set, and the causal convolution decomposition is performed on the node performance feature matrix by using the causal filling method. The long-range temporal dependency is extracted by performing multi-layer convolution operations and increasing the number of convolution kernels layer by layer, and the long-range temporal dependency features are obtained by adding residual connections between convolution layers.
[0015] Construct an adjacency matrix between nodes and determine edge weights according to real-time network delays, process the adjacency matrix through multi-layer graph attention convolution, concatenate the outputs of each layer using skip connections to obtain dynamic correlation features, and fuse the long-range temporal dependency features with the dynamic correlation features through a gating mechanism to obtain a node state vector;
[0016] Constructing a multi-layer heterogeneous connection pool according to the node state vector, calculating the objective function gradient and dual variable of the connection weight, optimizing the primary variable and the dual variable alternately based on the gradient difference between the original problem and the dual problem, setting the update step of the dual variable to the inverse of the update step of the original variable, and attenuating the learning rate in proportion to the degree of convergence of the optimization objective function;
[0017] Randomly sample the node state vector and calculate the conditional mutual information, introduce Gaussian white noise to construct comparison samples, obtain the node association matrix by maximizing and minimizing the mutual information, perform variational inference on the association matrix through diagonal Gaussian distribution, iteratively optimize the distribution parameters through the stochastic gradient method until convergence, and obtain the posterior probability distribution;
[0018] A random sampling sequence is generated based on the posterior probability distribution, the state values in the sampling sequence are mapped to node attributes in the probabilistic graph model, directed edges are constructed according to the temporal correlation of the node states, the edge weights are determined based on the state transition probability, a Markov chain of the node states is established, the state transition matrix of the Markov chain is combined with the evolution law of the node performance indicator, and a node performance probabilistic graph evolution model that supports dynamic prediction is constructed.
[0019] In an optional embodiment,
[0020] The node state vector is randomly sampled and the conditional mutual information is calculated. Gaussian white noise is introduced to construct a comparison sample. The node association matrix is obtained by maximizing and minimizing the mutual information. The association matrix is subjected to variational inference through diagonal Gaussian distribution. The distribution parameters are iteratively optimized through the stochastic gradient method until convergence. The posterior probability distribution is obtained, including:
[0021] Constructing positive sample pairs from a set of node state vectors based on a sliding window mechanism, setting the time step length of the sliding window to an even number, selecting state vectors in the first half of the time step and the second half of the time step of the sliding window to form positive sample pairs, obtaining time series dimension features of the state vectors, using the positive sample pairs as basic training data for state association, and using the time series dimension features of the state vectors as benchmark data for state evolution;
[0022] Applying Gaussian white noise of different intensities to the state vector in the positive sample pair, obtaining a noise intensity parameter by uniformly sampling within a preset noise intensity range, multiplying the noise intensity parameter by the state vector to generate multiple groups of comparison samples, using the comparison samples as negative sample training data, and constructing a noise robustness verification mechanism based on the negative sample training data;
[0023] Obtaining the node load state and the network topology as conditional variables, combining the conditional variables with the positive sample pairs to calculate the joint probability distribution, using the probability distribution of the positive sample pairs on the conditional variables as the marginal probability distribution, calculating the conditional mutual information of the positive sample pairs according to the ratio of the joint probability distribution to the marginal probability distribution, and using the conditional mutual information as a measurement indicator of state correlation;
[0024] Calculating the conditional mutual information between the positive sample pair and the comparison sample, constructing an optimization objective function based on the mutual information, setting a maximization term for the mutual information of the positive sample pair and a minimization term for the mutual information of the comparison sample in the optimization objective function, and constructing an inter-node association matrix by optimizing the objective function;
[0025] Taking the diagonal Gaussian distribution as an approximate representation of the posterior distribution, taking the mean vector and variance vector of the diagonal Gaussian distribution as variational parameters, initializing the mean vector with random sampling values of the standard normal distribution, initializing the variance vector with a unit vector, calculating the variational lower bound through the Monte Carlo sampling method, and constructing a gradient update mechanism for the distribution parameters based on the variational lower bound;
[0026] An adaptive moment estimation optimization algorithm is used to iteratively optimize the variational parameters. In each round of iteration, the gradient value of the variational lower bound with respect to the distribution parameter is calculated. The optimization step size is dynamically adjusted according to a preset learning rate update strategy. The variational parameters are updated based on the gradient value and the optimization step size. The change of the variational lower bound during continuous iterations is monitored. When the change is lower than a preset threshold, the optimization process is determined to have converged. The converged variational parameters are output as posterior probability distributions.
[0027] In an optional embodiment,
[0028] Predict the performance change trend of each node based on the probabilistic graph evolution model, receive database access requests and parse them into heterogeneous query graphs, iteratively update node representations through edge attention message passing mechanism, generate hierarchical query plans, semantically encode the hierarchical query plans through relative position encoding and bidirectional cross attention mechanism, integrate the performance change trend to generate query resource demand representation and construct a meta-learning-based policy network, optimize policy parameters through policy gradient and entropy regularization to perform resource evaluation, input resource evaluation results into a multi-agent reinforcement learning framework, solve the global optimal scheduling solution based on hierarchical soft maximum and minimum game and perform Monte Carlo tree search to expand the scheduling decision tree, determine the target node through upper confidence bound sampling, extract the discriminative features in the historical scheduling strategy through comparative distillation and migrate them to the online decision model, generate initial scheduling decisions and analyze heterogeneous causal effects based on dual machine learning methods, calibrate the decision results and forward the database access request to the target node, including:
[0029] Receiving a database access request, parsing the database access request into a heterogeneous query graph, setting node type identification, data table size, and filter selection rate attribute information in the nodes of the heterogeneous query graph, and calculating the weight of the edge between the nodes based on the cardinality of the connection key and the data distribution characteristics;
[0030] Perform edge attention message passing on the heterogeneous query graph, perform linear mapping on the attribute information of each node to obtain an initial feature vector, aggregate the adjacent node features and edge weight information of the node to form a node message, calculate the attention weight through the node feature vector, perform weighted aggregation on the attention weight and the node message to iteratively update the node feature representation;
[0031] Generate a hierarchical query plan using the data table nodes in the heterogeneous query graph as leaf nodes, construct table scan nodes, filter nodes, connection nodes, and aggregation nodes from the bottom up in the hierarchical query plan, and record the scale, computational complexity, and resource demand statistics of input and output data of different nodes;
[0032] The hierarchical query plan is semantically encoded by using relative position encoding and bidirectional cross-attention mechanism, the internal dependency between nodes is captured by self-attention mechanism, the cross-attention mechanism is used to realize the fusion of forward information and reverse information, the query semantic vector is generated by combining residual connection and layer normalization, the performance change trend is fused with the query semantic vector to generate the query resource demand representation, a meta-learning-based policy network is constructed to output the resource allocation probability distribution and scheme evaluation results, and the policy parameters are optimized by policy gradient and entropy regularization to perform resource evaluation;
[0033] The resource evaluation results are input into the multi-agent reinforcement learning framework, and the node resource quota is determined at the resource allocation layer using the hierarchical soft maximum and minimum game mechanism. The global optimal scheduling solution is solved at the task scheduling layer, and the Monte Carlo tree search is performed to select, expand, simulate, and return operations to expand the scheduling decision tree, and the target node is determined by upper confidence bound sampling.
[0034] Extract discriminative features from historical scheduling strategies and migrate the discriminative features to an online decision model through a comparative distillation method, generate an initial scheduling decision based on the migrated features, use a dual machine learning method to analyze the heterogeneous causal effects of the initial scheduling decision, calibrate the initial scheduling decision according to the analysis results, and forward the database access request to the target node based on the calibrated initial scheduling decision.
[0035] In an optional embodiment,
[0036] Extracting discriminative features from historical scheduling strategies and migrating the discriminative features to an online decision model through a comparative distillation method, generating an initial scheduling decision based on the migrated features, analyzing the heterogeneous causal effects of the initial scheduling decision using a dual machine learning method, calibrating the initial scheduling decision according to the analysis results, and forwarding the database access request to the target node based on the calibrated initial scheduling decision includes:
[0037] Obtaining historical scheduling strategies from a historical scheduling strategy database, extracting resource allocation ratio parameters, task segmentation granularity parameters, and parallelism configuration parameters in the historical scheduling strategies, constructing the extracted parameters into a discriminative feature vector, and generating a discriminative feature set;
[0038] Training a teacher model based on the discriminative feature set, calculating a similarity score between a current query and a historical case, selecting a historical decision parameter with the highest similarity score, calculating a contrast loss between the decision parameter output by the teacher model and the historical decision parameter, updating the parameters of the online decision model by minimizing the contrast loss, and migrating the decision experience in the teacher model to the online decision model;
[0039] Receive a database access request, parse the database access request to obtain a query type identifier, a data scale value, and a resource status parameter, add the query type identifier, a data scale value, and a resource status parameter to the online decision model, generate an initial scheduling decision including a memory allocation ratio value, a parallelism configuration value, and a data shard size value, construct a treatment group and a control group, use the memory allocation ratio value, the parallelism configuration value, and the data shard size value in the initial scheduling decision as processing variables, use the query execution time value and the resource utilization value as result variables, train causal inference models in the treatment group and the control group, respectively, and calculate the prediction difference between the treatment group and the control group as a heterogeneous causal effect value;
[0040] According to the heterogeneous causal effect value, the memory allocation ratio value and the parallelism configuration value in the initial scheduling decision are adjusted to generate calibrated scheduling decision parameters, and the resource allocation ratio, number of task slices and execution parallelism of the target node are configured according to the calibrated scheduling decision parameters, and the database access request is sent to the target node to obtain the execution effect data and write it into the historical scheduling strategy database.
[0041] In an optional embodiment,
[0042] The performance deviation value of each node is calculated based on the performance change trend and the real-time execution status. If the performance deviation value is greater than a preset performance deviation threshold, the current node is taken as a fault node, and the recent work information of the fault node is obtained to construct a fault feature sequence. The fault feature sequence is segmented according to a preset time window, and the statistical features in each time window are calculated and a feature matrix is constructed and sorted. The timing pattern is extracted through a long short-term memory network, the fault development trend is predicted, and the unfinished access requests are divided into multiple priority groups through a clustering algorithm, and a weighted directed graph is constructed based on the network bandwidth matrix and the current load level between nodes. The minimum spanning tree algorithm is applied to select the optimal migration path, the performance indicators of the target node are collected, and the performance fluctuations are monitored. The probability graph evolution model is updated according to the status information of the fault node after the migration is completed, including:
[0043] Based on the real-time execution status, the processor usage rate, memory occupancy rate, disk waiting time and network throughput are collected to obtain a performance vector and set a corresponding performance vector weight. The relative deviation between the actual value of each performance indicator in the performance vector and the predicted value corresponding to the performance change trend is multiplied by the corresponding weight and then summed to obtain the performance deviation value of each node;
[0044] If the performance deviation value is greater than a preset performance deviation threshold, the current node is marked as a faulty node, recent work information of the faulty node is obtained and a fault feature sequence is constructed, wherein the recent work information includes request type, response time and error code, a fault feature sequence is constructed, and for the fault feature sequence, the fault feature sequence is segmented according to a preset time window, and the maximum, minimum and average value of the number of request processing in each time window are counted, the response time quantile value is calculated, the fluctuation range of the processor utilization rate and the memory utilization rate is recorded, the change pattern of the disk read and write rate is analyzed, and a feature matrix is generated and sorted in descending order according to the size of the eigenvalues;
[0045] Calculate the correlation coefficient between the features of each column of the sorted feature matrix and the degree of performance degradation, determine the correlation coefficient corresponding to the feature values in the top 10% based on the sorting results, take the features whose correlation coefficients are greater than a preset correlation threshold as strong correlation features, perform dimensionality reduction operations on the strong correlation features until a preset information retention rate is reached, generate a strong correlation sequence, and for the strong correlation sequence, extract the long-term dependency and short-term variation rules between the features through a long short-term memory network to obtain the time series pattern, perform prediction based on the time series pattern, and obtain the fault development trend;
[0046] Extract the execution time, resource usage, and data access volume of the unfinished access requests in the faulty node, combine them with the predicted node available time, construct the request feature vector after standardization, set the density clustering radius and the minimum number of samples, randomly select the core request object, expand the cluster cluster with the core request object as the center, calculate the similarity between requests, divide similar requests larger than the density clustering radius into priority groups according to the urgency of the fault development trend, and repeat the clustering expansion until all requests are accessed;
[0047] Construct a network bandwidth matrix between nodes, convert bandwidth values into transmission costs, construct a weighted directed graph, select the edge with the minimum transmission cost, construct a minimum spanning tree, and select the optimal migration path. If the performance indicator fluctuation of the target node exceeds the preset performance fluctuation threshold, select an alternative migration path, execute migration, and update the probability graph evolution model according to the status information of the faulty node after the migration is completed.
[0048] In an optional embodiment,
[0049] Construct a network bandwidth matrix between nodes, convert bandwidth values into transmission costs, construct a weighted directed graph, select the edge with the minimum transmission cost, construct a minimum spanning tree, and select the optimal migration path. If the performance index fluctuation of the target node exceeds the preset performance fluctuation threshold, select an alternative migration path including:
[0050] Obtain a cloud data center topology map where the faulty node is located, collect real-time network bandwidth values of all edges in the cloud data center topology map based on a network monitoring system, and construct a network bandwidth adjacency matrix, in which the measured bandwidth values are recorded at corresponding positions of nodes with direct connections, and infinity is recorded at corresponding positions of nodes without direct connections;
[0051] Obtaining the processor usage and memory usage of each node in the network bandwidth adjacency matrix, converting the network bandwidth adjacency matrix into a transmission cost matrix, wherein the transmission cost of each edge in the transmission cost matrix is calculated by the inverse of the bandwidth value, and when the processor usage of the target node is greater than a preset processor usage threshold or the memory usage is greater than a preset memory usage threshold, increasing the transmission cost of the corresponding edge by a preset ratio;
[0052] Constructing a weighted directed graph based on the transmission cost matrix, taking the faulty node as the root node, selecting a node corresponding to the minimum transmission cost edge connected to the root node as the first access node, adding the first access node to the minimum spanning tree, selecting a node corresponding to the minimum transmission cost edge connected to the visited node from the unvisited nodes as a new access node, adding the new access node to the minimum spanning tree, and repeating the process until all nodes are visited;
[0053] A request migration path is generated from the faulty node to the target node based on the minimum spanning tree, wherein the total transmission cost of the request migration path is minimal and the node does not pass through nodes whose load is higher than a preset load threshold. The performance indicators of each node on the request migration path are monitored. When the load of the node on the request migration path exceeds the load threshold or the bandwidth is lower than the preset bandwidth threshold, the transmission cost of each edge is recalculated and the request migration path is updated.
[0054] A second aspect of an embodiment of the present invention provides a service proxy system based on a Doris front-end node, comprising:
[0055] The first unit is used to collect the performance indicator sequence of the front-end node, map the performance indicator into a feature vector by calculating the multi-head recursive attention weight matrix, generate a node performance feature matrix and perform causal convolution decomposition on the time series data in the node performance feature matrix, extract the long-range time series dependency features, and perform graph attention convolution operation to extract the dynamic association features between nodes, combine the dynamic association features and the long-range time series dependency features to obtain the node state vector, construct a heterogeneous multi-layer connection pool, use an online learning algorithm based on dual gradient to iteratively calculate the dynamic parameters of each layer of the connection pool, establish the node association matrix by maximizing the conditional mutual information between the node representation vectors and introducing contrast noise constraints, perform variational inference to obtain the posterior probability distribution, and construct a probabilistic graph evolution model of node performance based on Monte Carlo sampling;
[0056] The second unit is used to predict the performance change trend of each node based on the probabilistic graph evolution model, receive database access requests and parse them into heterogeneous query graphs, iteratively update node representations through the edge attention message passing mechanism, generate a hierarchical query plan, semantically encode the hierarchical query plan through relative position encoding and bidirectional cross attention mechanism, integrate the performance change trend to generate query resource demand representation and construct a meta-learning-based policy network, optimize policy parameters through policy gradient and entropy regularization to perform resource evaluation, input the resource evaluation results into the multi-agent reinforcement learning framework, solve the global optimal scheduling plan based on the hierarchical soft maximum and minimum game and perform Monte Carlo tree search to expand the scheduling decision tree, determine the target node through upper confidence bound sampling, extract the discriminative features in the historical scheduling strategy through comparative distillation and migrate them to the online decision model, generate an initial scheduling decision and analyze heterogeneous causal effects based on the dual machine learning method, calibrate the decision results and forward the database access request to the target node;
[0057] The third unit is used to calculate the performance deviation value of each node based on the performance change trend and the real-time execution status. If the performance deviation value is greater than a preset performance deviation threshold, the current node is regarded as a faulty node, and the recent work information of the faulty node is obtained to construct a fault feature sequence. The fault feature sequence is segmented according to a preset time window, and the statistical features in each time window are calculated, and a feature matrix is constructed and sorted. The timing pattern is extracted through a long short-term memory network, the fault development trend is predicted, and the unfinished access requests are divided into multiple priority groups through a clustering algorithm, and a weighted directed graph is constructed based on the network bandwidth matrix between nodes and the current load level. The minimum spanning tree algorithm is applied to select the optimal migration path, the performance indicators of the target node are collected, and the performance fluctuations are monitored, and the probability graph evolution model is updated according to the status information of the faulty node after the migration is completed.
[0058] According to a third aspect of the embodiments of the present invention,
[0059] An electronic device is provided, comprising:
[0060] processor;
[0061] a memory for storing processor-executable instructions;
[0062] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0063] A fourth aspect of the embodiments of the present invention is:
[0064] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0065] In the present invention, by combining technologies such as node performance prediction, query analysis, resource evaluation and multi-agent reinforcement learning, a globally optimal scheduling plan can be generated, and requests can be forwarded to the most appropriate node, thereby improving query efficiency and overall system performance. By real-time monitoring of node performance and fault prediction, potential faulty nodes can be discovered in advance, and migration measures can be taken in time to ensure the continuity of database services. Based on node performance prediction and resource evaluation, database resources can be allocated and utilized more effectively, avoiding resource waste and bottlenecks, and improving the overall efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 It is a flow chart of a service proxy method based on Doris front-end node according to an embodiment of the present invention;
[0067] Figure 2 It is a structural diagram of a service proxy system based on Doris front-end node according to an embodiment of the present invention. DETAILED DESCRIPTION
[0068] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0069] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0070] Figure 1FIG. 1 is a flow chart of a service proxy method based on a Doris front-end node according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0071] S1. Collect the performance indicator sequence of the front-end node, map the performance indicator to a feature vector by calculating the multi-head recursive attention weight matrix, generate a node performance feature matrix and perform causal convolution on the time series data in the node performance feature matrix, extract the long-range time series dependency features, and perform graph attention convolution operations to extract the dynamic association features between nodes. Combine the dynamic association features and the long-range time series dependency features to obtain the node state vector, build a heterogeneous multi-layer connection pool, use a dual gradient-based online learning algorithm to iteratively calculate the dynamic parameters of each layer of the connection pool, establish the node association matrix by maximizing the conditional mutual information between the node representation vectors and introducing contrast noise constraints, perform variational inference to obtain the posterior probability distribution, and build a probabilistic graph evolution model of node performance based on Monte Carlo sampling;
[0072] The front-end node refers to the node closest to the user or data source in a distributed system or network architecture, which is responsible for processing user requests, data collection or other preliminary operations, and usually serves as the starting point or access point of the data flow. The node performance feature matrix is a multidimensional data matrix representing the performance of each node, in which each element contains the performance information of a specific node under a specific task, time point or condition, and is widely used to optimize the scheduling and resource management of networks or distributed systems. The causal convolution decomposition is a method of decomposing the convolution process in a complex system or network through causal reasoning, with the aim of revealing the causal relationship between different factors and effectively extracting features, which is often used in time series data or image data analysis. The long-range time series dependency feature refers to the correlation feature with a longer time span in time series data. Features are usually difficult to capture, but are of great significance for predicting future states or modeling dynamic changes. The graph attention convolution is a model that combines graph convolution networks and attention mechanisms. The attention mechanism is introduced into graph structured data, so that the model can automatically adjust the weights of adjacent nodes during the convolution process, thereby enhancing the model's attention to important nodes. The dual gradient-based online learning algorithm is an online learning algorithm based on a dual optimization method. It optimizes the model by calculating gradients and iteratively updating parameters. It can handle large-scale data streams and maintain real-time performance during training. The Monte Carlo sampling is a numerical calculation method based on random sampling. It estimates expected values, integrals or other mathematical quantities by randomly sampling from a probability distribution and performing statistical analysis based on these samples.
[0073] In an optional embodiment,
[0074] The performance indicator sequence of the front-end node is collected, and the performance indicator is mapped into a feature vector by calculating the multi-head recursive attention weight matrix, a node performance feature matrix is generated, and the time series data in the node performance feature matrix is causally convolved and decomposed to extract the long-range time series dependency features, and the graph attention convolution operation is performed to extract the dynamic association features between nodes, and the node state vector is obtained by combining the dynamic association features and the long-range time series dependency features, and a heterogeneous multi-layer connection pool is constructed. The dynamic parameters of each layer of the connection pool are iteratively calculated using an online learning algorithm based on dual gradients, and the node association matrix is established by maximizing the conditional mutual information between the node representation vectors and introducing contrast noise constraints. The variational inference is performed to obtain the posterior probability distribution and a probability graph evolution model of the node performance is constructed based on Monte Carlo sampling, including:
[0075] Continuously record the CPU usage rate, memory occupancy rate, disk IO read / write rate and network bandwidth usage rate according to the preset sampling interval, wherein the CPU usage rate includes user state occupancy, system state occupancy and IO waiting time, the memory occupancy rate includes physical memory usage, virtual memory usage and cache usage, the disk IO read / write rate includes read rate, write rate and average response time, and the network bandwidth usage rate includes inbound traffic, outbound traffic and number of TCP connections, to obtain the performance indicator sequence;
[0076] The performance indicator sequence is divided into multiple performance indicator subsequences according to the indicator type, an attention head is constructed for each performance indicator subsequence, an attention weight is recursively calculated by a hyperbolic tangent function to generate a feature vector, the feature vectors output by each attention head are concatenated to form a node performance feature vector, and the node performance feature vectors of each time step are combined to generate a node performance feature matrix;
[0077] The convolution kernel size and step size parameters are set, and the causal convolution decomposition is performed on the node performance feature matrix by using the causal filling method. The long-range temporal dependency is extracted by performing multi-layer convolution operations and increasing the number of convolution kernels layer by layer, and the long-range temporal dependency features are obtained by adding residual connections between convolution layers.
[0078] Construct an adjacency matrix between nodes and determine edge weights according to real-time network delays, process the adjacency matrix through multi-layer graph attention convolution, concatenate the outputs of each layer using skip connections to obtain dynamic correlation features, and fuse the long-range temporal dependency features with the dynamic correlation features through a gating mechanism to obtain a node state vector;
[0079] Constructing a multi-layer heterogeneous connection pool according to the node state vector, calculating the objective function gradient and dual variable of the connection weight, optimizing the primary variable and the dual variable alternately based on the gradient difference between the original problem and the dual problem, setting the update step of the dual variable to the inverse of the update step of the original variable, and attenuating the learning rate in proportion to the degree of convergence of the optimization objective function;
[0080] Randomly sample the node state vector and calculate the conditional mutual information, introduce Gaussian white noise to construct comparison samples, obtain the node association matrix by maximizing and minimizing the mutual information, perform variational inference on the association matrix through diagonal Gaussian distribution, iteratively optimize the distribution parameters through the stochastic gradient method until convergence, and obtain the posterior probability distribution;
[0081] A random sampling sequence is generated based on the posterior probability distribution, the state values in the sampling sequence are mapped to node attributes in the probabilistic graph model, directed edges are constructed according to the temporal correlation of the node states, the edge weights are determined based on the state transition probability, a Markov chain of the node states is established, the state transition matrix of the Markov chain is combined with the evolution law of the node performance indicator, and a node performance probabilistic graph evolution model that supports dynamic prediction is constructed.
[0082] The hyperbolic tangent function is a common activation function with an output value range between -1 and 1. It is used in neural networks to introduce nonlinearity so that the model can learn complex patterns. The jump connection is a common structure in neural networks. It directly passes the output of a layer to the subsequent layers while bypassing the intermediate layers, thereby effectively alleviating the gradient vanishing problem and accelerating training. The Gaussian white noise is a statistical noise with a power spectrum that is uniform over the entire frequency range and obeys a normal distribution. It is widely used in signal processing, image analysis, and physical models. The variational inference is a technique for approximate reasoning, especially in Bayesian inference, by optimizing the variational lower bound to approximate the true posterior distribution. It is used for reasoning problems of high-dimensional data and complex models. The stochastic gradient method is an iterative method for optimizing model parameters. Each iteration only uses a small batch of samples from the training set to update the model. It has a small amount of computation and is suitable for processing large-scale data.
[0083] At preset time intervals, such as every 1 second, various performance indicators of the front-end nodes are continuously collected. These indicators include CPU usage (divided into user state usage, system state usage and IO waiting time), memory usage (divided into physical memory usage, virtual memory usage and cache usage), disk IO read and write rate (divided into read rate, write rate and average response time) and network bandwidth usage (divided into inbound traffic, outbound traffic and number of TCP connections). The collected data are arranged in chronological order to form a performance indicator sequence. For example, the data collected in a certain second may be: CPU user state usage 60%, system state usage 10%, IO waiting time 5%, physical memory usage 8GB, virtual memory usage 16GB, cache usage 2GB, disk read rate 100MB / s, write rate 50MB / s, average response time 10ms, network inbound traffic 2Mbps, outbound traffic 1Mbps, TCP connection number 100.
[0084] The collected performance indicator sequence is divided into multiple subsequences according to the indicator type, such as the CPU usage subsequence, the memory occupancy subsequence, etc. For each performance indicator subsequence, one or more attention heads are constructed. Each attention head recursively calculates the correlation between the indicator values at different time steps in the subsequence to generate a weight matrix. This weight matrix reflects the degree of influence of the indicator values at different time steps on the current time step. Using this weight matrix, the indicator values in the subsequence are weighted and summed to obtain a feature vector. The feature vectors output by all attention heads are concatenated together to form the performance feature vector of the node. Then, the performance feature vectors generated at each time step are combined to form a node performance feature matrix.
[0085] Perform causal convolution decomposition on the generated node performance feature matrix to extract long-range temporal dependency features. Set the convolution kernel size and step size parameters, for example, the convolution kernel size is 3 and the step size is 1. Use causal filling, that is, only fill in past data to avoid future information leakage. Through multi-layer convolution operations, increase the number of convolution kernels layer by layer, for example, use 8 convolution kernels in the first layer, 16 convolution kernels in the second layer, and so on, to extract temporal features of different scales. Add residual connections between convolution layers and add the input directly to the output to avoid gradient vanishing and gradient exploding problems, and finally obtain long-range temporal dependency features.
[0086] Construct an adjacency matrix between nodes to represent the connection relationship between nodes. The elements in the matrix represent the connection strength between nodes, and the edge weight can be determined based on the real-time network delay. For example, the smaller the delay, the greater the weight. The adjacency matrix is processed by multi-layer graph attention convolution to extract the dynamic correlation features between nodes. Using skip connections, the outputs of each layer of graph attention convolution are spliced together to obtain dynamic correlation features. Then, through a gating mechanism, such as a sigmoid function, the long-range temporal dependency features and dynamic correlation features are fused to obtain the node state vector.
[0087] Construct a multi-layer heterogeneous connection pool, each of which contains multiple neurons. Use the node state vector as input to calculate the objective function gradient and dual variable of the connection weight. Based on the gradient difference between the original problem and the dual problem, alternately optimize the primary variable (connection weight) and the dual variable. Set the update step size of the dual variable to the inverse of the original variable update step size to ensure the stability of the algorithm. Decrease the learning rate proportionally according to the degree of convergence of the optimization objective function to increase the convergence speed.
[0088] The node state vectors are randomly sampled and the conditional mutual information is calculated. Gaussian white noise is introduced to construct contrast samples. The node association matrix is obtained by maximizing the mutual information between the node state vector and the positive sample and minimizing the mutual information with the negative sample. Variational inference is performed on the association matrix through diagonal Gaussian distribution, and the distribution parameters are iteratively optimized using the stochastic gradient method until convergence to obtain the posterior probability distribution.
[0089] Generate a random sampling sequence based on the posterior probability distribution, and map the state values in the sampling sequence to the node attributes in the probabilistic graph model. Construct directed edges based on the temporal correlation of node states, determine edge weights based on state transition probabilities, and establish a Markov chain of node states. Combine the state transition matrix of the Markov chain with the evolution law of node performance indicators to build a node performance probabilistic graph evolution model that supports dynamic prediction.
[0090] In this embodiment, through the multi-head recursive attention mechanism and causal convolution decomposition, the temporal dependency and long-range trend of node performance indicators can be more accurately captured, thereby improving the prediction accuracy. The introduction of graph attention convolution and contrastive learning can effectively extract the dynamic correlation features between nodes and enhance the robustness of the model to noise and outliers. The probabilistic graph evolution model can provide the probability distribution of node performance evolution, enhance the interpretability and credibility of the model, and help to better understand the laws of node performance changes.
[0091] In an optional embodiment,
[0092] The node state vector is randomly sampled and the conditional mutual information is calculated. Gaussian white noise is introduced to construct a comparison sample. The node association matrix is obtained by maximizing and minimizing the mutual information. The association matrix is subjected to variational inference through diagonal Gaussian distribution. The distribution parameters are iteratively optimized through the stochastic gradient method until convergence. The posterior probability distribution is obtained, including:
[0093] Constructing positive sample pairs from a set of node state vectors based on a sliding window mechanism, setting the time step length of the sliding window to an even number, selecting state vectors in the first half of the time step and the second half of the time step of the sliding window to form positive sample pairs, obtaining time series dimension features of the state vectors, using the positive sample pairs as basic training data for state association, and using the time series dimension features of the state vectors as benchmark data for state evolution;
[0094] Applying Gaussian white noise of different intensities to the state vector in the positive sample pair, obtaining a noise intensity parameter by uniformly sampling within a preset noise intensity range, multiplying the noise intensity parameter by the state vector to generate multiple groups of comparison samples, using the comparison samples as negative sample training data, and constructing a noise robustness verification mechanism based on the negative sample training data;
[0095] Obtaining the node load state and the network topology as conditional variables, combining the conditional variables with the positive sample pairs to calculate the joint probability distribution, using the probability distribution of the positive sample pairs on the conditional variables as the marginal probability distribution, calculating the conditional mutual information of the positive sample pairs according to the ratio of the joint probability distribution to the marginal probability distribution, and using the conditional mutual information as a measurement indicator of state correlation;
[0096] Calculating the conditional mutual information between the positive sample pair and the comparison sample, constructing an optimization objective function based on the mutual information, setting a maximization term for the mutual information of the positive sample pair and a minimization term for the mutual information of the comparison sample in the optimization objective function, and constructing an inter-node association matrix by optimizing the objective function;
[0097] Taking the diagonal Gaussian distribution as an approximate representation of the posterior distribution, taking the mean vector and variance vector of the diagonal Gaussian distribution as variational parameters, initializing the mean vector with random sampling values of the standard normal distribution, initializing the variance vector with a unit vector, calculating the variational lower bound through the Monte Carlo sampling method, and constructing a gradient update mechanism for the distribution parameters based on the variational lower bound;
[0098] An adaptive moment estimation optimization algorithm is used to iteratively optimize the variational parameters. In each round of iteration, the gradient value of the variational lower bound with respect to the distribution parameter is calculated. The optimization step size is dynamically adjusted according to a preset learning rate update strategy. The variational parameters are updated based on the gradient value and the optimization step size. The change of the variational lower bound during continuous iterations is monitored. When the change is lower than a preset threshold, the optimization process is determined to have converged. The converged variational parameters are output as posterior probability distributions.
[0099] The positive sample pair refers to a pair of samples used to train the model during the training process, one of which is the target category (positive class) and the other is another instance of the category or an instance similar to the target category. The conditional mutual information is a concept in information theory, which represents the amount of residual information between two random variables under the condition that certain variables are known, and is used to quantify the dependence of the two variables. The diagonal Gaussian distribution is a special Gaussian distribution whose covariance matrix is diagonal, indicating the independence of each dimension. It is often used in machine learning for high-dimensional data modeling and dimensionality reduction. The adaptive moment estimation optimization algorithm is an algorithm that dynamically adjusts the step size or other hyperparameters in the optimization process by estimating the moment of the data (such as mean, variance, etc.) in real time. It is often used to solve non-stationary problems and improve optimization efficiency. The variational lower bound is a technique for approximately inferring the posterior distribution by introducing an auxiliary variational distribution and minimizing the difference between it and the true distribution. It is often used in variational reasoning and Bayesian deep learning.
[0100] From the set of node state vectors, construct positive sample pairs based on the sliding window mechanism. For example, set the time step length of the sliding window to 6, and from the time series 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, we can construct positive sample pairs such as (1, 4), (2, 5), (3, 6), (4, 7), (5, 8), (6, 9), (7, 10). Record the characteristics of these state vectors in the time series dimension, such as the sequence number of the time step, as the benchmark data for state evolution. These positive sample pairs will serve as the basic training data for state association.
[0101] Gaussian white noise of different intensities is applied to the state vector in the positive sample pair to generate multiple groups of comparison samples as negative sample training data. For example, if the state vector is [0.1, 0.2, 0.3], and the noise intensity parameters are 0.01, 0.05, and 0.1, the generated comparison samples are [0.101, 0.201, 0.301], [0.105, 0.205, 0.305], and [0.11, 0.21, 0.31]. The noise intensity parameter is obtained by uniform sampling within a preset noise intensity interval (e.g., [0.01, 0.1]). These negative sample data will be used to build a noise robustness verification mechanism.
[0102] Obtain node load status, network topology, etc. as conditional variables. For example, the node load status can be CPU utilization, and the network topology can be the connection relationship between nodes. Combine the conditional variables with the positive sample pairs to calculate the joint probability distribution. Then use the probability distribution of the positive sample pairs on these conditional variables as the marginal probability distribution. Calculate the conditional mutual information of the positive sample pairs by the ratio of the joint probability distribution to the marginal probability distribution as a measure of state correlation. For example, assuming the joint probability is 0.8 and the marginal probability is 0.2, the conditional mutual information is 0.8 / 0.2=4.
[0103] Calculate the conditional mutual information between the positive sample pair and the comparison sample, and construct an optimization objective function based on the mutual information. The objective function contains the maximization term of the mutual information of the positive sample pair and the minimization term of the mutual information of the comparison sample. By optimizing the objective function, the node association matrix is constructed.
[0104] Use the diagonal Gaussian distribution as an approximation of the posterior distribution, and use the mean vector and variance vector of the diagonal Gaussian distribution as variational parameters. Initialize the mean vector with random sampling values from the standard normal distribution, and initialize the variance vector with a unit vector. For example, a three-dimensional mean vector can be initialized to [0.2, -0.5, 1.2], and the variance vector can be initialized to [1, 1, 1]. The variational lower bound is calculated by the Monte Carlo sampling method, and the gradient update mechanism of the distribution parameters is constructed based on the variational lower bound.
[0105] The variational parameters are iteratively optimized using an adaptive moment estimation optimization algorithm. In each iteration, the gradient value of the variational lower bound to the distribution parameter is calculated, and the optimization step size is dynamically adjusted according to the preset learning rate update strategy. The variational parameters are updated based on the gradient value and the optimization step size. The change in the variational lower bound is continuously monitored. When the change is lower than a preset threshold (e.g. 0.001), the optimization process is determined to have converged, and the converged variational parameters are output as the posterior probability distribution.
[0106] In this embodiment, by introducing conditional mutual information and comparison samples, the correlation between node states can be captured more accurately, and the impact of noise can be reduced, thereby improving the accuracy of state correlation analysis. By introducing Gaussian white noise to construct comparison samples and incorporating them into the optimization objective function, the robustness of the model to noise can be enhanced, making the model more stable and reliable in practical applications. Through the variational reasoning method, the posterior probability distribution can be effectively approximated, avoiding complex integral calculations, thereby improving the reasoning efficiency.
[0107] S2. Predict the performance change trend of each node based on the probabilistic graph evolution model, receive database access requests and parse them into heterogeneous query graphs, iteratively update node representations through the edge attention message passing mechanism, generate a hierarchical query plan, semantically encode the hierarchical query plan through relative position encoding and bidirectional cross-attention mechanism, integrate the performance change trend to generate query resource demand representation and construct a meta-learning-based policy network, optimize policy parameters through policy gradient and entropy regularization for resource evaluation, input the resource evaluation results into the multi-agent reinforcement learning framework, solve the global optimal scheduling plan based on the hierarchical soft maximum-minimum game and perform Monte Carlo tree search to expand the scheduling decision tree, determine the target node through upper confidence bound sampling, extract the discriminative features in the historical scheduling strategy through comparative distillation and migrate them to the online decision model, generate the initial scheduling decision and analyze the heterogeneous causal effects based on the dual machine learning method, calibrate the decision results and forward the database access request to the target node;
[0108] The performance change trend refers to the change in system or model performance over a period of time or under different conditions. It is usually used to evaluate the stability and effect of algorithms or equipment and help identify potential optimization space or bottlenecks. The heterogeneous query graph refers to a graph structure containing different types of nodes and edges, which is used to represent queries and relationships between a variety of different resources or entities. It is widely used in complex multi-source data analysis and information retrieval. The relative position encoding is a coding strategy in deep learning, which is used to represent the relative position of elements in the data rather than the absolute position. Especially in sequence modeling and graph neural networks, it can effectively capture the relationship between elements. The bidirectional cross attention mechanism is a mechanism that combines bidirectional information flow and cross attention. It can extract important features from two inputs at the same time and enhance the processing of mutually related information through the attention mechanism. It is widely used in multimodal tasks. The hierarchical soft maximum and minimum game is a game Game theory methods, by making maximum and minimum decisions at different levels to solve the optimal strategy, are usually used in game models and optimization problems in multi-layer decision-making systems. The upper confidence bound sampling is an exploration strategy commonly used in reinforcement learning. By using the upper confidence bound in the exploration process to balance exploration and utilization, the optimal action can be selected. The contrastive distillation is a model compression technology. By comparing the output differences between the original model and the distilled model, the parameters of the distilled model are optimized so that the performance of the original model can be retained while maintaining a smaller model. The dual machine learning method is a strategy that combines two different models (such as supervised learning and reinforcement learning) to deal with complex problems. It is often used to deal with data analysis with high dimensional or nonlinear characteristics. The heterogeneous causal effect refers to the fact that in causal reasoning, causal relationships may present different effects for different groups or conditions. It is usually used in personalized intervention, treatment effect analysis and other fields.
[0109] In an optional embodiment,
[0110] Predict the performance change trend of each node based on the probabilistic graph evolution model, receive database access requests and parse them into heterogeneous query graphs, iteratively update node representations through edge attention message passing mechanism, generate hierarchical query plans, semantically encode the hierarchical query plans through relative position encoding and bidirectional cross attention mechanism, integrate the performance change trend to generate query resource demand representation and construct a meta-learning-based policy network, optimize policy parameters through policy gradient and entropy regularization to perform resource evaluation, input resource evaluation results into a multi-agent reinforcement learning framework, solve the global optimal scheduling solution based on hierarchical soft maximum and minimum game and perform Monte Carlo tree search to expand the scheduling decision tree, determine the target node through upper confidence bound sampling, extract the discriminative features in the historical scheduling strategy through comparative distillation and migrate them to the online decision model, generate initial scheduling decisions and analyze heterogeneous causal effects based on dual machine learning methods, calibrate the decision results and forward the database access request to the target node, including:
[0111] Receiving a database access request, parsing the database access request into a heterogeneous query graph, setting node type identification, data table size, and filter selection rate attribute information in the nodes of the heterogeneous query graph, and calculating the weight of the edge between the nodes based on the cardinality of the connection key and the data distribution characteristics;
[0112] Perform edge attention message passing on the heterogeneous query graph, perform linear mapping on the attribute information of each node to obtain an initial feature vector, aggregate the adjacent node features and edge weight information of the node to form a node message, calculate the attention weight through the node feature vector, perform weighted aggregation on the attention weight and the node message to iteratively update the node feature representation;
[0113] Generate a hierarchical query plan using the data table nodes in the heterogeneous query graph as leaf nodes, construct table scan nodes, filter nodes, connection nodes, and aggregation nodes from the bottom up in the hierarchical query plan, and record the scale, computational complexity, and resource demand statistics of input and output data of different nodes;
[0114] The hierarchical query plan is semantically encoded by using relative position encoding and bidirectional cross-attention mechanism, the internal dependency between nodes is captured by self-attention mechanism, the cross-attention mechanism is used to realize the fusion of forward information and reverse information, the query semantic vector is generated by combining residual connection and layer normalization, the performance change trend is fused with the query semantic vector to generate the query resource demand representation, a meta-learning-based policy network is constructed to output the resource allocation probability distribution and scheme evaluation results, and the policy parameters are optimized by policy gradient and entropy regularization to perform resource evaluation;
[0115] The resource evaluation results are input into the multi-agent reinforcement learning framework, and the node resource quota is determined at the resource allocation layer using the hierarchical soft maximum and minimum game mechanism. The global optimal scheduling solution is solved at the task scheduling layer, and the Monte Carlo tree search is performed to select, expand, simulate, and return operations to expand the scheduling decision tree, and the target node is determined by upper confidence bound sampling.
[0116] Extract discriminative features from historical scheduling strategies and migrate the discriminative features to an online decision model through a comparative distillation method, generate an initial scheduling decision based on the migrated features, use a dual machine learning method to analyze the heterogeneous causal effects of the initial scheduling decision, calibrate the initial scheduling decision according to the analysis results, and forward the database access request to the target node based on the calibrated initial scheduling decision.
[0117] The Monte Carlo tree search is a search algorithm based on random sampling and tree structure, which is widely used in game theory and decision-making problems. It evaluates the effect of decisions by simulating a large number of possibilities and has high flexibility and scalability. The scheduling decision tree is a scheduling algorithm based on a decision tree structure. It splits the input features and gradually generates scheduling decisions, thereby optimizing resource allocation, task scheduling and other issues. It is widely used in the fields of industrial and cloud computing scheduling.
[0118] Receive database access requests and parse them into a heterogeneous query graph, in which each node contains a node type identifier, and calculate the weight of the edge between nodes according to the cardinality of the connection key and the data distribution characteristics (for example, the weight of the connection edge is 1);
[0119] The edge attention message passing mechanism is performed on the heterogeneous query graph to update the node representation. The attribute information of each node (node type, data table size, filter selection rate) is linearly mapped into an initial feature vector. For example, the initial feature vector of the "users" node can be [1, 1000000, 0.01]. Then, the adjacent node features and edge weight information of each node are aggregated to form the node message. For example, the message of the "users" node is the product of the feature vector of its adjacent node "orders" and the connecting edge weight. The attention weight is calculated by the node feature vector. For example, the attention weight is obtained by calculating the dot product of the feature vector of the "users" node and the feature vector of the "orders" node. Finally, the attention weight is weightedly aggregated with the node message, and the node feature representation is iteratively updated.
[0120] Generate a hierarchical query plan using the data table nodes in the heterogeneous query graph as leaf nodes. Construct table scan nodes, filter nodes, join nodes, and aggregation nodes from the bottom up. For example, for the above SQL query, scan nodes for accessing the "users" table and the "orders" table will be generated first, then filter nodes for "city='Beijing'" and "age>25" will be generated, then join nodes will be generated, and finally root nodes will be generated. At the same time, the scale, computational complexity, and resource requirement statistics of the input and output data of different nodes are recorded. For example, the output data scale of the "users" table scan node is 1 million rows, and the output data scale of the "city='Beijing'" filter node is 10,000 rows.
[0121] Relative position encoding and bidirectional cross-attention mechanism are used to semantically encode hierarchical query plans. The internal dependencies between nodes are captured by the self-attention mechanism. For example, the filtering node depends on the table scanning node. The cross-attention mechanism is used to achieve the fusion of forward information and reverse information. The query semantic vector is generated by combining residual connection and layer normalization. At the same time, the performance change trend of each node is predicted based on the probabilistic graph evolution model (for example, it is predicted that the access delay of the "users" table will increase by 20% in the next hour). The performance change trend is fused with the query semantic vector to generate a query resource demand representation. Then, a meta-learning-based policy network is constructed to output the resource allocation probability distribution and scheme evaluation results, and the policy parameters are optimized through policy gradient and entropy regularization for resource evaluation.
[0122] The resource evaluation results are input into the multi-agent reinforcement learning framework. The hierarchical soft maximum and minimum game mechanism is used to determine the node resource quota at the resource allocation layer and solve the global optimal scheduling solution at the task scheduling layer. Monte Carlo tree search is performed to expand the scheduling decision tree through selection, expansion, simulation, and feedback operations, and the target node is determined by upper confidence bound sampling.
[0123] Extract discriminative features from historical scheduling policies and migrate them to the online decision model through contrast distillation. Generate initial scheduling decisions based on the migrated features. Use dual machine learning methods to analyze the heterogeneous causal effects of the initial scheduling decisions and calibrate the initial scheduling decisions based on the analysis results. Forward database access requests to the target node based on the calibrated initial scheduling decisions. For example, forward query requests to nodes with sufficient computing resources and low access latency.
[0124] In this embodiment, the probabilistic graph evolution model is used to predict the trend of node performance changes, and multi-agent reinforcement learning is combined for resource scheduling, which can effectively allocate query requests to appropriate nodes, thereby reducing query delays and improving query efficiency. The hierarchical soft maximum-minimum game mechanism can effectively allocate cluster resources, avoid resource waste, and improve resource utilization. By comparing the application of distillation and dual machine learning methods, the accuracy and stability of the online decision-making model can be improved, the robustness of the system can be enhanced, and it can better adapt to complex and changeable database environments.
[0125] In an optional embodiment,
[0126] Extracting discriminative features from historical scheduling strategies and migrating the discriminative features to an online decision model through a comparative distillation method, generating an initial scheduling decision based on the migrated features, analyzing the heterogeneous causal effects of the initial scheduling decision using a dual machine learning method, calibrating the initial scheduling decision according to the analysis results, and forwarding the database access request to the target node based on the calibrated initial scheduling decision includes:
[0127] Obtaining historical scheduling strategies from a historical scheduling strategy database, extracting resource allocation ratio parameters, task segmentation granularity parameters, and parallelism configuration parameters in the historical scheduling strategies, constructing the extracted parameters into a discriminative feature vector, and generating a discriminative feature set;
[0128] Training a teacher model based on the discriminative feature set, calculating a similarity score between a current query and a historical case, selecting a historical decision parameter with the highest similarity score, calculating a contrast loss between the decision parameter output by the teacher model and the historical decision parameter, updating the parameters of the online decision model by minimizing the contrast loss, and migrating the decision experience in the teacher model to the online decision model;
[0129] Receive a database access request, parse the database access request to obtain a query type identifier, a data scale value, and a resource status parameter, add the query type identifier, a data scale value, and a resource status parameter to the online decision model, generate an initial scheduling decision including a memory allocation ratio value, a parallelism configuration value, and a data shard size value, construct a treatment group and a control group, use the memory allocation ratio value, the parallelism configuration value, and the data shard size value in the initial scheduling decision as processing variables, use the query execution time value and the resource utilization value as result variables, train causal inference models in the treatment group and the control group, respectively, and calculate the prediction difference between the treatment group and the control group as a heterogeneous causal effect value;
[0130] According to the heterogeneous causal effect value, the memory allocation ratio value and the parallelism configuration value in the initial scheduling decision are adjusted to generate calibrated scheduling decision parameters, and the resource allocation ratio, number of task slices and execution parallelism of the target node are configured according to the calibrated scheduling decision parameters, and the database access request is sent to the target node to obtain the execution effect data and write it into the historical scheduling strategy database.
[0131] Build a historical scheduling strategy database. This database records the scheduling information of previous database access requests, including query type identifier, data size value, resource status parameter, resource allocation ratio parameter, task segmentation granularity parameter, parallelism configuration parameter, and the final query execution time and resource utilization. For example, a record may contain the following information: the query type is an aggregate query, the data size is 1TB, the resource status is 80% CPU usage, 60% memory usage, the resource allocation ratio is 20% CPU, 80% memory, the task segmentation granularity is 1GB, the parallelism configuration is 10, the query execution time is 10 minutes, and the resource utilization is 90%.
[0132] Extract discriminative features from the historical scheduling strategy database. For each historical record, extract the resource allocation ratio parameter, task segmentation granularity parameter, and parallelism configuration parameter, and combine these parameters into a discriminative feature vector. For example, the discriminative feature vector of the above record can be expressed as [20%, 80%, 1GB, 10]. The discriminative feature vectors of all historical records constitute a discriminative feature set.
[0133] The teacher model is trained based on a set of discriminative features. The goal of the teacher model is to learn from the experience in historical scheduling strategies in order to provide initial scheduling decisions for new database access requests. The input of the teacher model is the discriminative feature vector of the current query, and the output is the corresponding scheduling decision parameters, including the memory allocation ratio value, the parallelism configuration value, and the data shard size value. For example, for a new aggregation query with a data scale of 500GB, a resource status of 50% CPU utilization and 40% memory utilization, the teacher model may output a memory allocation ratio of 70%, a parallelism configuration of 8, and a data shard size of 50GB.
[0134] In order to transfer the experience of the teacher model to the online decision model, a contrastive distillation method is used to calculate the similarity score between the current query and the historical case. For example, cosine similarity can be used to measure the similarity between the feature vectors of the current query and the historical record. Then, the historical decision parameter with the highest similarity score is selected. The decision parameters output by the teacher model are compared with the selected historical decision parameters to calculate the contrastive loss. For example, the mean square error can be used to calculate the loss. The parameters of the online decision model are updated by minimizing the contrastive loss, so that the decision experience of the teacher model is transferred to the online decision model.
[0135] When a new database access request is received, the request is parsed to obtain the query type identifier, data scale value, and resource status parameters. These parameters are input into the online decision model to generate an initial scheduling decision including the memory allocation ratio value, parallelism configuration value, and data shard size value.
[0136] In order to evaluate the effect of the initial scheduling decision and calibrate it, a dual machine learning method is used to analyze its heterogeneous causal effect. Database access requests are randomly divided into a treatment group and a control group. The treatment group adopts the initial scheduling decision generated by the online decision model, while the control group adopts other randomly selected scheduling strategies. The memory allocation ratio value, parallelism configuration value, and data shard size value in the initial scheduling decision are used as treatment variables, and the query execution time value and resource utilization value are used as outcome variables. The causal inference model is trained in the treatment group and the control group respectively. The prediction difference between the treatment group and the control group is calculated as the heterogeneous causal effect value.
[0137] According to the heterogeneous causal effect value, the memory allocation ratio, parallelism configuration value, and data shard size value in the initial scheduling decision are adjusted to generate calibrated scheduling decision parameters. According to the calibrated scheduling decision parameters, the resource allocation ratio, number of task shards, and execution parallelism of the target node are configured, and the database access request is sent to the target node for execution. The execution effect data, including query execution time and resource utilization, is obtained and written into the historical scheduling strategy database for subsequent model training and optimization.
[0138] In this embodiment, by learning from the experience in historical scheduling strategies and combining the characteristics of the current query and the status of system resources, a better scheduling decision is generated, thereby reducing the query execution time. By dynamically adjusting the resource allocation ratio, the number of task shards and the execution parallelism, the system resources are fully utilized, resource waste is avoided, and the overall resource utilization is improved. By continuously accumulating new scheduling data and adding it to the historical scheduling strategy database, the teacher model and the online decision-making model can be continuously optimized to make the scheduling strategy more intelligent and efficient.
[0139] S3. Calculate the performance deviation value of each node based on the performance change trend and the real-time execution status. If the performance deviation value is greater than a preset performance deviation threshold, take the current node as a faulty node, obtain the recent work information of the faulty node to construct a fault feature sequence, segment the fault feature sequence according to a preset time window, calculate the statistical features within each time window, construct a feature matrix, and sort the feature matrix. Extract the timing pattern through a long short-term memory network, predict the fault development trend, and divide the unfinished access requests into multiple priority groups through a clustering algorithm, and construct a weighted directed graph based on the network bandwidth matrix between nodes and the current load level. Apply the minimum spanning tree algorithm to select the optimal migration path, collect the performance indicators of the target node, monitor the performance fluctuations, and update the probability graph evolution model according to the status information of the faulty node after the migration is completed.
[0140] The fault development trend refers to the evolution process after the fault occurs, which is usually expressed as the severity of the fault, the expansion speed and its impact on the system, so as to predict the potential impact of the fault on the performance of the entire system, thereby providing decision support for maintenance and repair. The clustering algorithm is a type of machine learning algorithm, which is used to divide data points into several groups (clusters) according to certain similarity metrics, where the data points in the same group are as similar as possible, and the data points between different groups are quite different. Common clustering algorithms include K-means, hierarchical clustering, etc. The priority group refers to a method of grouping tasks or requests according to their urgency or importance. In scenarios such as task scheduling and resource allocation, the order of task processing is optimized by setting different priorities. The weighted directed graph is a graph structure in which the edges have different weights, indicating information such as the strength or cost of the relationship between nodes. It is widely used in optimization problems and network flow problems, and is often used in tasks such as path finding and resource scheduling. The minimum spanning tree algorithm is a graph algorithm used to find a minimum weight tree containing all nodes in a weighted undirected graph. Common algorithms include Kruskal algorithm and Prim algorithm, which are widely used in network design, resource connection and other problems.
[0141] In an optional embodiment,
[0142] The performance deviation value of each node is calculated based on the performance change trend and the real-time execution status. If the performance deviation value is greater than a preset performance deviation threshold, the current node is taken as a fault node, and the recent work information of the fault node is obtained to construct a fault feature sequence. The fault feature sequence is segmented according to a preset time window, and the statistical features in each time window are calculated and a feature matrix is constructed and sorted. The timing pattern is extracted through a long short-term memory network, the fault development trend is predicted, and the unfinished access requests are divided into multiple priority groups through a clustering algorithm, and a weighted directed graph is constructed based on the network bandwidth matrix and the current load level between nodes. The minimum spanning tree algorithm is applied to select the optimal migration path, the performance indicators of the target node are collected, and the performance fluctuations are monitored. The probability graph evolution model is updated according to the status information of the fault node after the migration is completed, including:
[0143] Based on the real-time execution status, the processor usage rate, memory occupancy rate, disk waiting time and network throughput are collected to obtain a performance vector and set a corresponding performance vector weight. The relative deviation between the actual value of each performance indicator in the performance vector and the predicted value corresponding to the performance change trend is multiplied by the corresponding weight and then summed to obtain the performance deviation value of each node;
[0144] If the performance deviation value is greater than a preset performance deviation threshold, the current node is marked as a faulty node, recent work information of the faulty node is obtained and a fault feature sequence is constructed, wherein the recent work information includes request type, response time and error code, a fault feature sequence is constructed, and for the fault feature sequence, the fault feature sequence is segmented according to a preset time window, and the maximum, minimum and average value of the number of request processing in each time window are counted, the response time quantile value is calculated, the fluctuation range of the processor utilization rate and the memory utilization rate is recorded, the change pattern of the disk read and write rate is analyzed, and a feature matrix is generated and sorted in descending order according to the size of the eigenvalues;
[0145] Calculate the correlation coefficient between the features of each column of the sorted feature matrix and the degree of performance degradation, determine the correlation coefficient corresponding to the feature values in the top 10% based on the sorting results, take the features whose correlation coefficients are greater than a preset correlation threshold as strong correlation features, perform dimensionality reduction operations on the strong correlation features until a preset information retention rate is reached, generate a strong correlation sequence, and for the strong correlation sequence, extract the long-term dependency and short-term variation rules between the features through a long short-term memory network to obtain the time series pattern, perform prediction based on the time series pattern, and obtain the fault development trend;
[0146] Extract the execution time, resource usage, and data access volume of the unfinished access requests in the faulty node, combine them with the predicted node available time, construct the request feature vector after standardization, set the density clustering radius and the minimum number of samples, randomly select the core request object, expand the cluster cluster with the core request object as the center, calculate the similarity between requests, divide similar requests larger than the density clustering radius into priority groups according to the urgency of the fault development trend, and repeat the clustering expansion until all requests are accessed;
[0147] Construct a network bandwidth matrix between nodes, convert bandwidth values into transmission costs, construct a weighted directed graph, select the edge with the minimum transmission cost, construct a minimum spanning tree, and select the optimal migration path. If the performance indicator fluctuation of the target node exceeds the preset performance fluctuation threshold, select an alternative migration path, execute migration, and update the probability graph evolution model according to the status information of the faulty node after the migration is completed.
[0148] The degree of performance degradation refers to the degree of decline in the performance of a system, device or algorithm compared to its initial state after long-term operation or in a harsh environment. It is usually used to evaluate system reliability and lifespan. The core request object refers to the most important, frequent or critical request object in certain systems or networks. It is usually the key point of user demand or system resource allocation. Its performance and response time directly affect the overall performance of the system.
[0149] Collect the real-time execution status of the node, including processor usage, memory occupancy, disk waiting time, and network throughput. For example, collect these indicators every 1 second to form a performance vector. Set corresponding weights for each performance indicator, such as 0.4 for processor usage, 0.3 for memory occupancy, 0.2 for disk waiting time, and 0.1 for network throughput. At the same time, based on historical data and current trends, predict the future value of each performance indicator, such as predicting the performance indicator value in the next 5 minutes. Multiply the relative deviation between the actual value and the predicted value of each performance indicator by the corresponding weight and sum them to obtain the performance deviation value of each node.
[0150] Compare the performance deviation value with the preset performance deviation threshold. Assume that the performance deviation threshold is set to 0.2. If the performance deviation value of a node is greater than 0.2, mark the node as a faulty node. Obtain the recent work information of the faulty node, including request type, response time, and error code. For example, collect all request information processed by the node in the past 1 hour. Construct this information into a fault feature sequence.
[0151] Segment the fault feature sequence according to the preset time window. For example, with a time window of 10 minutes, divide the fault feature sequence of the past hour into 6 segments. For each time window, count the highest, lowest, and average values of the number of request processing, calculate the quantile value of the response time (for example, 25% quantile, 50% quantile, 75% quantile), record the fluctuation range of processor utilization and memory utilization, and analyze the change pattern of disk read and write rates. Combine these statistical features into a feature matrix and sort them in descending order according to the size of the eigenvalues.
[0152] Calculate the correlation coefficient between the features of each column of the sorted feature matrix and the degree of performance degradation. Assume that the degree of performance degradation is represented by a performance deviation value. Determine the correlation coefficient corresponding to the eigenvalues in the top 10%. Take the features whose correlation coefficient is greater than a preset correlation threshold (for example, the threshold is set to 0.8) as strongly correlated features. Perform dimensionality reduction operations on strongly correlated features, such as principal component analysis, until a preset information retention rate (for example, 95%) is reached to generate a strongly correlated sequence.
[0153] The long short-term memory network is used to extract the long-term dependency and short-term change rules between the features in the strongly correlated sequence to obtain the time series pattern. Based on the time series pattern, prediction is performed to obtain the fault development trend, such as predicting the performance deviation value change trend of the node in the next hour.
[0154] Extract the execution time, resource usage, and data access information of the unfinished access requests in the faulty node. Combined with the predicted node availability time, standardize this information and construct a request feature vector. Set the density clustering radius and the minimum number of samples, for example, the radius is set to 0.5 and the minimum number of samples is set to 5. Randomly select a core request object and expand the cluster cluster with the core request object as the center. Calculate the similarity between requests, for example, using cosine similarity. Divide similar requests that are larger than the density clustering radius into priority groups according to the urgency of the fault development trend. Repeat cluster expansion until all unfinished access requests are divided into priority groups.
[0155] Construct a network bandwidth matrix between nodes and convert the bandwidth value into transmission cost. For example, the larger the bandwidth value, the smaller the transmission cost. Construct a weighted directed graph based on the transmission cost. Apply the minimum spanning tree algorithm to select the optimal migration path and migrate the access requests of the high-priority group to nodes with good performance. If the performance index fluctuation of the target node exceeds the preset performance fluctuation threshold (for example, 0.1), select an alternative migration path. Perform the migration operation, and use the status information of the faulty node after the migration is completed to update the probability graph evolution model to better predict the failure probability of future nodes.
[0156] In this embodiment, by predicting the development trend of faults and intelligently migrating access requests, service interruptions caused by faulty nodes can be effectively avoided, thereby improving the reliability and availability of the system. Request migration is performed according to the real-time performance and load conditions of the nodes, which can balance the load, optimize resource utilization, avoid resource waste, and dynamically adjust the migration strategy according to the operating status of the system, thereby enhancing the system's adaptability and robustness.
[0157] In an optional embodiment,
[0158] Construct a network bandwidth matrix between nodes, convert bandwidth values into transmission costs, construct a weighted directed graph, select the edge with the minimum transmission cost, construct a minimum spanning tree, and select the optimal migration path. If the performance index fluctuation of the target node exceeds the preset performance fluctuation threshold, select an alternative migration path including:
[0159] Obtain a cloud data center topology map where the faulty node is located, collect real-time network bandwidth values of all edges in the cloud data center topology map based on a network monitoring system, and construct a network bandwidth adjacency matrix, in which the measured bandwidth values are recorded at corresponding positions of nodes with direct connections, and infinity is recorded at corresponding positions of nodes without direct connections;
[0160] Obtaining the processor usage and memory usage of each node in the network bandwidth adjacency matrix, converting the network bandwidth adjacency matrix into a transmission cost matrix, wherein the transmission cost of each edge in the transmission cost matrix is calculated by the inverse of the bandwidth value, and when the processor usage of the target node is greater than a preset processor usage threshold or the memory usage is greater than a preset memory usage threshold, increasing the transmission cost of the corresponding edge by a preset ratio;
[0161] Constructing a weighted directed graph based on the transmission cost matrix, taking the faulty node as the root node, selecting a node corresponding to the minimum transmission cost edge connected to the root node as the first access node, adding the first access node to the minimum spanning tree, selecting a node corresponding to the minimum transmission cost edge connected to the visited node from the unvisited nodes as a new access node, adding the new access node to the minimum spanning tree, and repeating the process until all nodes are visited;
[0162] A request migration path is generated from the faulty node to the target node based on the minimum spanning tree, wherein the total transmission cost of the request migration path is minimal and the node does not pass through nodes whose load is higher than a preset load threshold. The performance indicators of each node on the request migration path are monitored. When the load of the node on the request migration path exceeds the load threshold or the bandwidth is lower than the preset bandwidth threshold, the transmission cost of each edge is recalculated and the request migration path is updated.
[0163] The transmission cost refers to the cost required to transmit data in a computer network or communication system, usually including time delay, bandwidth consumption, power consumption or monetary cost, etc., and is often used to optimize network design and traffic management.
[0164] Get the cloud data center topology map where the faulty node is located. The topology map describes all nodes in the data center and the connection relationship between them. For example, the topology map can be represented as a list, where each element represents a node and contains the node information directly connected to it. Assume that there are four nodes A, B, C, and D in the data center. A is connected to B and C, B is connected to A and D, C is connected to A and D, and D is connected to B and C.
[0165] Based on the network monitoring system, the real-time network bandwidth values of all edges in the topology map are collected to construct a network bandwidth adjacency matrix. In this matrix, if there is a direct connection between two nodes, the measured bandwidth value is recorded at the corresponding position; otherwise, a large value is recorded, representing infinity. For example, suppose the bandwidth between A and B is 10Gbps, the bandwidth between A and C is 5Gbps, the bandwidth between B and D is 20Gbps, and the bandwidth between C and D is 15Gbps;
[0166] Get the processor usage and memory usage of each node in the adjacency matrix. Convert the network bandwidth adjacency matrix into a transmission cost matrix. The transmission cost of each edge in the transmission cost matrix is calculated by the inverse of the bandwidth value. For example, if the bandwidth between A and B is 10Gbps, the transmission cost is 1 / 10. In addition, when the processor usage of the target node is greater than the preset processor usage threshold (e.g., 80%) or the memory usage is greater than the preset memory usage threshold (e.g., 90%), the transmission cost of the corresponding edge is increased by a preset proportion (e.g., 50%). Assuming that the processor usage of node B is 90%, which exceeds the threshold, the transmission cost from A to B becomes (1 / 10)*1.5=0.15.
[0167] Construct a weighted directed graph based on the transmission cost matrix. Take the faulty node as the root node and use the minimum spanning tree algorithm to construct a minimum spanning tree. For example, if node A fails, select the node corresponding to the minimum transmission cost edge connected to A. Assuming that the transmission cost from A to C is the smallest, add C to the minimum spanning tree. Next, select the node corresponding to the minimum transmission cost edge connected to the visited nodes (A and C) from the unvisited nodes. Repeat until all nodes are visited and finally obtain a minimum spanning tree.
[0168] Generate a request migration path from the faulty node to the target node based on the minimum spanning tree. The total transmission cost of this path is the smallest and does not pass through nodes whose load is higher than the preset load threshold. Assuming the target node is D, the generated migration path may be A->C->D. Continuously monitor the performance indicators of each node on the migration path, such as load and bandwidth. When the load of a node on the path exceeds the preset load threshold or the bandwidth is lower than the preset bandwidth threshold, recalculate the transmission cost of each edge and update the migration path to select a better path.
[0169] In this embodiment, by migrating the requests of the faulty node to other normal nodes, the continuity of the service can be guaranteed and the overall service interruption caused by a single point failure can be avoided. Taking factors such as the node load and network bandwidth into consideration, the requests can be migrated to nodes with lower resource utilization to avoid excessive resource use and improve the overall resource utilization efficiency. The migration path can be dynamically adjusted according to the real-time network status and node load, thereby enhancing the robustness and adaptability of the system and enabling it to better cope with various emergencies.
[0170] Figure 2 FIG. 1 is a schematic diagram of the structure of a service proxy system based on a Doris front-end node according to an embodiment of the present invention. Figure 2 As shown, the system comprises:
[0171] The first unit is used to collect the performance indicator sequence of the front-end node, map the performance indicator into a feature vector by calculating the multi-head recursive attention weight matrix, generate a node performance feature matrix and perform causal convolution decomposition on the time series data in the node performance feature matrix, extract the long-range time series dependency features, and perform graph attention convolution operation to extract the dynamic association features between nodes, combine the dynamic association features and the long-range time series dependency features to obtain the node state vector, construct a heterogeneous multi-layer connection pool, use an online learning algorithm based on dual gradient to iteratively calculate the dynamic parameters of each layer of the connection pool, establish the node association matrix by maximizing the conditional mutual information between the node representation vectors and introducing contrast noise constraints, perform variational inference to obtain the posterior probability distribution, and construct a probabilistic graph evolution model of node performance based on Monte Carlo sampling;
[0172] The second unit is used to predict the performance change trend of each node based on the probabilistic graph evolution model, receive database access requests and parse them into heterogeneous query graphs, iteratively update node representations through the edge attention message passing mechanism, generate a hierarchical query plan, semantically encode the hierarchical query plan through relative position encoding and bidirectional cross attention mechanism, integrate the performance change trend to generate query resource demand representation and construct a meta-learning-based policy network, optimize policy parameters through policy gradient and entropy regularization to perform resource evaluation, input the resource evaluation results into the multi-agent reinforcement learning framework, solve the global optimal scheduling plan based on the hierarchical soft maximum and minimum game and perform Monte Carlo tree search to expand the scheduling decision tree, determine the target node through upper confidence bound sampling, extract the discriminative features in the historical scheduling strategy through comparative distillation and migrate them to the online decision model, generate an initial scheduling decision and analyze heterogeneous causal effects based on the dual machine learning method, calibrate the decision results and forward the database access request to the target node;
[0173] The third unit is used to calculate the performance deviation value of each node based on the performance change trend and the real-time execution status. If the performance deviation value is greater than a preset performance deviation threshold, the current node is regarded as a faulty node, and the recent work information of the faulty node is obtained to construct a fault feature sequence. The fault feature sequence is segmented according to a preset time window, and the statistical features in each time window are calculated, and a feature matrix is constructed and sorted. The timing pattern is extracted through a long short-term memory network, the fault development trend is predicted, and the unfinished access requests are divided into multiple priority groups through a clustering algorithm, and a weighted directed graph is constructed based on the network bandwidth matrix between nodes and the current load level. The minimum spanning tree algorithm is applied to select the optimal migration path, the performance indicators of the target node are collected, and the performance fluctuations are monitored, and the probability graph evolution model is updated according to the status information of the faulty node after the migration is completed.
[0174] According to a third aspect of the embodiments of the present invention,
[0175] An electronic device is provided, comprising:
[0176] processor;
[0177] a memory for storing processor-executable instructions;
[0178] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0179] A fourth aspect of the embodiments of the present invention is:
[0180] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0181] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A service proxy method based on Doris front-end node, characterized in that: include: Collect the performance indicator sequence of the front-end node, map the performance indicator to a feature vector by calculating the multi-head recursive attention weight matrix, generate a node performance feature matrix and perform causal convolution on the time series data in the node performance feature matrix, extract the long-range time series dependency features, and perform graph attention convolution operations to extract the dynamic association features between nodes. Combine the dynamic association features and the long-range time series dependency features to obtain the node state vector, build a heterogeneous multi-layer connection pool, use a dual gradient-based online learning algorithm to iteratively calculate the dynamic parameters of each layer of the connection pool, establish the node association matrix by maximizing the conditional mutual information between the node representation vectors and introducing contrast noise constraints, perform variational inference to obtain the posterior probability distribution, and build a probabilistic graph evolution model of node performance based on Monte Carlo sampling; Predict the performance change trend of each node based on the probabilistic graph evolution model, receive database access requests and parse them into heterogeneous query graphs, iteratively update node representations through the edge attention message passing mechanism, generate a hierarchical query plan, semantically encode the hierarchical query plan through relative position encoding and bidirectional cross attention mechanism, integrate the performance change trend to generate query resource demand representation and build a meta-learning-based policy network, optimize policy parameters through policy gradient and entropy regularization to perform resource evaluation, input the resource evaluation results into the multi-agent reinforcement learning framework, solve the global optimal scheduling solution based on the hierarchical soft maximum and minimum game and perform Monte Carlo tree search to expand the scheduling decision tree, determine the target node through upper confidence bound sampling, extract the discriminative features in the historical scheduling strategy through comparative distillation and migrate them to the online decision model, generate an initial scheduling decision and analyze heterogeneous causal effects based on the dual machine learning method, calibrate the decision results and forward the database access request to the target node; The performance deviation value of each node is calculated based on the performance change trend and the real-time execution status. If the performance deviation value is greater than a preset performance deviation threshold, the current node is taken as a faulty node, and the recent work information of the faulty node is obtained to construct a fault feature sequence. The fault feature sequence is segmented according to a preset time window, and the statistical features within each time window are calculated, and a feature matrix is constructed and sorted. The timing pattern is extracted through a long short-term memory network, the fault development trend is predicted, and the unfinished access requests are divided into multiple priority groups through a clustering algorithm, and a weighted directed graph is constructed based on the network bandwidth matrix between nodes and the current load level. The minimum spanning tree algorithm is applied to select the optimal migration path, the performance indicators of the target node are collected, and the performance fluctuations are monitored. The probability graph evolution model is updated according to the status information of the faulty node after the migration is completed.
2. The method according to claim 1, characterized in that: The performance indicator sequence of the front-end node is collected, and the performance indicator is mapped into a feature vector by calculating the multi-head recursive attention weight matrix, a node performance feature matrix is generated, and the time series data in the node performance feature matrix is causally convolved and decomposed to extract the long-range time series dependency features, and the graph attention convolution operation is performed to extract the dynamic association features between nodes, and the node state vector is obtained by combining the dynamic association features and the long-range time series dependency features, and a heterogeneous multi-layer connection pool is constructed. The dynamic parameters of each layer of the connection pool are iteratively calculated using an online learning algorithm based on dual gradients, and the node association matrix is established by maximizing the conditional mutual information between the node representation vectors and introducing contrast noise constraints. The variational inference is performed to obtain the posterior probability distribution and a probability graph evolution model of the node performance is constructed based on Monte Carlo sampling, including: Continuously record the CPU usage rate, memory occupancy rate, disk IO read / write rate and network bandwidth usage rate according to the preset sampling interval, wherein the CPU usage rate includes user state occupancy, system state occupancy and IO waiting time, the memory occupancy rate includes physical memory usage, virtual memory usage and cache usage, the disk IO read / write rate includes read rate, write rate and average response time, and the network bandwidth usage rate includes inbound traffic, outbound traffic and number of TCP connections, to obtain the performance indicator sequence; The performance indicator sequence is divided into multiple performance indicator subsequences according to the indicator type, an attention head is constructed for each performance indicator subsequence, an attention weight is recursively calculated by a hyperbolic tangent function to generate a feature vector, the feature vectors output by each attention head are concatenated to form a node performance feature vector, and the node performance feature vectors of each time step are combined to generate a node performance feature matrix; The convolution kernel size and step size parameters are set, and the causal convolution decomposition is performed on the node performance feature matrix by using the causal filling method. The long-range temporal dependency is extracted by performing multi-layer convolution operations and increasing the number of convolution kernels layer by layer, and the long-range temporal dependency features are obtained by adding residual connections between convolution layers. Construct an adjacency matrix between nodes and determine edge weights according to real-time network delays, process the adjacency matrix through multi-layer graph attention convolution, concatenate the outputs of each layer using skip connections to obtain dynamic correlation features, and fuse the long-range temporal dependency features with the dynamic correlation features through a gating mechanism to obtain a node state vector; Constructing a multi-layer heterogeneous connection pool according to the node state vector, calculating the objective function gradient and dual variable of the connection weight, optimizing the primary variable and the dual variable alternately based on the gradient difference between the original problem and the dual problem, setting the update step of the dual variable to the inverse of the update step of the original variable, and attenuating the learning rate in proportion to the degree of convergence of the optimization objective function; Randomly sample the node state vector and calculate the conditional mutual information, introduce Gaussian white noise to construct comparison samples, obtain the node association matrix by maximizing and minimizing the mutual information, perform variational inference on the association matrix through diagonal Gaussian distribution, iteratively optimize the distribution parameters through the stochastic gradient method until convergence, and obtain the posterior probability distribution; A random sampling sequence is generated based on the posterior probability distribution, the state values in the sampling sequence are mapped to node attributes in the probabilistic graph model, directed edges are constructed according to the temporal correlation of the node states, the edge weights are determined based on the state transition probability, a Markov chain of the node states is established, the state transition matrix of the Markov chain is combined with the evolution law of the node performance indicator, and a node performance probabilistic graph evolution model that supports dynamic prediction is constructed.
3. The method according to claim 2, characterized in that The node state vector is randomly sampled and the conditional mutual information is calculated. Gaussian white noise is introduced to construct a comparison sample. The node association matrix is obtained by maximizing and minimizing the mutual information. The association matrix is subjected to variational inference through diagonal Gaussian distribution. The distribution parameters are iteratively optimized through the stochastic gradient method until convergence. The posterior probability distribution is obtained, including: Constructing positive sample pairs from a set of node state vectors based on a sliding window mechanism, setting the time step length of the sliding window to an even number, selecting state vectors in the first half of the time step and the second half of the time step of the sliding window to form positive sample pairs, obtaining time series dimension features of the state vectors, using the positive sample pairs as basic training data for state association, and using the time series dimension features of the state vectors as benchmark data for state evolution; Applying Gaussian white noise of different intensities to the state vector in the positive sample pair, obtaining a noise intensity parameter by uniformly sampling within a preset noise intensity range, multiplying the noise intensity parameter by the state vector to generate multiple groups of comparison samples, using the comparison samples as negative sample training data, and constructing a noise robustness verification mechanism based on the negative sample training data; Obtaining the node load state and the network topology as conditional variables, combining the conditional variables with the positive sample pairs to calculate the joint probability distribution, using the probability distribution of the positive sample pairs on the conditional variables as the marginal probability distribution, calculating the conditional mutual information of the positive sample pairs according to the ratio of the joint probability distribution to the marginal probability distribution, and using the conditional mutual information as a measurement indicator of state correlation; Calculating the conditional mutual information between the positive sample pair and the comparison sample, constructing an optimization objective function based on the mutual information, setting a maximization term for the mutual information of the positive sample pair and a minimization term for the mutual information of the comparison sample in the optimization objective function, and constructing an inter-node association matrix by optimizing the objective function; Taking the diagonal Gaussian distribution as an approximate representation of the posterior distribution, taking the mean vector and variance vector of the diagonal Gaussian distribution as variational parameters, initializing the mean vector with random sampling values of the standard normal distribution, initializing the variance vector with a unit vector, calculating the variational lower bound through the Monte Carlo sampling method, and constructing a gradient update mechanism for the distribution parameters based on the variational lower bound; An adaptive moment estimation optimization algorithm is used to iteratively optimize the variational parameters. In each round of iteration, the gradient value of the variational lower bound with respect to the distribution parameter is calculated. The optimization step size is dynamically adjusted according to a preset learning rate update strategy. The variational parameters are updated based on the gradient value and the optimization step size. The change of the variational lower bound during continuous iterations is monitored. When the change is lower than a preset threshold, the optimization process is determined to have converged. The converged variational parameters are output as posterior probability distributions.
4. The method according to claim 1, characterized in that Predict the performance change trend of each node based on the probabilistic graph evolution model, receive database access requests and parse them into heterogeneous query graphs, iteratively update node representations through edge attention message passing mechanism, generate hierarchical query plans, semantically encode the hierarchical query plans through relative position encoding and bidirectional cross attention mechanism, integrate the performance change trend to generate query resource demand representation and construct a meta-learning-based policy network, optimize policy parameters through policy gradient and entropy regularization to perform resource evaluation, input resource evaluation results into a multi-agent reinforcement learning framework, solve the global optimal scheduling solution based on hierarchical soft maximum and minimum game and perform Monte Carlo tree search to expand the scheduling decision tree, determine the target node through upper confidence bound sampling, extract the discriminative features in the historical scheduling strategy through comparative distillation and migrate them to the online decision model, generate initial scheduling decisions and analyze heterogeneous causal effects based on dual machine learning methods, calibrate the decision results and forward the database access request to the target node, including: Receiving a database access request, parsing the database access request into a heterogeneous query graph, setting node type identification, data table size, and filter selection rate attribute information in the nodes of the heterogeneous query graph, and calculating the weight of the edge between the nodes based on the cardinality of the connection key and the data distribution characteristics; Perform edge attention message passing on the heterogeneous query graph, perform linear mapping on the attribute information of each node to obtain an initial feature vector, aggregate the adjacent node features and edge weight information of the node to form a node message, calculate the attention weight through the node feature vector, perform weighted aggregation on the attention weight and the node message to iteratively update the node feature representation; Generate a hierarchical query plan using the data table nodes in the heterogeneous query graph as leaf nodes, construct table scan nodes, filter nodes, connection nodes, and aggregation nodes from the bottom up in the hierarchical query plan, and record the scale, computational complexity, and resource demand statistics of input and output data of different nodes; The hierarchical query plan is semantically encoded by using relative position encoding and bidirectional cross-attention mechanism, the internal dependency between nodes is captured by self-attention mechanism, the cross-attention mechanism is used to realize the fusion of forward information and reverse information, the query semantic vector is generated by combining residual connection and layer normalization, the performance change trend is fused with the query semantic vector to generate the query resource demand representation, a meta-learning-based policy network is constructed to output the resource allocation probability distribution and scheme evaluation results, and the policy parameters are optimized by policy gradient and entropy regularization to perform resource evaluation; The resource evaluation results are input into the multi-agent reinforcement learning framework, and the node resource quota is determined at the resource allocation layer using the hierarchical soft maximum and minimum game mechanism. The global optimal scheduling solution is solved at the task scheduling layer, and the Monte Carlo tree search is performed to select, expand, simulate, and return operations to expand the scheduling decision tree, and the target node is determined by upper confidence bound sampling. Extract discriminative features from historical scheduling strategies and migrate the discriminative features to an online decision model through a comparative distillation method, generate an initial scheduling decision based on the migrated features, use a dual machine learning method to analyze the heterogeneous causal effects of the initial scheduling decision, calibrate the initial scheduling decision according to the analysis results, and forward the database access request to the target node based on the calibrated initial scheduling decision.
5. The method according to claim 4, characterized in that Extracting discriminative features from historical scheduling strategies and migrating the discriminative features to an online decision model through a comparative distillation method, generating an initial scheduling decision based on the migrated features, analyzing the heterogeneous causal effects of the initial scheduling decision using a dual machine learning method, calibrating the initial scheduling decision according to the analysis results, and forwarding the database access request to the target node based on the calibrated initial scheduling decision includes: Obtaining historical scheduling strategies from a historical scheduling strategy database, extracting resource allocation ratio parameters, task segmentation granularity parameters, and parallelism configuration parameters in the historical scheduling strategies, constructing the extracted parameters into a discriminative feature vector, and generating a discriminative feature set; Training a teacher model based on the discriminative feature set, calculating a similarity score between a current query and a historical case, selecting a historical decision parameter with the highest similarity score, calculating a contrast loss between the decision parameter output by the teacher model and the historical decision parameter, updating the parameters of the online decision model by minimizing the contrast loss, and migrating the decision experience in the teacher model to the online decision model; Receive a database access request, parse the database access request to obtain a query type identifier, a data scale value, and a resource status parameter, add the query type identifier, a data scale value, and a resource status parameter to the online decision model, generate an initial scheduling decision including a memory allocation ratio value, a parallelism configuration value, and a data shard size value, construct a treatment group and a control group, use the memory allocation ratio value, the parallelism configuration value, and the data shard size value in the initial scheduling decision as processing variables, use the query execution time value and the resource utilization value as result variables, train causal inference models in the treatment group and the control group, respectively, and calculate the prediction difference between the treatment group and the control group as a heterogeneous causal effect value; According to the heterogeneous causal effect value, the memory allocation ratio value and the parallelism configuration value in the initial scheduling decision are adjusted to generate calibrated scheduling decision parameters, and the resource allocation ratio, number of task slices and execution parallelism of the target node are configured according to the calibrated scheduling decision parameters, and the database access request is sent to the target node to obtain the execution effect data and write it into the historical scheduling strategy database.
6. The method according to claim 1, characterized in that The performance deviation value of each node is calculated based on the performance change trend and the real-time execution status. If the performance deviation value is greater than a preset performance deviation threshold, the current node is taken as a fault node, and the recent work information of the fault node is obtained to construct a fault feature sequence. The fault feature sequence is segmented according to a preset time window, and the statistical features in each time window are calculated and a feature matrix is constructed and sorted. The timing pattern is extracted through a long short-term memory network, the fault development trend is predicted, and the unfinished access requests are divided into multiple priority groups through a clustering algorithm, and a weighted directed graph is constructed based on the network bandwidth matrix and the current load level between nodes. The minimum spanning tree algorithm is applied to select the optimal migration path, the performance indicators of the target node are collected, and the performance fluctuations are monitored. The probability graph evolution model is updated according to the status information of the fault node after the migration is completed, including: Based on the real-time execution status, the processor usage rate, memory occupancy rate, disk waiting time and network throughput are collected to obtain a performance vector and set a corresponding performance vector weight. The relative deviation between the actual value of each performance indicator in the performance vector and the predicted value corresponding to the performance change trend is multiplied by the corresponding weight and then summed to obtain the performance deviation value of each node; If the performance deviation value is greater than a preset performance deviation threshold, the current node is marked as a faulty node, recent work information of the faulty node is obtained and a fault feature sequence is constructed, wherein the recent work information includes request type, response time and error code, a fault feature sequence is constructed, and for the fault feature sequence, the fault feature sequence is segmented according to a preset time window, and the maximum, minimum and average value of the number of request processing in each time window are counted, the response time quantile value is calculated, the fluctuation range of the processor utilization rate and the memory utilization rate is recorded, the change pattern of the disk read and write rate is analyzed, and a feature matrix is generated and sorted in descending order according to the size of the eigenvalues; Calculate the correlation coefficient between the features of each column of the sorted feature matrix and the degree of performance degradation, determine the correlation coefficient corresponding to the feature values in the top 10% based on the sorting results, take the features whose correlation coefficients are greater than a preset correlation threshold as strong correlation features, perform dimensionality reduction operations on the strong correlation features until a preset information retention rate is reached, generate a strong correlation sequence, and for the strong correlation sequence, extract the long-term dependency and short-term variation rules between the features through a long short-term memory network to obtain the time series pattern, perform prediction based on the time series pattern, and obtain the fault development trend; Extract the execution time, resource usage, and data access volume of the unfinished access requests in the faulty node, combine them with the predicted node available time, construct the request feature vector after standardization, set the density clustering radius and the minimum number of samples, randomly select the core request object, expand the cluster cluster with the core request object as the center, calculate the similarity between requests, divide similar requests larger than the density clustering radius into priority groups according to the urgency of the fault development trend, and repeat the clustering expansion until all requests are accessed; Construct a network bandwidth matrix between nodes, convert bandwidth values into transmission costs, construct a weighted directed graph, select the edge with the minimum transmission cost, construct a minimum spanning tree, and select the optimal migration path. If the performance indicator fluctuation of the target node exceeds the preset performance fluctuation threshold, select an alternative migration path, execute migration, and update the probability graph evolution model according to the status information of the faulty node after the migration is completed.
7. The method according to claim 6, characterized in that Construct a network bandwidth matrix between nodes, convert bandwidth values into transmission costs, construct a weighted directed graph, select the edge with the minimum transmission cost, construct a minimum spanning tree, and select the optimal migration path. If the performance index fluctuation of the target node exceeds the preset performance fluctuation threshold, select an alternative migration path including: Obtain a cloud data center topology map where the faulty node is located, collect real-time network bandwidth values of all edges in the cloud data center topology map based on a network monitoring system, and construct a network bandwidth adjacency matrix, in which the measured bandwidth values are recorded at corresponding positions of nodes with direct connections, and infinity is recorded at corresponding positions of nodes without direct connections; Obtaining the processor usage and memory usage of each node in the network bandwidth adjacency matrix, converting the network bandwidth adjacency matrix into a transmission cost matrix, wherein the transmission cost of each edge in the transmission cost matrix is calculated by the inverse of the bandwidth value, and when the processor usage of the target node is greater than a preset processor usage threshold or the memory usage is greater than a preset memory usage threshold, increasing the transmission cost of the corresponding edge by a preset ratio; Constructing a weighted directed graph based on the transmission cost matrix, taking the faulty node as the root node, selecting a node corresponding to the minimum transmission cost edge connected to the root node as the first access node, adding the first access node to the minimum spanning tree, selecting a node corresponding to the minimum transmission cost edge connected to the visited node from the unvisited nodes as a new access node, adding the new access node to the minimum spanning tree, and repeating the process until all nodes are visited; A request migration path is generated from the faulty node to the target node based on the minimum spanning tree, wherein the total transmission cost of the request migration path is minimal and the node does not pass through nodes whose load is higher than a preset load threshold. The performance indicators of each node on the request migration path are monitored. When the load of the node on the request migration path exceeds the load threshold or the bandwidth is lower than the preset bandwidth threshold, the transmission cost of each edge is recalculated and the request migration path is updated.
8. A service proxy system based on Doris front-end nodes, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to collect the performance indicator sequence of the front-end node, map the performance indicator into a feature vector by calculating the multi-head recursive attention weight matrix, generate a node performance feature matrix and perform causal convolution decomposition on the time series data in the node performance feature matrix, extract the long-range time series dependency features, and perform graph attention convolution operation to extract the dynamic association features between nodes, combine the dynamic association features and the long-range time series dependency features to obtain the node state vector, construct a heterogeneous multi-layer connection pool, use an online learning algorithm based on dual gradient to iteratively calculate the dynamic parameters of each layer of the connection pool, establish the node association matrix by maximizing the conditional mutual information between the node representation vectors and introducing contrast noise constraints, perform variational inference to obtain the posterior probability distribution, and construct a probabilistic graph evolution model of node performance based on Monte Carlo sampling; The second unit is used to predict the performance change trend of each node based on the probabilistic graph evolution model, receive database access requests and parse them into heterogeneous query graphs, iteratively update node representations through the edge attention message passing mechanism, generate a hierarchical query plan, semantically encode the hierarchical query plan through relative position encoding and bidirectional cross attention mechanism, integrate the performance change trend to generate query resource demand representation and construct a meta-learning-based policy network, optimize policy parameters through policy gradient and entropy regularization to perform resource evaluation, input the resource evaluation results into the multi-agent reinforcement learning framework, solve the global optimal scheduling plan based on the hierarchical soft maximum and minimum game and perform Monte Carlo tree search to expand the scheduling decision tree, determine the target node through upper confidence bound sampling, extract the discriminative features in the historical scheduling strategy through comparative distillation and migrate them to the online decision model, generate an initial scheduling decision and analyze heterogeneous causal effects based on the dual machine learning method, calibrate the decision results and forward the database access request to the target node; The third unit is used to calculate the performance deviation value of each node based on the performance change trend and the real-time execution status. If the performance deviation value is greater than a preset performance deviation threshold, the current node is regarded as a faulty node, and the recent work information of the faulty node is obtained to construct a fault feature sequence. The fault feature sequence is segmented according to a preset time window, and the statistical features in each time window are calculated, and a feature matrix is constructed and sorted. The timing pattern is extracted through a long short-term memory network, the fault development trend is predicted, and the unfinished access requests are divided into multiple priority groups through a clustering algorithm, and a weighted directed graph is constructed based on the network bandwidth matrix between nodes and the current load level. The minimum spanning tree algorithm is applied to select the optimal migration path, the performance indicators of the target node are collected, and the performance fluctuations are monitored, and the probability graph evolution model is updated according to the status information of the faulty node after the migration is completed.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Flow distribution system based on redundant switch
CN120128542A
Medical data abnormal access identification method and system
CN120185928A
Data structured processing method and system based on structured query tree
CN120256436A
Multi-dimensional dynamic carbon emission factor modeling method and system
CN120337799A
Disk scheduling method
CN120428922A